Concepts

The vocabulary Shimmy uses everywhere — in these docs, in the dashboard, and in the API.

Workflows and Steps

A Workflow is one thing you build on LLMs: an AI app — a chat assistant, a RAG feature, an endpoint that summarizes or extracts — or an agentic workflow that plans and calls tools across many steps. Shimmy treats them the same way. Each Workflow has its own API key; a call made with that key belongs to that Workflow.

A Step is one LLM call site within it, named by you: classify, rewrite_query, answer, draft_reply. An app may have two Steps and an agent twenty; either way each gets its own model. The name is the Step’s identity for its whole life, and everything Shimmy does keys on it — the Recordings it replays, the search that picks its model, the line in shimmy.lock.

Name it in code rather than letting Shimmy infer it from the prompt, because a content-derived identity changes when the content does: rewording a system prompt would make a new Step and throw away everything learned about the old one. A declared name survives prompt edits. Shimmy also keeps a Shape hash of each Step — its masked prompt, tool names and response format — so it notices when the request behind a name has changed, and re-checks the Step’s model instead of trusting a result measured on a different prompt.

Runs

A run is one execution of a Workflow — an agent handling a ticket, a job processed, or one request or conversation in an app. Steps inside a run are linked into a graph, which is what the dashboard draws.

You rarely name a run: its id resolves from an active OpenTelemetry trace, or a fresh uuid. Pass one when a request is not one run — a nightly job over 10,000 tickets is 10,000 runs — or when several requests are: pass a chat app’s conversation id and every turn reads as one run.

run.idstep.id
Answerswhich execution?which call site?
Changesevery runnever
You supply itrarelyalways

Phases and Stages

A Phase is how a process behaves right now, declared by that process:

PhaseWhat a call does
DevReplays the Step’s Recording; records one on a miss.
TuningReal call on a model under test, graded, on Shimmy’s keys.
ProductionManaged: through Shimmy, on your keys. Passthrough: straight to your provider.

A Stage is where a Workflow stands as a whole, computed by Shimmy from its traffic and Report: Developing → Tuning → Report ready → In production. One Workflow can have processes in several Phases at once — CI replaying Recordings while staging tunes — but it has one Stage.

Searching

Baseline search. Pin a model — the one you’d ship today — and it becomes the Step’s Baseline: the quality reference. Only cheaper models are tried, and if none holds up, the Step settles on the Baseline itself.

Discovery search. Send model: "discover" and there’s no Baseline: the search anchors on a frontier model and works down to the cheapest model that matches it.

A Mode decides which models a search may try:

ModeCandidates
Single-providerOne provider’s models — the Baseline’s, or the one you allow.
Multi-providerEvery provider you allow.
Open-weights migrationOnly open-weights models under a commercial-use license, whatever your Baseline is.

Allowed providers narrows any Mode to an explicit list, and failover honors it too. A Model is tried once even when several hosts offer it; its Offerings — one host’s price each — compete on price alone.

A Step is settled when its search converges on a winner.

The Report and the Backup chain

The Report is the result of Tuning, per Workflow: for every Step, the winner, its Backup chain, and every Model tried with its samples, score, cost and the answers it gave. It fills in as Steps settle and is ready when every Expected Step has settled — exclude a Step that only runs on rare paths so it doesn’t hold the Report up. Reports are versioned: a new version is saved when the Report becomes ready, when an answer changes, or when you finalize it.

The Backup chain is what to serve when the winner can’t: the winner, the best passing runner-up on a different allowed provider (an outage rarely takes out two providers), then the Baseline if there was one. Passthrough and managed Production both walk it.

The Report exports as shimmy.lock: each Step’s model, provider and Backup chain, as a file you commit. See files and CLI.

Recordings

A Recording is a real response captured once for a Step and replayed in Dev. It is matched by Step and by a hash of the request’s meaningful parts — messages, tool names, JSON mode — so the same input replays the same answer. Recordings are kept on the backend and, pulled or recorded locally, in your repo under shimmy/recordings/. See Recordings.

Drift

Drift is a settled Step’s winner no longer passing the quality bar it once passed — a provider updated the model, or your inputs moved. Managed Production grades a small sample of calls to catch it; when it does, the Step’s next backup serves immediately and the Step is re-searched. See managed or passthrough.

Outcomes

An outcome is what actually happened to a call’s output: the JSON parsed, the tool ran, a person accepted it. To grade a model, the search otherwise pays an LLM judge to approximate exactly these facts. Reporting them is free evidence that settles Steps sooner. See report outcomes.