Running the search
In the Tuning Phase, every call is an experiment: it’s served by a model under test, graded, and counted toward that Step’s search. Run the Workflow the way it will really be used — a test suite, last week’s traffic replayed, a staging environment — and each Step settles on the cheapest model that holds its quality.
SHIMMY_PHASE=tuning python run_eval_set.py Baseline or Discovery
How a Step is searched depends on what you send as the model:
// Baseline search: pin the model you'd ship today.
openai.chat.completions.create({ model: 'gpt-5.5', messages });
// Discovery search: no Baseline — Shimmy anchors on a frontier model.
openai.chat.completions.create({ model: 'discover', messages }); Baseline search — you pinned a model. It is the quality reference, and only cheaper models are tried against it. Every candidate’s answer is graded against the Baseline’s on the same live input; a candidate that keeps matching it becomes the winner, and the cheapest such candidate wins. If no cheaper model holds up, the Step settles on its Baseline — “keep what you have” is a real answer, and the Report says so.
Discovery search — "discover", no pin. The Step’s first call anchors on a
frontier model (within your allowed providers), and later calls run every
surviving candidate against the same input, graded by the same judge. Each round
halves the field by cost per quality point; a Step settles when one survivor has
matched the anchor across enough rounds (three, at 90%, by default). Discovery is
faster to settle and better at finding a model you’d never have tried, and costs
more per run while it searches.
You can also ask per Step, whatever the request’s model: step(..., mode="discover") searches that Step, mode="off" never searches it.
What the search may try
Your account’s Mode and Allowed providers (dashboard → Settings, or config as code) bound every search:
| Mode | Candidates |
|---|---|
| Single-provider | The Baseline’s provider only — a Claude Step is searched among Claude models. |
| Multi-provider | Every allowed provider. |
| Open-weights migration | Only open-weights models with a commercial-use license, served by open-weights hosts. Your Baseline can be anything — this is how you move off a closed model. |
Tuning runs on Shimmy’s provider keys, so you can try a provider you have no account with. A model offered by several hosts is tried once; the hosts compete on price.
When a Step settles
A settled Step stops being searched; its winner and Backup chain are fixed in the Report. A Step re-opens on its own when:
- its request changes shape — you edited the prompt, added a tool, turned on JSON mode. The settled model keeps serving while the search re-runs against the new prompt, and the Report gets a new version if the answer changes.
- it drifts in Production — see managed Production.
The Workflow is Report ready when every Expected Step has settled. A Step that only runs on a rare path (an error handler, an escalation) can hold the Report up; exclude it on the Workflow’s Steps tab, or finalize the Report early.
Speeding it up
- Report outcomes. Every
shimmy.report(schema_valid=True)is evidence the search would otherwise pay a judge for. See report outcomes. - Declare
kind.step("classify", kind="classification")tells the search what the Step is for, so it starts in the right place. - Optimization runs. Settings → advanced search can try several candidates per call for a few runs — fewer runs to settle, more spent per run.
What it costs
Each Tuning call is charged to your prepaid Wallet at what the providers charge Shimmy, plus a margin — grading and the candidates’ calls included. Nothing is charged for a call that ran on your own key.
See billing for prices.