Running the search

In the Tuning Phase, every call is an experiment: it’s served by a model under test, graded, and counted toward that Step’s search. Run the Workflow the way it will really be used — a test suite, last week’s traffic replayed, a staging environment — and each Step settles on the cheapest model that holds its quality.

SHIMMY_PHASE=tuning python run_eval_set.py

Baseline or Discovery

How a Step is searched depends on what you send as the model:

// Baseline search: pin the model you'd ship today.
openai.chat.completions.create({ model: 'gpt-5.5', messages });

// Discovery search: no Baseline — Shimmy anchors on a frontier model.
openai.chat.completions.create({ model: 'discover', messages });

Baseline search — you pinned a model. It is the quality reference, and only cheaper models are tried against it. Every candidate’s answer is graded against the Baseline’s on the same live input; a candidate that keeps matching it becomes the winner, and the cheapest such candidate wins. If no cheaper model holds up, the Step settles on its Baseline — “keep what you have” is a real answer, and the Report says so.

Discovery search — "discover", no pin. The Step’s first call anchors on a frontier model (within your allowed providers), and later calls run every surviving candidate against the same input, graded by the same judge. Each round halves the field by cost per quality point; a Step settles when one survivor has matched the anchor across enough rounds (three, at 90%, by default). Discovery is faster to settle and better at finding a model you’d never have tried, and costs more per run while it searches.

You can also ask per Step, whatever the request’s model: step(..., mode="discover") searches that Step, mode="off" never searches it.

What the search may try

Your account’s Mode and Allowed providers (dashboard → Settings, or config as code) bound every search:

ModeCandidates
Single-providerThe Baseline’s provider only — a Claude Step is searched among Claude models.
Multi-providerEvery allowed provider.
Open-weights migrationOnly open-weights models with a commercial-use license, served by open-weights hosts. Your Baseline can be anything — this is how you move off a closed model.

Tuning runs on Shimmy’s provider keys, so you can try a provider you have no account with. A model offered by several hosts is tried once; the hosts compete on price.

When a Step settles

A settled Step stops being searched; its winner and Backup chain are fixed in the Report. A Step re-opens on its own when:

  • its request changes shape — you edited the prompt, added a tool, turned on JSON mode. The settled model keeps serving while the search re-runs against the new prompt, and the Report gets a new version if the answer changes.
  • it drifts in Production — see managed Production.

The Workflow is Report ready when every Expected Step has settled. A Step that only runs on a rare path (an error handler, an escalation) can hold the Report up; exclude it on the Workflow’s Steps tab, or finalize the Report early.

Speeding it up

  • Report outcomes. Every shimmy.report(schema_valid=True) is evidence the search would otherwise pay a judge for. See report outcomes.
  • Declare kind. step("classify", kind="classification") tells the search what the Step is for, so it starts in the right place.
  • Optimization runs. Settings → advanced search can try several candidates per call for a few runs — fewer runs to settle, more spent per run.

What it costs

Each Tuning call is charged to your prepaid Wallet at what the providers charge Shimmy, plus a margin — grading and the candidates’ calls included. Nothing is charged for a call that ran on your own key.

See billing for prices.