Shimmy · how it works

One step’s life,
from first call to production.

Shimmy follows each step of your AI app or agent through three Phases — recorded while you build, searched while you tune, served on the answer once you ship. Here’s what happens to one of them.

The lifecycle

A single Step, stage by stage.

Dev makes the Step cheap to build against. Tuning finds the model it should run on. Production keeps it there.

1
declare

Name the Step

Wrap your OpenAI or Anthropic client with the Shimmy SDK and give each LLM call a name — classify, plan, draft_reply. The name is the Step’s identity for its whole life: rewording its prompt never resets what Shimmy has learned, and a hash of its shape tells Shimmy when the request behind the name has really changed.

2
record

Dev: the first call is real, every later one is replayed

In the Dev Phase the Step’s first call goes to a real model and is saved as a Recording, matched by a hash of the messages, tools and response format. Tests replay it from your repo — no network, no spend — and in CI a test whose input changed fails instead of passing on a stale answer.

3
anchor

Tuning: measure against what you’d ship

In the Tuning Phase the model you pinned is the Baseline, the quality reference. With no pin, a Discovery search anchors on a frontier model instead. Candidates come from the providers your Mode allows — one provider, several, or open-weights models only — on Shimmy’s keys.

4
grade

Every candidate answers your real input

A cheaper candidate answers the same live call as the Baseline, and the two are compared: prose by an LLM judge, tool calls structurally against the tools you declared. Outcomes your code reports — the JSON parsed, the tool ran — are counted as evidence too, and free.

5
settle

The cheapest model that keeps matching wins

A candidate that fails is dropped; one that keeps matching the Baseline across enough calls — 20 by default, at 90% — settles the Step. A Discovery search races several candidates per call and halves the field each round. If nothing cheaper holds, the Step keeps its Baseline, and the Report says so.

6
report

The Report, and its lockfile

When every expected Step has settled, the Report is ready: each Step’s winner, its Backup chain — the best runner-up on a different provider, then the Baseline — and every model tried, with cost, score and the answers it gave. It exports as shimmy.lock.

7
serve

Production: on the answer, with or without Shimmy

Passthrough: the SDK calls your provider directly with each Step’s winner and walks the Backup chain when a call fails. Managed: calls go through Shimmy on your own keys, with failover, hedging for latency-sensitive Steps, and a response cache.

8
watch

Drift: the answer stays right

In managed Production a small sample of each Step’s calls is graded in the background. When the winner stops passing, its backup serves immediately and the Step is re-searched among providers you hold keys for. You get an alert, and the Report a new version.

The evidence

Why you can trust a cheaper model.

A search is only as good as the evidence it settles on. Every verdict is measured on your inputs, against your own Baseline — and the Report keeps every answer, so you can check it yourself.

judge

Graded against your own Baseline

A candidate isn’t scored in the abstract; it is compared with the answer your Baseline gave to the same input. “As good as what I’d ship” is the bar, measured on your traffic.

structure

Tool calls verified, not opined on

A tool-calling Step is checked mechanically: every call must name a declared tool with arguments that parse. A model that invents tools is eliminated before it can break an agent loop.

outcomes

Your code’s evidence counts

What your program already knows — schema valid, tool executed, a person accepted the answer — is stronger than any judge, and free. A hard failure always outranks a generous score.

requirements

Never a model that can’t do the job

Requirements are inferred from the request — tools, JSON mode, images, context size — or declared ahead of time, and models that can’t meet them are never tried.

Get in touch

Want to see it on your app or agent?

Wrap your client, name your steps, and run one tuning pass. Recordings are free; the dashboard shows each step settle as it happens.