The right model for every step.
Found before you ship.
Shimmy records each step of your AI app or agent while you build, searches for the best model on cost and quality while you tune, and hands you a Report — and a lockfile — when every step has settled.
Python · TypeScript · Rust · OpenAI and Anthropic clients
Already have an account? Log in →
- classify gpt-5.4-nano was gpt-5.5 −96%
- retrieve_plan gpt-5.4-mini was gpt-5.5 −81%
- draft_reply claude-sonnet-4-6 was claude-opus-4-8 −40%
- policy_check gpt-5.5 nothing cheaper held — keep it ±0%
Wrap your client. Name your steps.
The SDK wraps the client you already use. Each LLM call gets a name, and one environment variable says which Phase the process is in — the same code replays in CI, searches in staging and serves in production.
# SHIMMY_PHASE=dev | tuning | production
shimmy = Shimmy()
client = shimmy.wrap(OpenAI())
with shimmy.run("triage"):
with shimmy.step("classify"):
client.chat.completions.create(
model="gpt-5.5", # your Baseline
messages=messages,
)Every LLM call is its own model decision.
A classifier doesn’t need the model your planner needs. Picking by hand is guessing, and checking the guess means an eval harness per step. Shimmy does that work in the course of building — and stays out of your way once you ship.
Record once. Replay forever.
Each Step’s call is made for real once and saved as a Recording in your repo. Unit and integration tests replay it — offline, deterministic, free. Assert on which Steps ran, and inject the timeouts, rate limits and refusals production will throw at you.
Search every Step for its best model.
Run your app or agent on real inputs. Shimmy tries cheaper models against the one you pinned — or works down from a frontier model — grades every answer, and settles each Step on the cheapest model that holds its quality. On our keys, from a prepaid Wallet.
See the answer, and why.
For every Step: the winner, its backups on other providers, and every model tried — what it cost, how it scored, what it actually said for your inputs. Exported as shimmy.lock, a file you commit and review like any other.
Ship it, with us or without us.
Passthrough: the SDK calls your provider directly with each Step’s winner, and falls back down its Backup chain on failure — free. Managed: drift detection that re-searches when a model gets worse, failover, caching and spend caps, for a flat fee.
Keep the answer right after launch.
Models change underneath you. Keep calls going through Shimmy — on your own provider keys — and the answer Tuning found stays the right one.
- 1 drift
a sample of calls is graded; a model that slips is replaced by its backup and re-searched
- 2 failover
outage, 429, refused key, overflow or refusal → the next model in the Backup chain
- 3 hedging
a Step with a latency budget races its first backup when the winner is slow
- 4 cache
exact repeats answered without a provider call
- 5 spend caps
a monthly cap and a warning threshold, per account
- 6 alerts
drift, failover and spend — by email or signed webhook
Pay for the search, not for shipping.
Replays cost nothing. A free Dev allowance covers making Recordings.
What the providers charge, from a prepaid Wallet. It never spends more than you put in.
Passthrough is free. Managed is $4.99 per Workflow per month, on your own keys.
Bring us the app or agent you’re about to ship.
We’ll tune it with you — every step, every model worth trying — and you’ll launch on an answer instead of a guess.