Shimmy · model discovery for AI apps and agents

The right model for every step.
Found before you ship.

Shimmy records each step of your AI app or agent while you build, searches for the best model on cost and quality while you tune, and hands you a Report — and a lockfile — when every step has settled.

Python · TypeScript · Rust · OpenAI and Anthropic clients

Already have an account? Log in →

report · support-triage report ready
  • classify gpt-5.4-nano was gpt-5.5 −96%
  • retrieve_plan gpt-5.4-mini was gpt-5.5 −81%
  • draft_reply claude-sonnet-4-6 was claude-opus-4-8 −40%
  • policy_check gpt-5.5 nothing cheaper held — keep it ±0%
4 of 4 Steps settled shimmy.lock ↓
The integration

Wrap your client. Name your steps.

The SDK wraps the client you already use. Each LLM call gets a name, and one environment variable says which Phase the process is in — the same code replays in CI, searches in staging and serves in production.

# SHIMMY_PHASE=dev | tuning | production
shimmy = Shimmy()
client = shimmy.wrap(OpenAI())

with shimmy.run("triage"):
    with shimmy.step("classify"):
        client.chat.completions.create(
            model="gpt-5.5",   # your Baseline
            messages=messages,
        )
How it works

Every LLM call is its own model decision.

A classifier doesn’t need the model your planner needs. Picking by hand is guessing, and checking the guess means an eval harness per step. Shimmy does that work in the course of building — and stays out of your way once you ship.

01 / dev

Record once. Replay forever.

Each Step’s call is made for real once and saved as a Recording in your repo. Unit and integration tests replay it — offline, deterministic, free. Assert on which Steps ran, and inject the timeouts, rate limits and refusals production will throw at you.

02 / tuning

Search every Step for its best model.

Run your app or agent on real inputs. Shimmy tries cheaper models against the one you pinned — or works down from a frontier model — grades every answer, and settles each Step on the cheapest model that holds its quality. On our keys, from a prepaid Wallet.

03 / report

See the answer, and why.

For every Step: the winner, its backups on other providers, and every model tried — what it cost, how it scored, what it actually said for your inputs. Exported as shimmy.lock, a file you commit and review like any other.

04 / production

Ship it, with us or without us.

Passthrough: the SDK calls your provider directly with each Step’s winner, and falls back down its Backup chain on failure — free. Managed: drift detection that re-searches when a model gets worse, failover, caching and spend caps, for a flat fee.

Managed production

Keep the answer right after launch.

Models change underneath you. Keep calls going through Shimmy — on your own provider keys — and the answer Tuning found stays the right one.

  1. 1
    drift

    a sample of calls is graded; a model that slips is replaced by its backup and re-searched

  2. 2
    failover

    outage, 429, refused key, overflow or refusal → the next model in the Backup chain

  3. 3
    hedging

    a Step with a latency budget races its first backup when the winner is slow

  4. 4
    cache

    exact repeats answered without a provider call

  5. 5
    spend caps

    a monthly cap and a warning threshold, per account

  6. 6
    alerts

    drift, failover and spend — by email or signed webhook

Pricing

Pay for the search, not for shipping.

dev Free

Replays cost nothing. A free Dev allowance covers making Recordings.

tuning Cost + 25%

What the providers charge, from a prepaid Wallet. It never spends more than you put in.

production $0 or $4.99

Passthrough is free. Managed is $4.99 per Workflow per month, on your own keys.

Talk to us

Bring us the app or agent you’re about to ship.

We’ll tune it with you — every step, every model worth trying — and you’ll launch on an answer instead of a guess.