Shimmy · FAQ

Questions, answered.

How Shimmy finds the right model for every step of your AI app or agent, what it costs, and how you run on the answer.

General

General.

What is Shimmy?

Shimmy finds the best LLM for every step of an AI app or agentic workflow on cost and quality, before you ship it. While you build, it records each step once so tests replay offline for free. While you tune, it searches each step — cheaper models measured against the one you would ship, on your real inputs — until every step settles. You get a Report with the winner for each step, its backups and every model tried, and a shimmy.lock file to ship on.

Why per step, instead of one model for the whole app?

Because the steps are different jobs. A classifier picking one of five labels, a query rewrite before retrieval, a planner breaking down a request and a writer drafting the reply need very different models, and an app or agent run on one model is either overpaying for its easy steps or underpowered for its hard ones. Shimmy measures each step on its own inputs.

Does Shimmy work for an AI app, not just an agent?

Yes. A chat assistant, a RAG feature or an endpoint that summarizes or extracts is a Workflow like any agent: each LLM call site is a Step, and each request or conversation is a run. An app with two Steps tunes and settles the same way an agent with twenty does, usually faster.

How does Shimmy decide a cheaper model is good enough?

Each candidate answers the same live input as your baseline model and is compared with it — prose by an LLM judge, tool calls structurally against the tools you declared. It has to keep matching across many calls (20 by default, at 90%) before a step settles on it. Outcomes your own code reports, like a schema that validated, count too. If nothing cheaper holds, the step keeps your model, and the Report says so.

What if I don’t know which model to start from?

Use a Discovery search: send model "discover" and Shimmy anchors the step on a frontier model, races several cheaper candidates per call, and halves the field each round until one has matched across enough rounds. It settles in a few runs.

How long does it take to settle?

It depends on how many steps you have and how often each runs: every Tuning call moves its step’s search along. A Discovery search usually settles a step in a handful of runs; a Baseline search needs enough graded calls for each candidate it tries. Reporting outcomes from your code settles steps sooner.

Pricing

Pricing.

How much does Shimmy cost?

Development is free: recordings replay at no cost, and a free Dev allowance covers making them. Tuning is billed at what the model providers charge plus 25%, from a prepaid wallet. In production, passthrough is free and managed production is $4.99 per workflow per month.

Why a prepaid wallet?

Because what Tuning costs depends on your workflow, and that is hard to know before you start. A prepaid wallet makes the ceiling yours: Tuning spends what you put in and no more, and pauses when it runs out. The Tuning estimator gives a range for sizing the first top-up.

Is there a free tier?

Development is free, including a $5 Dev allowance for making recordings — no card required. Passthrough production is free forever. You pay for Tuning, and for managed production if you want it.

Do credits expire?

Wallet credits expire 365 days after purchase and are not refundable. The oldest credit is spent first.

Production

Production.

Do I have to send production traffic through Shimmy?

No. Export the lockfile and run the SDK in passthrough: it calls your provider directly with each step’s winner, and falls back down that step’s backup chain on an outage, a rate limit or a refusal. Shimmy is not in the path.

What does managed production add?

Drift detection: a small sample of each step’s calls is graded in the background, and when a model gets worse its backup serves immediately while the step is re-searched. Plus failover across providers, hedged requests for latency-sensitive steps, a response cache, monthly spend caps, and alerts by email or signed webhook.

What happens when a model provider goes down?

Each settled step has a backup chain: its winner, the best passing runner-up on a different provider, then your baseline. Both passthrough and managed production walk it, so an outage at one provider fails over to another automatically.

Technical

Technical.

Which languages and providers are supported?

SDKs for Python, TypeScript and Rust, wrapping the OpenAI and Anthropic clients. Searches can span OpenAI, Anthropic, Perplexity, Mistral, Cloudflare and open-weights hosts like Ollama Cloud — limited to the providers you allow. An open-weights migration mode tries only commercially licensed open-weights models.

Do you store my API keys?

Tuning runs on Shimmy’s own provider keys, so you need none. For managed production you store your provider keys with us; they are encrypted at rest and never shown again. Shimmy API keys are stored as hashes — the plaintext is shown once.

Do you store my prompts and responses?

Tuning keeps them for the life of the Report — they are what the Report shows you, model by model. Outside Tuning, storing content is a setting you can turn off. Sensitive values such as emails, cards and keys are masked in what we store by default.

Do recordings and tests work offline?

Yes. Under pytest, vitest, jest, node --test or cargo test, the SDK replays recordings from your repo with no network. In CI, a test whose input no longer matches a recording fails rather than calling a model.

Can I self-host?

Yes. A self-hosted build ships as a Docker image with Postgres. Traffic and keys never leave your VPC; pricing is a flat enterprise license.

Comparison

How we compare.

How is Shimmy different from an AI gateway like LiteLLM or Portkey?

A gateway routes your production traffic; it is in the path of every call, forever. Shimmy does its work before you ship: it finds the right model for each step during development and tuning, and hands you the answer as a file. You can keep it in the path for drift protection — or not.

How is it different from an eval tool?

An eval tool scores the models you choose to test, on datasets you assemble. Shimmy chooses the candidates, runs them on your real inputs as your app or agent executes, grades them against your own baseline per step, and stops when each step has settled.

How is it different from mocking the LLM in tests?

A mock returns what you wrote; a recording returns what the model actually said, captured once from a real call. Shimmy also logs every call your code makes, so a test can assert on which steps ran and in what order, and inject the failures providers really produce.

Still have questions?

Talk to us.

Whether you’re tuning an app or agent of your own or want us to build one with you, we’d like to hear from you.