Shimmy
Find the best model for every LLM call in your AI app or agentic workflow — before you ship it.
Every LLM call is a separate decision. A chat app’s query rewrite doesn’t need the model its answer needs; an agent’s classifier doesn’t need the model its planner needs. Picking them by hand means guessing, and checking the guess means building an eval harness for every call. Shimmy does that work while you build.
Three Phases
Your code declares which Phase it is in — one environment variable, no code changes between them.
Dev. Each Step’s call is made once for real and saved as a Recording. After that, unit and integration tests replay it: offline, deterministic and free. Recordings live in your repo, next to the tests that use them.
Tuning. Run your app or workflow on real inputs. Shimmy searches each Step: it tries cheaper models against the one you pinned (or, with no pin, works down from a frontier model), grades every answer, and settles each Step on the cheapest model that holds its quality. Tuning runs on Shimmy’s provider keys, paid from a prepaid Wallet — you need no account with the providers being tried.
Production. When every Step has settled, the Report says which model to use where, with every model tried, what it cost and what it answered. Then either:
- Passthrough — commit the exported
shimmy.lock. The SDK calls your provider directly with each Step’s winner, and falls back down its Backup chain if the winner fails. Shimmy is out of the path. - Managed — keep calls going through Shimmy on your own provider keys, for drift detection that re-searches when a model gets worse, failover, caching, spend caps and alerts. A flat monthly fee per Workflow.
What a Workflow looks like to Shimmy
A Workflow is one AI app or agentic workflow — a support chatbot, a RAG
search feature, a document pipeline, a tool-calling agent. A Step is one LLM
call site in it, named by you — rewrite_query, answer, classify, draft_reply. Naming Steps is the one thing Shimmy
needs from your code: it’s what lets a Recording replay the right answer, a search
accumulate evidence for the right call, and a lockfile pin the right model. Read concepts for the rest of the vocabulary, and instrument an AI app or instrument an agent for the shape of
yours.
Install
npm install @rfa-labs/shimmy Then follow the quickstart.
Everything here
- Start
- Overview — What Shimmy is: the best model for every step, found before you ship.
- Quickstart — Wrap a client, declare a Step, and take it from Dev to a Report.
- Concepts — Workflows, Steps, Phases, Modes, the Report and the Backup chain.
- Dev Phase
- Recordings — Record each Step once, replay it forever — offline and free.
- Testing — pytest, vitest, jest, node --test, cargo test: assertions and fault injection.
- Instrument an AI app — A chat endpoint or RAG feature: one run per conversation, its calls as Steps.
- Instrument an agent — Runs and Steps, including the parallel case inference gets wrong.
- Framework adapters — LangChain, Vercel AI SDK, LlamaIndex.
- Tuning Phase
- Running the search — Baseline or Discovery search, Modes, and when a Step settles.
- Report outcomes — Free evidence that settles Steps sooner and cheaper.
- The Report — Reading it, versions, finalizing, and exporting shimmy.lock.
- Production
- Managed or passthrough — Keep Shimmy in the path for drift and failover, or ship the lockfile alone.
- Alerts and webhooks — Drift, failover and spend — by email or signed webhook.
- Reference
- Client — run, step, wrap, report, verify — and the Phase options.
- Configuration — Every SHIMMY_* variable, its default, and how test runners change it.
- Files and CLI — shimmy.lock, Recording files, the request hash, and the shimmy command.
- Control plane — Workflows, Reports, billing, settings — every endpoint.
- Config as code — Mode, allowed providers and guardrails in git, not a dashboard.
- Types — Step kinds, outcome signals, annotations.
- Admin API — Model catalog and custom providers, managed at runtime.
- Billing