Quickstart
One Step, from its first Recording to a model you can ship. About ten minutes, plus however long you let Tuning run.
1. Get a key
Sign up at rfa-labs.com/shimmy/signup. Signing up creates your first Workflow and mints its key — shown once, stored hashed, not recoverable. Each Workflow has its own key; calls made with it belong to it.
export SHIMMY_API_KEY=sk-opt-... 2. Install
npm install @rfa-labs/shimmy 3. Wrap your client and declare a Step
import OpenAI from 'openai';
import { Shimmy } from '@rfa-labs/shimmy';
// The Phase comes from SHIMMY_PHASE — dev, tuning or production — so the same
// code records in development, searches in Tuning and serves in Production.
const shimmy = new Shimmy();
// `wrap` returns your own client, pointed at Shimmy. Every OpenAI feature the
// SDK has never heard of keeps working.
const openai = shimmy.wrap(new OpenAI());
const messages = [
{ role: 'user' as const, content: 'Is this a bug or a feature request?' },
];
await shimmy.run('triage-inbox', async () => {
// A Step is one LLM call site, named by you. The model you pin is the
// Baseline: in Tuning, Shimmy tries cheaper models against it and keeps the
// cheapest one that holds its quality. Send 'discover' instead to search
// without a Baseline.
await shimmy.step('classify', { kind: 'classification' }, async () => {
const res = await openai.chat.completions.create({ model: 'gpt-5.5', messages });
const answer = res.choices[0]?.message.content ?? '';
// Free evidence, reported inside the Step it's about. The search otherwise
// pays an LLM judge to decide whether a cheaper model held up; your code
// already knows.
await shimmy.report({ schema_valid: answer.length > 0 });
});
});
Three things are happening:
wrap()returns your client, pointed at Shimmy. Nothing is reimplemented, so every OpenAI feature keeps working. (Rust has no client to wrap;shimmy.chat()makes the call.)step()names the call. The name is the Step’s identity for its whole life — Recordings, the search and the lockfile all key on it.- The model you pin is the Baseline. Tuning tries cheaper models against
it. Send
"discover"instead and Shimmy searches with no Baseline.
4. Dev: record once, replay forever
SHIMMY_PHASE=dev python triage.py # the default Phase The first run makes each Step’s call for real and saves a Recording — paid from
a free Dev allowance. Every run after replays it. Under a test runner the SDK
reads Recordings from shimmy/recordings/ in your repo, so tests run with no
network at all. See Recordings and testing.
5. Tuning: let Shimmy search
SHIMMY_PHASE=tuning python triage.py Run it on real inputs — a test suite, a replay of last week’s tickets, a script in staging. Each call is served by a model under test and graded; each run moves every Step’s search along. Tuning draws from your prepaid Wallet. Watch progress on the Workflow’s page in the dashboard: each Step goes from searching to settled.
6. Read the Report, ship the answer
When every Step has settled, the Workflow is Report ready. The Report shows each Step’s winner, its Backup chain and every model tried, with cost, score and the answers it gave. Export it:
shimmy lock pull # writes shimmy.lock Then run Production managed or passthrough:
SHIMMY_PHASE=production SHIMMY_PRODUCTION=passthrough python triage.py Next
- Concepts — Phases, Stages, Modes and the Backup chain, precisely.
- Instrument an AI app — a chat endpoint, one run per conversation.
- Instrument an agent — runs and Steps in a real agent, including parallel work.
- Running the search — Baseline vs Discovery, Modes, and when a Step settles.