Managed or passthrough
The Report told you which model each Step should use. Production is how you run on that answer — with Shimmy still in the path, or without it.
| Passthrough | Managed | |
|---|---|---|
| Calls go | straight to your provider | through Shimmy, on your provider keys |
| Each Step’s model | from shimmy.lock | from the Report, kept current |
| Outage, 429, refusal | SDK walks the Backup chain | Shimmy walks the Backup chain |
| A model gets worse | you won’t know | drift is detected and re-searched |
| Cache, spend caps, alerts | — | yes |
| Cost | free | flat monthly fee per Workflow |
Either way, Production calls run on your provider keys, not Shimmy’s.
Passthrough
Commit the shimmy.lock exported from the Report and run with SHIMMY_PHASE=production SHIMMY_PRODUCTION=passthrough:
import OpenAI from 'openai';
import { Shimmy } from '@rfa-labs/shimmy';
const shimmy = new Shimmy({
phase: 'production',
production: 'passthrough',
lockfile: 'shimmy.lock',
});
// In passthrough the client is NOT repointed: it keeps its own base URL and
// provider key. The SDK only rewrites each Step's model from the lockfile, and
// walks the Step's Backup chain if the winner is down, rate-limited or refuses.
const openai = shimmy.wrap(new OpenAI());
await shimmy.run('triage-inbox', async () => {
const res = await shimmy.step('classify', () =>
openai.chat.completions.create({
// Replaced by the lockfile's winner for `classify`. A Step the lockfile
// doesn't know is sent with this model, and the SDK warns once.
model: 'gpt-5.5',
messages: [{ role: 'user', content: 'Is this a bug or a feature request?' }],
}),
);
console.log(res.model, res.choices[0]?.message.content);
});
The SDK leaves your client alone — its own base URL, its own key — and only
replaces each Step’s model with the lockfile’s winner. If a call fails in a way
another model might not (the provider is down, a 429, an auth failure, the context
didn’t fit, a content refusal) it retries down the Step’s Backup chain. A request
every model would reject — a malformed one — fails at once.
A Step that isn’t in the lockfile is sent with the model you wrote, and the SDK
warns once. Rust’s chat() calls any OpenAI-compatible endpoint
(provider_base_url, defaulting to OPENAI_BASE_URL).
Managed
Keep SHIMMY_PHASE=production (managed is the default) and start a Production
subscription for the Workflow — dashboard → the Workflow → Production. Add your
provider keys under Settings → provider keys; managed calls run on them.
What managed Production does on every call:
- Serves the settled model, and on an outage, rate limit, refused key, context overflow or content refusal, fails over down the Backup chain — the backup on a different provider first. A context overflow skips straight past backups whose window is no larger.
- Hedges a Step that declared a latency budget: if the winner hasn’t answered by half the budget, the first backup starts too, and the first answer wins.
- Caches: an exact repeat of a request is answered from a persistent cache (30 days by default) without calling a provider.
- Watches for drift: a small sample of each Step’s calls — about 2%, and at least one a day — is graded in the background. When the winner stops passing, the next model in the Backup chain serves immediately and the Step is re-searched among providers you hold keys for, at no extra charge. The Report gets a new version when it settles.
- Enforces your spend cap and sends alerts.
Switching
The Phase is per process, so switching is a deploy, not a migration. Moving from
managed to passthrough means exporting a fresh shimmy.lock (a re-search may have
changed a winner) and setting SHIMMY_PRODUCTION=passthrough. Going back to
Tuning — a new Step, a new Mode — is SHIMMY_PHASE=tuning in staging while
Production keeps running on the current answer.