Managed or passthrough

The Report told you which model each Step should use. Production is how you run on that answer — with Shimmy still in the path, or without it.

PassthroughManaged
Calls gostraight to your providerthrough Shimmy, on your provider keys
Each Step’s modelfrom shimmy.lockfrom the Report, kept current
Outage, 429, refusalSDK walks the Backup chainShimmy walks the Backup chain
A model gets worseyou won’t knowdrift is detected and re-searched
Cache, spend caps, alerts—yes
Costfreeflat monthly fee per Workflow

Either way, Production calls run on your provider keys, not Shimmy’s.

Passthrough

Commit the shimmy.lock exported from the Report and run with SHIMMY_PHASE=production SHIMMY_PRODUCTION=passthrough:

passthrough.ts
import OpenAI from 'openai';
import { Shimmy } from '@rfa-labs/shimmy';

const shimmy = new Shimmy({
  phase: 'production',
  production: 'passthrough',
  lockfile: 'shimmy.lock',
});

// In passthrough the client is NOT repointed: it keeps its own base URL and
// provider key. The SDK only rewrites each Step's model from the lockfile, and
// walks the Step's Backup chain if the winner is down, rate-limited or refuses.
const openai = shimmy.wrap(new OpenAI());

await shimmy.run('triage-inbox', async () => {
  const res = await shimmy.step('classify', () =>
    openai.chat.completions.create({
      // Replaced by the lockfile's winner for `classify`. A Step the lockfile
      // doesn't know is sent with this model, and the SDK warns once.
      model: 'gpt-5.5',
      messages: [{ role: 'user', content: 'Is this a bug or a feature request?' }],
    }),
  );
  console.log(res.model, res.choices[0]?.message.content);
});
Compiled in CI.

The SDK leaves your client alone — its own base URL, its own key — and only replaces each Step’s model with the lockfile’s winner. If a call fails in a way another model might not (the provider is down, a 429, an auth failure, the context didn’t fit, a content refusal) it retries down the Step’s Backup chain. A request every model would reject — a malformed one — fails at once.

A Step that isn’t in the lockfile is sent with the model you wrote, and the SDK warns once. Rust’s chat() calls any OpenAI-compatible endpoint (provider_base_url, defaulting to OPENAI_BASE_URL).

Managed

Keep SHIMMY_PHASE=production (managed is the default) and start a Production subscription for the Workflow — dashboard → the Workflow → Production. Add your provider keys under Settings → provider keys; managed calls run on them.

What managed Production does on every call:

  • Serves the settled model, and on an outage, rate limit, refused key, context overflow or content refusal, fails over down the Backup chain — the backup on a different provider first. A context overflow skips straight past backups whose window is no larger.
  • Hedges a Step that declared a latency budget: if the winner hasn’t answered by half the budget, the first backup starts too, and the first answer wins.
  • Caches: an exact repeat of a request is answered from a persistent cache (30 days by default) without calling a provider.
  • Watches for drift: a small sample of each Step’s calls — about 2%, and at least one a day — is graded in the background. When the winner stops passing, the next model in the Backup chain serves immediately and the Step is re-searched among providers you hold keys for, at no extra charge. The Report gets a new version when it settles.
  • Enforces your spend cap and sends alerts.

Switching

The Phase is per process, so switching is a deploy, not a migration. Moving from managed to passthrough means exporting a fresh shimmy.lock (a re-search may have changed a winner) and setting SHIMMY_PRODUCTION=passthrough. Going back to Tuning — a new Step, a new Mode — is SHIMMY_PHASE=tuning in staging while Production keeps running on the current answer.