Files and CLI

Two files live in your repo — Recordings and shimmy.lock — and the shimmy command keeps them current.

The CLI

Shipped with the Python SDK (shimmy) and the TypeScript SDK (npx shimmy):

bash
export SHIMMY_API_KEY=sk-opt-…  SHIMMY_WORKFLOW_ID=…
npx shimmy recordings pull    # backend Recordings → shimmy/recordings/
npx shimmy recordings diff --exit-code
npx shimmy lock pull          # the Report's shimmy.lock
CommandDoes
shimmy recordings pullWrites the Workflow’s backend Recordings under --dir (default shimmy/recordings), redacted.
shimmy recordings diffLists what a pull would add or change. --exit-code exits 1 if anything would.
shimmy lock pullWrites the Report’s lockfile to -o (default shimmy.lock).

Every command takes --workflow <id>, or reads SHIMMY_WORKFLOW_ID, and authenticates with SHIMMY_API_KEY. --no-redact writes Recordings exactly as stored.

shimmy.lock

{
  "shimmy_lock": 1,
  "workflow": "099730af-41ae-44d2-8885-e0d2b10ae305",
  "workflow_name": "Support triage",
  "report_version": 3,
  "generated_at": "2026-10-02T20:06:31Z",
  "steps": {
    "classify": {
      "step_fingerprint": "da4f6f8aa1bc7385",
      "model": "gpt-4.1",
      "provider": "openai",
      "backups": [{ "model": "claude-haiku-4-5", "provider": "anthropic" }],
      "baseline": "gpt-5.5"
    }
  }
}

Steps are keyed by their declared id. A Step that was never declared is keyed fp:<fingerprint> — it’s in the file for the record, but passthrough can only match declared ids, so declare every Step you ship. backups is the Backup chain after the winner, in order; baseline is the pinned model, or null for a Discovery search.

Recording files

<recordings dir>/<step id>/<request hash>.json, with the hash’s : written as - — for example shimmy/recordings/classify/v1-814066…df33a.json:

{
  "step_id": "classify",
  "request_hash": "v1:814066310e3bb4010fa5863e66b5c12a514fbe909bea7756e312e02de66df33a",
  "request": { "messages": [ … ], "tools": [], "json": false },
  "completion": {
    "content": "bug",
    "model": "gpt-5.5",
    "tool_calls": null,
    "usage": { "prompt_tokens": 18, "completion_tokens": 1 }
  },
  "created_at": "2026-10-02T20:05:03Z"
}

Only completion is replayed. request is there so a reviewer can see what the Recording answers; editing it changes nothing.

The request hash

The hash that matches a call to its Recording is v1: followed by the SHA-256 of this JSON, serialized with object keys sorted and no whitespace:

{
  "messages": [{ "role": "…", "content": "…", "tool_call_id": null, "tool_calls": null }],
  "tools": ["sorted", "unique", "tool names"],
  "json": false
}

A message’s content is its text, with the text parts of a multi-part message joined by newlines. json is true for JSON mode (response_format: json_object). Everything else — the model, temperature, max tokens — is left out on purpose, so those can change without invalidating Recordings.