Instrument an AI app

Shimmy isn’t only for agents. A chat assistant, a RAG feature, an endpoint that summarizes, classifies or extracts — any app that calls an LLM — is a Workflow, and each place it calls one is a Step. An app may have two Steps where an agent has twenty; each still gets its own model, and the whole lifecycle — Recordings, Tuning, the Report, Production — works the same way.

ai-app.ts
import OpenAI from 'openai';
import { Shimmy } from '@rfa-labs/shimmy';

const shimmy = new Shimmy();
const openai = shimmy.wrap(new OpenAI());

const DOCS: Record<string, string> = {
  refund: 'Refunds are issued to the original card within 5 business days.',
  shipping: 'Orders ship within 24 hours; tracking arrives by email.',
};

// Your retrieval — a vector store, full-text search, an internal API. It makes
// no LLM call, so it is not a Step.
function searchDocs(query: string): string {
  return Object.entries(DOCS)
    .filter(([key]) => query.toLowerCase().includes(key))
    .map(([, text]) => text)
    .join('\n');
}

/** The request handler. Call it from Express, Next.js, Hono — anything. */
export async function handleMessage(conversationId: string, message: string): Promise<string> {
  // One request is one run. Keyed by the conversation, every turn of a chat
  // reads as one run, across restarts and replicas.
  return shimmy.run(
    'support-chat',
    async () => {
      // A Step is one LLM call site. An app usually has a handful, and each
      // gets its own model: rewriting a query is not answering it.
      const query = await shimmy.step('rewrite_query', { kind: 'formatting' }, async () => {
        const res = await openai.chat.completions.create({
          model: 'gpt-5.5',
          messages: [
            { role: 'system', content: 'Rewrite the question as a search query.' },
            { role: 'user', content: message },
          ],
        });
        return res.choices[0]?.message.content ?? message;
      });

      const context = searchDocs(query);

      return shimmy.step('answer', { kind: 'question_answering' }, async () => {
        const res = await openai.chat.completions.create({
          model: 'gpt-5.5',
          messages: [
            { role: 'system', content: `Answer only from:\n${context}` },
            { role: 'user', content: message },
          ],
        });
        const answer = res.choices[0]?.message.content ?? '';
        // Free evidence, per request: your app knows when an answer is empty.
        await shimmy.report({ schema_valid: answer.length > 0 });
        return answer;
      });
    },
    { id: conversationId },
  );
}

/** A thumbs up/down, arriving later in a separate request. */
export async function handleFeedback(conversationId: string, helpful: boolean): Promise<void> {
  // Outside the Step, so name it — and the run, by the same conversation id.
  await shimmy.report(
    { human_verdict: helpful ? 'accepted' : 'rejected' },
    { stepId: 'answer', runId: conversationId },
  );
}

console.log(await handleMessage('conv-42', 'How long does a refund take?'));
await handleFeedback('conv-42', true);
Compiled in CI against the real SDK.

Mapping your app

In your appIn Shimmy
The app or featureOne Workflow, with its own API key
One LLM call site — the query rewrite, the answerOne Step, named in code
One request, or one conversationOne run — pass the conversation id to keep every turn together
Retrieval, database reads, business logicNothing: only LLM calls are Steps

Two features with nothing in common — a support chat and an invoice extractor — are better as two Workflows: each gets its own Report, its own Stage and its own shimmy.lock.

A Step opened outside any run still works: it gets a single-step run of its own. Wrap the handler in a run when it makes more than one call, so the dashboard draws them as one request.

Dev: test the handler, not the model

Your request handler is a plain function, so test it like one. In the Dev Phase each Step’s call is recorded once and replayed after that — the test runs offline, deterministically and for free, and can assert which Steps ran. See Recordings and testing.

Tuning: replay real requests

Tuning needs inputs that look like production. For an app the easiest source is requests you already have: last week’s questions from your logs, a support-ticket export, a staging environment with real users. Send them through the same handler with SHIMMY_PHASE=tuning and each Step settles on its cheapest model that holds quality.

Apps often have a Step that only runs on some requests — an escalation, a refusal rewrite. The Report waits for every Expected Step; exclude a rare one on the Workflow’s Steps tab, or finalize the Report early. See running the search.

Free evidence from your users

An app knows things about its answers that a judge would charge to guess: the answer was empty, the JSON parsed, the user clicked thumbs up. Report it inside the Step, or later — a feedback endpoint names the Step and the conversation’s run, as handle_feedback does above. See report outcomes.

Shipping

When the Report is ready, ship on shimmy.lock in passthrough, or keep Shimmy in the path with managed Production — drift detection, failover and caching for the app, at a flat monthly fee per Workflow.

Next