Instrument an AI app
Shimmy isn’t only for agents. A chat assistant, a RAG feature, an endpoint that summarizes, classifies or extracts — any app that calls an LLM — is a Workflow, and each place it calls one is a Step. An app may have two Steps where an agent has twenty; each still gets its own model, and the whole lifecycle — Recordings, Tuning, the Report, Production — works the same way.
import OpenAI from 'openai';
import { Shimmy } from '@rfa-labs/shimmy';
const shimmy = new Shimmy();
const openai = shimmy.wrap(new OpenAI());
const DOCS: Record<string, string> = {
refund: 'Refunds are issued to the original card within 5 business days.',
shipping: 'Orders ship within 24 hours; tracking arrives by email.',
};
// Your retrieval — a vector store, full-text search, an internal API. It makes
// no LLM call, so it is not a Step.
function searchDocs(query: string): string {
return Object.entries(DOCS)
.filter(([key]) => query.toLowerCase().includes(key))
.map(([, text]) => text)
.join('\n');
}
/** The request handler. Call it from Express, Next.js, Hono — anything. */
export async function handleMessage(conversationId: string, message: string): Promise<string> {
// One request is one run. Keyed by the conversation, every turn of a chat
// reads as one run, across restarts and replicas.
return shimmy.run(
'support-chat',
async () => {
// A Step is one LLM call site. An app usually has a handful, and each
// gets its own model: rewriting a query is not answering it.
const query = await shimmy.step('rewrite_query', { kind: 'formatting' }, async () => {
const res = await openai.chat.completions.create({
model: 'gpt-5.5',
messages: [
{ role: 'system', content: 'Rewrite the question as a search query.' },
{ role: 'user', content: message },
],
});
return res.choices[0]?.message.content ?? message;
});
const context = searchDocs(query);
return shimmy.step('answer', { kind: 'question_answering' }, async () => {
const res = await openai.chat.completions.create({
model: 'gpt-5.5',
messages: [
{ role: 'system', content: `Answer only from:\n${context}` },
{ role: 'user', content: message },
],
});
const answer = res.choices[0]?.message.content ?? '';
// Free evidence, per request: your app knows when an answer is empty.
await shimmy.report({ schema_valid: answer.length > 0 });
return answer;
});
},
{ id: conversationId },
);
}
/** A thumbs up/down, arriving later in a separate request. */
export async function handleFeedback(conversationId: string, helpful: boolean): Promise<void> {
// Outside the Step, so name it — and the run, by the same conversation id.
await shimmy.report(
{ human_verdict: helpful ? 'accepted' : 'rejected' },
{ stepId: 'answer', runId: conversationId },
);
}
console.log(await handleMessage('conv-42', 'How long does a refund take?'));
await handleFeedback('conv-42', true);
Mapping your app
| In your app | In Shimmy |
|---|---|
| The app or feature | One Workflow, with its own API key |
| One LLM call site — the query rewrite, the answer | One Step, named in code |
| One request, or one conversation | One run — pass the conversation id to keep every turn together |
| Retrieval, database reads, business logic | Nothing: only LLM calls are Steps |
Two features with nothing in common — a support chat and an invoice extractor —
are better as two Workflows: each gets its own Report, its own Stage and its own shimmy.lock.
A Step opened outside any run still works: it gets a single-step run of its own. Wrap the handler in a run when it makes more than one call, so the dashboard draws them as one request.
Dev: test the handler, not the model
Your request handler is a plain function, so test it like one. In the Dev Phase each Step’s call is recorded once and replayed after that — the test runs offline, deterministically and for free, and can assert which Steps ran. See Recordings and testing.
Tuning: replay real requests
Tuning needs inputs that look like production. For an app the easiest source is
requests you already have: last week’s questions from your logs, a support-ticket
export, a staging environment with real users. Send them through the same
handler with SHIMMY_PHASE=tuning and each Step settles on its cheapest model
that holds quality.
Apps often have a Step that only runs on some requests — an escalation, a refusal rewrite. The Report waits for every Expected Step; exclude a rare one on the Workflow’s Steps tab, or finalize the Report early. See running the search.
Free evidence from your users
An app knows things about its answers that a judge would charge to guess: the
answer was empty, the JSON parsed, the user clicked thumbs up. Report it inside
the Step, or later — a feedback endpoint names the Step and the conversation’s
run, as handle_feedback does above. See report outcomes.
Shipping
When the Report is ready, ship on shimmy.lock in passthrough, or keep Shimmy in the path with managed
Production — drift detection, failover and caching for the app, at a flat monthly
fee per Workflow.
Next
- Instrument an agent — when the app grows parallel or looping Steps.
- Framework adapters — if LangChain, LlamaIndex or the Vercel AI SDK owns your client.