Testing
Unit and integration tests for an AI app or agent, without mocking the LLM client: the Dev Phase replays Recordings, and the SDK logs every call so a test can assert on what your code did.
import { test } from 'node:test';
import assert from 'node:assert/strict';
import OpenAI from 'openai';
import { Shimmy, testing } from '@rfa-labs/shimmy';
const shimmy = new Shimmy();
const openai = shimmy.wrap(new OpenAI());
/** The code under test: two Steps, in order. */
async function triage(ticket: string): Promise<string> {
return shimmy.run('triage', async () => {
const label = await shimmy.step('classify', () =>
openai.chat.completions.create({
model: 'gpt-5.5',
messages: [{ role: 'user', content: `Classify: ${ticket}` }],
}),
);
const reply = await shimmy.step('draft_reply', () =>
openai.chat.completions.create({
model: 'gpt-5.5',
messages: [
{ role: 'user', content: `Reply to a ${label.choices[0]?.message.content} ticket.` },
],
}),
);
return reply.choices[0]?.message.content ?? '';
});
}
test('triage classifies, then drafts', async () => {
testing.reset();
const reply = await triage('The export button does nothing');
assert.ok(reply.length > 0);
testing.assertStepsInOrder(['classify', 'draft_reply']);
testing.assertStepCalled('classify', 1);
});
test('triage surfaces a rate limit instead of hiding it', async () => {
testing.reset();
// The next `classify` call fails the way the provider does: openai's own
// RateLimitError, so this exercises your real error handling.
await testing.withFault('classify', 'rate_limit', () =>
assert.rejects(triage('The export button does nothing'), OpenAI.RateLimitError),
);
testing.assertStepNotCalled('draft_reply');
});
Setting up your test runner
The SDK notices it is under a test runner — pytest, vitest, jest, node --test,
or the Rust crate’s testing feature — and defaults to replaying from shimmy/recordings/ in your repo. When CI is set, a missing Recording fails the
test instead of recording one.
// vitest.config.ts — resets the call log and faults before every test.
// (jest: setupFilesAfterEnv. node --test: call testing.reset() yourself.)
import { defineConfig } from 'vitest/config';
export default defineConfig({
test: { setupFiles: ['@rfa-labs/shimmy/testing-setup'] },
}); The pytest plugin is installed with the SDK; it adds the --shimmy-record, --shimmy-strict and --shimmy-recordings-dir options, resets the call log
between tests, and provides a shimmy_calls fixture.
Asserting on Steps
| Python | TypeScript | Rust | |
|---|---|---|---|
| Every call so far | testing.calls() | testing.calls() | testing::calls() |
| Called (n times) | assert_step_called(s, times=) | assertStepCalled(s, n) | assert_step_called(s, Some(n)) |
| Not called | assert_step_not_called(s) | assertStepNotCalled(s) | assert_step_not_called(s) |
| Order | assert_steps_in_order([...]) | assertStepsInOrder([...]) | assert_steps_in_order(&[...]) |
| Clear | reset() | reset() | reset() |
Each logged call carries its Step, run, Phase, model and source — replay, fallback (a changed input replayed the latest Recording), recorded, remote, passthrough or fault — so a test can also check that nothing went to the
network.
Making a Step fail
LLM calls fail in production in a handful of ways. Inject them to test your error
handling — the next times calls of the Step (or any Step) fail like a provider
would:
| Fault | Python / TypeScript | Rust |
|---|---|---|
timeout | openai’s timeout error | ShimmyError::Transport |
rate_limit | openai’s RateLimitError (429) | ShimmyError::Api { status: 429 } |
server_error | openai’s InternalServerError (500) | ShimmyError::Api { status: 500 } |
refusal | a completion that declines | the same |
malformed_json | a completion with broken JSON | the same |
Python: with testing.inject_fault("classify", "rate_limit"):. TypeScript: await testing.withFault('classify', 'rate_limit', fn). Rust: let _guard = testing::inject_fault(Some("classify"), FaultKind::RateLimit, 1); —
the fault is removed when the guard drops.
Integration tests against Shimmy
To prove the integration with Shimmy itself — keys, Steps, annotations — without paying for real answers, point the Dev source at the backend:
SHIMMY_DEV_SOURCE=remote pytest tests/integration Calls then go to Shimmy, which replays its stored Recordings.