Testing

Unit and integration tests for an AI app or agent, without mocking the LLM client: the Dev Phase replays Recordings, and the SDK logs every call so a test can assert on what your code did.

triage.test.ts
import { test } from 'node:test';
import assert from 'node:assert/strict';
import OpenAI from 'openai';
import { Shimmy, testing } from '@rfa-labs/shimmy';

const shimmy = new Shimmy();
const openai = shimmy.wrap(new OpenAI());

/** The code under test: two Steps, in order. */
async function triage(ticket: string): Promise<string> {
  return shimmy.run('triage', async () => {
    const label = await shimmy.step('classify', () =>
      openai.chat.completions.create({
        model: 'gpt-5.5',
        messages: [{ role: 'user', content: `Classify: ${ticket}` }],
      }),
    );
    const reply = await shimmy.step('draft_reply', () =>
      openai.chat.completions.create({
        model: 'gpt-5.5',
        messages: [
          { role: 'user', content: `Reply to a ${label.choices[0]?.message.content} ticket.` },
        ],
      }),
    );
    return reply.choices[0]?.message.content ?? '';
  });
}

test('triage classifies, then drafts', async () => {
  testing.reset();
  const reply = await triage('The export button does nothing');

  assert.ok(reply.length > 0);
  testing.assertStepsInOrder(['classify', 'draft_reply']);
  testing.assertStepCalled('classify', 1);
});

test('triage surfaces a rate limit instead of hiding it', async () => {
  testing.reset();
  // The next `classify` call fails the way the provider does: openai's own
  // RateLimitError, so this exercises your real error handling.
  await testing.withFault('classify', 'rate_limit', () =>
    assert.rejects(triage('The export button does nothing'), OpenAI.RateLimitError),
  );
  testing.assertStepNotCalled('draft_reply');
});
Compiled in CI. Run Python with pytest, TypeScript with node --test.

Setting up your test runner

The SDK notices it is under a test runner — pytest, vitest, jest, node --test, or the Rust crate’s testing feature — and defaults to replaying from shimmy/recordings/ in your repo. When CI is set, a missing Recording fails the test instead of recording one.

// vitest.config.ts — resets the call log and faults before every test.
// (jest: setupFilesAfterEnv. node --test: call testing.reset() yourself.)
import { defineConfig } from 'vitest/config';

export default defineConfig({
  test: { setupFiles: ['@rfa-labs/shimmy/testing-setup'] },
});

The pytest plugin is installed with the SDK; it adds the --shimmy-record, --shimmy-strict and --shimmy-recordings-dir options, resets the call log between tests, and provides a shimmy_calls fixture.

Asserting on Steps

PythonTypeScriptRust
Every call so fartesting.calls()testing.calls()testing::calls()
Called (n times)assert_step_called(s, times=)assertStepCalled(s, n)assert_step_called(s, Some(n))
Not calledassert_step_not_called(s)assertStepNotCalled(s)assert_step_not_called(s)
Orderassert_steps_in_order([...])assertStepsInOrder([...])assert_steps_in_order(&[...])
Clearreset()reset()reset()

Each logged call carries its Step, run, Phase, model and source — replay, fallback (a changed input replayed the latest Recording), recorded, remote, passthrough or fault — so a test can also check that nothing went to the network.

Making a Step fail

LLM calls fail in production in a handful of ways. Inject them to test your error handling — the next times calls of the Step (or any Step) fail like a provider would:

FaultPython / TypeScriptRust
timeoutopenai’s timeout errorShimmyError::Transport
rate_limitopenai’s RateLimitError (429)ShimmyError::Api { status: 429 }
server_erroropenai’s InternalServerError (500)ShimmyError::Api { status: 500 }
refusala completion that declinesthe same
malformed_jsona completion with broken JSONthe same

Python: with testing.inject_fault("classify", "rate_limit"):. TypeScript: await testing.withFault('classify', 'rate_limit', fn). Rust: let _guard = testing::inject_fault(Some("classify"), FaultKind::RateLimit, 1); — the fault is removed when the guard drops.

Integration tests against Shimmy

To prove the integration with Shimmy itself — keys, Steps, annotations — without paying for real answers, point the Dev source at the backend:

SHIMMY_DEV_SOURCE=remote pytest tests/integration

Calls then go to Shimmy, which replays its stored Recordings.