Shimmy · compared

Not another gateway.

LiteLLM, Portkey and Cloudflare AI Gateway route production traffic, and do it well. Shimmy does its work before you ship — finding the model each step should run on — and can stay out of your production path entirely. Here’s an honest look at what each does.

Feature by feature

The capabilities, side by side.

✓ = built in · ~ = partial / basic · — = not offered. We try hard to keep the competitor columns accurate; if we got one wrong, tell us.

CapabilityShimmyLiteLLMPortkeyCloudflare AI Gateway
Wraps your existing OpenAI / Anthropic client✓✓✓✓
Recorded responses replayed in tests (offline, no spend)✓~——
Per-step model search, graded against your own baseline✓———
Tool-call verification (declared tool + valid args)✓———
Open-weights migration (search only commercially licensed open models)✓———
A report of every model tried — cost, score, answers✓———
Lockfile you commit; production without a proxy in the path✓———
Fallback across providers on outages and rate limits✓✓✓~
Drift detection with automatic re-search✓———
Response caching✓✓✓✓
Spend caps and alerts✓✓✓~
Request logging and analytics dashboards~✓✓✓
Self-hosted VPC deployment✓✓~—
Open source—✓~—
Pricing modelDev free · Tuning at provider cost + 25% · managed production $4.99/workflow/mo, passthrough freeFree, open source (enterprise plan available)Subscription + usage tiersPer-request / Cloudflare plan
The honest version

When to pick them — and how they fit with Shimmy.

LiteLLM

An excellent open-source gateway and SDK for calling many providers through one interface, with routing, fallbacks, budgets and caching. It routes the models you configure; it doesn’t search for which model each step of your agent should use, or show you why. The two fit together: tune with Shimmy, then put the winning models in your LiteLLM config.

Portkey

A strong production gateway with observability — logs, analytics, guardrails, prompt management. If what you need is to see and govern live traffic, Portkey is solid. Shimmy answers an earlier question: which model each step should run on in the first place, measured on your inputs before launch.

Cloudflare AI Gateway

Caching, rate limiting and analytics at the edge, close to your users. It’s infrastructure for traffic you already route; it doesn’t choose models. Complementary: Shimmy’s passthrough calls whatever endpoint your client points at, Cloudflare included.

Try it on your app or agent

Every model it tried, and every answer it got.

Run one tuning pass and read the Report: the winner for each step, what it cost, and the answers behind the verdict.