One API. Every model.
Route requests to OpenAI, Anthropic, Gemini, Azure, Bedrock, or Ollama. Sub-millisecond overhead. Single binary. No runtime.
Install · one line
cargo install crabllm crabctl
00 · Install
Fig. 00 · The Router. Requests enter as one format, pass through the gateway membrane, and fan out to provider zones. Hover a zone to simulate failover. Click for a burst of traffic.
“Same OpenAI/Anthropic format, any provider.”
What you get
Six capabilities the gateway gives you out of the box. Each one a plate; each plate a sentence; the sheet reads in any order you like.
01 / 06
Send OpenAI format. CrabLLM translates to Anthropic, Gemini, Bedrock, and Azure automatically.
02 / 06
Weighted random selection across providers. Exponential backoff retry. Automatic failover.
03 / 06
SSE proxied without buffering. Per-chunk extension hooks. Keep-alive pings.
04 / 06
Per-key model access control. Rate limiting, usage tracking, and budget enforcement.
05 / 06
SHA-256 response cache. Per-key RPM and TPM limits. Sliding window enforcement.
06 / 06
Per-key spend limits in USD. Automatic cost tracking from token usage and pricing config.
Performance
Gateway overhead at 5,000 concurrent requests per second. All gateways run with identical resource limits (2 CPUs, 512 MB) against a mock backend with instant responses. Full results →
| Gateway | P50 | P99 |
|---|---|---|
| CrabLLM | 0.26ms | 0.54ms |
| Bifrost | 0.61ms | 1.26ms |
| LiteLLM | 159ms | 227ms |
02-A · Latency
| Gateway | Peak RSS |
|---|---|
| CrabLLM | 34.9 MB |
| Bifrost | 171.7 MB |
| LiteLLM | 541.8 MB |
02-B · Memory
How it works
1 Configure
listen = "0.0.0.0:8080"
[providers.openai]
kind = "openai"
api_key = "${OPENAI_API_KEY}"
models = ["gpt-4o"]
[providers.anthropic]
kind = "anthropic"
api_key = "${ANTHROPIC_API_KEY}"
models = ["claude-sonnet-4-20250514"]
03-A · Config
2 Run
crabllm --config crabllm.toml
03-B · Launch
3 Send requests
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "claude-sonnet-4-20250514",
"messages": [{"role": "user", "content": "Hello!"}]}'
03-C · Request
Same OpenAI format, any provider. CrabLLM translates automatically.
Frequently asked questions
What's the overhead?
0.26ms P50 at 5,000 RPS. Rust with Tokio — no GC pauses, no interpreter. The gateway is not the bottleneck.
How is this different from LiteLLM?
Does it support streaming?
What about CrabTalk?
How much does it cost?
One API. Every model.
0.26ms · single binary