06Cost control
Gateway
Lower model bills, no code changes.
An OpenAI-compatible proxy: point your SDK's base_url at it and it caches identical requests, routes simple prompts to cheaper models, and meters spend — same request shape in and out.
Drop-in
OpenAI-compatible /v1/chat/completions. Change one base_url; everything else stays the same.
Cache
Identical requests are served from cache — repeat calls cost zero tokens and return instantly.
Smart routing
Short, non-complex prompts route to a cheaper model; quality-sensitive prompts keep the strong one.
Guardrails
Per-key rate limits, monthly spend caps, and X-Sentinel-* headers that report every decision.
Cost per 1,000 requests
Direct to provider0%
+ cheap-model routing0%
+ cache on repeats0%
Illustrative — a mixed workload with cache hits + routing.