Score every request before a single token is spent.
A hybrid classifier reads intent, context, tools, and risk. Rules first, embeddings second, an LLM only when the case is genuinely hard.
Steeros routes every prompt to the cheapest model that can handle it. Up to 85% lower spend, same quality. One proxy, no code changes.
Pick a flagship and the door flies open for everything. Teams burn 60-70% of their spend on prompts a small model handles perfectly.
Like taking a Ferrari to buy milk. Every single time.
A hybrid classifier reads intent, context, tools, and risk. Rules first, embeddings second, an LLM only when the case is genuinely hard.
Four modes, one policy: cheap, fast, balanced, quality. You set per-tier cost ceilings; the router spends inside them, every time.
Token caps get clamped, thinking blocks stripped, mid-chat system roles demoted. Each transform is logged so nothing surprises you.
Provider spend, cache reads and writes, classifier calls, compaction. Measured against a fair baseline, session and all-time.
The smallest model that clears the bar wins. You set the ceilings; the router spends inside them. One proxy in front of DeepSeek, Anthropic, OpenAI, and Gemini.
51% of a typical mix lands here
32% of a typical mix lands here
17% of a typical mix lands here
Every request is priced the way the provider actually bills it: cache reads, cache writes, full inputs. Our own traffic banked $22.79 in cache savings.
Savings are measured against the same token bundle on a fixed model, at that model’s own cache rates. No flattering math.
Held-out eval across chat, math, mcq, and code, scored blind against a judge and ground truth. Small samples on the non-chat categories.
Spend and carbon fall together. Fewer flagship tokens means fewer GPU-hours.
fewer flagship tokens, same output
Long sessions get compacted; past answers get reused, so repeat prompts skip straight to the right tier. 399 calls on our traffic, $0.12 total.
Monthly caps that degrade to the light tier. Automatic escalation when a cheap model fails. PII scrubbing on the roadmap.
Watch every route, token, and cent in real time, on your own machine at localhost:3000.
Slide your team size and see spend before Steeros, spend after, and what stays in your budget.
1 to 100 seats
10 to 300 prompts
Same output, roughly 65% fewer flagship tokens, so energy falls with the bill.
Modeled on our pilot mix: 51% light at $0.004, 32% balanced at $0.020, 17% flagship at $0.090 per request. Your distribution will differ.