Proof before routing
Most routers pick a cheaper model from a published average. Averages hide the cases your users see. Gimbix proves a route on your own requests first.
An OpenAI-compatible gateway
Point your client's base URL at Gimbix and keep your code. Requests are priced, repaired so provider prompt caches can serve them, and routed only where an approval exists.
Prompt-cache repair
When a prompt changes a line near the top on every call, the provider's cache misses. Gimbix moves that line after the user turn, so the stable part is cached. In our guarded runs, 92–93% of input was then served from cache with identical answers.
The rule a route must pass
The candidate must score at least as well as your current model on the selection half of your traffic. On the confirmation half it may make at most three answers worse, and no more than one net. Both halves are fixed by a salted hash.
Receipts
Each saving is tied to the approval that allowed it and the signed statement that recorded it. Receipts contain counts and digests only, so they can live where prompts cannot.
What the evidence shows so far
| Workload | Route and test | Answers | Cost | Verdict |
|---|---|---|---|---|
| Support triage | Luna answers first, Sol takes low-confidence cases (selection, 400 held-out) | 391 vs 390 | −95.9% | selected |
| Support triage | Same route on 148 fresh cases | 142 vs 141, 0 worse | −97.4% | confirmed |
| Support triage | Through the real gateway | 1 worse, 2 better | −97.8% | confirmed |
| Support triage | Guarded cloud run with an independent meter | 31 of 31 correct in both arms | −95.2% | pass |
| BANKING77 intents | Cheapest route matching GPT-6 Sol on 400 fresh messages | 16 worse, 15 better | −98.0% | not confirmed |
| BANKING77 intents | Cheapest route matching Gemini 3.8 Flash on 400 fresh messages | 21 worse, 16 better | −97.8% | not confirmed |
| BANKING77 intents | Cheapest route matching Claude Opus 5.5 (selection) | 354 vs 354, 259 of 400 escalated | −31.9% | selected |