What this measures. Hopscotch's pitch against OpenRouter and Vercel AI Gateway is not speed, it is accounting: the model you name is the model that runs, the price is the upstream list price with no markup, and every response says which upstream answered, how many attempts it took and whether a backup route served it. This watcher sends the same streaming prompt to the five most-used models on a schedule and checks every one of those promises per call, alongside the usual speed numbers. The prompt is five fixed questions with known answers, so every probe also yields a quality score: a cheap tripwire for silent quantization or a swapped model that the headers alone would not show.
| time | model | outcome | answers | fault | served by | provider | attempts | route | ttft | total | provider | outside | out tok | reasoning | cost | finish | request id |
|---|
Legend and method
- Fidelity
- The
x-hopscotch-served-byheader names the exact model id that answered. Fidelity is the share of served calls where it equals the id we named. Hopscotch promises this is always 100%: a named model is never swapped for a stand-in. - Attempts
x-hopscotch-retry-attempt-count: extra upstream attempts before this answer. 0 means the first try served it.- Route
x-hopscotch-served-via: backupis sent only when a backup route answered. Otherwise the ordinary route did.- Outcome
- ok: full stream with
[DONE]and a finish reason other thanlength. refused: the gateway said no before calling a provider (rate limit, credit, model not available). truncated: tokens arrived but the answer is not whole: the stream ended with an error frame, the provider stopped atmax_tokens, or a 200 came back with no visible content. These are billed like successes, which is why they are counted separately. error: no tokens and a fault. timeout: our client gave up. - Answers
- The prompt is five fixed questions with unambiguous answers (17×23, the symbol for gold, the year of the Moon landing, Australia's capital, bits in a byte). Each probe is scored out of five. A drop on one model while the served-by header still says the right name is the signature of a quantized or degraded deployment.
- Judge
- When
TYPESAFE_API_KEYis set, each answer is graded by TypeSafe's System One model (a yes/no probability per question, plus a 0–2 format score), and every failed call is classified by fault owner (gateway, provider, account, request) and whether a retry would likely succeed. Without the key, a regex per question does the grading and faults are not classified. - List price
- Hourly, the catalog rate from
GET /v1/modelsis compared with the vendor's own published price (hand-maintained inconfig.json, with source and date). A difference beyond 0.5% is a markup or a stale table; either way it is flagged. - Outside the provider
- Wall time of the whole call minus
provider_elapsed_msfrom the gateway's usage line. It bundles the gateway's own processing with the network between this watcher and the gateway, so read it as a ceiling on gateway overhead, not an exact figure. - Cost
- Provider-reported token counts priced at the rate in
GET /v1/models, which Hopscotch says is the upstream list price with no markup. Reasoning tokens bill as output. - Tiers
- good / warning / serious / critical, from the worst of: last outcome, fidelity below 100%, success rate, mean known-answer score, median TTFT, median outside-provider time, any backup route or more than 10% multi-attempt calls.