Hopscotch Watcher

every call keeps its row
Claims

What this measures. Hopscotch's pitch against OpenRouter and Vercel AI Gateway is not speed, it is accounting: the model you name is the model that runs, the price is the upstream list price with no markup, and every response says which upstream answered, how many attempts it took and whether a backup route served it. This watcher sends the same streaming prompt to the five most-used models on a schedule and checks every one of those promises per call, alongside the usual speed numbers. The prompt is five fixed questions with known answers, so every probe also yields a quality score: a cheap tripwire for silent quantization or a swapped model that the headers alone would not show.

Time to first tokenms, per probe, as the client saw it
Time outside the providerms, wall time minus provider_elapsed_ms
Visible tokens per secondfirst visible token to last
Cost per probeUSD at the catalog rate, same prompt every time
Known-answer score% of the five fixed questions answered correctly
timemodeloutcomeanswersfaultserved byproviderattemptsroute ttfttotalprovideroutside out tokreasoningcostfinishrequest id
Legend and method
Fidelity
The x-hopscotch-served-by header names the exact model id that answered. Fidelity is the share of served calls where it equals the id we named. Hopscotch promises this is always 100%: a named model is never swapped for a stand-in.
Attempts
x-hopscotch-retry-attempt-count: extra upstream attempts before this answer. 0 means the first try served it.
Route
x-hopscotch-served-via: backup is sent only when a backup route answered. Otherwise the ordinary route did.
Outcome
ok: full stream with [DONE] and a finish reason other than length. refused: the gateway said no before calling a provider (rate limit, credit, model not available). truncated: tokens arrived but the answer is not whole: the stream ended with an error frame, the provider stopped at max_tokens, or a 200 came back with no visible content. These are billed like successes, which is why they are counted separately. error: no tokens and a fault. timeout: our client gave up.
Answers
The prompt is five fixed questions with unambiguous answers (17×23, the symbol for gold, the year of the Moon landing, Australia's capital, bits in a byte). Each probe is scored out of five. A drop on one model while the served-by header still says the right name is the signature of a quantized or degraded deployment.
Judge
When TYPESAFE_API_KEY is set, each answer is graded by TypeSafe's System One model (a yes/no probability per question, plus a 0–2 format score), and every failed call is classified by fault owner (gateway, provider, account, request) and whether a retry would likely succeed. Without the key, a regex per question does the grading and faults are not classified.
List price
Hourly, the catalog rate from GET /v1/models is compared with the vendor's own published price (hand-maintained in config.json, with source and date). A difference beyond 0.5% is a markup or a stale table; either way it is flagged.
Outside the provider
Wall time of the whole call minus provider_elapsed_ms from the gateway's usage line. It bundles the gateway's own processing with the network between this watcher and the gateway, so read it as a ceiling on gateway overhead, not an exact figure.
Cost
Provider-reported token counts priced at the rate in GET /v1/models, which Hopscotch says is the upstream list price with no markup. Reasoning tokens bill as output.
Tiers
good / warning / serious / critical, from the worst of: last outcome, fidelity below 100%, success rate, mean known-answer score, median TTFT, median outside-provider time, any backup route or more than 10% multi-attempt calls.