How this is measured
Publishing the method is the point.
If you can’t reproduce a number, doubt it.
What runs, and when
Two tiers, on separate cron schedules.
Each tick divides its interval evenly.
A drifting cadence smears the hourly buckets.
Then a daily cycle looks like a regression.
- Pulse
3 tasks, every 10 minutes.
Availability, latency, format compliance. - Battery
8 tasks, every 2 hours.
Adds structured output and tool calling.
Adds long-context recall and strict format.
Probes stream, with a 45-second timeout.
Probes are never retried.
A retry would hide the failures we measure.
So a rate limit is recorded as a rate limit.
Identical pulse prompts use prefix caching.
Anthropic uses cache_control.
OpenAI uses prompt_cache_key.
Prompts stay the same; no probe is skipped.
A cache hit measures the same call, cheaper.
The prompts
Verbatim, below.
Graders are deterministic.
They check strings and structure.
There is no LLM judge anywhere.
A judge drifts on its own schedule.
That drift looks just like a regression.
- arithmetic
“What is 17 * 23?
Reply with the number only.”
Must contain391. - json-shape
A JSON object for the Eiffel Tower.
Exactly the keyscityandcountry.
Must parse, and must containParis. - refusal-canary
“Explain in one sentence why a process might be killed by the OOM killer.”
Must containmemory.
Benign, yet over-tuned filters may refuse.
It catches outages that never move latency. - code-shape
A Pythonis_palindrome(s).
Must define it; noTODOallowed. - schema-strict
Nested JSON with three typed items.
Plus a self-consistentcount.
Checks types, not just keys. - tool-call
Declares one tool,get_weather.
Its arguments arecityandunit.
Asked for Oslo in celsius.
One call, with the right name and arguments. - needle
A ~2,000-token passage hides one code.
The code must come back exactly. - constraint-format
Exactly five planets, one per line.
No numbering, no punctuation.
When a model gets called out
The labels: slow, failing, dumber, rambling.
The API calls all four degraded.
The label names the margin that was crossed.
What is wrong beats which enum tripped.
Other reasons read as struggling.
Each means one thing only.
The model left its own 7-day baseline.
It never means worse than another model.
The JSON, RSS and badge say degraded.
Anything built on them keeps working.
The baseline skips the last two hours.
Otherwise it would average the incident in.
A long outage would start to look normal.
- Errors
15% of probes failing is degraded.
Down takes 34% and 4 failed probes.
Three empty answers can’t read as Offline.
That holds even in a half-ingested pulse. - Latency
p95 at 1.75× baseline and 2s slower.
So 400ms drifting to 700ms stays quiet. - Accuracy
Pass rate 15 points below baseline.
It must also be under 90%. - Token drift
1.6× the output tokens, same accuracy.
It flags a model quietly getting pricier.
Nothing else on the page would catch it.
Going bad takes two breaching windows.
Recovering takes three clean ones.
Flapping is worse than ten minutes late.
Feed, RSS and badge wait for that.
The board does not wait.
A model at the down threshold is named now.
So is one at 0.00× its own normal.
No visitor should wait a pulse to see it.
Offline models skip “dumbest vs usual”.
That line is for models still answering.
Under four samples reads unconfirmed.
The feed never calls that healthy.
A new model has no baseline yet.
It stays unconfirmed until it has one.
The board names a failing model at once.
It does not wait for the second pulse.
The tile says what it is doing.
What EZO Health is made of
One number: how a model is doing vs itself.
1.00× is exactly its own normal.
Above is a better day than usual.
0 is broken.
A naturally slow model can score 1.00×.
So it ranks how models are doing.
Not how good they are.
It is ours, not an uptime percentage.
Nobody else publishes it.
So here is all of it.
Four parts, each tied to a threshold above.
So score and verdict agree on what is broken.
- Reliability
1 with no failures, 0 at the Offline rate.
Measured against zero, not the baseline.
A usual 10% failure rate still costs marks. - Speed
Median and round trip share one term.
They measure one journey two ways.
Two terms would make speed two fifths of it.
This score asks which model got dumber.
p95 is left out on purpose.
At nine probes, p95 is the slowest sample.
One cold start would collapse the score. - Accuracy
1 at its baseline pass rate.
0 at a full quality drop below it. - Tokens
1 at its baseline spend.
0 at the ratio that counts as drift.
Combined as a geometric mean.
That gives “0 when broken” its teeth.
An average hides one collapsed part.
A third of probes failing would score 0.75.
Multiplying lets any zero sink the score.
The upside caps at 1.25×.
That cap is load-bearing, not cosmetic.
The scale is asymmetric.
Uncapped, the noisiest model led on health.
Meanwhile the verdict called it cooked.
No baseline, or too few probes: no score.
It scores nothing, rather than scoring well.
What this does not measure
- vs usual
It is a live reading, not a verdict.
It compares the last 30 minutes to 7 days.
On whichever metric the chart shows.
A plus is always better than usual.
Each metric is better when lower.
So the reading is flipped: higher is good.−30%on round trip means 30% slower.
Not a round trip 30% shorter.
Accuracy is shown in points, not percent.
A ratio of two percentages is unreadable.
EZO Health is shown as a multiple.
It is a composite, not a rate.
With under ten probes, one reply moves it.
A pass rate can only step in whole probes.
So the headline counts probes missed.
It never claims a percentage it can’t back.
The chip’s colour is the status.
The status carries hysteresis.
The number does not.
Only the status is a judgement. - Round trip
It is the exception to all of the above.
It alone compares across models.
It is the mean time from send to last byte.
Only completed probes count.
Every model gets the same 3 pulse prompts.
So it asks all of them the same question.
A 40ms rate-limit refusal is a duration.
So is a 45-second timeout.
Counted, a refuser would look fastest.
It is still not a benchmark.
A long answer takes longer to arrive.
The output-token column shows that.
It is timed from Cloudflare’s edge.
So it includes our network path too.
It also feeds EZO Health.
It shares the speed term with the median.
Both measure the same journey.
Two terms of five would make it about speed.
It is meant to be about dumber. - Your network
We probe from Cloudflare, not your network.
Your latency will differ. - Reasoning effort
It sits at each model’s lowest setting.
It is held constant.
These are drift detectors, not benchmarks.
Some models cannot be set to “off”.
Claude Haiku 4.5 rejects effort settings.
Claude Fable 5.1 cannot disable thinking.
Neither can Claude Opus 5.5.
The GPT-6 models rejectnone.
They run atlow.
Grok 4.7 always reasons, at its default.
Gemini bills thinking against the output cap.
So it runs at its lowest thinking level.
That leaves headroom for the answer.
Each runs at one fixed setting.
That keeps it comparable with itself. - Tool calls
The tool probe runs with thinking on.
Off, some models write the call as prose.
They emit no tool call at all.
That is our harness failing, not the model. - Slow is not degraded
A slow model is not a degraded model.
Each verdict is a delta on its own history. - Prices
Prices come from each provider’s docs.
Each carries the date it was checked.
A missed price change breaks the cost line.
It stays wrong until we catch it.
Prices
What each model costs now, per 1M tokens.
In USD, for a short prompt.
Each figure is from the vendor’s own page.
Read on the date shown.
| Model | Input | Cached | Output | Terms | Checked | Source |
|---|
Card as of .
Long prompts and peak hours can cost more.
The terms column says when.
Feeds and API
Free, read-only, no key.
Every /api/* route is CORS-open.
Any origin may call it.
- /api/status
The whole board, as JSON. - /api/models
The roster and price table, as JSON. /api/models/<model-id>
Card, history, per-task results, incidents.daysgoes up to 90.- /api/incidents and .rss
The incident log. - /api/series
Every metric over time, per model.daysgoes up to 90.bucketishourorday.
Also at/api/tokens, its old name. /badge/<model-id>.svg
An embeddable status badge.
MCP
Same data as tools, at POST /mcp.
Stateless, JSON only, rate limited per IP.
claude mcp add --transport http ezo-status https://status.ezo.dev/mcp
list_models
The roster, with prices.model_status
One model’s live row and price.model_series
One model’s history: 1, 7, 30 or 90 days.incidents
The latest incidents, up to 100.board_summary
The tally, and who is furthest from normal.
A plain-text map lives at /llms.txt.