[ 01 ] Methodology

How this is measured

Publishing the method is the point.
If you can’t reproduce a number, doubt it.

What runs, and when

Two tiers, on separate cron schedules.
Each tick divides its interval evenly.
A drifting cadence smears the hourly buckets.
Then a daily cycle looks like a regression.

Probes stream, with a 45-second timeout.
Probes are never retried.
A retry would hide the failures we measure.
So a rate limit is recorded as a rate limit.
Identical pulse prompts use prefix caching.
Anthropic uses cache_control.
OpenAI uses prompt_cache_key.
Prompts stay the same; no probe is skipped.
A cache hit measures the same call, cheaper.

The prompts

Verbatim, below.
Graders are deterministic.
They check strings and structure.
There is no LLM judge anywhere.
A judge drifts on its own schedule.
That drift looks just like a regression.

When a model gets called out

The labels: slow, failing, dumber, rambling.
The API calls all four degraded.
The label names the margin that was crossed.
What is wrong beats which enum tripped.
Other reasons read as struggling.
Each means one thing only.
The model left its own 7-day baseline.
It never means worse than another model.
The JSON, RSS and badge say degraded.
Anything built on them keeps working.

The baseline skips the last two hours.
Otherwise it would average the incident in.
A long outage would start to look normal.

Going bad takes two breaching windows.
Recovering takes three clean ones.
Flapping is worse than ten minutes late.
Feed, RSS and badge wait for that.
The board does not wait.
A model at the down threshold is named now.
So is one at 0.00× its own normal.
No visitor should wait a pulse to see it.
Offline models skip “dumbest vs usual”.
That line is for models still answering.

Under four samples reads unconfirmed.
The feed never calls that healthy.
A new model has no baseline yet.
It stays unconfirmed until it has one.
The board names a failing model at once.
It does not wait for the second pulse.
The tile says what it is doing.

What EZO Health is made of

One number: how a model is doing vs itself.
1.00× is exactly its own normal.
Above is a better day than usual.
0 is broken.
A naturally slow model can score 1.00×.
So it ranks how models are doing.
Not how good they are.
It is ours, not an uptime percentage.
Nobody else publishes it.
So here is all of it.

Four parts, each tied to a threshold above.
So score and verdict agree on what is broken.

Combined as a geometric mean.
That gives “0 when broken” its teeth.
An average hides one collapsed part.
A third of probes failing would score 0.75.
Multiplying lets any zero sink the score.
The upside caps at 1.25×.
That cap is load-bearing, not cosmetic.
The scale is asymmetric.
Uncapped, the noisiest model led on health.
Meanwhile the verdict called it cooked.
No baseline, or too few probes: no score.
It scores nothing, rather than scoring well.

What this does not measure

Prices

What each model costs now, per 1M tokens.
In USD, for a short prompt.
Each figure is from the vendor’s own page.
Read on the date shown.

Model prices in USD per million tokens
Model Input Cached Output Terms Checked Source

Card as of .
Long prompts and peak hours can cost more.
The terms column says when.

Feeds and API

Free, read-only, no key.
Every /api/* route is CORS-open.
Any origin may call it.

MCP

Same data as tools, at POST /mcp.
Stateless, JSON only, rate limited per IP.

claude mcp add --transport http ezo-status https://status.ezo.dev/mcp

A plain-text map lives at /llms.txt.