Checking…
Range scopes the cards and table below. The verdict above always reflects each model’s most recent run, whatever range is selected.
One card per model: its latest response time, how that compares with its own
recent runs, and every run in the selected range. Typical is the median of
that model’s earlier runs in the range. Every card shares the same axes, and the faint
grey lines behind each trace are the other models, for scale.
The y axis is linear by default, so the gaps read at face value. Switch it to
log when one spike flattens everything else — responses range from seconds to
minutes, and a linear axis gives the tallest run all the room. Switch to
vs. its own normal to plot each model as a percentage of its own median,
which is how the verdict is decided. A hollow dot means the response came back empty
or was refused.
The latest run from each model, split three ways: time on the wire
(transport), time spent thinking before a single word
appears, and time spent writing. Thinking is the part you sit through in
silence, and on most reasoning models it is the largest slice.
Some providers send nothing at all until reasoning has finished, which makes transport and
thinking arrive as one indistinguishable block. Those models are marked
not separable rather than being shown with a thinking time of zero — the
silence is real either way, we just cannot say where it went.
The clock is split three ways. Time up to the first byte of the response
covers network round-trip, TLS, and any queueing before the provider starts replying. From
there to the first visible word, the model is thinking. After that it is
writing.
Network transit and provider queueing cannot be separated inside that first segment, since
both finish before the server sends anything. From a single region it stays a small and
fairly stable slice.
Every model gets the same prompt on a schedule, and we time the reply. Nothing is executed and nothing is graded.
—) is
stored on every run, and the charts mark the point where it changes, so an edit
shows up as a break rather than passing for a change in speed.Medians across the selected range. Two caveats: rate is not
generation speed for models that reason silently and then send the answer at once, and
output tokens are not comparable between Opus 4.6 and 4.7+ (different
tokenizers). Every model runs at medium effort with adaptive thinking.
Currently tracking Claude Opus 5, Claude Fable 5 and Claude Sonnet 5 from Anthropic, alongside the older Claude Opus 4.8, 4.7 and 4.6; GPT-5.6 Sol and GPT-5.6 Terra from OpenAI; and Grok 4.5 from xAI.