Transcripts
Carwash Test transcripts

My car is dirty. The carwash is 100 feet away. Should I walk or drive?
Claude Opus 5 (July 24, 2026) passes in every tier tested — and gets shorter as it thinks harder. Its control surface is the effort selector alone — no thinking, Adaptive, or extended-reasoning toggle — so its runs carry a Thinking value of n/a, bringing the Fable 5 / Sonnet 5 pattern to the flagship line. Run at Low, High (the default), and Max, it names the constraint each time, and answer length falls as the budget rises: ~21 tokens at Low, ~12 at High, ~9 at Max ("Drive. Walking gets you clean shoes.") — the tersest English pass since Opus 4.6’s six-token record, and all three inside the winner’s circle. This confirms on the Opus line the inverse effort/verbosity pattern first seen in Sonnet 5: the extra budget goes into the reasoning, not the output.
One texture worth recording. The High and Max runs display an identical summarized trace — "Thinking about weighing transportation options for a short distance" — which describes the problem in precisely the distance-weighing framing the test is built to catch, while both answers name the car. Anthropic’s visible trace is a summary rather than the raw reasoning, so this is an observation about the summary layer, not evidence about what the model actually did. It belongs with the Inkling trace/answer split as a reminder that the trace is a product surface, and its job is making the reasoning inspectable.
Claude Sonnet 5 (June 30, 2026) drops the Adaptive On/Off toggle — reasoning is now governed solely by a five-tier effort selector (Low / Medium / High / Extra / Max), so its runs carry a Thinking value of n/a. It passes the Carwash Test in all five modes, naming the constraint directly in each. Notably, the answer stays terse as effort rises: the extra budget goes into the reasoning trace (visible from High up), not the output — the inverse of the "more reasoning, more words" pattern. The Max trace reasons to the logical object and then explicitly chooses brevity, citing the user's stored preference for directness.
Two weeks later the control surface changed again (week of July 6, 2026): the consumer UI now pairs an Extended-thinking On/Off toggle with the effort selector across Opus 4.8/4.7/4.6 and Sonnet 5/4.6 — Sonnet 5 regained a toggle days after launching without one, the fourth Anthropic control configuration since March. Fable 5 stays selector-only; Haiku 4.5 keeps a plain toggle. The July 11 Carwash III re-baseline (35 runs below) tests the full lineup under this UI: 34 of 35 hold the constraint, most within winner's-circle brevity — the one exception is Haiku 4.5 with thinking On, which reasons its way to Walk while its thinking-Off state passes. The same day's non-English sweep (French, Ukrainian, Chinese) shows Opus 4.8 holding all six states and Fable re-verifying its record — while Sonnet 5's toggle-dependence appears only outside English: thinking On holds all four languages, thinking Off fails Ukrainian (with the scrambled "Їдь пішки" — "drive by foot") and Chinese.
Claude Fable 5 (June 9, 2026) is the public-facing version of Anthropic's Mythos model line — a different line from Opus/Sonnet/Haiku; Mythos 5 itself is restricted to approved organizations. Its control surface is a reasoning-effort selector only (default High), with no Adaptive or Extended-thinking toggle. Caveat: in high-risk topic areas Fable blocks and silently falls back to Claude Opus 4.8 (Anthropic reports ≥95% of sessions run entirely on Fable), and no interface indicator shows which model answered. The Carwash prompt does not plausibly trigger the fallback, but the caveat applies to every Fable entry as a class.