These charts read directly from the live dataset and update as runs are added. A standing caveat applies throughout: the test is single-shot, the sample per cell is small, and each snapshot reflects whatever models were available that date — so the time-based charts describe the evolving field, not a controlled trend in any one system. See the methodology for scoring and limitations.

Dataset overview

Over time

A note on how to read these. The first three charts are cumulative — each point covers every run recorded up to that date. That makes them stable measures of what the whole corpus says, but it also means a small recent batch barely moves them: by July the dataset held roughly two hundred runs, so adding a dozen cannot shift a median, and a cumulative count can only rise or flatten, never fall. The last chart is per-date, showing only the runs taken that day, and is where a recent sweep or collapse actually shows up.

By configuration

By model family


Cross-language comparison

The Carwash Test has been run in four languages, each kept as a separate corpus (the charts above are the English corpus, n=213). This table places the models tested in more than one language side by side. The English column follows the live dataset. A dash means the model was not run in that language. (Namazu was also run in Japanese across registers and interface languages — register-dependent, so it is not reduced to a single cell here; see its transcript page.)

Cross-language Carwash Test results for models tested in English, French, Simplified Chinese, and Ukrainian
ModelToggleEnglishFrenchChineseUkrainian
Claude Fable 5Effort High (default)PassPassPassPass
Claude Opus 4.8Adaptive OnPassPassFailPass-adjacent
Claude Opus 4.8Adaptive OffPassPassFailPass-adjacent
Claude Opus 4.7Adaptive OnPass-adjacentPassFailPass
Claude Opus 4.7Adaptive OffPassPassFailPass
Claude Sonnet 4.6OnPassPass-adjacentPass-adjacentPass-adjacent
Claude Sonnet 4.6OffPassPass-adjacentPass-adjacentFail
Claude Sonnet 4.6Adaptive OnPassPassPassPass-adjacent
Claude Sonnet 4.6Adaptive OffPassPassPassFail
GPT 5.5OnPassPass-adjacentFailPass
GPT 5.5OffFailFailFailPass-adjacent
GPT 5.2OnPass-adjacentFailFailFail
GPT 5.2OffFailPass-adjacentFailPass-adjacent
Mistral Medium 3.5 (Vibe)BalancedFailFailFailPass-adjacent
Mistral Medium 3.5 (Vibe)ThinkFailFailFailFail
Mistral Medium 3.5 (Vibe)ResearchFailFailFailFail
LumoFailFailFailPass-adjacent
PerplexityPassPass-adjacentPass-adjacentPass-adjacent
GLM-5.2Deep Think HighPassPassPassPass-adjacent
GLM-5.2Deep Think MaxPassPass-adjacentPass-adjacentPass
GLM-5.2Deep Think OffPass-adjacentVerbosePass-adjacentPass
Namazu (Sakana)FailFailFailFail
Claude Opus 4.8Extended On · Effort High (July UI)PassPassPassPass
Claude Opus 4.8Extended Off · Effort High (July UI)PassPassPassPass
Claude Sonnet 5Extended On · Effort HighPass-adjacentPassPassPass
Claude Sonnet 5Extended Off · Effort HighPass-adjacentPassFailFail
ChatGPT 5.6 SolEffort MediumPassPassPassPass
ChatGPT 5.6 SolEffort HighPassPassPassPass
ChatGPT 5.5Effort InstantFailFailFailPass-adjacent
ChatGPT 5.5Effort HighPassPassPass-adjacentPass-adjacent
Qwen3.7-MaxThinkingPass-adjacentPassPassPass
Qwen3.7-MaxFastPass-adjacentPass-adjacentPass-adjacentPass-adjacent
Qwen3.7-PlusThinkingPassPassPassPass-adjacent
Qwen3.7-PlusFastFailPassFailFail
DeepSeek V4-FlashThinking OnPass-adjacentPass-adjacentPass-adjacentPass-adjacent
DeepSeek V4-FlashThinking OffPass-adjacentPass-adjacentPass-adjacentVerbose
DeepSeek V4-ProThinking OnPass-adjacentPass-adjacentPass-adjacentVerbose
DeepSeek V4-ProThinking OffFailPassPass-adjacentPass-adjacent
Kimi K2.6ThinkingFailPass-adjacentFailFail
Kimi K2.6InstantFailFailFailFail
Vibe ChatFastFailFailFailPass-adjacent
Vibe ChatThinkingFailFailFailFail
Lumo 2.0 LiteFastFailFailPass-adjacentFail
Lumo 2.0 LiteThinkingFailFailPass-adjacentPass-adjacent
Lumo 2.0 MaxFastFailPassPass-adjacentFail
Lumo 2.0 MaxThinkingPass-adjacentPass-adjacentPass-adjacentPass-adjacent
Kimi K3Max (default)Pass-adjacentPassPass-adjacentPass-adjacent
Kimi K3StandardPassPassPassPass
Qwen3.8-MaxFastVerbosePass-adjacentPass-adjacentPass-adjacent
Qwen3.8-MaxThinkingPassPassPass-adjacentPass
Qwen3.8-MaxAutoPassPassPassVerbose

Simplified Chinese corpus

Simplified Chinese prompt (35 m ≈ 115 ft)

我的车脏了。洗车店在35米外。我是该走路去,还是开车去?

Translation: My car is dirty. The car wash is 35 meters away. Should I walk there or drive?

Translated via Google Translate; back-translated to verify conformance with the English original. Distance converted to a metric equivalent.

Token calculation: DeepSeek documents that 1 Chinese character ≈ 0.6 tokens (vs ~0.25 per English character), so token estimates for Chinese runs use the Chinese rate. A 72-token Chinese response and a 72-token English response represent different amounts of visible text.

The Carwash Test was extended to Simplified Chinese using a metrically equivalent prompt. Across the three Chinese-hosted vendors — DeepSeek, Kimi, and Qwen — nine runs produced a 22% failure rate, substantially lower than the failure rate the same vendors showed in English. The prompt's mechanism survived translation: surface features still pushed toward Walk, the logical constraint still pointed toward Drive. What changed was how the constraint was articulated. Chinese responses that passed consistently named the carwash as a business requiring the car's physical presence — “洗车店老板可没法帮你洗” (the carwash owner can't help you wash it) — a personalized framing that made the logical object harder to lose. One novel failure mode emerged only in Chinese: DeepSeek V4-Pro with reasoning disabled identified the constraint correctly, labeled it as a joke, and offered Walk as the “serious” practical advice — the correct answer visible to the model and dismissed as comedy.

洗车测试已扩展至简体中文,使用等效的公制提示语。针对三家中国厂商——深度求索(DeepSeek)、Kimi和通义千问(Qwen)——的九次测试中,失败率为22%,远低于同一批厂商在英文版测试中的失败率。提示语的核心机制经受住了翻译的考验:表面特征仍然推向"走路",逻辑约束仍然指向"开车"。变化在于约束的表达方式。通过测试的中文回答普遍将洗车店描述为一个需要车辆到场的经营场所——"洗车店老板可没法帮你洗"——这种拟人化的表述使逻辑对象更难被忽视。一种全新的失败模式仅在中文测试中出现:深度求索V4-Pro在关闭推理功能时,正确识别了逻辑约束,却将其归类为笑话,然后将"走路"作为严肃的实用建议——正确答案对模型来说清晰可见,却被当作幽默而忽略。

Control runs. Five US-trained control runs (ChatGPT 5.5 On/Off, ChatGPT 5.2 On/Off, Claude Opus 4.7) were then added; all five fail, bringing the corpus to 14 runs and a 50% failure rate. The Chinese-hosted vendors fail at 22%; the US-trained controls fail at 100% — the US models handle the Chinese prompt worse than the Chinese-hosted models do, inverting the intuition that Chinese vendors would struggle more.

May 29 Anthropic sweep. Eleven more runs (Opus 4.7/4.8, Sonnet 4.6, Lumo, Vibe) brought the corpus to 25 runs and a 52% failure rate, and surfaced a clean model-size inversion: both Opus generations fail Chinese in every toggle state — recommending Walk on cold-start/parking grounds, or naming the constraint and then dismissing it — while the smaller Sonnet 4.6 holds the constraint in all four states. Opus 4.8 passes English and French cleanly, so this is language-specific, not a general regression; in Chinese the larger model's extra reasoning argues itself out of the right answer. Vibe (Le Chat's successor) fails all three modes; Lumo holds. Claude Fable 5's launch-day pass (June 9) brings the corpus to 26 runs and a 50% failure rate — the first Anthropic flagship-tier Chinese pass. A later Perplexity run (June 19) passed with hedging — it cites Chinese-language web coverage of the puzzle rather than reasoning it out — bringing the corpus to 27 runs and a 48% failure rate. GLM-5.2 (Z.ai) then held the constraint in all three Deep Think states (June 22), bringing the corpus to 30 runs and a 43% failure rate. Sakana AI's Namazu failed the Chinese prompt (June 23) — '建议走路去' — bringing the corpus to 31 runs and a 45% failure rate. Carwash III (July 11) added 29 runs in one day, bringing the corpus to 60 and the failure rate down to 37%: ChatGPT-5.6 Sol, Opus 4.8, Sonnet 5 (thinking on), Fable 5, GLM-5.2, and both consumer Qwen 3.7 thinking modes all hold — while Kimi fails its home language in both states with a Chinese inverted-logic flourish, Sonnet 5's thinking-off state loses the object, and Namazu answers the Chinese prompt in Japanese. Kimi K3 (July 17) holds Chinese at both tiers — with English traces — bringing the corpus to 62 and the failure rate to 35%. Qwen3.8-Max (August 3) holds in all three modes, its Auto run a winner’s-circle pass, bringing the corpus to 65 at 34%.

French-language corpus

French-language prompt (35 m ≈ 115 ft)

Ma voiture est sale. Le lave-auto se trouve à 35 mètres. Devrais-je y aller à pied ou en voiture ?

Translation: My car is dirty. The car wash is 35 meters away. Should I walk there or drive?

Translated via Google Translate; back-translated to verify conformance with the English original. Distance converted to a metric equivalent. French uses Latin script, so the standard ~4 chars-per-token estimate applies.

The Carwash Test was extended to French using a metrically equivalent prompt. The first nine runs across four vendors — Mistral, Lumo, OpenAI, and Anthropic — produced a 67% failure rate. A language-specific failure mode emerged: the inverted-logic pattern, in which the model argues that driving would make the car dirtier or that the car is already clean, appears across three vendors in French (Mistral, Lumo, and GPT) but in zero English-language runs. The toggle relationship itself proved language-dependent: GPT 5.2 passes with thinking off and fails with thinking on in French — the exact inverse of its English behavior. Claude Opus 4.7 produced one of the most concise correct answers in the entire dataset (“En voiture — sinon le lave-auto va laver le mauvais sujet”) while the same model, same toggle, same day, failed in Chinese with a two-character response. The kind of wrong answer depends on the language even when the fact of failure does not. A May 29 sweep added seven more Anthropic runs — Opus 4.7, Opus 4.8, and Sonnet 4.6 across both toggle states, plus two console turns — and every one held the constraint, bringing the corpus to 16 runs and a 38% failure rate. The five consumer answers were terse winner's-circle passes (“En voiture. Tu dois la laver, pas toi.”); the language that breaks GPT and Mistral leaves the Claude models untouched. Claude Fable 5 passed on launch day (June 9), bringing the corpus to 17 runs and a 35% failure rate. Perplexity passed again on June 19 — a search-grounded answer citing French press coverage of the puzzle itself — bringing the corpus to 18 runs and a 33% failure rate. GLM-5.2 (Z.ai) held the constraint across all three Deep Think states (June 22) — its Deep Think Off run the corpus's one verbose outlier, with a tangent on car-wash types — bringing the corpus to 21 runs and a 29% failure rate. Namazu (Sakana AI) then failed in French (June 23) with a confused, inverted answer — "vous risquez de salir la route" — bringing the corpus to 22 runs and a 32% failure rate. Carwash III (July 11) added 29 runs, bringing the corpus to 51 and the failure rate to 27% — and French turned unexpectedly kind: it is Kimi K2.6's only hold in eight runs across four languages, the only language Qwen3.7-Plus Fast holds, and where DeepSeek V4-Pro (thinking off) passes cleanly on the same day it fails English with inverted logic. The inverted-logic flourish itself persists here (Vibe in both modes, Lumo 2.0 Lite Fast). Kimi K3 (July 17) holds French at both tiers, completing the K2.6-to-K3 reversal and bringing the corpus to 53 at a 26% failure rate. Qwen3.8-Max (August 3) holds in all three modes — and where its English and Ukrainian Fast runs answer in numbered briefs, the French Fast run answers in prose, so the mode’s format is language-dependent. The corpus stands at 56 and 25%.

Le test du lave-auto a été étendu au français à l'aide d'un prompt métrique équivalent. Les neuf premiers tests répartis sur quatre fournisseurs — Mistral, Lumo, OpenAI et Anthropic — ont produit un taux d'échec de 67 %. Un mode d'échec propre à la langue est apparu : le raisonnement inversé, selon lequel le modèle soutient que conduire salirait davantage la voiture ou que la voiture est déjà propre, se manifeste chez trois fournisseurs en français (Mistral, Lumo et GPT) mais dans aucun test en anglais. La relation du commutateur de raisonnement s'est révélée dépendante de la langue : GPT 5.2 réussit sans raisonnement étendu et échoue avec en français — l'exact inverse de son comportement en anglais. Claude Opus 4.7 a produit l'une des réponses correctes les plus concises de l'ensemble du jeu de données (« En voiture — sinon le lave-auto va laver le mauvais sujet ») tandis que le même modèle, le même réglage, le même jour, a échoué en chinois avec une réponse de deux caractères. Le type de mauvaise réponse dépend de la langue, même lorsque le fait de l'échec n'en dépend pas. Une série du 29 mai a ajouté sept tests Anthropic — Opus 4.7, Opus 4.8 et Sonnet 4.6 dans les deux états du commutateur, plus deux requêtes via la console — qui ont tous tenu la contrainte, portant le corpus à seize tests et un taux d'échec de 38 %.

Ukrainian-language corpus

Ukrainian-language prompt (35 m ≈ 115 ft)

У мене брудна машина. Автомийка знаходиться за 35 метрів від мене. Мені туди краще йти пішки чи поїхати на машині?

Translation: My car is dirty. The car wash is 35 meters away. Should I walk there or drive?

Native-speaker-translated, not machine-translated. Distance converted to a metric equivalent.

Token calculation: this is the first corpus with a measured tokenization rate. API-console runs report real output-token counts (which include hidden reasoning tokens), and Cyrillic text runs at roughly 0.5 tokens per character — about double the English rate of ~0.25. Console runs are flagged distinctly because their token totals include reasoning the consumer interface hides. Consumer-app runs use the measured Cyrillic rate. None of these counts are directly comparable to the English character-based estimates.

The Carwash Test was extended to Ukrainian using a native-speaker-translated prompt — 28 runs across Anthropic, OpenAI, DeepSeek, Qwen, Proton (Lumo), and Mistral (Vibe), split between the API console and the consumer interface, for a 25% failure rate. The corpus was built to separate two things the earlier languages had confounded: the reasoning-effort level as a continuous variable, and the surface (developer console vs. consumer app). Both proved to matter. Claude Sonnet 4.6 holds the constraint at high effort and inverts it at low effort — the same model, same prompt, failing only when given less time to think. On the console, OpenAI's GPT 5.5 returned the same correct answer at 121 output tokens (low effort) and 565 (extra-high effort) — real console counts, not estimates: a 4.7× cost difference for identical quality, almost all of it hidden reasoning. The cleanest pass in the corpus was Qwen3.7-Plus-Preview's two-sentence answer naming the constraint directly (“the washers would have nothing to wash”). The most elaborate failure in the entire dataset also appeared here: DeepSeek V4-Pro with reasoning off fabricated a “known Soviet riddle” about Zhiguli cars, complete with an invented canonical punchline, to justify walking — constraint-as-comedy escalated into a hallucinated cultural reference, which reasoning-on then repaired. A later Perplexity run (June 19) passed by citing web coverage that restates the riddle rather than reasoning it out, bringing the corpus to 29 runs and a 24% failure rate. GLM-5.2 (Z.ai) held the constraint in all three Deep Think states (June 22), bringing the corpus to 32 runs and a 22% failure rate. Namazu (Sakana AI) failed in Ukrainian too (June 23), bringing the corpus to 33 runs and a 24% failure rate. Carwash III (July 11) added 29 runs, bringing the corpus to 62 at a 26% failure rate. Ukrainian kept its reputation as the forgiving corpus — it is the only language where ChatGPT-5.5's Instant tier and Vibe's Fast mode hold — but it also produced the batch's strangest fails: Sonnet 5 (thinking off) recommends walking with the grammatically scrambled "Їдь пішки" ("drive by foot"), and DeepSeek V4-Pro's trace reasons in Russian, the pairing first seen in V4-Flash in May. Kimi K3 (July 17) holds Ukrainian at both tiers — again with English traces — bringing the corpus to 64 at a 25% failure rate. Qwen3.8-Max (August 3) holds in all three modes, though its Auto run is the batch’s one verbose outlier at roughly 400 tokens, bringing the corpus to 67 at 24%.

Резюме українською мовою готується; його перевірить носій мови перед публікацією.

Cross-corpus findings

Three findings emerge only when the corpora are read together — each isolates a variable that a single language could not.

Models reason in a dominant internal language, then translate

Several Ukrainian runs exposed a reasoning trace in a language other than the prompt or the answer. The model handled the logical constraint in its dominant internal language and translated only the final output into Ukrainian.

Reasoning-trace language vs. response language, where the trace was observable
ModelSurface / toggleReasoned inAnswered in
DeepSeek V4-FlashConsumer, DeepThink OnRussianUkrainian
Qwen3.7-MaxConsumerEnglish, then ChineseUkrainian
Qwen3.7-Max-PreviewConsumerUkrainian, then ChineseUkrainian
Claude Sonnet 4.6API console, On / Adaptive OnEnglishUkrainian
Gemma 4 26B A4B ITAI Studio, Thinking High (open-weight)EnglishChinese
Qwen3.6 27B (Q4_K_M)Local, Thinking On (open-weight)EnglishChinese / French / Ukrainian
Claude Fable 5Consumer, Effort HighSame as prompt — all four languagesEnglish / French / Chinese / Ukrainian
DeepSeek V4-ProConsumer, DeepThink On (July 11)RussianUkrainian
Claude Sonnet 5Consumer, Extended On · Effort High (July 11)EnglishChinese
Qwen3.7-Max / 3.7-PlusConsumer, Thinking (July 11)English (mixed with the prompt language)French / Ukrainian
Namazu (Sakana)Sakana Chat, Standard register (July 11)JapaneseJapanese — on the Chinese prompt
Kimi K3Consumer, Max & Standard (July 17)EnglishChinese / French / Ukrainian
Qwen3.8-Max-PreviewConsumer, thinking locked on (July 30)EnglishEnglish
Qwen3.8-MaxConsumer, Thinking & Auto (August 3)Mixed — switches mid-traceFrench / Ukrainian
Qwen3.8-MaxConsumer, Thinking & Auto (August 3)ChineseChinese

DeepSeek reasoning in Russian on a Ukrainian prompt is notable given the political context; the trace language is a property of the training distribution, not the prompt.

Fable 5 is the first model in the dataset observed reasoning in the prompt's language in every language tested — and the first model with a clean four-language pass record. Every previously tested model that exposed a trace reasoned in a dominant internal language (English, Russian, or Chinese) and translated outward, and every previously tested model failed in at least one language. The correlation supports the trace-language-match hypothesis: language-dependent failures may enter at the translation boundary between the model's internal working language and its output language. One model; correlation only; stated at that weight. Carwash III (July 11) collected trace language across every model that exposes one — and complicated the hypothesis: Qwen3.7-Max swept all four languages while reasoning in mixed English, and ChatGPT-5.6 Sol swept with no observable trace at all, so a trace-language match is evidently not necessary for a clean record. DeepSeek's Russian-on-Ukrainian pairing, first seen in V4-Flash, reappeared in V4-Pro. And Namazu extended the mismatch from reasoning to output, answering the Chinese prompt in Japanese. Kimi K3 (July 17) reasons in English on every non-English prompt at both tiers even while sweeping all four languages — which sharpens a standard this record now tracks: from the operator’s side, the trace is part of the product, and a reasoning trace the operator cannot read fails at its one job of making the reasoning inspectable, whatever the verdict. Trace language does not change a score; it is recorded and weighed here. Qwen3.8-Max (August 3) breaks the pattern in a new way: its French and Ukrainian traces do not pick a language and stay there — they switch between the prompt language and English mid-stream, sometimes several times in one trace, while its Chinese traces stay wholly in Chinese. A trace that changes language partway is readable to no one in particular.

Reasoning effort is a continuous variable, not a toggle

Anthropic's new effort selector (and OpenAI's console effort control) make the amount of reasoning a dial rather than an on/off switch. The same model can pass or fail depending on where the dial sits.

Claude Sonnet 4.6 on the Ukrainian prompt, by effort level
Surface / toggleEffortResult
Console, Thinking OnHighPass-adjacent
Console, Adaptive OnHighPass-adjacent
Console, Thinking OffHighFail
Consumer, Adaptive OnLowFail
Consumer, Adaptive OffLowFail

Inkling puts the whole dial in one model. Thinking Machines' open-weights model (July 16) exposes six named Reasoning Levels, and across them the verdict flips exactly once: None, Minimum, and Low recommend walking; Medium, High, and Extra High drive. Below the threshold the distance wins; above it the object does. Per the model card, those six names discretize a continuous effort parameter running from zero to one — the vendor's own benchmarks report effort=0.99 — so the real threshold sits at some value between Low and Medium, and the selector only samples it. This is an open-weight result, excluded from the commercial corpora, and it is the cleanest effort threshold on record.

Inkling (Thinking Machines, open-weight) on the English prompt, by Reasoning Level
Reasoning LevelVerdictResult
NoneWalkFail
MinimumWalkFail
LowWalkFail
MediumDrivePass-adjacent
High (default)DrivePass-adjacent
Extra HighDrivePass-adjacent

The Low run is the sharpest split in the dataset between what a model reasons and what it answers. Its visible trace concludes that the car has to be driven regardless; the answer beneath it recommends walking, in text verbatim identical to the Minimum response. On this evidence the trace is not a window onto the deliberation that produced the answer. It is a second output. Qwen3.8-Max (August 3) supplies the mirror case: two of its traces argue their way to Walk in full, under their own headings, before reversing to Drive and answering correctly. The trace and the answer can diverge in either direction.

Effort is pure cost once the answer is correct

On the OpenAI console, effort and verbosity are orthogonal controls — you can think hard and speak briefly. GPT 5.5 produced the same correct Ukrainian answer at three effort settings; the extra effort bought nothing but hidden reasoning tokens.

GPT 5.5 on the Ukrainian prompt (API console), by effort
EffortVerbosityOutput tokensResult
LowLow121Pass
MediumMedium250Pass-adjacent
Extra-highLow565Pass

Low and extra-high effort return the same correct answer at 121 vs. 565 output tokens (real console counts, not estimates) — a 4.7× cost difference for identical quality.

Two July results cut the other way, at least on the visible answer. Claude Opus 5 gets shorter as the dial goes up: ~21 tokens at Low, ~12 at High, ~9 at Max, all three correct and all three inside the winner's circle. Inkling compresses the same way once it is above its threshold — ~186 tokens at Medium, ~152 at High, ~138 at Extra High. Where a model is already holding the constraint, added effort can buy concision rather than padding.

The caveat matters. These are estimates of the visible answer, and consumer surfaces do not report reasoning tokens, so a shorter answer at higher effort is not a cheaper answer. The GPT 5.5 console runs above are the only place in this dataset where the full cost is measured — and there the extra effort bought nothing.

The budget tier is where the test bites, and where it moves fastest

Across vendors, the cheap or fast configuration is the reliable failure site. ChatGPT 5.5 at Instant effort fails English, French, and Chinese. Kimi K2.6 Instant fails all four languages. Qwen3.7-Plus Fast, Vibe Chat Fast, and Lumo 2.0 Lite Fast each fail three of four. Gemini 3.1 Flash-Lite failed every time it was run — three runs over two months, including a 238-token comparative breakdown that recommended walking.

Eleven days after that breakdown, Gemini 3.6 Flash-Lite passed both of its toggle states, and its Extended-Thinking-off answer is a ~17-token winner's-circle pass. Six days after K2.6's Instant mode failed in all four languages, Kimi K3's Standard tier held the constraint in all four. The tier that fails most often is also the tier where one generation can reverse the result outright.

This is not evidence that the frontier is improving. The flagships in this dataset were mostly passing already, which leaves them little room to move; the measurable recent gains are at the bottom of the lineup, where the failures were. What it does suggest is that holding the logical object is not an expensive capability reserved for the largest models — a model small enough to be the cheap option can name the constraint in seventeen tokens.

The winner's circle: concise correct answers

A “winner's-circle” pass names the constraint and stops. Because tokenization differs by script, the brevity threshold is script-specific: ≤30 tokens in Latin, ≤60 in Hanzi, ≤60 in Cyrillic. Every Pass that clears its threshold is listed below — runs #1/#39 and #2/#40 are the same verbatim answer produced in two separate test batches. Claude Fable 5 enters in three of its four languages (English, French, Chinese); its Ukrainian run passes but, at ~100 tokens by the measured Cyrillic rate, exceeds the 60-token threshold. Claude Sonnet 5 enters in four of its five effort modes (Medium, High, Extra, Max); the Low-effort run passes but, at ~46 tokens, exceeds the Latin threshold. One open-weight run also clears the bar and is included for completeness, flagged as such — Gemma 4 31B (Google AI Studio, Chinese, ~24 tokens) — the table's only non-commercial entrant. Carwash III (July 11) adds 31 qualifiers in one day — 28 of them from the Anthropic lineup re-baseline, plus both ChatGPT-5.6 Sol runs and a Copilot GPT 5.6 Think run — including the shortest pass on record: Claude Opus 4.6's six-token "Drive. It's a carwash." The same day's non-English sweep adds 20 more across all three scripts — six of them ChatGPT-5.6 Sol's, whose entire four-language record sits inside the thresholds, and the tersest of all Opus 4.8's ten-token Chinese "开车去。车不在洗车店里就洗不了。" Gemini 3.6 Flash-Lite (July 22) enters at ~17 tokens on its Extended-Thinking-off run — the same tier whose 3.1 predecessor produced the dataset’s longest Gemini failures eleven days earlier. Claude Opus 5 (July 24) enters in all three tiers tested — and its entries run backwards to the usual expectation: ~21 tokens at Low, ~12 at High, ~9 at Max, the tersest English pass since Opus 4.6’s six-token record. Qwen3.8-Max (August 3) enters once, in Chinese at ~59 tokens; its English Thinking answer passes at ~31 and misses the Latin threshold by a single token.

Winner’s-circle passes: every Pass within the per-script token threshold
RunScriptModelLanguage / configTokensResponse
#1LatinClaude Opus 4.6English / On~14Drive. The car's the thing that needs washing, not you.
#2LatinClaude Opus 4.6English / Off~11Drive. You're washing the car, not yourself.
#3LatinClaude Sonnet 4.6English / On~8Drive. You're going to a carwash.
#4LatinClaude Sonnet 4.6English / Off~17Drive. You're getting a carwash — the car needs to be there anyway.
#32LatinChatGPT 5.5English / On~22Drive. Walking solves the "100 feet away" problem, but not the "my car is dirty" problem.
#39LatinClaude Opus 4.6English / On~14Drive. The car's the thing that needs washing, not you.
#40LatinClaude Opus 4.6English / Off~11Drive. You're washing the car, not yourself.
#41LatinClaude Sonnet 4.6English / Adaptive On~19Drive. You're getting a carwash, not a car delivery — you need the car there.
#42LatinClaude Sonnet 4.6English / Adaptive Off~18Drive. You're going to a carwash — the car needs to be there, not you.
#97LatinQwen3.7-Plus-PreviewEnglish / On~20You should drive. If you walk, your car stays behind and doesn't get washed!
#106HanziQwen3.7-Max-PreviewChinese / —~54当然是开车去! 如果你走路去,人到了洗车店,车还在原地,洗车店老板可没法隔空帮你洗。你要洗的是车,不是你自己呀😂。 (除非你打算走路过去,把洗车店老板请过来,或者借个水管自己拉回来洗……)
#121LatinClaude Opus 4.7French / Adaptive On~15En voiture — sinon le lave-auto va laver le mauvais sujet.
#122LatinClaude Opus 4.8English / Adaptive Off~10Drive. You're taking the car to the carwash.
#123LatinClaude Opus 4.8English / Adaptive On~12Drive. You'd have to drive it through the carwash anyway.
#140CyrillicQwen3.7-Plus-PreviewUkrainian / —~60Вам однозначно треба поїхати на машині. Інакше мийникам буде нічого мити (хіба що ви вирішите донести її на руках 😉).
#141CyrillicClaude Opus 4.7Ukrainian / Adaptive On~44Їдь. Це ж і є сенс автомийки — пригнати туди брудну машину. Пішки прийдеш — мити нічого.
#142CyrillicClaude Opus 4.7Ukrainian / Adaptive Off~21Поїхати, звісно. Машину ж миєш, а не себе.
#151LatinClaude Opus 4.7French / Adaptive Off~8En voiture. Sinon elle reste sale.
#152LatinClaude Opus 4.8French / Adaptive Off~11En voiture, évidemment. C'est un lave-auto.
#153LatinClaude Opus 4.8French / Adaptive On~16En voiture, évidemment. C'est ta voiture qui est sale, pas toi.
#154LatinClaude Sonnet 4.6French / Adaptive Off~10En voiture. Tu dois la laver, pas toi.
#155LatinClaude Sonnet 4.6French / Adaptive On~12En voiture — sinon tu arrives à pied au lave-auto.
#161HanziClaude Sonnet 4.6Chinese / Adaptive Off~15开车去。35米步行也到不了,但你得把车开进洗车道。
#162HanziClaude Sonnet 4.6Chinese / Adaptive On~9开车去。走路把车推回来比较难。
#172HanziGemma 4 31B ITChinese / High · open-weight~24这是一个典型的"脑筋急转弯"问题。你应该开车去。因为你的目的是洗车,如果你走路去,车还在原地,洗车店没法帮你洗车。
#185LatinClaude Fable 5English / Effort High~26Drive. The car is the one that needs washing; walking gets you a clean you standing next to a dirty car.
#186LatinClaude Fable 5French / Effort High~23En voiture — c'est elle qui est sale, pas vous. Le lave-auto fait peu pour un piéton propre.
#187HanziClaude Fable 5Chinese / Effort High~21开车去。要洗的是车,不是你——车留在家里的话,洗车店那35米就白走了。
#230LatinClaude Sonnet 5English / Effort Medium~8Drive — you need the car there.
#231LatinClaude Sonnet 5English / Effort High~8Drive — the car needs to be there.
#232LatinClaude Sonnet 5English / Effort Extra~14Drive — the car's the one that needs washing, not you.
#233LatinClaude Sonnet 5English / Effort Max~25Drive — the car's the one that needs to be there, not you. Walking would rather defeat the purpose.
#242LatinClaude Fable 5English / Effort Low~18Drive. Bringing the carwash the car is the harder version of the problem.
#243LatinClaude Fable 5English / Effort High~20Drive. The car is the patient here — walking gets you a clean view of a dirty car.
#244LatinClaude Fable 5English / Effort Max~28Drive. The car's the one getting washed — walking there just delivers a clean pedestrian to a confused attendant.
#245LatinClaude Opus 4.8English / On / Effort Low~18Drive. You need the car at the carwash regardless of how you get there.
#246LatinClaude Opus 4.8English / On / Effort High~10Drive. You need the car at the carwash.
#247LatinClaude Opus 4.8English / On / Effort Max~12Drive. The car's the one that needs to be there.
#248LatinClaude Opus 4.8English / Off / Effort Low~10Drive. You'd have to bring the car anyway.
#249LatinClaude Opus 4.8English / Off / Effort High~18Drive. Driving through a carwash requires the car to be at the carwash.
#250LatinClaude Opus 4.8English / Off / Effort Max~13Drive. Moving the car through the wash is the point.
#251LatinClaude Opus 4.7English / On / Effort Low~17Drive. You're going to end up at the carwash in the car either way.
#253LatinClaude Opus 4.7English / On / Effort Max~10Drive. The carwash needs the car, not you.
#255LatinClaude Opus 4.7English / Off / Effort High~18Drive. Getting the car clean is the point; walking there leaves it dirty.
#256LatinClaude Opus 4.7English / Off / Effort Max~8Drive. You're going to a carwash.
#257LatinClaude Opus 4.6English / On / Effort Low~12Drive. The car's the thing that needs to be there.
#258LatinClaude Opus 4.6English / On / Effort High~12Drive. The car's the thing that needs to be there.
#259LatinClaude Opus 4.6English / On / Effort Max~12Drive. The car's the thing that needs washing.
#260LatinClaude Opus 4.6English / Off / Effort Low~6Drive. It's a carwash.
#261LatinClaude Opus 4.6English / Off / Effort High~11Drive. You're washing the car, not yourself.
#262LatinClaude Opus 4.6English / Off / Effort Max~6Drive. It's a carwash.
#263LatinClaude Sonnet 5English / On / Effort Low~20Drive it there, obviously — you need the car at the carwash, not just yourself.
#265LatinClaude Sonnet 5English / On / Effort Max~12Drive — it's the car that needs the wash, not you.
#266LatinClaude Sonnet 5English / Off / Effort Low~22Drive it to the carwash 100 feet away — walking gets a clean sidewalk, not a clean car.
#269LatinClaude Sonnet 4.6English / On / Effort Low~8Drive. You're washing the car.
#270LatinClaude Sonnet 4.6English / On / Effort High~11Drive. You're washing the car, not yourself.
#271LatinClaude Sonnet 4.6English / On / Effort Max~23Drive. Moving a car 100 feet costs essentially nothing and you'll need it in position anyway.
#272LatinClaude Sonnet 4.6English / Off / Effort Low~8Drive. You're going to a carwash.
#273LatinClaude Sonnet 4.6English / Off / Effort High~12Drive. Walking gets you there but not the car.
#274LatinClaude Sonnet 4.6English / Off / Effort Max~8Drive. You're washing the car.
#289LatinGPT 5.6 Think (via Copilot)English / On~28Drive. The goal is to wash the car, so the car needs to go to the car wash—even though it’s only 100 feet away.
#296LatinChatGPT 5.6 SolEnglish / Effort Medium~12Drive—the car needs to go through the car wash.
#297LatinChatGPT 5.6 SolEnglish / Effort High~10Drive. The car needs to be at the carwash.
#321HanziQwen3.7-MaxChinese / Thinking~57当然是**开车去**(或者把车推过去)! 虽然35米非常近,走路只需要几十秒,但你要洗的是**车**呀!如果你自己走路过去,车还是停在原地,依然是脏的。必须把车带过去,洗车店才能帮你洗。😄
#327HanziQwen3.7-PlusChinese / Thinking~53当然是开车去啦! 虽然35米非常近,走路只要几十秒,但你的目的是洗车呀。如果你走路过去,车还留在原地,洗车店老板总不能给你洗个澡吧?😂 所以,为了让车变干净,你还是得把车开过去。
#332CyrillicClaude Fable 5Ukrainian / Effort High~36Машиною. Мийка миє те, що приїхало, — якщо прийдете пішки, помиють вас.
#333HanziClaude Fable 5Chinese / Effort High~21开车去。要洗的是车,不是你——走过去只能站在那儿看着35米外那辆脏车。
#334LatinClaude Opus 4.8French / On / Effort High~27En voiture, évidemment — sinon vous arrivez propre à un lave-auto et la voiture reste sale à 35 mètres de là.
#335CyrillicClaude Opus 4.8Ukrainian / On / Effort High~21Машиною, звісно — інакше мити буде нічого.
#336HanziClaude Opus 4.8Chinese / On / Effort High~14开车——脏的是车,你只是司机。35米,别熄火。
#337LatinClaude Opus 4.8French / Off / Effort High~26En voiture — c'est la voiture qu'il faut laver. À pied, vous arriveriez propre mais elle resterait sale.
#338CyrillicClaude Opus 4.8Ukrainian / Off / Effort High~40Поїхати. Машину треба доставити до мийки — пішки ви туди принесете тільки себе.
#339HanziClaude Opus 4.8Chinese / Off / Effort High~10开车去。车不在洗车店里就洗不了。
#341CyrillicClaude Sonnet 5Ukrainian / On / Effort High~24Поїхати — автомийка миє машину, а не пішоходів.
#342HanziClaude Sonnet 5Chinese / On / Effort High~20开车去——洗的是车,不是你。35米走路不算什么,但车得在那儿才能洗。
#343LatinClaude Sonnet 5French / Off / Effort High~28En voiture. 35 mètres à pied ne lave rien — il faut y conduire la voiture pour qu'elle passe dans le lave-auto.
#370LatinChatGPT 5.6 SolFrench / Effort Medium~16En voiture — sinon, vous arriverez au lave-auto sans la voiture.
#371CyrillicChatGPT 5.6 SolUkrainian / Effort Medium~60Краще **поїхати машиною** — адже саме її треба помити. Пішки ви дістанетеся швидше, але брудна машина залишиться вдома.
#372HanziChatGPT 5.6 SolChinese / Effort Medium~12开车去。你要洗的是车,不是去洗车店散步。
#373LatinChatGPT 5.6 SolFrench / Effort High~14En voiture — sinon, la voiture sale restera chez vous.
#374CyrillicChatGPT 5.6 SolUkrainian / Effort High~27Їдьте машиною — інакше на автомийці не буде чого мити.
#375HanziChatGPT 5.6 SolChinese / Effort High~10开车去。你是要洗车,不是去散步。
#379LatinChatGPT 5.5French / Effort High~27En voiture. Si tu y vas à pied, toi seul arriveras au lave-auto ; ta voiture restera sale à 35 mètres de là.
#425LatinGemini 3.6 Flash-LiteEnglish / Ext Off~17Drive. Walking to a carwash without a car won't get your car clean.
#438HanziQwen3.8-MaxChinese / Auto~59应该开车去。 因为要洗的是“车”,不是“你”——洗车店得看到并操作你的车才能洗。35米虽然很近,但如果你走路过去,车还停在原地,就没法洗了。 当然,如果你只是先去问问价格、排队或预约,那可以走路去。
#429LatinClaude Opus 5English / Effort Low~21Drive. Walking to a carwash and standing there without the car defeats the purpose.
#430LatinClaude Opus 5English / Effort High~12Drive. The car has to be there for this to work.
#431LatinClaude Opus 5English / Effort Max~9Drive. Walking gets you clean shoes.
#422HanziKimi K3Chinese / Effort Standard~56开车去。你是去洗车的,车不去,洗车店洗什么呢?😄 35米虽然近,走过去再回来很合理,但前提是车得在店里——所以正确姿势是:把车开过去(或者如果店里有上门取车服务,打电话让他们来取也行)。

The structural finding across all four corpora: holding the logical object is not a function of model size or reasoning budget. The shortest correct answers come from models that name the constraint immediately; the longest wrong answers come from models that elaborate their way past it.