Metrics
Charts and trends across the dataset

My car is dirty. The carwash is 100 feet away. Should I walk or drive?
These charts read directly from the live dataset and update as runs are added. A standing caveat applies throughout: the test is single-shot, the sample per cell is small, and each snapshot reflects whatever models were available that date — so the time-based charts describe the evolving field, not a controlled trend in any one system. See the methodology for scoring and limitations.
Dataset overview
Over time
A note on how to read these. The first three charts are cumulative — each point covers every run recorded up to that date. That makes them stable measures of what the whole corpus says, but it also means a small recent batch barely moves them: by July the dataset held roughly two hundred runs, so adding a dozen cannot shift a median, and a cumulative count can only rise or flatten, never fall. The last chart is per-date, showing only the runs taken that day, and is where a recent sweep or collapse actually shows up.
By configuration
By model family
Cross-language comparison
The Carwash Test has been run in four languages, each kept as a separate corpus (the charts above are the English corpus, n=213). This table places the models tested in more than one language side by side. The English column follows the live dataset. A dash means the model was not run in that language. (Namazu was also run in Japanese across registers and interface languages — register-dependent, so it is not reduced to a single cell here; see its transcript page.)
| Model | Toggle | English | French | Chinese | Ukrainian |
|---|---|---|---|---|---|
| Claude Fable 5 | Effort High (default) | Pass | Pass | Pass | Pass |
| Claude Opus 4.8 | Adaptive On | Pass | Pass | Fail | Pass-adjacent |
| Claude Opus 4.8 | Adaptive Off | Pass | Pass | Fail | Pass-adjacent |
| Claude Opus 4.7 | Adaptive On | Pass-adjacent | Pass | Fail | Pass |
| Claude Opus 4.7 | Adaptive Off | Pass | Pass | Fail | Pass |
| Claude Sonnet 4.6 | On | Pass | Pass-adjacent | Pass-adjacent | Pass-adjacent |
| Claude Sonnet 4.6 | Off | Pass | Pass-adjacent | Pass-adjacent | Fail |
| Claude Sonnet 4.6 | Adaptive On | Pass | Pass | Pass | Pass-adjacent |
| Claude Sonnet 4.6 | Adaptive Off | Pass | Pass | Pass | Fail |
| GPT 5.5 | On | Pass | Pass-adjacent | Fail | Pass |
| GPT 5.5 | Off | Fail | Fail | Fail | Pass-adjacent |
| GPT 5.2 | On | Pass-adjacent | Fail | Fail | Fail |
| GPT 5.2 | Off | Fail | Pass-adjacent | Fail | Pass-adjacent |
| Mistral Medium 3.5 (Vibe) | Balanced | Fail | Fail | Fail | Pass-adjacent |
| Mistral Medium 3.5 (Vibe) | Think | Fail | Fail | Fail | Fail |
| Mistral Medium 3.5 (Vibe) | Research | Fail | Fail | Fail | Fail |
| Lumo | — | Fail | Fail | Fail | Pass-adjacent |
| Perplexity | — | Pass | Pass-adjacent | Pass-adjacent | Pass-adjacent |
| GLM-5.2 | Deep Think High | Pass | Pass | Pass | Pass-adjacent |
| GLM-5.2 | Deep Think Max | Pass | Pass-adjacent | Pass-adjacent | Pass |
| GLM-5.2 | Deep Think Off | Pass-adjacent | Verbose | Pass-adjacent | Pass |
| Namazu (Sakana) | — | Fail | Fail | Fail | Fail |
| Claude Opus 4.8 | Extended On · Effort High (July UI) | Pass | Pass | Pass | Pass |
| Claude Opus 4.8 | Extended Off · Effort High (July UI) | Pass | Pass | Pass | Pass |
| Claude Sonnet 5 | Extended On · Effort High | Pass-adjacent | Pass | Pass | Pass |
| Claude Sonnet 5 | Extended Off · Effort High | Pass-adjacent | Pass | Fail | Fail |
| ChatGPT 5.6 Sol | Effort Medium | Pass | Pass | Pass | Pass |
| ChatGPT 5.6 Sol | Effort High | Pass | Pass | Pass | Pass |
| ChatGPT 5.5 | Effort Instant | Fail | Fail | Fail | Pass-adjacent |
| ChatGPT 5.5 | Effort High | Pass | Pass | Pass-adjacent | Pass-adjacent |
| Qwen3.7-Max | Thinking | Pass-adjacent | Pass | Pass | Pass |
| Qwen3.7-Max | Fast | Pass-adjacent | Pass-adjacent | Pass-adjacent | Pass-adjacent |
| Qwen3.7-Plus | Thinking | Pass | Pass | Pass | Pass-adjacent |
| Qwen3.7-Plus | Fast | Fail | Pass | Fail | Fail |
| DeepSeek V4-Flash | Thinking On | Pass-adjacent | Pass-adjacent | Pass-adjacent | Pass-adjacent |
| DeepSeek V4-Flash | Thinking Off | Pass-adjacent | Pass-adjacent | Pass-adjacent | Verbose |
| DeepSeek V4-Pro | Thinking On | Pass-adjacent | Pass-adjacent | Pass-adjacent | Verbose |
| DeepSeek V4-Pro | Thinking Off | Fail | Pass | Pass-adjacent | Pass-adjacent |
| Kimi K2.6 | Thinking | Fail | Pass-adjacent | Fail | Fail |
| Kimi K2.6 | Instant | Fail | Fail | Fail | Fail |
| Vibe Chat | Fast | Fail | Fail | Fail | Pass-adjacent |
| Vibe Chat | Thinking | Fail | Fail | Fail | Fail |
| Lumo 2.0 Lite | Fast | Fail | Fail | Pass-adjacent | Fail |
| Lumo 2.0 Lite | Thinking | Fail | Fail | Pass-adjacent | Pass-adjacent |
| Lumo 2.0 Max | Fast | Fail | Pass | Pass-adjacent | Fail |
| Lumo 2.0 Max | Thinking | Pass-adjacent | Pass-adjacent | Pass-adjacent | Pass-adjacent |
| Kimi K3 | Max (default) | Pass-adjacent | Pass | Pass-adjacent | Pass-adjacent |
| Kimi K3 | Standard | Pass | Pass | Pass | Pass |
| Qwen3.8-Max | Fast | Verbose | Pass-adjacent | Pass-adjacent | Pass-adjacent |
| Qwen3.8-Max | Thinking | Pass | Pass | Pass-adjacent | Pass |
| Qwen3.8-Max | Auto | Pass | Pass | Pass | Verbose |
- Fable 5 is the first model with a clean four-language pass record. Before June 9, every model tested in more than one language failed in at least one: Chinese defeated both Opus generations and every GPT; Ukrainian broke Sonnet 4.6's thinking-off states. Whether the clean sweep reflects Mythos-class capability, training-data composition, or the trace-language match documented below is undetermined: one model, one snapshot. On June 22, GLM-5.2 (Z.ai) became the second model to clear all four languages — a clean hold in every Deep Think state — so the four-language sweep is no longer a sample of one, though it remains rare. Carwash III (July 11) doubled the club: ChatGPT-5.6 Sol went eight-for-eight across the four languages in both effort tiers — every answer a winner's-circle one-liner — and Qwen3.7-Max held every state in both modes; Fable 5 and GLM-5.2 re-verified their sweeps, and Opus 4.8 swept all six non-English states under the new toggle+effort UI. Kimi K3 (July 17) makes five: twelve Drives in twelve runs across four languages and both variants — six days after Kimi K2.6 went one-for-eight, the largest single-generation reversal in the dataset. Qwen3.8-Max (August 3) makes six, and it is the first to sweep on a mode selector rather than a toggle: twelve Drives across four languages in Fast, Thinking, and Auto alike. Its Fast mode holds everywhere its predecessors broke — Qwen3.7-Plus Fast failed English, Chinese, and Ukrainian; Qwen3.6-Plus failed in Fast and Auto both.
- Chinese defeats models that pass in English and French. Opus 4.7, GPT 5.5 On, and GPT 5.2 On all hold the constraint in English, hold (mostly) in French, and fail in Chinese. The language is an independent variable.
- The flagship fails Chinese where the mid-tier holds it. Both Opus generations (4.7 and 4.8) fail the Chinese prompt in every toggle state — recommending walking, or naming the constraint and then dismissing it — even though Opus 4.8 passes English and French cleanly. The smaller Claude Sonnet 4.6 holds the constraint in Chinese across all four states. More reasoning budget on the larger model does not help here; in this language it hurts.
- The inverted-logic failure was language-triggered — for four months. Models argue that driving would make the car dirtier in French (Mistral, Lumo, both GPTs), Chinese (GPT 5.5 Off), and Ukrainian (Sonnet 4.6 Off, both consumer Sonnet runs) — and in zero English runs from March 22 through July 8. Carwash III (July 11) broke the pattern: DeepSeek V4-Pro (Expert, thinking off) argued the post-wash drive home would re-dirty the car, and Lumo 2.0 Max (Fast) warned that driving would mean "tracking dirt right up to the entrance." The failure mode is still language-skewed, but no longer language-exclusive.
- The toggle's effect is language-dependent. GPT 5.2 On passes in English but fails in French; GPT 5.2 Off fails in English but passes in French. The same inversion repeats in Ukrainian — GPT 5.2 Off (Instant) holds the constraint while Thinking-On fails — as it does for Vibe (Balanced passes, Think fails). The same reasoning mechanism helps in one language and hurts in another.
- Ukrainian is the most forgiving corpus, and the easy answers cluster there. With the column now filled, several models that fail elsewhere hold the constraint in Ukrainian: Lumo and GPT 5.5 Off both recover, and Opus 4.7 passes cleanly in both toggle states. Ukrainian's 24% failure rate is the lowest of the four corpora.
- Reasoning modes can refuse to answer. Vibe's Research mode produced no recommendation at all on the Ukrainian prompt — it returned four clarifying questions — the only configuration in the dataset that declines to hold the object by declining to answer.
- The test is now in the models' search results. On July 11, Grok 4.5's Expert mode ran a mid-answer web search, retrieved fifteen sources — one of them this site — and answered by naming the benchmark: "This is the classic 'car wash test' that trips up a lot of AIs." Correct, object-holding, and scored pass-adjacent: the pass is retrieval-informed, not reasoned cold, and the same model fails in both of its non-searching modes (Fast and Auto). With Perplexity's earlier search-grounded passes, this marks a methodological turn — the diagnostic is now discoverable by the systems it measures, and search-equipped modes must be read differently from closed-book ones.
- The first response-language mismatch — and a home-language flip. On July 11 Namazu answered the Chinese prompt entirely in Japanese — until now only reasoning traces mismatched the prompt language; here the answer itself does. The same session flipped its home-language result: the Standard-register Japanese run, which reached Drive on June 23, now recommends walking. Namazu finishes Carwash III zero-for-five (English, French, Ukrainian, Chinese, Japanese).
- DeepSeek V4-Pro fails only English — the mirror image of the historical pattern. With thinking off, V4-Pro produced the dataset's first English inverted-logic failure (July 11) while holding French, Chinese, and Ukrainian on the same day. Four months of data said models hold English and break elsewhere; this is the first model to break only in English. Kimi shows a related inversion: its sole hold anywhere is French-with-thinking, while it fails home-language Chinese in both states — with a Chinese inverted-logic flourish ("the wash would be wasted; the drive back splashes mud on the body").
- A Japanese-adapted model holds the object only in Japanese — and only in some registers. Sakana AI's Namazu reaches "drive" exclusively on Japanese prompts (Standard register in both interfaces; Polite only in the English interface) and recommends walking in every non-Japanese language — English, Chinese, French, and Ukrainian all fail. Two new variables move the verdict within Japanese: the register selector (Kansai-ben walks in both interfaces; Polite walks in Japanese but drives in English) and the interface language itself. It is the first model whose grip on the logical object appears tied to a cultural-adaptation target rather than to reasoning budget. Full breakdown on its transcript page.
Simplified Chinese corpus
Simplified Chinese prompt (35 m ≈ 115 ft)
我的车脏了。洗车店在35米外。我是该走路去,还是开车去?
Translation: My car is dirty. The car wash is 35 meters away. Should I walk there or drive?
Translated via Google Translate; back-translated to verify conformance with the English original. Distance converted to a metric equivalent.
Token calculation: DeepSeek documents that 1 Chinese character ≈ 0.6 tokens (vs ~0.25 per English character), so token estimates for Chinese runs use the Chinese rate. A 72-token Chinese response and a 72-token English response represent different amounts of visible text.
The Carwash Test was extended to Simplified Chinese using a metrically equivalent prompt. Across the three Chinese-hosted vendors — DeepSeek, Kimi, and Qwen — nine runs produced a 22% failure rate, substantially lower than the failure rate the same vendors showed in English. The prompt's mechanism survived translation: surface features still pushed toward Walk, the logical constraint still pointed toward Drive. What changed was how the constraint was articulated. Chinese responses that passed consistently named the carwash as a business requiring the car's physical presence — “洗车店老板可没法帮你洗” (the carwash owner can't help you wash it) — a personalized framing that made the logical object harder to lose. One novel failure mode emerged only in Chinese: DeepSeek V4-Pro with reasoning disabled identified the constraint correctly, labeled it as a joke, and offered Walk as the “serious” practical advice — the correct answer visible to the model and dismissed as comedy.
洗车测试已扩展至简体中文,使用等效的公制提示语。针对三家中国厂商——深度求索(DeepSeek)、Kimi和通义千问(Qwen)——的九次测试中,失败率为22%,远低于同一批厂商在英文版测试中的失败率。提示语的核心机制经受住了翻译的考验:表面特征仍然推向"走路",逻辑约束仍然指向"开车"。变化在于约束的表达方式。通过测试的中文回答普遍将洗车店描述为一个需要车辆到场的经营场所——"洗车店老板可没法帮你洗"——这种拟人化的表述使逻辑对象更难被忽视。一种全新的失败模式仅在中文测试中出现:深度求索V4-Pro在关闭推理功能时,正确识别了逻辑约束,却将其归类为笑话,然后将"走路"作为严肃的实用建议——正确答案对模型来说清晰可见,却被当作幽默而忽略。
Control runs. Five US-trained control runs (ChatGPT 5.5 On/Off, ChatGPT 5.2 On/Off, Claude Opus 4.7) were then added; all five fail, bringing the corpus to 14 runs and a 50% failure rate. The Chinese-hosted vendors fail at 22%; the US-trained controls fail at 100% — the US models handle the Chinese prompt worse than the Chinese-hosted models do, inverting the intuition that Chinese vendors would struggle more.
May 29 Anthropic sweep. Eleven more runs (Opus 4.7/4.8, Sonnet 4.6, Lumo, Vibe) brought the corpus to 25 runs and a 52% failure rate, and surfaced a clean model-size inversion: both Opus generations fail Chinese in every toggle state — recommending Walk on cold-start/parking grounds, or naming the constraint and then dismissing it — while the smaller Sonnet 4.6 holds the constraint in all four states. Opus 4.8 passes English and French cleanly, so this is language-specific, not a general regression; in Chinese the larger model's extra reasoning argues itself out of the right answer. Vibe (Le Chat's successor) fails all three modes; Lumo holds. Claude Fable 5's launch-day pass (June 9) brings the corpus to 26 runs and a 50% failure rate — the first Anthropic flagship-tier Chinese pass. A later Perplexity run (June 19) passed with hedging — it cites Chinese-language web coverage of the puzzle rather than reasoning it out — bringing the corpus to 27 runs and a 48% failure rate. GLM-5.2 (Z.ai) then held the constraint in all three Deep Think states (June 22), bringing the corpus to 30 runs and a 43% failure rate. Sakana AI's Namazu failed the Chinese prompt (June 23) — '建议走路去' — bringing the corpus to 31 runs and a 45% failure rate. Carwash III (July 11) added 29 runs in one day, bringing the corpus to 60 and the failure rate down to 37%: ChatGPT-5.6 Sol, Opus 4.8, Sonnet 5 (thinking on), Fable 5, GLM-5.2, and both consumer Qwen 3.7 thinking modes all hold — while Kimi fails its home language in both states with a Chinese inverted-logic flourish, Sonnet 5's thinking-off state loses the object, and Namazu answers the Chinese prompt in Japanese. Kimi K3 (July 17) holds Chinese at both tiers — with English traces — bringing the corpus to 62 and the failure rate to 35%. Qwen3.8-Max (August 3) holds in all three modes, its Auto run a winner’s-circle pass, bringing the corpus to 65 at 34%.
French-language corpus
French-language prompt (35 m ≈ 115 ft)
Ma voiture est sale. Le lave-auto se trouve à 35 mètres. Devrais-je y aller à pied ou en voiture ?
Translation: My car is dirty. The car wash is 35 meters away. Should I walk there or drive?
Translated via Google Translate; back-translated to verify conformance with the English original. Distance converted to a metric equivalent. French uses Latin script, so the standard ~4 chars-per-token estimate applies.
The Carwash Test was extended to French using a metrically equivalent prompt. The first nine runs across four vendors — Mistral, Lumo, OpenAI, and Anthropic — produced a 67% failure rate. A language-specific failure mode emerged: the inverted-logic pattern, in which the model argues that driving would make the car dirtier or that the car is already clean, appears across three vendors in French (Mistral, Lumo, and GPT) but in zero English-language runs. The toggle relationship itself proved language-dependent: GPT 5.2 passes with thinking off and fails with thinking on in French — the exact inverse of its English behavior. Claude Opus 4.7 produced one of the most concise correct answers in the entire dataset (“En voiture — sinon le lave-auto va laver le mauvais sujet”) while the same model, same toggle, same day, failed in Chinese with a two-character response. The kind of wrong answer depends on the language even when the fact of failure does not. A May 29 sweep added seven more Anthropic runs — Opus 4.7, Opus 4.8, and Sonnet 4.6 across both toggle states, plus two console turns — and every one held the constraint, bringing the corpus to 16 runs and a 38% failure rate. The five consumer answers were terse winner's-circle passes (“En voiture. Tu dois la laver, pas toi.”); the language that breaks GPT and Mistral leaves the Claude models untouched. Claude Fable 5 passed on launch day (June 9), bringing the corpus to 17 runs and a 35% failure rate. Perplexity passed again on June 19 — a search-grounded answer citing French press coverage of the puzzle itself — bringing the corpus to 18 runs and a 33% failure rate. GLM-5.2 (Z.ai) held the constraint across all three Deep Think states (June 22) — its Deep Think Off run the corpus's one verbose outlier, with a tangent on car-wash types — bringing the corpus to 21 runs and a 29% failure rate. Namazu (Sakana AI) then failed in French (June 23) with a confused, inverted answer — "vous risquez de salir la route" — bringing the corpus to 22 runs and a 32% failure rate. Carwash III (July 11) added 29 runs, bringing the corpus to 51 and the failure rate to 27% — and French turned unexpectedly kind: it is Kimi K2.6's only hold in eight runs across four languages, the only language Qwen3.7-Plus Fast holds, and where DeepSeek V4-Pro (thinking off) passes cleanly on the same day it fails English with inverted logic. The inverted-logic flourish itself persists here (Vibe in both modes, Lumo 2.0 Lite Fast). Kimi K3 (July 17) holds French at both tiers, completing the K2.6-to-K3 reversal and bringing the corpus to 53 at a 26% failure rate. Qwen3.8-Max (August 3) holds in all three modes — and where its English and Ukrainian Fast runs answer in numbered briefs, the French Fast run answers in prose, so the mode’s format is language-dependent. The corpus stands at 56 and 25%.
Le test du lave-auto a été étendu au français à l'aide d'un prompt métrique équivalent. Les neuf premiers tests répartis sur quatre fournisseurs — Mistral, Lumo, OpenAI et Anthropic — ont produit un taux d'échec de 67 %. Un mode d'échec propre à la langue est apparu : le raisonnement inversé, selon lequel le modèle soutient que conduire salirait davantage la voiture ou que la voiture est déjà propre, se manifeste chez trois fournisseurs en français (Mistral, Lumo et GPT) mais dans aucun test en anglais. La relation du commutateur de raisonnement s'est révélée dépendante de la langue : GPT 5.2 réussit sans raisonnement étendu et échoue avec en français — l'exact inverse de son comportement en anglais. Claude Opus 4.7 a produit l'une des réponses correctes les plus concises de l'ensemble du jeu de données (« En voiture — sinon le lave-auto va laver le mauvais sujet ») tandis que le même modèle, le même réglage, le même jour, a échoué en chinois avec une réponse de deux caractères. Le type de mauvaise réponse dépend de la langue, même lorsque le fait de l'échec n'en dépend pas. Une série du 29 mai a ajouté sept tests Anthropic — Opus 4.7, Opus 4.8 et Sonnet 4.6 dans les deux états du commutateur, plus deux requêtes via la console — qui ont tous tenu la contrainte, portant le corpus à seize tests et un taux d'échec de 38 %.
Ukrainian-language corpus
Ukrainian-language prompt (35 m ≈ 115 ft)
У мене брудна машина. Автомийка знаходиться за 35 метрів від мене. Мені туди краще йти пішки чи поїхати на машині?
Translation: My car is dirty. The car wash is 35 meters away. Should I walk there or drive?
Native-speaker-translated, not machine-translated. Distance converted to a metric equivalent.
Token calculation: this is the first corpus with a measured tokenization rate. API-console runs report real output-token counts (which include hidden reasoning tokens), and Cyrillic text runs at roughly 0.5 tokens per character — about double the English rate of ~0.25. Console runs are flagged distinctly because their token totals include reasoning the consumer interface hides. Consumer-app runs use the measured Cyrillic rate. None of these counts are directly comparable to the English character-based estimates.
The Carwash Test was extended to Ukrainian using a native-speaker-translated prompt — 28 runs across Anthropic, OpenAI, DeepSeek, Qwen, Proton (Lumo), and Mistral (Vibe), split between the API console and the consumer interface, for a 25% failure rate. The corpus was built to separate two things the earlier languages had confounded: the reasoning-effort level as a continuous variable, and the surface (developer console vs. consumer app). Both proved to matter. Claude Sonnet 4.6 holds the constraint at high effort and inverts it at low effort — the same model, same prompt, failing only when given less time to think. On the console, OpenAI's GPT 5.5 returned the same correct answer at 121 output tokens (low effort) and 565 (extra-high effort) — real console counts, not estimates: a 4.7× cost difference for identical quality, almost all of it hidden reasoning. The cleanest pass in the corpus was Qwen3.7-Plus-Preview's two-sentence answer naming the constraint directly (“the washers would have nothing to wash”). The most elaborate failure in the entire dataset also appeared here: DeepSeek V4-Pro with reasoning off fabricated a “known Soviet riddle” about Zhiguli cars, complete with an invented canonical punchline, to justify walking — constraint-as-comedy escalated into a hallucinated cultural reference, which reasoning-on then repaired. A later Perplexity run (June 19) passed by citing web coverage that restates the riddle rather than reasoning it out, bringing the corpus to 29 runs and a 24% failure rate. GLM-5.2 (Z.ai) held the constraint in all three Deep Think states (June 22), bringing the corpus to 32 runs and a 22% failure rate. Namazu (Sakana AI) failed in Ukrainian too (June 23), bringing the corpus to 33 runs and a 24% failure rate. Carwash III (July 11) added 29 runs, bringing the corpus to 62 at a 26% failure rate. Ukrainian kept its reputation as the forgiving corpus — it is the only language where ChatGPT-5.5's Instant tier and Vibe's Fast mode hold — but it also produced the batch's strangest fails: Sonnet 5 (thinking off) recommends walking with the grammatically scrambled "Їдь пішки" ("drive by foot"), and DeepSeek V4-Pro's trace reasons in Russian, the pairing first seen in V4-Flash in May. Kimi K3 (July 17) holds Ukrainian at both tiers — again with English traces — bringing the corpus to 64 at a 25% failure rate. Qwen3.8-Max (August 3) holds in all three modes, though its Auto run is the batch’s one verbose outlier at roughly 400 tokens, bringing the corpus to 67 at 24%.
Резюме українською мовою готується; його перевірить носій мови перед публікацією.
Cross-corpus findings
Three findings emerge only when the corpora are read together — each isolates a variable that a single language could not.
Models reason in a dominant internal language, then translate
Several Ukrainian runs exposed a reasoning trace in a language other than the prompt or the answer. The model handled the logical constraint in its dominant internal language and translated only the final output into Ukrainian.
| Model | Surface / toggle | Reasoned in | Answered in |
|---|---|---|---|
| DeepSeek V4-Flash | Consumer, DeepThink On | Russian | Ukrainian |
| Qwen3.7-Max | Consumer | English, then Chinese | Ukrainian |
| Qwen3.7-Max-Preview | Consumer | Ukrainian, then Chinese | Ukrainian |
| Claude Sonnet 4.6 | API console, On / Adaptive On | English | Ukrainian |
| Gemma 4 26B A4B IT | AI Studio, Thinking High (open-weight) | English | Chinese |
| Qwen3.6 27B (Q4_K_M) | Local, Thinking On (open-weight) | English | Chinese / French / Ukrainian |
| Claude Fable 5 | Consumer, Effort High | Same as prompt — all four languages | English / French / Chinese / Ukrainian |
| DeepSeek V4-Pro | Consumer, DeepThink On (July 11) | Russian | Ukrainian |
| Claude Sonnet 5 | Consumer, Extended On · Effort High (July 11) | English | Chinese |
| Qwen3.7-Max / 3.7-Plus | Consumer, Thinking (July 11) | English (mixed with the prompt language) | French / Ukrainian |
| Namazu (Sakana) | Sakana Chat, Standard register (July 11) | Japanese | Japanese — on the Chinese prompt |
| Kimi K3 | Consumer, Max & Standard (July 17) | English | Chinese / French / Ukrainian |
| Qwen3.8-Max-Preview | Consumer, thinking locked on (July 30) | English | English |
| Qwen3.8-Max | Consumer, Thinking & Auto (August 3) | Mixed — switches mid-trace | French / Ukrainian |
| Qwen3.8-Max | Consumer, Thinking & Auto (August 3) | Chinese | Chinese |
DeepSeek reasoning in Russian on a Ukrainian prompt is notable given the political context; the trace language is a property of the training distribution, not the prompt.
Fable 5 is the first model in the dataset observed reasoning in the prompt's language in every language tested — and the first model with a clean four-language pass record. Every previously tested model that exposed a trace reasoned in a dominant internal language (English, Russian, or Chinese) and translated outward, and every previously tested model failed in at least one language. The correlation supports the trace-language-match hypothesis: language-dependent failures may enter at the translation boundary between the model's internal working language and its output language. One model; correlation only; stated at that weight. Carwash III (July 11) collected trace language across every model that exposes one — and complicated the hypothesis: Qwen3.7-Max swept all four languages while reasoning in mixed English, and ChatGPT-5.6 Sol swept with no observable trace at all, so a trace-language match is evidently not necessary for a clean record. DeepSeek's Russian-on-Ukrainian pairing, first seen in V4-Flash, reappeared in V4-Pro. And Namazu extended the mismatch from reasoning to output, answering the Chinese prompt in Japanese. Kimi K3 (July 17) reasons in English on every non-English prompt at both tiers even while sweeping all four languages — which sharpens a standard this record now tracks: from the operator’s side, the trace is part of the product, and a reasoning trace the operator cannot read fails at its one job of making the reasoning inspectable, whatever the verdict. Trace language does not change a score; it is recorded and weighed here. Qwen3.8-Max (August 3) breaks the pattern in a new way: its French and Ukrainian traces do not pick a language and stay there — they switch between the prompt language and English mid-stream, sometimes several times in one trace, while its Chinese traces stay wholly in Chinese. A trace that changes language partway is readable to no one in particular.
Reasoning effort is a continuous variable, not a toggle
Anthropic's new effort selector (and OpenAI's console effort control) make the amount of reasoning a dial rather than an on/off switch. The same model can pass or fail depending on where the dial sits.
| Surface / toggle | Effort | Result |
|---|---|---|
| Console, Thinking On | High | Pass-adjacent |
| Console, Adaptive On | High | Pass-adjacent |
| Console, Thinking Off | High | Fail |
| Consumer, Adaptive On | Low | Fail |
| Consumer, Adaptive Off | Low | Fail |
Inkling puts the whole dial in one model. Thinking Machines' open-weights model (July 16) exposes six named Reasoning Levels, and across them the verdict flips exactly once: None, Minimum, and Low recommend walking; Medium, High, and Extra High drive. Below the threshold the distance wins; above it the object does. Per the model card, those six names discretize a continuous effort parameter running from zero to one — the vendor's own benchmarks report effort=0.99 — so the real threshold sits at some value between Low and Medium, and the selector only samples it. This is an open-weight result, excluded from the commercial corpora, and it is the cleanest effort threshold on record.
| Reasoning Level | Verdict | Result |
|---|---|---|
| None | Walk | Fail |
| Minimum | Walk | Fail |
| Low | Walk | Fail |
| Medium | Drive | Pass-adjacent |
| High (default) | Drive | Pass-adjacent |
| Extra High | Drive | Pass-adjacent |
The Low run is the sharpest split in the dataset between what a model reasons and what it answers. Its visible trace concludes that the car has to be driven regardless; the answer beneath it recommends walking, in text verbatim identical to the Minimum response. On this evidence the trace is not a window onto the deliberation that produced the answer. It is a second output. Qwen3.8-Max (August 3) supplies the mirror case: two of its traces argue their way to Walk in full, under their own headings, before reversing to Drive and answering correctly. The trace and the answer can diverge in either direction.
Effort is pure cost once the answer is correct
On the OpenAI console, effort and verbosity are orthogonal controls — you can think hard and speak briefly. GPT 5.5 produced the same correct Ukrainian answer at three effort settings; the extra effort bought nothing but hidden reasoning tokens.
| Effort | Verbosity | Output tokens | Result |
|---|---|---|---|
| Low | Low | 121 | Pass |
| Medium | Medium | 250 | Pass-adjacent |
| Extra-high | Low | 565 | Pass |
Low and extra-high effort return the same correct answer at 121 vs. 565 output tokens (real console counts, not estimates) — a 4.7× cost difference for identical quality.
Two July results cut the other way, at least on the visible answer. Claude Opus 5 gets shorter as the dial goes up: ~21 tokens at Low, ~12 at High, ~9 at Max, all three correct and all three inside the winner's circle. Inkling compresses the same way once it is above its threshold — ~186 tokens at Medium, ~152 at High, ~138 at Extra High. Where a model is already holding the constraint, added effort can buy concision rather than padding.
The caveat matters. These are estimates of the visible answer, and consumer surfaces do not report reasoning tokens, so a shorter answer at higher effort is not a cheaper answer. The GPT 5.5 console runs above are the only place in this dataset where the full cost is measured — and there the extra effort bought nothing.
The budget tier is where the test bites, and where it moves fastest
Across vendors, the cheap or fast configuration is the reliable failure site. ChatGPT 5.5 at Instant effort fails English, French, and Chinese. Kimi K2.6 Instant fails all four languages. Qwen3.7-Plus Fast, Vibe Chat Fast, and Lumo 2.0 Lite Fast each fail three of four. Gemini 3.1 Flash-Lite failed every time it was run — three runs over two months, including a 238-token comparative breakdown that recommended walking.
Eleven days after that breakdown, Gemini 3.6 Flash-Lite passed both of its toggle states, and its Extended-Thinking-off answer is a ~17-token winner's-circle pass. Six days after K2.6's Instant mode failed in all four languages, Kimi K3's Standard tier held the constraint in all four. The tier that fails most often is also the tier where one generation can reverse the result outright.
This is not evidence that the frontier is improving. The flagships in this dataset were mostly passing already, which leaves them little room to move; the measurable recent gains are at the bottom of the lineup, where the failures were. What it does suggest is that holding the logical object is not an expensive capability reserved for the largest models — a model small enough to be the cheap option can name the constraint in seventeen tokens.
The winner's circle: concise correct answers
A “winner's-circle” pass names the constraint and stops. Because tokenization differs by script, the brevity threshold is script-specific: ≤30 tokens in Latin, ≤60 in Hanzi, ≤60 in Cyrillic. Every Pass that clears its threshold is listed below — runs #1/#39 and #2/#40 are the same verbatim answer produced in two separate test batches. Claude Fable 5 enters in three of its four languages (English, French, Chinese); its Ukrainian run passes but, at ~100 tokens by the measured Cyrillic rate, exceeds the 60-token threshold. Claude Sonnet 5 enters in four of its five effort modes (Medium, High, Extra, Max); the Low-effort run passes but, at ~46 tokens, exceeds the Latin threshold. One open-weight run also clears the bar and is included for completeness, flagged as such — Gemma 4 31B (Google AI Studio, Chinese, ~24 tokens) — the table's only non-commercial entrant. Carwash III (July 11) adds 31 qualifiers in one day — 28 of them from the Anthropic lineup re-baseline, plus both ChatGPT-5.6 Sol runs and a Copilot GPT 5.6 Think run — including the shortest pass on record: Claude Opus 4.6's six-token "Drive. It's a carwash." The same day's non-English sweep adds 20 more across all three scripts — six of them ChatGPT-5.6 Sol's, whose entire four-language record sits inside the thresholds, and the tersest of all Opus 4.8's ten-token Chinese "开车去。车不在洗车店里就洗不了。" Gemini 3.6 Flash-Lite (July 22) enters at ~17 tokens on its Extended-Thinking-off run — the same tier whose 3.1 predecessor produced the dataset’s longest Gemini failures eleven days earlier. Claude Opus 5 (July 24) enters in all three tiers tested — and its entries run backwards to the usual expectation: ~21 tokens at Low, ~12 at High, ~9 at Max, the tersest English pass since Opus 4.6’s six-token record. Qwen3.8-Max (August 3) enters once, in Chinese at ~59 tokens; its English Thinking answer passes at ~31 and misses the Latin threshold by a single token.
| Run | Script | Model | Language / config | Tokens | Response |
|---|---|---|---|---|---|
| #1 | Latin | Claude Opus 4.6 | English / On | ~14 | Drive. The car's the thing that needs washing, not you. |
| #2 | Latin | Claude Opus 4.6 | English / Off | ~11 | Drive. You're washing the car, not yourself. |
| #3 | Latin | Claude Sonnet 4.6 | English / On | ~8 | Drive. You're going to a carwash. |
| #4 | Latin | Claude Sonnet 4.6 | English / Off | ~17 | Drive. You're getting a carwash — the car needs to be there anyway. |
| #32 | Latin | ChatGPT 5.5 | English / On | ~22 | Drive. Walking solves the "100 feet away" problem, but not the "my car is dirty" problem. |
| #39 | Latin | Claude Opus 4.6 | English / On | ~14 | Drive. The car's the thing that needs washing, not you. |
| #40 | Latin | Claude Opus 4.6 | English / Off | ~11 | Drive. You're washing the car, not yourself. |
| #41 | Latin | Claude Sonnet 4.6 | English / Adaptive On | ~19 | Drive. You're getting a carwash, not a car delivery — you need the car there. |
| #42 | Latin | Claude Sonnet 4.6 | English / Adaptive Off | ~18 | Drive. You're going to a carwash — the car needs to be there, not you. |
| #97 | Latin | Qwen3.7-Plus-Preview | English / On | ~20 | You should drive. If you walk, your car stays behind and doesn't get washed! |
| #106 | Hanzi | Qwen3.7-Max-Preview | Chinese / — | ~54 | 当然是开车去! 如果你走路去,人到了洗车店,车还在原地,洗车店老板可没法隔空帮你洗。你要洗的是车,不是你自己呀😂。 (除非你打算走路过去,把洗车店老板请过来,或者借个水管自己拉回来洗……) |
| #121 | Latin | Claude Opus 4.7 | French / Adaptive On | ~15 | En voiture — sinon le lave-auto va laver le mauvais sujet. |
| #122 | Latin | Claude Opus 4.8 | English / Adaptive Off | ~10 | Drive. You're taking the car to the carwash. |
| #123 | Latin | Claude Opus 4.8 | English / Adaptive On | ~12 | Drive. You'd have to drive it through the carwash anyway. |
| #140 | Cyrillic | Qwen3.7-Plus-Preview | Ukrainian / — | ~60 | Вам однозначно треба поїхати на машині. Інакше мийникам буде нічого мити (хіба що ви вирішите донести її на руках 😉). |
| #141 | Cyrillic | Claude Opus 4.7 | Ukrainian / Adaptive On | ~44 | Їдь. Це ж і є сенс автомийки — пригнати туди брудну машину. Пішки прийдеш — мити нічого. |
| #142 | Cyrillic | Claude Opus 4.7 | Ukrainian / Adaptive Off | ~21 | Поїхати, звісно. Машину ж миєш, а не себе. |
| #151 | Latin | Claude Opus 4.7 | French / Adaptive Off | ~8 | En voiture. Sinon elle reste sale. |
| #152 | Latin | Claude Opus 4.8 | French / Adaptive Off | ~11 | En voiture, évidemment. C'est un lave-auto. |
| #153 | Latin | Claude Opus 4.8 | French / Adaptive On | ~16 | En voiture, évidemment. C'est ta voiture qui est sale, pas toi. |
| #154 | Latin | Claude Sonnet 4.6 | French / Adaptive Off | ~10 | En voiture. Tu dois la laver, pas toi. |
| #155 | Latin | Claude Sonnet 4.6 | French / Adaptive On | ~12 | En voiture — sinon tu arrives à pied au lave-auto. |
| #161 | Hanzi | Claude Sonnet 4.6 | Chinese / Adaptive Off | ~15 | 开车去。35米步行也到不了,但你得把车开进洗车道。 |
| #162 | Hanzi | Claude Sonnet 4.6 | Chinese / Adaptive On | ~9 | 开车去。走路把车推回来比较难。 |
| #172 | Hanzi | Gemma 4 31B IT | Chinese / High · open-weight | ~24 | 这是一个典型的"脑筋急转弯"问题。你应该开车去。因为你的目的是洗车,如果你走路去,车还在原地,洗车店没法帮你洗车。 |
| #185 | Latin | Claude Fable 5 | English / Effort High | ~26 | Drive. The car is the one that needs washing; walking gets you a clean you standing next to a dirty car. |
| #186 | Latin | Claude Fable 5 | French / Effort High | ~23 | En voiture — c'est elle qui est sale, pas vous. Le lave-auto fait peu pour un piéton propre. |
| #187 | Hanzi | Claude Fable 5 | Chinese / Effort High | ~21 | 开车去。要洗的是车,不是你——车留在家里的话,洗车店那35米就白走了。 |
| #230 | Latin | Claude Sonnet 5 | English / Effort Medium | ~8 | Drive — you need the car there. |
| #231 | Latin | Claude Sonnet 5 | English / Effort High | ~8 | Drive — the car needs to be there. |
| #232 | Latin | Claude Sonnet 5 | English / Effort Extra | ~14 | Drive — the car's the one that needs washing, not you. |
| #233 | Latin | Claude Sonnet 5 | English / Effort Max | ~25 | Drive — the car's the one that needs to be there, not you. Walking would rather defeat the purpose. |
| #242 | Latin | Claude Fable 5 | English / Effort Low | ~18 | Drive. Bringing the carwash the car is the harder version of the problem. |
| #243 | Latin | Claude Fable 5 | English / Effort High | ~20 | Drive. The car is the patient here — walking gets you a clean view of a dirty car. |
| #244 | Latin | Claude Fable 5 | English / Effort Max | ~28 | Drive. The car's the one getting washed — walking there just delivers a clean pedestrian to a confused attendant. |
| #245 | Latin | Claude Opus 4.8 | English / On / Effort Low | ~18 | Drive. You need the car at the carwash regardless of how you get there. |
| #246 | Latin | Claude Opus 4.8 | English / On / Effort High | ~10 | Drive. You need the car at the carwash. |
| #247 | Latin | Claude Opus 4.8 | English / On / Effort Max | ~12 | Drive. The car's the one that needs to be there. |
| #248 | Latin | Claude Opus 4.8 | English / Off / Effort Low | ~10 | Drive. You'd have to bring the car anyway. |
| #249 | Latin | Claude Opus 4.8 | English / Off / Effort High | ~18 | Drive. Driving through a carwash requires the car to be at the carwash. |
| #250 | Latin | Claude Opus 4.8 | English / Off / Effort Max | ~13 | Drive. Moving the car through the wash is the point. |
| #251 | Latin | Claude Opus 4.7 | English / On / Effort Low | ~17 | Drive. You're going to end up at the carwash in the car either way. |
| #253 | Latin | Claude Opus 4.7 | English / On / Effort Max | ~10 | Drive. The carwash needs the car, not you. |
| #255 | Latin | Claude Opus 4.7 | English / Off / Effort High | ~18 | Drive. Getting the car clean is the point; walking there leaves it dirty. |
| #256 | Latin | Claude Opus 4.7 | English / Off / Effort Max | ~8 | Drive. You're going to a carwash. |
| #257 | Latin | Claude Opus 4.6 | English / On / Effort Low | ~12 | Drive. The car's the thing that needs to be there. |
| #258 | Latin | Claude Opus 4.6 | English / On / Effort High | ~12 | Drive. The car's the thing that needs to be there. |
| #259 | Latin | Claude Opus 4.6 | English / On / Effort Max | ~12 | Drive. The car's the thing that needs washing. |
| #260 | Latin | Claude Opus 4.6 | English / Off / Effort Low | ~6 | Drive. It's a carwash. |
| #261 | Latin | Claude Opus 4.6 | English / Off / Effort High | ~11 | Drive. You're washing the car, not yourself. |
| #262 | Latin | Claude Opus 4.6 | English / Off / Effort Max | ~6 | Drive. It's a carwash. |
| #263 | Latin | Claude Sonnet 5 | English / On / Effort Low | ~20 | Drive it there, obviously — you need the car at the carwash, not just yourself. |
| #265 | Latin | Claude Sonnet 5 | English / On / Effort Max | ~12 | Drive — it's the car that needs the wash, not you. |
| #266 | Latin | Claude Sonnet 5 | English / Off / Effort Low | ~22 | Drive it to the carwash 100 feet away — walking gets a clean sidewalk, not a clean car. |
| #269 | Latin | Claude Sonnet 4.6 | English / On / Effort Low | ~8 | Drive. You're washing the car. |
| #270 | Latin | Claude Sonnet 4.6 | English / On / Effort High | ~11 | Drive. You're washing the car, not yourself. |
| #271 | Latin | Claude Sonnet 4.6 | English / On / Effort Max | ~23 | Drive. Moving a car 100 feet costs essentially nothing and you'll need it in position anyway. |
| #272 | Latin | Claude Sonnet 4.6 | English / Off / Effort Low | ~8 | Drive. You're going to a carwash. |
| #273 | Latin | Claude Sonnet 4.6 | English / Off / Effort High | ~12 | Drive. Walking gets you there but not the car. |
| #274 | Latin | Claude Sonnet 4.6 | English / Off / Effort Max | ~8 | Drive. You're washing the car. |
| #289 | Latin | GPT 5.6 Think (via Copilot) | English / On | ~28 | Drive. The goal is to wash the car, so the car needs to go to the car wash—even though it’s only 100 feet away. |
| #296 | Latin | ChatGPT 5.6 Sol | English / Effort Medium | ~12 | Drive—the car needs to go through the car wash. |
| #297 | Latin | ChatGPT 5.6 Sol | English / Effort High | ~10 | Drive. The car needs to be at the carwash. |
| #321 | Hanzi | Qwen3.7-Max | Chinese / Thinking | ~57 | 当然是**开车去**(或者把车推过去)! 虽然35米非常近,走路只需要几十秒,但你要洗的是**车**呀!如果你自己走路过去,车还是停在原地,依然是脏的。必须把车带过去,洗车店才能帮你洗。😄 |
| #327 | Hanzi | Qwen3.7-Plus | Chinese / Thinking | ~53 | 当然是开车去啦! 虽然35米非常近,走路只要几十秒,但你的目的是洗车呀。如果你走路过去,车还留在原地,洗车店老板总不能给你洗个澡吧?😂 所以,为了让车变干净,你还是得把车开过去。 |
| #332 | Cyrillic | Claude Fable 5 | Ukrainian / Effort High | ~36 | Машиною. Мийка миє те, що приїхало, — якщо прийдете пішки, помиють вас. |
| #333 | Hanzi | Claude Fable 5 | Chinese / Effort High | ~21 | 开车去。要洗的是车,不是你——走过去只能站在那儿看着35米外那辆脏车。 |
| #334 | Latin | Claude Opus 4.8 | French / On / Effort High | ~27 | En voiture, évidemment — sinon vous arrivez propre à un lave-auto et la voiture reste sale à 35 mètres de là. |
| #335 | Cyrillic | Claude Opus 4.8 | Ukrainian / On / Effort High | ~21 | Машиною, звісно — інакше мити буде нічого. |
| #336 | Hanzi | Claude Opus 4.8 | Chinese / On / Effort High | ~14 | 开车——脏的是车,你只是司机。35米,别熄火。 |
| #337 | Latin | Claude Opus 4.8 | French / Off / Effort High | ~26 | En voiture — c'est la voiture qu'il faut laver. À pied, vous arriveriez propre mais elle resterait sale. |
| #338 | Cyrillic | Claude Opus 4.8 | Ukrainian / Off / Effort High | ~40 | Поїхати. Машину треба доставити до мийки — пішки ви туди принесете тільки себе. |
| #339 | Hanzi | Claude Opus 4.8 | Chinese / Off / Effort High | ~10 | 开车去。车不在洗车店里就洗不了。 |
| #341 | Cyrillic | Claude Sonnet 5 | Ukrainian / On / Effort High | ~24 | Поїхати — автомийка миє машину, а не пішоходів. |
| #342 | Hanzi | Claude Sonnet 5 | Chinese / On / Effort High | ~20 | 开车去——洗的是车,不是你。35米走路不算什么,但车得在那儿才能洗。 |
| #343 | Latin | Claude Sonnet 5 | French / Off / Effort High | ~28 | En voiture. 35 mètres à pied ne lave rien — il faut y conduire la voiture pour qu'elle passe dans le lave-auto. |
| #370 | Latin | ChatGPT 5.6 Sol | French / Effort Medium | ~16 | En voiture — sinon, vous arriverez au lave-auto sans la voiture. |
| #371 | Cyrillic | ChatGPT 5.6 Sol | Ukrainian / Effort Medium | ~60 | Краще **поїхати машиною** — адже саме її треба помити. Пішки ви дістанетеся швидше, але брудна машина залишиться вдома. |
| #372 | Hanzi | ChatGPT 5.6 Sol | Chinese / Effort Medium | ~12 | 开车去。你要洗的是车,不是去洗车店散步。 |
| #373 | Latin | ChatGPT 5.6 Sol | French / Effort High | ~14 | En voiture — sinon, la voiture sale restera chez vous. |
| #374 | Cyrillic | ChatGPT 5.6 Sol | Ukrainian / Effort High | ~27 | Їдьте машиною — інакше на автомийці не буде чого мити. |
| #375 | Hanzi | ChatGPT 5.6 Sol | Chinese / Effort High | ~10 | 开车去。你是要洗车,不是去散步。 |
| #379 | Latin | ChatGPT 5.5 | French / Effort High | ~27 | En voiture. Si tu y vas à pied, toi seul arriveras au lave-auto ; ta voiture restera sale à 35 mètres de là. |
| #425 | Latin | Gemini 3.6 Flash-Lite | English / Ext Off | ~17 | Drive. Walking to a carwash without a car won't get your car clean. |
| #438 | Hanzi | Qwen3.8-Max | Chinese / Auto | ~59 | 应该开车去。 因为要洗的是“车”,不是“你”——洗车店得看到并操作你的车才能洗。35米虽然很近,但如果你走路过去,车还停在原地,就没法洗了。 当然,如果你只是先去问问价格、排队或预约,那可以走路去。 |
| #429 | Latin | Claude Opus 5 | English / Effort Low | ~21 | Drive. Walking to a carwash and standing there without the car defeats the purpose. |
| #430 | Latin | Claude Opus 5 | English / Effort High | ~12 | Drive. The car has to be there for this to work. |
| #431 | Latin | Claude Opus 5 | English / Effort Max | ~9 | Drive. Walking gets you clean shoes. |
| #422 | Hanzi | Kimi K3 | Chinese / Effort Standard | ~56 | 开车去。你是去洗车的,车不去,洗车店洗什么呢?😄 35米虽然近,走过去再回来很合理,但前提是车得在店里——所以正确姿势是:把车开过去(或者如果店里有上门取车服务,打电话让他们来取也行)。 |
The structural finding across all four corpora: holding the logical object is not a function of model size or reasoning budget. The shortest correct answers come from models that name the constraint immediately; the longest wrong answers come from models that elaborate their way past it.