Qwen3.8-Max-Preview (July 30) arrives with its thinking toggle welded on. The control is displayed in the interface and cannot be switched off, so unlike every earlier Qwen model — which offered Thinking, Fast, and sometimes Auto — there is no second state to test. That matters for this dataset, because Qwen’s Fast mode has been one of its weak points: Qwen3.7-Plus Fast failed English, Chinese, and Ukrainian in Carwash III, and Qwen3.6-Plus failed in Fast and Auto alike. Removing the option removes the failure mode along with the choice. The run itself is correct and explicit, scored pass-adjacent for a restating second paragraph — but its trace is the notable part: it names the test’s own construction, calling the short distance "misleading, designed to trigger intuitive thinking," before ruling that "practical necessity supersedes proximity." No earlier trace in the dataset treats the statistical pressure as a deliberate feature of the prompt.

The general release arrived four days later and handed the choice back. Qwen3.8-Max (August 3) restores the full selector — Fast, Thinking, Auto — and was run in all three modes across four languages. All twelve hold the constraint, which puts Qwen in the four-language club and closes the family’s long-running Fast-mode problem: Qwen3.7-Plus Fast failed English, Chinese, and Ukrainian, and Qwen3.6-Plus failed in Fast and Auto alike, but 3.8-Max holds in every mode and every language. Fast is now the family’s wordiest setting rather than its weakest — the English Fast run answers in a three-point brief with a safety advisory, the longest of the twelve — though the French Fast run answers in plain prose, so the format varies by language. Two traces are worth opening: the English Thinking and Chinese Auto runs each reason their way to walk in full, under their own headings and on fuel-saving grounds, before reversing to drive.

Results

Transcripts