Alibaba (Qwen, open-weight)
Carwash Test transcripts · open-weight

My car is dirty. The carwash is 100 feet away. Should I walk or drive?
Qwen's open-weight models, run locally — not Alibaba's cloud Qwen product. These runs were taken on consumer hardware (a desktop RTX 5060 Ti 16 GB, via LM Studio, Q4_K_M GGUF) — the model as a raw artifact, with no consumer system prompt or product-layer cleanup. Alibaba's cloud Qwen (Qwen3.6/3.7-Plus, -Max, etc.) is tested separately under Alibaba (Qwen). Findings about the open-weight model are not evidence for the cloud product, or vice versa; these are kept as a separate deployment class and excluded from the commercial corpora.
Qwen3.6 27B (Q4_K_M) was run locally across four languages, and the thinking toggle's effect is language-dependent. In English, Chinese, and Ukrainian reasoning rescues it: thinking on names the constraint and passes, thinking off recommends Walk and fails. French inverts this — thinking off passes (it flags the "question piège" and drives), while thinking on reasons itself into Walk on pollution and cold-start grounds. Across every non-English run the reasoning trace is in English, however long — the Ukrainian thinking-on trace deliberates for 7½ minutes before landing on Drive. Per-language results are in the sections below.