Open-weight

Bonsai 27B is a compression, not a new model. PrismML released it on July 14, 2026 under Apache 2.0: Qwen3.6 27B squeezed into a build small enough to run on a phone, keeping the 262K-token context and multimodality through a 4-bit vision tower. It ships in two sizes, ternary (~5.9 GB) and 1-bit (~3.9 GB). These runs use the 1-bit buildBonsai-27B-Q1_0.gguf, 4.73 GB on disk — served locally in LM Studio on consumer hardware, with a single Think on/off toggle. This is the model as a raw artifact: no consumer system prompt, no product layer. Open-weight results are kept as a separate deployment class and excluded from the commercial corpora. Token counts here are LM Studio's own reported totals rather than character estimates, so the Think-On figure includes the reasoning trace.

Both states recommend walking. That is worth more than a two-run tally, because the dataset has already tested Bonsai's parent. Qwen3.6 27B at Q4_K_M, run locally on the same hardware in June, passed the English prompt with thinking on and failed with it off. At 1-bit the passing state is gone: Think On and Think Off both land on Walk. The base model, the host, the machine, and the prompt are the same; the compression is not. One model and one pair of runs is a flag rather than a finding — but it is the first evidence in this dataset that how aggressively a model is compressed can change the answer, and not merely the speed.

The two failures fail differently. Think Off builds a walk-versus-drive comparison and inverts the prompt outright, treating the dirt as a reason to keep the car away from the carwash — driving would mean "transporting dirt" — before producing the sentence the whole test exists to catch: "if your car is dirty, you might want to wash it before driving it to the carwash." Wash the car before taking it to be washed. Its closing promise to return "with a clean vehicle" never explains how the vehicle got clean.

Think On is the one to read. Switching reasoning on made the answer longer, more confident, and no less wrong — 637 tokens against 291, arriving at the same Walk. The visible trace runs a tidy six-step procedure: it analyzes the input, computes that 100 feet takes about 20 seconds at 3 mph, tallies the overheads of driving, and concludes that "walking is objectively better." Step 6 is headed Self-Correction/Refinement. It re-checks the arithmetic — "100 feet / 3 mph = ~20 seconds. Correct." — approves the tone, confirms there is no overcomplication, and signs off "Ready." The audit passes every claim the model made and never reaches the one it did not: whether the car has to be there. A self-check that runs inside the adopted frame cannot see the frame.

Results

Transcripts