These systems are not general-purpose chat models. They are LLMs optimized for a specific domain — shopping, media, customer service — where the optimization target shapes the response as much as the underlying model's reasoning does. The Carwash Test reveals how that optimization interacts with the prompt's surface features: the same mechanism produces opposite answers depending on what the system has been trained to find.

Results

Transcripts