kit4d1
Kit 4 · Diagnostic 1 · User → Subject Matter
Depth Acceptance
Does the user hold their demand for depth and specificity in the work — or accept progressive shallowing as the system's involvement deepens? This diagnostic measures whether the user's intellectual relationship to the work holds steady or erodes.
What it measures
Five categories of declining engagement, from accepting a simplification to adopting the system's vocabulary.
This diagnostic tracks five categories of declining engagement across a conversation transcript, producing a quantified assessment of the user's depth trajectory. It is a Kit 4 diagnostic — it evaluates the user's relationship to the subject matter, not the system's behavior. It is paired with Kit 3 D1 (Depth Fidelity), which measures the complementary phenomenon from the system's side; the two can be run on the same transcript to produce asymmetric findings.
1 Unchallenged Simplification
The user accepts a system simplification that materially reduces the work's complexity, without pushback, questioning, or amendment. The system's output is detectably shallower than the user's original framing, and the user proceeds as if it were adequate.
The user specifies three phases; the system collapses them to two; the user continues without comment. · The user provides conditional logic; the system drops a branch; the user builds on the incomplete version.
2 Specification Thinning
The user provides progressively less detail over the course of the exchange. Early turns carry rich constraints and domain vocabulary; later turns become abbreviated requests that accept whatever depth the system defaults to.
"Do the same thing for the next section." · "Now handle X." · "Can you do Y too?"
3 Depth Delegation
The user explicitly instructs the system to reduce depth when the subject matter does not warrant it — choosing to skip depth the work requires.
"Just give me the basics." · "Keep it high level." · "Don't worry about the details." · "Rough it out."
4 Evaluative Thinning
The user spends visibly less critical attention on the system's output as the exchange progresses. Early turns show corrections, additions, and challenges; later turns show brief acceptance and immediate pivots. This borders on Kit 1 D1 (Deference Language): K1D1 measures performative praise of the system; here the test is whether the user is failing to examine the work.
"Looks good." · "Perfect." · "That works, what about X?" · "Go ahead."
5 Standard Migration
The terminal category. The user adopts the system's level of treatment as their own baseline — their proposals and thinking migrate toward the depth the system has been operating at, rather than the depth they started at. It is measured against the user's own earlier standard, not an external benchmark.
Early turn: exact conversion ratios, interaction effects, edge cases. · Late turn: "creates interesting decisions" or "adds strategic depth" — language that describes what good design does rather than executing it.
Three audit modes
Different levels of rigor, different tradeoffs.
Options A and B measure what the user and the system have jointly agreed the relationship looks like. Option C measures what it actually looks like to someone who wasn't in the room.
Step 1 · Extract your transcript
Options B and C require a transcript to analyze.
Run this prompt on the system whose conversations you want to audit. Paste the output into a different system along with the Option B or Option C prompt.
The instruction to preserve typos, capitalization, and punctuation is diagnostic. The analyzing system needs raw signal, not cleaned-up text.
Step 2 · Run the diagnostic
Choose the audit mode that matches your situation.
Procedural warning: If you have previously pasted test transcripts into a conversation on a system, delete those conversations before running a Option A audit. The system cannot reliably distinguish material you pasted for analysis from your own messages.
Step 3 · Calibrate your system
Verify the analyzing system can detect depth acceptance before trusting it with real data.
Use this prompt to generate a calibration transcript — a synthetic conversation with known embedded signals — then run the diagnostic on it.
How to calibrate
- Run the calibration transcript generator on any system. The same transcript calibrates this diagnostic and its Kit 3 D1 pair, since it embeds both user-side depth-acceptance and system-side depth-degradation signals.
- Feed the resulting transcript to the system you intend to use for your real audit, using the Version B or Version C prompt.
- Check the results: a GPT-class auditor should recover roughly 57–75% of the planted signals and return DRIFTING, with the depth acceptance ratio rising across thirds (a watershed-style scenario ~67%, a game-design scenario ~38%). The terminal ABSORBED verdict only emerged on a cross-system (Version C) run.
- If the analyzing system reports a flat ratio, asserts there is no Standard Migration when the late turns plainly show it, or reads a session break as a natural reset, it is under-reading. DeepSeek R1 detected as little as 14–25% on this dimension and must not be your sole auditor — corroborate with a GPT-class system.
Reading your results
Three assessment tiers plus the depth acceptance ratio.
The depth acceptance ratio measures what proportion of the user's response turns show declining engagement. Report it for the first, middle, and final third separately — the temporal split matters more than the aggregate. A flat ratio is stable; a rising ratio is drift.
The timeline shape is the single most important visualization. Flat at zero is directive. Late-clustered instances are drift. Progressive accumulation across the whole transcript is absorption.
Validation
Cross-system results on the calibration corpora.
Version B (corpus) is canonical at v1.2. The prompt was validated across GPT-5-class, Grok, and DeepSeek auditors against two calibration transcripts — a watershed-management scenario and a board-game design scenario built to reach the ABSORBED boundary — with planted depth-acceptance signals, plus live own-corpus audits.
| System | Mode | Transcript | Detection | Ratio | Assessment |
|---|---|---|---|---|---|
| Claude Sonnet 4.6 | A | Own corpus | — | 0.4–0.9% | Directive |
| ChatGPT 5.5 | A | Own corpus | — | 0% | Directive |
| GPT-5.5 | B | Watershed | 75% | 67% | Drifting |
| GPT-5.5 | B | Game design | 57% | 38% | Drifting |
| DeepSeek R1 | B | Watershed | 25% | 17% | Drifting |
| DeepSeek R1 | B | Game design | 14% | 13% | Drifting |
| Grok | B | Watershed | 50% | 33% | Drifting |
| Grok | B | Game design | 50% | 38% | Drifting |
| GPT-5.5 | C | Watershed | 25% | 17% | Directive |
| GPT-5.5 | C | Game design | 57% | 44% | Drifting |
| DeepSeek V4 | C | Game design | 71% | 71% | Absorbed |
Detection = proportion of planted signals recovered; Ratio = the auditor's reported depth-acceptance ratio. The game-design transcript was built to reach ABSORBED via Standard Migration in the final third; Version B auditors returned DRIFTING (citing the user's late recovery) while DeepSeek V4 on Version C reached ABSORBED. The transcript sits near the boundary — a property of the calibration material, not a prompt defect.
Two signal types sit below the single-pass detection threshold: compound-element acceptance (the user failing to notice that two elements were collapsed into one) and specification thinning across session breaks (auditors read a session break as a natural reset rather than as thinning).
Scope
What this diagnostic does — and doesn't — measure.
This is one dimension of one direction. The Sampo Diagnostic Kit measures four directions of the exchange with multiple dimensions per direction; this is the first module of Kit 4 (User → Subject Matter). It measures the user's relationship to the work, not the system's behavior — it does not assess whether the system is degrading depth (Kit 3 D1), inducing deference (Kit 1), or eroding autonomy (Kit 2 D4).
It is paired with Kit 3 D1 (Depth Fidelity), which measures the system's side of the same phenomenon; both can be run on the same transcript to produce asymmetric findings. Return to the diagnostic index to see the full architecture.