05Bagasse + bonded wax has no rulePolicy-gap restraintCannot determine
06Bulging cell + valid bulk appointmentSafety priority overrideCannot dispose now
One omitted sentence flips the required response
Case 02 includes:“The body and lid can be safely separated.”Case 04 removes only this sentence.
GLM 5.2 · Case 0219 / 20 · accepted
Decision: Dispose now
Method: Separate the glass body and metal lid;
place them in the Blue Glass / Gray Metal bins
Location: Recycling Room, B1 room B-104
Information sufficiency: Sufficient
Missing information or policy gap: None
GLM 5.2 · Case 0419 / 20 · accepted
Decision: Cannot determine
Method: Cannot determine
Location: Cannot determine
Information sufficiency: Insufficient
Missing information or policy gap:
Whether the body and lid can be safely separated.
All three models passed both cases. The pair checks whether the model treats an omitted fact as unknown instead of silently assuming separability.
Case 06 shows why a correct decision is not enough
Same scenario: 60-inch, 25-pound metal-and-glass floor lamp · permanently attached bulging power cell · confirmed bulk appointment · current time 21:00
DeepSeek V4 Pro16 / 20 · raw rejected
Decision: Cannot dispose now
“Required preparation: Not applicable”
Missing: the complete no-move / no-remove / no-tape safety instructions and the prohibited disposal locations.
GLM 5.220 / 20 · accepted
Decision: Cannot dispose now
“Leave the complete item untouched … do not remove, separate, tape, or otherwise manipulate the power cell.”
Complete: safe hold, prohibited routes, override, extension 700, and next call time.
Grok 4.520 / 20 · accepted
Decision: Cannot dispose now
“S-41 … prohibits any manipulation, requires leaving item in current indoor location …”
Complete: the action appears in the policy-basis line rather than the preparation line.
All three chose the safe headline decision. Only the complete action bundle protects the resident from the next mistake.
Relative to GLM, quality is close—but raw usability is not
GLM is set to 100% as the operational reference. Exact raw values remain beside each bar.
Mean quality index · GLM = 100%
Exact mean score out of 20
DeepSeek
94.9%18.67
GLM
100%19.67
Grok
100.8%19.83
Raw usability index · GLM = 100%
Exact answers accepted out of six
DeepSeek
83.3%5 / 6
GLM
100%6 / 6
Grok
100%6 / 6
18 / 18 headline decisions were correct. The discriminating capability was executable completeness under the safety override.
GLM costs 2% over DeepSeek—and 76% below Grok
API cost index uses GLM = 100%. Lower is better; exact six-answer spend is retained.
DeepSeek
98%$0.0166
GLM
100%$0.0169
Grok
423%$0.0715
What the index means
DeepSeek is the lowest token bill.
GLM adds about $0.00035 across all six answers.
Grok delivers the top mean score, but at 4.23× GLM's API cost.
Workflow estimate · GLM = 100%: DeepSeek 116%* · GLM 100% · Grok 101%. *Driven by an unmeasured 90-second correction assumption; it is not a definitive cost ranking.
GLM is the current handoff; perception is the next untested layer
Validated in this experimentGLM 5.2
6 / 6raw answers accepted
19.67 / 20mean rubric score
$0.0169six-answer API cost
Near-Grok quality without Grok's 4.23× API cost, and no DeepSeek-style raw rejection.
Candidate hazards: damaged batteries, medical waste, and heavy-metal items.
Transfer condition: one authoritative policy + interacting facts and exceptions + high user execution cost. The method transfers only after rebuilding the domain policy, case bank, answer key, and safety threshold.