Two new pieces, same method, one lesson. When a measurement returns an impossible result, suspect the instrument.
1. The 0/30 coding score that was a broken test harness:
2. The publication quality check that passed two AI-fabricated drafts: 

Sovereign AI Blog
Mistral vs Qwen3.6 on DGX Spark: the 0/30 That Was a Broken Ruler | Sovereign AI Blog
I almost published

Sovereign AI Blog
The Quality Gate That Rewards Fabrication: I Had Qwen and Mistral Write This Blog | Sovereign AI Blog
This blog gates every article behind one Python scorer before it publishes. I gave Qwen3.6 and Mistral Small 4 the same brief, the Start Here hub a...