Thread
The Fabrication Problem
Most AI numbers are unverifiable. Source material fixes it. Self-checking catches only the surface. Trust signals are backwards.
Answered Source material collapses unsourced numbers to single digits. Prohibition outperforms monitoring 5x. Self-checking is unreliable because the same process generates and evaluates: it catches formatting and surface errors, not substance. Trust signals (citations, confidence, specificity) are higher in fabricated output than sourced output.
Open Does the fix work beyond reformulation tasks? Reasoning shows improvement (75% vs 38%), but strategy and creative untested.
Most AI Numbers Are Unverifiable
77 to 100 percent of AI-generated numbers are temporally unstable. Source material fixes it. Prompts don't.
How to Stop AI from Making Up Numbers
Source material drops unsourced numbers from roughly half to single digits. Three steps.
Three AIs, No Source, the Same Answer
Same model, same prompt. The source you paste, not the prompt you write, decides whether the numbers are real.
Why AI Can't Verify Its Own Work
The agent reported clean. The output was wrong. Same process generating and evaluating.
The Output That Feels Most Trustworthy Is Often the Least Reliable
The signals you use to judge AI trustworthiness are the same signals fabrication produces.