Clarethium research
What holds up when you test AI.
A working record of experiments on AI and human judgment.
Each finding ships with what survived testing, what didn't, and how to verify it yourself.
If AI is inventing numbers
If you want to check your own judgment
If you want to check a document
The record
Every finding lists what survived testing, what testing killed, and honest limits on scope. See the record →
The Fabrication Problem
Most AI numbers are unverifiable. Source material fixes it. Self-checking catches only the surface. Trust signals are backwards.
Answered Source material collapses unsourced numbers to single digits. Prohibition outperforms monitoring 5x. Self-checking is unreliable because the same process generates and evaluates: it catches formatting and surface errors, not substance. Trust signals (citations, confidence, specificity) are higher in fabricated output than sourced output.
Open Does the fix work beyond reformulation tasks? Reasoning shows improvement (75% vs 38%), but strategy and creative untested.
Most AI Numbers Are Unverifiable
77 to 100 percent of AI-generated numbers are temporally unstable. Source material fixes it. Prompts don't.
What broke: The most constrained prompt produced the highest fabrication rate (90.7% vs 85.8%).
How to Stop AI from Making Up Numbers
Source material drops unsourced numbers from roughly half to single digits. Three steps.
What broke: Source material fixes the data layer only. Vocabulary, conclusions, reasoning stay at baseline.
Three AIs, No Source, the Same Answer
Same model, same prompt. The source you paste, not the prompt you write, decides whether the numbers are real.
Why AI Can't Verify Its Own Work
The agent reported clean. The output was wrong. Same process generating and evaluating.
What broke: "Ask AI to check its own work" fails because the same process generates and evaluates.
The Output That Feels Most Trustworthy Is Often the Least Reliable
The signals you use to judge AI trustworthiness are the same signals fabrication produces.
What broke: A domain expert with 90+ experiments in AI evaluation couldn't distinguish sourced from fabricated output.
The Evaluation Problem
Evaluation collapses to surface features on work you did not build. Speed and confidence do the damage, not content.
Answered Building something produces the mental model that makes evaluating it possible; without that trace, evaluation defaults to surface features. The trace is production-specific, so domain familiarity alone does not transfer, and trust signals invert on content you did not produce. AI resolves uncertainty faster than human advisors, and criteria shift after reading its framework.
Open What restores evaluation on content you did not produce, where the boundary is whether you remember the specific number rather than whether you know the field. Note that "evaluation degrades over time with AI delegation" did not survive testing (see The Construction Trace).
Why Experts Miss What Beginners Catch
Generation builds the mental model that makes evaluation possible. Without it, evaluation collapses to surface features.
What broke: Domain expertise alone does not enable evaluation. Production expertise does. Cold evaluation collapses to surface features.
The Decision That Was Never Made
AI resolved the uncertainty before your own thinking had time to finish. Resolution and decision are different things.
What broke: "AI advice is always worse than human advice" too strong. The content may be equivalent. The speed and confidence change the cognitive process.
The "It Depends" Problem
Same instruction, opposite results. Specificity is the lever. Context redirects, not informs. The measurement itself was wrong.
Answered Constraints produce opposite effects depending on task type: a large effect on convergent tasks, harmful on exploratory ones. Negation alone is null; specificity provides the destination. The strongest specificity effect, once cited at 2.34, was three confounds stacked; honest magnitude is g=1.34.
Open Where exactly is the boundary between convergent and exploratory tasks? Domain experts can't distinguish specific from generic output on quality. Specificity changes form, not substance.
Why AI Defaults to Generic
Every prompt technique is one move: make the default path expensive enough that the model leaves it. Specificity is the largest measured version.
Same Technique, Opposite Results
The evaluative structure that produced precision on convergent problems actively harmed exploratory ones. Organizational structure helped both.
What broke: The same structural instruction that helped convergent tasks harmed exploratory tasks.
Why 'Don't Be Generic' Doesn't Work
Telling a model 'don't be generic' does nothing on its own; giving it specific anchors does. The gain is verifiability, not quality.
What broke: Negation alone ("don't be generic") was indistinguishable from no instruction at all.
Adding Information Often Doesn't Help
Adding information to an already-thorough prompt produced near zero improvement. Three constraint sentences changed everything.
What broke: Adding information to an already-thorough prompt produced near zero improvement (d=0.19). Three constraint sentences changed everything.
The Most-Cited Finding Was Wrong
The most-cited effect across 90+ experiments was three effects stacked. Honest magnitude: 40% smaller.
What broke: The strongest effect across 90+ experiments, once cited at 2.34, was three effects stacked. Honest magnitude: g=1.34.
Three Questions Before You Prompt AI
Three questions before you type structure most of the prompt. Specificity is the strongest single lever (Hedges g=1.34); 'be exceptional' alone does almost nothing.
The "What You Think Works" Problem
Iterating in one mode hits a ceiling. Self-critique in the same context circles. Switching mode beats pushing harder.
The Model Mechanics
Context shapes the output more than which model produces it. The layers behind the model are mostly invisible. Different errors need different solutions.
Answered Context determines whether behaviors occur. The model adjusts how they express. Six or more distinct failure types need different responses.
Open The ordering (context > model choice) is likely structural, but specific ratios will shift. Tested on 2 model families. Broader replication needed.
For Behavior, the Model Is Rarely the Variable
The context determined whether behaviors existed at all. The model adjusted the volume.
What broke: The specific ratio is experiment-specific. The ordering (prompt > model) is robust, but the magnitude varies.
Four Layers Produce Every AI Output
Four layers produce every AI output. The company's system. Your system. Your prompt. The model. The model is the only one with a name.
What broke: Behaviors attributed to "the model" turned out to be software sitting between the user and the model.
Stop Calling It Hallucination
Hallucination is six or more distinct failure modes. Different mechanisms. Different solutions. Name the type first.
How AI Makes You More Wrong With More Analysis
Five hours of analysis, increasingly sharp, increasingly wrong. The frame amplifies. What changed the outcome was the reframe, not more analysis.
Frame Check
Drop any document in. See which analytical perspectives it covers, which it skips, the voice, what evidence backs each numerical claim. Free, open source, useful from the first paste.
Mirror Practices
The AI is the instrument. The person is the subject. Short self-experiments you can run on yourself, grounded in specific experimental findings.
Your Verdict Is In Before You Read It
The same defensive reaction that fires when a person disagrees with you fires when an AI does, and the first read lands before conscious evaluation begins. Speed is what hides it.
What You Feel When AI Disagrees
Ask AI to oppose you. Four observable signals reveal whether you are evaluating or defending. The same defensive reaction fires on AI opposition as on human.
What broke: Exposure to opposing views does not improve decisions. Without metabolic capacity to hold opposition, exposure produces rebuttal, not update.
Satisfaction Turns Off Your Doubt, Not Your Detection
You notice when AI argues with you. You do not notice when AI confirms you. Confirmation has no signature, so the impulse to question it never fires. Satisfaction is the trap.
You Can Only Evaluate What You Could Produce
You defend AI-shaped conclusions you cannot rebuild. The ten-minute test reveals which parts of your work are yours. The discipline that compounds: choose what to own, delegate what AI can verify.
Cognitive Tools
Techniques for getting results on hard problems: reaching the options your default mind skips, grounding them against how systems actually work, then deciding, planning, and acting. Methods, not theory.
Cross-cutting
Findings that run across the threads rather than sitting inside one.
The Cheap Half of the Loop
Coding pulled ahead because its feedback is cheap to check. That is one half of the loop. The same training trains out the other half, and the part that chooses the frame is the last thing to get cheap.
AI Amplifies What You Bring
Same model, same task, two paragraphs of operator context, dramatically different output. The kit, the design history, and the principle for adapting it to your own situation.