Skip to content

Clarethium research

What holds up when you test AI.

A working record of experiments on AI and human judgment.

Each finding ships with what survived testing, what didn't, and how to verify it yourself.

Latest: Jul 12, 2026

The record

Every finding lists what survived testing, what testing killed, and honest limits on scope. See the record →

The Fabrication Problem

Most AI numbers are unverifiable. Source material fixes it. Self-checking catches only the surface. Trust signals are backwards.

Answered Source material collapses unsourced numbers to single digits. Prohibition outperforms monitoring 5x. Self-checking is unreliable because the same process generates and evaluates: it catches formatting and surface errors, not substance. Trust signals (citations, confidence, specificity) are higher in fabricated output than sourced output.

Open Does the fix work beyond reformulation tasks? Reasoning shows improvement (75% vs 38%), but strategy and creative untested.

The Evaluation Problem

Evaluation collapses to surface features on work you did not build. Speed and confidence do the damage, not content.

Answered Building something produces the mental model that makes evaluating it possible; without that trace, evaluation defaults to surface features. The trace is production-specific, so domain familiarity alone does not transfer, and trust signals invert on content you did not produce. AI resolves uncertainty faster than human advisors, and criteria shift after reading its framework.

Open What restores evaluation on content you did not produce, where the boundary is whether you remember the specific number rather than whether you know the field. Note that "evaluation degrades over time with AI delegation" did not survive testing (see The Construction Trace).

The "It Depends" Problem

Same instruction, opposite results. Specificity is the lever. Context redirects, not informs. The measurement itself was wrong.

Answered Constraints produce opposite effects depending on task type: a large effect on convergent tasks, harmful on exploratory ones. Negation alone is null; specificity provides the destination. The strongest specificity effect, once cited at 2.34, was three confounds stacked; honest magnitude is g=1.34.

Open Where exactly is the boundary between convergent and exploratory tasks? Domain experts can't distinguish specific from generic output on quality. Specificity changes form, not substance.

Mechanism4 minJun 22, 2026

Why AI Defaults to Generic

Every prompt technique is one move: make the default path expensive enough that the model leaves it. Specificity is the largest measured version.

Experiment4 minMar 24, 2026

Same Technique, Opposite Results

The evaluative structure that produced precision on convergent problems actively harmed exploratory ones. Organizational structure helped both.

What broke: The same structural instruction that helped convergent tasks harmed exploratory tasks.

Experiment3 minJul 12, 2026

Why 'Don't Be Generic' Doesn't Work

Telling a model 'don't be generic' does nothing on its own; giving it specific anchors does. The gain is verifiability, not quality.

What broke: Negation alone ("don't be generic") was indistinguishable from no instruction at all.

Experiment4 minApr 20, 2026

Adding Information Often Doesn't Help

Adding information to an already-thorough prompt produced near zero improvement. Three constraint sentences changed everything.

What broke: Adding information to an already-thorough prompt produced near zero improvement (d=0.19). Three constraint sentences changed everything.

Arc4 minMar 23, 2026

The Most-Cited Finding Was Wrong

The most-cited effect across 90+ experiments was three effects stacked. Honest magnitude: 40% smaller.

What broke: The strongest effect across 90+ experiments, once cited at 2.34, was three effects stacked. Honest magnitude: g=1.34.

Recipe5 minJul 5, 2026

Three Questions Before You Prompt AI

Three questions before you type structure most of the prompt. Specificity is the strongest single lever (Hedges g=1.34); 'be exceptional' alone does almost nothing.

The "What You Think Works" Problem

Iterating in one mode hits a ceiling. Self-critique in the same context circles. Switching mode beats pushing harder.

The Model Mechanics

Context shapes the output more than which model produces it. The layers behind the model are mostly invisible. Different errors need different solutions.

Answered Context determines whether behaviors occur. The model adjusts how they express. Six or more distinct failure types need different responses.

Open The ordering (context > model choice) is likely structural, but specific ratios will shift. Tested on 2 model families. Broader replication needed.

Mirror Practices

The AI is the instrument. The person is the subject. Short self-experiments you can run on yourself, grounded in specific experimental findings.

Cognitive Tools

Techniques for getting results on hard problems: reaching the options your default mind skips, grounding them against how systems actually work, then deciding, planning, and acting. Methods, not theory.

Cross-cutting

Findings that run across the threads rather than sitting inside one.