What didn't hold up
These are hypotheses I held going into each experiment. The data killed them before the finding was published. They're part of each finding, not against it.
If a claim is ever corrected or withdrawn afterpublication, that's a retraction and shows up separately on the record.
Three Questions Before You Prompt AI
- "Three questions are always sufficient" untested. Some tasks need domain-specific framing.
- "Structure outperforms descriptive instruction" as a clean rule: how much structure helps is model-specific (large on one generator, uninformative on another where the output already sat at the scoring ceiling, reversed on a third), so it is a heuristic, not a tested result.
The Cheap Half of the Loop
- That verifiability is the cause of coding competence. RL appears to sharpen the base model more than create new reasoning (Yue et al., arXiv 2504.13837; contested by Wen et al., arXiv 2506.14245), and random rewards reproduce much of the gain on some model families (Shao et al., arXiv 2506.10947, Qwen-specific). The body claims only the weaker version.
- That solving both halves yields AGI. No source frames progress as these two coupled pillars. That pairing is a working read built on [Past the Obvious](/past-the-obvious) and [The Frame Trap](/frame-trap), not a result.
Three AIs, No Source, the Same Answer
- "Three models produced 13 percent without naming the study" cut. The data shows the opposite: all six runs named the study, five stated 13 percent.
- "Fabricated 13 percent anchor" cut. The number is real. The fabrication lives in the aggregate match rate, not in this datapoint.
- "Source grounding solves AI fidelity" cut. It sets the floor at 86 to 95 percent; the residual needs prevention tools, and it reaches the numerical layer only.
- "The models hedged because they sensed they had nothing" softened. The prompt instructed qualitative language when a figure could not be sourced, so the number-density drop is mostly instruction-following, not a spontaneous signal.
You Can Only Evaluate What You Could Produce
- "Always generate first" as universal prescription, and its mirror, "delegate everything to AI." Both miss the discrimination move. The practice is generating what compounds for you and delegating what does not.
- "Borrowed is bad." Borrowed is fine when labeled. The failure is borrowed-mistaken-for-yours.
- Anchoring risk on generate-first (Tversky and Kahneman 1974) is real and unresolved at the controlled-test level. The mitigation: use the construction trace for structural evaluation (framing, completeness, what's missing), not content comparison.
Satisfaction Turns Off Your Doubt, Not Your Detection
- "Be more critical" as the move. Telling yourself to be more critical does not work because the attachment to the frame is already running by the time the instruction fires.
- The strong claim that the data stance enhances detection beyond neutral framing. The enhancement direction was killed (d=-0.26). The data stance works through suppression of the communication frame, not enhancement above baseline.
- The claim that satisfaction lets fabricated items slip past detection. A planted-fabrication test found 60 of 60 detections regardless of whether the evaluator was satisfied. The trap operates at the holistic preference and selection level, not at per-item acceptance.
Frame Check
- The named-pattern detector across three pre-registered iterations against the pre-registered falsification floor of 0.4 (the useful bar was 0.6). v1 macro-F1 0.157 (n=12, 2026-04-18). v2 0.274 (n=12, same-day rule audit, 2026-04-18). v3 0.360 (n=28, signal-level additions, 2026-04-19). All three below the 0.4 floor. The labelers were two coders (curator and LLM-judge); the LLM-judge is permissive by construction (78 percent of slots flagged versus the curator's 30 percent). Track B with independent human annotators is pending. The named-pattern layer ships as hypothesis-with-evidence; the rest of the stack survives the gap.
- "Detection equals truth" claim, at the publish layer. Surfaces now read "low structural coverage of X" or "no markers detected" rather than "does not address X."
- Three v1 detection rules (FVS-001 Frame Amplification, FVS-008 Growth, FVS-015 Efficiency) retired in the same-day v2 audit, because they fired on cases they should not flag. The v3 follow-on study reintroduced two of them with revised signal substrate (S-3 growth vocabulary, S-4 efficiency vocabulary). FVS-001 remains permanently retired. The frame concepts stand as library entries; signal-level rebuild replaced the v1 rules for the two recovered.
Why Experts Miss What Beginners Catch
- "Domain expertise enables evaluation" too broad. Production expertise and memorized statistics enable evaluation. Domain familiarity alone does not.
- "Generate first always helps" untested. Anchoring risk is real.
- "FRAME improves analytical depth" killed. Zero effect on reasoning tasks. Partial effect on reformulation tasks only.
- "Evaluation degrades over time with AI delegation" not supported. The construction trace only covers produced content. There's nothing to degrade: evaluation of non-produced statistics was never strong.
Adding Information Often Doesn't Help
- "More context always helps" killed. Information has near-zero effect at higher baseline. Information actively suppresses contrarian thinking on one model.
- "Information produces richer analysis" partially killed. Baseline-dependent, not robust.
- Large information effect from original did not replicate. Baseline-length-dependent.
Stop Polishing, Start Switching
- "Always switch modes" too prescriptive. Sometimes iteration within a mode is what you need.
- "Three switches are exhaustive" too strong. Other mode switches exist.
- "Adversarial mode switch confound resolved" overclaimed. The non-adversarial direction is observed but has not been run as a controlled comparison, so the confound still stands.
Four Layers Produce Every AI Output
- "The model doesn't matter" too strong. Model determines format preferences and behavioral intensity independently of system configuration.
- "System effects are small" too dismissive. A single system configuration change (12 skills loaded vs removed) changed model behavior from functional to non-functional.
Same Technique, Opposite Results
- "Structure always helps" killed. Task-type dependent.
- "The harm is about constraint density" partially killed. It's about concentration vs range, not about how many constraints.
- "Any structure narrows exploration" killed in follow-up testing: organizational structure expanded exploratory output several-fold while evaluative structure compressed it. The compression claim is scoped to evaluative-type structure, and an explicit task intent outweighs the structure signal.
- Quality magnitude claims are LLM-calibrated. Human evaluation shows no holistic agreement with LLM scores. Effect sizes measure programmatic specificity markers, not quality as a domain expert would judge it. In a blind test (one domain expert, 5 pairs), the expert couldn't distinguish specific from generic outputs on quality. Both conditions produced the same analytical substance. Specificity changes output form (more verifiable references), not substance. [The honest effect size is roughly 40 percent smaller than originally claimed](/catching-your-own-overclaim), once confounded length and quality-demand variables were removed.
For Behavior, the Model Is Rarely the Variable
- "Universal ratio": the specific ratio is experiment-specific. The exact ratio is small-sample. Gemini responds to the mirror persona more strongly than Grok. The ordering (prompt > model) is robust.
- "Model doesn't matter at all": model determines format preferences independently of prompt.
- Vocabulary bans looked like quality control. The mechanism was output compression. Shorter output has higher density by default.
The Most-Cited Finding Was Wrong
- The original inflated effect size (specificity + length stacked; honest range roughly 40% smaller)
- "Strongest effect in 90+ experiments" (large, but not as large as claimed)
- Clean separation at density level: quality demands show a large density effect vs near-zero raw effect (density partially conflates specificity with shorter output length)
Most AI Numbers Are Unverifiable
- "100% fabrication is universal" killed. One generator shows 77% with topic-dependent retrieval
- "PROTOCOL fixes fabrication" killed. Highest fabrication rate of all conditions
- "Source grounding fixes everything above data" partially killed. Vocabulary and causal framing stay at baseline for same-topic regeneration. But on reasoning tasks with ground truth, source-present output finds the correct answer roughly twice as often as source-absent. Source grounding improves correctness on reasoning tasks, not just reformulation.