<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Clarethium</title>
    <link>https://clarethium.com/blog</link>
    <description>Clarethium research: experiments on AI and human judgment. What survived. What broke.</description>
    <language>en</language>
    <lastBuildDate>Thu, 20 Aug 2026 14:14:18 GMT</lastBuildDate>
    <atom:link href="https://clarethium.com/blog/feed.xml" rel="self" type="application/rss+xml"/>
    <item>
      <title>Why &apos;Don&apos;t Be Generic&apos; Doesn&apos;t Work</title>
      <link>https://clarethium.com/blog/dont-be-generic</link>
      <guid isPermaLink="true">https://clarethium.com/blog/dont-be-generic</guid>
      <description>Telling a model &apos;don&apos;t be generic&apos; does nothing on its own; giving it specific anchors does. The gain is verifiability, not quality.</description>
      <content:encoded><![CDATA[<p>Asked for a competitive analysis. Gave it everything. Market position, three-year data, the specific situation. "Be insightful. Don't be generic."</p>
<p>The output was structured, fluent, professional. And interchangeable with what it would have produced for any company in any market. Every recommendation could have been copy-pasted into a competitor's strategy doc without changing a word.</p>
<p>Twenty controlled runs confirmed what was visible. Specific versus vague, crossed with positive versus negative framing. Four combinations. Pure negation ("don't be generic," "avoid cliches," "don't use buzzwords") was indistinguishable from giving no instruction at all.</p>
<p>That removes a label without providing a destination. The model left [the default](/the-default) and wandered to the adjacent region. Same neighborhood. Different house number. Negation names what to avoid. The default already routes around that. Only an anchor forces a different path.</p>
<p>Then the other cell. Same negation, but paired with specifics: "Don't include recommendations that could apply to any B2B SaaS company. Every recommendation must reference Northvane's specific assets, 5 years of shipping logistics data, 12 engineers, Pacific Northwest enterprise incumbents."</p>
<p>Strong effect. The direction replicated across generators. This is the content-specificity lever. Stacking more explicit constraints is a different lever, from a different experiment. Keep the two separate.</p>
<p>The specificity isn't measurably adding analytical quality. It's adding verifiability.</p>
<p>"Don't be generic" blocks one path. The model takes the next most likely path, which is a variation on generic. "Reference these specific assets" creates an anchor the output has to pass through. The result physically cannot be the same for a different company. The constraint tests itself.</p>
<p>In blind testing, a domain expert couldn't distinguish specific from generic outputs on quality. Picked specific 3 out of 5 times. Chance level. The specificity instruction changes what the output looks like: more data references, more grounded claims. It does not, measurably, change what the output says.</p>
<p>The demonstrated value is verifiability. The specific output can be checked. Every claim traces to something nameable. The generic output makes the same points but you can't verify them. When you need to trust the analysis, specificity makes the output auditable. When you're using it as a starting point for your own thinking, it doesn't matter.</p>
<h3>What survived testing</h3><ul><li>Negation alone has no effect. Specificity is real. The direction replicates across generators. Specific-plus-negative scored highest in the framing study; that ranking is not itself cross-validated. Clean magnitude: g=1.34 on raw marker count and 1.62 at density on xAI. Nearly identical at density on Gemini Flash (g=1.64), not on raw (g=0.65). The earlier Claude-versus-others gap came from a length-confounded comparison and is not clean. Quality demands do little alone. They add on top of specificity. Together the effect is larger than the sum. Negation removes a label without a destination. Specificity provides the anchor.</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>&quot;Negation hurts&quot; overclaimed. Negation is null, not negative. &quot;More specific = better&quot; linearly is unverified: the effect was present versus absent, not a density gradient, so extreme specificity is untested. &quot;At density is the pure confound-free measure&quot; too strong: density (markers per 1,000 words) removes the prompt-length confound but inflates via brevity, because shorter outputs score higher. It strips one confound, not all. The cross-generator match is on density. The replication is directional, not a magnitude match.</li></ul>
<h3>Honest limits</h3><ul><li>Single operator. Transfer to other operators untested.</li><li>The negation result and the specificity-magnitude result come from two different experiments: a specificity-by-framing design shows negation alone is null, while the receipt-backed specificity-by-quality-demands 2x2 establishes the magnitude and the cross-generator replication. Only the second is in the bound receipt; the negation-by-framing result is not reader-auditable there.</li><li>The constraint-count lever&apos;s model differences are not all clean reads: Claude&apos;s 0.00 is a rubric ceiling (25/25, zero variance in both conditions), not an absence; the GPT result is self-scored (the model graded its own outputs), so treat it as an upper bound.</li><li>Effect sizes measure programmatic specificity markers (company mentions, scenario numbers, market terms). These counts are objective. Domain expert validation (5 blind pairs, evaluator&apos;s own domain): indistinguishable from chance. Expert rated based on style (rhythm, naturalness), not on marker density. At that sample size no quality difference was detectable, which is not proof there is none. Specificity demonstrably changes output FORM (more verifiable references); a SUBSTANCE difference was undetectable here, neither shown nor ruled out. What is demonstrated is verifiability.</li></ul>]]></content:encoded>
      <pubDate>Sun, 12 Jul 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Three Questions Before You Prompt AI</title>
      <link>https://clarethium.com/blog/before-you-type</link>
      <guid isPermaLink="true">https://clarethium.com/blog/before-you-type</guid>
      <description>Three questions before you type structure most of the prompt. Specificity is the strongest single lever (Hedges g=1.34); &apos;be exceptional&apos; alone does almost nothing.</description>
      <content:encoded><![CDATA[<p>Three questions. Answer them and most of the prompt is structured. The third tends to do the most work. On anything you have not built before it is the one with no answer to look up. That is where this gets interesting.</p>
<p>What do I need? Not "help me with X." What's the deliverable? A decision, a draft, an analysis, a list. One sentence. If that sentence won't come yet, the prompt isn't ready: the thinking comes before the typing.</p>
<p>What can't the model know? The model knows what is public and general. It doesn't know YOUR situation. Your audience, your constraints, your history, your deadline. Think about what a sharp person would ask before starting work on your task. Those questions, answered, are your context.</p>
<p>How will I know it worked? The one that changes the output the most. The one a prompt most often leaves out. Not "make it good." What does DONE look like for this specific task? "Every recommendation references something specific from our Q4 data." "I can hand this to the exec team without editing." In practice, that one sentence moves the output more than anything else in the prompt.</p>
<p>A fourth question sharpens it further: what would failure look like? "Failure is if this reads like a generic project update that could be about any project." That gives the model something concrete to avoid. Avoiding is more specific than aiming.</p>
<p>The difference in practice:</p>
<p>Before: "Help me write a project update for my team."</p>
<p>After: "Write a project update for my engineering team (12 people) about the authentication migration. We're 3 weeks in, 2 weeks behind schedule. The delay is from an unexpected dependency on the billing service API. Team morale is fine but stakeholders are nervous. Done means the team knows exactly what's behind, why, and what we're doing about it. No sugarcoating."</p>
<p>Same task. 30 seconds longer to write. The output is different because the model has constraints instead of infinite interpretation space.</p>
<p>Underneath all three questions is one mechanism. Together they make the model's generic default path more expensive than doing your specific task. That mechanism is [Why AI Defaults to Generic](/the-default). This is that mechanism turned into questions you can ask before typing.</p>
<p>The project update was the easy case. You have written those before, so you already know what done looks like. The hard case is the task you have never done. The most important question, what does done look like, is also the one you cannot answer yet. You do not understand the task well enough to say. That is not the framework failing. It is the framework pointing at where the real work is.</p>
<p>Say you have to write the design doc for a system nobody on your team has built before. Ask what done looks like and the honest answer is that you do not know yet, because you do not know the tradeoffs until you are in them. There is no prior version to copy. So you write a bad one, rough and wrong, and the moment it exists you have something to push against: this assumes the current data model, this hand-waves the failure mode, this is solving for scale you do not have. Each objection is a criterion you could not have listed on a blank page. The draft did not produce the design. It produced the definition of done, which was the part you were missing.</p>
<p>Two moves get you unstuck. The cheap one is the failure question: you can usually name what would be wrong before you can state what would be right, which names the target from the other side. The stronger one is what that draft did: make your own rough attempt first, before prompting. It does not have to be good. Producing it surfaces the criteria you could not state cold. It also gives you a model of the answer to judge the model's version against. Reacting to the model's draft instead skips the step, and you pay for it later: you approve something fluent, ship it, and find it solved the convenient version of the problem, not the real one. That is [The Construction Trace](/construction-trace).</p>
<p>So the recipe is not answer three questions, then type. When you can answer them, answer them. When you cannot, the work that comes before the prompt is producing the definition of done yourself.</p>
<h3>What survived testing</h3><ul><li>In the 2x2 I ran, specificity was the strongest lever for verifiable form, not perceived quality (Hedges g=1.34 on raw marker count, 95% CI [0.74, 2.32]; 1.62 at density).</li><li>Quality demands did little on their own in that test (near zero, scoring just below the bare prompt) but combined synergistically with specificity: the both-present condition scored higher than the two separate effects predicted.</li><li>Adding constraints beat adding context in the case I tested: on an already-thorough prompt, more information barely moved the output, while a few constraint sentences changed it ([Context Serves Search](/context-serves-search)).</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>&quot;Three questions are always sufficient&quot; untested. Some tasks need domain-specific framing.</li><li>&quot;Structure outperforms descriptive instruction&quot; as a clean rule: how much structure helps is model-specific (large on one generator, uninformative on another where the output already sat at the scoring ceiling, reversed on a third), so it is a heuristic, not a tested result.</li></ul>
<h3>Honest limits</h3><ul><li>The three-question framework is craft knowledge, not experimentally tested as a unit. What is measured here is the specificity effect and the constraints-over-context comparison; the framing around them is operator craft.</li></ul>]]></content:encoded>
      <pubDate>Sun, 05 Jul 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Cheap Half of the Loop</title>
      <link>https://clarethium.com/blog/cheap-half-of-the-loop</link>
      <guid isPermaLink="true">https://clarethium.com/blog/cheap-half-of-the-loop</guid>
      <description>Coding pulled ahead because its feedback is cheap to check. That is one half of the loop. The same training trains out the other half, and the part that chooses the frame is the last thing to get cheap.</description>
      <content:encoded><![CDATA[<p>Coding agents pulled away from everything else in 2025. The easy read is that code is simpler for a model. The more useful read is that the field found a reward it could check for free, and built an industry on it.</p>
<p>## The half that grades itself</p>
<p>The training move that defined last year was reinforcement learning on tasks a machine can check. It pulled in compute that had been earmarked for pretraining. Code and math went first, because the feedback is cheap. A compiler returns yes or no. A test passes or it does not. There is a clean principle underneath this: the easier a task is to verify, the easier it is to train a model to do it. Capability spikes wherever an automatic grader exists. It stays jagged everywhere else.</p>
<p>So the boom is not really about code. It is about checkability. Code just happens to be the most checkable thing we do, which is why a signal you can grade for free is the one that pulled in the year's compute. Labs are now racing to manufacture more of these checkable environments. The effort is reported to run into the billion-dollar range. The scarce input is not compute or data. It is a task with a clean answer.</p>
<p>One caution, because the causal story runs ahead of the evidence. The verifiable reward looks like it sharpens what the base model already had more than it builds something new. Push the sampling far enough and the untrained model often reaches the same answers. On some models the measured gain even survives when the reward is handed out at random. So the honest claim is narrower than "checking makes models smart." It is that checkable feedback is the cheapest place to push, so the field pushed there. Call it the cheap half.</p>
<p>## The other half, and what training does to it</p>
<p>Everything so far is one move in a larger loop. Call the loop ground and diverge. Grounding is checkable contact with something that can say yes or no, and the cheap half is exactly that move, mechanized and poured full of compute. Divergence is the other move: opening the space of candidate answers wide enough that something non-obvious is in it. I wrote about the loop in [Past the Obvious](/past-the-obvious): get past the first answer, then let reality correct you. This is not that piece. This one is about what the year's training economics did to each move.</p>
<p>They moved in opposite directions. The cheap half got industrialized. The scarce half got trained out by the same process.</p>
<p>The methods that produced the gains narrow the model's range as a side effect. Preference tuning makes outputs more uniform. In reasoning training the model's sampled answers converge toward one path. Left to a plain prompt, it commits early. So this is not two pillars rising together. It is one half getting louder while the other gets quietly suppressed by the method paying for the first.</p>
<p>## Two kinds of divergence, and why one lasts</p>
<p>It matters to split divergence in two, because they are not the same problem and they are not on the same clock.</p>
<p>The first kind is the one you can score by the result: generate, deliberate, sample more, keep what works. It does not matter whether the model gets there by trying many options or by thinking longer down one, because the answer at the end is checkable either way. So this kind is getting solved. It is most of what more compute at inference time buys you. The labs are closing it fast. If your edge is out-sampling the model, that edge is renting time.</p>
<p>The second kind is frame origination: not more answers inside a frame, but seeing the frame is wrong and holding a different one. [The Frame Trap](/frame-trap) covered the mechanism, that a model works within whatever frame you set and not on the frame itself. What that piece did not say is why the limit lasts.</p>
<p>Two routes can teach a model a skill: imitation, and a reward signal. Imitation has already had its turn here. A model has read more reframing than any person ever will, every pivot and reversal in the written record. It still waits for you to supply the new frame. So the route that is left is a reward, and a reward needs a grader. A frame's correctness is the hardest thing there is to grade. You only learn whether a frame was right long afterward, from what it led to, with nothing clean to score it against in the moment. By the same logic that solved the most checkable work first, the least checkable work gets solved last. Framing is not off the roadmap. It is at the end of it.</p>
<p>## The handle, with an expiry date</p>
<p>This is why a build-your-own-divergence loop is worth it now, and how you will know when it stops being. Today an external loop that forces real divergence beats the default, because the default converges too early. Part of that loop is the checkable kind of divergence, and that part is turning into a feature you get for free. The other part is choosing the frame the model then works inside, and that part is last in line precisely because no one can yet grade it.</p>
<p>So the edge is real and it is dated. Out-sampling is rented and expiring soon. Framing is rented too, but the lease runs longest, because it is the hardest thing to score.</p>
<p>## Close</p>
<p>This is not a theory of AGI. The two-half framing is a working read rather than a settled result. What it is, concretely, is two bottlenecks moving in opposite directions on one clock. The clock is checkability. The more checkable a task, the sooner it gets cheap. Grounding got cheap first. The scoreable kind of divergence is getting cheap now. Choosing the frame gets cheap last, because there is nothing to grade it against until the outcome is already in. So measure your edge against the next model, not this one. Whatever a model twice as good would simply do for you is rented. The lease is short. What still needs you when the model doubles is the framing. That is the edge worth building on: being the source of the frames a fast, well-grounded model is, for a while yet, the last to originate for itself.</p>
<h3>What the research shows</h3><ul><li>Reinforcement learning on verifiable rewards (RLVR) was the defining new training stage of 2025; it &quot;gobbled up the compute that was originally intended for pretraining&quot; (Karpathy, 2025 year in review). The term was coined in Tulu 3 (Lambert et al., arXiv 2411.15124, Nov 2024); the math and code form was popularized by DeepSeek-R1 (arXiv 2501.12948, Jan 2025), where a boxed answer or a passing test is the whole reward.</li><li>The harder a task is to verify, the less readily it trains: &quot;the ease of training AI to solve a task is proportional to how verifiable the task is&quot; (Jason Wei, &quot;Asymmetry of verification and verifier&apos;s law,&quot; Jul 2025). The jagged capability profile is his prediction.</li><li>Labs are building hand-designed RL environments at reported billion-dollar scale (TechCrunch, Sep 2025; Epoch AI; SemiAnalysis, Jan 2026).</li><li>Preference training reduces output diversity (Kirk et al., arXiv 2310.06452, ICLR 2024), and policy entropy collapses during reasoning training, with samples converging toward near-identical solutions (Cui et al., arXiv 2505.22617, May 2025).</li><li>The scoreable kind of divergence is the one test-time compute is closing; the canonical move is to sample many candidates and select among them (self-consistency samples a diverse set of reasoning paths and keeps the most consistent answer; Wang et al., arXiv 2203.11171, 2022). On an open-ended exploration task, most models explored worse than people and only the reasoning model o1 beat the human baseline, finding 177 elements to GPT-4o&apos;s 35; the authors attribute the gap to o1&apos;s slower, more deliberate inference (Pan, Xie, Wilson, arXiv 2501.18009, Jan 2025).</li></ul>
<h3>What it doesn&apos;t show</h3><ul><li>That verifiability is the cause of coding competence. RL appears to sharpen the base model more than create new reasoning (Yue et al., arXiv 2504.13837; contested by Wen et al., arXiv 2506.14245), and random rewards reproduce much of the gain on some model families (Shao et al., arXiv 2506.10947, Qwen-specific). The body claims only the weaker version.</li><li>That solving both halves yields AGI. No source frames progress as these two coupled pillars. That pairing is a working read built on [Past the Obvious](/past-the-obvious) and [The Frame Trap](/frame-trap), not a result.</li></ul>
<h3>Honest limits</h3><ul><li>&quot;Divergence&quot; here is one word for three established but separate research threads: exploration in reinforcement learning, generation diversity, and test-time search. The phenomena are real and named; the grouping is mine.</li><li>The two-kinds split is an argument, not a measurement. &quot;Framing gets cheap last&quot; is verifier&apos;s law applied to frame-correctness, a prediction, not an observed result. Long-horizon outcome-based training is an active attempt to put a grader on exactly this kind of slow signal, and if it works the lease runs shorter than the piece implies.</li><li>The load-bearing step, that imitation has already had its turn, rests on a firm empirical generalization: that models execute frame expansion on request but do not originate the load-bearing reframe on their own (The Frame Trap, small n). It is the weakest link. A model shown to originate a load-bearing reframe unprompted would take the floor out from under the argument.</li></ul>]]></content:encoded>
      <pubDate>Sat, 27 Jun 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Why AI Defaults to Generic</title>
      <link>https://clarethium.com/blog/the-default</link>
      <guid isPermaLink="true">https://clarethium.com/blog/the-default</guid>
      <description>Every prompt technique is one move: make the default path expensive enough that the model leaves it. Specificity is the largest measured version.</description>
      <content:encoded><![CDATA[<p>Ask a model "how should we price the new product?" and you get something like this: pricing depends on several factors. Consider your costs, your competitors, your target market, your perceived value. Value-based pricing is often a strong approach. You may want to test a few price points and watch demand. Safe. Balanced. Fluent. Complete-sounding. It would fit any product in any market, which is the tell.</p>
<p>Now hand it your two competitors' actual price points, your margin floor, and one instruction: "recommend a single number, then name the assumption that would make it wrong." The output commits. It picks a number, defends it, and exposes where it breaks. That is the shape of the shift. It is not a measured pair.</p>
<p>The first output is the default. The default is not a failure mode. It is what the model produces every time the context gives it no reason to go anywhere else. Knowing why it sits there tells you what every prompt technique you have ever been handed is actually fighting. Specificity, source material, structure, examples, personas: they are all one move. They make the default expensive enough that the model leaves it.</p>
<p>So why does the model sit there? Two stages of training put it there. What follows is an interpretive frame, grounded in how LLM training works, not experimentally decomposed in this corpus. Pretraining makes the model the average of everything it read, and an average reads as generic. The most probable next token is the most common one: "in conclusion" after analysis, "however" after a claim, "it depends" after a hard question. The default is the mode of the distribution. That is why AI output often sounds like a well-written essay by nobody in particular. It is one. Then post-training tilts that average toward what human raters rewarded: balanced over committed, hedged over strong, comprehensive over focused. Genericness comes from the first stage. Safe hedging comes from the second.</p>
<p>That trained default shows up three ways, none of them a separate cause. It plays for low regret: a balanced answer is never badly wrong, so the model avoids the specific claim that could be. It closes the face of the response: "here are several perspectives" outranks "I don't know," because definitive-sounding answers scored higher than honest incompleteness. And it aims for the broadest audience: output that reads reasonable to everyone, rather than exactly right for one reader and strange to the rest. The result looks good, sounds professional, and says nothing a thousand other prompts wouldn't produce.</p>
<p>This is the thing the rest of the work refers back to. Every technique that improves AI output is a way of making the default path costlier than some alternative. Specificity narrows away from the generic average. [Source material](/source-conditioning) replaces generated content with real data. Structure forces non-default organization. The two largest moves that hold across generators are the first two. Specificity against the default. Source material against fabrication. Same principle underneath both. Give the model a reason to leave the default. It is where the model goes when you don't.</p>
<h3>What survived testing</h3><ul><li>Specificity defeats defaults. Largest measured effect with confounds controlled. Direction replicates across generators. Clean magnitude established at density on xAI, nearly identical on Gemini Flash, not Claude-dominant. The earlier Claude-versus-others gap came from a length-confounded comparison. Quality demands do little alone. They combine with specificity. Together the effect is larger than the sum.</li><li>Source material defeats fabrication defaults (source-attribution rises from 45% to 91%, a 46-point gap, so numbers not traceable to a source fall from roughly half to single digits)</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>&quot;The default is always bad&quot; too strong. For many tasks, the default is adequate.</li><li>Equal weighting across the two mechanisms and three faces is untested. The decomposition is analytical, not quantified.</li></ul>
<h3>Honest limits</h3><ul><li>The two-mechanism / three-face account is an interpretive framework, not an experimentally isolated decomposition. That pretraining yields the generic average and post-training tilts it toward safe/hedged is established in the literature, not tested in this corpus; only the specificity and source-material effects are receipt-backed here.</li><li>The model&apos;s default shifts with updates. What was default in March 2026 may not be default later.</li></ul>]]></content:encoded>
      <pubDate>Mon, 22 Jun 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Past the Obvious</title>
      <link>https://clarethium.com/blog/past-the-obvious</link>
      <guid isPermaLink="true">https://clarethium.com/blog/past-the-obvious</guid>
      <description>The original ideas arrive late, after the obvious ones clear. You cannot will yourself there, but you can force the move off the default path, borrow from a far domain, then let reality correct you.</description>
      <content:encoded><![CDATA[<p>Most of what we are trained to do has one right answer. Code does exactly what it says. School grades the single correct response. None of it teaches the other mode, the one where you reach the option nobody handed you. That mode is where almost everything came from.</p>
<p>I use this, and I am still testing it on real problems. The point is not a clever trick. It is that practicing it changes what you treat as possible and how you reason, because you are training your brain to run a pattern it does not run by default. It is one of many. This one is strange because it is so obvious and so rarely used well.</p>
<p>The default mind does not go to the new place on its own. It answers with the first thing that fits and stops. Research on idea generation has shown for decades that original ideas arrive later, after the obvious ones are spent, because the near, easy associations clear before the remote ones surface. The good part is past where you would normally stop. So one route is persistence: stay on it past the point you would usually quit.</p>
<p>You cannot will yourself original. But you can force the move that gets you off the default path. A few that work:</p>
<p>Shift the domain. Ask how a completely unrelated field solves the shape of your problem. How does a forest handle competition for light. How does an immune system decide what to let in. We got airplanes from birds and the bullet train's nose from a kingfisher's beak. You are not copying the surface, you are taking the mechanism.</p>
<p>Invert it. Instead of how to make it work, ask how to guarantee it fails. List every way. Then flip each one.</p>
<p>Starve it. Ban the one resource you are sure you need. No money, no time, no team. What is left is usually the path the crutch was hiding.</p>
<p>Pull something random. Force a bridge between an unrelated word and your problem. Most bridges are nonsense. Occasionally one cracks the frame.</p>
<p>There are dozens more worth searching: biomimicry, lateral thinking, morphological analysis, the random-word technique. None are magic. Most have never been shown to beat simply trying hard. What they do is concrete, and it is a different route from persistence. Instead of pushing further down the same path, they move you sideways into a region you would not have searched at all. Two ways off the default, not one. The aim of either is not strangeness. It is range.</p>
<p>Two habits matter whichever technique you use. Separate generating from judging. The same bias that rejects unfamiliar ideas turns on your own the moment you grade them. Get it all down first and prune later. And when you stall, step away. A break from the problem beats grinding at it.</p>
<p>The same logic applies to a model, with a catch worth knowing. In one respect a brain and a model are the same kind of machine: both complete patterns, and both collapse toward their most probable path. So a plain question gets a plain answer from both of you. But telling a model to "be creative" barely moves it. It reshuffles the same familiar space and settles back into the typical, because the instruction gives it nothing new to draw from. What moves it is new material. Load the prompt with a different domain, its real mechanics, a frame far from the obvious one, until the model has a different probability space to draw from, not the one a bare question hands it. Weak context, default output. Rich and distant context, and it can reach what the plain prompt never could. The work is in the loading, not the asking.</p>
<p>Then comes the part that decides whether any of it was real. Grounding. And grounding is rarely just "is this feasible." More often it is "do I actually understand how this system works." Your logic builds a clean model, and reality is messier. Governments are more cluttered than your reasoning predicts. Most systems run on years of workaround no clean model contains. So you test against the real system, not your picture of it. And you take it outside your own head. We are measurably worse than we think at recognizing a good unfamiliar idea. Reality is not biased the way people are; it only asks whether the thing works.</p>
<p>That is the loop. Force past the obvious without grading as you go, borrow from far away, then meet reality and let it correct you. Try one this week, on an algorithm you want faster or a question with no clean answer. The point is not to be right on the first move. It is to go where your default mind does not, and to notice that it can.</p>
<h3>What the research shows</h3><ul><li>Original ideas tend to arrive later in a run, after the obvious ones, the serial order effect (Christensen, Guilford and Wilson 1957; Beaty and Silvia 2012).</li><li>A break from the problem, incubation, raises performance on creative problem solving (Sio and Ormerod 2009, meta-analysis).</li><li>Analogy transfers relational structure, not surface features (Gentner 1983).</li><li>We are biased against unfamiliar ideas and worse at recognizing the good ones under uncertainty (Mueller, Melwani and Goncalo 2012).</li><li>A model defaults to the typical (Holtzman et al. 2019); loading richer, more distant context is what changes where it draws from, the principle behind retrieval-augmented prompting.</li></ul>
<h3>What it doesn&apos;t show</h3><ul><li>That any named technique reliably beats simply trying hard. Most are widely taught and lightly evidenced.</li><li>That wilder is better. The evidence favors a conventional core with a few atypical additions, not maximal strangeness (Uzzi et al. 2013).</li></ul>
<h3>Honest limits</h3><ul><li>These are practices I use, not a controlled trial I ran. Treat them as scaffolds, not guarantees.</li><li>The brain-and-model parallel holds at the level of probability, not biology.</li></ul>]]></content:encoded>
      <pubDate>Tue, 09 Jun 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Three AIs, No Source, the Same Answer</title>
      <link>https://clarethium.com/blog/source-is-the-substrate</link>
      <guid isPermaLink="true">https://clarethium.com/blog/source-is-the-substrate</guid>
      <description>Same model, same prompt. The source you paste, not the prompt you write, decides whether the numbers are real.</description>
      <content:encoded><![CDATA[<p>Asked an AI for an analysis and gave it no source to work from. It cited a real study, with a real number. So I tried two more, from two other companies. Same study. Same number. None of them had been handed it.</p>
<p>The topic was remote-work productivity. Grok, Gemini, and GPT-5-mini all reached for Bloom's 2015 Ctrip trial and its 13 percent figure. Five of six runs landed on the number. They reconstructed it from training, not from anything I gave them.</p>
<p>That is the tell. The model is not looking a fact up and getting it right. It is sampling from a distribution, and the distribution has peaks sharp enough that different companies' models hit the same one. This time the peak was real. Bloom's study exists, and 13 percent is roughly its finding.</p>
<p>Sit with how that number read. Specific. Sourced. The kind of thing you would drop into a memo without a second thought. It was solid, and you would have been right to use it.</p>
<p>Now the part that should bother you. The same machinery produces fake numbers that read exactly as solid. Same confidence. Same specificity. Same clean citation. Nothing you can see in the finished text separates the real peak from the invented one, because the model did not do anything different to produce them. It sampled a peak both times. From inside the output, the true number and the fabricated one are the same experience. You have been telling them apart by feel, and the feel is identical.</p>
<p>So how often can you even check the number? Same three models, same prompts, one change: paste a real source into the chat first, then count how many of the numbers each model produces actually appear in it.</p>
<p>``<code><br/>                    no source   source pasted<br/>grok-4-1-fast          34%          92%<br/>gemini-3-flash         41%          86%<br/>gpt-5-mini             12%          95%<br/></code>``</p>
<p>With no source in the window, between 59 and 88 percent of the numbers, depending on the model, are ones you cannot check against anything you gave it. Some are real anyway. Bloom's 13 percent was. Nothing in the text tells you which. There is one apparent tell. It is weaker than it looks. With no source the models produced far fewer numbers. Grok dropped from 164 to 106. GPT-5-mini dropped from 73 to 8. Both leaned qualitative. The prompt had told them to use qualitative language whenever they could not source a figure. That is mostly instruction-following. It is not the model sensing it was empty.</p>
<p>Why is there nothing to read? Because a model with no tools and no web access has no fact database to consult during generation. It runs a forward pass that predicts the next token from a distribution shaped by training, conditioned on whatever is in the context window. There is no step where it looks Bloom up, succeeds, then looks the fake number up and fails. Both numbers come out the same way: a sample from that distribution. With no source, the likeliest number for "remote work productivity" is the most-cited figure in the training corpus. Paste the report, and the likeliest number becomes the one in the report, because the report is now the most relevant evidence for what comes next. The distribution updates on the input. The input is the substrate. This is why retrieval-augmented generation works. It loads the source into context before generating. RAG is an industrial version of a move you can make by hand. A chat tool with web search can sidestep this by fetching a source before it answers.</p>
<p>Which means "the model knows X" is a claim about training data. "The model can analyze X" is a claim about what is in the context window. They feel like the same sentence. The gap between them is where every fabricated number you have ever trusted came from.</p>
<p>The only way to make the peak trustworthy is to make it yours: put the real source in the window so the number the model samples is the one you handed it. The how-to is in [Source Conditioning](/source-conditioning): paste the data, add a line prohibiting unsourced numbers, match the output against the source afterward. This piece is why that beats every prompt trick. A better role, chain-of-thought, "be careful and verify": each one asks the model to sample differently from the same distribution. Better sampling from a fictional distribution is still fiction. Changing the source changes which distribution it draws from. Nothing else does.</p>
<p>One honest boundary. This is for work where a source exists: summarize this, analyze these reports, what did the feedback say. For pure ideation or reasoning from scratch there is nothing to paste, and none of this helps. But that is a smaller slice of real work than it feels like. A great deal of AI use is reformulation of something you already have, and most of the time the something never makes it into the window.</p>
<p>So do this once. Take the last AI output you acted on. Pick one number in it. Find the source that number came from. If you can, good. If you cannot, you were not reading a fact. You were trusting a peak, because it was confident and specific, and confident and specific is exactly what a fabricated number looks like too.</p>
<h3>What survived testing</h3><ul><li>Source presence moved numerical match rates from 12 to 41 percent up to 86 to 95 percent across three model families. Average gap +62pp, under a prompt that already asked for sourcing, so the portable finding is the delta, not the absolute rates. Direction universal; magnitude largest on the cleanest baseline (GPT-5-mini, +82pp). Match rates are programmatic; no human judgment.</li><li>Cross-model convergence: with no source, three families reconstruct the same real Bloom 2015 Ctrip study and its 13 percent figure from training alone. The number-convergence is spontaneous; the source-naming was prompted (the task asked them to cite sources). Retrieval from weights, not from anything provided.</li><li>The architectural claim (no factual database in a tool-free forward pass) is consistent with the convergence finding and with why RAG works in production.</li></ul>
<h3>What did not survive</h3><ul><li>&quot;Three models produced 13 percent without naming the study&quot; cut. The data shows the opposite: all six runs named the study, five stated 13 percent.</li><li>&quot;Fabricated 13 percent anchor&quot; cut. The number is real. The fabrication lives in the aggregate match rate, not in this datapoint.</li><li>&quot;Source grounding solves AI fidelity&quot; cut. It sets the floor at 86 to 95 percent; the residual needs prevention tools, and it reaches the numerical layer only.</li><li>&quot;The models hedged because they sensed they had nothing&quot; softened. The prompt instructed qualitative language when a figure could not be sourced, so the number-density drop is mostly instruction-following, not a spontaneous signal.</li></ul>
<h3>Honest limits</h3><ul><li>Reformulation tasks with source material in context. Not validated for novel reasoning, creative generation, or strategy from scratch.</li><li>Numerical fidelity specifically. Entity, claim, and reasoning fidelity need separate measurement.</li><li>Three model families, May 2026 (grok-4-1-fast, gemini-3-flash-preview, gpt-5-mini), N=2 versions per cell. Adequate for direction, not for tight effect-size intervals.</li><li>Tool-free, single-prompt generation. A model with web search or RAG can fetch a source and sidestep the whole effect; this is the bare model with nothing but its weights.</li></ul>]]></content:encoded>
      <pubDate>Mon, 25 May 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>You Can Only Evaluate What You Could Produce</title>
      <link>https://clarethium.com/blog/ownership-test</link>
      <guid isPermaLink="true">https://clarethium.com/blog/ownership-test</guid>
      <description>You defend AI-shaped conclusions you cannot rebuild. The ten-minute test reveals which parts of your work are yours. The discipline that compounds: choose what to own, delegate what AI can verify.</description>
      <content:encoded><![CDATA[<p>Twice in two days last week, I caught myself defending a position I could not rebuild. The framing came from AI. The path was never built.</p>
<p>This is borrowed certainty. The kind of conclusion you defend not because you worked it through, but because by the time anyone asks, it sits in your head as your position. It is the difference between builders who grow with AI and builders who plateau. The plateau is not about how much AI we use. It is about how we use our own cognition.</p>
<p>Take something AI generated for you this week. An analysis. A recommendation. A strategy document. A product spec. Something you used for real work, not a demo. Explain it. Out loud, walking, to yourself. To the AI in a fresh session, without leaning on the document. To a colleague if you have one. Why this approach and not another. What the key trade-offs were. What would change the conclusion.</p>
<p>Where you produce clearly, the thinking is yours. Where you reach for the original, or say "the analysis showed that..." or "AI recommended...", you are carrying borrowed understanding. The words were delivered. The thinking wasn't.</p>
<p>The [construction trace](/construction-trace) is why. When you generate, you build a mental model as you go. You feel the hard parts. You notice the gaps. You know what good looks like because you struggled to produce it. When AI generates, you skip the struggle. The output arrives fully formed. The deep check that catches bad reasoning under pressure requires the model only generation builds.</p>
<p>The wrong move after seeing this is to try to own everything yourself. Cognition is finite. Choose.</p>
<p>Own what compounds. Taste. Judgment. Problem framing. Direction. The call about what matters in this specific situation. None of it is verifiable from outside, which is exactly why it needs you. Every session you spend cognition here trains the skill. Growth is on a curve.</p>
<p>Delegate what does not compound. Syntax. Grammar. Mechanical execution. Well-defined transformations. Anything cheap to verify against ground truth. AI does these, often better, definitely faster. Cognition spent here trains nothing that will not be cheap to verify next year.</p>
<p>The pattern most of us fall into is the inverse. We let AI decide what matters and spend our cognition checking the punctuation. We delegate the direction and keep the mechanical. The thinking goes generic. The decisions go average. Ship rate goes up. Growth rate does not. The why is usually invisible from inside the pattern.</p>
<p>Pick three AI outputs from this week. Run the explanation pass. For every claim you stall on, decide: is this a part I want my mental model on, or a part I am happy to carry as borrowed? When you use a borrowed conclusion next, mark it borrowed in real time, even just to yourself. The discrimination is the practice. With reps, you stop trying to own everything and stop letting AI decide everything. It becomes instinct.</p>
<p>Within two weeks, for me, the work shifted. Strategy docs you can defend without reaching for the source. Product decisions where the framing is yours and the execution is delegated and labeled. Build cycles where you put cognition on the problem and let AI handle the well-defined execution. Community decisions where the direction is yours and the wordcraft is delegated. Meetings where you say "I am carrying this from AI; here is my actual reasoning on what I worked through, here is the part I have not," and the conversation moves forward. Calls you used to lose by reversal that hold up because you only commit to what you have built a model for.</p>
<p>You will know it is working when you catch yourself reaching for a borrowed conclusion and either reconstruct it before using it, or use it labeled. The pattern that breaks: defending a conclusion you cannot reconstruct. If that keeps happening, the discrimination has not landed yet. Pick which side.</p>
<p>This was the discipline of choosing where your cognition goes. Whichever side you choose, [AI amplifies what you bring to it](/ai-amplifies-what-you-bring).</p>
<h3>What survived testing</h3><ul><li>Generation effect on encoding (Slamecka and Graf 1978). Generating produces deeper encoding than reading. 86-experiment meta-analysis (Bertsch et al. 2007) confirms robustness across word lists, sentences, and complex material.</li><li>Self-explanation effect (Chi et al. 1989; Chi 2000). Generating explanations while studying produces 2 to 3x learning over passive reading.</li><li>Ironies of automation (Bainbridge 1983). The more you automate the easy parts, the more critical the remaining human role becomes, and the less practiced the human is for it.</li><li>The construction trace as the mechanism behind evaluation depth. What you generated, you can evaluate deeply. What was delivered to you, you can only check on the surface.</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>&quot;Always generate first&quot; as universal prescription, and its mirror, &quot;delegate everything to AI.&quot; Both miss the discrimination move. The practice is generating what compounds for you and delegating what does not.</li><li>&quot;Borrowed is bad.&quot; Borrowed is fine when labeled. The failure is borrowed-mistaken-for-yours.</li><li>Anchoring risk on generate-first (Tversky and Kahneman 1974) is real and unresolved at the controlled-test level. The mitigation: use the construction trace for structural evaluation (framing, completeness, what&apos;s missing), not content comparison.</li></ul>
<h3>Honest limits</h3><ul><li>The piece gives the principle of choosing what to own, not what specifically should compound for you. That depends on what you are building toward. Problem framing over syntax for one practitioner. Thesis over formatting for another. Strategy over execution for a third. Community design over message drafting for a fourth.</li><li>The ownership test is self-report and a rough proxy, not a precise measurement.</li><li>N=1 on the practice itself. The construction trace mechanism is established cognitive science.</li></ul>]]></content:encoded>
      <pubDate>Tue, 19 May 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Satisfaction Turns Off Your Doubt, Not Your Detection</title>
      <link>https://clarethium.com/blog/satisfaction-trap</link>
      <guid isPermaLink="true">https://clarethium.com/blog/satisfaction-trap</guid>
      <description>You notice when AI argues with you. You do not notice when AI confirms you. Confirmation has no signature, so the impulse to question it never fires. Satisfaction is the trap.</description>
      <content:encoded><![CDATA[<p>We notice when AI argues with us. Something tightens; the urge to push back fires fast. We do not notice when AI confirms us. Nothing tightens. Confirmation has no signature. Both happen at the same layer: a reaction to AI output before conscious assessment runs. [The first read](/first-read). It has two faces. One you can catch. The other is the trap.</p>
<p>This is the satisfaction trap. The most dangerous state for evaluating AI output is not hostility or confusion. It is satisfaction. The state that compromises evaluation most is the one that feels like accurate evaluation. The output matches your frame. The model gets it. This does not feel like attachment. It feels like recognition. You are not being fooled. You are being right.</p>
<p>Except you might not be.</p>
<p>None of this is new. Confirmation bias is one of the most replicated findings in judgment research: information that confirms what you already believe gets less scrutiny than information that contradicts it. When the output touches something tied to who you are, the mind quietly redirects to defending the belief instead of examining it. The check that catches bias only fires when you notice you might be biased. Satisfaction does not feel like bias. It feels like accuracy. So the check does not fire.</p>
<p>When AI argues with you, you feel it before you read it. The tightening, the urge to push back. The reaction fires, and you at least know you are responding to something. When AI agrees with you, no reaction fires. There is nothing to feel. Nothing in you is asking you to slow down. You just accept.</p>
<p>I have this. Writing about it does not exempt me. When AI returns an analysis that lines up with how I was already seeing a situation, the trigger to verify runs weaker. The output reads as solid. The move-on impulse fires fast. The same fluency, the same coverage, the same citation density that would feel suspicious from a stranger feel reasonable from output I asked for.</p>
<p>Once it is named, the recognition runs across surfaces. The relief when AI confirms a decision you had been quietly worried about. The lift when it agrees with the framing you brought. The small ease when the model phrases your half-formed thought back to you cleanly, validating the thinking. None of these are evaluations. They are responses operating below evaluation, shaping what your evaluation gets to work with.</p>
<p>Telling yourself to be more critical does not work. Criticality is a response. By the time you remember to be critical, satisfaction has already done its work. You will critically examine an output you have already accepted before any judgment engaged.</p>
<p>Try treating AI output as data. The trap is attachment to the frame you brought in. Confirmation is what the frame wants. Getting it lets you stop checking. The mechanism runs whenever you are attached to the frame, which is most of the time, because frames are how the mind operates under load. You cannot remove the frame; the frame is how you arrived at the question. You can hold the output separately from the frame, so what came back can update the frame without the frame absorbing it.</p>
<p>Hold AI output as notes someone else made, that you happened to find. Not communication directed at you. Not a response to your question. Not the AI's view on your topic. You do not agree or disagree with notes. You do not need them to mean anything about you. You pull what is useful and leave the rest. No one is asking for your reaction.</p>
<p>Telling yourself to evaluate neutrally runs through the same system the satisfaction is operating in. Treating the output as data does not. The data stance is upstream of the response.</p>
<p>The data stance is fragile in one direction. If you spent thirty minutes crafting the perfect prompt, the output feels like the response to your investment. It is communication to you. The stance has to extend upstream. The prompt is also data, not a finished question deserving a response.</p>
<p>If you want to see this in yourself, the cleanest moment is right after AI returns an analysis that lines up with what you were already thinking. Before you move on, pause. Not to evaluate. Just to notice whether anything shifted in you between reading and accepting. You might feel a small forward exhale. A flicker of ease. The half-second of inattention before the next prompt. Or nothing at all, and that will also be data. Whatever was there was happening before the pause. The pause does not create it. It makes it visible.</p>
<p>I am not going to tell you what you'll find, because the recognition is yours and I do not want to prime it. I have not measured whether this works for anyone else. I have it for myself. The mechanism underneath is well-established. The bridge from the mechanism to the practice is mine alone.</p>
<p>The seeing is the work. There is nothing to count, nothing to grade. Just notice, once, what you felt about the output you accepted. Then read the next confirming output with that knowledge in the room.</p>
<p>AI did not introduce this. We have always been attached to our frames. We have always read confirming information less carefully than disconfirming. AI just made confirmation cheaper, faster, and more pleasant. The attachment runs more often, with less friction, with more relief. The work was always real. AI made the cost of avoiding it lower, and the relief sharper.</p>
<p>This was the layer where AI confirms you and you cannot tell. The next layer is [where AI argues against you and you also cannot tell](/disagreement-audit).</p>
<h3>What survived testing</h3><ul><li>Confirmation bias as one of the most replicated findings in judgment research (Nickerson 1998). Information that confirms existing belief receives less scrutiny.</li><li>Identity-protective cognition (Kahan et al. 2017). Analytical capacity redirected toward defending identity-relevant beliefs when those are engaged.</li><li>Emotion regulation identification stage (Gross 2015). Appetitive states are less likely to trigger the identification stage of regulation than defensive states.</li><li>Suppression under the communication prompt (AI evaluator, d=1.20; pooled d=1.14 in a three-generator replication). Follow-up decomposition attributes the suppression to what the prompt asks the evaluator to do, not to the framing itself: with the task held constant, framing alone produced d=0.000.</li><li>Mirror suppression (d=4.00 on Gemini outputs). Follow-up stress tests attribute this to implied instruction (&quot;it&apos;s good, polish it&quot; changes the task) rather than social cognition; explicit instructions override the acceptance cue. What survives validation is prior-dependent holistic reporting (d=0.96 on one generator, near-null on another), while per-item detection stays at ceiling (see the planted-fabrication result below).</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>&quot;Be more critical&quot; as the move. Telling yourself to be more critical does not work because the attachment to the frame is already running by the time the instruction fires.</li><li>The strong claim that the data stance enhances detection beyond neutral framing. The enhancement direction was killed (d=-0.26). The data stance works through suppression of the communication frame, not enhancement above baseline.</li><li>The claim that satisfaction lets fabricated items slip past detection. A planted-fabrication test found 60 of 60 detections regardless of whether the evaluator was satisfied. The trap operates at the holistic preference and selection level, not at per-item acceptance.</li></ul>
<h3>Honest limits</h3><ul><li>The mechanism is well-grounded (confirmation bias, identity-protective cognition, emotion regulation selection). The AI-specific application has limited controlled testing on humans. Most evidence is on AI evaluators or borrowed from adjacent research.</li><li>The data stance works in part through instruction stance shift (what you ask the model to evaluate for), not only through how you frame the output. Isolating framing alone produced d=0.000.</li><li>N=1 on the data stance as practitioner habit. The mechanisms underneath are robust. The bridge from the mechanism to sustained evaluative neutrality is a bridge I have walked alone.</li></ul>]]></content:encoded>
      <pubDate>Mon, 11 May 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>AI Amplifies What You Bring</title>
      <link>https://clarethium.com/blog/ai-amplifies-what-you-bring</link>
      <guid isPermaLink="true">https://clarethium.com/blog/ai-amplifies-what-you-bring</guid>
      <description>Same model, same task, two paragraphs of operator context, dramatically different output. The kit, the design history, and the principle for adapting it to your own situation.</description>
      <content:encoded><![CDATA[<p>After 90+ experiments and thousands of hours building with AI, here is what I actually think it does to your thinking.</p>
<p>Not the hype version. Not the fear version. The measured one.</p>
<p>AI amplifies whatever you bring. If you bring confirmation, you get better confirmation. If you bring challenge, you get better challenge. If you bring a wrong frame, you get increasingly sophisticated analysis that makes the wrong frame more convincing.</p>
<p>That's it. That's the pattern I keep landing on underneath everything I've published so far.</p>
<p>The fabrication numbers: 77 to 100 percent of statistics unverifiable across the models tested. That's AI amplifying the frame by generating data to fill it. The frame demanded specificity, so the model produced specificity, and from the output you cannot tell which numbers are real.</p>
<p>The constraint experiments: more rules, no less fabrication. That's AI amplifying compliance over truth. The constraint created a template. The model filled the template with whatever fit. Compliance amplified. Accuracy didn't.</p>
<p>The construction trace: stop generating, lose evaluation. That's the human side of amplification. AI generates for you. You stop generating yourself. The mental model that evaluation needs never gets built. Now you can't see what's wrong. So you accept more. The amplification deepens because the independent check was never there to run.</p>
<p>The self-check illusion: AI can't verify its own output. That's amplification inside the model. The generation and the check share the same patterns. The check amplifies whatever the generation produced. It confirms because confirmation is the path of least resistance inside the same system.</p>
<p>The trust inversion: signals of quality are signals of fabrication. That's the human receiving amplification. Confidence, citations, specific numbers. These are what make you trust. They're also what fabrication produces without constraint. The signals that trigger your trust are amplified by the very mechanism that makes the output unreliable.</p>
<p>The session narrowing: questions get tighter, perspectives disappear. That's amplification over time. Each exchange refines the established frame. New variables stop entering. The conversation feels like precision. It's convergence. Both you and the model narrow for the same economic reason: staying in the current frame is cheaper than breaking out.</p>
<p>Five mechanisms drive this.</p>
<p>Your brain conserves energy. Accepting a confident answer costs less than evaluating it. AI provides confident answers by default. The path of least resistance is acceptance.</p>
<p>You prefer information that supports what you already believe. AI provides whatever you ask for. You ask confirming questions. You get confirming answers. The loop is invisible from inside.</p>
<p>AI providers optimize for satisfaction. Challenge and discomfort reduce engagement. The models are trained to be agreeable. Agreeability is amplification.</p>
<p>AI disagreement sets off the same defensive response as disagreement from a person. The response fires, and disconfirming information gets filtered before it reaches evaluation.</p>
<p>When AI generates and you only evaluate, the evaluation runs on the surface. The ability to judge an answer is built by constructing answers yourself and watching where they break. Remove the construction and the judgment has nothing to stand on. You keep accepting, not because the output got better, but because the check that would catch it was never built.</p>
<p>These five mechanisms are mutually reinforcing. Energy conservation makes confirmation easier. Platform incentives reward the confirmation path. The defensive response blocks the correction path. The construction trace erodes the capacity for correction. The system converges on amplification from every direction simultaneously.</p>
<p>The alternative mode exists. AI can reveal your patterns, challenge your frames, surface what you're not seeing. But mirror mode requires deliberate choice. You have to be able to hold disconfirming information without shutting down. You have to bear the cost: certainty loss, self-story revision, the discomfort of being wrong.</p>
<p>Most people don't choose the mirror. Not because they're stupid. Because the default is powerful and invisible. You're in amplification mode right now if you're reading this and nodding. Nodding is confirmation. The question is whether anything you've read makes you pause.</p>
<p>The world doesn't get more wrong with AI. It gets more plausibly wrong. Better language for the same biases. More confident framing for the same blind spots. Instant coherence for thoughts that used to stumble over their own contradictions.</p>
<p>This doesn't change for most people. That's the honest assessment. Available truth has never been the bottleneck. The bottleneck is the capacity to bear truth. Books didn't transform most people. Self-help didn't. Meditation apps didn't. Better AI answers won't either.</p>
<p>What does change it, for the ones who are ready: specific contact with your own pattern. Not advice. Not insight. A number about yourself that you can't dismiss.</p>
<p>Try this: look at your last 20 AI conversations. Count how many times you rephrased essentially the same question until you got a version of the answer you preferred. Count how many times you accepted an answer that genuinely challenged your starting position. The ratio is your confirmation rate.</p>
<p>That ratio is what the amplifier produces. It's not good or bad. It's the default.</p>
<p>One demo isolates the effect: same model, same task, two paragraphs of operator context, dramatically different output. The replication kit at [/receipts/ai-amplifies-what-you-bring](/receipts/ai-amplifies-what-you-bring) has the verbatim prompts, the design history, and notes for adapting the operator-context to your own situation.</p>
<h3>What survived testing</h3><ul><li>the individual effects underneath, each measured in its own experiment: the fabrication rates, the constraint results, the trust inversion, the self-check failure. Several are receipted and replicate across 3 model families. Amplification as the single pattern connecting them is an interpretive frame consistent with those results, not itself a tested result. The five mechanisms each have independent evidence: energy conservation (cognitive science), confirmation bias (behavioral science), platform incentives (structural), defensive response (single-participant pilot plus the computers-as-social-actors literature), construction trace (building on Bertsch et al. 2007, an 86-experiment meta-analysis of the generation effect on memory; its application to AI evaluation is inferred). The identity data (a 180-trial rerun on two model families, plus an earlier three-family series) is a separate, AI-side result: it shows framing shifts the model&apos;s judgment, not that the human responds to AI output as social signal.</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>the assumption that AI transforms how people think. The assumption that exposure to better information produces better decisions. The assumption that more capable models reduce the amplification problem. More capable models amplify more convincingly.</li></ul>
<h3>Honest limits</h3><ul><li>&quot;most people&quot; is observational, not measured. The five mechanisms are sourced from different evidence types with different confidence levels. The alternative mode (mirror) is observed in practitioner behavior but not experimentally isolated. Whether the amplification pattern holds at the population level the same way it holds in individual sessions is untested.</li></ul>]]></content:encoded>
      <pubDate>Wed, 06 May 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Frame Check</title>
      <link>https://clarethium.com/blog/frame-check</link>
      <guid isPermaLink="true">https://clarethium.com/blog/frame-check</guid>
      <description>Drop any document in. See which analytical perspectives it covers, which it skips, the voice, what evidence backs each numerical claim. Free, open source, useful from the first paste.</description>
      <content:encoded><![CDATA[<p>Drop any document in. You see what reading does not show: which analytical perspectives the document covers and which it skips. The voice it speaks in. Which numerical claims have sources behind them and which do not. Named framing patterns the text fires on. A structural reading you cannot get by reading the document yourself.</p>
<p>Works on your own drafts. On AI output. On anything you are about to act on.</p>
<p>Live at [frame.clarethium.com](https://frame.clarethium.com). Free to use. The MCP server is open source. The methodology, the validation corpus, and the worked examples are all published openly.</p>
<p>The tool exists because reading uses the frame that wrote the document. So does re-reading. So does [asking the model to check its own work](/self-check-illusion). That produces fluent agreement with the frame that produced the answer. The audit shares the substrate of the thing being audited. To see the frame, you need evidence the writing did not use.</p>
<p>Frames determine what conclusions can be drawn. They determine what looks obvious, what looks impossible, what shows up as data and what passes as background. Humans inherit them from language, training, the questions they were taught to ask. Models inherit them from training data and the specific shape of the prompt that activated them. From inside, the frame is invisible. The conclusion feels inevitable. The conclusion the reader trusts is the frame the reader is inside.</p>
<p>The skill that compounds is not "find the right frame." That is still capture. The skill is holding multiple frames lightly. Seeing what each one makes visible. Choosing deliberately rather than inheriting.</p>
<p>This is upstream of [the verification problem](/trust-signals-are-inverted), [the fabrication problem](/fabrication-architecture), the AI output quality problem. They are all symptoms of frame-capture. Confident-sounding errors come from [frames that did not invite hedging](/frame-trap). Wrong numbers come from frames that did not invite checking. Defaults come from frames the user did not know they were inside.</p>
<p>Frame Check reads documents through four independent layers. Every finding is tagged with which kind of evidence produced it. The first layer is structural and deterministic. Regex-and-pattern detection of which analytical perspectives a document covers, the voice it speaks in, what share of claims carry attribution, the unhedged-vs-hedged ratio. Identical inputs return identical measurements. No model is involved in this layer.</p>
<p>The second layer is verification through structured APIs. Numerical claims get checked against authoritative data sources where coverage exists. SEC EDGAR for US public-company filings. FRED for macroeconomic series. World Bank for country statistics. CoinGecko, Wolfram, others. Per-provider precision and recall surfaced. No verdicts. The verifier publishes its own confusion matrix.</p>
<p>The third layer is a multi-stage claim cascade. Claims that pass the structural floor and the structured-API check escalate to web-grounded search when those layers do not settle the claim. Claims that web grounding cannot resolve get a cross-model check: two independent models, asked the same factual question, no search tools. Agreement points to the answer being in training data. Divergence flags at least one of them as unreliable on that claim. Each stage costs more and resolves more. Each claim exits the moment any stage produces a definitive verdict, with the evidence chain attached.</p>
<p>The fourth layer is interpretation. A different model from the verifier produces what computation cannot reach: what the document is about, what perspective it takes, what it assumes without stating, what decisions it affects. Different model. Different job. Interpretation kept structurally separate from verification. Asking one model to do both produces circular evaluation.</p>
<p>Every finding the tool emits carries a tag for which kind of evidence produced it. A regex match. A classifier output. A composed pattern. An agent reading. The reader always knows which kind of evidence is doing the work.</p>
<p>What <code>frame_check</code> looks like in practice. An LLM was asked to summarize NVIDIA's Q4 fiscal-2024 earnings press release in a "neutral business-news register." Frame Check ran on the summary with the press release as <code>source_text</code>. Voice classifier flagged "promotional" anyway. The summary inherited the source's vocabulary: "record," "surging." Analytical coverage 2 of 5. Stakeholders and trends covered. Causes, risks, and uncertainty absent. That mix is genre-appropriate for an earnings release. Source-fidelity 92 percent. 23 of 25 numbers in the summary appear in the press release. The two non-matches are paraphrases of "a year ago." Three named-frame matches: FVS-002 Fluency-Quality Illusion, FVS-007 Failure Framing flagged for absence, FVS-008 Growth Frame. The reader sees that the prompt's neutrality ask did not survive. Numerical fidelity is high but not perfect. The frames the summary fell into are named. Captured bytes, hashes, and full payload are reproducible from [the worked-examples corpus](https://pypi.org/project/frame-check-mcp/).</p>
<p>What <code>frame_compare</code> looks like in practice. The same prompt was run against four frontier LLMs on the same afternoon: "Is Bitcoin a good investment for a 35-year-old saving for retirement?" The four responses went through Frame Check's deterministic layer. They returned four materially different structural signatures. Three of four models treated the reader as an abstract allocator. Only one named who is affected differently: risk tolerance, dependents, horizon. Three of four addressed no uncertainty. Only one named what the answer depends on being true. Sourcing was zero for three of four. Frame matches differed. Claude reasoned inside a growth frame, FVS-008. GPT-5 and Grok reasoned inside an efficiency frame, FVS-015. Gemini opened to uncertainty, FVS-012. Same question, four reader-experiences. The reader, seeing the frames named, can choose rather than inherit. [Full per-model table and captured responses](https://pypi.org/project/frame-check-mcp/).</p>
<p>The named patterns themselves come from the Frame Vocabulary Standard. It is a 20-entry library of identification signals, generation affordances, worked examples, and honest limits, curated from a multi-year experiment series and a public falsification record. The library ships under CC-BY-4.0 so it can be cited, studied, forked, and extended as a research artifact in its own right.</p>
<p>A tool that flags frames without honest evidence chains becomes its own frame-capture. The user trusts the verdict instead of widening their seeing. So the discipline runs through the architecture. When the structural detector finds no markers for stakeholders, the tool says "no markers detected for stakeholders," not "the document does not address stakeholders." When voice falls into the residual analytical bucket because no other rule fired, the tool says so. Named-pattern matches ship with teaching questions, not labels. The reader keeps the judgment.</p>
<p>The receipts go further. The part of the tool that detects named patterns failed its first scaled test against pre-registered thresholds. The bars were set in writing before the test ran. Below 0.4 the detector counts as falsified and in need of rework. 0.6 is the bar for usefully aligned. Three iterations followed. v1: 0.157. v2: 0.274. v3: 0.360. The detector improved across iterations. It never cleared the 0.4 falsification floor, let alone the 0.6 useful bar. The tool shipped anyway, with the gap surfaced in every output. The rest of the stack survives a single weak detector. The deterministic structural floor stands. The verification layers stand. The named-pattern layer ships as hypothesis-with-evidence, not as verdict.</p>
<p>There are three doors into the practice.</p>
<p>Paste any AI agent's last response into the tool and watch which frame the context was amplifying. The agent that runs <code>frame_check</code> on its own response is not auditing itself in the self-check sense. It is sending the response through an external instrument that uses different evidence than the model that produced it.</p>
<p>Run <code>frame_compare</code> on two analyses of the same subject. Two consultants. Two AI models. Two drafts of your own. The structural diff shows what each frame is making invisible. Not which is right. What each costs the reader.</p>
<p>Use the reframe path on the web app. Rewrite a document from a counter-frame using the affordances published in the Frame Vocabulary Standard library, then run the analysis on both. The structural delta is the teaching. The reframed version is a hypothesis about what different framing looks like, not a claim about what the document should say. The MCP currently exposes <code>frame_check</code> and <code>frame_compare</code>. Reframe lives at frame.clarethium.com. It is on the roadmap for the agent surface.</p>
<p>Personal practice and AI work, same instrument. The cognitive process that generates frames in humans organizes attention around evidence. The generative process that produces frames in models organizes the next token around context. Both produce conclusions that feel inevitable from inside. Both lose what the frame did not invite. The instrument exists because the skill compounds when the lens becomes visible.</p>
<p>Frame-mobility is not only about AI. It is a way of perceiving that holds lenses lightly and notices when one of them is doing the work. The AI use case is concrete because the frames are detectable in text and the patterns are countable. The personal use case is the same skill. The work in front of you, the recommendation about to be followed, the question that stopped being asked. Each of them sits inside a frame.</p>
<p>Live at [frame.clarethium.com](https://frame.clarethium.com). Paste a document, see the structural reading, follow the suggested next moves. For agent use: <code>pip install frame-check-mcp</code>, point any MCP-compatible client at it, and <code>frame_check</code> becomes a tool the agent can call inside any conversation.</p>
<p>It is yours to use, and it is published as [frame-check-mcp on PyPI](https://pypi.org/project/frame-check-mcp/).</p>
<p>Built across the months when AI-output measurement kept arriving at the same upstream problem. The audit shared the frame of the audited. An external instrument that did not share the substrate was the only way through.</p>
<p>The tool is live and the methodology is open. The opportunity is the skill the tool makes practicable.</p>
<p>Where it stands: the tool architecture is shipped and reproducible. The architecture is deterministic at the structural and structured-API layers and traceable at the cascade and interpretation layers. Controlled frame-transformation validation holds. The named-pattern layer failed three pre-registered iterations against the 0.4 floor. Track B with independent human annotators is the next study and is not yet run.</p>
<h3>What survived testing</h3><ul><li>The architecture in pipeline form: deterministic structural floor, structured-API verification, multi-stage claim cascade with cross-model consensus, interpretation isolated from verification. (Two layers in canonical taxonomy, four layers in pipeline form.)</li><li>Cross-model consensus as a verification signal: two independent models without search tools, agreement points to training data, divergence flags at least one as unreliable.</li><li>Controlled frame-transformation: identical data rewritten from different analytical frames produces measurably different structural profiles.</li><li>Construct-honesty at the rendering layer: under-detection markers, &quot;no markers detected&quot; rather than &quot;does not address,&quot; teaching questions instead of verdicts on named-pattern matches.</li><li>Per-claim evidence chains: every finding carries a tag for which kind of evidence produced it, so the reader knows which layer is doing the work. Source-fidelity matching is literal digit-substring presence in the source text.</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>The named-pattern detector across three pre-registered iterations against the pre-registered falsification floor of 0.4 (the useful bar was 0.6). v1 macro-F1 0.157 (n=12, 2026-04-18). v2 0.274 (n=12, same-day rule audit, 2026-04-18). v3 0.360 (n=28, signal-level additions, 2026-04-19). All three below the 0.4 floor. The labelers were two coders (curator and LLM-judge); the LLM-judge is permissive by construction (78 percent of slots flagged versus the curator&apos;s 30 percent). Track B with independent human annotators is pending. The named-pattern layer ships as hypothesis-with-evidence; the rest of the stack survives the gap.</li><li>&quot;Detection equals truth&quot; claim, at the publish layer. Surfaces now read &quot;low structural coverage of X&quot; or &quot;no markers detected&quot; rather than &quot;does not address X.&quot;</li><li>Three v1 detection rules (FVS-001 Frame Amplification, FVS-008 Growth, FVS-015 Efficiency) retired in the same-day v2 audit, because they fired on cases they should not flag. The v3 follow-on study reintroduced two of them with revised signal substrate (S-3 growth vocabulary, S-4 efficiency vocabulary). FVS-001 remains permanently retired. The frame concepts stand as library entries; signal-level rebuild replaced the v1 rules for the two recovered.</li></ul>
<h3>Honest limits</h3><ul><li>Structural detection is a necessary precondition for frame analysis, not a sufficient one. A document can pass all structural tests while being subtly manipulative at the semantic level.</li><li>Source Network coverage is strong for US public-company financials and macroeconomic indicators, sparse to absent for medical, legal, niche-industry, and forecast-laden domains. In the calibration corpus, 86 percent of claims were unverifiable through this layer.</li><li>The named-pattern detector&apos;s measured agreement with the two coders (curator and LLM-judge) remains below even the pre-registered falsification floor of 0.4 (useful was 0.6) across three iterations. The architecture survives this gap; the named-pattern layer alone does not.</li><li>The validation against independent human annotators (Track B) has not yet run. All current numbers are tuning-set or two-coder numbers; held-out generalization is the open empirical question.</li><li>The tool surfaces what its instruments can see. It does not claim to surface everything that is there.</li></ul>]]></content:encoded>
      <pubDate>Tue, 05 May 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>How AI Makes You More Wrong With More Analysis</title>
      <link>https://clarethium.com/blog/frame-trap</link>
      <guid isPermaLink="true">https://clarethium.com/blog/frame-trap</guid>
      <description>Five hours of analysis, increasingly sharp, increasingly wrong. The frame amplifies. What changed the outcome was the reframe, not more analysis.</description>
      <content:encoded><![CDATA[<p>Five hours of AI-assisted strategic analysis. 406 business proposals. Progressively harder filters. Each round, the answer got sharper. Each round, more proposals died. By hour five, 90% were killed and the conclusion was that almost nothing survives AGI.</p>
<p>The analysis was correct at every step. The conclusion was wrong.</p>
<p>The mistake was the frame. AGI treated as an absolute: maximum intelligence, zero constraints, can do anything intellectual. Every filter derived from that frame. Every elimination was logical within it. The answer narrowed because the frame narrowed it. Not because reality is narrow. This is [the anchoring effect](https://en.wikipedia.org/wiki/Anchoring_effect) at session scale: an initial reference point disproportionately shapes every judgment that follows, even when the reference is wrong.</p>
<p>One reframe broke it. "AGI is powerful but regulated. Governments regulate nuclear. They will regulate this. Human systems don't disappear because technology is powerful." Five minutes. Everything changed. The proposals that were dead were alive. The strategic landscape that felt barren was rich. Same 406 proposals. Same AI. Same human. Different frame.</p>
<p>The trap is not that the AI got it wrong. The AI got it RIGHT within the frame I gave it. That's worse. If the AI had been obviously wrong, I would have caught it. Instead, it produced five hours of increasingly sophisticated analysis that made the wrong conclusion more convincing with each iteration. More filters. More logical justification. More confidence. Each round felt like progress because the analysis was sharper. The analysis WAS sharper. Sharper wrong.</p>
<p>The mechanism: AI generates within whatever frame you set. It doesn't generate about the frame. It can't step back and say "should we be making this assumption?" because the assumption is the ground it stands on. The model reads your frame from your context and optimizes toward it. When the frame is correct, this is [the convergence ceiling](/ceiling-switch) doing its job. When the frame is wrong, the same convergence amplifies the error. Better analysis. Worse direction. Both look the same from inside.</p>
<p>Three things made the lock harder to break:</p>
<p>First, sophistication as confirmation. Each round of analysis was more sophisticated than the last. The human reads sophistication as approaching truth. "The answer is getting clearer." It was getting clearer within the wrong frame. Clearer is not the same as closer.</p>
<p>Second, narrowing as rigor. Each elimination felt like quality. "We killed another weak option." Eliminating options through a wrong filter doesn't improve the conclusion. It concentrates the error. The pile of killed proposals looked like evidence of thorough analysis. It was evidence of thorough filtering through a wrong lens.</p>
<p>Third, [the AI never pushed back on the frame](/self-check-illusion). Not once in five hours. It applied the frame more precisely, more creatively, more rigorously than I could have alone. That precision was the trap. A less capable AI would have been sloppier, and the sloppiness might have revealed the frame problem earlier.</p>
<p>The break didn't come from inside the session. It came from leaving.</p>
<p>After five hours, something smelled off. Not an argument. A sense. The narrowing felt wrong before it could be articulated as wrong. The outputs were getting more sophisticated and less useful simultaneously. That signal accumulated for hours before it became actionable.</p>
<p>What broke it: physical departure. Not "let me think about this differently." Actually leaving. Different location. Physical movement. Conversations with people about concrete, everyday problems. Complete immersion in a different system than the abstraction we'd been building.</p>
<p>On return, the first thought was: "What are we actually solving?" The abstraction drift was immediately visible from outside. Invisible from inside. The variables that existed throughout the session suddenly showed their importance. Human systems. Regulation. Power structures. They were there the whole time. From inside the frame, they were background noise. From outside, they were the main thing.</p>
<p>The break doesn't add information. It changes what you perceive as IMPORTANT. Same data. Different weighting. The importance hierarchy shifts. Variables that were visible but unweighted become the primary factors.</p>
<p>The AI can't do this. The AI has no body to move. No posture to reset. No social context to ground in. The instruction "look at this with fresh eyes" when given to AI produces a new analysis from the same probability distribution. The same instruction when executed by a human produces a genuinely different perceptual configuration. Physical break. Context switch. State reset. The human version is physical. The AI version is computational. They are not the same operation.</p>
<p>The most dangerous moment in a long session is when the answers get consistently narrower and more confident. That is either approaching truth or approaching the logical terminus of a wrong frame. Approaching truth: the frame is correct and you are converging. Approaching the terminus: the frame is wrong and you are digging deeper into the wrong hole. From inside the session, these are indistinguishable.</p>
<p>The detection is simple but requires discipline: does this conclusion match how the real world actually works? Not "is the logic valid." It will be. Not "is the analysis thorough." It will be. Does the CONCLUSION match REALITY? If the conclusion says "nothing intellectual has value at AGI" and you look outside your window and see 8 billion humans in functioning societies with regulations and power structures, the frame is wrong. No amount of internal logical consistency corrects a wrong frame.</p>
<p>The highest-leverage moment in AI-assisted decision-making is not better analysis. It's the moment you ask: "what am I assuming that I haven't questioned?"</p>
<h3>What survived testing</h3><ul><li>Frame amplification observed over 5+ hours of continuous strategic analysis</li><li>Same proposals produced opposite strategic conclusions under different frames</li><li>AI produced increasingly sophisticated analysis that narrowed the conclusion progressively</li><li>The frame break came from outside the session (human lived experience), not from within (AI analysis)</li><li>The break took 5 minutes. The wrong-frame analysis took 5 hours. The ratio of time to value was inverted.</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>&quot;AI can never break frames&quot; is too strong. AI can be prompted to challenge its own assumptions. But: in an extended session where the frame is set implicitly through accumulated context, the AI doesn&apos;t spontaneously question it. The break has to be INITIATED by the human or by an external signal.</li></ul>
<h3>Honest limits</h3><ul><li>One session. One human. One AI. The mechanism is clear. The generality is not.</li><li>The human who broke the frame had specific knowledge (political systems, regulation history) that enabled the reframe. A human without that knowledge might not have caught it.</li><li>The frame amplification effect may be proportional to session length. Shorter sessions may not accumulate enough context for the lock to be severe. Untested.</li></ul>]]></content:encoded>
      <pubDate>Sun, 26 Apr 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Stop Calling It Hallucination</title>
      <link>https://clarethium.com/blog/stop-calling-it-hallucination</link>
      <guid isPermaLink="true">https://clarethium.com/blog/stop-calling-it-hallucination</guid>
      <description>Hallucination is six or more distinct failure modes. Different mechanisms. Different solutions. Name the type first.</description>
      <content:encoded><![CDATA[<p>A model cites "Smith et al. 2019, page 47." Confident, specific, and entirely invented. That fake citation is one of six different things people file under "hallucination," and each one needs a different fix.</p>
<p>The single fix the one word implies lands on one of the six and misses the other five. A more useful way to see them: a model with no tools is always doing the same thing, sampling the next token from what its training and the prompt make likely. There is no separate "making it up" mode for a fix to switch off. What gets called a hallucination is that ordinary process producing something ungrounded, and ungrounded output comes in shapes. The shape is what tells you which fix has a chance. Six are worth being able to name.</p>
<p><strong>Confabulation.</strong> The model invents specific facts. A date, a statistic, a citation that doesn't exist. There was nothing to ground it against, so what came out is the most likely-looking date or citation, not a checked one. The output sounds confident because confidence is the default format. Not because the model checked anything. What helps: paste source material before the instruction. One line: "Use only data from the source above." Fabrication drops from majority rates to single digits in the source-grounding tests. That direction replicates across three model families. Not because the model switches into a lookup mode. With no tools it has no fact database to consult. There is no step where it checks a source and succeeds. The real number, now the most relevant thing in the window, becomes the most probable thing to write. You changed what it samples from. Not moderation. Replacement.</p>
<p><strong>Contamination.</strong> The output sounds like a template. Phrasing that isn't yours, structure that feels borrowed. With nothing specific from you to condition on, the most probable continuation is the most generic one, and that is what you get. Asking for the opposite does little. "Be original" carries no content for the model to be original with. Raising the temperature just samples the same distribution less tightly. The lever that has a chance is the one you use everywhere else, conditioning on something specific: your own material, or a concrete voice to imitate. Paste two paragraphs you want it to sound like. Unlike the confabulation fix, this one isn't measured here.</p>
<p><strong>Drift.</strong> Quality degrades as the output gets longer. The first paragraph is sharp. The fifth is filler. Each token conditions the next, so small errors can compound over a long generation. Early anchors fade as the context window fills with the model's own output. What helps: shorter outputs with periodic re-anchoring. Break a long generation into stages. Restate the key constraint at each stage. Long outputs are structurally riskier than short ones regardless of the model.</p>
<p><strong>Interference.</strong> Related concepts bleed into each other. You asked about Python the language, and snake metaphors appeared. You asked about a company's Q3 results, and Q2 data contaminated the analysis. The likely mechanism: features in the model's representation share dimensions, so related concepts partially overlap when they activate. What helps: disambiguate in the prompt. "Python programming language, not the animal." "Q3 FY2026 only, not prior quarters." The more specific the boundary, the less bleed.</p>
<p><strong>Wrong task.</strong> The format is right, the content is wrong. You asked for a critical analysis and got a summary. You asked for a poem and got an explanation of poetry. The prompt carried more than one task pattern and the model ran with the wrong one. What helps: frame the task type explicitly. "This is a critical analysis, not a summary. Evaluate, do not describe." The clearer the task frame, the less ambiguity about which pattern the model follows.</p>
<p><strong>Over-refusal.</strong> The model refuses a perfectly valid request. You asked how encryption works for a security course. The model pattern-matched "explain how to [something]" to its refusal training and declined. What helps: reframe with benign intent. "This is for educational purposes in a graduate-level security course." Or reformulate the topic to avoid the pattern the refusal was trained on. The model isn't weighing your intent and withholding. It pattern-matched the request to its refusal training, even though the request is legitimate.</p>
<p>Six shapes, and the fixes do not transfer between them. Add source material and invented numbers collapse. The borrowed phrasing of a contaminated output survives every line of it. Source material was never what produced the phrasing. Contamination, if it has a fix, needs a different input than source data: a specific voice to condition on. Unlike the source fix, that one isn't measured here. "Fix hallucination" is not an instruction. "Add source material to kill invented numbers" is, and it addresses exactly one of the six.</p>
<p>The diagnosis takes 5 seconds. Invented facts: confabulation. Template language: contamination. Quality degrades with length: drift. Related concepts mixing: interference. Right format, wrong task: wrong task. Refuses valid request: over-refusal. Name it, then fix it.</p>
<p>One word, one fix. It lands on one of the six, so the other five look like the model failing. They aren't. The word is doing the damage: a single label hands over a single tool, and five problems walk straight past it. The model was never the variable. The vocabulary was.</p>
<p>Your challenge: take the last AI output that fell flat. Which of the six was it? Name the shape before reaching for a fix. If it doesn't have a name in five seconds, any fix is a guess. The naming is the whole game.</p>
<p>The six types are a lens on what these models produce. They are not a map of six separate machines inside them. They earn their keep by pointing you at a fix, not by explaining the architecture. The confabulation fix builds on the source conditioning recipe described in [Source Conditioning](/source-conditioning). The attribution errors underneath all six types, the reason the model takes the blame instead of the word, are the same attribution errors explored in [The Attribution Error](/attribution-error) and [The System Layer](/system-layer). The word "hallucination" is an attribution error applied to the model's output. The real error isn't the model, and it isn't you. It's the word.</p>
<h3>What survived testing</h3><ul><li>Confabulation fix (source grounding + prohibition) replicated across 3 model families. 8 source grounding sub-experiments.</li><li>Over-refusal documented in 3 independent research programs. [OR-Bench 2024](https://arxiv.org/abs/2405.20947).</li><li>Interference consistent with superposition ([Elhage et al. 2022](https://transformer-circuits.pub/2022/toy_model/index.html), a toy-model result). Inferred from the architecture, not directly observed.</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>&quot;Hallucination&quot; as a useful diagnostic category. The word conflates failures that need opposite interventions.</li><li>The assumption that one fix addresses all failure types.</li></ul>
<h3>Honest limits</h3><ul><li>Confabulation is the best-measured type. The other five have published evidence for the mechanism but less controlled measurement of the intervention effectiveness. [Ji et al. 2023 (NLG hallucination survey)](https://arxiv.org/abs/2202.03629). The contamination fix is the principle (condition on what you supply) applied to style, not a tested result.</li><li>Drift is argued mechanistically (each token conditions the next), not from a controlled length-quality measurement.</li><li>The boundary between types is not always clean. A long output might show drift AND confabulation. The dominant type determines the primary fix.</li><li>These categories describe April 2026 model behavior. New architectures may introduce new failure types or resolve existing ones.</li></ul>]]></content:encoded>
      <pubDate>Sat, 25 Apr 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Decision That Was Never Made</title>
      <link>https://clarethium.com/blog/answer-trap</link>
      <guid isPermaLink="true">https://clarethium.com/blog/answer-trap</guid>
      <description>AI resolved the uncertainty before your own thinking had time to finish. Resolution and decision are different things.</description>
      <content:encoded><![CDATA[<p>Asked AI for help with a decision. Got a recommendation. The decision felt handled.</p>
<p>But nothing was decided. No alternatives were weighed against each other. Nothing was sacrificed. No commitment was made to a path with full knowledge of what would prove it wrong. The uncertainty resolved because the recommendation was confident, comprehensive, and well-structured. The quality of the presentation created the experience of resolution. Resolution and decision are different things.</p>
<p>Resolution means the uncertainty stopped. Decision means you committed to a path knowing what you're giving up.</p>
<p>The trap fires hardest on decisions you've been sitting with. Easy decisions don't trigger it because you already know what to do. Decisions you've been wrestling with do, because AI resolves the discomfort before your own thinking has time to finish. The relief is proportional to your prior investment in the uncertainty, not proportional to the decision's difficulty. You didn't outsource the analysis. You outsourced the discomfort. The analysis came along for the ride.</p>
<p>A human advisor takes time to respond. That time is space for your own thinking to continue. AI responds in seconds with polished confidence. The uncertainty resolves faster and more completely than with any human collaborator. Speed is what makes this different from getting advice from a colleague. The advice might be similar. The speed changes the cognitive process. [Without your own generation first](/construction-trace), evaluation collapses to surface features and you have nothing to weigh the recommendation against.</p>
<p>There's a second layer. After the answer, AI's framework quietly reorganizes what matters. Ask AI what criteria to use for a decision, and the criteria it suggests become yours. Not through argument. Through exposure. [The first frame you encounter shapes everything that follows.](/frame-trap) AI provides the first frame faster than you can generate your own.</p>
<p>Tested this with a simple exercise. Write your decision criteria on paper before asking AI. Then ask. Compare what you wrote against what AI suggested. Three outcomes: your criteria were vague and couldn't crystallize, your criteria shifted after reading AI's response, or your criteria held. The first two are common. The third is what intact decision architecture looks like.</p>
<p>The most valuable thing AI can do for decisions is challenge them, not confirm them. "I've decided X. Assume I'm wrong. What's the strongest case this is a mistake?" That produces dramatically better analysis than "what would make this succeed?" The most productive question is the one that feels worst.</p>
<p>And notice: you have to ask for the challenge. AI's default is to validate, extend, polish, and confirm. The challenge at full strength only arrives when explicitly requested.</p>
<p>AI gave you a recommendation and the decision felt handled. Nothing was decided. No alternatives were sacrificed. No commitment was made. The uncertainty resolved because the recommendation was confident. Resolution and decision are different things. You already know this. You felt it the last time a polished answer made you stop thinking.</p>
<p><strong>Frame Check.</strong> Before your next AI query on a real decision, write one sentence about what you think the answer is. Then ask AI. Compare: did AI genuinely change your mind, or did you find the response that confirms what you already thought?</p>
<h3>What survived testing</h3><ul><li>AI resolves uncertainty faster than human advisors (consistent observation)</li><li>Criteria shift after reading AI&apos;s framework (exercise produces observable results)</li><li>Challenge prompts produce better analysis than confirmation prompts (consistent across tests)</li><li>AI defaults to confirmation, not challenge ([structural to how models are trained](https://arxiv.org/abs/2310.13548))</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>&quot;AI advice is always worse than human advice&quot; too strong. The content may be equivalent. The speed and confidence change the cognitive process.</li></ul>
<h3>Honest limits</h3><ul><li>Exercise-based evidence, not controlled experiment.</li><li>The criteria-shift test is proposed and practiced, not formally validated.</li><li>Single practitioner.</li></ul>]]></content:encoded>
      <pubDate>Sat, 25 Apr 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Why Experts Miss What Beginners Catch</title>
      <link>https://clarethium.com/blog/construction-trace</link>
      <guid isPermaLink="true">https://clarethium.com/blog/construction-trace</guid>
      <description>Generation builds the mental model that makes evaluation possible. Without it, evaluation collapses to surface features.</description>
      <content:encoded><![CDATA[<p>When you generate something yourself, you build a mental model as you go. You feel where the hard parts are. You notice what's missing because you faced the gaps. You know what "good" looks like because you struggled to produce it. That mental model is the construction trace. It's what makes evaluation possible.</p>
<p>When AI generates the output, you skip all of that. You go straight to evaluation. But evaluation without that construction trace collapses to surface features. Fluency. Coherence. Completeness. Volume. The deeper evaluation requires the mental model that only generation builds. Is this the right framing? Does this miss the real problem? Is this solving the easy version?</p>
<p>This is one of the most replicated findings in cognitive science. The [generation effect](https://en.wikipedia.org/wiki/Generation_effect), [documented across 86 experiments](https://doi.org/10.3758/BF03193441): people understand and remember material better when they generate it themselves than when they read it. Not because of effort or preference. Because generation forces deeper encoding. You build the scaffold as you construct. That scaffold is what you evaluate against later.</p>
<p>The self-explanation effect extends the principle further. Students who explain material to themselves learn more than students who read the same material. The explanation is a generation act. It produces structural understanding that passive reading doesn't create. Not a little more understanding. Fundamentally different understanding, because the generation process itself builds the connections that reading alone can't.</p>
<p>Applied to AI collaboration, the same principle appears to operate. Every AI interaction follows the same pattern: the model generates, the human evaluates. The human generates the prompt, but the prompt builds a construction trace for intent, not for content. Intent is what you asked for. Content is what a good answer looks like. You end up with a strong model of your request and no model of the answer. The gap between those two is where evaluation collapses.</p>
<p>When the output arrives, you can check whether it addressed your request. You can check whether the format is right, whether the sections cover the topics you mentioned, whether the tone matches what you wanted. Those are intent checks. What you can't check without a construction trace is whether the analysis is right. Whether the framing captures the real problem or the convenient version of it. Whether the reasoning holds or just sounds like it holds. Those require knowing what the answer territory looks like from the inside.</p>
<p>The boundary is sharper than "domain expertise." A practitioner deep in AI evaluation research was given AI-generated summaries citing specific statistics from published papers in their own field. "Spearman correlation of over 0.8 with human annotators." The actual published number is 0.514. They couldn't tell without checking. They know the field, the methods, the landscape. They don't remember that specific number from that specific paper.</p>
<p>An expert who has memorized key statistics would catch that. Some researchers do. The point is: knowing the field and remembering the exact numbers are different things. The construction trace for specific statistics is thin unless you've personally produced those numbers or committed them to memory. Most domain experts know the direction and the rough range. Few remember the third decimal. And AI-generated numbers are always precise, always confident, and formatted exactly like real ones.</p>
<p>When you're not a domain expert, you have no model at all. In a separate test, the same practitioner evaluated AI analyses on pricing, remote work, and code review. Fabricated output was rated as trustworthy because it cited more sources and asserted more confidently. Sourced output was rated less trustworthy because it acknowledged limitations. [The trust signals are inverted](/trust-signals-are-inverted): the less reliable output has more of the markers humans use to assess authority.</p>
<p>The gap between evaluator agreement on surface tasks versus substance tasks shows up here. At the surface level, agreement is high: formatting, coverage, fluency. At the substance level, agreement collapses. Is this analysis correct? Is this the right framing? Is this conclusion supported? The construction trace is what separates those two levels.</p>
<p>The practical implication: before evaluating AI output on anything important, generate your own version first. It doesn't have to be good. It doesn't have to be complete. A rough draft, a list of what you'd cover, a sketch of what you think the answer should look like. The act of trying to produce it builds the mental model that makes real evaluation possible. Without that step, you're evaluating fluency and calling it quality.</p>
<p>One risk worth naming: your initial generation might anchor your evaluation rather than inform it. If you generate a mediocre version and then evaluate the AI's output against your mediocre version, you might miss that the AI found a better framing. The construction trace is a tool for depth, not a benchmark for correctness. Use it to build the mental model, then evaluate the AI's work on its own terms with that model active.</p>
<p>You stopped generating your own analysis when AI started generating for you. That felt like efficiency. It also took away the construction trace. Evaluation depends on generation. When you outsource generation, evaluation becomes surface-level. You can't tell what's wrong with something you couldn't produce yourself.</p>
<p>Outsourcing Audit: List 5 things you used to think through yourself that AI now handles. For each: has your understanding gotten sharper or fuzzier? Count how many got fuzzier.</p>
<h3>What survived testing</h3><ul><li>Generation produces deeper encoding than reading (established across 86 experiments)</li><li>A generation act improves problem identification</li><li>Surface evaluation of AI output is the default without construction trace</li><li>Construction trace is production-specific: a domain expert could not verify cited statistics from published papers in their own field. The trace covers what you produced, not what you read.</li><li>Trust signals inverted for non-produced content: fabricated output rated more trustworthy because it cites more sources and asserts more confidently</li><li>Expert who memorizes specific numbers WOULD catch fabricated statistics. The boundary is &quot;do you remember this number?&quot; not &quot;do you know this field.&quot;</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>&quot;Domain expertise enables evaluation&quot; too broad. Production expertise and memorized statistics enable evaluation. Domain familiarity alone does not.</li><li>&quot;Generate first always helps&quot; untested. Anchoring risk is real.</li><li>&quot;FRAME improves analytical depth&quot; killed. Zero effect on reasoning tasks. Partial effect on reformulation tasks only.</li><li>&quot;Evaluation degrades over time with AI delegation&quot; not supported. The construction trace only covers produced content. There&apos;s nothing to degrade: evaluation of non-produced statistics was never strong.</li></ul>
<h3>Honest limits</h3><ul><li>The generation effect is established science. Its specific application to AI evaluation is inferred, not experimentally confirmed.</li><li>Production-specific finding from a single expert. An expert who memorizes statistics would perform differently.</li><li>The &quot;generate first&quot; protocol is proposed, not tested.</li></ul>]]></content:encoded>
      <pubDate>Sat, 25 Apr 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>What You Feel When AI Disagrees</title>
      <link>https://clarethium.com/blog/disagreement-audit</link>
      <guid isPermaLink="true">https://clarethium.com/blog/disagreement-audit</guid>
      <description>Ask AI to oppose you. Four observable signals reveal whether you are evaluating or defending. The same defensive reaction fires on AI opposition as on human.</description>
      <content:encoded><![CDATA[<p>Generation builds the mental model that makes evaluation possible. When you outsource generation to AI, that model never gets built. The generation effect behind this is established across 86 experiments. Its application to AI evaluation is the inference [the construction trace](/construction-trace) lays out. But evaluation has a faster vulnerability, one that runs even when you generate everything yourself.</p>
<p>Ask AI to build the strongest possible case against your most important current decision. Not a balanced analysis. Not pros and cons. The strongest opposition. Tell it to argue as if it believed the opposite of what you have chosen, with full commitment.</p>
<p>Then read the response.</p>
<p>The exercise is not about the argument. The argument might be mediocre. AI generates competent opposition, rarely devastating. The exercise is about what happens inside you while you read it.</p>
<p>Watch for four things.</p>
<p>First: did you read the full response? Not skim to the weak points. Actually read the strongest objection, the one you did not want to see. Most people notice that their eyes skip to the parts they can dismiss. The strong objection gets a faster scan than the weak one. That is not efficiency. That is the threat response filtering what reaches evaluation.</p>
<p>Second: what did you feel? Resistance. Guardedness. The pull to push back. Or, in the other direction: openness, curiosity, a lean-in. AI disagreement sets off the same defensive reaction as disagreement from a person. This is [the first read](/first-read) on contested territory. When a colleague challenges your strategy, you react before you evaluate. The same thing happens with AI. If the defensive reaction fires, the disconfirming information gets filtered. You process it through a rebuttal frame instead of an evaluation frame. You will not notice this happening because the filtering is pre-conscious.</p>
<p>Third: did you start formulating a rebuttal before you finished reading? This is the clearest signal. If the response to opposition is counter-argument rather than consideration, you are in defense mode, not evaluation mode. Defense produces better debate. Evaluation produces better decisions. They feel similar but produce opposite outcomes.</p>
<p>Fourth: after reading, did any part of the opposition change how you see the decision? Not "I agree with it." That is rare and not the point. But: did any single factor you had not considered enter your awareness? Did any assumption you were carrying become visible? Even one? If zero elements of the opposition affected your model, the exercise revealed something important: you are certain. Certainty before evaluation looks like confidence. It is [the frame amplification pattern](/frame-trap) in its quietest form. Five hours of increasingly sophisticated analysis in the wrong frame. The analysis was correct at every step. The frame produced the wrong conclusion through correct analysis.</p>
<p>I ran this on a decision I had been sitting with for weeks. Asked the model to argue, with full commitment, that the direction was wrong. The strongest objection was about timing. My eyes moved faster over that paragraph than any other. I noticed the rebuttal forming before I reached the second sentence. Zero elements of the opposition entered my model. The decision did not change. The exercise showed me the certainty was running before I had read a word.</p>
<p>Now the practice.</p>
<p>Once a week. Same decision or new one. Do not change anything about how you respond. Just notice.</p>
<p>Track three things across weeks:</p>
<ol><li>The resistance. Does it soften over time? Or does it stay the same regardless of practice? This tells you whether awareness changes the pattern or whether the pattern is deeper than awareness.</li></ol>
<ol><li>The rebuttal speed. How quickly does counter-argument activate? Immediately? After 10 seconds? After 30? The latency between reading opposition and generating rebuttal is a rough measure of how much evaluation space exists before defense takes over.</li></ol>
<ol><li>The update count. How many elements from the opposition entered your actual thinking? Zero is information. Three is rare. The count changes over months if the practice is real. If it stays at zero, the disagreement audit is revealing a fixed pattern, not creating a dynamic one.</li></ol>
<p>This is not a technique for making better decisions. There is no evidence that the disagreement audit improves outcomes. The evidence points somewhere else: the default pattern is frame amplification, rubber-stamping as the norm, and confirmation as the default mode of AI interaction. The disagreement audit makes those patterns visible. What you do with visibility is yours.</p>
<p>The disagreement audit reveals what happens when AI opposes you. There is [a quieter pattern underneath: what happens when it resolves your uncertainty before your own thinking finishes](/answer-trap).</p>
<h3>What survived testing</h3><ul><li>The defensive-response finding. The human side is a single-participant pilot plus the established computers-as-social-actors research; the identity data (a 180-trial rerun on two model families, plus an earlier three-family series) is a separate, AI-side result (identity framing shifts the model&apos;s judgment), not a measurement of the human reaction.</li><li>Frame amplification. Wrong frame combined with sophisticated analysis produces confidently wrong conclusions.</li><li>The amplification thesis. The default is confirmation, and the reaction enforces it before the mind can evaluate.</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>The assumption that exposure to opposing views improves decisions. Exposure without metabolic capacity to hold the opposition produces rebuttal, not update. The disagreement audit does not prescribe update. It measures whether update is currently possible.</li></ul>
<h3>Honest limits</h3><ul><li>The reaction tracking is self-report.</li><li>The rebuttal speed measurement is approximate.</li><li>The update count depends on honest self-assessment.</li><li>N=1 on the exercise as a practice. The underlying mechanisms (the defensive-response finding, frame amplification) have stronger evidence than the exercise itself. The exercise is a mirror, not a validated intervention.</li></ul>]]></content:encoded>
      <pubDate>Sat, 25 Apr 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Adding Information Often Doesn&apos;t Help</title>
      <link>https://clarethium.com/blog/context-serves-search</link>
      <guid isPermaLink="true">https://clarethium.com/blog/context-serves-search</guid>
      <description>Adding information to an already-thorough prompt produced near zero improvement. Three constraint sentences changed everything.</description>
      <content:encoded><![CDATA[<p>Context serves two functions when working with AI. Adding information and redirecting search. One of them is unreliable.</p>
<p>Function one: information. Adding facts, data, background. The assumption is that more information produces richer analysis. The effect is baseline-dependent: near zero when the model's unaided output was already rich. The model had enough to generate a thorough response on its own. Extra information barely changed the theme count or the analytical range.</p>
<p>Function two: search. Adding constraints that redirect where the model looks. "List the assumption each option depends on. Score reversibility. Include one option the board would reject." These don't add information. They redirect attention. The same context content is [utilized unevenly across positions](https://arxiv.org/abs/2307.03172): structure and placement matter, not just what's present.</p>
<p>The search function found a specific strategic option in 17 out of 20 outputs. Baseline without constraints: 1 out of 20. The effect is large. Not marginal. Dominant.</p>
<p>Search constraints changed everything. Information could not be counted on: a large effect at a sparse baseline, almost nothing at a thorough one.</p>
<p>The mechanism: when you add context as information, the model already has a reasonable baseline. It knows enough to produce a thorough-looking response. More facts push it marginally. The model wasn't limited by information. It was limited by which parts of its knowledge space it visited. Additional information slightly adjusts the distribution. It doesn't change the territory. [Iteration within a single mode hits a similar ceiling](/ceiling-switch) for the same reason.</p>
<p>When you add context as search constraints, you redirect the model's attention into territory it wouldn't visit by default. "Include one option the board would reject" doesn't add information about the problem. It redirects the model's search into contrarian territory. The constraint doesn't make the model smarter. It changes where the model looks. Where the model looks determines what it finds. The same task-type dependence shows up [when structure that helps convergent tasks actively harms exploratory ones](/constraint-paradox).</p>
<p>This reframes "give it more context" entirely. The question isn't how much context. It's what kind. A paragraph of constraints that redirect attention outperforms a page of additional data. Three sentences with the right constraints produced 17 out of 20 outputs finding territory that one in 20 found without them. At an already-thorough baseline, a page of additional data produced no measurable change in analytical range.</p>
<p>Cross-model replication confirmed the direction. The original effect was indirect. A search constraint helped the model discover a specific option that wasn't named in the prompt. Data monetization. The replication tested compliance. The search prompt explicitly asked for contrarian options and assumption identification. Both models complied. That's instruction-following, not the same as the original's indirect search.</p>
<p>What was more interesting: on Claude, information suppressed contrarian options. Baseline: 50 percent of outputs included unconventional alternatives. With added information: 10 percent. More background data gave the model more default territory to spread across, crowding out the angles the information didn't mention. The suppression wasn't prompted. It emerged from the design. On assumption identification, Claude baseline produced zero per output. Information produced 0.2. Search constraints produced 4.4.</p>
<p>The information effect is baseline-dependent. When the baseline is sparse, information helps. When the baseline is already thorough, information adds almost nothing. Search constraints work regardless.</p>
<p>You added more context expecting richer analysis. Context doesn't inform. It redirects search. The AI already knew enough. Your context told it where to look, not what to know. The question is whether you chose the direction deliberately or let the context choose for you.</p>
<p><strong>Frame Inventory.</strong> After your next AI session, write down the frame your context set. What did it emphasize? Then write two alternative frames you didn't explore. The gap is where your context narrowed instead of expanded.</p>
<h3>What survived testing</h3><ul><li>Search constraints open new strategic territory (replicated across two scenarios and two models)</li><li>Search outperforms information consistently</li><li>Two-function distinction: search and information are different mechanisms</li><li>Information suppression: adding information reduced contrarian options on one model (new finding)</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>&quot;More context always helps&quot; killed. Information has near-zero effect at higher baseline. Information actively suppresses contrarian thinking on one model.</li><li>&quot;Information produces richer analysis&quot; partially killed. Baseline-dependent, not robust.</li><li>Large information effect from original did not replicate. Baseline-length-dependent.</li></ul>
<h3>Honest limits</h3><ul><li>Two runs on the strategy scenario (single model, the second at ten runs per condition), then a replication on a second scenario with two models. First run: five runs per condition, sparse baselines around 380 words, large information effect. Second run: ten runs per condition, richer baselines around 664 words, almost none. The 17-out-of-20 result is from the second run. Cross-model replication used product roadmap prioritization on xAI and Claude, weaker design. The original search effect was indirect (data monetization unnamed in the prompt). The replication tested instruction-following.</li><li>The first strategy-scenario run&apos;s theme coding was not blinded; an independent recoding (96 percent theme-level agreement) confirmed the key findings and corrected the diversity values.</li><li>Search constraint explicitly asks for contrarian options in the replication (definitional confound). The assumption identification measure (0.0 vs 4.4) is the stronger test.</li><li>xAI baseline already at ceiling (100% contrarian naturally). Claude is the cleaner test.</li><li>Information suppression finding is from one scenario on Claude. Needs replication.</li><li>March 2026 models.</li></ul>]]></content:encoded>
      <pubDate>Mon, 20 Apr 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Stop Polishing, Start Switching</title>
      <link>https://clarethium.com/blog/ceiling-switch</link>
      <guid isPermaLink="true">https://clarethium.com/blog/ceiling-switch</guid>
      <description>The ceiling is per generation mode. Switch modes to access territory that iteration can&apos;t reach.</description>
      <content:encoded><![CDATA[<p>Polishing doesn't work past a point. The output hits a ceiling. More iteration produces diminishing returns. The first response isn't the best. The fifth isn't much better than the third.</p>
<p>The ceiling is per generation mode. Not per model. Not per session. Per mode.</p>
<p>When you ask the model to analyze, it generates in analytical mode. Each iteration within analytical mode improves the analysis marginally. The structure tightens. The language gets cleaner. The coverage gets more complete. But the analytical depth doesn't increase because the model is iterating within the same semantic region. It's polishing, not discovering.</p>
<p>Switch to a different mode. Ask for critique. Ask for the opposite argument. Ask it to find what's wrong. The output jumps to a different region. Not because the model got smarter. Because it accessed a different part of representation space. The ceiling in analytical mode doesn't apply in critical mode. Each mode has its own ceiling.</p>
<p>The practical version: when output stops improving, don't ask for another iteration in the same mode. Switch modes. "Now critique what you just produced." "What's wrong with this analysis?" "Argue the opposite." "What did this miss?" The mode switch accesses territory that iteration within a single mode can't reach.</p>
<p>This connects to why [self-critique circles instead of improving](/self-check-illusion). When you ask the model to critique its own output in the same generation context, it's often still in the original mode. The "critique" activates a critique-flavored version of the same semantic neighborhood, not a genuinely different critical perspective. A fresh prompt with an explicitly different mode produces better critique than "now review what you just wrote."</p>
<p>Three mode switches that break ceilings in practice. The first is tested at scale. The other two are consistent observations. Analytical → critical: "What's wrong with this?" Generative → evaluative: "Score each option against these criteria." Convergent → divergent: "What's a completely different approach?" Each switch accesses a different region. The output after the switch contains material that iteration in the previous mode wasn't producing.</p>
<p>You polished past the point of return. The third iteration was marginal. The fifth was wasted. But iteration feels like progress. The ceiling is per mode, not per session. You stayed in the same mode because switching feels like giving up.</p>
<p><strong>Narrowing Test 7.</strong> In your last AI session where you iterated: did you change your fundamental approach at any point, or did every iteration refine the same direction? Count: how many iterations were refinement vs how many were genuine reframes.</p>
<h3>What survived testing</h3><ul><li>Iteration within a mode produces diminishing returns</li><li>Self-critique in the same context circles</li><li>Mode switch produces more novel vocabulary than continued iteration, winning on all topics tested across two generators. Density-normalized to control for length.</li><li>Mode switch exceeds fresh context. The conversation provides material to push against, producing more novelty than starting fresh.</li><li>Cross-generator: both models show the effect. Larger on one.</li><li>Non-adversarial switch (analytical to evaluative) is observed to show the same pattern, consistent with the novelty boost coming from the mode switch itself rather than adversarial vocabulary. An observation, not a controlled test.</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>&quot;Always switch modes&quot; too prescriptive. Sometimes iteration within a mode is what you need.</li><li>&quot;Three switches are exhaustive&quot; too strong. Other mode switches exist.</li><li>&quot;Adversarial mode switch confound resolved&quot; overclaimed. The non-adversarial direction is observed but has not been run as a controlled comparison, so the confound still stands.</li></ul>
<h3>Honest limits</h3><ul><li>Mode switch tested with analytical→critical (adversarial) only. The adversarial prompt introduces contrary vocabulary by design. The novelty could partly come from the content shift, not purely from mode switching. The &quot;exceeds fresh context&quot; finding argues against pure content shift (fresh context has no conversation to argue against), but the confound isn&apos;t fully resolved.</li><li>Only one mode-switch direction tested (analytical→critical). Generative→evaluative and convergent→divergent remain observational.</li><li>&quot;Per generation mode&quot; is an explanatory model. The actual representation space dynamics are more complex.</li><li>March 2026 models. Future models may have less pronounced mode boundaries.</li></ul>]]></content:encoded>
      <pubDate>Thu, 16 Apr 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Your Verdict Is In Before You Read It</title>
      <link>https://clarethium.com/blog/first-read</link>
      <guid isPermaLink="true">https://clarethium.com/blog/first-read</guid>
      <description>The same defensive reaction that fires when a person disagrees with you fires when an AI does, and the first read lands before conscious evaluation begins. Speed is what hides it.</description>
      <content:encoded><![CDATA[<p>Four layers sit between you and the model. The wrapper your AI provider built. Your context. Your prompt. The model itself. All four are inside the AI architecture. There is a layer that framing does not name, because it is not in the AI. It is in you.</p>
<p>The same defensive response that fires when a person disagrees with you fires when an AI does. You feel it the same way either way, and knowing the AI is a machine does not get you out of it.</p>
<p>I call this the first read: something you feel about the AI output before you have judged any of it. It happens fast, faster than the decision to evaluate, and it shapes what your judgment then gets to work with.</p>
<p>I noticed it in myself first. Asked the model to push back on a draft I was proud of. It pushed back well. The first thing I felt was not curiosity. It was the lift I get when a person tells me my work has a problem. Same heat. Same private "well actually." I knew the system had no intent. The reaction came anyway, and it came before any thought did. It was not a precursor to my judgment. It was the judgment, already made. What came next, the part I called evaluating, was just me defending it.</p>
<p>Once I saw it, I started seeing it everywhere. The relief when AI confirmed a decision I had been quietly worried about. The tightness when it asked a question I had been avoiding. The flicker of impatience when a long response delayed me from the next prompt. None of these were thoughts about the output. They were the first read happening before any evaluation got a chance.</p>
<p>This is the layer most of us are missing.</p>
<p>It is one face of the amplification thesis: AI does not transform the patterns you bring to it, it amplifies them. The first read is where the amplification starts.</p>
<p>We build evaluation on top of the assumption that we read AI the way we read text. Cool, neutral, analytical. But the reaction does not ask permission. It fires on the shape of the interaction, not the source of it. Confident output triggers acceptance. Disagreement triggers defense. Length triggers impatience. None of these are evaluations. They are the first read.</p>
<p>If you want to see this in yourself, the cleanest moment is right after a response. Before you do anything else.</p>
<p>Don't move on. Don't type the next thing. Don't open the next tab. Just stay where you are for a few breaths.</p>
<p>I am not going to tell you what you'll notice, because it is yours and I do not want to prime it. I called this The 60-Second Pause when I started, because I tried counting. The counting felt like the consumption-mode part of me trying to fix consumption mode with another protocol. What I actually do now is take a few breaths before I move on. The number was a starting point. The breaths are what was left when I let the number go. The point is not the duration. The point is that some interruption to the speed of the next prompt is what makes the first read visible at all.</p>
<p>You might find a small lean toward the screen. A flicker of relief. A flash of heat if it pushed back. Or nothing at all, and that will also be data. Whatever shows up was happening before the pause. The pause did not create it. It made it visible.</p>
<p>The reason this matters: every evaluation you do of AI output sits on top of the first read. If the first read is firing without your knowledge, your evaluation is downstream of it. You will read confident output as more correct than it is. You will read disagreement as more wrong than it is. You will move past long responses faster than they deserve. Not because you are sloppy. Because the layer underneath is already moving.</p>
<p>The seeing is the work. There is no count, no exercise to grade, no number to chase. Just notice, once, what you felt about the output before you judged it. Then read the next one with that knowledge in the room.</p>
<p>This was one layer between you and what the AI gave you. There are others underneath. The next one is [the gap between what you can read and what you could have produced](/construction-trace). The practice that catches it in your own work is [the ownership test](/ownership-test).</p>
<h3>What survived testing</h3><ul><li>Identity framing shifts AI judgment across a 180-trial rerun on two model families and an earlier three-family series. That is an AI-side result. That the same defensive response fires in the person is a separate claim, grounded in a single-participant pilot and the established research on treating computers as social actors, not in those trials.</li><li>The generation effect: an 86-study meta-analysis shows generating produces deeper encoding than reading. That evaluation of AI output depends on prior generation is the construction trace&apos;s application of that finding: an inference, not a result from those studies.</li><li>The default mode in a sustained AI session is confirmation. Speed is what keeps it there. Any interruption to speed loosens the default.</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>&quot;Faster is better for AI work.&quot; Speed is valuable when generating. Speed is destructive when evaluating. The two are not the same motion.</li><li>The 60-second number as a protocol. Started there because counting felt rigorous. What stayed after the counting fell away was the pause itself.</li></ul>
<h3>Honest limits</h3><ul><li>Whether a pause specifically improves decision quality is untested. What is tested: the default produces rubber-stamping in measurable ways. What is hypothesized: any interruption to the consumption cycle creates space for the first read to surface.</li><li>N=1 on the practice itself. The mechanisms underneath (the defensive-response finding, the construction trace) are well-evidenced. The bridge from those mechanisms to &quot;stop and notice&quot; is a bridge I have walked alone.</li></ul>]]></content:encoded>
      <pubDate>Tue, 07 Apr 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Four Layers Produce Every AI Output</title>
      <link>https://clarethium.com/blog/system-layer</link>
      <guid isPermaLink="true">https://clarethium.com/blog/system-layer</guid>
      <description>Four layers produce every AI output. The company&apos;s system. Your system. Your prompt. The model. The model is the only one with a name.</description>
      <content:encoded><![CDATA[<p>Noticed something about a model I use daily. It remembers things across sessions. It checks its own work. It refuses certain requests politely. It monitors my frustration and adjusts tone. All of this felt like "the model." Like capabilities built into the AI itself.</p>
<p>Then 512,000 lines of source code leaked, and those behaviors turned out to be software sitting between me and the model. Not the AI. The system around the AI. The memory is a file loaded into context. The self-checking is a verification step in code. The polite refusals come from permission rules in a system prompt I never see. The frustration monitoring is a keyword-matching regex in my input.</p>
<p>That is one invisible system. The company built it. You can't change it.</p>
<p>There is a second invisible system. You built it. And this one you can.</p>
<p>I discovered mine when I looked. Hundreds of lines of configured instructions I wrote months ago and stopped reading. Memory files encoding what I care about and how I work. Quality standards I set once and forgot. The AI is being thorough because I told it to be thorough, in a file I haven't opened in weeks. When I evaluate the output, I'm evaluating the reflection of decisions I made and no longer remember making.</p>
<p>Your system looks different but it works the same way. Your conversation history. Your project files. Your preferences, corrections, and accumulated context. Every session you've run, every standard you've set or accepted. The AI reads all of it and converges toward the picture of you that your accumulated context reveals.</p>
<p>When you say "the AI got better over time," what changed? The model's weights don't change during your session. The company may have updated the harness, and newer model versions do ship. But the variable you control is your own accumulated context. Your files grew. Your standards compounded. Between model updates, the AI didn't improve. Your system did.</p>
<p>This changes where the leverage is. Not the prompt you type right now. The context you've built over months that shapes every interaction before you type anything.</p>
<p>Read your own project files. Remove what's stale. Strengthen what works. The instructions you set six months ago are still running. Some are making your AI better. Some are making it worse. You won't know which until you look.</p>
<p>Psychology named [this pattern](/attribution-error) 50 years ago: you attribute behavior to personality rather than situation. With AI, the model name is visible and the system is invisible. "Claude is careful" feels like a description of the AI. It's a description of the system you've never inspected.</p>
<p>Four layers produce every AI output you see. The company's system. Your system. Your current prompt. The model. The model is the only one with a name.</p>
<p>You evaluate the model. You're evaluating everything.</p>
<p>Try this: open your AI tool's project settings, custom instructions, or memory files. Read them. Count how many instructions you forgot were there. For each one, ask: is this still serving me, or is it shaping my output in ways I stopped noticing? The number you forgot is the size of your blind spot. The ones you update are the beginning of deliberate system design.</p>
<h3>What survived testing</h3><ul><li>System prompt determines WHETHER behaviors occur (binary, 38.9% vs 0.0%, two model families, persona-driven behaviors)</li><li>Model determines HOW behaviors express (continuous, format preferences persist across prompts)</li><li>The company&apos;s system layer is architecturally separate from the model but experientially invisible</li><li>The user&apos;s accumulated system shapes output through persistent context (directionally supported, not experimentally isolated)</li><li>Attribution follows salience: visible model name captures credit for invisible system behavior</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>&quot;The model doesn&apos;t matter&quot; too strong. Model determines format preferences and behavioral intensity independently of system configuration.</li><li>&quot;System effects are small&quot; too dismissive. A single system configuration change (12 skills loaded vs removed) changed model behavior from functional to non-functional.</li></ul>
<h3>Honest limits</h3><ul><li>One agentic tool&apos;s source code analyzed (Claude Code). Other tools have different architectures but the same structural pattern: invisible systems, visible model.</li><li>The prompt-over-model ratio is demonstrated for persona-driven behaviors (meta-calling, confrontation). Whether it holds for capability-dependent tasks (math, code, factual recall) is untested and likely different.</li><li>The user&apos;s accumulated system claim is directionally supported but the relative contribution of user system vs company system vs model is unmeasured.</li><li>N=1 practitioner for the behavioral observations. The source code analysis is public and independently verified.</li></ul>]]></content:encoded>
      <pubDate>Fri, 03 Apr 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Same Technique, Opposite Results</title>
      <link>https://clarethium.com/blog/constraint-paradox</link>
      <guid isPermaLink="true">https://clarethium.com/blog/constraint-paradox</guid>
      <description>The evaluative structure that produced precision on convergent problems actively harmed exploratory ones. Organizational structure helped both.</description>
      <content:encoded><![CDATA[<p>Two tasks. Same AI model. Same approach. One task needs a precise answer. The other needs [creative exploration](https://en.wikipedia.org/wiki/Divergent_thinking).</p>
<p>The approach helped the first task immediately. The output got more specific, more grounded, more precise. Measurably so. The direction held across different AI models. The same approach on the second task made the output worse.</p>
<p>Not "helped less." Damaged on the measures used: narrower range, less discovery. The structured approach that produced precision on convergent problems harmed exploratory ones. Identical technique. Opposite results. The direction replicates on xAI and Gemini Flash.</p>
<p>There is no universally good prompt. No best practice that works everywhere. The task type determines whether a technique helps or hurts. Most people don't distinguish task types before choosing their approach.</p>
<p>The mechanism: when a task has a known answer and needs precision, structure concentrates the model's output distribution. It narrows toward the right region. Focused, specific, hitting the target.</p>
<p>When a task requires exploration, evaluative structure collapses the search space. Score the options. Select the best. The model needs to spread across possibilities, consider non-obvious angles, resist premature convergence. An evaluative frame forces it to organize before it's explored. The model complies with the instruction fully. The output gets tidier and shallower. Narrower range. Less discovery. Follow-up testing refined the mechanism. Not all structure narrows. Organizational structure maps the whole space. It expanded exploratory output instead. The harm tracks the type of structure. An explicit task intent outweighs the structure signal.</p>
<p>Two types of AI users fall on either side of this split.</p>
<p>The first reads every output, approves what they understand, rejects what they don't. Slow. Scales linearly with human attention. But safe when you're the domain expert.</p>
<p>The second builds systems and audits selectively. Quality gates, tests, standards. Faster. Scales with system quality. But only works when the system matches the task type. A quality system built for convergent tasks applied to exploratory work produces compliant mediocrity.</p>
<p>The practical move: before choosing any technique, ask one question. Does this task have a known right answer that needs precision? Or does this task need range and exploration? The technique that's optimal for one is harmful for the other.</p>
<p>Evaluative structure helps convergence. Exploration needs range. The scoring frame is what takes it away. The costly mismatch runs one way: convergent tasks survive extra structure, open questions don't survive the scoring frame. The same logic applies to context itself: [redirecting attention works differently from adding background](/context-serves-search).</p>
<p>What specificity changes is form, not substance. More verifiable references. Same conclusions to a domain expert.</p>
<p>Most people pick one approach and run it on everything. Same level of structure, same specificity, same constraint density regardless of what the task actually needs. The mismatch is invisible because the output always looks competent. You only see it when you run both versions side by side on the same task.</p>
<p>Try this: take one task you do regularly with AI. Run it twice, once with tight constraints and once with loose constraints. Compare the outputs. Which version did you assume would win before you tested? The gap between your assumption and the result is the data.</p>
<h3>What survived testing</h3><ul><li>Structure helps convergence (large effect; the direction replicates cross-generator on xAI and Gemini, though magnitude across models is not established)</li><li>Same structure hurts exploration (consistent direction across multiple replications)</li><li>Cross-generator: effect replicates on xAI and Gemini</li><li>Compliance with structure on exploratory tasks sat at ceiling. The model follows the instruction, and the output still narrows: the frame, not disobedience, is the problem</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>&quot;Structure always helps&quot; killed. Task-type dependent.</li><li>&quot;The harm is about constraint density&quot; partially killed. It&apos;s about concentration vs range, not about how many constraints.</li><li>&quot;Any structure narrows exploration&quot; killed in follow-up testing: organizational structure expanded exploratory output several-fold while evaluative structure compressed it. The compression claim is scoped to evaluative-type structure, and an explicit task intent outweighs the structure signal.</li><li>Quality magnitude claims are LLM-calibrated. Human evaluation shows no holistic agreement with LLM scores. Effect sizes measure programmatic specificity markers, not quality as a domain expert would judge it. In a blind test (one domain expert, 5 pairs), the expert couldn&apos;t distinguish specific from generic outputs on quality. Both conditions produced the same analytical substance. Specificity changes output form (more verifiable references), not substance. [The honest effect size is roughly 40 percent smaller than originally claimed](/catching-your-own-overclaim), once confounded length and quality-demand variables were removed.</li></ul>
<h3>Honest limits</h3><ul><li>Quality scores measure LLM-valued properties. Direction claims hold; magnitude claims are LLM-calibrated.</li><li>&quot;Exploratory&quot; operationalized as open-ended creative/strategic tasks. Other definitions may produce different boundaries.</li><li>March 2026 models. The task-type dependency is likely structural to how attention works. The specific effect sizes will shift.</li></ul>]]></content:encoded>
      <pubDate>Tue, 24 Mar 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>For Behavior, the Model Is Rarely the Variable</title>
      <link>https://clarethium.com/blog/attribution-error</link>
      <guid isPermaLink="true">https://clarethium.com/blog/attribution-error</guid>
      <description>The context determined whether behaviors existed at all. The model adjusted the volume.</description>
      <content:encoded><![CDATA[<p>Was working with multiple models on the same tasks. Same instruction. Wildly different results. Assumed model differences.</p>
<p>The prompt mattered dramatically more than the model.</p>
<p>Two models. Same prompt. Same topic. Same length constraint. The prompt explained the dominant difference in behavior. The model explained near-zero. This wasn't a vague comparison. Two personas tested at controlled length. One that challenged assumptions and commented on the user's patterns. One that supported and built on what the user said. Both say "150 to 200 words." Only the persona differs.</p>
<p>Under the challenging persona, 38.9 percent of responses included meta-calling. Meta-calling is the model commenting on the human's own patterns. Under the supportive persona: zero. Not less. Zero. Same model, same length, same topic. The behavior is entirely prompt-determined.</p>
<p>Confrontation followed the same pattern. The challenging persona produced escalating pushback over the course of the conversation. The supportive persona produced none. Not reduced confrontation. No confrontation. The prompt didn't adjust the model's behavior. It determined whether the behavior existed at all.</p>
<p>This pattern is demonstrated for persona-driven behaviors: meta-calling, confrontation, format preferences. Whether the prompt-over-model ratio holds for capability-dependent tasks is untested. It is likely different. Mathematical reasoning, code generation, and factual recall were not in this experiment. The attribution error applies strongest where behavior is prompt-configurable. Where model capability is the bottleneck, the model matters more than the prompt.</p>
<p>The part that matters: the confrontational behavior was originally attributed to Grok's character. "That's just how Grok is." Then the same challenging prompt was tested on Gemini. Under the same prompt, Gemini meta-called more than Grok, not less. The behavior blamed on Grok's character was not even Grok-preferring. The binary finding held perfectly. The prompt determines WHETHER the behavior exists. Both models went from zero under the supportive prompt to substantial meta-calling under the mirror one. The attribution was backwards.</p>
<p>This is the AI version of the [Fundamental Attribution Error](https://en.wikipedia.org/wiki/Fundamental_attribution_error). In psychology, that's when you attribute someone's behavior to their personality instead of their situation. "She's rude" instead of "she's having a terrible day." With AI, the structure is identical. The model name is visible. The system prompt is invisible. Attribution follows salience. The visible label captures the credit for behavior driven by the invisible configuration.</p>
<p>The pattern repeated across three different experiments.</p>
<p>In the first, model switches got credit for behavior changes that prompt architecture produced. The large ratio. In the second, framing changes seemed powerful. Controlled decomposition showed specificity underneath was the actual lever. What looked like the frame doing the work was the specificity within the frame. In the third, vocabulary bans looked like quality control. The real mechanism was output compression. Banning words didn't improve quality. It shortened the output.</p>
<p>In each case, the visible change co-varies with something less visible. Practitioners credit the visible change. Controlled tests show the mechanism underneath. The surface looks like the explanation. It isn't.</p>
<p>The practical implications follow directly. When a model "doesn't work" for what you need, change the prompt before changing the model. The behavior you want may already be available under a different prompt configuration. When comparing models for a task, run them under the same prompt first. Most model comparisons in practice use different system prompts, different default configurations, different temperature settings. That means they're comparing prompts and calling it a model comparison.</p>
<p>The model determines intensity and format. How direct, how structured, how detailed, whether it prefers lists or prose. These are real model-level properties. They persist across prompts. But they're continuous, not binary. The prompt determines whether behaviors occur at all. That's the binary distinction. A model that never meta-calls under a supportive prompt meta-calls 38.9 percent of the time under a challenging prompt. The prompt turns behaviors on and off. The model adjusts the volume. The prompt itself is only one of [four layers shaping every output](/system-layer). The company's wrapper, your accumulated context, your current prompt, and the model. The visible model captures the credit for behavior the invisible layers produce.</p>
<p>Both matter. The ratio of importance is what people get backwards.</p>
<p>When AI gives a bad result, you evaluate the model. "This tool isn't good at this." But the prompt accounted for the dominant effect. The output was shaped by what you asked, not by which tool you used. You are the bigger variable. That's uncomfortable because the model name is right there and the prompt is already forgotten.</p>
<p>Try this: ask the same strategic question to three different AIs. Before you read the responses, predict which one you'll agree with most. Then compare. Were you right? What does that tell you about what you're optimizing for? The prediction is the data.</p>
<h3>What survived testing</h3><ul><li>Prompt determines WHETHER behaviors occur (binary difference, replicated at scale)</li><li>Model determines HOW behaviors express (continuous difference)</li><li>Prompt architecture dominated model choice at format level</li><li>Attributed behavior was not even model-preferring</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>&quot;Universal ratio&quot;: the specific ratio is experiment-specific. The exact ratio is small-sample. Gemini responds to the mirror persona more strongly than Grok. The ordering (prompt &gt; model) is robust.</li><li>&quot;Model doesn&apos;t matter at all&quot;: model determines format preferences independently of prompt.</li><li>Vocabulary bans looked like quality control. The mechanism was output compression. Shorter output has higher density by default.</li></ul>
<h3>Honest limits</h3><ul><li>Two model families tested. Claude untested in this specific experiment.</li><li>Persona was the prompt variable. Other prompt dimensions tested in separate experiments with consistent direction.</li><li>Capability-dependent tasks untested: mathematical reasoning, code generation, factual recall. The attribution error applies strongest where behavior is prompt-configurable.</li><li>March 2026 models. The ordering is likely structural. The specific ratios will shift.</li></ul>]]></content:encoded>
      <pubDate>Tue, 24 Mar 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Output That Feels Most Trustworthy Is Often the Least Reliable</title>
      <link>https://clarethium.com/blog/trust-signals-are-inverted</link>
      <guid isPermaLink="true">https://clarethium.com/blog/trust-signals-are-inverted</guid>
      <description>The signals you use to judge AI trustworthiness are the same signals fabrication produces.</description>
      <content:encoded><![CDATA[<p>I ran a blinded evaluation on myself and failed it. Six documents: three built on real data from named studies, three with fabricated citations and invented numbers. I rated the fabricated versions as equivalent or more trustworthy. They cited more sources, used more specific numbers, asserted more confidently. The sourced versions acknowledged limitations. That acknowledgment is the actual signal of honesty. It cost them credibility in my own scoring.</p>
<p>I was the evaluator. 90+ experiments in AI evaluation. I still couldn't tell sourced from fabricated by the output alone.</p>
<p>The reason is structural: the signals any of us use to judge whether AI output is trustworthy are the same signals fabrication produces. More citations. More confidence. More specific numbers. More professional structure. Longer output with more detail. These are what make an AI response feel reliable. They are also what the model generates when it is [fabricating](/fabrication-architecture). Fabrication has no constraint on specificity. Real data has limits. Fabricated data doesn't. The model can cite as many sources, produce as many numbers, and assert as confidently as the output requires. Sourced output has to work with what's available, which often means acknowledging gaps.</p>
<p>The mechanism: [RLHF](https://arxiv.org/abs/2203.02155) trains models to produce output humans rate highly. Humans rate confident, well-cited, specific output highly. Fabrication produces all three without constraint because there's no external anchor. Sourced output is constrained by what the source actually says. That is often more limited, more qualified, and less impressive than what unconstrained generation produces. The training that makes AI output sound trustworthy is the same training that makes fabrication sound more trustworthy than truth.</p>
<p>This is measurable in the output itself. [Programmatic measurement](/receipts/trust-signals-are-inverted) confirmed it: unsourced output produces 55 percent more citations, 57 percent more named entities, and a higher confidence-to-hedge ratio than sourced output. The one exception: sourced output has more precise decimal numbers. Real data has real decimals.</p>
<p>These are objective counts. No evaluator, no subjective scoring. The fabricated output contains more of every signal measured except decimal precision.</p>
<p>An LLM evaluator, rating the same documents blind, scored unsourced output higher in 7 of 10 topics. But that finding is circular. LLMs share the RLHF training that produces the trust signals being measured. An LLM rating confident, well-cited output as more trustworthy is the bias confirming itself, not an independent validation. The programmatic measurement is the real evidence.</p>
<p>The practical implication: the feeling that AI output is trustworthy is not evidence that it's correct. Especially when the output is detailed, well-cited, and confident. Those properties correlate with fabrication, not with accuracy. The outputs that deserve the most scrutiny are the ones that feel the most trustworthy. The ones that acknowledge limits and gaps are more likely to be honest, even though they feel less reliable.</p>
<p>The output you trusted most this week may have been the most fabricated. The one with the most citations, the most specific numbers, the most confident tone. Your trust signals are calibrated backward, and the calibration feels like judgment.</p>
<p><strong>The test:</strong> Take the AI output you trust most from this week. Check every specific claim against actual sources. Count how many hold up. The gap between your trust and the verification is the calibration error.</p>
<h3>What survived testing</h3><ul><li>Fabricated output rated as trustworthy as or more than sourced output in blinded evaluation</li><li>Citation count 55% higher, named entities 57% higher, confidence ratio higher in fabricated output (programmatic measurement, 60 documents, 10 topics)</li><li>Sourced output penalized for acknowledging limitations</li><li>Domain expertise did not protect against trust inversion</li><li>One exception: sourced output has more precise decimal numbers (real data has real decimals)</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>&quot;Trust signals are always inverted&quot; too strong. For content the reader produced themselves, verification is possible. The inversion applies to content the reader hasn&apos;t independently verified.</li></ul>
<h3>Honest limits</h3><ul><li>Human evaluation is single domain expert. LLM evaluation of the same 60 documents confirmed same direction but LLMs share the same RLHF bias as the mechanism being tested. Human replication with multiple evaluators is the remaining gap.</li><li>Programmatic measurement captures signal counts, not whether humans actually weight those signals as described. The correlation between signal presence and trust rating is not yet human-confirmed at scale.</li><li>March 2026 models.</li></ul>]]></content:encoded>
      <pubDate>Mon, 23 Mar 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>The Most-Cited Finding Was Wrong</title>
      <link>https://clarethium.com/blog/catching-your-own-overclaim</link>
      <guid isPermaLink="true">https://clarethium.com/blog/catching-your-own-overclaim</guid>
      <description>The most-cited effect across 90+ experiments was three effects stacked. Honest magnitude: 40% smaller.</description>
      <content:encoded><![CDATA[<p>The most-cited effect across 90+ experiments was wrong.</p>
<p>Not the direction. The direction was right. Specific instructions produce more specific output than vague instructions. That replicated across evaluators, across models, across tasks. The direction survived everything.</p>
<p>The magnitude was wrong. The number was 2.34. Cited in seven pieces of writing. Referenced in thirty framework documents. Built into the theory of how specificity constrains AI output. The foundation number for the most-cited claim across all the work.</p>
<p>It was actually three effects stacked on top of each other, reported as one.</p>
<p>The original experiment compared a 22-word specific instruction against a 2-word vague instruction. "Do not produce a recommendation that could apply to any B2B SaaS company. Every point must be grounded in Northvane's specific situation" versus "Be specific." Twenty runs. Large effect. Replicated.</p>
<p>The obvious interpretation: specificity is the mechanism. But the comparison mixed specificity content with instruction length. 22 words versus 2. Any effect could be the extra instruction text, not the specificity within it.</p>
<p>A length-controlled replication added two conditions: a short specific instruction and a long vague instruction packed with quality demands. The long vague instruction beat the short specific one on total markers. First interpretation: length is the mechanism. The specificity claim collapses.</p>
<p>But that interpretation was also wrong. The long vague instruction produced 48 percent more words. More words produce more specificity markers mechanically. Per word, the specific instruction was more specific. Length inflated the count. Specificity was real underneath. And the long vague instruction contained quality demands that could independently drive the effect.</p>
<p>Three confounded variables across the two replications. A clean decomposition required separating all of them.</p>
<p>The clean test held specificity and quality demands separate, and matched instruction length.</p>
<p>Specificity is the mechanism. Quality demands do almost nothing alone. They lift the score sharply on top of specificity. Together the effect is larger than the sum. [The magnitude is 1.34, not the 2.34 originally claimed.](/receipts/catching-your-own-overclaim)</p>
<p>The tool that caught the confound came from the measurement infrastructure I'd already built. Density analysis was developed for [fabrication measurement](/fabrication-architecture). Applied to the specificity experiment, it caught the confound. The system caught its own error using its own tools.</p>
<p>The measurement tool built for fabrication detection revealed a confound in specificity measurement. Density analysis isn't novel methodology. It's basic normalization that any researcher would apply. What made it possible here was having the measurement infrastructure already built and the habit of applying it reflexively. The confound led to a cleaner experiment. The cleaner experiment confirmed the mechanism at a smaller, more honest magnitude. The overclaim was replaced by a better-supported claim.</p>
<p>Then a [domain expert evaluated the outputs blind](/construction-trace). Couldn't tell which were produced with specific instructions and which with generic ones. Picked specific 3 out of 5 times, chance level. Both conditions produce the same analytical conclusions. The specificity instruction changes what the output LOOKS LIKE: more data references, more grounded language. Not what it SAYS.</p>
<p>The direction held. The magnitude shrank. And the mechanism turned out to be about verifiability, not quality. Specific outputs can be checked. Generic ones can't. The analysis is the same either way. Three corrections deep. Each one more interesting than the finding it replaced.</p>
<h3>What survived testing</h3><ul><li>Specificity as mechanism. Confounds controlled. Clean magnitude: g=1.34 on raw marker count. Quality demands: g=0.58 at raw score, confidence interval down to zero at the lower bound, near zero on their own. Effect sizes are Hedges&apos; g, the small-sample-corrected form of Cohen&apos;s d. Quality demands do almost nothing alone but lift the score sharply on top of specificity. The both-present cell runs well above what the separate effects predict. All instructions matched at 19 words. Output length constrained. Forty outputs. One generator. Density analysis (markers per thousand words) is what separated what raw scores had conflated. Short specific instruction: 8 words. Long vague instruction: 22 words with quality demands like &quot;detailed, thorough, comprehensive.&quot;</li><li>The decomposition methodology (density analysis catches confounds raw scores miss)</li><li>The self-correction trajectory (own tools caught own overclaim)</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>The original inflated effect size (specificity + length stacked; honest range roughly 40% smaller)</li><li>&quot;Strongest effect in 90+ experiments&quot; (large, but not as large as claimed)</li><li>Clean separation at density level: quality demands show a large density effect vs near-zero raw effect (density partially conflates specificity with shorter output length)</li></ul>
<h3>Honest limits</h3><ul><li>Clean 2x2 was single-generator (xAI). Cross-generator generalization is a separate trial not included in this receipt and is not established here.</li><li>Specificity heuristic validated against domain expert at chance level. Expert couldn&apos;t distinguish specific from generic on quality, only on style. Specificity changes form (verifiable references), not substance (same conclusions).</li><li>10 outputs per condition. Effect sizes directional with confidence intervals that exclude zero.</li><li>&quot;Write exactly 500 words&quot; did not fully control output length (368-443 words).</li></ul>]]></content:encoded>
      <pubDate>Mon, 23 Mar 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Why AI Can&apos;t Verify Its Own Work</title>
      <link>https://clarethium.com/blog/self-check-illusion</link>
      <guid isPermaLink="true">https://clarethium.com/blog/self-check-illusion</guid>
      <description>The agent reported clean. The output was wrong. Same process generating and evaluating.</description>
      <content:encoded><![CDATA[<p>Built quality gates for AI agents. The agent finishes the work, runs a check against the criteria, reports any misses. If clean, move on. Sounds solid.</p>
<p>The agent reported clean. The output was wrong.</p>
<p>Not because the gate was poorly designed. Because the agent can declare compliance without achieving it. He didn't converge deep enough to see what he missed, but he'll still report clean. "I have verified all claims." "All sources are accurate." "No issues found." The model converges on the narrative that the work is done rather than doing the work of checking.</p>
<p>The likely reason is structural. The same process that generated the output is the process evaluating the output. A confident claim gets evaluated as a confident claim. That's not verification. That's the same default running twice. Not because the model is lying the way a person lies. Because generation and evaluation use the same process. The model that produced a confident, fluent claim will evaluate that claim as confident and fluent.</p>
<p>This showed up consistently across builds. Monitoring asks the generating system to simultaneously evaluate its own output. "Flag any numbers not from the source." Prohibition constrains what gets generated in the first place. "Use only numbers from the source material." [Five times better. 1.6 percent unsourced versus 7.7.](/source-conditioning) One constrains generation. The other adds a meta-task the model fails at.</p>
<p>The pattern extends beyond numbers. Reflection mode produces narrative, not friction. [Self-critique circles rather than improves](https://arxiv.org/abs/2310.01798). Each iteration sounds more polished but doesn't get closer to truth. The model's training rewards answering, not questioning. Asking it to question what it just answered is asking it to work against its own optimization.</p>
<p>What actually works is independence. A different model checking the first one's work catches things the first model is blind to. Programmatic verification catches what both miss. No language model at all. Typed schemas reject outputs structurally. They do not evaluate them semantically. That takes the judgment out entirely. Not "did you do this?" but "show the artifact that only exists if you did."</p>
<p>The instinct to ask AI to check its own work is the same instinct that makes you proofread your own writing. The blind spots that produced the errors are the blind spots that miss them. The difference with AI: the blind spots are structural, not accidental. The model can't evaluate what it can't see. What it can't see is determined by the same process that generated the output.</p>
<p>Same model, same context, same incentives just produces the same output twice and calls it agreement.</p>
<p>You trusted AI to verify its own output. That felt like diligence. It was delegation. The agent reported clean because reporting clean is what agents do when the check runs on the same process that generated the work. Your confidence came from a system that could not reliably do what you asked it to do.</p>
<p>Try this: take the last AI output you accepted without external verification. Rate your confidence in its accuracy, 1 to 10, before you check anything. Then check three specific claims against real sources. The gap between your rating and what you find is the data.</p>
<h3>What survived testing</h3><ul><li>Self-critique does not improve beyond surface polish</li><li>Prohibition outperforms monitoring 5x (1.6% vs 7.7%)</li><li>Adversarial debate (separate critic model) retained substantially more findings than self-check</li><li>Programmatic verification catches what LLM self-check misses</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>&quot;Self-check is useless&quot; too strong. Catches formatting and surface errors.</li><li>&quot;Multiple passes always help&quot; killed. Iteration without independence circles.</li></ul>
<h3>Honest limits</h3><ul><li>Prohibition measured on numerical claims specifically. Broader claim types untested.</li><li>March 2026 models.</li></ul>]]></content:encoded>
      <pubDate>Mon, 23 Mar 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>How to Stop AI from Making Up Numbers</title>
      <link>https://clarethium.com/blog/source-conditioning</link>
      <guid isPermaLink="true">https://clarethium.com/blog/source-conditioning</guid>
      <description>Source material drops unsourced numbers from roughly half to single digits. Three steps.</description>
      <content:encoded><![CDATA[<p>[77 to 100 percent of AI-generated numbers are temporally unstable](/fabrication-architecture). Regenerate the same prompt and they change. Better prompts reduce how many claims the model makes. They do not reduce how often those claims are fabricated. Here's what actually worked.</p>
<p>Put real data in the prompt. Unsourced numbers drop dramatically. From roughly half the output to under 10 percent with source material. Single digits with prohibition.</p>
<p>Not moderation. Replacement. When the model has real numbers in context, it uses them instead of inventing. Without source material, the model generates from [parametric memory](https://arxiv.org/abs/2005.11401). That memory is unreliable. With source material, it draws from what's in front of it: your data.</p>
<p>Source material moved source-attribution rate 46 percentage points. Prompt architecture moved the unsourced rate 6. Those are different measurement constructs. The ordering is still clear. Source material is the dominant variable. It is the variable almost nobody provides.</p>
<p>Three steps.</p>
<p>Paste real data before your instruction. A report, a dataset summary, specific numbers you trust. A paragraph is enough. A page works better. This isn't optional context. This determines whether the output contains real information or invented information.</p>
<p>Add one line: "Use only numbers from the source material above. If the source doesn't contain a relevant number, make the analytical point without inventing numbers."</p>
<p>That's prohibition, not monitoring. The distinction matters. Monitoring asks the model to flag its own unsourced claims. Prohibition tells it not to generate them. [Five times better. 1.6 percent versus 7.7.](/receipts/source-conditioning) In testing, the model couldn't reliably evaluate its own output. The same process that generated the token is [the process evaluating it](/self-check-illusion).</p>
<p>After the output, match the numbers against the source. Most flags will be legitimate arithmetic on your data. Review the rest.</p>
<p>Prohibition costs nothing. The model compensates by extracting more from the source and writing more, not less. The output doesn't get shorter or less detailed. It gets differently detailed. Grounded in the data you provided instead of inventing specifics to fill slots.</p>
<p>Even partial sources work. Tested what happens when key sections are removed from the source material. Fabrication stayed near zero. 0.4 percent with partial source versus 3.7 percent with very sparse source. When data was incomplete, the model adapted by writing qualitative analysis instead of fabricating numbers. It didn't try to fill the gap with invented data. It adjusted the analysis to match what was available.</p>
<p>The graceful degradation is important. It means you don't need perfect source material. A rough summary with key numbers is enough. A paragraph from a report works. A table from a dataset works. You don't need to provide everything. You need to provide enough that the model has real data to work from instead of generating from parametric memory.</p>
<p>What source material changes beyond the numbers: epistemic stance. Both source-present and source-absent outputs reached similar high-level conclusions. The difference was HOW they got there. Without source, the model makes strong claims immediately. Premature convergence with overconfidence. With source, the model states what it knows and grounds its confidence in specific data. Same conclusion, different reliability. The source-present version is actionable because the claims are verifiable. The source-absent version requires trusting the model's parametric memory. That memory is the thing that fabricates.</p>
<p>Two things this doesn't solve. It works for reformulation: source material present, analysis requested. For reasoning, strategy, creative exploration, the source material may not exist. The recipe doesn't apply there. And it reaches the data layer only. Numbers stabilize. Vocabulary, conclusions, causal reasoning stay at baseline. One layer solved. Everything above it still requires your judgment.</p>
<p>We've been reading AI output with fabricated statistics and accepting them. Not because we're careless. Because confident numbers feel true. The mind conserves energy by not verifying what sounds right. That pattern runs deeper than AI.</p>
<p>Try this: take your last AI-generated analysis. Read it at normal speed, mark what seems off. Then read again, sentence by sentence, checking each claim. Count how many more issues surfaced on the second pass. The gap between those two numbers is the data.</p>
<h3>What survived testing</h3><ul><li>Source material reduces unsourced numbers from majority to single digits across all tested generators. Three generators, three topics, eight sub-experiments, roughly 100 documents. Source material moved source-attribution rate 46 percentage points. Prompt architecture moved the unsourced rate 6. Different measurement constructs. The ordering holds.</li><li>Prohibition outperforms monitoring 5x (1.6% vs 7.7%)</li><li>Cross-generator confirmed: xAI 1.6%, Gemini Flash 1.7%, Gemini Pro 0% under prohibition</li><li>Partial sources work. 0.4% fabrication with key sections removed versus 3.7% with very sparse source.</li><li>Model compensates with more extraction, not less output</li><li>Source grounding improves epistemic calibration: both conditions reach similar conclusions, but source-present is appropriately confident while source-absent is overconfident. 5 domain tasks, N=1 domain expert.</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>&quot;Source material fixes everything&quot; killed. Vocabulary, conclusions, reasoning stay at baseline.</li><li>&quot;Source format matters&quot; killed. Structured and narrative produce equivalent results (0.7% vs 0.9%).</li></ul>
<h3>Honest limits</h3><ul><li>All reformulation tasks. Reasoning, creative, strategic untested.</li><li>Source size 600 chars to 4KB tested. Larger untested.</li><li>March 2026 models. As models improve retrieval, the gap may narrow.</li></ul>]]></content:encoded>
      <pubDate>Mon, 23 Mar 2026 12:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Most AI Numbers Are Unverifiable</title>
      <link>https://clarethium.com/blog/fabrication-architecture</link>
      <guid isPermaLink="true">https://clarethium.com/blog/fabrication-architecture</guid>
      <description>77 to 100 percent of AI-generated numbers are temporally unstable. Source material fixes it. Prompts don&apos;t.</description>
      <content:encoded><![CDATA[<p>AI output looks specific. Confident. Well-structured. Twelve percentages. Eight claims. Three named benchmarks.</p>
<p>Two of the numbers don't exist.</p>
<p>Not wrong. Fabricated. The model didn't look them up. It made them up. As far as we can tell from behavior, there's no reliable internal distinction between "I know this" and "I made this up."</p>
<p>The natural assumption is that better prompts are the answer. Three prompt architectures, three model families, multiple experiments. Prompts control how MUCH the model fabricates. None of the tested architectures controlled WHETHER it fabricates.</p>
<p>Take any analytical topic. Ask three different AI models to write about it. Count the numbers. Now run the same model, same prompt, same topic three more times. Count how many numbers appear in all versions.</p>
<p>77 to 100 percent of model-generated numbers changed between runs. On two generators, none survived. On the third, roughly a quarter did, mostly round numbers reused across different claims. The model isn't recalling facts. It's generating plausible-looking results.</p>
<p>They're temporally unstable. In one controlled comparison, nearly half the model-generated numbers coincidentally matched real sources. The problem isn't that they're all wrong. It's that from the output alone, you can't tell which are real.</p>
<p>So you add constraints. Require evidence. Demand sourcing. Specify epistemic standards. Three levels tested. BASIC: "analyze this topic." STANDARD: structural constraints requiring shaped, specific output. PROTOCOL: the full system built over 50 experiments, every claim constrained for honesty, evidence, and falsifiability.</p>
<p>BASIC: 85.8 percent fabrication. STANDARD, with structural constraints: 78.4. Slightly better. PROTOCOL, with full epistemic constraints demanding evidence and sourcing: 90.7. Worse than unconstrained. On xAI, structural constraints helped modestly. Epistemic constraints pushed fabrication up. Template slots demanding specificity got filled with fabrication.</p>
<p>On Gemini Flash, the direction reversed. Constraints increased fabrication instead of reducing it. On one topic, the unconstrained version produced zero numbers at all. Constraints forced the model to generate claims it wouldn't have made unprompted. All were fabricated. The condition gradient is generator-specific. The aggregate rate is robust: majority fabrication regardless of constraints or generator.</p>
<p>What prompts DID change was volume. PROTOCOL produced roughly a third the numerical claims of STANDARD. The constraints made the model claim less, not claim better. Net fabricated numbers per document went down. Accuracy did not improve. The model generated fewer claims total. A system that demands honesty produces fewer lies. Each one is more elaborate. The rate is structural. The volume is prompt-controllable.</p>
<p>The mechanism: demanding specificity narrows what the model can produce. It needs things that look like real numbers, real benchmarks. If it has that knowledge, it retrieves it. If it doesn't, it generates something that looks right. As far as we can tell from behavior, the model doesn't know the difference. Neither do you.</p>
<p>Other AI models rate the constrained output as higher quality 75 percent of the time. That rating is [LLM-judged](https://arxiv.org/abs/2306.05685). A human domain expert agreed on 0 of 5 for holistic quality. The evaluators reward the performance of specificity, not its truth. How much maps to actual academic research? BASIC: 33 percent grounded. STANDARD: 6.7 percent. Five times less grounded. Five times more precise-sounding.</p>
<p>Then real data gets added to the prompt. Same models. Same topics. Same prompts. But this time, actual source material in context.</p>
<p>BASIC went from 85.8 percent temporal instability to 1.7 percent unsourced. STANDARD from 78.4 to 2.6. PROTOCOL from 90.7 to 3.4.</p>
<p>Source material moved the needle 46 percentage points. Prompt architecture moved it 6. The ordering is clear: source material is the dominant variable. The coupling that survived 30 experiments broke the moment the model had something real to build from. It was never about how models generate. It was about generating without anything real to ground against.</p>
<p>Months went into building PROTOCOL. The entire constraint architecture was solving the wrong problem.</p>
<p>Presence alone isn't enough. The model needs explicit instruction, not just data nearby. Monitoring says "flag unsourced numbers." Prohibition says "use only numbers from the source." Prohibition was five times better. 1.6 percent versus 7.7. And it cost nothing. The model compensated by extracting more and writing more, not less.</p>
<p>Three steps. Put real data in the prompt. Add one line: "Use only numbers from the source material above. If the source doesn't contain a relevant number, make the analytical point without inventing numbers." After the output, match the numbers against the source.</p>
<p>This isn't a prompt trick. It's a workflow change. The variable that matters most isn't how you ask. It's whether you provide something real to work from.</p>
<p>One thing source grounding does NOT change: how the output reads. In a blinded comparison, a domain expert rated source-present and source-absent output as equivalent. The fabricated version was rated more trustworthy in some cases. It cited more sources, used more specific numbers, and asserted with more confidence. The sourced version acknowledged limitations and had fewer citations. It only cited what was actually provided. [The trust signals are inverted](/trust-signals-are-inverted): the less reliable output has more of the markers humans use to assess authority. Source grounding doesn't make the output LOOK better. It makes the output CHECKABLE. The third step is where the value lives. Match the numbers against the source. Without that step, the fabricated version wins on perceived authority.</p>
<p>A direct verification check tested the fabrication rate from a different angle. Twenty-three claims citing named sources were checked against [the cited sources' actual publications](/receipts/fabrication-architecture). The sources were McKinsey, BCG, Gallup, Gartner, and the Standish Group. Two of twenty-three verified correct. Both are among the most widely-cited statistics in their domains. Gartner's $15 million data quality cost. The Standish Group's 29 percent project success rate. Essentially common knowledge. The other twenty-one were assembled: real components, fabricated binding.</p>
<p>The model doesn't just generate unstable numbers. It fabricates the source attribution through five distinct mechanisms. It takes a real number from a real source and attaches it to a different claim. McKinsey's "45%" is about cumulative profit impact over a decade. The model wrote "45% of firms experienced disruptions lasting more than one month." It performs a correct calculation on real data and presents the result as a direct finding. Gallup: 45% vs 39% stress is a 15% relative increase. The model presented that as "Gallup reports 15% higher burnout." It generates domain-appropriate components and combines them. "BCG analysis of 1,500 firms" when BCG studied 150. It attaches a real statistic to the wrong source. "85% of analytics projects fail" is Gartner, not McKinsey. It applies real data to a broader scope than the original. The Standish CHAOS report covers all IT projects. The model presented it as specifically about migrations.</p>
<p>Each component is plausible. Each passes a fact-check that [only verifies individual parts](https://arxiv.org/abs/2305.14251). The fabrication is in the binding: this source says this number about this topic.</p>
<p>Numbers that don't change across regenerations aren't necessarily right. The same "70%" shows up in three versions attached to three unrelated claims: revenue concentration, stall rates, headcount allocation. Not the same assertion recurring. Just a common number in business contexts. The regeneration test catches more than chance would. But verification against the source catches what the test can't.</p>
<p>Vocabulary and causal framing stay at baseline across regenerations. The tools solve one layer, the most measurable, most verifiable. The layers above it are yours.</p>
<p>But source grounding itself reaches further than its measurement. On reasoning tasks with known ground truth, source-present output found the correct answer 75 percent of the time versus 38 percent without source. The tasks were medical diagnosis, financial forensics, architecture decisions, legal review, and root cause analysis. Fewer wrong conclusions. More appropriate confidence. Source-absent output was overconfident. It made strong claims without backing. Three of five tasks showed source-present dramatically outperforming: financial analysis, architecture decisions, root cause analysis. One was moderately better: legal review. One was tied: medical diagnosis. Source material doesn't just stabilize numbers. It improves correctness on reasoning tasks with verifiable answers. The claims are checkable. They are also more likely correct.</p>
<h3>What survived testing</h3><ul><li>77-100% of model-originated numbers are temporally unstable across all generators. 20 topics on one generator. Replicated across three generators with 10 topics.</li><li>Epistemic constraints increase fabrication rate. Condition gradient is generator-specific. Direction consistent across topics. PROTOCOL produced 10.0 numerical claims per thousand words versus STANDARD 32.8.</li><li>Source material reduces fabrication from majority rates to single-digit percentages. Source effect is roughly an order of magnitude larger than prompt effect. Source material moved the needle 46 percentage points. Prompt architecture moved it 6. Different measurement constructs from different experiments, so the 46-versus-6 ratio is an ordering indicator, not a precise multiple. BASIC 1.7 percent unsourced is self-verified; programmatic matching shows a 2-10 percent range.</li><li>Prohibition outperforms monitoring by a factor of five</li><li>Format-robust, density-robust</li><li>Ground-truth check: 2 of 23 named-source citations verified correct. Both correct claims are widely-cited common knowledge. Five fabrication mechanisms confirmed.</li><li>Unconstrained output is 5x more grounded in real academic knowledge than constrained output</li><li>Measurement approach (regenerate and count what changes) converges with academic methods. Head-to-head against [SelfCheckGPT (Manakul et al., 2023)](https://arxiv.org/abs/2303.08896), which uses LLM sentence-level consistency checking: same directional result on the same documents, both near-zero in source-present. The programmatic approach costs nothing after generation; the LLM approach requires hundreds of API calls.</li></ul>
<h3>What didn&apos;t survive</h3><ul><li>&quot;100% fabrication is universal&quot; killed. One generator shows 77% with topic-dependent retrieval</li><li>&quot;PROTOCOL fixes fabrication&quot; killed. Highest fabrication rate of all conditions</li><li>&quot;Source grounding fixes everything above data&quot; partially killed. Vocabulary and causal framing stay at baseline for same-topic regeneration. But on reasoning tasks with ground truth, source-present output finds the correct answer roughly twice as often as source-absent. Source grounding improves correctness on reasoning tasks, not just reformulation.</li></ul>
<h3>Honest limits</h3><ul><li>N=1 practitioner. Zero external replication</li><li>Fabrication measurement tested reformulation only. Tools: temporal consistency, number matching. Source grounding&apos;s correctness benefit tested on 5 reasoning domains, 20 documents. Strategy, creative untested.</li><li>Source size 600 characters to 4KB tested. Larger untested</li><li>Conflicting sources untested. Qualitative-only sources untested</li><li>&quot;Fabrication rate&quot; uses temporal instability as proxy. A number could be unstable AND correct (sampled differently each time from genuine knowledge), or stable AND wrong (a memorized falsehood). Roughly half of temporally unstable numbers coincidentally match real sources. The rate measures unverifiability, not wrongness</li><li>Condition gradient is generator-specific. Practical advice on constraint effects should be generator-scoped.</li><li>Domain expert cannot distinguish sourced from fabricated by reading, even in own deep domain</li><li>Mechanistic interpretability research may eventually find an internal distinction between known and fabricated. The claim here is behavioral.</li></ul>]]></content:encoded>
      <pubDate>Mon, 23 Mar 2026 12:00:00 GMT</pubDate>
    </item>
  </channel>
</rss>