Skip to content
All findings
ExperimentTested March 2026 · xAI / Gemini· 3 min

Why 'Don't Be Generic' Doesn't Work

By Lovro Lucic ·

The "It Depends" Problem · 3 of 6

Asked for a competitive analysis. Gave it everything. Market position, three-year data, the specific situation. "Be insightful. Don't be generic."

The output was structured, fluent, professional. And interchangeable with what it would have produced for any company in any market. Every recommendation could have been copy-pasted into a competitor's strategy doc without changing a word.

Twenty controlled runs confirmed what was visible. Specific versus vague, crossed with positive versus negative framing. Four combinations. Pure negation ("don't be generic," "avoid cliches," "don't use buzzwords") was indistinguishable from giving no instruction at all.

That removes a label without providing a destination. The model left the default and wandered to the adjacent region. Same neighborhood. Different house number. Negation names what to avoid. The default already routes around that. Only an anchor forces a different path.

Then the other cell. Same negation, but paired with specifics: "Don't include recommendations that could apply to any B2B SaaS company. Every recommendation must reference Northvane's specific assets, 5 years of shipping logistics data, 12 engineers, Pacific Northwest enterprise incumbents."

Strong effect. The direction replicated across generators. This is the content-specificity lever. Stacking more explicit constraints is a different lever, from a different experiment. Keep the two separate.

The specificity isn't measurably adding analytical quality. It's adding verifiability.

"Don't be generic" blocks one path. The model takes the next most likely path, which is a variation on generic. "Reference these specific assets" creates an anchor the output has to pass through. The result physically cannot be the same for a different company. The constraint tests itself.

In blind testing, a domain expert couldn't distinguish specific from generic outputs on quality. Picked specific 3 out of 5 times. Chance level. The specificity instruction changes what the output looks like: more data references, more grounded claims. It does not, measurably, change what the output says.

The demonstrated value is verifiability. The specific output can be checked. Every claim traces to something nameable. The generic output makes the same points but you can't verify them. When you need to trust the analysis, specificity makes the output auditable. When you're using it as a starting point for your own thinking, it doesn't matter.

Test this yourself

Next time you write "don't be X," finish the sentence: "instead, Y with Z criteria." Run both versions. Measure the difference.

What survived testing

  • Negation alone has no effect. Specificity is real. The direction replicates across generators. Specific-plus-negative scored highest in the framing study; that ranking is not itself cross-validated. Clean magnitude: g=1.34 on raw marker count and 1.62 at density on xAI. Nearly identical at density on Gemini Flash (g=1.64), not on raw (g=0.65). The earlier Claude-versus-others gap came from a length-confounded comparison and is not clean. Quality demands do little alone. They add on top of specificity. Together the effect is larger than the sum. Negation removes a label without a destination. Specificity provides the anchor.Copy link

What didn't survive

  • "Negation hurts" overclaimed. Negation is null, not negative. "More specific = better" linearly is unverified: the effect was present versus absent, not a density gradient, so extreme specificity is untested. "At density is the pure confound-free measure" too strong: density (markers per 1,000 words) removes the prompt-length confound but inflates via brevity, because shorter outputs score higher. It strips one confound, not all. The cross-generator match is on density. The replication is directional, not a magnitude match.Copy link

Honest limits

  • Single operator. Transfer to other operators untested.Copy link
  • The negation result and the specificity-magnitude result come from two different experiments: a specificity-by-framing design shows negation alone is null, while the receipt-backed specificity-by-quality-demands 2x2 establishes the magnitude and the cross-generator replication. Only the second is in the bound receipt; the negation-by-framing result is not reader-auditable there.Copy link
  • The constraint-count lever's model differences are not all clean reads: Claude's 0.00 is a rubric ceiling (25/25, zero variance in both conditions), not an absence; the GPT result is self-scored (the model graded its own outputs), so treat it as an upper bound.Copy link
  • Effect sizes measure programmatic specificity markers (company mentions, scenario numbers, market terms). These counts are objective. Domain expert validation (5 blind pairs, evaluator's own domain): indistinguishable from chance. Expert rated based on style (rhythm, naturalness), not on marker density. At that sample size no quality difference was detectable, which is not proof there is none. Specificity demonstrably changes output FORM (more verifiable references); a SUBSTANCE difference was undetectable here, neither shown nor ruled out. What is demonstrated is verifiability.Copy link
Receipts

Next in The "It Depends" Problem

Adding Information Often Doesn't Help

New findings when they land.

No spam. Just what held up.