Skip to content

🌐 日本語

Prompt Sensitivity — Same Meaning, Different Results

NOTE

In short: LLMs generate significantly different outputs for semantically equivalent prompts. Merely changing the few-shot formatting has been reported to cause differences of up to 76 percentage points in accuracy (Sclar et al. 2023). This is not merely instability, but a reflection of the model's reliance on statistical token patterns — though the magnitude of the observed difference also depends on the evaluation method (see below).

What is Prompt Sensitivity?

Prompt Sensitivity is the phenomenon where LLMs produce substantially different outputs even when given semantically identical prompts, if the wording differs.

For example:

  • "Please refactor this function"
  • "Please improve this function"
  • "Please clean up this function"

Although these are semantically nearly equivalent, an LLM may generate different outputs for each.

Why Does It Occur?

Mathematical Explanation (a conceptual first-order approximation)

The output's sensitivity to a small change in the input can, conceptually, be estimated by a first-order approximation (a Taylor expansion with a Cauchy-Schwarz upper bound):

Output Difference ≲ Gradient Norm × Embedding Difference Norm

NOTE

This is not a theorem from any specific paper, but a conceptual heuristic grounded in gradient-based saliency (saliency = ‖∇(output logit)‖, Lu et al. 2024). Because Transformers are strongly nonlinear (Attention, FFN), a first-order approximation is only a local indicator of sensitivity, with limited global explanatory power.

The point to note: in embedding space, semantically similar inputs are clustered. Sensitivity still arises because small embedding differences are amplified by downstream nonlinear transformations. The accurate framing is "the meaning is close, yet the effect on the output distribution can be large."

Impact of Surface Form

LLMs respond largely to statistical patterns in tokens rather than meaning. As a result:

  • Imperative vs. interrogative sentences produce different results
  • Bullet points vs. free text produce different results
  • Technical terminology vs. plain language produce different results

Quantitative Evidence

  • Merely changing the few-shot formatting (surface-level, spurious features such as separators, symbols, and casing) produces a difference of up to 76 accuracy points on LLaMA-2-13B (Sclar et al. 2023)
  • Note that this "76 points" is a difference due to formatting changes, not semantically equivalent paraphrases. Treat it as distinct from sensitivity to meaning-preserving rephrasing
  • The magnitude of sensitivity varies greatly by task, model, and evaluation method

NOTE

A substantial part of the observed sensitivity is an artifact of brittle evaluation metrics (log-likelihood scoring and rigid answer matching overlooking semantically correct answers expressed through alternative phrasings); under appropriate evaluation design, modern LLMs are more robust than previously reported (Hua et al. 2025). So do not over-generalize "prompt sensitivity is an inevitable structural constraint of the Transformer." The effect is real, but the observed magnitude depends on the evaluation method. The practical implication (ambiguous instructions are unstable) is unchanged, but the numbers must be read together with the benchmark setup.

Underspecification — When an Axis Is Left Unstated, the Prior Takes Over

A twin problem of Prompt Sensitivity is Underspecification. If Prompt Sensitivity is "changing the wording of an already-specified prompt changes the output," Underspecification is "leaving an axis unstated entirely lets the model fill it from its prior distribution." Underspecification is the limiting case of sensitivity — for an axis with zero specification, the output is decided not by reasoning but by the most frequent pattern in the training data.

NOTE

Framing Underspecification as a twin / limiting case of Prompt Sensitivity, and the connection to the sister site below, is this site's own framing (none of the individual cited papers claim this correspondence).

Why the Model Cannot Decide on Its Own

An LLM's output is a sample from the conditional probability distribution P(output | token sequence). The prompt is merely the token sequence that conditions that distribution.

  • If the prompt specifies an axis (role, output format, success criteria, etc.), the distribution is sharply narrowed along that axis.
  • If it does not, the conditioning on that axis stays weak, and the model fills it by sampling from its prior — the pattern most frequent in the training data.

So the model does not "fail to decide." It mechanically fills the unspecified axis from a statistical prior rather than by reasoning. Because that prior shifts with context and token sequence, the same request gets filled differently across sessions — and that is exactly the source of nondeterminism (the parts that drift each time).

IMPORTANT

When we say "without a stated role, the model cannot decide which perspective to answer from," strictly speaking it is not deciding. It is merely filling a weakly-specified axis with the mode of its prior. So the remedy is not "get the model to decide well" but to explicitly specify the axes you do not want to vary, sharpening the conditioning.

RequestUnspecified axisWhat the model fills from its prior
"Write tests"Test frameworkWhatever is most frequent in training data (Jest, etc., depending on the project)
"Document this function"Output format (JSDoc / Markdown / comments)The most frequent style per language
"Review this"Lens (bugs / design / style) and strictnessA generic, "safe" lens

Connection to the Sister Site

The sister site ai-agent-architecture organizes the seven conditions of a well-formed prompt (Role, Premise, Objective, Input, Process/Constraints, Output Format, Examples) as "independent axes along which output can vary." This section carries the why behind it — the principle by which weakly-specified axes get filled from the prior. For the design decision to externalize each axis into a layer instead of re-filling it in every prompt, see there.

Impact on Coding

  • Rules written ambiguously in CLAUDE.md are less likely to be followed
  • Vague Skills descriptions lead to failed automatic invocations
  • The quality of generated code varies depending on how users phrase their natural language requests

Mitigation in Claude Code

Mitigation StrategyMechanismWhy It Works
CLAUDE.md writing styleConcrete, imperative language with code examplesEliminates ambiguous expressions, improves compliance rate
Skills description designInclude diverse user natural language expressionsSimilar to SEO principles, improves matching accuracy across varied phrasings
Conditional injection via .claude/rules/Reduces number of simultaneously active instructionsPrevents sensitivity degradation (effect increases with more instructions)
Hooks and testsExternal validation independent of prompt wordingVerifies results regardless of how the prompt is written
Plugins / MarketplacesDistribute verified prompts as installable packagesSee Appendix: Plugins & Marketplaces — team-wide calibration instead of per-engineer trial and error

Writing Effective CLAUDE.md

markdown
# ❌ Ambiguous (high sensitivity)

- Please write good tests
- I want clean code

# ✅ Concrete (low sensitivity)

- Create Jasmine tests for all public methods
- Place test files in *.spec.ts
- Use describe/it structure in test writing

Writing Effective Skills Descriptions

yaml
# ❌ Ambiguous (auto-invocation often fails)
description: Component-related tasks

# ✅ Concrete (covers diverse expressions)
description: >
  Create new Angular components. Generate scaffolding with OnPush
  change detection, NgRx Store integration, and Jasmine tests.
  Use for requests like "create a component", "add a new screen", etc.

Relationship to Other Structural Problems

Prompt Sensitivity bidirectionally amplifies with other problems.

TIP

Solid arrows (→): Direction in which each problem amplifies Prompt Sensitivity / Dashed arrows (⇢): Feedback loops where Prompt Sensitivity worsens each problem

References

  • Sclar, M., Choi, Y., Tsvetkov, Y., & Suhr, A. (2023). "Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design." arXiv:2310.11324. arXiv — Up to 76 accuracy points on LLaMA-2-13B from few-shot formatting (spurious feature) changes. The source of this page's "76 points"
  • Zhuo, J., Zhang, S., Fang, X., Duan, H., Lin, D., & Chen, K. (2024). "ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs." EMNLP 2024 Findings. ACL Anthology — Empirical sensitivity assessment via PromptSensiScore (PSS) and decoding confidence (no Taylor-expansion formulation appears in this paper)
  • Lu, S., Schuff, H., & Gurevych, I. (2024). "How are Prompts Different in Terms of Sensitivity?" NAACL 2024. ACL Anthology — Analyzes prompt sensitivity via gradient-based saliency (‖∇output‖). The grounding for this page's first-order approximation
  • Hua, A., Tang, K., Gu, C., Gu, J., Wong, E., & Qin, Y. (2025). "Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs." EMNLP 2025. arXiv — A substantial part of observed sensitivity is an artifact of brittle evaluation metrics; under proper evaluation design, LLMs are more robust than reported

Previous: Knowledge Boundary

Next: Instruction Decay

Discussion: #12 Prompt Sensitivity

Released under the CC BY 4.0 License.