🛰️ Daily AI Frontier
‹ back to 2026-07-21

It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief

Research LLMs & Foundation Models

Merged summary

TL;DR - This paper introduces a linguistically grounded benchmark for measuring how expressions of belief influence whether LLMs follow user-provided context or prior knowledge. Results show systematic sensitivity to phrasing, raising robustness and prompt-engineering concerns.

  • The typology covers 17 expression types across form, evidentiality, epistemic stance, and tone.
  • Controlled belief-query pairs isolate linguistic framing effects while using world-knowledge facts.
  • Evaluation spans 16 Llama3, Qwen3, and Gemma3 models from 1B to 30B parameters.
  • Larger and instruction-tuned models were generally less context-following, while some expression types were consistently more persuasive than others.

Sources (1)

It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief

arXiv cs.CL Kevin Du, Clara Kümpel, Michelle Wastl, Alex Warstadt 2026-07-20 arXiv:2607.18232 doi:10.18653/v1/2026.acl-long.142

TL;DR - This paper introduces a linguistically grounded benchmark for measuring how expressions of belief influence whether LLMs follow user-provided context or prior knowledge. Results show systematic sensitivity to phrasing, raising robustness and prompt-engineering concerns.

  • The typology covers 17 expression types across form, evidentiality, epistemic stance, and tone.
  • Controlled belief-query pairs isolate linguistic framing effects while using world-knowledge facts.
  • Evaluation spans 16 Llama3, Qwen3, and Gemma3 models from 1B to 30B parameters.
  • Larger and instruction-tuned models were generally less context-following, while some expression types were consistently more persuasive than others.
item →