It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief
Merged summary
TL;DR - This paper introduces a linguistically grounded benchmark for measuring how expressions of belief influence whether LLMs follow user-provided context or prior knowledge. Results show systematic sensitivity to phrasing, raising robustness and prompt-engineering concerns.
- The typology covers 17 expression types across form, evidentiality, epistemic stance, and tone.
- Controlled belief-query pairs isolate linguistic framing effects while using world-knowledge facts.
- Evaluation spans 16 Llama3, Qwen3, and Gemma3 models from 1B to 30B parameters.
- Larger and instruction-tuned models were generally less context-following, while some expression types were consistently more persuasive than others.
Sources (1)
It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief
TL;DR - This paper introduces a linguistically grounded benchmark for measuring how expressions of belief influence whether LLMs follow user-provided context or prior knowledge. Results show systematic sensitivity to phrasing, raising robustness and prompt-engineering concerns.
- The typology covers 17 expression types across form, evidentiality, epistemic stance, and tone.
- Controlled belief-query pairs isolate linguistic framing effects while using world-knowledge facts.
- Evaluation spans 16 Llama3, Qwen3, and Gemma3 models from 1B to 30B parameters.
- Larger and instruction-tuned models were generally less context-following, while some expression types were consistently more persuasive than others.