Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models
Merged summary
TL;DR - Controlled experiments across five LLMs show that instruction adherence collapses as rule counts grow, while long-context recall degrades near context limits. Prompt format has model-specific effects, with no consistent markdown advantage.
- Perfect-response rates fell to zero by 80 simultaneous instructions across all models, formats, and instruction placements.
- System-prompt versus user-turn placement affected compliance at least as much as formatting in most models at 160 rules.
- Recall remained strong through 64K–128K tokens before declining sharply and format-dependently near model context limits.
- Fabrication was absent and sycophancy negligible; failures near context ceilings primarily appeared as refusals, reaching 79–90%.
Sources (1)
Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models
TL;DR - Controlled experiments across five LLMs show that instruction adherence collapses as rule counts grow, while long-context recall degrades near context limits. Prompt format has model-specific effects, with no consistent markdown advantage.
- Perfect-response rates fell to zero by 80 simultaneous instructions across all models, formats, and instruction placements.
- System-prompt versus user-turn placement affected compliance at least as much as formatting in most models at 160 rules.
- Recall remained strong through 64K–128K tokens before declining sharply and format-dependently near model context limits.
- Fabrication was absent and sycophancy negligible; failures near context ceilings primarily appeared as refusals, reaching 79–90%.