🛰️ Daily AI Frontier
‹ back to 2026-08-18

Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning

Research LLMs & Foundation Models

Ranking

Overall 84
Content 95
Popularity 60

Observed public metrics from 1 member.

Representative image for Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning

Merged summary

TL;DR - ACA-RL trains reasoning models to handle underspecified questions by asking for missing information, giving conditional answers, or abstaining. It improves missing-premise reasoning while retaining competitive performance on well-posed tasks.

  • Generates training examples by removing premises from well-posed problems and annotating the resulting gaps.
  • Uses structured rewards covering five observable response behaviors.
  • Introduces MPB, a 274-instance, human-verified benchmark spanning mathematical, logical, and real-world problems.
  • Consistently improves Qwen3 and Llama models on MPB; code, benchmark, and training data are released.

Sources (1)

Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning

arXiv cs.CL Yongqi Tong, Zhenyu Zhang, Zimi Liu, Kewei Fu, Mingli Song, Haofei Zhang, Junshao Zhang, Hong Zhu, Jiang-Ming Yang, Xin Zhang, Jianshe Li 2026-08-17 arXiv:2608.16554
Public signals Hugging Face upvotes 1
Providers: Hugging Face · Upvotes 1 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-17 14:33:10.391592 UTC

TL;DR - ACA-RL trains reasoning models to handle underspecified questions by asking for missing information, giving conditional answers, or abstaining. It improves missing-premise reasoning while retaining competitive performance on well-posed tasks.

  • Generates training examples by removing premises from well-posed problems and annotating the resulting gaps.
  • Uses structured rewards covering five observable response behaviors.
  • Introduces MPB, a 274-instance, human-verified benchmark spanning mathematical, logical, and real-world problems.
  • Consistently improves Qwen3 and Llama models on MPB; code, benchmark, and training data are released.
item →