🛰️ Daily AI Frontier
‹ back to 2026-08-17

Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling

Research LLM Agents

Ranking

Overall 79
Content 95
Popularity 41

Observed public metrics from 1 member.

Representative image for Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling

Merged summary

TL;DR - AdmitOR is a label-free gate for deciding which optimization-modeling skills an LLM agent should retain. It improves admission precision and downstream benchmark accuracy, while exposing limitations when benchmark descriptions misrepresent labeled instances.

  • Compares value-function traces across resampled parameter instances, model families, prompts, and solver stacks.
  • Achieves 0.927 admission precision versus 0.871 for majority vote and 0.726 for execution-only checks.
  • Its smaller skill library reaches 58.4 macro accuracy across five benchmarks, outperforming majority vote at 54.8 and a ground-truth-labeled library at 53.9.
  • Its calibrated false-discovery criterion fails on a wild stream, largely because some benchmark texts do not faithfully encode their labeled instances.

Sources (1)

Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling

arXiv cs.AI Junbo Jacob Lian, Huiling Chen, Hanzhang Qin, Chung-Piaw Teo 2026-08-16 arXiv:2608.15565
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-14 14:22:27.085293 UTC

TL;DR - AdmitOR is a label-free gate for deciding which optimization-modeling skills an LLM agent should retain. It improves admission precision and downstream benchmark accuracy, while exposing limitations when benchmark descriptions misrepresent labeled instances.

  • Compares value-function traces across resampled parameter instances, model families, prompts, and solver stacks.
  • Achieves 0.927 admission precision versus 0.871 for majority vote and 0.726 for execution-only checks.
  • Its smaller skill library reaches 58.4 macro accuracy across five benchmarks, outperforming majority vote at 54.8 and a ground-truth-labeled library at 53.9.
  • Its calibrated false-discovery criterion fails on a wild stream, largely because some benchmark texts do not faithfully encode their labeled instances.
item →