论文 | AutoDesign:面向长程智能体设计的元Harness优化
Ranking
Overall
83
Content
90
Popularity
66
Observed public metrics from 1 member.
Merged summary
TL;DR - AutoDesign is a long-horizon agent framework that improves the external tools, context, runtime, orchestration, and feedback around a fixed model rather than updating model weights. On paper-to-poster generation, this reusable “DesignHarness” substantially improved automated and human-rated quality.
- A meta-harness iteratively updates one DesignHarness component at a time, accepting changes only when training performance improves without degrading a held-out development set.
- PosterBench evaluates 100 papers across fidelity, coverage, density, visual evidence, layout, readability, and aesthetics.
- AutoDesign scored 78.32, 7.45 points above Claude Design; across seven controlled configurations, the harness raised the average score from 54.99 to 67.39.
- The system produces editable HTML artifacts and combines traceable evidence, rule-based validation, visual critique, local refinement, and fallback selection.
Sources (1)
论文 | AutoDesign:面向长程智能体设计的元Harness优化
Public signals
Hugging Face upvotes 63
TL;DR - AutoDesign is a long-horizon agent framework that improves the external tools, context, runtime, orchestration, and feedback around a fixed model rather than updating model weights. On paper-to-poster generation, this reusable “DesignHarness” substantially improved automated and human-rated quality.
- A meta-harness iteratively updates one DesignHarness component at a time, accepting changes only when training performance improves without degrading a held-out development set.
- PosterBench evaluates 100 papers across fidelity, coverage, density, visual evidence, layout, readability, and aesthetics.
- AutoDesign scored 78.32, 7.45 points above Claude Design; across seven controlled configurations, the harness raised the average score from 54.99 to 67.39.
- The system produces editable HTML artifacts and combines traceable evidence, rule-based validation, visual critique, local refinement, and fallback selection.