Magenta: Closing the Loop Between Mathematical Reasoning and Lean Verification
Ranking
Overall
78
Content
95
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - Magenta is a training-free agentic pipeline that converts natural-language math solutions into Lean 4 statements and machine-checked proofs. It closes the informal–formal reasoning loop with judges that detect faulty formalizations and route errors for targeted correction.
- Generates an answer, formalizes it in Lean 4, and constructs a verified proof from only a natural-language problem.
- Uses a statement judge to check semantic fidelity and an error-attribution judge to choose between mathematical re-derivation and local Lean repair.
- Reports 100% accuracy on evaluated olympiad benchmarks, including AIME 2025/2026 and HMMT February 2026.
- Paired with the open-weight K2-Horizon-7B reasoner, it reportedly solves all six IMO 2026 problems.
Sources (1)
Magenta: Closing the Loop Between Mathematical Reasoning and Lean Verification
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - Magenta is a training-free agentic pipeline that converts natural-language math solutions into Lean 4 statements and machine-checked proofs. It closes the informal–formal reasoning loop with judges that detect faulty formalizations and route errors for targeted correction.
- Generates an answer, formalizes it in Lean 4, and constructs a verified proof from only a natural-language problem.
- Uses a statement judge to check semantic fidelity and an error-attribution judge to choose between mathematical re-derivation and local Lean repair.
- Reports 100% accuracy on evaluated olympiad benchmarks, including AIME 2025/2026 and HMMT February 2026.
- Paired with the open-weight K2-Horizon-7B reasoner, it reportedly solves all six IMO 2026 problems.