RT by @huggingface: Today, we’re excited to open-source TwiL-LM3, the first formal reasoning model…
TL;DR - webAI Intelligence Lab open-sourced TwiL-LM3, a 3B-parameter "formal reasoning" model claimed to beat GPT-OSS-120B on 4 of 5 formal reasoning benchmarks while running on consumer/edge hardware. It matters as another data point that curated training pipelines, not just scale, can drive reasoning performance.
- Claims 40× fewer parameters and 2.6× faster inference than GPT-OSS-120B, with wins on 4 of 5 formal reasoning benchmarks (specific benchmarks and numbers not given in the post).
- Trained via a proprietary reasoning pipeline on webAI-owned, verified datasets rather than scraped web data — the stated thesis is that data/pipeline quality beats raw model size.
- Targeted at downstream reliability tasks: tool calling, code generation, structured outputs, and agents; positioned for on-device deployment from Raspberry Pi to iPhone, avoiding cloud dependence.
- First open-source release from the lab, announced via a HuggingFace retweet; claims are vendor-reported and not independently verified in the provided content.