Rank 41 · Content 30 · Popularity 65
TL;DR — A cluster of Chinese-language WeChat commentary centered on Wang Hong (王虹) and Joshua Zahl's 127-page proof of the three-dimensional Kakeya (挂谷) conjecture, used as a case study for where today's AI hits its ceiling in frontier research: AI is an "outstanding student" for delegated technical subtasks, not yet a "master" that creates paradigms. Adjacent pieces (Terence Tao's ICM 2026 lecture, the MLS-Bench benchmark, Daphne Koller on drug discovery) independently corroborate the same boundary from evaluation, workflow, and applied-science angles.
- Why 3D Kakeya is the test case: unlike the flexible 2D setting, the 3D problem has rigid geometry, infinitely nested multi-scale/chaotic structure, "sticky" Kakeya configurations, and outruns existing Fourier/harmonic-analysis tooling. Wang and Zahl's contribution is framed as system-level innovation — abandoning prior research paths and building new machinery — spanning four branches of mathematics across 127 pages.
- The four claimed AI limits: reuse of existing paradigms instead of inventing frameworks; no global research-strategy judgment; inability to sustain very long, multi-layer, cross-domain logical chains; and no mathematical intuition for structures without precedent or template. AI is positioned as an accelerator for verification, gap-hunting, symbolic simplification, and pruning dead ends, and as dominant on bounded, competition-style problems.
- Empirical echo (MLS-Bench): 140 real research tasks across 12 ML areas, each with a real codebase, ≥3 transfer conditions, and ≥3 reproduced strong human methods (anchored scoring: weakest baseline = 0, strongest = 50, theoretical bound = 100). All five frontier models failed to beat the strongest human method; expert review found they mostly recombine losses/modules/tricks from provided baselines, "optimize/debug" prompts beat "discover a new method," and removing the parameter-count guard let models win purely by scaling — unisolated variables masquerading as discovery. With more compute freedom they performed worse, failing to allocate budget to high-information experiments. (Full run ≈700 H100-hours per candidate; a 30-task subset ≈100 H100-hours is used in official frontier-model release evals.)
- The downstream bottleneck (Tao): AI plus formal tools like Lean accelerate only proof generation and verification, congesting the remaining stages — exposition, publication, digestion, canonization — creating "proof indigestion." Cited: First Proof reported 7 of 10 novel research-level problems solved to publishable quality at $10–$1,000 each, while erdosproblems.com accumulates AI-submitted proofs without enough qualified verifiers. Prescriptions: disclose AI assistance, reward refereeing/surveys/translation over priority races, and apply a "talk test."
- Same ceiling in applied science (Koller): insitro's founder argues AI's drug-development bottleneck is identifying correct disease mechanisms, not generating molecules — >90% of trials fail on wrong mechanisms; 38 targets each carry 50+ programs while new targets/year fell from ~100 (2015) to ~30 (2024). Cell atlases cover a tiny slice of perturbation space and lack causal data; agentic self-driving labs need fast, cheap, objective feedback that human clinical trials cannot provide.
Emphasis differs by source: 图灵人工智能 argues the boundary conceptually (Kakeya) and procedurally (Tao's proof lifecycle), 学术头条 supplies the quantitative benchmark evidence, 智药局 transfers the argument to biomedicine, and the 专知 post is an unrelated content-marketing index of 100+ "AI + military" reports with no results or original analysis.