数学大佬在前面拓荒,AI研究员在后面捡宝,菲尔兹奖还能拿来破AI「黑盒」?
TL;DR - GeoLAN applies geometric constraints inspired by “sticky Kakeya sets” to reduce representation and attention collapse during LLM training. The approach aims to make latent concepts more separable and interpretable without sacrificing task performance.
- KT-CW penalizes semantic concentration along a few directions, encouraging more isotropic latent representations.
- KT-Attn promotes diversity among attention heads to mitigate attention-rank collapse.
- Reported tests on Llama-3-8B and Gemma-3 models show improved representation uniformity, modest benchmark gains, and reduced bias or greater semantic stability.
- The benefits reportedly concentrate in medium-sized models; overly strict constraints can hurt smaller or much larger models.