智谱唐杰:万亿参数是行业早期的探索
Ranking
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR - Zhipu AI co-founder Tang Jie argues that scaling laws should optimize parameters, data, inference cost, compute depth, and post-training together—not prioritize trillion-parameter models. GLM-5.3 illustrates this approach by retaining GLM-5.2’s base model and parameter counts while improving performance through expanded long-horizon environments and reinforcement learning.
- Tang calls the industry’s early pursuit of trillion-parameter models a detour, citing Chinchilla’s evidence for balancing model size with training data.
- Lifecycle inference costs favor smaller models trained on more tokens, while MoE architectures require distinguishing total parameters from activated parameters and effective compute depth.
- Total parameters mainly expand knowledge capacity; activated compute and depth are more important for sustained multi-step reasoning.
- Zhipu spent one month scaling GLM-5.3’s long-horizon task environments and reinforcement learning, while leaving architecture and parameter counts unchanged.
Sources (1)
智谱唐杰:万亿参数是行业早期的探索
TL;DR - Zhipu AI co-founder Tang Jie argues that scaling laws should optimize parameters, data, inference cost, compute depth, and post-training together—not prioritize trillion-parameter models. GLM-5.3 illustrates this approach by retaining GLM-5.2’s base model and parameter counts while improving performance through expanded long-horizon environments and reinforcement learning.
- Tang calls the industry’s early pursuit of trillion-parameter models a detour, citing Chinchilla’s evidence for balancing model size with training data.
- Lifecycle inference costs favor smaller models trained on more tokens, while MoE architectures require distinguishing total parameters from activated parameters and effective compute depth.
- Total parameters mainly expand knowledge capacity; activated compute and depth are more important for sustained multi-step reasoning.
- Zhipu spent one month scaling GLM-5.3’s long-horizon task environments and reinforcement learning, while leaving architecture and parameter counts unchanged.