🛰️ Daily AI Frontier
‹ back to 2026-08-10

Is SwiGLU's Open Positive Tail Necessary? Evidence from Closed-Tail Gating with MemGLU

Research LLMs & Foundation Models

Ranking

Overall 66
Content 75
Popularity 43

Observed public metrics from 1 member.

Merged summary

TL;DR - An arXiv preprint testing whether SwiGLU's unbounded positive tail is actually required in decoder-only LLM feed-forward networks, using a closed-tail gating alternative called MemGLU. At small pretraining scales it isn't: MemGLU matches SwiGLU within ~0.1% validation NLL, suggesting activation-function design has more slack than commonly assumed.

  • MemGLU is introduced as a closed-tail comparator derived from a memristive branch geometry, contrasting with SwiGLU's open (unbounded) positive tail.
  • Evaluation used paired pretraining runs at 9M and 30M parameters with three seeds; MemGLU stayed within roughly 0.1% of SwiGLU in validation negative log-likelihood.
  • Trained SwiGLU checkpoints degrade under positive-tail suppression, and mechanism diagnostics show the two models use their gates differently despite comparable loss — implying models adapt to whatever gate geometry is present during pretraining.
  • Claims are explicitly scoped to "the tested scales" (9M/30M); no evidence is offered for larger models or downstream task performance.

Sources (1)

Is SwiGLU's Open Positive Tail Necessary? Evidence from Closed-Tail Gating with MemGLU

arXiv cs.LG Yuting Ge, Pengju Yang, Mingkai Nie 2026-08-07 arXiv:2608.07323
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-23 14:16:26.121760 UTC

TL;DR - An arXiv preprint testing whether SwiGLU's unbounded positive tail is actually required in decoder-only LLM feed-forward networks, using a closed-tail gating alternative called MemGLU. At small pretraining scales it isn't: MemGLU matches SwiGLU within ~0.1% validation NLL, suggesting activation-function design has more slack than commonly assumed.

  • MemGLU is introduced as a closed-tail comparator derived from a memristive branch geometry, contrasting with SwiGLU's open (unbounded) positive tail.
  • Evaluation used paired pretraining runs at 9M and 30M parameters with three seeds; MemGLU stayed within roughly 0.1% of SwiGLU in validation negative log-likelihood.
  • Trained SwiGLU checkpoints degrade under positive-tail suppression, and mechanism diagnostics show the two models use their gates differently despite comparable loss — implying models adapt to whatever gate geometry is present during pretraining.
  • Claims are explicitly scoped to "the tested scales" (9M/30M); no evidence is offered for larger models or downstream task performance.
item →