🛰️ Daily AI Frontier
‹ back to 2026-09-03

LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style Updates

Research LLMs & Foundation Models

Ranking

Overall 78
Content 95
Popularity 37

Observed public metrics from 1 member.

Merged summary

TL;DR - LoRA-TSD is a geometry-aware optimizer that applies Muon-style spectral descent within the tangent space of the fixed-rank LoRA update. It improves benchmark performance across several model scales while reducing retraction cost and providing convergence guarantees.

  • Treats each LoRA update as a tangent vector on a fixed-rank matrix manifold rather than optimizing its two factors independently.
  • Uses a LoRA-native retraction that is up to 2.8Ă— cheaper than the truncated-SVD retraction used in prior manifold methods.
  • Establishes the tangent-projected gradient as a stationarity measure and gives the first global convergence guarantees under this measure for LoRA-TSD and LoRA-Pro.
  • Outperforms competing LoRA optimizers across six commonsense and natural-language-inference benchmarks using Llama and Qwen models, while remaining robust to adapter rank.

Sources (1)

LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style Updates

arXiv cs.LG Dmitrii Andriianov, Andrey Veprikov, Aleksandr Beznosikov 2026-09-02 arXiv:2609.02734
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-19 14:18:50.181318 UTC

TL;DR - LoRA-TSD is a geometry-aware optimizer that applies Muon-style spectral descent within the tangent space of the fixed-rank LoRA update. It improves benchmark performance across several model scales while reducing retraction cost and providing convergence guarantees.

  • Treats each LoRA update as a tangent vector on a fixed-rank matrix manifold rather than optimizing its two factors independently.
  • Uses a LoRA-native retraction that is up to 2.8Ă— cheaper than the truncated-SVD retraction used in prior manifold methods.
  • Establishes the tangent-projected gradient as a stationarity measure and gives the first global convergence guarantees under this measure for LoRA-TSD and LoRA-Pro.
  • Outperforms competing LoRA optimizers across six commonsense and natural-language-inference benchmarks using Llama and Qwen models, while remaining robust to adapter rank.
item →