Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs
Ranking
Overall
75
Content
90
Popularity
40
Observed public metrics from 1 member.
Merged summary
TL;DR - SmartVL jointly adapts visual-token count and LLM computation to each input and inference budget. Coordinated resource allocation improves the accuracy-efficiency tradeoff over methods that optimize these dimensions independently.
- Uses separate vision-token and LLM-compute controllers linked by a shared budget encoding.
- Dynamically selects informative visual tokens and adjusts language-model computation.
- Employs a differentiable latency estimator for end-to-end, budget-aware training.
- Demonstrates stronger accuracy-efficiency Pareto frontiers across multiple multimodal benchmarks.
Sources (1)
Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - SmartVL jointly adapts visual-token count and LLM computation to each input and inference budget. Coordinated resource allocation improves the accuracy-efficiency tradeoff over methods that optimize these dimensions independently.
- Uses separate vision-token and LLM-compute controllers linked by a shared budget encoding.
- Dynamically selects informative visual tokens and adjusts language-model computation.
- Employs a differentiable latency estimator for end-to-end, budget-aware training.
- Demonstrates stronger accuracy-efficiency Pareto frontiers across multiple multimodal benchmarks.