Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs
Merged summary
TL;DR - SmartVL jointly adapts visual-token count and LLM computation to each input and inference budget. Coordinated resource allocation improves the accuracy-efficiency tradeoff over methods that optimize these dimensions independently.
- Uses separate vision-token and LLM-compute controllers linked by a shared budget encoding.
- Dynamically selects informative visual tokens and adjusts language-model computation.
- Employs a differentiable latency estimator for end-to-end, budget-aware training.
- Demonstrates stronger accuracy-efficiency Pareto frontiers across multiple multimodal benchmarks.
Sources (1)
Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs
TL;DR - SmartVL jointly adapts visual-token count and LLM computation to each input and inference budget. Coordinated resource allocation improves the accuracy-efficiency tradeoff over methods that optimize these dimensions independently.
- Uses separate vision-token and LLM-compute controllers linked by a shared budget encoding.
- Dynamically selects informative visual tokens and adjusts language-model computation.
- Employs a differentiable latency estimator for end-to-end, budget-aware training.
- Demonstrates stronger accuracy-efficiency Pareto frontiers across multiple multimodal benchmarks.