英伟达Vera Rubin的Agent吞吐暴涨最高提升30倍,却不只靠GPU | Hot Chips 2026
Ranking
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR - Nvidia says its Vera Rubin NVL72 platform delivers up to 30× higher agent-workload throughput per megawatt than GB300 NVL72 by optimizing the entire heterogeneous inference stack, not just the GPU. The design divides long-context processing, low-latency decoding, tool execution, and networking among specialized components.
- The preliminary, SemiAnalysis-pending AgentX benchmark used realistic multi-turn traces with median input contexts above 140,000 tokens and measured DeepSeek V4-Pro at 160 tokens per second per user.
- Rubin GPUs handle large-scale model computation, while Groq 3 LPX accelerators target low-latency token generation through prefill/decode partitioning, attention/FFN splitting, or speculative decoding.
- Vera CPUs handle orchestration, code execution, data processing, and other work between model calls, extending optimization beyond inference kernels.
- Spectrum-X, BlueField-4, and DOCA address infrastructure traffic, cluster scaling, and resilience as tool, storage, and data access increasingly affect end-to-end agent performance.
Sources (1)
英伟达Vera Rubin的Agent吞吐暴涨最高提升30倍,却不只靠GPU | Hot Chips 2026
TL;DR - Nvidia says its Vera Rubin NVL72 platform delivers up to 30× higher agent-workload throughput per megawatt than GB300 NVL72 by optimizing the entire heterogeneous inference stack, not just the GPU. The design divides long-context processing, low-latency decoding, tool execution, and networking among specialized components.
- The preliminary, SemiAnalysis-pending AgentX benchmark used realistic multi-turn traces with median input contexts above 140,000 tokens and measured DeepSeek V4-Pro at 160 tokens per second per user.
- Rubin GPUs handle large-scale model computation, while Groq 3 LPX accelerators target low-latency token generation through prefill/decode partitioning, attention/FFN splitting, or speculative decoding.
- Vera CPUs handle orchestration, code execution, data processing, and other work between model calls, extending optimization beyond inference kernels.
- Spectrum-X, BlueField-4, and DOCA address infrastructure traffic, cluster scaling, and resilience as tool, storage, and data access increasingly affect end-to-end agent performance.