🛰️ Daily AI Frontier
‹ back to 2026-08-29

Sora 为什么输给 Codex?

雷峰网 (AI科技评论) Efficiency & Systems 2026-08-28
Representative image for Sora 为什么输给 Codex?

TL;DR - OpenAI reportedly prioritized Codex over Sora not because coding agents inherently use less compute, but because their fragmented workloads let shared GPU infrastructure serve more concurrent tasks. The comparison shows how workload architecture and scheduling efficiency can shape an AI product’s scalability and resource allocation.

  • Sora’s video diffusion repeatedly processes large spatiotemporal latent states, creating long, compute-heavy jobs with limited opportunities to reuse prior computation.
  • Codex alternates model inference with tool execution, testing, compilation, and I/O, allowing GPUs to serve other agents during those pauses.
  • Prompt caching, continuous batching, paged KV caches, and chunked or disaggregated prefill/decode can improve Codex’s effective throughput despite long contexts and repeated inference calls.
  • Capacity planning therefore depends on GPU-seconds per task, cache occupancy and hit rates, latency, and SLO-compliant throughput—not GPU utilization alone.

view merged work →