🛰️ Daily AI Frontier
‹ back to 2026-08-03

Google 牺牲了图片分辨率、编辑能力和准确性换取了 「它」 的极致性价比!

雷峰网 (AI科技评论) Multimodal & Generative 2026-08-03
Representative image for Google 牺牲了图片分辨率、编辑能力和准确性换取了 「它」 的极致性价比!

TL;DR — Google released Nano Banana 2 Lite (Gemini 3.1 Flash-Lite Image) and opened Gemini Omni Flash to developers, pairing a 4-second, $0.034-per-1K-image generator with an image-to-video model to form a chainable visual production pipeline. It matters because it reframes image generation from a human creative act into a cheap, callable function inside automated agent/ad/e-commerce workflows.

  • The tradeoff: Lite caps output at 1K (1024×1024), drops Search grounding, and trims most of the "thinking" deliberation step (Thinking On remains, but abbreviated), buying 4s latency and ~half the cost of NB2 ($0.067); batch tier halves it again to $0.0168.
  • Benchmarks: text-to-image scores 1251 vs NB2's 1270, beating Flux Klein (1069), Grok Imagine (1174), and SeedreamLite (1132); editing is the weak spot at 1308, losing to NB2 (1387) and Grok Imagine (1329). Nano Banana PRO was pointedly absent from Google's comparison charts.
  • Six-scenario hands-on: Lite held up for "consumable" images (promo creatives, WeChat cover art, thumbnails) but failed accuracy-critical ones — garbled the rare character 嵴 on a mitochondria diagram and fabricated numbers with blurry small text on an Nvidia revenue chart. Dividing line: decoration → Lite, information carrier or brand asset → PRO.
  • Pipeline framing: Omni Flash ($0.10/sec, matching Veo 3.1 Fast) turns Lite output into 10s video with up to three conversational edit passes; the authors argue the real bottleneck shifts from generation to QA, proposing generate → automated validation (OCR, rule checks, similarity, moderation) → model routing → human sign-off, since verification and error costs can dwarf the per-image price.

view merged work →