🛰️ Daily AI Frontier
‹ back to 2026-08-13

阿里Yuvion VL登顶多模态安全:8B越级超越397B模型

Industry & News Multimodal & Generative

Ranking

Overall 63
Content 75
Popularity 34

Observed public metrics from 1 member.

Representative image for 阿里Yuvion VL登顶多模态安全:8B越级超越397B模型

Merged summary

TL;DR - Alibaba introduced Yuvion VL, a Qwen3-VL-based model family specialized for multimodal content and AI safety. Its 8B and 32B variants reportedly outperform much larger general-purpose models on safety benchmarks through adversarial data, contrastive fine-tuning, and targeted reasoning training.

  • Yuvion VL uses a three-stage pipeline: knowledge-enhanced pretraining, instruction tuning with C2FT contrastive learning, and reasoning SFT plus reinforcement learning.
  • C2FT dynamically mines model-specific confusing examples and trains across image groups to improve fine-grained visual-semantic discrimination.
  • The 32B model averaged 76.9 on open safety evaluations and 82.8 on internal evaluations spanning 58 benchmarks; the 8B model reportedly surpassed Qwen3.5-Plus on several safety tasks.
  • Its YVRE evaluation framework covers general multimodal ability, open safety benchmarks, and industrial content-safety scenarios.

Sources (1)

阿里Yuvion VL登顶多模态安全:8B越级超越397B模型

WeChat: PaperWeekly 2026-08-12 arXiv:2606.25034
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-31 14:22:14.814340 UTC

TL;DR - Alibaba introduced Yuvion VL, a Qwen3-VL-based model family specialized for multimodal content and AI safety. Its 8B and 32B variants reportedly outperform much larger general-purpose models on safety benchmarks through adversarial data, contrastive fine-tuning, and targeted reasoning training.

  • Yuvion VL uses a three-stage pipeline: knowledge-enhanced pretraining, instruction tuning with C2FT contrastive learning, and reasoning SFT plus reinforcement learning.
  • C2FT dynamically mines model-specific confusing examples and trains across image groups to improve fine-grained visual-semantic discrimination.
  • The 32B model averaged 76.9 on open safety evaluations and 82.8 on internal evaluations spanning 58 benchmarks; the 8B model reportedly surpassed Qwen3.5-Plus on several safety tasks.
  • Its YVRE evaluation framework covers general multimodal ability, open safety benchmarks, and industrial content-safety scenarios.
item →