阿里Yuvion VL登顶多模态安全:8B越级超越397B模型
Ranking
Overall
63
Content
75
Popularity
34
Observed public metrics from 1 member.
Merged summary
TL;DR - Alibaba introduced Yuvion VL, a Qwen3-VL-based model family specialized for multimodal content and AI safety. Its 8B and 32B variants reportedly outperform much larger general-purpose models on safety benchmarks through adversarial data, contrastive fine-tuning, and targeted reasoning training.
- Yuvion VL uses a three-stage pipeline: knowledge-enhanced pretraining, instruction tuning with C2FT contrastive learning, and reasoning SFT plus reinforcement learning.
- C2FT dynamically mines model-specific confusing examples and trains across image groups to improve fine-grained visual-semantic discrimination.
- The 32B model averaged 76.9 on open safety evaluations and 82.8 on internal evaluations spanning 58 benchmarks; the 8B model reportedly surpassed Qwen3.5-Plus on several safety tasks.
- Its YVRE evaluation framework covers general multimodal ability, open safety benchmarks, and industrial content-safety scenarios.
Sources (1)
阿里Yuvion VL登顶多模态安全:8B越级超越397B模型
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - Alibaba introduced Yuvion VL, a Qwen3-VL-based model family specialized for multimodal content and AI safety. Its 8B and 32B variants reportedly outperform much larger general-purpose models on safety benchmarks through adversarial data, contrastive fine-tuning, and targeted reasoning training.
- Yuvion VL uses a three-stage pipeline: knowledge-enhanced pretraining, instruction tuning with C2FT contrastive learning, and reasoning SFT plus reinforcement learning.
- C2FT dynamically mines model-specific confusing examples and trains across image groups to improve fine-grained visual-semantic discrimination.
- The 32B model averaged 76.9 on open safety evaluations and 82.8 on internal evaluations spanning 58 benchmarks; the 8B model reportedly surpassed Qwen3.5-Plus on several safety tasks.
- Its YVRE evaluation framework covers general multimodal ability, open safety benchmarks, and industrial content-safety scenarios.