🛰️ Daily AI Frontier
‹ back to 2026-07-16

Peak-End-Net: A Peak-End Rule Inspired Framework for Generalizable Video Aesthetic Assessment

Research Multimodal & Generative

Ranking

Overall 62
Content 60
Popularity 67

Observed public metrics from 1 member.

Merged summary

TL;DR — Peak-End-Net is a lightweight, interpretable framework for video aesthetic assessment (VAA) that borrows the psychological "peak-end rule" to weight salient moments and endings, achieving state-of-the-art results on two benchmarks.

  • Transfers knowledge from image aesthetic assessment (IAA) via a pretrained IAA head that produces frame-wise aesthetic priors, used to identify salient moments and guide peak-end-based temporal aggregation.
  • Adds an aesthetic rhythm encoder to model how a video evolves aesthetically over time, plus a dynamic gated fusion mechanism for robustness under distribution shift.
  • Built on a frozen ViT with only a small number of trainable parameters, making it parameter-efficient and scalable.
  • Reports state-of-the-art performance on in-domain VADB and cross-domain DIVIDE-3K; code and models released on GitHub.

Sources (1)

Peak-End-Net: A Peak-End Rule Inspired Framework for Generalizable Video Aesthetic Assessment

arXiv cs.CV Geng Li, Haiwen Li, Rui Chen, Jing Tang, Lei Sun, Xiangxiang Chu 2026-07-15 arXiv:2607.13941
Public signals Hugging Face upvotes 2
Providers: Hugging Face · Upvotes 2 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-08-15 14:34:19.035848 UTC

TL;DR — Peak-End-Net is a lightweight, interpretable framework for video aesthetic assessment (VAA) that borrows the psychological "peak-end rule" to weight salient moments and endings, achieving state-of-the-art results on two benchmarks.

  • Transfers knowledge from image aesthetic assessment (IAA) via a pretrained IAA head that produces frame-wise aesthetic priors, used to identify salient moments and guide peak-end-based temporal aggregation.
  • Adds an aesthetic rhythm encoder to model how a video evolves aesthetically over time, plus a dynamic gated fusion mechanism for robustness under distribution shift.
  • Built on a frozen ViT with only a small number of trainable parameters, making it parameter-efficient and scalable.
  • Reports state-of-the-art performance on in-domain VADB and cross-domain DIVIDE-3K; code and models released on GitHub.
item →