Can LLMs Design Video Coding Tools? A Case Study on Planar Mode
Yingwen Zhang, Meng Wang, Liqiang He, Shiqi Wang
Abstract
This paper explores whether large language models (LLMs) can design video coding tools, a highly challenging task due to the intricate algorithmic coupling of tool modifications. In particular, we present an empirical case study on the Planar mode, a long-standing intra prediction tool in video coding standards. Our experiments operate within a generation-and-evaluation loop, with the LLM generating new Planar predictors, encoder trials evaluating their coding performance, and the LLM re-generating refined implementations based on the evaluation feedback. We first examine directly replacing the default Planar mode in the Fraunhofer Versatile Video Encoder (VVenC) under its faster preset. Experimental results demonstrate that the LLM-generated mode can outperform the conventional Planar mode on this lightweight toolset, achieving 0.18% bitrate savings with 0.4% complexity overhead on the standard benchmark. We further extend our evaluation to the Enhanced Compression Model (ECM). Leveraging newly introduced directional Planar modes, we investigate two integration strategies: directly replacing them, and introducing the LLM-generated predictor as an additional prediction mode with new syntax elements. The empirical results suggest that both strategies can yield coding gains under a constrained low-resolution setting. Overall, this study offers preliminary evidence and practical insights, highlighting both the potential and open challenges of LLM-based coding tool design.
Create a lesson
Related papers
The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation
Yichen Liu, Quanwei Zhang, Haozhe Wang et al.
Transfer Safety Awareness for Cross-Modal Safety Drift in Multimodal Large Language Models
Tianqi Xiao, Shiyao Cui, Minghao Zhang et al.
MuSP-Bench: Advanced Multimodal Benchmarking of Music Understanding across Score and Performance
Milan Liessens Dujardin, Song-Ze Yu, Kevin Miao
Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities
Tianfu Wang, Zhezheng Hao, Xilin Xia et al.
A Mixed-Behavior Vote Model for Multimedia Subjective Quality Votes, Means, and Variances
Jaden Pieper, Stephen D. Voran
How AI Experiences Art: Emergent Aesthetic Structure in a Self-Supervised Multimodal Embedding Space
Corey D. C. Heath