AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
Abstract
While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate visual transformations on silent clips or isolated audio editing, leaving complex audio-visual editing and cross-modal consistency underexplored. We introduce AVE-Compass, a comprehensive benchmark with 145 curated source videos, 196 audio-visually coupled editing instructions, and 2,688 fine-grained checklist items. It evaluates Instruction Following, Fidelity Preserving, Realism, and Editing Intent through checklist-based MLLM judging and a dedicated realism rubric, complemented by automated cross-modal, video, and audio metrics. Extensive evaluation shows that state-of-the-art models still struggle to execute cross-modal instructions while preserving non-target content. We further propose AVE-Agent, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively improves editing results through self-reflection and evaluator feedback. AVE-Agent improves instruction execution, Fidelity Preserving, and audio-visual alignment in joint editing while maintaining competitive perceptual quality.
Community
AVE-Compass provides a comprehensive and realistic benchmark for evaluating whether audio-visual editing systems can follow instructions while preserving unedited content, cross-modal consistency, and perceptual quality.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing (2026)
- MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation (2026)
- KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation (2026)
- AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning (2026)
- CoT-Edit: Let CoT Guide Instruction Video Editing (2026)
- MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control (2026)
- Empowering Long-form Omni-modal Understanding with Robust Audio Perception (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.24821 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper