KE-VUM: Knowledge-enhanced video understanding model for fine-grained movie description.
Journal:
Neural networks : the official journal of the International Neural Network Society
Published Date:
Feb 24, 2026
Abstract
Fine-grained movie description is challenging due to limited annotations and persistent cross-modal semantic gaps, especially for clips with complex and long-range narrative structures. Existing video description models rarely exploit structured knowledge from movie scripts, resulting in shallow, subtitle-like semantics and frequent character confusion. We propose KE-VUM, a knowledge-enhanced video understanding model that performs collaborative reasoning between a video-based multimodal backbone and a script knowledge base. The backbone first produces a draft description, which is refined by a script-guided two-stage optimization strategy: entity and relationship correction using a script-derived knowledge graph, followed by narrative reconstruction aligned with script-level plot progressions to improve temporal and causal coherence. To support evaluation, we construct the MovieClip benchmark of narrative-rich movie clips and develop VD-Eval, an LLM-based framework that jointly measures Semantic Accuracy Score (SAS) and Narrative Coherence Score (NCS). Experiments on MovieClip show that KE-VUM consistently outperforms strong baselines in both descriptive accuracy and narrative coherence, achieving a 6.7% absolute improvement in BERTScore F1 and maintaining robust performance in delayed-narrative scenarios.
Authors
Keywords
No keywords available for this article.