OmniCapBench: Breaking Audio-Visual Captioning into Verifiable Atomic Units
OmniCapBench scores audio-visual captions as atomic, verifiable units in three tracks: entities, visual shots, audio events. Its 786 annotated videos expose identity drift and cross-modal misalignment in frontier MLLMs.
Multimodal large language models are moving away from the era of looking at one image and answering one question. They are now expected to listen and watch at the same time, to align sound with picture on a common timeline, and to remember the same person or object across tens of seconds or several minutes. Capability is rising quickly. Evaluation has not kept pace. Audio-visual captioning is an attractive diagnostic task, because a good caption must cover who is present, how the scene is cut, and when each sound occurs. Yet it is hard to measure well. A caption is free text, and one scene admits countless valid phrasings, so checking it item by item is difficult and reproducing a score is harder still. OmniCapBench, posted to arXiv on 8 October 2026 under cs.CV, targets exactly this gap.
The authors argue that existing benchmarks are caught in a coupled trade-off. The first family scores the whole caption with a single number. It offers broad coverage, but it cannot say where the model went wrong, because omissions, hallucinations and misalignments are folded into one score. The second family uses local probes, which ask about one specific detail. Localization is clear, but coverage is thin, and a different set of probes might reach a different verdict. A third problem comes from the scorer itself. An unconstrained LLM judge can swing with the prompt, the caption length and small changes in wording. Coverage, localization and stability are each available alone. Getting all three together has been the unsolved part. The central move in OmniCapBench is to change the prediction target. Free-form text stops being the final product. The evaluation object becomes a set of atomic, verifiable units, organized into three tracks: entity references, visual shots and audio events. The entity track asks whether the model keeps pointing at the same person or object over time. The shot track examines how the picture changes and what each shot contains. The audio track examines which sounds occur and when. Scoring also has two layers. Constraints that can be decided by rule, such as whether an event appears or whether two events are in the right order, use deterministic checks. Parts that truly need meaning go to a localized LLM comparison, which confines the judge to one small unit. The benchmark rests on 786 densely annotated videos.
This deep-structured design has an engineering benefit beyond cleaner scoring. Free text is hard to score reproducibly because the judge must decide inside an open space of possible phrasings. Atomic units cut that open space into many closed questions: does this person reappear in the second shot, does the explosion sound occur before or after the character turns. A closed question can often be settled by rule. Only what rules cannot settle goes to semantic comparison, and even then the judge reads a short passage with little context, so run-to-run variation should be smaller. A public benchmark needs exactly this property, because teams must be able to repeat each other's numbers. The second large benefit is that errors become sortable. According to the paper, the benchmark separates four kinds of MLLM perception error. Temporal grounding failures place an event at the wrong moment. Identity drift treats the same person as someone else later in the clip. Cross-modal misalignment pairs a sound with the wrong picture. Hallucinated descriptions invent content that the video does not contain. In a whole-caption score these problems blur together, and a developer learns only that the model is somewhat worse than a rival. With units tied to tracks, each failure mode can be tracked on its own, and teams can decide whether to change data, architecture or training objectives.
The evaluation of frontier models yields a useful picture: strong local perception, weak long-horizon audio-visual reasoning. A model can say what appears in one shot and what sound is heard at one instant, yet it struggles to stay consistent across a longer span. Identity drift and cross-modal misalignment are named as the most pronounced weaknesses. This matches a common experience among builders. Cut a long video into clips and each clip is described well, but the stitched account contradicts itself. The abstract gives no scores, so the size of the gaps between models, and which error type dominates, cannot be stated here. Readers who cite the result should go to the tables in the paper. For developers, a practical question is how to use such a benchmark. One approach is to read the three tracks separately instead of merging them into a ranking. If a model leads on visual shots but trails on entity references, the visual encoder may be sound while memory and binding across time are the problem. If it trails on audio events, audio encoding or audio-video alignment training may be thin. This breakdown gives ablation studies a clear metric and lets separate groups reproduce one another's conclusions on the same units. From an industry view, the value of such a benchmark is a roadmap. As omnimodal models move into video understanding, live assistance, accessibility captioning and robot perception, long-span identity consistency and sound-picture alignment matter more than single-frame recognition. Structured evaluation lets a team direct a limited research budget at the weakest link. Some caution is still wise. A set of 786 videos is modest in size. The way atoms are split may favor content that is easy to enumerate. Localized LLM comparison may still carry residual bias. Even so, replacing one total score with units that can be located and rechecked is a sound direction, and evaluation and model teams should keep watching it.
Sources
FAQ
What separates OmniCapBench from earlier audio-visual captioning benchmarks?
It does not score one free-form caption with one number. It makes the prediction target a set of atomic, verifiable evaluation units across three tracks: entity references, visual shots and audio events. This keeps coverage, adds localization of errors, and limits reliance on an unconstrained LLM judge.
Which kinds of model errors can the benchmark tell apart?
The abstract names four: temporal grounding failures, identity drift, cross-modal misalignment and hallucinated descriptions. Frontier models show strong local perception but weak long-horizon audio-visual reasoning, and identity drift and cross-modal misalignment stand out as the weakest areas.
Does scoring depend fully on an LLM judge?
No. Scoring has two layers. Deterministic constraint checks handle everything that can be decided by rule. Localized LLM-based semantic comparisons handle only the parts that need meaning. The judge sees a small unit, not a whole caption, so results should be more stable.