OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning

Published · AI Daily — AI-assisted deep research, methodology & disclosure

OmniSeek turns an Omni Large Language Model into an active, multi-turn reasoning agent that natively calls tools to look or listen on demand, rather than processing an entire audio-visual sequence in one forward pass. It iteratively retrieves sparse evidence windows and appends them back into context. Cold-started on a synthesized 170K-trajectory multi-hop chain-of-thought corpus, the policy is then refined via two-stage reinforcement learning with verifiable rewards, plus an Audio-Visual Necessity objective that penalizes single-modality shortcuts, yielding consistent gains across audio-visual reasoning benchmarks.

Background and Problem Definition

Most Omni-LLMs that jointly handle audio, video, and text today still process a full audio-visual sequence in a single forward pass — the entire recording is encoded once and handed to the model for one-shot reasoning. This passive paradigm runs into two structural problems once sequences get long.

First, the signal is diluted: the handful of seconds that actually carry the decisive evidence are buried inside hours of audio or thousands of video frames, yet compute scales with the whole sequence regardless of where the evidence sits. Second, the model has no mechanism to go back and check itself. If the first pass under-samples a modality, or the correct answer genuinely requires combining audio and visual cues, a single-pass model simply commits to an answer built on incomplete evidence — there is no retry, no re-examination, no way to ask "did I actually look at the right place?" OmniSeek reframes audio-visual understanding as an active evidence-acquisition problem: the model should form a hypothesis, identify what evidence is still missing, decide whether to look or listen and over which window, and repeat until the evidence is sufficient, rather than consuming the entire input passively in one shot.

Architectural Core and Technical Principles

OmniSeek wraps an Omni-LLM into a multi-turn reasoning agent with native tool use. Instead of a single forward pass, reasoning proceeds as an iterative protocol: at each turn the model reasons over the current context, decides whether to invoke a "look" or "listen" tool, and specifies the temporal window of the segment it wants retrieved. The retrieved raw audio or visual segment is appended back into the context, and the cycle repeats until the model judges the evidence sufficient to answer.

To cold-start this multi-turn tool-use behavior, the authors built a data engine that synthesizes OmniTraj-170K, a large corpus of multi-hop chain-of-thought trajectories with evidence-retrieval and reasoning steps interleaved across modalities. Training proceeds in two stages: supervised fine-tuning on these synthesized trajectories first instills the pattern of "reason, decide what to retrieve, splice the evidence back in, reason again," and a subsequent two-stage reinforcement learning phase with verifiable rewards further sharpens the retrieval-and-reasoning policy against task-level correctness signals rather than imitation alone. The most distinctive design choice is the Audio-Visual Necessity objective, which explicitly rewards trajectories whose correct conclusion genuinely depends on evidence from both modalities — directly discouraging the model from exploiting single-modality shortcuts that happen to get the right answer without doing the cross-modal reasoning the task actually demands.

Practical Evaluation and Applications

The resulting capability is adaptive cross-modal evidence seeking: rather than treating the whole input uniformly, the model decides how many retrieval rounds to take, which temporal windows to target, and whether to prioritize audio or visual inspection based on the difficulty and evidence distribution of the specific question.

Across a broad suite of audio-visual reasoning benchmarks, OmniSeek consistently outperforms baselines that ingest the entire sequence in one pass. Crucially, the gains are not an artifact of improving one modality in isolation — they concentrate on cases that genuinely require combining evidence from both audio and video, which is the direct empirical signature the Audio-Visual Necessity objective was designed to produce, and it supports the claim that the model learned a real retrieval-based reasoning skill rather than a single-modality shortcut.

Industry Impact and Outlook

OmniSeek signals a meaningful shift in multimodal agent engineering: it imports the native tool-use paradigm, already mature in text-only agents, into audio-visual understanding, letting the model itself decide when it needs more perceptual input instead of relying on a fixed external preprocessing pipeline to pre-trim or summarize long recordings.

This has direct relevance for long-video understanding, meeting-transcript analysis, surveillance and security review, and fact-checking workflows that require cross-modal corroboration — all settings characterized by sparse evidence embedded in long context, exactly the regime OmniSeek targets. Looking forward, plausible extensions include widening the tool palette beyond look/listen to richer perception and retrieval operations (external retrieval-augmented lookups, dedicated object-tracking or speech-recognition sub-tools), and coupling this multi-turn evidence-seeking loop with longer, real-time streaming input for online audio-visual reasoning.

Sources

FAQ

What is the core mechanism that distinguishes OmniSeek from passive single-pass Omni-LLMs?

OmniSeek turns reasoning into an iterative multi-turn protocol where the model natively calls 'look' or 'listen' tools to retrieve specific temporal windows of audio or video evidence and appends them back into context, rather than encoding the whole sequence once for a single forward pass.

How is the model cold-started to perform multi-turn tool use?

The authors build a data engine that synthesizes OmniTraj-170K, a corpus of multi-hop chain-of-thought trajectories with interleaved audio-visual evidence retrieval and reasoning steps, and use it for supervised fine-tuning before reinforcement learning.

What does the Audio-Visual Necessity objective prevent?

It prevents the model from exploiting single-modality shortcuts — trajectories that happen to reach the right answer using only one modality's cues without performing the cross-modal reasoning the task genuinely requires are not rewarded.