Keep pulling the thread on Junyang Lin.
OmniAgent achieves state-of-the-art performance among open-source models across ten benchmarks, including VideoMME and LVBench.
The 7B parameter version of OmniAgent scored 50.5% on the LVBench benchmark, outperforming the 72B parameter Qwen2.5-VL-72B model's score of 47.3%.
OmniAgent is a native omni-modal agent that formulates video understanding as a POMDP-based iterative Observation-Thought-Action cycle.
OmniAgent uses on-demand actions to distill audio-visual cues into a textual memory, decoupling its reasoning complexity from video duration.
OmniAgent exhibits positive test-time scaling, where its performance improves as the number of reasoning turns increases.