“OmniAgent uses on-demand actions to distill audio-visual cues into a textual memory, decoupling its reasoning complexity from video duration.”