TREE NEWS update: Alibaba’s Qwen team released Qwen3.8-Omni-Flash on September 18, its first omni-modal model built around agentic capabilities, integrating native audio-video understanding, reasoning and tool calling into a single model. It supports a 1 million token context and can actively locate key segments in long videos. On OmniVideoBench, its agentic perception mode cut token consumption by 51.8% versus static understanding, with video input cost down about 89% from Qwen3.5-Omni-Plus.
Alibaba’s Qwen Releases Qwen3.8-Omni-Flash, Cuts Video Input Cost ~89%
The headline number is cost, but the more consequential shift is architectural: folding audio-video understanding, reasoning and tool calling into one agentic model reframes video from something to be ingested into something to be navigated. Token savings from agentic perception matter because context length, not raw capability, has been the binding constraint on long-video work. The open question is whether selective perception holds up on messy real-world footage rather than benchmark segments, and whether cheaper video input pulls more RWA and media workflows toward multimodal agents.
Generated by AI for reference only.
Share on WeChat
Open WeChat → Scan → then tap "…" to send to a chat or Moments.
Tap "…" in the top-right corner to send to a chat or share to Moments.