Retrieved article excerpt
Open article · Retrieved 2026-09-18T02:21:59.303269+00:00
Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery.
## Introduction[#](https://qwen.ai/blog?id=qwen3.8-omni-flash#lFNA4)
Today, we are launching **[Qwen3.8-Omni-Flash](https://www.alibabacloud.com/help/en/model-studio/qwen-omni)**, our next-generation native omnimodal model. Its core objective is to strengthen agent capabilities in real-world productivity scenarios, advancing omnimodal models from “understanding omnimodal content” to “planning tasks, calling tools, and completing creative work.” Building on general agentic capabilities in coding, text-based knowledge work, and GUI operation, Qwen3.8-Omni-Flash further extends agentic applications centered on audio and video, delivering strong results across workflows such as video editing, music video creation, film production and commentary, audio-visual summarization, and real-time conversations.
Qwen3.8-Omni-Flash and its applications in production
Figure 1. Qwen3.8-Omni-Flash and its applications in production.
- **Qwen3.8-Omni-Flash** — now available on the [Qianwen AI Platform](https://www.qianwenai.com/):
- Text, image, audio, and video inputs with a **1M-token context window**.
Qwen3.8-Omni-Flash supports a 1M-token context window while maintaining text performance comparable to a text-only model of the same size and delivering significant improvements in omnimodal capabilities. Across 29 evaluations1, its average score improves by more than 25% over Qwen3.5-Omni-Plus; the API price per hour of audio input decreases by more than 98%, and the price per hour of audio-visual input decreases by more than 93%2. For audio-visual agents, coding, and long-horizon tasks, the model improves by 36.5 points on WildClawBench-MM and 22.3 points on AgenticVBench, while scoring a strong 69.6 on UniClawBench. Its core capabilities also improve significantly in long-form audio and audio-visual understanding, audio-visual reasoning, audio-visual captioning, and multi-speaker recognition. For example, it gains 8.3 points on LongAudioSpan and 9.6 points on OmniVideoBench; its OmniCap-IF CSR and ISR improve by 8.5 and 14.1 points, respectively; and its AliMeeting DER and cpWER decrease from 88.11 / 89.61 to 3.35 / 17.18. By scaling data, context, and agentic environments, Qwen3.8-Omni-Flash achieves **audio-visual performance close to Gemini 3.8 Flash and overall audio performance that exceeds Gemini 3.8 Flash**. These advances also mean that audio and video are evolving from perceptual inputs into core media through which agents understand their environment, reason, and execute tasks.
1. The scope includes audio reasoning benchmarks: AliMeeting-test, AISHELL-4, MagicData-RAMC, MLC-SLM (en), WenetSpeech (Net | Meeting), FLEURS-60 ASR, FLEURS-60 S2TT, SpotSoundBench, MMAU, MMAR, MMSU, MuchoMusic-RUL, HumMusQA, MusTBench, Audio-MultiChallenge, WildSpeech, and VoiceBench; audio-visual reasoning benchmarks: DailyOmni, WorldSense, AVUT, JointAVBench, OmniCloze, OmniCap-IF, QIVD, OmniVideoBench, and StreamingBench; and audio-visual agent benchmarks: WildClawBench-MM, UniClawBench, and OmniGAIA.
2. Pricing methodology: hourly audio or audio-visual input prices are estimated as 30 times the input cost of two minutes of source material; audio-visual input uses 720p at 1 fps. Gemini 3.8 Flash uses media\_resolution=high, Seed 2.0 Lite uses max\_frame\_tokens=384, and all other API parameters use their default values. Text input and output prices are in CNY per 1M tokens; Gemini and Muse prices are converted from USD at an exchange rate of 1 USD = 6.7191 CNY.
Qwen3.8-Omni-Flash pricing and representative benchmark comparison
Audio and video are important media for bringing agents into real-world productivity scenarios, but they also introduce a new set of system-level challenges. Long-form audio and video are costly to store, transmit, and process across multiple rounds of inference; existing agent harness frameworks lack native support for these modalities; and workflows that connect omnimodal understanding with end-to-end task execution are still at an early stage. Addressing these challenges requires models, harness tools, and runtime environments to evolve together.
To address these challenges, we use Qwen3.8-Omni-Flash to explore how to connect source understanding, task planning, tool execution, and result delivery into a complete pipeline. It supports end-to-end, long-horizon workflows such as video editing, translation, film commentary, and content creation, advancing Omni from audio-visual understanding toward autonomous action and task completion.
To this end, we have further expanded Qwen-MM-Plugins with on-demand perception, tool use, and workflow execution for long-form audio and video. We have also open-sourced Qwen-Live Harness as a native runtime for continuous, real-time omnimodal interaction. Together, they address long-horizon workflows and real-time interaction while continuing to expand the capabilities of omnimodal agents alongside the model.
[PLUGINQwen-MM-PluginsThe gateway to audio-visual productivity—connecting multimodal understanding, content creation, and agent harnesses.Get started →](https://qwen.ai/blog?id=qwen3.8-omni-flash#IbE99)[HARNESSQwen-Live HarnessThe gateway to real-time interaction—connecting audio-visual conversations, task delegation, memory, and context management.Get started →](https://qwen.ai/blog?id=qwen3.8-omni-flash#iTpXF)
## Long-Form Audio-Visual Understanding[#](https://qwen.ai/blog?id=qwen3.8-omni-flash#yejLW)
Qwen3.8-Omni-Flash brings a major upgrade to long-form audio-visual understanding—from controllable descriptions and agentic evidence gathering, to understanding meetings and advancing follow-up tasks, and finally to producing video-centered deep research reports. It does not merely process longer content, but finds relevant evidence more precisely, reasons more deeply, and acts more efficiently.
### Controllable Audio-Visual Captioning[#](https://qwen.ai/blog?id=qwen3.8-omni-flash#LQzOj)
There is no single answer to how a video should be described. Content creation prioritizes narrative, footage retrieval focuses on specific segments, and asset management depends on structure. Different applications need different video descriptions. In Qwen3.8-Omni-Flash, we have upgraded video captioning from answering "what the model saw" to understanding "what the user wants to know." Users can freely specify the subject, time range, level of detail, and output format. For the same video, the model can provide an overview, locate key segments, or **analyze character actions, camera shots, lighting, and sound in depth, producing structured results as needed**. You define what to look at, how closely to look, and how to present it.
[](https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen3.8-Omni-Flash/demo-06-audio-video-caption.en.mp4)
00:00
/
00:00
### Agentic Long-Form Audio-Visual Understanding[#](https://qwen.ai/blog?id=qwen3.8-omni-flash#ctcGg)
For videos lasting several hours, conventional approaches require the model to process the entire recording from beginning to end, even when the answer appears in only a few minutes of footage. The native Qwen3.8-Omni-Flash agent starts from the question, **independently decides what to watch and listen to, and locates key information through multiple rounds of coarse-to-fine evidence gathering**. Without processing every frame, it can focus limited compute and token budgets on the relevant segments, enabling more efficient long-form video understanding. On OmniVideoBench, Agentic Understanding **improves accuracy from 63.4 to 67.8** while reducing token consumption from 145,736 to 79,117, **a reduction of approximately 45.7%**. The table below compares the accuracy and token consumption of Static Understanding and Agentic Understanding on OmniVideoBench:
Note: Agentic mode preserves context across turns.
| | Static Understanding | Agentic Understanding |
| --- | --- | --- |
| Accuracy (↑) | 63.4 | 67.8 |
| Tokens per query (↓) | 145,736 | 79,117 |
[](https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen3.8-Omni-Flash/demo-01-long-video-qa.en.mp4?v=20260917-1116)
00:00
/
00:00
### Long Meetings: From Minutes to Action[#](https://qwen.ai/blog?id=qwen3.8-omni-flash#S0bBD)
Multi-participant meetings are among the most complex audio-visual understanding scenarios: speakers take turns and overlap, while identities, references, and discussion topics continuously change. Qwen3.8-Omni-Flash **jointly recognizes speakers across audio and video** and natively supports **up to one hour** of audio-visual input. It can **perform speaker segmentation, content transcription, and identity alignment end to end**. Given a complete meeting video and a request, the model can map participant relationships, generate meeting minutes, identify action items, and analyze project risks, using visual information to resolve references and entity ambiguity in the audio. Combined with agents and tool use, it can also send emails, organize tasks, and even begin coding in response to meeting requirements—moving from understanding a meeting to acting on it.
[](https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen3.8-Omni-Flash/demo-07-long-meeting.en.mp4)
00:00
/
00:00
### Conducting Deep Research with Audio and Video[#](https://qwen.ai/blog?id=qwen3.8-omni-flash#vaZFL)
When users watch a video with a specific question in mind, the answer often extends beyond the video itself. Qwen3.8-Omni-Flash combines the user’s needs with the video content to identify questions worth deeper investigation, organize the key material, and search multimodal sources across the web—including images, videos, and documents. It then produces a video-centered, richly illustrated research report that helps users understand the content and solve practical problems. For example, when a user encounters color fringing around a Photoshop hair cutout, the model can break down the tutorial steps, study the principles behind Multiply and Screen blend modes, compare alternative edge-repair techniques, and explain which approach best fits the user’s situation.
Qwen3.8-Omni-Flash Conducting Deep Research with Audio and Video Demo
Expand
## Audio-Visual Production and Editing[#](https://qwen.ai/blog?id=qwen3.8-omni-flash#PBrMh)
Qwen3.8-Omni-Flash is taking audio-visual agents into a new stage: **from understanding sounds and images to independently planning, calling tools, and delivering finished videos**, bringing omnimodal intelligence into professional audio-visual content production workflows.
### Music2MV[#](https://qwen.ai/blog?id=qwen3.8-omni-flash#qIrHQ)
For music video (MV) creation, Qwen3.8-Omni-Flash can understand the structure, rhythm, mood, vocals, and instrumental changes of a user-provided song in fine detail, informing the design of characters, scenes, and shots. It can also output line-level lyrics with timestamps to align singing, subtitles, and visuals. Combined with creative tools such as Qwen-MM-Plugins, the model supports the complete workflow from music understanding and creative planning to final quality review, demonstrating strong audio-visual understanding, reasoning, and creation capabilities.
Expand all demos
Demo1 Workflow
1 / 4
[](https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen3.8-Omni-Flash/demo-03-music2mv.en.mp4)
00:00
/
00:00
### Short Drama Translation[#](https://qwen.ai/blog?id=qwen3.8-omni-flash#uyAwo)
Traditional video translation often requires repeatedly switching between transcription, translation, dubbing, and editing platforms. This complicates API calls and workflow coordination and makes it difficult to maintain consistency across character voices, dialogue duration, and visual pacing. With an agent built on Qwen3.8-Omni-Flash, users can describe their needs in a single sentence to perform speaker-aware dialogue recognition, conversational translation, character voice cloning and dubbing, audio remixing, and final quali