2026-10-11 16:38 UTC

Alibaba's Qwen Team claims its released Qwen3.8-Omni-Flash combines 1M-token multimodal context and stronger audiovisual agent performance with over 98% lower hourly audio-input pricing than Qwen3.5-Omni-Plus, potentially making long-form media and realtime agent workflows substantially cheaper.

state: watchingheat: lowuncertainty: highconvergesscott: mediumfrontier-models multimodal-agents inference-economics agent-harnessesAlibabaQwen Team

What is this?

The case describes Alibaba’s Qwen Team announcing Qwen3.8-Omni-Flash, claiming million-token multimodal context, improved audiovisual agent performance, and over 98% lower hourly audio-input pricing than Qwen3.5-Omni-Plus. The supplied Alibaba Cloud snippet documents Qwen3.8-Flash—not explicitly the Omni variant—with million-token context and coding, agentic, and video-understanding capabilities; separate reporting describes Qwen3.5-Omni’s text, image, audio, and video support and realtime APIs. These snippets establish related offerings but do not verify the exact Qwen3.8-Omni-Flash release, its performance gains, or the claimed hourly audio-price reduction; the listed per-token prices do not establish that comparison.

Why it matters to Scott

The claimed combination of long multimodal context and cheaper audio input converges directionally with Scott’s Inference Field argument for keeping a task world resident, and warrants a cost/quality comparison against his Bulk Transcribe YouTube pipeline—not an assumption that native media beats deterministic extraction. This is a conditional evaluation opportunity: the grounding does not verify the exact Omni release or price reduction, and the radar’s Alibaba Qwen and Qwen3.8 model-cycle pages do not establish that it already tracks this specific development.
ip:source.the-inference-field-ebookip:source.how-to-read-a-youtube-video-ebookdev:project.bulk-transcribe-youtube-videos-from-playlistradar:person.alibaba-qwenradar:qwen3-8-model-cycle-validationradar:concept.inference-economicsradar:concept.long-contextradar:concept.multimodal-agents
queries asked of Scott's wikis
  • audio video ingestion economics long-form knowledge extraction
  • realtime multimodal agents latency cost constraints
  • million-token context versus RAG retrieval architecture
  • agent harness model selection capability cost evaluation
  • Qwen Alibaba API integration projects

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 569h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-18 02:27 (minted)⭐ origin echo-reconstructedAnnounces Qwen3.8-Omni-Flash with text, image, audio, and video inputs, a 1M-token context window, claimed capability gains and input-price
Qwen Team on blog (echo) · attributed from hn.story.49747925 · published time unknown
—
09-17 23:05first on hacker news · published · lag ?Alibaba releases Qwen 3.8 Omni Flash
jjcm
—
09-18 03:15first on r/LocalLLaMA · published · lag ?Qwen 3.8 Omni Flash announced
Mr_Moonsilver
—
09-17 23:05amplified on hacker news 👑hn.story.49747925
jjcm
peak 346 · 138 comments · 97% of case engagement
09-18 03:15amplified on r/LocalLLaMAreddit.post.1wjeqnn
Mr_Moonsilver
peak 12 · 18 comments · 3% of case engagement
09-18 02:20our radar first saw it · lag ?discovery anchor: hn.story.49747925—
pace: p84 vs 1032 stories at the 336h mark (now 569h old) — ahead of fractal-blt-nvme-moe-runtime (1.0x), behind r9700-nvfp4-mxfp4-fast-path (1.0x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnAlibaba releases Qwen 3.8 Omni Flash
Retrieved article excerpt

Open article · Retrieved 2026-09-18T02:21:59.303269+00:00

Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery.

## Introduction[#](https://qwen.ai/blog?id=qwen3.8-omni-flash#lFNA4)

Today, we are launching **[Qwen3.8-Omni-Flash](https://www.alibabacloud.com/help/en/model-studio/qwen-omni)**, our next-generation native omnimodal model. Its core objective is to strengthen agent capabilities in real-world productivity scenarios, advancing omnimodal models from “understanding omnimodal content” to “planning tasks, calling tools, and completing creative work.” Building on general agentic capabilities in coding, text-based knowledge work, and GUI operation, Qwen3.8-Omni-Flash further extends agentic applications centered on audio and video, delivering strong results across workflows such as video editing, music video creation, film production and commentary, audio-visual summarization, and real-time conversations.

Qwen3.8-Omni-Flash and its applications in production

Figure 1. Qwen3.8-Omni-Flash and its applications in production.

- **Qwen3.8-Omni-Flash** — now available on the [Qianwen AI Platform](https://www.qianwenai.com/):
  - Text, image, audio, and video inputs with a **1M-token context window**.

Qwen3.8-Omni-Flash supports a 1M-token context window while maintaining text performance comparable to a text-only model of the same size and delivering significant improvements in omnimodal capabilities. Across 29 evaluations1, its average score improves by more than 25% over Qwen3.5-Omni-Plus; the API price per hour of audio input decreases by more than 98%, and the price per hour of audio-visual input decreases by more than 93%2. For audio-visual agents, coding, and long-horizon tasks, the model improves by 36.5 points on WildClawBench-MM and 22.3 points on AgenticVBench, while scoring a strong 69.6 on UniClawBench. Its core capabilities also improve significantly in long-form audio and audio-visual understanding, audio-visual reasoning, audio-visual captioning, and multi-speaker recognition. For example, it gains 8.3 points on LongAudioSpan and 9.6 points on OmniVideoBench; its OmniCap-IF CSR and ISR improve by 8.5 and 14.1 points, respectively; and its AliMeeting DER and cpWER decrease from 88.11 / 89.61 to 3.35 / 17.18. By scaling data, context, and agentic environments, Qwen3.8-Omni-Flash achieves **audio-visual performance close to Gemini 3.8 Flash and overall audio performance that exceeds Gemini 3.8 Flash**. These advances also mean that audio and video are evolving from perceptual inputs into core media through which agents understand their environment, reason, and execute tasks.

1. The scope includes audio reasoning benchmarks: AliMeeting-test, AISHELL-4, MagicData-RAMC, MLC-SLM (en), WenetSpeech (Net | Meeting), FLEURS-60 ASR, FLEURS-60 S2TT, SpotSoundBench, MMAU, MMAR, MMSU, MuchoMusic-RUL, HumMusQA, MusTBench, Audio-MultiChallenge, WildSpeech, and VoiceBench; audio-visual reasoning benchmarks: DailyOmni, WorldSense, AVUT, JointAVBench, OmniCloze, OmniCap-IF, QIVD, OmniVideoBench, and StreamingBench; and audio-visual agent benchmarks: WildClawBench-MM, UniClawBench, and OmniGAIA.  
2. Pricing methodology: hourly audio or audio-visual input prices are estimated as 30 times the input cost of two minutes of source material; audio-visual input uses 720p at 1 fps. Gemini 3.8 Flash uses media\_resolution=high, Seed 2.0 Lite uses max\_frame\_tokens=384, and all other API parameters use their default values. Text input and output prices are in CNY per 1M tokens; Gemini and Muse prices are converted from USD at an exchange rate of 1 USD = 6.7191 CNY.

Qwen3.8-Omni-Flash pricing and representative benchmark comparison

Audio and video are important media for bringing agents into real-world productivity scenarios, but they also introduce a new set of system-level challenges. Long-form audio and video are costly to store, transmit, and process across multiple rounds of inference; existing agent harness frameworks lack native support for these modalities; and workflows that connect omnimodal understanding with end-to-end task execution are still at an early stage. Addressing these challenges requires models, harness tools, and runtime environments to evolve together.

To address these challenges, we use Qwen3.8-Omni-Flash to explore how to connect source understanding, task planning, tool execution, and result delivery into a complete pipeline. It supports end-to-end, long-horizon workflows such as video editing, translation, film commentary, and content creation, advancing Omni from audio-visual understanding toward autonomous action and task completion.

To this end, we have further expanded Qwen-MM-Plugins with on-demand perception, tool use, and workflow execution for long-form audio and video. We have also open-sourced Qwen-Live Harness as a native runtime for continuous, real-time omnimodal interaction. Together, they address long-horizon workflows and real-time interaction while continuing to expand the capabilities of omnimodal agents alongside the model.

[PLUGINQwen-MM-PluginsThe gateway to audio-visual productivity—connecting multimodal understanding, content creation, and agent harnesses.Get started →](https://qwen.ai/blog?id=qwen3.8-omni-flash#IbE99)[HARNESSQwen-Live HarnessThe gateway to real-time interaction—connecting audio-visual conversations, task delegation, memory, and context management.Get started →](https://qwen.ai/blog?id=qwen3.8-omni-flash#iTpXF)

## Long-Form Audio-Visual Understanding[#](https://qwen.ai/blog?id=qwen3.8-omni-flash#yejLW)

Qwen3.8-Omni-Flash brings a major upgrade to long-form audio-visual understanding—from controllable descriptions and agentic evidence gathering, to understanding meetings and advancing follow-up tasks, and finally to producing video-centered deep research reports. It does not merely process longer content, but finds relevant evidence more precisely, reasons more deeply, and acts more efficiently.

### Controllable Audio-Visual Captioning[#](https://qwen.ai/blog?id=qwen3.8-omni-flash#LQzOj)

There is no single answer to how a video should be described. Content creation prioritizes narrative, footage retrieval focuses on specific segments, and asset management depends on structure. Different applications need different video descriptions. In Qwen3.8-Omni-Flash, we have upgraded video captioning from answering "what the model saw" to understanding "what the user wants to know." Users can freely specify the subject, time range, level of detail, and output format. For the same video, the model can provide an overview, locate key segments, or **analyze character actions, camera shots, lighting, and sound in depth, producing structured results as needed**. You define what to look at, how closely to look, and how to present it.

[](https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen3.8-Omni-Flash/demo-06-audio-video-caption.en.mp4)

00:00

/

00:00

### Agentic Long-Form Audio-Visual Understanding[#](https://qwen.ai/blog?id=qwen3.8-omni-flash#ctcGg)

For videos lasting several hours, conventional approaches require the model to process the entire recording from beginning to end, even when the answer appears in only a few minutes of footage. The native Qwen3.8-Omni-Flash agent starts from the question, **independently decides what to watch and listen to, and locates key information through multiple rounds of coarse-to-fine evidence gathering**. Without processing every frame, it can focus limited compute and token budgets on the relevant segments, enabling more efficient long-form video understanding. On OmniVideoBench, Agentic Understanding **improves accuracy from 63.4 to 67.8** while reducing token consumption from 145,736 to 79,117, **a reduction of approximately 45.7%**. The table below compares the accuracy and token consumption of Static Understanding and Agentic Understanding on OmniVideoBench:

Note: Agentic mode preserves context across turns.

|  | Static Understanding | Agentic Understanding |
| --- | --- | --- |
| Accuracy (↑) | 63.4 | 67.8 |
| Tokens per query (↓) | 145,736 | 79,117 |

[](https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen3.8-Omni-Flash/demo-01-long-video-qa.en.mp4?v=20260917-1116)

00:00

/

00:00

### Long Meetings: From Minutes to Action[#](https://qwen.ai/blog?id=qwen3.8-omni-flash#S0bBD)

Multi-participant meetings are among the most complex audio-visual understanding scenarios: speakers take turns and overlap, while identities, references, and discussion topics continuously change. Qwen3.8-Omni-Flash **jointly recognizes speakers across audio and video** and natively supports **up to one hour** of audio-visual input. It can **perform speaker segmentation, content transcription, and identity alignment end to end**. Given a complete meeting video and a request, the model can map participant relationships, generate meeting minutes, identify action items, and analyze project risks, using visual information to resolve references and entity ambiguity in the audio. Combined with agents and tool use, it can also send emails, organize tasks, and even begin coding in response to meeting requirements—moving from understanding a meeting to acting on it.

[](https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen3.8-Omni-Flash/demo-07-long-meeting.en.mp4)

00:00

/

00:00

### Conducting Deep Research with Audio and Video[#](https://qwen.ai/blog?id=qwen3.8-omni-flash#vaZFL)

When users watch a video with a specific question in mind, the answer often extends beyond the video itself. Qwen3.8-Omni-Flash combines the user’s needs with the video content to identify questions worth deeper investigation, organize the key material, and search multimodal sources across the web—including images, videos, and documents. It then produces a video-centered, richly illustrated research report that helps users understand the content and solve practical problems. For example, when a user encounters color fringing around a Photoshop hair cutout, the model can break down the tutorial steps, study the principles behind Multiply and Screen blend modes, compare alternative edge-repair techniques, and explain which approach best fits the user’s situation.

Qwen3.8-Omni-Flash Conducting Deep Research with Audio and Video Demo

Expand

## Audio-Visual Production and Editing[#](https://qwen.ai/blog?id=qwen3.8-omni-flash#PBrMh)

Qwen3.8-Omni-Flash is taking audio-visual agents into a new stage: **from understanding sounds and images to independently planning, calling tools, and delivering finished videos**, bringing omnimodal intelligence into professional audio-visual content production workflows.

### Music2MV[#](https://qwen.ai/blog?id=qwen3.8-omni-flash#qIrHQ)

For music video (MV) creation, Qwen3.8-Omni-Flash can understand the structure, rhythm, mood, vocals, and instrumental changes of a user-provided song in fine detail, informing the design of characters, scenes, and shots. It can also output line-level lyrics with timestamps to align singing, subtitles, and visuals. Combined with creative tools such as Qwen-MM-Plugins, the model supports the complete workflow from music understanding and creative planning to final quality review, demonstrating strong audio-visual understanding, reasoning, and creation capabilities.

Expand all demos

Demo1 Workflow

1 / 4

[](https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen3.8-Omni-Flash/demo-03-music2mv.en.mp4)

00:00

/

00:00

### Short Drama Translation[#](https://qwen.ai/blog?id=qwen3.8-omni-flash#uyAwo)

Traditional video translation often requires repeatedly switching between transcription, translation, dubbing, and editing platforms. This complicates API calls and workflow coordination and makes it difficult to maintain consistency across character voices, dialogue duration, and visual pacing. With an agent built on Qwen3.8-Omni-Flash, users can describe their needs in a single sentence to perform speaker-aware dialogue recognition, conversational translation, character voice cloning and dubbing, audio remixing, and final quali
jjcm346138
🟧 echo.blog ⭐Announces Qwen3.8-Omni-Flash with text, image, audio, and video inputs, a 1M-token context window, claimed capability gains and input-price Qwen Team——
🟠 redditQwen 3.8 Omni Flash announced
LocalLLaMA
Mr_Moonsilver1218

Interpretation history

Decision trace