Qwen has introduced Qwen3.8-Omni-Flash, a new multimodal artificial intelligence model designed for complex audio-visual tasks.

The model supports a context window of up to one million tokens. It is now available through the Qianwen AI Platform.

Qwen3.8 is designed to work across text, audio, video and other multimodal inputs.

Its potential applications include video editing, film commentary and audio-visual summarization. It can also support music video production and real-time conversations.

Qwen3.8 Brings Major Performance Improvements

Qwen says its latest model delivers significant improvements over Qwen3.5-Omni-Plus.

According to the company, average performance across 29 evaluations improved by more than 25%.

Qwen also reports major reductions in input costs.

The company says audio input costs have dropped by more than 98%. Audio-visual input costs have declined by over 93%.

These reductions are based on Qwen’s own pricing methodology.

The company also compared Qwen3.8 with Gemini 3.8 Flash. Qwen claims its model offers similar audio-visual performance.

It also claims stronger overall audio performance. However, these comparisons are based on Qwen’s internal testing.

The model scored 71.0 on WildClawBench-MM and 69.6 on UniClawBench.

It achieved 85.1 on DailyOmni and 80.8 on StreamingBench.

Other reported scores include 63.3 on SWE-bench Pro and 92.6 on LiveCodeBench v6.

The model also recorded a score of 75.3 on CoWorkBench.

AI Agent Can Analyze Important Video Sections

One major feature involves the way the model handles long videos.

Instead of processing every frame equally, its agent can decide which sections require closer examination.

The system can begin with a user’s question and identify relevant sections of a recording.

It then gathers evidence progressively to generate its response.

Qwen says this method increased OmniVideoBench accuracy from 63.4 to 67.8.

At the same time, token usage fell by approximately 45.7%.

The model can examine characters, camera shots, lighting and sound when requested.

Meetings Can Be Converted Into Tasks

The model can process up to one hour of audio-visual meeting content.

It can identify speakers and produce transcripts from discussions.

The system can also prepare meeting minutes, identify action items and analyze potential project risks.

When connected with external tools, it can perform additional tasks.

These may include sending emails, organizing assignments or beginning coding work based on meeting instructions.

Qwen3.8 Supports Video Editing and Production

Qwen is also positioning the model for professional media workflows.

It can analyze music and use that information when planning music videos.

Its video translation workflow can support transcription, translation, voice cloning and dubbing.

Audio mixing and final review can also form part of the workflow.

For longer films, the system can assist with plot extraction and script planning.

Qwen says agents can also handle voiceovers, music, editing, rendering and quality checks.

Qwen3.8-Omni-Flash-Realtime Introduced

Qwen has also launched a real-time version called Qwen3.8-Omni-Flash-Realtime.

This version can process live audio and video while generating responses and interacting with tools.

The company says the model can combine visual information with spatial audio.

This allows it to analyze where sounds are coming from and assist with navigation or localization.

Speech recognition supports 74 languages, including Urdu and Punjabi. Speech generation is available across 29 languages.

Qwen has also expanded Qwen-MM-Plugins and open-sourced Qwen-Live Harness.

These technologies support multimodal agents, task delegation, long-term memory and real-time interactions.

In other news read more about: TikTok Introduces New Tools to Help Users Detect AI-Generated Content

Overall, Qwen3.8 expands Qwen’s focus beyond traditional chatbot functions. The model targets workflows involving long videos, meetings, coding, media production and live multimodal interaction.