Qwen ships an agent-focused Omni-Flash model at a fifth of Gemini Flash's input price and a LiveTranslate model with 2.3-second lag
Qwen3.8-Omni-Flash handles audio and video together with a million-token context, while Qwen3.8-LiveTranslate interprets live speech across 60 languages as a hosted API.

Alibaba's Qwen team released two multimodal models within a day. Qwen3.8-Omni-Flash is described by The Decoder as Qwen's first multimodal model built for AI agents. It processes audio and video together, reasons over them, and uses tools on its own for tasks such as editing vlogs, translating short videos or summarizing films, with a context window of one million tokens. Qwen says it comes close to Gemini 3.8 Flash on audio-video benchmarks. API pricing is $0.15 per million input tokens and $0.47 per million output tokens. Qwen estimates audio input at under $0.01 per hour and 720p video with audio at one frame per second at about $0.20, excluding response costs. Gemini 3.8 Flash charges $0.75 for input and $3.75 for output per million tokens at an introductory rate that is set to double on January 1, 2027. The model is available through Qwen Studio, Qwen Cloud and the API, and open-source Qwen-MM-Plugins add video editing, speaker recognition and PDF video notes to agents including Claude Code, Gemini CLI and Qwen Code. A Qwen-Live Harness supports real-time interaction through a camera and microphone. Separately, MarkTechPost reports Qwen3.8-LiveTranslate, a simultaneous interpretation model that returns translated text and speech while the speaker is still talking. A new Interleave architecture cuts average lagging, measured by LAAL, from 2.8 to 2.3 seconds, and adds real-time speaker diarization, synchronized bilingual display and long-context disambiguation. It runs as a hosted WebSocket API on Alibaba Cloud Model Studio and QwenCloud.