Multimodal
Qwen3-Omni-Flash-Realtime
Realtime Qwen3 Omni Flash model for efficient multimodal understanding and speech generation.
Model details
Model code
qwen3-omni-flash-realtime
Category
Multimodal
Family
Qwen3 Omni
Capability
Realtime omni-modal
Modality
Text / image / audio / video -> Text / audio
Release / status
2025-12-04
Snapshot
Current model code
Source region
Console
Official detail price
Input Text: $0.52 / 1M tokens · Output Text (When input contains only text): $1.99 / 1M tokens
Input
Text: $0.52 / 1M tokens
Input
Audio: $4.57 / 1M tokens
Input
Vision: $0.94 / 1M tokens
Output
Text (When input contains only text): $1.99 / 1M tokens
Output
Text (When input contains images/audio/video): $3.67 / 1M tokens
Source region: International. Current highlighted entries were rechecked on July 21, 2026; long-tail entries retain their earlier official detail evidence where applicable. Final quotes still require official console confirmation for region, account route, quota, promotions, taxes, and current availability.
Buyer review
Questions to confirm before purchase
Source note
Current highlighted model coverage was checked against the official Alibaba Cloud Model Studio model page on 2026-07-21. Existing long-tail detail summaries retain their earlier console evidence date where noted. Availability, region, account route, quota, taxes, promotions, and official terms must be confirmed before purchase.
Open official console source