Back to model catalog

Multimodal

Qwen3.5-Omni-Flash-Realtime

Realtime omni-modal model for voice assistants, multimedia analysis, interruptions, and tool calls.

Model details

Model code

qwen3.5-omni-flash-realtime

Category

Multimodal

Family

Qwen3.5 Omni

Capability

Realtime omni-modal

Modality

Text / image / audio / video -> Text / audio

Release / status

2026-03-26

Snapshot

Current model code

Source region

Console

Official detail price

Input Audio: $4.5 / 1M tokens · Output Text&Audio (Output text is not charged): $17.7 / 1M tokens

Input

Audio: $4.5 / 1M tokens

input:Text/Image/Video

$0.55 / 1M tokens

Output

Text&Audio (Output text is not charged): $17.7 / 1M tokens

Output

Text: $3.3 / 1M tokens

search_strategy:agent

$10 / 1K calls

Detail checked

Source region: International. Current highlighted entries were rechecked on July 21, 2026; long-tail entries retain their earlier official detail evidence where applicable. Final quotes still require official console confirmation for region, account route, quota, promotions, taxes, and current availability.

Buyer review

Questions to confirm before purchase

Does this exact model code support the buyer's region?
Is this for Token Plan, direct API, or both?
Does the workload need text, image, video, audio, or embeddings?
Are context length, rate limits, and quota enough for production?
Are official usage cost and ModelSmarter service fee separated?
Is there a lower-cost fallback model if usage grows?

Source note

Current highlighted model coverage was checked against the official Alibaba Cloud Model Studio model page on 2026-07-21. Existing long-tail detail summaries retain their earlier console evidence date where noted. Availability, region, account route, quota, taxes, promotions, and official terms must be confirmed before purchase.

Open official console source