Build what's next

AI-Native Cloud with models, tools, and apps, ready out of the box.

In collaboration with

Alibaba Cloud Elite Partner
README

Send this prompt to your agent to add Skills.

> Read https://qwen.materia-logic.com/skills.md and follow the instructions to install qwencloud skills for me.
1,000,000+
Model Users
Scaling
In Seconds
99.9%
Uptime
Ultra Low
Latency
Access featured models

Featured
models for
every
modality

Experiment with leading Qwen models across text, image, video and speech.

ReasoningVUTG

Qwen3.6-Plus

The Qwen3.6 native vision-language Plus series models demonstrate exceptional performance on par with the current state-of-the-art models, with a significant improvement in overall results compared to the 3.5 series. The models have been markedly enhanced in code-related capabilities such as agentic coding, front-end programming, and Vibe coding, as well as in multi-modal general object recognition, OCR, and object localization.

Input:
$0.05–2.5/M tokens
Output:
$3–6/M tokens

1.0M

Context

65.5K

Max Out

PDFQ4_Market_Report.pdfScanning pages
QWDocument Insight

Executive recap

  • Revenue grew 18% year-over-year, led by cloud infrastructure and AI services.
  • Customer retention remained above 93%, with enterprise expansion the largest contributor.
  • Key risks include supply-chain volatility and slower recovery in overseas channels.
VG

Wan-T2V

Creates video sequences from text descriptions with smooth motion generation and cinematic aesthetics control. Features precise instruction adherence for frame-level artistic direction.

Price:
$0.1/second

300

RPM

5

Concurrency

4K - Detailed

Abstraction

Resolution

ReasoningVUTG

Qwen3.5-Open-Source

The Qwen3.5 series of open-source models is a native vision-language model designed with a hybrid architecture, integrating a linear attention mechanism and a sparse mixture-of-experts model to achieve higher inference efficiency.

Input:
$0.3/M tokens
Output:
$2.4/M tokens

262.1K

Context

65.5K

Max Out

TTS

CosyVoice

Based on a new generation of generative speech models, CosyVoice deeply integrates text understanding and speech generation technologies. It can accurately parse and interpret various text contents, converting them into natural speech like real people, bringing a highly humanized natural speech synthesis experience.

Price:
$0.26/10,000 characters

180

RPM

Multimodal-Omni

Qwen3-Omni-Flash

Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages and speech in 20 languages, generating human-like speech for precise cross-lingual communication.

Input:
$0.43/M tokens
Output:
$1.66/M tokens

60

RPM

Why Qwen Cloud?

Enterprise-Grade Security

Isolated VPCs for every deployment with dedicated infrastructure.

Global Compliance

150+ compliance certifications across the Asia-Pacific region.

Stable Performance

Guaranteed P95 first-packet latency for predictable, consistent inference.

Integrated Toolchain

Model evaluation, rapid experimentation and deployment monitoring in one suite.

Contact our
sales team

Get tailored help fast. Tell us about your use case and goals, and our sales team will get in touch.

Features

  • Access to the latest Qwen models
  • Dedicated onboarding & support
  • Flexible enterprise pricing
  • Security and compliance controls
*Required fields

Get 70M+ free tokens and start building instantly.

View now

Build what's next