MiMo-V2.6-Flash

An efficiency-balanced model in Xiaomi's MiMo-V2.6 series, built for agentic workloads.
$0.003 cached, $0.14 in, $0.28 out / 1M
FAST
Flexible AI inference, ready when you are.
Announcement · May 2026
Co-led by 500 Global and Georges Harik, with participation from A.Capital Ventures, Crescent Cove, Felicis, NVIDIA, Peak6, Samsung Next, Supermicro, and Upper90.
Low pay-as-you-go pricing—no long-term contracts, no hidden fees, no surprises. Startup? Enterprise? We can scale. We're here for you with simple APIs and hands-on technical support.
An inference partner that meets your needs. Whether you're optimizing for cost, latency, throughput or scale, we design the solution around your priorities. Deepinfra provides 100+ models to cover your needs.
With our zero retention policy, your inputs, outputs and user data stay private. Deepinfra is SOC 2 and ISO 27001 certified. We follow best practices in information security and privacy.
Deepinfra runs on our own cutting-edge, inference-optimized infrastructure in secure US-based data centers. Better performance and reliability for you.
Models

An efficiency-balanced model in Xiaomi's MiMo-V2.6 series, built for agentic workloads.
$0.003 cached, $0.14 in, $0.28 out / 1M

The flagship MiMo-V2.6 model for demanding long-horizon coding and multi-tool agents.
$0.004 cached, $0.435 in, $0.87 out / 1M

A multimodal mixture-of-experts model with support for long contexts.
$0.004 cached, $0.14 in, $0.42 out / 1M

A large-scale reasoning model for complex software engineering and long tasks.
$0.125 cached, $0.563 in, $2.50 out / 1M

A fast model suited to efficient coding and long-horizon agent tasks.
$0.015 cached, $0.075 in, $0.25 out / 1M

Moonshot AI's open-weight multimodal reasoning model, built for complex coding.
$0.285 cached, $2.85 in, $14.25 out / 1M

An open-weight sparse mixture-of-experts model from Qwen.
$0.20 cached, $2.00 in, $6.00 out / 1M

A DeepSeek-V4-Pro release for demanding reasoning and agentic work.
$0.20 cached, $1.30 in, $2.60 out / 1M

An efficient DeepSeek-V4-Flash release for everyday inference.
$0.015 cached, $0.06 in, $0.18 out / 1M

A flagship model for long-horizon tasks and agentic applications.
$0.105 cached, $0.563 in, $1.80 out / 1M

Built for frontier reasoning, orchestration, coding agents and deep research.
$0.10 cached, $0.50 in, $2.20 out / 1M

An efficiency-focused mixture-of-experts model for responsive applications.
$0.018 cached, $0.09 in, $0.18 out / 1M
The measures that matter for speed, scale, stability and spend
Tokens per second
Time to first token
Requests per second
Infrastructure at scale
DeepCluster
Dedicated hardware, procured and operated by Deepinfra. Full ownership, Tier 3 datacenter, 99.982% uptime SLA.
Explore GPU clustersNVIDIA B300 · 5-year term
vs $6.50 /GPU-hr on public cloud