FAST

deepinfra AI Inference

Flexible AI inference, ready when you are.

Try the inference cloud

Start with an idea or choose an example.

Pay as you go · no long-term contracts

OR CHOOSE AN EXAMPLE

DeepSeek · NVIDIA · Qwen · and more

VercelAdobeUBISOFTIconic HeartsLATITUDECupidBot.ai

Announcement · May 2026

We've raised $107M Series B to scale the inference cloud

Co-led by 500 Global and Georges Harik, with participation from A.Capital Ventures, Crescent Cove, Felicis, NVIDIA, Peak6, Samsung Next, Supermicro, and Upper90.

Explore what's possible Built for what comes next.
$107MSeries B funding
25xtoken volume since Series A

Scale to trillions of tokens without breaking the bank

Low pay-as-you-go pricing—no long-term contracts, no hidden fees, no surprises. Startup? Enterprise? We can scale. We're here for you with simple APIs and hands-on technical support.

Isometric blue and turquoise inference platform with a glowing compute cube, small servers and flowing data across a dark digital landscape.
Isometric luminous blue computing tile with modular blocks and fine violet data paths on a dark background.

Inference Tailored to You

An inference partner that meets your needs. Whether you're optimizing for cost, latency, throughput or scale, we design the solution around your priorities. Deepinfra provides 100+ models to cover your needs.

Zero Retention. Compliant. Secure.

With our zero retention policy, your inputs, outputs and user data stay private. Deepinfra is SOC 2 and ISO 27001 certified. We follow best practices in information security and privacy.

Glowing cyan privacy shield above a precise blue circuit-board network, shown as an isometric technical illustration against deep navy.
Cluster of white and electric-blue server cabinets on an isometric data-center platform with turquoise lights and a dark navy background.

Our Hardware. Our Data Centers. Your Performance Edge.

Deepinfra runs on our own cutting-edge, inference-optimized infrastructure in secure US-based data centers. Better performance and reliability for you.

Models

Explore our Featured Models

View All text-generation text-to-speech text-to-image
XiaomiMiMo/text-generation

MiMo-V2.6-Flash

MiMo-V2.6-Flash model cover

An efficiency-balanced model in Xiaomi's MiMo-V2.6 series, built for agentic workloads.

$0.003 cached, $0.14 in, $0.28 out / 1M

ZDRfp81024k
XiaomiMiMo/text-generation

MiMo-V2.6-Pro

MiMo-V2.6-Pro model cover

The flagship MiMo-V2.6 model for demanding long-horizon coding and multi-tool agents.

$0.004 cached, $0.435 in, $0.87 out / 1M

ZDRfp81024k
deepseek-ai/text-generation

DeepSeek-V4.1-Flash

DeepSeek-V4.1-Flash model cover

A multimodal mixture-of-experts model with support for long contexts.

$0.004 cached, $0.14 in, $0.42 out / 1M

ZDRPriorityFlex1024k
zai-org/text-generation

GLM-5.3

GLM-5.3 model cover

A large-scale reasoning model for complex software engineering and long tasks.

$0.125 cached, $0.563 in, $2.50 out / 1M

ZDRFlex1024k
zai-org/text-generation

GLM-5.3-Flash

GLM-5.3-Flash model cover

A fast model suited to efficient coding and long-horizon agent tasks.

$0.015 cached, $0.075 in, $0.25 out / 1M

ZDRfp41024k
moonshotai/text-generation

Kimi-K3

Kimi-K3 model cover

Moonshot AI's open-weight multimodal reasoning model, built for complex coding.

$0.285 cached, $2.85 in, $14.25 out / 1M

ZDRCache retention1024k
Qwen/text-generation

Qwen3.8-2.4T-A95B

Qwen3.8-2.4T-A95B model cover

An open-weight sparse mixture-of-experts model from Qwen.

$0.20 cached, $2.00 in, $6.00 out / 1M

ZDRfp41024k
deepseek-ai/text-generation

DeepSeek-V4-Pro-0813

DeepSeek-V4-Pro-0813 model cover

A DeepSeek-V4-Pro release for demanding reasoning and agentic work.

$0.20 cached, $1.30 in, $2.60 out / 1M

ZDRPriorityFlex
deepseek-ai/text-generation

DeepSeek-V4-Flash-0731

DeepSeek-V4-Flash-0731 model cover

An efficient DeepSeek-V4-Flash release for everyday inference.

$0.015 cached, $0.06 in, $0.18 out / 1M

ZDRPriorityFlex
zai-org/text-generation

GLM-5.2

GLM-5.2 model cover

A flagship model for long-horizon tasks and agentic applications.

$0.105 cached, $0.563 in, $1.80 out / 1M

ZDRPriorityFlex
nvidia/text-generation

NVIDIA-Nemotron-3-Ultra-550B-A55B

NVIDIA Nemotron 3 Ultra model cover

Built for frontier reasoning, orchestration, coding agents and deep research.

$0.10 cached, $0.50 in, $2.20 out / 1M

ZDRPriorityFlex
deepseek-ai/text-generation

DeepSeek-V4-Flash

DeepSeek-V4-Flash model cover

An efficiency-focused mixture-of-experts model for responsive applications.

$0.018 cached, $0.09 in, $0.18 out / 1M

ZDRPriorityFlex

AI Inference Metrics

The measures that matter for speed, scale, stability and spend

Throughput

Tokens per second

Responsiveness

Time to first token

Capacity

Requests per second

Compute

Infrastructure at scale

DeepCluster

Your own NVIDIA B300
GPU cluster

Dedicated hardware, procured and operated by Deepinfra. Full ownership, Tier 3 datacenter, 99.982% uptime SLA.

Explore GPU clusters

NVIDIA B300 · 5-year term

$1.98/GPU-hr

vs $6.50 /GPU-hr on public cloud

70%cheaper than cloud
288 GBHBM3e per GPU
256–5,000GPUs available
Try AI models
Try AI models