Simple, transparent pricing

Pay only for what you use. $1 free credit on signup. All rates are pulled live from the same database that bills your account.

Rates last verified Jul 16, 2026 (UTC)

Inference API

Pay per token across all platform models. OpenAI-compatible endpoints, multi-region routing.

Language

ModelInputOutputNotes
DeepSeek-V4-Flash
DeepSeek-V4-Flash
DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens.
$0.16/1M$0.33/1M
GLM-5.2
GLM-5.2
GLM-5.2 marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: Solid 1M Context: A solid 1M-token context that stably sustains long-horizon work Advanced Coding with Flexible Effort: Stronger coding capabilities with multiple thinking effort levels to balance performance and latency Improved Architecture: We propose IndexShare, which reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9× at a 1M context length. We also improve GLM-5.2’s MTP layer for speculative decoding, increasing the acceptance length by up to 20% Pure Open: An MIT open-source license — no regional limits, technical access without borders
$1.00/1M$4.00/1M
gpt-oss-20b
gpt-oss-20b
gpt-oss-20b, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases. for lower latency, and local or specialized use cases (21B parameters with 3.6B active parameters)
$0.05/1M$0.18/1MContext: 125K
Llama-3.1-8B-Instruct
llama-3.1-8b-instruct
The Meta Llama 3.1 collection of multilingual large language models (LLMs) is a collection of pretrained and instruction tuned generative models in 8B, 70B and 405B sizes (text in/text out). The Llama 3.1 instruction tuned text only models (8B, 70B, 405B) are optimized for multilingual dialogue use cases and outperform many of the available open source and closed chat models on common industry benchmarks.
$0.10/1M$0.10/1MContext: 125K
Qwen2.5-7B-Instruct
qwen2.5-7b-instruct
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2: Significantly more knowledge and has greatly improved capabilities in coding and mathematics, thanks to our specialized expert models in these domains. Significant improvements in instruction following, generating long texts (over 8K tokens), understanding structured data (e.g, tables), and generating structured outputs especially JSON. More resilient to the diversity of system prompts, enhancing role-play implementation and condition-setting for chatbots. Long-context Support up to 128K tokens and can generate up to 8K tokens. Multilingual support for over 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai, Arabic, and more.
$0.20/1M$0.20/1MContext: 16K
Qwen3-235B-A22B
Qwen3-235B-A22B
Qwen3-235B-A22B has the following features: Type: Causal Language Models Training Stage: Pretraining & Post-training Number of Parameters: 235B in total and 22B activated Number of Paramaters (Non-Embedding): 234B Number of Layers: 94 Number of Attention Heads (GQA): 64 for Q and 4 for KV Number of Experts: 128 Number of Activated Experts: 8 Context Length: 32,768 natively and 131,072 tokens with YaRN.
$0.33/1M$1.33/1M
qwen3-coder-30b-a3b-instruct
qwen3-coder-30b-a3b-instruct
Qwen3-Coder-30B-A3B-Instruct has the following features: Type: Causal Language Models Training Stage: Pretraining & Post-training Number of Parameters: 30.5B in total and 3.3B activated Number of Layers: 48 Number of Attention Heads (GQA): 32 for Q and 4 for KV Number of Experts: 128 Number of Activated Experts: 8 Context Length: 262,144 natively.
$0.10/1M$0.30/1MContext: 32K

Vision

ModelInputOutputNotes
Gemma-4-31B-IT
gemma-4-31b-it
Gemma 4 models are designed to deliver frontier-level performance at each size, targeting deployment scenarios from mobile and edge devices (E2B, E4B) to consumer GPUs and workstations (26B A4B, 31B). They are well-suited for reasoning, agentic workflows, coding, and multimodal understanding. The models employ a hybrid attention mechanism that interleaves local sliding window attention with full global attention, ensuring the final layer is always global. This hybrid design delivers the processing speed and low memory footprint of a lightweight model without sacrificing the deep awareness required for complex, long-context tasks. To optimize memory for long contexts, global layers feature unified Keys and Values, and apply Proportional RoPE (p-RoPE).
$0.50/1M$0.50/1MContext: 32K
Kimi-K2.6
Kimi-K2.6
Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration. Key Features Long-Horizon Coding: K2.6 achieves significant improvements on complex, end-to-end coding tasks, generalizing robustly across programming languages (Rust, Go, Python) and domains spanning front-end, DevOps, and performance optimization. Coding-Driven Design: K2.6 is capable of transforming simple prompts and visual inputs into production-ready interfaces and lightweight full-stack workflows, generating structured layouts, interactive elements, and rich animations with deliberate aesthetic precision. Elevated Agent Swarm: Scaling horizontally to 300 sub-agents executing 4,000 coordinated steps, K2.6 can dynamically decompose tasks into parallel, domain-specialized subtasks, delivering end-to-end outputs from documents to websites to spreadsheets in a single autonomous run. Proactive & Open Orchestration: For autonomous tasks, K2.6 demonstrates strong performance in powering persistent, 24/7 background agents that proactively manage schedules, execute code, and orchestrate cross-platform operations without human oversight.
$1.08/1M$4.50/1M
qwen3-omni-30b-a3b-instruct
qwen3-omni-30b-a3b-instruct
$0.40/1M$0.80/1M
qwen3-vl-8b-instruct
qwen3-vl-8b-instruct
Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date. This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities. Available in Dense and MoE architectures that scale from edge to cloud, with Instruct and reasoning‑enhanced Thinking editions for flexible, on‑demand deployment. Key Enhancements: Visual Agent: Operates PC/mobile GUIs—recognizes elements, understands functions, invokes tools, completes tasks. Visual Coding Boost: Generates Draw.io/HTML/CSS/JS from images/videos. Advanced Spatial Perception: Judges object positions, viewpoints, and occlusions; provides stronger 2D grounding and enables 3D grounding for spatial reasoning and embodied AI. Long Context & Video Understanding: Native 256K context, expandable to 1M; handles books and hours-long video with full recall and second-level indexing. Enhanced Multimodal Reasoning: Excels in STEM/Math—causal analysis and logical, evidence-based answers. Upgraded Visual Recognition: Broader, higher-quality pretraining is able to “recognize everything”—celebrities, anime, products, landmarks, flora/fauna, etc. Expanded OCR: Supports 32 languages (up from 19); robust in low light, blur, and tilt; better with rare/ancient characters and jargon; improved long-document structure parsing. Text Understanding on par with pure LLMs: Seamless text–vision fusion for lossless, unified comprehension.
$0.15/1M$0.50/1MContext: 32K
Qwen3.5-35B-A3B
qwen3.5-35b-a3b
Model Overview Type: Causal Language Model with Vision Encoder Training Stage: Pre-training & Post-training Language Model Number of Parameters: 35B in total and 3B activated Hidden Dimension: 2048 Token Embedding: 248320 (Padded) Number of Layers: 40 Hidden Layout: 10 × (3 × (Gated DeltaNet → MoE) → 1 × (Gated Attention → MoE)) Gated DeltaNet: Number of Linear Attention Heads: 32 for V and 16 for QK Head Dimension: 128 Gated Attention: Number of Attention Heads: 16 for Q and 2 for KV Head Dimension: 256 Rotary Position Embedding Dimension: 64 Mixture Of Experts Number of Experts: 256 Number of Activated Experts: 8 Routed + 1 Shared Expert Intermediate Dimension: 512 LM Output: 248320 (Padded) MTP: trained with multi-steps Context Length: 262,144 natively and extensible up to 1,010,000 tokens.
$0.40/1M$0.40/1M

Embedding

ModelInputOutputNotes
Jina-Embeddings-V3
jina-embeddings-v3
jina-embeddings-v3 is a multilingual multi-task text embedding model designed for a variety of NLP applications. Based on the Jina-XLM-RoBERTa architecture, this model supports Rotary Position Embeddings to handle long input sequences up to 8192 tokens.
$0.02/1M
Jina-Embeddings-V4
jina-embeddings-v4
jina-embeddings-v4 is a universal embedding model for multimodal and multilingual retrieval. The model is specially designed for complex document retrieval, including visually rich documents with charts, tables, and illustrations.
$0.12/1M
qwen3-embedding-0.6b
qwen3-embedding-0.6b
Qwen3-Embedding-0.6B has the following features: Model Type: Text Embedding Supported Languages: 100+ Languages Number of Parameters: 0.6B Context Length: 32k Embedding Dimension: Up to 1024, supports user-defined output dimensions ranging from 32 to 1024
$0.12/1M
qwen3-embedding-4b
qwen3-embedding-4b
Qwen3-Embedding-4B has the following features: Model Type: Text Embedding Supported Languages: 100+ Languages Number of Paramaters: 4B Context Length: 32k Embedding Dimension: Up to 2560, supports user-defined output dimensions ranging from 32 to 2560
$0.12/1M
qwen3-embedding-8b
qwen3-embedding-8b
Qwen3-Embedding-8B has the following features: Model Type: Text Embedding Supported Languages: 100+ Languages Number of Paramaters: 8B Context Length: 32k Embedding Dimension: Up to 4096, supports user-defined output dimensions ranging from 32 to 4096
$0.12/1M

Reranker

ModelInputOutputNotes
BGE-Reranker-V2-M3
bge-reranker-v2-m3
bge-reranker-v2-m3,Different from embedding model, reranker uses question and document as input and directly output similarity instead of embedding. You can get a relevance score by inputting query and passage to the reranker. And the score can be mapped to a float value in [0,1] by sigmoid function.
$0.03/1M

Image

ModelInputOutputNotes
FLUX.2 Klein
flux2-klein
The FLUX.2 [klein] model family are our fastest image models to date. FLUX.2 [klein] unifies generation and editing in a single compact architecture, delivering state-of-the-art quality with end-to-end inference in as low as under a second. Built for applications that require real-time image generation without sacrificing quality. FLUX.2 [klein] 9B is a 9 billion parameter rectified flow transformer capable of generating images from text descriptions and supports multi-reference editing capabilities. Our flagship small model. Defines the Pareto frontier for quality vs. latency across text-to-image, single-reference editing, and multi-reference generation. Matches or exceeds models 5x its size—in under half a second. Built on a 9B flow model with 8B Qwen3 text embedder, step-distilled to 4 inference steps.
$20.00/1M out tokTokens / item: 1,000
Qwen-Image
qwen-image
Qwen-Image, an image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing. Experiments show strong general capabilities in both image generation and editing, with exceptional performance in text rendering, especially for Chinese.
$30.00/1M out tokTokens / item: 1,000
Z-Image-Turbo
z-image-turbo
Z-Image is a powerful and highly efficient image generation model family with 6B parameters. Currently there are four variants: 🚀 Z-Image-Turbo – A distilled version of Z-Image that matches or exceeds leading competitors with only 8 NFEs (Number of Function Evaluations). It offers ⚡️sub-second inference latency⚡️ on enterprise-grade H800 GPUs and fits comfortably within 16G VRAM consumer devices. It excels in photorealistic image generation, bilingual text rendering (English & Chinese), and robust instruction adherence. 🎨 Z-Image – The foundation model behind Z-Image-Turbo. Z-Image focuses on high-quality generation, rich aesthetics, strong diversity, and controllability, well-suited for creative generation, fine-tuning, and downstream development. It supports a wide range of artistic styles, effective negative prompting, and high diversity across identities, poses, compositions, and layouts. 🧱 Z-Image-Omni-Base – The versatile foundation model capable of both generation and editing tasks. By releasing this checkpoint, we aim to unlock the full potential for community-driven fine-tuning and custom development, providing the most "raw" and diverse starting point for the open-source community. ✍️ Z-Image-Edit – A variant fine-tuned on Z-Image specifically for image editing tasks. It supports creative image-to-image generation with impressive instruction-following capabilities, allowing for precise edits based on natural language prompts.
$10.00/1M out tokTokens / item: 1,000

Speech

ModelInputOutputNotes
Chatterbox
chatterbox
Chatterbox Multilingual V3 is the latest general-purpose multilingual TTS model in the Chatterbox family. It keeps the same 0.5B model size while improving speaker similarity, reducing hallucinations, and producing more natural, conversational speech across languages. V3 is designed for broad language coverage like V2, but with stronger stability and more expressive generation. It is the recommended multilingual model for users who want one voice cloning model that works across many languages.
$2.00/1M out tokTokens / sec audio: 100
Fun-ASR-Nano
fun-asr-nano
LLM-Powered Speech Recognition — 31 Languages, Dialects & Accents End-to-end ASR trained on tens of millions of hours of data. Supports Chinese (+ dialects), English, Japanese, Korean, French, German, Spanish, and 24 more languages.
$0.05/1M in tokTokens / sec audio: 1,000
Kokoro-82M
kokoro-82m
Fast, lightweight text-to-speech model
$1.00/1M out tokTokens / sec audio: 100
Qwen3-ASR-1.7B
qwen3-asr-1-7b
The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. Experiments show that the 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs.
$0.05/1M in tokTokens / sec audio: 1,000
Qwen3-TTS
qwen3-tts
Qwen3-TTS covers 10 major languages (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian) as well as multiple dialectal voice profiles to meet global application needs. In addition, the models feature strong contextual understanding, enabling adaptive control of tone, speaking rate, and emotional expression based on instructions and text semantics, and they show markedly improved robustness to noisy input text.
$2.00/1M out tokTokens / sec audio: 100
Whisper-Large-V3
whisper-large-v3
Whisper is a Transformer based encoder-decoder model, also referred to as a sequence-to-sequence model. There are two flavours of Whisper model: English-only and multilingual. The English-only models were trained on the task of English speech recognition. The multilingual models were trained simultaneously on multilingual speech recognition and speech translation. For speech recognition, the model predicts transcriptions in the same language as the audio. For speech translation, the model predicts transcriptions to a different language to the audio.
$0.10/1M in tokTokens / sec audio: 1,000
Whisper-Large-V3-Turbo
whisper-large-v3-turbo
Whisper large-v3-turbo is a finetuned version of a pruned Whisper large-v3. In other words, it's the exact same model, except that the number of decoding layers have reduced from 32 to 4. As a result, the model is way faster, at the expense of a minor quality degradation.
$0.10/1M in tokTokens / sec audio: 1,000

GPU Compute

Pay per GPU-hour. Same rate for dedicated inference, workspace instances, and clusters. Billed per second of running time.

GPUVRAMOn-Demand
RTX Pro 600096 GB$1.89/GPU-hr

Storage

Persistent storage attached to your GPU instances and clusters. Cloud Drives are single-instance (ReadWriteOnce); Shared Filesystems mount across multiple instances (ReadWriteMany). Billed per second of attached time.

TypeSizeMonthly
Cloud Drive50 GB$5.00/mo
Cloud Drive100 GB$10.00/mo
Cloud Drive200 GB$20.00/mo
Cloud Drive300 GB$30.00/mo
Cloud Drive400 GB$40.00/mo
Cloud Drive500 GB$50.00/mo
Shared Filesystem50 GB$5.00/mo
Shared Filesystem100 GB$10.00/mo
Shared Filesystem200 GB$20.00/mo
Shared Filesystem300 GB$30.00/mo
Shared Filesystem400 GB$40.00/mo
Shared Filesystem500 GB$50.00/mo

$1 free credit on signup. No minimum commitment.