AI models & pricing
All models on the Apertus platform with current token prices – usage-based billing via an OpenAI-compatible API.
Last updated: September 29, 2026
- Open-weight models
- 31
- Chat, embedding & reranking
- New in the last month
- 3
- incl. DeepSeek V4.1 Flash, Qwen 3.8 Flash Next
- All-rounder from
- €0.50
- DeepSeek V4.1 Flash · per 1M input tokens
- Context up to
- 1M
- in 7 models
Our recommendations
Not sure which model fits? You can't go wrong with these.
DeepSeek V4.1 Flash
DeepSeek
Our everyday all-rounder: capable, fast, multimodal and affordable.
Multimodal MoE model with 552B parameters and up to 1M tokens of context – natively processes images and text.
- Context
- 1M
- Input
- €0.50
- Output
- €1.50
per 1M tokens
Kimi K3
Moonshot AI
Frontier level for demanding coding, agents and long contexts.
Frontier-grade coding and agentic performance, a full 1M-token context and native multimodal input (text, image, video).
- Context
- 1M
- Input
- €3.00
- Output
- €15.00
per 1M tokens
GLM 5.3 Flash
Z.ai
Multimodal and powerful – at the price of a small model.
The first natively multimodal model in the GLM-5 series. Outperforms GLM 5.2 across benchmarks and real-world workloads at a fraction of the price.
- Context
- 1M
- Input
- €0.20
- Output
- €0.50
per 1M tokens
SwissAI Apertus 1.5 70B
Swiss AI
Fully open, multilingual model from Switzerland – for sovereign AI in Europe.
Switzerland's first large-scale, open and multilingual language model – ideal for sovereign AI applications.
- Context
- 65K
- Input
- €0.70
- Output
- €2.80
per 1M tokens
Chat & reasoning
Language models for chat, coding, agents, vision and document analysis.
| Model | Context | Input per 1M tokens | Cached input per 1M tokens | Output per 1M tokens |
|---|---|---|---|---|
DeepSeek V4.1 Flash
Recommended: all-rounder
New
Vision deepseek-v4.1-flash · DeepSeekMultimodal MoE model with 552B parameters and up to 1M tokens of context – natively processes images and text. | 1M | €0.50 | €0.13 | €1.50 |
![]() Qwen 3.8 Flash Next
New
Vision qwen3.8-flash-next · Qwen (Alibaba)Ultra-efficient multimodal model with elite coding from just 6B active parameters. | 250K | €0.20 | €0.05 | €0.50 |
GLM 5.3 Flash
Favorite: best value
New
Vision glm-5.3-flash · Z.aiThe first natively multimodal model in the GLM-5 series. Outperforms GLM 5.2 across benchmarks and real-world workloads at a fraction of the price. | 1M | €0.20 | €0.05 | €0.50 |
Kimi K3
Favorite: top performance
Vision kimi-k3 · Moonshot AIFrontier-grade coding and agentic performance, a full 1M-token context and native multimodal input (text, image, video). | 1M | €3.00 | €0.75 | €15.00 |
GLM 5.2 glm-5.2 · Z.aiZ.ai's flagship for long-horizon tasks – a substantial leap over GLM 5.1, delivered on a solid 1M-token context. | 1M | €1.50 | €0.72 | €4.50 |
MiniMax M3
Vision minimax-m3 · MiniMaxOpen-weight model built for coding, long-context agent work and multimodal input. | 1M | €0.40 | €0.15 | €2.00 |
Gemma 4 31B
Vision gemma4-31b · Google DeepMindOpen multimodal model from Google DeepMind – takes text and image input and generates text. | 250K | €0.20 | – | €0.40 |
![]() Qwen3.6 27B qwen3.6-27b · Qwen (Alibaba)Compact Qwen 3.6 generation model for chat, text and coding tasks. | – | €0.40 | – | €2.70 |
DeepSeek V4 Pro 0813 deepseek-v4-pro · DeepSeek1.6T MoE with 49B active parameters and Hybrid Attention – a strong fit for complex coding, deep reasoning and long-running agentic workflows. | 1M | €2.00 | €0.50 | €4.00 |
Kimi K2.6
Vision kimi-k2.6 · Moonshot AINative multimodal, with Agent Swarm scaling to 300 specialized sub-agents and 4,000 coordinated steps per autonomous run. | 256K | €1.00 | €0.25 | €4.00 |
DeepSeek V4 Flash 0731 deepseek-v4-flash · DeepSeekTrained from scratch on the same data as V4 Pro. Well-suited for high-volume workloads where cost and speed matter – chat, classification, summarization. | 1M | €0.25 | €0.08 | €0.30 |
![]() Qwen3 VL 235B
Vision Qwen3-VL-235B-A22B-Instruct · Qwen (Alibaba)The most powerful vision-language model in the Qwen series – for OCR, document parsing, spatial grounding and code generation from images. | 256K | €2.00 | – | €2.00 |
SwissAI Apertus 1.5 70B
European open source swissai-apertus-70b · Swiss AISwitzerland's first large-scale, open and multilingual language model – ideal for sovereign AI applications. | 65K | €0.70 | – | €2.80 |
GPT-OSS-120B gpt-oss-120b · OpenAIOpenAI's most powerful open-weight model. | 131K | €0.48 | – | €0.48 |
![]() Qwen3 Coder 30B Qwen3-Coder-30B-A3B-Instruct · Qwen (Alibaba)Versatile code model for programming and agentic coding. | 250K | €0.264 | – | €0.264 |
![]() Qwen2.5 VL 72B Instruct
Vision Qwen2.5-VL-72B-Instruct · Qwen (Alibaba)Large Qwen vision-language model for image and document understanding. | 32K | €1.092 | – | €1.092 |
GPT-OSS-20B gpt-oss-20b · OpenAIOpenAI's small but efficient open-weight model. | 131K | €0.18 | – | €0.18 |
![]() Mistral 7B Instruct Mistral-7B-Instruct-v0.3 · Mistral AISmall and powerful Mistral model. | 127K | €0.12 | – | €0.12 |
![]() Llama 3.3 70B Instruct Meta-Llama-3_3-70B-Instruct · MetaPowerful reasoning, versatile, production-ready. | 131K | €0.804 | – | €0.804 |
![]() Qwen3 32B Qwen3-32B · Qwen (Alibaba)Advanced reasoning, multilingual, balanced capacity. | 32K | €0.276 | – | €0.276 |
![]() Mistral Small 3.2 24B Instruct Mistral-Small-3.2-24B-Instruct-2506 · Mistral AIFast inference, lightweight, instruction-optimized. | 125K | €0.336 | – | €0.336 |
![]() Mixtral 8x7B Instruct Mixtral-8x7B-Instruct-v0.1 · Mistral AIMixture of experts, high quality, efficient routing. | 32K | €0.756 | – | €0.756 |
![]() Mistral Nemo Instruct Mistral-Nemo-Instruct-2407 · Mistral AIBalanced performance, multilingual, instruction-tuned. | 118K | €0.156 | – | €0.156 |
![]() Llama 3.1 8B Instruct Llama-3.1-8B-Instruct · MetaLightweight and fast for general-purpose tasks. | 131K | €0.12 | – | €0.12 |
Embedding
Vector embeddings for semantic search, RAG and classification.
| Model | Context | Price per 1M tokens |
|---|---|---|
![]() Qwen3 Embedding 8B Qwen3-Embedding-8B · Qwen (Alibaba)Embedding model with state-of-the-art results across a wide range of retrieval and downstream benchmarks. | 8K | €0.084 |
BGE Multilingual Gemma 2 bge-multilingual-gemma2 · BAAIMultilingual BGE embedding model built on Gemma 2. | 8K | €0.012 |
BGE-M3 BGE-M3 · BAAIMultilingual embedding model (100+ languages) for dense, sparse and multi-vector retrieval. | 8K | €0.012 |
Reranking
Relevance ranking of search results for more precise RAG answers – currently free of charge.
| Model | Context | Price per 1M tokens |
|---|---|---|
![]() Qwen3 Reranker 0.6B qwen3-reranker-0.6b · Qwen (Alibaba)Multilingual reranker supporting over 100 languages. | 32K | Free |
BGE Reranker v2 M3 bge-reranker-v2-m3 · BAAILightweight multilingual reranker with fast inference. | – | Free |
Additional models on request
We enable these models for your account on request. Closed-source models in the EU region are also available on request.
| Model | Context | Input per 1M tokens | Cached input per 1M tokens | Output per 1M tokens |
|---|---|---|---|---|
![]() Devstral 2 123B devstral-2-123b-instruct-2512 · Mistral AIPurpose-built for agentic software engineering – 72.2% on SWE-Bench Verified. | 200K | €2.00 | – | €2.00 |
![]() Qwen3 235B A22B Instruct 2507 qwen3-235b-a22b-instruct-2507 · Qwen (Alibaba)Dedicated instruct model with improved instruction following, multilingual coverage and long-context understanding. | 250K | €2.00 | – | €2.00 |
No models found. Feel free to ask us about additional models.
Pricing notes · Last updated: September 29, 2026
- All prices in euros per 1 million tokens, net, excluding VAT.
- Cached input: reduced price for reused prompt content (prompt caching), where supported by the model.
- Billed monthly based on actual usage. The prices shown in the dashboard are binding.
- All models are GDPR-compliant unless marked otherwise.
- Closed-source models in the EU region, dedicated instances, on-premise installations and volume pricing on request.
Get started
Sign up on the platform and use all models right away – or talk to us about volume pricing, dedicated instances and models on request.
Learn more about our AI models

