LLM API Model Catalogue
This catalogue lists all models currently available via the Connected AI Platform (CAIP) LLM API, grouped by model type and region. For the most up-to-date list, always refer to the API documentation or endpoint.
Chat Completion Models
| Model Name | Provider | Region | Use Case | Status | Est. Price | TPM (Tokens/min) | RPM (Requests/min) |
|---|---|---|---|---|---|---|---|
| claude-3-haiku | AWS Bedrock | RoW | Fast, cost-effective; forwards to Claude Haiku 4.5 | Deprecated | ~$0.25/million in, $1.25/million out | N/A | N/A |
| claude-37-sonnet | AWS Bedrock | RoW | Enhanced Sonnet; forwards to Claude Sonnet 4.6 | Deprecated | N/A (Contact CAIP) | N/A | N/A |
| claude-4-sonnet | AWS Bedrock | RoW | Latest Claude; forwards to Claude Sonnet 4.6 | Deprecated | N/A (Contact CAIP) | N/A | N/A |
| claude-haiku-4.5 | AWS Bedrock | RoW | Fast, cost-efficient | Active | N/A (Contact CAIP) | N/A | N/A |
| claude-sonnet-4.6 | AWS Bedrock | RoW | Balanced; strong reasoning, content generation | Active | N/A (Contact CAIP) | N/A | N/A |
| claude-opus-4.5 | AWS Bedrock | RoW | High-end reasoning; complex tasks, planning, coding | Active | N/A (Contact CAIP) | N/A | N/A |
| claude-opus-4.6 | AWS Bedrock | RoW | Flagship reasoning model | Active | N/A (Contact CAIP) | N/A | N/A |
| nova-lite | AWS Bedrock | RoW | Cost/speed optimized; basic chat and automation | Active | N/A (Contact CAIP) | N/A | N/A |
| nova-micro | AWS Bedrock | RoW | Smallest Nova; low-latency, low-cost | Active | N/A (Contact CAIP) | N/A | N/A |
| nova-pro | AWS Bedrock | RoW | Enterprise Nova; advanced features | Active | N/A (Contact CAIP) | N/A | N/A |
| llama-32-1b | AWS Bedrock | RoW | Small Llama; lightweight tasks, experimentation | Upcoming deprecation | N/A (Contact CAIP) | N/A | N/A |
| llama-32-3b | AWS Bedrock | RoW | Larger Llama; multi-turn, more complex tasks | Upcoming deprecation | N/A (Contact CAIP) | N/A | N/A |
| deepseek-r1 | Azure OpenAI | RoW | Advanced reasoning model | Active | N/A (Contact CAIP) | N/A | N/A |
| deepseek-v3 | Azure OpenAI | RoW | Strong general-purpose model | Active | N/A (Contact CAIP) | N/A | N/A |
| deepseek-v3.1 | AWS Bedrock | RoW | Hybrid thinking/non-thinking model | Active | N/A (Contact CAIP) | N/A | N/A |
| qwen-plus | AWS Bedrock | RoW | Large-scale general-purpose | Upcoming deprecation | ~$0.80/million tokens (in+out) | N/A | N/A |
| qwen-flash | AWS Bedrock | RoW | Fast, cost-efficient | Active | N/A (Contact CAIP) | N/A | N/A |
| qwen3-14b | AWS Bedrock | RoW | Lightweight general-purpose | Upcoming deprecation | ~$1.20/million tokens (in+out) | N/A | N/A |
| qwen3-30b-a3b-instruct-2507 | AWS Bedrock | RoW | Code-focused lightweight | Active | ~$0.80/million tokens (in+out) | N/A | N/A |
| gpt-4o | Azure OpenAI | RoW | Strong reasoning; creative, coding, advanced chat | Active | ~$5/million in, $15/million out | 5M | 30K |
| gpt-4o-mini | Azure OpenAI | RoW | Smaller, faster GPT-4o; cost/latency sensitive | Active | N/A (Contact CAIP) | 10M | 100K |
| gpt-41 | Azure OpenAI | RoW | Enhanced GPT-4; improved accuracy and reasoning | Active | N/A (Contact CAIP) | 2M | 2K |
| gpt-41-mini | Azure OpenAI | RoW | Smaller GPT-4.1; optimized for speed and cost | Active | N/A (Contact CAIP) | 10M | 10K |
| gpt-41-nano | Azure OpenAI | RoW | Smallest GPT-4.1; lightweight applications | Active | N/A (Contact CAIP) | 10M | 10M |
| gpt-5 | Azure OpenAI | RoW | Flagship; advanced reasoning, multimodal, complex tasks | Active | ~$1.25/million in, $10/million out | 30M | 300K |
| gpt-5-mini | Azure OpenAI | RoW | Balanced; chat, summarization, extraction | Active | ~$0.25/million in, $2/million out | 10M | 10K |
| gpt-5-nano | Azure OpenAI | RoW | Fastest/cheapest; classification, routing, short replies | Active | ~$0.05/million in, $0.4/million out | 10K | 10K |
| gpt-5-2 | Azure OpenAI | RoW | Enhanced reasoning capabilities | Active | ~$1.75/million in, $14/million out | 10M | 100K |
| gpt-5.5 | Azure OpenAI | RoW | Most up-to-date model with latest features | Active | N/A (Contact CAIP) | N/A | N/A |
| o3 | Azure OpenAI | RoW | Reasoning-first; math/logic, planning, tool-use | Active | ~$2/million in, $8/million out | 10M | 10K |
| qwen3.8-max | Alibaba Cloud | China | Flagship Qwen 3.8 model for advanced chat and reasoning | Active | N/A (Contact CAIP) | N/A | N/A |
| qwen3.7-plus | Alibaba Cloud | China | Balanced Qwen 3.7 model for complex tasks | Active | N/A (Contact CAIP) | N/A | N/A |
| qwen3.7-max | Alibaba Cloud | China | Flagship Qwen 3.7 model for advanced chat and reasoning | Active | N/A (Contact CAIP) | N/A | N/A |
| qwen3.7-flash | Alibaba Cloud | China | Fast, cost-efficient for complex tasks | Active | N/A (Contact CAIP) | N/A | N/A |
| qwen3.6-plus | Alibaba Cloud | China | Balanced quality/latency for complex tasks | Active | N/A (Contact CAIP) | N/A | N/A |
| qwen3.6-flash | Alibaba Cloud | China | Fast, cost-efficient for complex tasks | Active | N/A (Contact CAIP) | N/A | N/A |
| qwen3.6-27b | Alibaba Cloud | China | Large-scale model for advanced tasks | Active | N/A (Contact CAIP) | N/A | N/A |
| qwen3.5-plus | Alibaba Cloud | China | Balanced quality/latency for complex tasks | Active | N/A (Contact CAIP) | N/A | N/A |
| qwen3.5-flash | Alibaba Cloud | China | Fast, cost-efficient for complex tasks | Active | N/A (Contact CAIP) | N/A | N/A |
| qwen3.5-27b | Alibaba Cloud | China | Large-scale model for advanced tasks | Active | N/A (Contact CAIP) | N/A | N/A |
| qwen3.5-omni-plus | Alibaba Cloud | China | Multimodal chat model for text/image/video/audio | Active | N/A (Contact CAIP) | N/A | N/A |
| qwen-plus | Alibaba Cloud | China | Balanced quality/latency for general chat | Active | ~$0.80/million tokens (in+out) | N/A | N/A |
| qwen-flash | Alibaba Cloud | China | Fast, cost-efficient for general chat | Active | N/A (Contact CAIP) | N/A | N/A |
| qwen-max | Alibaba Cloud | China | Flagship model for complex chat tasks | Active | N/A (Contact CAIP) | N/A | N/A |
| qwen-long | Alibaba Cloud | China | Long-context model for extended conversations | Active | N/A (Contact CAIP) | N/A | N/A |
| gui-plus | Alibaba Cloud | China | Balanced quality/latency for general chat with image input | Active | N/A (Contact CAIP) | N/A | N/A |
| tongyi-intent-detect-v3 | Alibaba Cloud | China | Intent detection for understanding user queries | Active | N/A (Contact CAIP) | N/A | N/A |
| qwen3-vl-plus | Alibaba Cloud | China | Multimodal chat for image and video understanding | Active | N/A (Contact CAIP) | N/A | N/A |
| deepseek-v4-pro | Alibaba Cloud | China | Advanced reasoning and coding capabilities | Active | N/A (Contact CAIP) | N/A | N/A |
| deepseek-v4-pro-0813 | Alibaba Cloud | China | Advanced reasoning and coding capabilities | Active | N/A (Contact CAIP) | N/A | N/A |
| deepseek-v4-flash | Alibaba Cloud | China | Fast, cost-efficient for reasoning and coding | Active | N/A (Contact CAIP) | N/A | N/A |
| deepseek-v4-flash-0731 | Alibaba Cloud | China | General-purpose model for reasoning and coding | Active | N/A (Contact CAIP) | N/A | N/A |
| glm-5.2-fast-preview | Alibaba Cloud | China | General-purpose model for chat and reasoning tasks | Active | N/A (Contact CAIP) | N/A | N/A |
| glm-5.2 | Alibaba Cloud | China | General-purpose model for chat and reasoning tasks | Active | N/A (Contact CAIP) | N/A | N/A |
| glm-5 | Alibaba Cloud | China | General-purpose model for chat and reasoning tasks | Active | N/A (Contact CAIP) | N/A | N/A |
| MiniMax-M3 | Alibaba Cloud | China | Large multimodal chat model for broad assistant workloads | Active | N/A (Contact CAIP) | N/A | N/A |
| MiniMax-M2.5 | Alibaba Cloud | China | Large multimodal chat model for broad assistant use | Active | ¥2.1/million in, ¥8.4/million out | N/A | N/A |
| kimi-k3 | Alibaba Cloud | China | Third-party model served via Alibaba Bailian | Active | N/A (Contact CAIP) | N/A | N/A |
| kimi-k2.7-code | Alibaba Cloud | China | Long-context model for code-heavy and document-centric chat | Active | N/A (Contact CAIP) | N/A | N/A |
| kimi-k2.6 | Alibaba Cloud | China | Long-context model for retrieval-heavy and document-centric chat | Active | N/A (Contact CAIP) | N/A | N/A |
| kimi-k2.5 | Alibaba Cloud | China | Long-context model for retrieval-heavy and document-centric chat | Active | ¥4/million in, ¥21/million out | N/A | N/A |
Realtime Models
The following models are available only in China through the CAIP Realtime WebSocket endpoint. They are not used with Chat Completions.
| Model Name | Provider | Region | Use Case | Status | Est. Price | TPM (Tokens/min) | RPM (Requests/min) |
|---|---|---|---|---|---|---|---|
| qwen3.5-omni-plus-realtime | Alibaba Cloud | China | High-quality bidirectional text and audio conversations | Active | N/A (Contact CAIP) | N/A | N/A |
| qwen3.5-omni-flash-realtime | Alibaba Cloud | China | Low-latency bidirectional text and audio conversations | Active | N/A (Contact CAIP) | N/A | N/A |
Connect through wss://llm.api.caip.bmwchina.cloud/v1/realtime?model={model} using your CAIP API key. See the Realtime API documentation for the protocol and examples.
- Legacy Claude (
claude-3-haiku,claude-37-sonnet,claude-4-sonnet): requests are internally forwarded to newer Claude models. Please migrate toclaude-haiku-4.5,claude-sonnet-4.6, orclaude-opus-4.6. - Llama 3.2 (
llama-32-1b,llama-32-3b): will be phased out in the near future. - CN models in RoW (
qwen-plus,qwen3-14b): will soon be replaced with newer alternatives.
Note:
- TPM (Tokens per minute) is the maximum number of tokens the API can process per minute.
- RPM (Requests per minute) is the maximum number of requests the API can accept per minute.
Pricing: CAIP pricing may differ. For details, contact the CAIP team.
Responses API Models
These models are only available via the /v1/responses endpoint and are not accessible through Chat Completions.
| Model Name | Provider | Region | Use Case | Est. Price | TPM (Tokens/min) | RPM (Requests/min) |
|---|---|---|---|---|---|---|
| gpt-5-codex | Azure OpenAI | RoW | Coding; generation, refactor, debugging, tests | ~$1.25/million in, $10/million out | 10M | 10K |
| gpt-5-pro | Azure OpenAI | RoW | Highest quality; deep reasoning, agentic tasks | ~$15/million in, $120/million out | 1.6M | 16K |
| o3-pro | Azure OpenAI | RoW | Best reasoning; hardest problems, high-stakes | N/A (Contact CAIP) | 10M | 1K |
Embeddings Models
| Model Name | Provider | Region | Use Case | Est. Price | TPM (Tokens/min) | RPM (Requests/min) |
|---|---|---|---|---|---|---|
| text-embedding-3-small | Azure OpenAI | RoW | Efficient embeddings; search, RAG, similarity | ~$0.02/1K tokens | 5M | 30K |
| text-embedding-3-large | Azure OpenAI | RoW | Higher accuracy; advanced search, KM | ~$0.13/1K tokens | 10M | 60K |
| titan-text-embeddings-v2 | AWS Bedrock | RoW | Amazon embeddings; search, recommendations | ~$0.13/1K tokens | N/A | N/A |
| qwen3.7-text-embedding | Alibaba Cloud | China | Latest embeddings; search, RAG, similarity (CN) | N/A (Contact CAIP) | N/A | N/A |
| text-embedding-v4 | Alibaba Cloud | China | Latest embeddings; search, RAG, similarity (CN) | N/A (Contact CAIP) | N/A | N/A |
Rerank Models
| Model Name | Provider | Region | Use Case | Est. Price | RPM (Requests/min) |
|---|---|---|---|---|---|
| cohere-rerank-v3-5 | AWS Bedrock | RoW | Semantic relevance reranking; RAG, search pipelines | ~$2/1,000 queries | N/A |
Image Models
| Model Name | Provider | Region | Use Case | Est. Price | RPM (Requests/min) |
|---|---|---|---|---|---|
| gpt-image-1 | Azure OpenAI | RoW | Text+image-to-image; image generation & editing | (1M tokens) Input text $5 (cached $1.25); input image $10 (cached $2.50); output image $40 | 45 |
| gpt-image-1-mini | Azure OpenAI | RoW | Text+image-to-image; fast, cost-efficient | (1M tokens) Input text $2 (cached $0.20); input image $2.50 (cached $0.25); output image $8 | 45 |
| gpt-image-1.5 | Azure OpenAI | RoW | Image generation, improved quality | N/A (Contact CAIP) | N/A |
| gpt-image-2 | Azure OpenAI | RoW | Image generation, latest model | N/A (Contact CAIP) | N/A |
| titan-image-generator-v2 | AWS Bedrock | RoW | Text-to-image; creative content, marketing | N/A (Contact CAIP) | N/A |
| qwen-image-3.0 | Alibaba Cloud (DashScope) | China | High-fidelity text-to-image model with strong realism and naturalness | N/A (Contact CAIP) | N/A |
| qwen-image-3.0-pro | Alibaba Cloud (DashScope) | China | High-fidelity text-to-image model with strong realism and naturalness | N/A (Contact CAIP) | N/A |
| qwen-image-2.0 | Alibaba Cloud (DashScope) | China | Faster generation and editing with flexible 2K output | N/A (Contact CAIP) | N/A |
| qwen-image-2.0-pro | Alibaba Cloud (DashScope) | China | Stronger text rendering and prompt following | N/A (Contact CAIP) | N/A |
| qwen-image | Alibaba Cloud (DashScope) | China | General-purpose image model with strong text rendering | N/A (Contact CAIP) | N/A |
| qwen-image-plus | Alibaba Cloud (DashScope) | China | Artistic text-to-image with stronger style diversity | N/A (Contact CAIP) | N/A |
| qwen-image-max | Alibaba Cloud (DashScope) | China | Higher-fidelity image generation with stronger realism | N/A (Contact CAIP) | N/A |
| wan2.7-image | Alibaba Cloud (DashScope) | China | Faster Wan model for image generation and editing | N/A (Contact CAIP) | N/A |
| wan2.7-image-pro | Alibaba Cloud (DashScope) | China | High-resolution Wan model for advanced image tasks | N/A (Contact CAIP) | N/A |
Video Models
| Model Name | Provider | Region | Use Case | Est. Price | RPM (Requests/min) |
|---|---|---|---|---|---|
| sora-2 | Azure OpenAI | RoW | Text-to-video; 8s or 16s, up to 1080p | N/A (Contact CAIP) | N/A |
OpenAI has announced the retirement of the sora-2 model beginning of May 2026. We are actively looking for alternatives for video generation.
Audio Models
Speech-to-Text / Transcription
| Model Name | Provider | Region | Use Case | Status | Est. Price |
|---|---|---|---|---|---|
| whisper | Azure OpenAI | RoW | General-purpose speech-to-text | Deprecated | N/A (Contact CAIP) |
| gpt-4o-mini-transcribe | Azure OpenAI | RoW | Fast speech-to-text | Active | N/A (Contact CAIP) |
| gpt-4o-transcribe | Azure OpenAI | RoW | Premium speech-to-text | Active | N/A (Contact CAIP) |
Text-to-Speech
| Model Name | Provider | Region | Use Case | Status | Est. Price |
|---|---|---|---|---|---|
| gpt-4o-mini-tts | Azure OpenAI | RoW | Expressive text-to-speech | Deprecated | N/A (Contact CAIP) |
Self-Hosted Models (Early Access)
Self-hosted models run directly inside the CAIP Kubernetes cluster on dedicated GPU infrastructure and never leave our VPC. They are available at zero additional cost and use the same standard LLMAPI endpoints. See the Self-Hosted Models documentation for full details.
| Model Name | Modality | Underlying Model | Description | Status | Est. Price |
|---|---|---|---|---|---|
qwen3-chat | Chat Completions | Qwen/Qwen3-4B | Lightweight, fast chat model | Active (Beta) | Free |
qwen3.6-35b-a3b | Chat / Coding (agentic) | Qwen/Qwen3.6-35B-A3B (FP8) | Coding & agentic model — tool calling, vision, 128K context (-think variant for reasoning) | Active (Beta) | Free |
qwen3-embedding | Embeddings | Qwen/Qwen3-Embedding-0.6B | Text embeddings for search, RAG, similarity | Active (Beta) | Free |
qwen3-asr | Speech-to-Text | Qwen/Qwen3-ASR-1.7B | Audio transcription | Active (Beta) | Free |
flux-image | Image Generation | FLUX.2-klein-4B | Text-to-image generation | Active (Beta) | Free |
wan-video | Video Generation | Wan2.2-TI2V-5B | Text/image-to-video generation | Active (Beta) | Free |
Self-hosted models are in early access with limited throughput. On the lightweight qwen3-chat model, some features (tool calling, vision input, certain request parameters) are not yet supported. The qwen3.6-35b-a3b coding model does support tool calling and vision — see Coding with the LLM API. For full details, see Known Limitations.
Model Selection Notes
- Not all models are available in every region.
- Some models (e.g., certain Qwen/DeepSeek variants) are only available in China.
- In RoW, DeepSeek and Qwen models are served via AWS Bedrock (not Alibaba Cloud).
- Self-hosted models run on CAIP-managed GPU infrastructure within our VPC.
- For streaming support, check the model documentation.
- Pricing is managed by CAIP and may differ from public cloud pricing. For details, contact the CAIP team.
For more details on model capabilities and usage, refer to the main LLM API Documentation.