Skip to main content

LLM API Model Catalogue

This catalogue lists all models currently available via the Connected AI Platform (CAIP) LLM API, grouped by model type and region. For the most up-to-date list, always refer to the API documentation or endpoint.

Chat Completion Models

Model NameProviderRegionUse CaseStatusEst. PriceTPM (Tokens/min)RPM (Requests/min)
claude-3-haikuAWS BedrockRoWFast, cost-effective; forwards to Claude Haiku 4.5Deprecated~$0.25/million in, $1.25/million outN/AN/A
claude-37-sonnetAWS BedrockRoWEnhanced Sonnet; forwards to Claude Sonnet 4.6DeprecatedN/A (Contact CAIP)N/AN/A
claude-4-sonnetAWS BedrockRoWLatest Claude; forwards to Claude Sonnet 4.6DeprecatedN/A (Contact CAIP)N/AN/A
claude-haiku-4.5AWS BedrockRoWFast, cost-efficientActiveN/A (Contact CAIP)N/AN/A
claude-sonnet-4.6AWS BedrockRoWBalanced; strong reasoning, content generationActiveN/A (Contact CAIP)N/AN/A
claude-opus-4.5AWS BedrockRoWHigh-end reasoning; complex tasks, planning, codingActiveN/A (Contact CAIP)N/AN/A
claude-opus-4.6AWS BedrockRoWFlagship reasoning modelActiveN/A (Contact CAIP)N/AN/A
nova-liteAWS BedrockRoWCost/speed optimized; basic chat and automationActiveN/A (Contact CAIP)N/AN/A
nova-microAWS BedrockRoWSmallest Nova; low-latency, low-costActiveN/A (Contact CAIP)N/AN/A
nova-proAWS BedrockRoWEnterprise Nova; advanced featuresActiveN/A (Contact CAIP)N/AN/A
llama-32-1bAWS BedrockRoWSmall Llama; lightweight tasks, experimentationUpcoming deprecationN/A (Contact CAIP)N/AN/A
llama-32-3bAWS BedrockRoWLarger Llama; multi-turn, more complex tasksUpcoming deprecationN/A (Contact CAIP)N/AN/A
deepseek-r1Azure OpenAIRoWAdvanced reasoning modelActiveN/A (Contact CAIP)N/AN/A
deepseek-v3Azure OpenAIRoWStrong general-purpose modelActiveN/A (Contact CAIP)N/AN/A
deepseek-v3.1AWS BedrockRoWHybrid thinking/non-thinking modelActiveN/A (Contact CAIP)N/AN/A
qwen-plusAWS BedrockRoWLarge-scale general-purposeUpcoming deprecation~$0.80/million tokens (in+out)N/AN/A
qwen-flashAWS BedrockRoWFast, cost-efficientActiveN/A (Contact CAIP)N/AN/A
qwen3-14bAWS BedrockRoWLightweight general-purposeUpcoming deprecation~$1.20/million tokens (in+out)N/AN/A
qwen3-30b-a3b-instruct-2507AWS BedrockRoWCode-focused lightweightActive~$0.80/million tokens (in+out)N/AN/A
gpt-4oAzure OpenAIRoWStrong reasoning; creative, coding, advanced chatActive~$5/million in, $15/million out5M30K
gpt-4o-miniAzure OpenAIRoWSmaller, faster GPT-4o; cost/latency sensitiveActiveN/A (Contact CAIP)10M100K
gpt-41Azure OpenAIRoWEnhanced GPT-4; improved accuracy and reasoningActiveN/A (Contact CAIP)2M2K
gpt-41-miniAzure OpenAIRoWSmaller GPT-4.1; optimized for speed and costActiveN/A (Contact CAIP)10M10K
gpt-41-nanoAzure OpenAIRoWSmallest GPT-4.1; lightweight applicationsActiveN/A (Contact CAIP)10M10M
gpt-5Azure OpenAIRoWFlagship; advanced reasoning, multimodal, complex tasksActive~$1.25/million in, $10/million out30M300K
gpt-5-miniAzure OpenAIRoWBalanced; chat, summarization, extractionActive~$0.25/million in, $2/million out10M10K
gpt-5-nanoAzure OpenAIRoWFastest/cheapest; classification, routing, short repliesActive~$0.05/million in, $0.4/million out10K10K
gpt-5-2Azure OpenAIRoWEnhanced reasoning capabilitiesActive~$1.75/million in, $14/million out10M100K
gpt-5.5Azure OpenAIRoWMost up-to-date model with latest featuresActiveN/A (Contact CAIP)N/AN/A
o3Azure OpenAIRoWReasoning-first; math/logic, planning, tool-useActive~$2/million in, $8/million out10M10K
qwen3.8-maxAlibaba CloudChinaFlagship Qwen 3.8 model for advanced chat and reasoningActiveN/A (Contact CAIP)N/AN/A
qwen3.7-plusAlibaba CloudChinaBalanced Qwen 3.7 model for complex tasksActiveN/A (Contact CAIP)N/AN/A
qwen3.7-maxAlibaba CloudChinaFlagship Qwen 3.7 model for advanced chat and reasoningActiveN/A (Contact CAIP)N/AN/A
qwen3.7-flashAlibaba CloudChinaFast, cost-efficient for complex tasksActiveN/A (Contact CAIP)N/AN/A
qwen3.6-plusAlibaba CloudChinaBalanced quality/latency for complex tasksActiveN/A (Contact CAIP)N/AN/A
qwen3.6-flashAlibaba CloudChinaFast, cost-efficient for complex tasksActiveN/A (Contact CAIP)N/AN/A
qwen3.6-27bAlibaba CloudChinaLarge-scale model for advanced tasksActiveN/A (Contact CAIP)N/AN/A
qwen3.5-plusAlibaba CloudChinaBalanced quality/latency for complex tasksActiveN/A (Contact CAIP)N/AN/A
qwen3.5-flashAlibaba CloudChinaFast, cost-efficient for complex tasksActiveN/A (Contact CAIP)N/AN/A
qwen3.5-27bAlibaba CloudChinaLarge-scale model for advanced tasksActiveN/A (Contact CAIP)N/AN/A
qwen3.5-omni-plusAlibaba CloudChinaMultimodal chat model for text/image/video/audioActiveN/A (Contact CAIP)N/AN/A
qwen-plusAlibaba CloudChinaBalanced quality/latency for general chatActive~$0.80/million tokens (in+out)N/AN/A
qwen-flashAlibaba CloudChinaFast, cost-efficient for general chatActiveN/A (Contact CAIP)N/AN/A
qwen-maxAlibaba CloudChinaFlagship model for complex chat tasksActiveN/A (Contact CAIP)N/AN/A
qwen-longAlibaba CloudChinaLong-context model for extended conversationsActiveN/A (Contact CAIP)N/AN/A
gui-plusAlibaba CloudChinaBalanced quality/latency for general chat with image inputActiveN/A (Contact CAIP)N/AN/A
tongyi-intent-detect-v3Alibaba CloudChinaIntent detection for understanding user queriesActiveN/A (Contact CAIP)N/AN/A
qwen3-vl-plusAlibaba CloudChinaMultimodal chat for image and video understandingActiveN/A (Contact CAIP)N/AN/A
deepseek-v4-proAlibaba CloudChinaAdvanced reasoning and coding capabilitiesActiveN/A (Contact CAIP)N/AN/A
deepseek-v4-pro-0813Alibaba CloudChinaAdvanced reasoning and coding capabilitiesActiveN/A (Contact CAIP)N/AN/A
deepseek-v4-flashAlibaba CloudChinaFast, cost-efficient for reasoning and codingActiveN/A (Contact CAIP)N/AN/A
deepseek-v4-flash-0731Alibaba CloudChinaGeneral-purpose model for reasoning and codingActiveN/A (Contact CAIP)N/AN/A
glm-5.2-fast-previewAlibaba CloudChinaGeneral-purpose model for chat and reasoning tasksActiveN/A (Contact CAIP)N/AN/A
glm-5.2Alibaba CloudChinaGeneral-purpose model for chat and reasoning tasksActiveN/A (Contact CAIP)N/AN/A
glm-5Alibaba CloudChinaGeneral-purpose model for chat and reasoning tasksActiveN/A (Contact CAIP)N/AN/A
MiniMax-M3Alibaba CloudChinaLarge multimodal chat model for broad assistant workloadsActiveN/A (Contact CAIP)N/AN/A
MiniMax-M2.5Alibaba CloudChinaLarge multimodal chat model for broad assistant useActive¥2.1/million in, ¥8.4/million outN/AN/A
kimi-k3Alibaba CloudChinaThird-party model served via Alibaba BailianActiveN/A (Contact CAIP)N/AN/A
kimi-k2.7-codeAlibaba CloudChinaLong-context model for code-heavy and document-centric chatActiveN/A (Contact CAIP)N/AN/A
kimi-k2.6Alibaba CloudChinaLong-context model for retrieval-heavy and document-centric chatActiveN/A (Contact CAIP)N/AN/A
kimi-k2.5Alibaba CloudChinaLong-context model for retrieval-heavy and document-centric chatActive¥4/million in, ¥21/million outN/AN/A

Realtime Models

The following models are available only in China through the CAIP Realtime WebSocket endpoint. They are not used with Chat Completions.

Model NameProviderRegionUse CaseStatusEst. PriceTPM (Tokens/min)RPM (Requests/min)
qwen3.5-omni-plus-realtimeAlibaba CloudChinaHigh-quality bidirectional text and audio conversationsActiveN/A (Contact CAIP)N/AN/A
qwen3.5-omni-flash-realtimeAlibaba CloudChinaLow-latency bidirectional text and audio conversationsActiveN/A (Contact CAIP)N/AN/A

Connect through wss://llm.api.caip.bmwchina.cloud/v1/realtime?model={model} using your CAIP API key. See the Realtime API documentation for the protocol and examples.

Deprecation Notice
  • Legacy Claude (claude-3-haiku, claude-37-sonnet, claude-4-sonnet): requests are internally forwarded to newer Claude models. Please migrate to claude-haiku-4.5, claude-sonnet-4.6, or claude-opus-4.6.
  • Llama 3.2 (llama-32-1b, llama-32-3b): will be phased out in the near future.
  • CN models in RoW (qwen-plus, qwen3-14b): will soon be replaced with newer alternatives.

Note:

  • TPM (Tokens per minute) is the maximum number of tokens the API can process per minute.
  • RPM (Requests per minute) is the maximum number of requests the API can accept per minute.

Pricing: CAIP pricing may differ. For details, contact the CAIP team.

Responses API Models

These models are only available via the /v1/responses endpoint and are not accessible through Chat Completions.

Model NameProviderRegionUse CaseEst. PriceTPM (Tokens/min)RPM (Requests/min)
gpt-5-codexAzure OpenAIRoWCoding; generation, refactor, debugging, tests~$1.25/million in, $10/million out10M10K
gpt-5-proAzure OpenAIRoWHighest quality; deep reasoning, agentic tasks~$15/million in, $120/million out1.6M16K
o3-proAzure OpenAIRoWBest reasoning; hardest problems, high-stakesN/A (Contact CAIP)10M1K

Embeddings Models

Model NameProviderRegionUse CaseEst. PriceTPM (Tokens/min)RPM (Requests/min)
text-embedding-3-smallAzure OpenAIRoWEfficient embeddings; search, RAG, similarity~$0.02/1K tokens5M30K
text-embedding-3-largeAzure OpenAIRoWHigher accuracy; advanced search, KM~$0.13/1K tokens10M60K
titan-text-embeddings-v2AWS BedrockRoWAmazon embeddings; search, recommendations~$0.13/1K tokensN/AN/A
qwen3.7-text-embeddingAlibaba CloudChinaLatest embeddings; search, RAG, similarity (CN)N/A (Contact CAIP)N/AN/A
text-embedding-v4Alibaba CloudChinaLatest embeddings; search, RAG, similarity (CN)N/A (Contact CAIP)N/AN/A

Rerank Models

Model NameProviderRegionUse CaseEst. PriceRPM (Requests/min)
cohere-rerank-v3-5AWS BedrockRoWSemantic relevance reranking; RAG, search pipelines~$2/1,000 queriesN/A

Image Models

Model NameProviderRegionUse CaseEst. PriceRPM (Requests/min)
gpt-image-1Azure OpenAIRoWText+image-to-image; image generation & editing(1M tokens) Input text $5 (cached $1.25); input image $10 (cached $2.50); output image $4045
gpt-image-1-miniAzure OpenAIRoWText+image-to-image; fast, cost-efficient(1M tokens) Input text $2 (cached $0.20); input image $2.50 (cached $0.25); output image $845
gpt-image-1.5Azure OpenAIRoWImage generation, improved qualityN/A (Contact CAIP)N/A
gpt-image-2Azure OpenAIRoWImage generation, latest modelN/A (Contact CAIP)N/A
titan-image-generator-v2AWS BedrockRoWText-to-image; creative content, marketingN/A (Contact CAIP)N/A
qwen-image-3.0Alibaba Cloud (DashScope)ChinaHigh-fidelity text-to-image model with strong realism and naturalnessN/A (Contact CAIP)N/A
qwen-image-3.0-proAlibaba Cloud (DashScope)ChinaHigh-fidelity text-to-image model with strong realism and naturalnessN/A (Contact CAIP)N/A
qwen-image-2.0Alibaba Cloud (DashScope)ChinaFaster generation and editing with flexible 2K outputN/A (Contact CAIP)N/A
qwen-image-2.0-proAlibaba Cloud (DashScope)ChinaStronger text rendering and prompt followingN/A (Contact CAIP)N/A
qwen-imageAlibaba Cloud (DashScope)ChinaGeneral-purpose image model with strong text renderingN/A (Contact CAIP)N/A
qwen-image-plusAlibaba Cloud (DashScope)ChinaArtistic text-to-image with stronger style diversityN/A (Contact CAIP)N/A
qwen-image-maxAlibaba Cloud (DashScope)ChinaHigher-fidelity image generation with stronger realismN/A (Contact CAIP)N/A
wan2.7-imageAlibaba Cloud (DashScope)ChinaFaster Wan model for image generation and editingN/A (Contact CAIP)N/A
wan2.7-image-proAlibaba Cloud (DashScope)ChinaHigh-resolution Wan model for advanced image tasksN/A (Contact CAIP)N/A

Video Models

Model NameProviderRegionUse CaseEst. PriceRPM (Requests/min)
sora-2Azure OpenAIRoWText-to-video; 8s or 16s, up to 1080pN/A (Contact CAIP)N/A
Sora-2 Retirement

OpenAI has announced the retirement of the sora-2 model beginning of May 2026. We are actively looking for alternatives for video generation.

Audio Models

Speech-to-Text / Transcription

Model NameProviderRegionUse CaseStatusEst. Price
whisperAzure OpenAIRoWGeneral-purpose speech-to-textDeprecatedN/A (Contact CAIP)
gpt-4o-mini-transcribeAzure OpenAIRoWFast speech-to-textActiveN/A (Contact CAIP)
gpt-4o-transcribeAzure OpenAIRoWPremium speech-to-textActiveN/A (Contact CAIP)

Text-to-Speech

Model NameProviderRegionUse CaseStatusEst. Price
gpt-4o-mini-ttsAzure OpenAIRoWExpressive text-to-speechDeprecatedN/A (Contact CAIP)

Self-Hosted Models (Early Access)

Self-hosted models run directly inside the CAIP Kubernetes cluster on dedicated GPU infrastructure and never leave our VPC. They are available at zero additional cost and use the same standard LLMAPI endpoints. See the Self-Hosted Models documentation for full details.

Model NameModalityUnderlying ModelDescriptionStatusEst. Price
qwen3-chatChat CompletionsQwen/Qwen3-4BLightweight, fast chat modelActive (Beta)Free
qwen3.6-35b-a3bChat / Coding (agentic)Qwen/Qwen3.6-35B-A3B (FP8)Coding & agentic model — tool calling, vision, 128K context (-think variant for reasoning)Active (Beta)Free
qwen3-embeddingEmbeddingsQwen/Qwen3-Embedding-0.6BText embeddings for search, RAG, similarityActive (Beta)Free
qwen3-asrSpeech-to-TextQwen/Qwen3-ASR-1.7BAudio transcriptionActive (Beta)Free
flux-imageImage GenerationFLUX.2-klein-4BText-to-image generationActive (Beta)Free
wan-videoVideo GenerationWan2.2-TI2V-5BText/image-to-video generationActive (Beta)Free
Self-Hosted Limitations

Self-hosted models are in early access with limited throughput. On the lightweight qwen3-chat model, some features (tool calling, vision input, certain request parameters) are not yet supported. The qwen3.6-35b-a3b coding model does support tool calling and vision — see Coding with the LLM API. For full details, see Known Limitations.

Model Selection Notes

  • Not all models are available in every region.
  • Some models (e.g., certain Qwen/DeepSeek variants) are only available in China.
  • In RoW, DeepSeek and Qwen models are served via AWS Bedrock (not Alibaba Cloud).
  • Self-hosted models run on CAIP-managed GPU infrastructure within our VPC.
  • For streaming support, check the model documentation.
  • Pricing is managed by CAIP and may differ from public cloud pricing. For details, contact the CAIP team.

For more details on model capabilities and usage, refer to the main LLM API Documentation.