Qwen3.8 27B
Qwen3.8 27B is an open-weight dense vision-language model from Qwen, suited for coding, professional workflows, research, multimodal interaction and long-running agent tasks, with flexible thinking that can be turned up or down. Qwen Open-Weight Vision-Language 27B, with a 1M token context window.
/v1/ai/chatOverview
Qwen3.8 27B is an open-weight dense vision-language model from Qwen, suited for coding, professional workflows, research, multimodal interaction and long-running agent tasks, with flexible thinking that can be turned up or down. Qwen Open-Weight Vision-Language 27B, with a 1M token context window.
| Property | Value |
|---|---|
| Model ID | qwen/qwen3.8-27b |
| Context Window | 1,000,000 tokens |
| Max Output | 131,072 tokens |
| Input Price | $0.62 / 1M tokens |
| Output Price | $4.43 / 1M tokens |
Usage
const res = await fetch('https://api.yepapi.com/v1/ai/chat', {
method: 'POST',
headers: {
'x-api-key': 'YOUR_API_KEY',
'Content-Type': 'application/json',
},
body: JSON.stringify({
model: 'qwen/qwen3.8-27b',
messages: [{ role: 'user', content: 'Explain this architecture diagram and point out the single points of failure.' }],
}),
});
const { data } = await res.json();
console.log(data.message.content);curl -X POST https://api.yepapi.com/v1/ai/chat \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "qwen/qwen3.8-27b", "messages": [{"role": "user", "content": "Explain this architecture diagram and point out the single points of failure."}]}'Request Body
| Parameter | Type | Required | Description | Default |
|---|---|---|---|---|
model | string | Yes | Model ID (e.g. qwen/qwen3.8-27b) | — |
messages | Message[] | Yes | Array of { role, content } objects | — |
maxTokens | number | No | Maximum tokens in the response | Model default |
temperature | number | No | Sampling temperature (0.0–2.0) | 1.0 |
topP | number | No | Nucleus sampling threshold | 1.0 |
frequencyPenalty | number | No | Penalize repeated tokens | 0 |
presencePenalty | number | No | Penalize tokens already present | 0 |
stream | boolean | No | Enable SSE streaming | false |
All AI models use the /v1/ai/chat endpoint. Specify the model with the model field.
Response
{
"ok": true,
"data": {
"model": "qwen/qwen3.8-27b",
"message": {
"role": "assistant",
"content": "Traffic enters through one load balancer into two app servers that share one Postgres primary with no replica and one Redis node. The load balancer, the database and Redis are each single points of failure; the app tier is not."
},
"usage": {
"promptTokens": 16,
"completionTokens": 245,
"totalTokens": 261
}
}
}Streaming
Set "stream": true to receive Server-Sent Events. Each chunk contains a delta object:
data: {"delta":{"content":"Use"},"model":"qwen/qwen3.8-27b","index":0}
data: [DONE]We handle auth, billing, and response normalization — you just send messages.
Qwen3.8 Omni Flash
Qwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio and video understanding. Qwen Omni-Modal Audio & Video Agent, with a 1M token context window.
Qwen3.7 Max
Alibaba's largest Qwen model — strong multilingual, reasoning, and coding performance with a 1M token context window.