GPT-6 Astra Pro
GPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with OpenAI's reasoning mode set to pro. OpenAI GPT-6 Astra Pro Reasoning Mode, with a 1M token context window.
/v1/ai/chatOverview
GPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with OpenAI's reasoning mode set to pro. OpenAI GPT-6 Astra Pro Reasoning Mode, with a 1M token context window.
| Property | Value |
|---|---|
| Model ID | openai/gpt-6-astra-pro |
| Context Window | 1,050,000 tokens |
| Max Output | 128,000 tokens |
| Input Price | $14.77 / 1M tokens |
| Output Price | $73.85 / 1M tokens |
Usage
const res = await fetch('https://api.yepapi.com/v1/ai/chat', {
method: 'POST',
headers: {
'x-api-key': 'YOUR_API_KEY',
'Content-Type': 'application/json',
},
body: JSON.stringify({
model: 'openai/gpt-6-astra-pro',
messages: [{ role: 'user', content: 'Find the bug in this concurrent cache implementation and explain why it only appears under load.' }],
}),
});
const { data } = await res.json();
console.log(data.message.content);curl -X POST https://api.yepapi.com/v1/ai/chat \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "openai/gpt-6-astra-pro", "messages": [{"role": "user", "content": "Find the bug in this concurrent cache implementation and explain why it only appears under load."}]}'Request Body
| Parameter | Type | Required | Description | Default |
|---|---|---|---|---|
model | string | Yes | Model ID (e.g. openai/gpt-6-astra-pro) | — |
messages | Message[] | Yes | Array of { role, content } objects | — |
maxTokens | number | No | Maximum tokens in the response | Model default |
temperature | number | No | Sampling temperature (0.0–2.0) | 1.0 |
topP | number | No | Nucleus sampling threshold | 1.0 |
frequencyPenalty | number | No | Penalize repeated tokens | 0 |
presencePenalty | number | No | Penalize tokens already present | 0 |
stream | boolean | No | Enable SSE streaming | false |
All AI models use the /v1/ai/chat endpoint. Specify the model with the model field.
Response
{
"ok": true,
"data": {
"model": "openai/gpt-6-astra-pro",
"message": {
"role": "assistant",
"content": "The get-or-compute path checks the map, misses, computes, then inserts \u2014 without holding the lock across the whole sequence. Under load two threads miss at the same time, both compute, and the second insert silently overwrites the first, so listeners registered on the first entry never fire. Compute under a per-key lock or use a compute-if-absent primitive."
},
"usage": {
"promptTokens": 16,
"completionTokens": 245,
"totalTokens": 261
}
}
}Streaming
Set "stream": true to receive Server-Sent Events. Each chunk contains a delta object:
data: {"delta":{"content":"Use"},"model":"openai/gpt-6-astra-pro","index":0}
data: [DONE]We handle auth, billing, and response normalization — you just send messages.
GPT-6 Luna
GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. OpenAI GPT-6 Fast & Low-Cost, with a 1M token context window.
GPT-6 Sol Pro
GPT-6 Sol Pro is the same underlying model as GPT-6 Sol, served with OpenAI's reasoning mode set to pro. OpenAI GPT-6 Sol Pro Reasoning Mode, with a 1M token context window.