YepAPI
AI Models

GPT-6 Astra Pro

GPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with OpenAI's reasoning mode set to pro. OpenAI GPT-6 Astra Pro Reasoning Mode, with a 1M token context window.

POST/v1/ai/chat
$0.01/call

Overview

GPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with OpenAI's reasoning mode set to pro. OpenAI GPT-6 Astra Pro Reasoning Mode, with a 1M token context window.

PropertyValue
Model IDopenai/gpt-6-astra-pro
Context Window1,050,000 tokens
Max Output128,000 tokens
Input Price$14.77 / 1M tokens
Output Price$73.85 / 1M tokens

Usage

const res = await fetch('https://api.yepapi.com/v1/ai/chat', {
  method: 'POST',
  headers: {
    'x-api-key': 'YOUR_API_KEY',
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({
    model: 'openai/gpt-6-astra-pro',
    messages: [{ role: 'user', content: 'Find the bug in this concurrent cache implementation and explain why it only appears under load.' }],
  }),
});
const { data } = await res.json();
console.log(data.message.content);
curl -X POST https://api.yepapi.com/v1/ai/chat \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "openai/gpt-6-astra-pro", "messages": [{"role": "user", "content": "Find the bug in this concurrent cache implementation and explain why it only appears under load."}]}'

Request Body

ParameterTypeRequiredDescriptionDefault
modelstringYesModel ID (e.g. openai/gpt-6-astra-pro)—
messagesMessage[]YesArray of { role, content } objects—
maxTokensnumberNoMaximum tokens in the responseModel default
temperaturenumberNoSampling temperature (0.0–2.0)1.0
topPnumberNoNucleus sampling threshold1.0
frequencyPenaltynumberNoPenalize repeated tokens0
presencePenaltynumberNoPenalize tokens already present0
streambooleanNoEnable SSE streamingfalse
Info

All AI models use the /v1/ai/chat endpoint. Specify the model with the model field.

Response

{
  "ok": true,
  "data": {
    "model": "openai/gpt-6-astra-pro",
    "message": {
      "role": "assistant",
      "content": "The get-or-compute path checks the map, misses, computes, then inserts \u2014 without holding the lock across the whole sequence. Under load two threads miss at the same time, both compute, and the second insert silently overwrites the first, so listeners registered on the first entry never fire. Compute under a per-key lock or use a compute-if-absent primitive."
    },
    "usage": {
      "promptTokens": 16,
      "completionTokens": 245,
      "totalTokens": 261
    }
  }
}

Streaming

Set "stream": true to receive Server-Sent Events. Each chunk contains a delta object:

data: {"delta":{"content":"Use"},"model":"openai/gpt-6-astra-pro","index":0}
data: [DONE]
Under the Hood

We handle auth, billing, and response normalization — you just send messages.

On this page