YepAPI
AI Models

MiniMax Speech 2.8 HD

MiniMax's premium speech tier — the highest-fidelity option on the platform.

POST/v1/media/queue
$0.2110/1K chars

Overview

MiniMax Speech 2.8 HD is the premium tier of MiniMax's speech family and the highest-fidelity text-to-speech model on YepAPI. It accepts arbitrary MiniMax voice IDs rather than a fixed catalogue, so any voice you have provisioned with MiniMax can be named directly.

PropertyValue
Model IDminimax/speech-2.8-hd
Aliasminimax-speech
Upstream Modelminimax/speech-2.8-hd
CategoryText to Speech
Languagesmultilingual
Outputmp3 (default) or pcm
Pricing$0.2110 per 1,000 characters

Usage

All media models use the async job queue. Submit a job, then poll for the result.

Step 1: Submit Job

const res = await fetch('https://api.yepapi.com/v1/media/queue', {
  method: 'POST',
  headers: {
    'x-api-key': 'YOUR_API_KEY',
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({
    model: 'minimax/speech-2.8-hd',
    prompt: 'Your text to speak goes here.',
    options: { voice: 'English_expressive_narrator' },
  }),
});
const { data } = await res.json();
// data.jobId — use this to poll for results
curl -X POST https://api.yepapi.com/v1/media/queue \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "minimax/speech-2.8-hd", "prompt": "Your text to speak goes here."}'

Step 2: Poll for Result

const status = await fetch(`https://api.yepapi.com/v1/media/status/${data.jobId}`, {
  headers: { 'x-api-key': 'YOUR_API_KEY' },
});
const { data: job } = await status.json();
// job.status — "pending" | "processing" | "completed" | "failed"
// job.result.audio — { mimeType, base64 } when completed
curl https://api.yepapi.com/v1/media/status/JOB_ID \
  -H "x-api-key: YOUR_API_KEY"

Write the audio to a file:

import { writeFileSync } from 'node:fs';

writeFileSync('speech.mp3', Buffer.from(job.result.audio.base64, 'base64'));

Request Body

ParameterTypeRequiredDescriptionDefault
modelstringYesminimax/speech-2.8-hd
promptstringYesThe text to speak (max 50,000 bytes)
options.voicestringNoVoice identifierEnglish_expressive_narrator
options.outputFormatstringNomp3 (default) or pcmmp3
options.speednumberNoPlayback speed multiplier. Honoured only by models that support it1.0

Voices

This model accepts any MiniMax voice ID rather than a fixed catalogue. An explicit voice is required upstream; YepAPI sends English_expressive_narrator when you omit it.

Common IDs:

  • English_expressive_narrator
  • male-qn-qingse
  • female-shaonv

Features

  • Highest-fidelity speech available on the platform
  • Accepts arbitrary MiniMax voice IDs
  • Multilingual output
  • An explicit voice is required — there is no provider default

Billing

Speech is billed on the length of the input text, in UTF-8 bytes — for English text one byte is one character. The cost is known before synthesis starts, so the balance check at submit time quotes the exact final charge. Jobs have a $0.01 minimum.

Under the Hood

Powered by OpenRouter's unified speech API.

On this page