YepAPI
AI Models

Qwen-Audio 3.0 TTS Plus

The higher-quality Qwen speech tier for when Flash isn't clean enough.

POST/v1/media/queue
$0.0422/1K chars

Overview

Qwen-Audio-3.0-TTS Plus is Alibaba's higher-quality text-to-speech tier, generating spoken audio via the DashScope Speech Synthesizer API. Use it when Flash's output is not clean enough for the surface you are shipping.

PropertyValue
Model IDqwen/qwen-audio-3.0-tts-plus
Aliasqwen-tts-plus
Upstream Modelqwen/qwen-audio-3.0-tts-plus
CategoryText to Speech
LanguagesChinese and English
Outputmp3 (default) or pcm
Pricing$0.0422 per 1,000 characters

Usage

All media models use the async job queue. Submit a job, then poll for the result.

Step 1: Submit Job

const res = await fetch('https://api.yepapi.com/v1/media/queue', {
  method: 'POST',
  headers: {
    'x-api-key': 'YOUR_API_KEY',
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({
    model: 'qwen/qwen-audio-3.0-tts-plus',
    prompt: 'Your text to speak goes here.',
    options: { voice: 'longanlingxin' },
  }),
});
const { data } = await res.json();
// data.jobId — use this to poll for results
curl -X POST https://api.yepapi.com/v1/media/queue \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "qwen/qwen-audio-3.0-tts-plus", "prompt": "Your text to speak goes here."}'

Step 2: Poll for Result

const status = await fetch(`https://api.yepapi.com/v1/media/status/${data.jobId}`, {
  headers: { 'x-api-key': 'YOUR_API_KEY' },
});
const { data: job } = await status.json();
// job.status — "pending" | "processing" | "completed" | "failed"
// job.result.audio — { mimeType, base64 } when completed
curl https://api.yepapi.com/v1/media/status/JOB_ID \
  -H "x-api-key: YOUR_API_KEY"

Write the audio to a file:

import { writeFileSync } from 'node:fs';

writeFileSync('speech.mp3', Buffer.from(job.result.audio.base64, 'base64'));

Request Body

ParameterTypeRequiredDescriptionDefault
modelstringYesqwen/qwen-audio-3.0-tts-plus
promptstringYesThe text to speak (max 50,000 bytes)
options.voicestringNoVoice identifierlonganlingxin
options.outputFormatstringNomp3 (default) or pcmmp3
options.speednumberNoPlayback speed multiplier. Honoured only by models that support it1.0

Voices

longanlingxin is used when options.voice is omitted. 2 voices available:

  • longanlingxin
  • longanlufeng

Features

  • Higher fidelity than the Flash tier
  • Strong Chinese-language coverage
  • Two built-in voices
  • Served via Alibaba DashScope

Billing

Speech is billed on the length of the input text, in UTF-8 bytes — for English text one byte is one character. The cost is known before synthesis starts, so the balance check at submit time quotes the exact final charge. Jobs have a $0.01 minimum.

Under the Hood

Powered by OpenRouter's unified speech API.

On this page