Streaming
Stream Vietnamese speech while it is being generated. Voixa voices the text clause by clause and sends each clause the moment it is ready, so a listener hears the first words in about two seconds and the rest follows without gaps. Built for voice bots, call centres and live apps.
How it works
Streaming takes two requests. The first goes through the API like every other call — your API key, your rate limit, your tokens — and returns a one-use ticket. The second trades the ticket for the audio stream.
POST /streamwith the voice and text. The API checks your tokens and answers withurlandticket(valid for two minutes, usable once).POST {url}with{ "ticket": "…" }. The response body is the audio, delivered as it is generated. Tokens are charged when the stream opens.
The first clause is kept short so it plays almost immediately; each next clause may be longer, and is voiced while the previous one plays.
Measured
| First audio, both requests included | about 2 s from Vietnam: 0.5 s for POST /stream, then 1.4–1.6 s to the first bytes |
| Longest pause while playing | 0.2 s |
| Cold start (first call after a quiet period) | 5–9 s |
Request
POST /stream
| Field | Type | |
|---|---|---|
voiceId | string | A Vietnamese voice: a built-in one or your clone |
text | string | Up to 5 000 characters, cleaned and counted like speak |
sampleRate | integer | 48000, 24000 (default), 16000 or 8000 |
format | string | pcm (default) or wav |
The answer:
{
"url": "https://….lambda-url.us-east-1.on.aws/",
"ticket": "4f1c…",
"expiresIn": 120,
"tokens": 105,
"sampleRate": 16000,
"format": "pcm",
"encoding": "pcm_s16le",
"channels": 1
}
Then:
curl -N -X POST "$URL" -H 'content-type: application/json' \
-d "{\"ticket\":\"$TICKET\"}" --output reply.pcm
The body is mono 16-bit little-endian PCM at the sample rate you asked for. With format: "wav" the same data comes behind a WAV header whose length is left open, which most players accept for streaming. The response headers repeat the format: x-voixa-encoding, x-voixa-sample-rate, x-voixa-channels, x-voixa-tokens.
With the SDK
SDK
import { Voixa } from '@vovix/voixa'
const vx = new Voixa({ apiKey: process.env.VOIXA_API_KEY })
const s = await vx.stream({ voiceId: 'vi-truc-ly', text: 'Xin chào, Voixa có thể giúp gì cho anh chị?', sampleRate: 16000 })
for await (const chunk of s.body) {
player.write(chunk) // raw PCM, 16 kHz mono
}
Price
Streaming costs 1.5 tokens per character (see Tokens); it runs on a faster tier than speak. A ticket you never use costs nothing. If the engine fails, the tokens are refunded.
Limits and errors
- Vietnamese voices only for now. Other languages answer
400; usespeak. - A cloned voice needs a one-off preparation for streaming. Right after you create it,
POST /streammay answer409 VOICE_NOT_READY— try again in about two minutes. - A used ticket answers
404 TICKET_USED, an old one410 TICKET_EXPIRED. AskPOST /streamfor a new one. - Not enough tokens answers
402 INSUFFICIENT_TOKENS. - Up to 10 streams run at the same time across the service; more wait for a free slot.