Streaming

Stream Vietnamese speech while it is being generated. Voixa voices the text clause by clause and sends each clause the moment it is ready, so a listener hears the first words in about two seconds and the rest follows without gaps. Built for voice bots, call centres and live apps.

How it works

Streaming takes two requests. The first goes through the API like every other call — your API key, your rate limit, your tokens — and returns a one-use ticket. The second trades the ticket for the audio stream.

  1. POST /stream with the voice and text. The API checks your tokens and answers with url and ticket (valid for two minutes, usable once).
  2. POST {url} with { "ticket": "…" }. The response body is the audio, delivered as it is generated. Tokens are charged when the stream opens.

The first clause is kept short so it plays almost immediately; each next clause may be longer, and is voiced while the previous one plays.

Measured

First audio, both requests includedabout 2 s from Vietnam: 0.5 s for POST /stream, then 1.4–1.6 s to the first bytes
Longest pause while playing0.2 s
Cold start (first call after a quiet period)5–9 s

Request

POST /stream

FieldType
voiceIdstringA Vietnamese voice: a built-in one or your clone
textstringUp to 5 000 characters, cleaned and counted like speak
sampleRateinteger48000, 24000 (default), 16000 or 8000
formatstringpcm (default) or wav

The answer:

{
  "url": "https://….lambda-url.us-east-1.on.aws/",
  "ticket": "4f1c…",
  "expiresIn": 120,
  "tokens": 105,
  "sampleRate": 16000,
  "format": "pcm",
  "encoding": "pcm_s16le",
  "channels": 1
}

Then:

curl -N -X POST "$URL" -H 'content-type: application/json' \
  -d "{\"ticket\":\"$TICKET\"}" --output reply.pcm

The body is mono 16-bit little-endian PCM at the sample rate you asked for. With format: "wav" the same data comes behind a WAV header whose length is left open, which most players accept for streaming. The response headers repeat the format: x-voixa-encoding, x-voixa-sample-rate, x-voixa-channels, x-voixa-tokens.

With the SDK

SDK

import { Voixa } from '@vovix/voixa'

const vx = new Voixa({ apiKey: process.env.VOIXA_API_KEY })

const s = await vx.stream({ voiceId: 'vi-truc-ly', text: 'Xin chào, Voixa có thể giúp gì cho anh chị?', sampleRate: 16000 })
for await (const chunk of s.body) {
  player.write(chunk) // raw PCM, 16 kHz mono
}

Price

Streaming costs 1.5 tokens per character (see Tokens); it runs on a faster tier than speak. A ticket you never use costs nothing. If the engine fails, the tokens are refunded.

Limits and errors

  • Vietnamese voices only for now. Other languages answer 400; use speak.
  • A cloned voice needs a one-off preparation for streaming. Right after you create it, POST /stream may answer 409 VOICE_NOT_READY — try again in about two minutes.
  • A used ticket answers 404 TICKET_USED, an old one 410 TICKET_EXPIRED. Ask POST /stream for a new one.
  • Not enough tokens answers 402 INSUFFICIENT_TOKENS.
  • Up to 10 streams run at the same time across the service; more wait for a free slot.

Was this page helpful?