Speak
speak turns text into a clip: a WAV in your library with a signed URL. Short texts come back ready in the same call; longer ones come back processing and finish a few seconds later.
Clips created with an API key are kept for 24 hours and then deleted, audio included; expiresAt on the clip says when. Download the WAV to your own storage, or share the clip to keep it. Clips made in Studio do not expire.
The clip model
Properties
- Name
clipId- Type
- string
- Description
Unique id of the clip.
- Name
voiceId- Type
- string
- Description
The voice it was read with;
voiceNameis its display name.
- Name
language- Type
- string
- Description
The voice's language.
- Name
text- Type
- string
- Description
The text as it was read, after cleaning (see Text cleaning).
- Name
chars- Type
- integer
- Description
Characters counted against your allowance.
- Name
project- Type
- string
- Description
The content group you passed, if any.
- Name
source- Type
- string
- Description
apifor calls with an API key,studiofor clips made in Studio.
- Name
status- Type
- string
- Description
processing,readyorfailed(withfailReason).
- Name
url- Type
- string
- Description
Signed WAV URL, valid for one hour. Present when
ready.
- Name
durationSec- Type
- number
- Description
Length of the audio, once ready.
sampleRateis its sample rate.
- Name
parts- Type
- integer
- Description
For long texts: how many parts were produced in parallel.
partsDonereports progress whileprocessing.
- Name
shareUrl- Type
- string
- Description
Public listening page, after sharing.
Create a clip
Reads text with the voice and stores the clip in your library. The API waits up to about 20 seconds for the audio, which covers most texts; otherwise the clip comes back processing and you poll GET /clips/{clipId} every second or two. Texts longer than about 250 characters are split into parts that run in parallel and are joined into one WAV — to you it is still one clip.
Required attributes
- Name
voiceId- Type
- string
- Description
A voice from
GET /voices.
- Name
text- Type
- string
- Description
Up to 5 000 characters (1 200 for Chinese, Japanese and Korean). Longer? Use a podcast.
Optional attributes
- Name
project- Type
- string
- Description
A content group such as
episode-12. Letters, digits, spaces and. _ -, up to 60.
- Name
wait- Type
- boolean
- Description
true(default) waits up to ~20 s for the audio;falseanswers at once withprocessing.
Request
curl -X POST https://api.voixa.vovix.io/v1/speak \
-H "x-api-key: $VOIXA_API_KEY" \
-H "content-type: application/json" \
-d '{
"voiceId": "vi-truc-ly",
"text": "Xin chào, cuộc gọi của bạn đang được kết nối.",
"project": "ivr"
}'
Response
{
"clip": {
"clipId": "e0f43e3b-e343-430d-be01-6b6d4309f8d8",
"voiceId": "vi-truc-ly",
"voiceName": "Trúc Ly",
"language": "vi",
"text": "Xin chào, cuộc gọi của bạn đang được kết nối.",
"chars": 45,
"project": "ivr",
"source": "api",
"status": "ready",
"durationSec": 3.04,
"sampleRate": 48000,
"url": "https://…"
}
}
Text cleaning
Before counting and reading, Voixa removes what would sound wrong read aloud: markdown symbols (*, #, >, list bullets, code fences), URLs, emoji, bracketed notes, and repeated punctuation (!!! becomes !). Line breaks and runs of spaces collapse to one space. Numbers, dates, times, percentages and money stay — they are read naturally (8h30, 1.250.000 đồng, 68,5%).