Python SDK
voixa wraps every endpoint, waits for clips, designs and transcripts for you, and has no runtime dependencies: only the standard library. Python 3.9 or newer.
Install
pip
pip install voixa
Quick start
from voixa import Voixa
vx = Voixa() # reads VOIXA_API_KEY; create a key in Studio -> API keys
# or: Voixa(api_key="...", base_url="https://api.voixa.vovix.io/v1", timeout=30)
voices = vx.voices(language="en")["voices"]
clip = vx.say(
voice_id=voices[0]["voiceId"],
text="The lighthouse keeper switched the lamp on at dusk.",
project="episode-12", # optional: group clips by content
)
print(clip["url"], clip["durationSec"]) # signed WAV URL, valid for 1 hour
say() returns when the clip is ready. Use speak(..., wait=False) + wait_for()
to manage the wait yourself: the API answers in under a second and you poll.
Every method returns what the API returns: plain dicts with the API's
camelCase keys (clipId, durationSec, ...). The package ships TypedDict
definitions for all of them (voixa.Clip, voixa.Voice, voixa.Transcript, ...)
so editors and type checkers know the fields. Request arguments are snake_case
keywords.
timeout (default 30 s) bounds each HTTP request. The wait_* helpers and the
methods that wait (say, make_podcast, create_voice, save_design,
transcribe_file, ...) take their own poll_interval and timeout, in seconds.
Streaming (Vietnamese)
Hear the first words in about two seconds: audio arrives clause by clause while the
rest is voiced. stream() returns a SpeechStream, an iterator of bytes chunks
(16-bit little-endian mono PCM, or WAV with format="wav"). The connection closes
when the iteration ends; use with to close it early.
# Play it as it arrives
for chunk in vx.stream(voice_id="vi-truc-ly", text="Xin chào, Voixa có thể giúp gì?", sample_rate=16000):
player.write(chunk)
# Or write it to a file
with vx.stream(voice_id="vi-truc-ly", text="Xin chào!", format="wav") as s:
print(s.sample_rate, s.tokens)
s.save("hello.wav") # or: open("hello.wav", "wb").write(s.read())
sample_rate: 48000, 24000 (default), 16000 or 8000; format: pcm or wav.
Costs 1.5 tokens per character, charged when the stream opens and refunded if
generation fails.
Tokens
Every request is priced in tokens: 1 token = 1 Vietnamese character. Free tokens refill daily; top-ups carry over.
from voixa import estimate_tokens
account = vx.account()["account"]
account["tokens"] # {available, freeToday, dailyFree, balance, usedToday, refillsAt, rates, unlimited}
rates = account["tokens"]["rates"]
estimate_tokens(rates, language="vi", chars=1200) # -> 1200
estimate_tokens(rates, design_options=3) # a voice design with three options
estimate_tokens(rates, audio_seconds=600) # ten minutes of transcription
page = vx.usage(limit=20) # every spend, refund and top-up
A request you cannot afford raises VoixaError with status 402 and code
INSUFFICIENT_TOKENS (plus needed, available, refills_at). Failed clips
and voices are refunded.
Voices and samples
languages = vx.languages()["languages"] # codes, labels, limits, clone-recording guides
samples = vx.samples(language="fr") # built-in voices with public sample URLs
voice = vx.get_voice("en-alba")["voice"]
Voice design: a voice from a description
No recording needed. Describe the voice, listen to up to three options, keep the
one you like. A saved design is an ordinary voice of your account: use it with
say, make_podcast and stream. Characters for dubbing, narrators, ads: one
voice per role.
design = vx.design_voice(
description="An old man in his seventies, hoarse and deep, speaking slowly",
language="vi", # any language `languages()` marks `designable`
count=3, # 1-3 options, `rates["design"]` tokens each
)["design"]
done = vx.wait_for_design(design["designId"]) # about a minute per option
for c in done["candidates"]:
print(c["index"], c["status"], c.get("url")) # listen, then pick
voice = vx.save_design(done["designId"], candidate=1, name="Grandpa Ba", gender="male")
vx.say(voice_id=voice["voiceId"], text="Ngày ấy, ông vẫn nhớ con đường làng.")
# Or in one call, keeping the first option:
narrator = vx.create_voice_from_description(
description="A calm female narrator, warm and clear", language="en", name="Narrator",
)
vx.list_designs(limit=10) # your designs, newest first
vx.get_design(design["designId"]) # one design with its options
vx.delete_design(design["designId"]) # saved voices are kept
Options are drafts kept for 23 hours (expiresAt). Saving one costs
rates["clone"] tokens.
Voice cloning and sharing
voice = vx.create_voice(
name="Studio narrator",
language="en",
gender="female",
audio="reference.wav", # 3-10 s of clean speech, WAV or MP3, max 10 MB
# audio accepts bytes, a path, or a binary file object: open("ref.mp3", "rb")
# format="mp3", # wav by default
# ref_text="...", # what the speaker says; helps Chinese, Japanese, Korean
)
vx.say(voice_id=voice["voiceId"], text="Hello from my own voice.")
vx.update_voice(voice["voiceId"], shared=True) # let every Voixa account use it
vx.update_voice(voice["voiceId"], name="Narrator", description=None) # None clears the description
vx.delete_voice(voice["voiceId"])
create_voice uploads the recording, registers the voice and waits for its
sample (check voice["status"]).
Voice cloning is available for Vietnamese, English, French, German, Italian,
Spanish and Portuguese. Chinese, Japanese and Korean use built-in voices only;
languages() reports cloneable per language and create_voice rejects those
languages with a 400.
Podcasts
Long scripts, up to 100 000 characters, become one episode: Voixa splits the text into sentences, produces the parts in parallel and joins them into a single WAV. Blank lines become short pauses.
episode = vx.make_podcast(voice_id="vi-truc-ly", title="Episode 12", text=script, project="my-show")
episode["url"] # one WAV, signed for an hour
episode["durationSec"] # e.g. 4800 for an 80-minute episode
# or without waiting:
clip = vx.create_podcast(voice_id="en-alba", text=script)["clip"]
clip["parts"], clip.get("partsDone") # progress while status is "processing"
Speech to text
Audio (or video: only the audio track is read) becomes a transcript with timed segments, SRT and VTT. Up to 5 GB and eight hours per file.
t = vx.transcribe_file(
audio="interview.mp3", # bytes, a path or a binary file object; files are streamed from disk
title="Interview with the lighthouse keeper",
# language="vi", # optional: Voixa detects it well on its own
# task="translate", # recognise any language, return English
# words=True, # per-word timestamps
# prompt="Voixa, Vovix", # spelling hints for names and jargon
)
print(t["language"], t["durationSec"], t["text"])
print(t["segments"][0]) # {"id", "start", "end", "text"}
print(t["srtUrl"]) # signed SRT/VTT/TXT/JSON URLs, valid for 1 hour
The format is taken from format=, else from the file name (file_name=, or the
name of the path or file you pass), else MP3.
Transcription runs on a machine that starts on demand, so it is never
synchronous: transcribe() returns status: "processing" as soon as the upload
finishes and you poll, while transcribe_file() does the polling for you (two to
three minutes before recognition starts on a cold queue, then roughly one minute
of work per five minutes of audio).
transcript = vx.transcribe(audio="talk.m4a")["transcript"]
done = vx.wait_for_transcript(transcript["transcriptId"])
# progress while processing: doneSec / durationSec, jobStatus
items = vx.list_transcripts(project="episode-12")["items"] # no text in lists
full = vx.get_transcript(items[0]["transcriptId"]) # text + segments + URLs
vx.delete_transcripts([transcript["transcriptId"]]) # also deletes the upload
Languages you may pin, the accepted formats and the limits come from
stt_languages(). Transcription is charged per second of audio once it finishes.
Library
projects = vx.projects()["projects"] # groups with clip counts
page = vx.list_clips(project="episode-12") # newest first
from_api = vx.list_clips(source="api") # clips made with an API key (vs. "studio")
podcasts = vx.list_clips(kind="podcast", limit=50)
more = vx.list_clips(project="episode-12", cursor=page.get("cursor"))
vx.delete_clips([c["clipId"] for c in page["items"]])
share = vx.share_clip(clip["clipId"]) # {"shareUrl": "https://voixa.vovix.io/s/...", "shared": True}
vx.unshare_clip(clip["clipId"]) # the link stops working
vx.get_clip(clip["clipId"])
vx.wait_for(clip["clipId"], timeout=300)
Errors and limits
Every API failure raises VoixaError with:
status: HTTP status codemessage: human-readable, safe to show (alsostr(err))code:NOT_APPROVED,INSUFFICIENT_TOKENS,CLONE_LIMIT,VOICE_NOT_READY,INVALID,AUDIO_TOO_LARGE,NO_AUDIO(QUOTAandAUDIO_QUOTAare legacy); the SDK itself setsTIMEOUT,SYNTH_FAILED,DESIGN_FAILED,TRANSCRIBE_FAILEDretry_after: seconds, from theRetry-Afterheaderneeded,available,refills_at(insufficient tokens);resets_at,used,limit(legacy quotas)body: the full JSON error body
from voixa import VoixaError
try:
vx.say(voice_id="vi-truc-ly", text=long_text)
except VoixaError as e:
if e.code == "INSUFFICIENT_TOKENS":
print(f"Need {e.needed}, have {e.available}; free tokens refill at {e.refills_at}")
elif e.code == "TIMEOUT":
... # still processing: poll later with vx.wait_for(...)
else:
raise
Requests the SDK can reject without calling the API (empty or too long text,
a design description under three characters, audio over 5 GB) raise the same
VoixaError with status 400. Network failures are not wrapped: they raise
the standard library's OSError subclasses (urllib.error.URLError, socket timeouts).
- One
speakcall takes up to 5 000 characters (fewer for Chinese, Japanese and Korean; seelanguages()). Line breaks and repeated spaces collapse to one space before counting, on your side and on the server. - One podcast takes up to 100 000 characters.
- One transcription takes up to 5 GB and eight hours of audio.
- API keys are rate-limited to 1 request/second and 1 000 calls/day; new keys activate after two to three minutes.
- Speech, streaming, transcription, clones and designs are paid in tokens (see
Tokens above); read the wallet with
account().
Method reference
| Group | Methods |
|---|---|
| Account | account(), usage(limit, cursor), languages(), samples(language) |
| Voices | voices(language), get_voice(id), create_voice(...), update_voice(id, ...), delete_voice(id), wait_for_voice(id) |
| Voice design | design_voice(...), get_design(id), list_designs(limit, cursor), wait_for_design(id), save_design(id, ...), delete_design(id), create_voice_from_description(...) |
| Speech | speak(...), say(...), stream(...) |
| Podcasts | create_podcast(...), make_podcast(...) |
| Library | get_clip(id), list_clips(...), projects(), delete_clips(ids), share_clip(id), unshare_clip(id), wait_for(id) |
| Speech to text | stt_languages(), transcribe(...), transcribe_file(...), get_transcript(id), list_transcripts(...), delete_transcripts(ids), wait_for_transcript(id) |
| Helpers | estimate_tokens(rates, ...), constants MAX_SPEAK_CHARS, MAX_PODCAST_CHARS, MAX_AUDIO_BYTES, MAX_AUDIO_MINUTES, DEFAULT_BASE_URL |
A signed-in user of your own web app can call the API with their ID token
instead of a key: Voixa(id_token=...).