Python SDK

voixa wraps every endpoint, waits for clips, designs and transcripts for you, and has no runtime dependencies: only the standard library. Python 3.9 or newer.

Install

pip

pip install voixa

Quick start

from voixa import Voixa

vx = Voixa()  # reads VOIXA_API_KEY; create a key in Studio -> API keys
# or: Voixa(api_key="...", base_url="https://api.voixa.vovix.io/v1", timeout=30)

voices = vx.voices(language="en")["voices"]
clip = vx.say(
    voice_id=voices[0]["voiceId"],
    text="The lighthouse keeper switched the lamp on at dusk.",
    project="episode-12",  # optional: group clips by content
)
print(clip["url"], clip["durationSec"])  # signed WAV URL, valid for 1 hour

say() returns when the clip is ready. Use speak(..., wait=False) + wait_for() to manage the wait yourself: the API answers in under a second and you poll.

Every method returns what the API returns: plain dicts with the API's camelCase keys (clipId, durationSec, ...). The package ships TypedDict definitions for all of them (voixa.Clip, voixa.Voice, voixa.Transcript, ...) so editors and type checkers know the fields. Request arguments are snake_case keywords.

timeout (default 30 s) bounds each HTTP request. The wait_* helpers and the methods that wait (say, make_podcast, create_voice, save_design, transcribe_file, ...) take their own poll_interval and timeout, in seconds.

Streaming (Vietnamese)

Hear the first words in about two seconds: audio arrives clause by clause while the rest is voiced. stream() returns a SpeechStream, an iterator of bytes chunks (16-bit little-endian mono PCM, or WAV with format="wav"). The connection closes when the iteration ends; use with to close it early.

# Play it as it arrives
for chunk in vx.stream(voice_id="vi-truc-ly", text="Xin chào, Voixa có thể giúp gì?", sample_rate=16000):
    player.write(chunk)

# Or write it to a file
with vx.stream(voice_id="vi-truc-ly", text="Xin chào!", format="wav") as s:
    print(s.sample_rate, s.tokens)
    s.save("hello.wav")  # or: open("hello.wav", "wb").write(s.read())

sample_rate: 48000, 24000 (default), 16000 or 8000; format: pcm or wav. Costs 1.5 tokens per character, charged when the stream opens and refunded if generation fails.

Tokens

Every request is priced in tokens: 1 token = 1 Vietnamese character. Free tokens refill daily; top-ups carry over.

from voixa import estimate_tokens

account = vx.account()["account"]
account["tokens"]  # {available, freeToday, dailyFree, balance, usedToday, refillsAt, rates, unlimited}
rates = account["tokens"]["rates"]
estimate_tokens(rates, language="vi", chars=1200)   # -> 1200
estimate_tokens(rates, design_options=3)             # a voice design with three options
estimate_tokens(rates, audio_seconds=600)            # ten minutes of transcription
page = vx.usage(limit=20)                            # every spend, refund and top-up

A request you cannot afford raises VoixaError with status 402 and code INSUFFICIENT_TOKENS (plus needed, available, refills_at). Failed clips and voices are refunded.

Voices and samples

languages = vx.languages()["languages"]   # codes, labels, limits, clone-recording guides
samples = vx.samples(language="fr")        # built-in voices with public sample URLs
voice = vx.get_voice("en-alba")["voice"]

Voice design: a voice from a description

No recording needed. Describe the voice, listen to up to three options, keep the one you like. A saved design is an ordinary voice of your account: use it with say, make_podcast and stream. Characters for dubbing, narrators, ads: one voice per role.

design = vx.design_voice(
    description="An old man in his seventies, hoarse and deep, speaking slowly",
    language="vi",  # any language `languages()` marks `designable`
    count=3,        # 1-3 options, `rates["design"]` tokens each
)["design"]
done = vx.wait_for_design(design["designId"])  # about a minute per option
for c in done["candidates"]:
    print(c["index"], c["status"], c.get("url"))  # listen, then pick

voice = vx.save_design(done["designId"], candidate=1, name="Grandpa Ba", gender="male")
vx.say(voice_id=voice["voiceId"], text="Ngày ấy, ông vẫn nhớ con đường làng.")

# Or in one call, keeping the first option:
narrator = vx.create_voice_from_description(
    description="A calm female narrator, warm and clear", language="en", name="Narrator",
)

vx.list_designs(limit=10)          # your designs, newest first
vx.get_design(design["designId"])  # one design with its options
vx.delete_design(design["designId"])  # saved voices are kept

Options are drafts kept for 23 hours (expiresAt). Saving one costs rates["clone"] tokens.

Voice cloning and sharing

voice = vx.create_voice(
    name="Studio narrator",
    language="en",
    gender="female",
    audio="reference.wav",  # 3-10 s of clean speech, WAV or MP3, max 10 MB
    # audio accepts bytes, a path, or a binary file object: open("ref.mp3", "rb")
    # format="mp3",          # wav by default
    # ref_text="...",        # what the speaker says; helps Chinese, Japanese, Korean
)
vx.say(voice_id=voice["voiceId"], text="Hello from my own voice.")

vx.update_voice(voice["voiceId"], shared=True)       # let every Voixa account use it
vx.update_voice(voice["voiceId"], name="Narrator", description=None)  # None clears the description
vx.delete_voice(voice["voiceId"])

create_voice uploads the recording, registers the voice and waits for its sample (check voice["status"]).

Voice cloning is available for Vietnamese, English, French, German, Italian, Spanish and Portuguese. Chinese, Japanese and Korean use built-in voices only; languages() reports cloneable per language and create_voice rejects those languages with a 400.

Podcasts

Long scripts, up to 100 000 characters, become one episode: Voixa splits the text into sentences, produces the parts in parallel and joins them into a single WAV. Blank lines become short pauses.

episode = vx.make_podcast(voice_id="vi-truc-ly", title="Episode 12", text=script, project="my-show")
episode["url"]          # one WAV, signed for an hour
episode["durationSec"]  # e.g. 4800 for an 80-minute episode

# or without waiting:
clip = vx.create_podcast(voice_id="en-alba", text=script)["clip"]
clip["parts"], clip.get("partsDone")  # progress while status is "processing"

Speech to text

Audio (or video: only the audio track is read) becomes a transcript with timed segments, SRT and VTT. Up to 5 GB and eight hours per file.

t = vx.transcribe_file(
    audio="interview.mp3",  # bytes, a path or a binary file object; files are streamed from disk
    title="Interview with the lighthouse keeper",
    # language="vi",        # optional: Voixa detects it well on its own
    # task="translate",     # recognise any language, return English
    # words=True,           # per-word timestamps
    # prompt="Voixa, Vovix",  # spelling hints for names and jargon
)
print(t["language"], t["durationSec"], t["text"])
print(t["segments"][0])  # {"id", "start", "end", "text"}
print(t["srtUrl"])       # signed SRT/VTT/TXT/JSON URLs, valid for 1 hour

The format is taken from format=, else from the file name (file_name=, or the name of the path or file you pass), else MP3.

Transcription runs on a machine that starts on demand, so it is never synchronous: transcribe() returns status: "processing" as soon as the upload finishes and you poll, while transcribe_file() does the polling for you (two to three minutes before recognition starts on a cold queue, then roughly one minute of work per five minutes of audio).

transcript = vx.transcribe(audio="talk.m4a")["transcript"]
done = vx.wait_for_transcript(transcript["transcriptId"])
# progress while processing: doneSec / durationSec, jobStatus

items = vx.list_transcripts(project="episode-12")["items"]   # no text in lists
full = vx.get_transcript(items[0]["transcriptId"])            # text + segments + URLs
vx.delete_transcripts([transcript["transcriptId"]])           # also deletes the upload

Languages you may pin, the accepted formats and the limits come from stt_languages(). Transcription is charged per second of audio once it finishes.

Library

projects = vx.projects()["projects"]                  # groups with clip counts
page = vx.list_clips(project="episode-12")            # newest first
from_api = vx.list_clips(source="api")                # clips made with an API key (vs. "studio")
podcasts = vx.list_clips(kind="podcast", limit=50)
more = vx.list_clips(project="episode-12", cursor=page.get("cursor"))
vx.delete_clips([c["clipId"] for c in page["items"]])

share = vx.share_clip(clip["clipId"])   # {"shareUrl": "https://voixa.vovix.io/s/...", "shared": True}
vx.unshare_clip(clip["clipId"])         # the link stops working
vx.get_clip(clip["clipId"])
vx.wait_for(clip["clipId"], timeout=300)

Errors and limits

Every API failure raises VoixaError with:

  • status: HTTP status code
  • message: human-readable, safe to show (also str(err))
  • code: NOT_APPROVED, INSUFFICIENT_TOKENS, CLONE_LIMIT, VOICE_NOT_READY, INVALID, AUDIO_TOO_LARGE, NO_AUDIO (QUOTA and AUDIO_QUOTA are legacy); the SDK itself sets TIMEOUT, SYNTH_FAILED, DESIGN_FAILED, TRANSCRIBE_FAILED
  • retry_after: seconds, from the Retry-After header
  • needed, available, refills_at (insufficient tokens); resets_at, used, limit (legacy quotas)
  • body: the full JSON error body
from voixa import VoixaError

try:
    vx.say(voice_id="vi-truc-ly", text=long_text)
except VoixaError as e:
    if e.code == "INSUFFICIENT_TOKENS":
        print(f"Need {e.needed}, have {e.available}; free tokens refill at {e.refills_at}")
    elif e.code == "TIMEOUT":
        ...  # still processing: poll later with vx.wait_for(...)
    else:
        raise

Requests the SDK can reject without calling the API (empty or too long text, a design description under three characters, audio over 5 GB) raise the same VoixaError with status 400. Network failures are not wrapped: they raise the standard library's OSError subclasses (urllib.error.URLError, socket timeouts).

  • One speak call takes up to 5 000 characters (fewer for Chinese, Japanese and Korean; see languages()). Line breaks and repeated spaces collapse to one space before counting, on your side and on the server.
  • One podcast takes up to 100 000 characters.
  • One transcription takes up to 5 GB and eight hours of audio.
  • API keys are rate-limited to 1 request/second and 1 000 calls/day; new keys activate after two to three minutes.
  • Speech, streaming, transcription, clones and designs are paid in tokens (see Tokens above); read the wallet with account().

Method reference

GroupMethods
Accountaccount(), usage(limit, cursor), languages(), samples(language)
Voicesvoices(language), get_voice(id), create_voice(...), update_voice(id, ...), delete_voice(id), wait_for_voice(id)
Voice designdesign_voice(...), get_design(id), list_designs(limit, cursor), wait_for_design(id), save_design(id, ...), delete_design(id), create_voice_from_description(...)
Speechspeak(...), say(...), stream(...)
Podcastscreate_podcast(...), make_podcast(...)
Libraryget_clip(id), list_clips(...), projects(), delete_clips(ids), share_clip(id), unshare_clip(id), wait_for(id)
Speech to textstt_languages(), transcribe(...), transcribe_file(...), get_transcript(id), list_transcripts(...), delete_transcripts(ids), wait_for_transcript(id)
Helpersestimate_tokens(rates, ...), constants MAX_SPEAK_CHARS, MAX_PODCAST_CHARS, MAX_AUDIO_BYTES, MAX_AUDIO_MINUTES, DEFAULT_BASE_URL

A signed-in user of your own web app can call the API with their ID token instead of a key: Voixa(id_token=...).

Was this page helpful?