Go SDK

voixa-go wraps every endpoint with typed requests and responses, takes a context.Context on every call, waits for clips, designs and transcripts for you, and uses only the standard library. Go 1.22 or newer.

Install

go

go get github.com/vovix-ai/voixa-go

Reference documentation: pkg.go.dev/github.com/vovix-ai/voixa-go.

Quick start

package main

import (
	"context"
	"fmt"
	"log"

	voixa "github.com/vovix-ai/voixa-go"
)

func main() {
	ctx := context.Background()
	vx := voixa.NewClient("") // reads VOIXA_API_KEY; create a key in Studio → API keys

	voices, err := vx.Voices(ctx, voixa.English)
	if err != nil {
		log.Fatal(err)
	}
	clip, err := vx.Say(ctx, voixa.SpeakRequest{
		VoiceID: voices[0].VoiceID,
		Text:    "The lighthouse keeper switched the lamp on at dusk.",
		Project: "episode-12", // optional: group clips by content
	})
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(clip.URL, clip.DurationSec) // signed WAV URL, valid for 1 hour
}

Say returns when the clip is ready. Use Speak with Wait: voixa.Bool(false) plus WaitForClip to manage the wait yourself — the API answers in under a second and you poll.

Every method takes a context.Context first; cancelling it stops a request or a wait immediately.

Client options

vx := voixa.NewClient(apiKey,
	voixa.WithBaseURL("https://api.voixa.vovix.io/v1"), // the default
	voixa.WithTimeout(60*time.Second),                   // per API call; default 30 s
	voixa.WithHTTPClient(myHTTPClient),                  // proxies, tracing, pacing
	voixa.WithUserAgent("my-app/1.2"),
)

An empty key falls back to the VOIXA_API_KEY environment variable. A client without credentials returns an error on every call instead of panicking.

Streaming (Vietnamese)

Hear the first words in about two seconds: audio arrives clause by clause while the rest is voiced. The stream is an io.ReadCloser; close it when done.

s, err := vx.Stream(ctx, voixa.StreamRequest{VoiceID: "vi-truc-ly", Text: "Xin chào, Voixa có thể giúp gì?", SampleRate: 16000})
if err != nil {
	log.Fatal(err)
}
defer s.Close()
io.Copy(player, s) // 16-bit little-endian mono PCM at s.SampleRate

SampleRate: 48000, 24000 (default), 16000 or 8000; Format: voixa.StreamPCM (default) or voixa.StreamWAV. Costs 1.5 tokens per character, charged when the stream opens (s.Tokens). The stream lasts as long as the context you pass; the per-call timeout does not cut it off.

Tokens

Every request is priced in tokens: 1 token = 1 Vietnamese character. Free tokens refill daily; top-ups carry over.

acc, _ := vx.Account(ctx)
w := acc.Tokens // Available, FreeToday, DailyFree, Balance, UsedToday, RefillsAt, Rates
w.Rates.EstimateSpeak(voixa.Vietnamese, 1200) // → 1200
w.Rates.EstimateStream(500)
w.Rates.EstimateTranscription(3600) // seconds of audio
w.Rates.EstimateClone()
w.Rates.EstimateDesign(3)

page, _ := vx.Usage(ctx, voixa.PageOptions{Limit: 20}) // every spend, refund and top-up

A request you cannot afford fails with an *voixa.Error of status 402 and code INSUFFICIENT_TOKENS (voixa.IsInsufficientTokens(err)); Needed, Available and RefillsAt tell you by how much and until when. Failed clips and voices are refunded.

Voices and samples

langs, _ := vx.Languages(ctx)                  // codes, labels, limits, clone-recording guides
samples, _ := vx.Samples(ctx, voixa.French)    // built-in voices with public sample URLs
voices, _ := vx.Voices(ctx, "")                // built-in voices, your clones, shared clones
voice, _ := vx.GetVoice(ctx, "vi-truc-ly")

Voice design: a voice from a description

No recording needed. Describe the voice, listen to up to three options, keep the one you like. A saved design is an ordinary voice of your account: use it with Say, MakePodcast and Stream. Characters for dubbing, narrators, ads — one voice per role.

d, err := vx.DesignVoice(ctx, voixa.DesignVoiceRequest{
	Description: "An old man in his seventies, hoarse and deep, speaking slowly",
	Language:    voixa.Vietnamese, // any language Languages() marks Designable
	Count:       3,                // 1–3 options, Rates.Design tokens each
})
done, err := vx.WaitForDesign(ctx, d.DesignID) // about a minute per option
for _, c := range done.Candidates {
	fmt.Println(c.Index, c.Status, c.URL) // listen, then pick
}

voice, err := vx.SaveDesign(ctx, done.DesignID, voixa.SaveDesignRequest{Candidate: 1, Name: "Grandpa Ba", Gender: voixa.GenderMale})
_, err = vx.Say(ctx, voixa.SpeakRequest{VoiceID: voice.VoiceID, Text: "Ngày ấy, ông vẫn nhớ con đường làng."})

// Or in one call, keeping the first option:
narrator, err := vx.CreateVoiceFromDescription(ctx, voixa.VoiceFromDescriptionRequest{
	Description: "A calm female narrator, warm and clear", Language: voixa.English, Name: "Narrator",
})

Options are drafts kept for 23 hours (ExpiresAt). Saving one costs Rates.Clone tokens. GetDesign, ListDesigns and DeleteDesign manage your designs; deleting a design keeps the voices saved from it.

Voice cloning and sharing

f, _ := os.Open("reference.wav") // 3–10 s of clean speech, WAV or MP3, max 10 MB
defer f.Close()
voice, err := vx.CreateVoice(ctx, voixa.CreateVoiceRequest{
	Name:     "Studio narrator",
	Language: voixa.English,
	Gender:   voixa.GenderFemale,
	Audio:    f, // any io.Reader; Size is optional for files and in-memory readers
})
_, err = vx.Say(ctx, voixa.SpeakRequest{VoiceID: voice.VoiceID, Text: "Hello from my own voice."})

err = vx.UpdateVoice(ctx, voice.VoiceID, voixa.VoiceUpdate{Shared: voixa.Bool(true)}) // let every Voixa account use it
err = vx.DeleteVoice(ctx, voice.VoiceID)

CreateVoice uploads the recording, registers the voice and waits for its sample. Voice cloning is available for Vietnamese, English, French, German, Italian, Spanish and Portuguese. Chinese, Japanese and Korean use built-in voices only; Languages reports Cloneable per language and CreateVoice rejects those languages with a 400.

Podcasts

Long scripts, up to 100 000 characters, become one episode: Voixa splits the text into sentences, produces the parts in parallel and joins them into a single WAV. Blank lines become short pauses.

episode, err := vx.MakePodcast(ctx, voixa.PodcastRequest{VoiceID: "vi-truc-ly", Title: "Episode 12", Text: script, Project: "my-show"})
episode.URL          // one WAV, signed for an hour
episode.DurationSec  // e.g. 4800 for an 80-minute episode

// or without waiting:
clip, err := vx.CreatePodcast(ctx, voixa.PodcastRequest{VoiceID: "en-alba", Text: script})
clip.Parts, clip.PartsDone // progress while Status is processing

MakePodcast waits up to two hours by default; check Status on the result.

Speech to text

Audio (or video — only the audio track is read) becomes a transcript with timed segments, SRT and VTT.

f, _ := os.Open("interview.mp3")
defer f.Close()
t, err := vx.TranscribeFile(ctx, voixa.TranscribeRequest{
	Audio:    f,
	FileName: "interview.mp3",
	Title:    "Interview with the lighthouse keeper",
	// Language: "vi",              // optional — Voixa detects it well on its own
	// Task: voixa.TaskTranslate,   // recognise any language, return English
	// Words: true,                 // per-word timestamps
	// Prompt: "Voixa, Vovix",      // spelling hints for names and jargon
})
fmt.Println(t.Language, t.DurationSec, t.Text)
fmt.Println(t.Segments[0])  // {ID Start End Text}
fmt.Println(t.SRTURL)       // signed SRT/VTT/TXT/JSON URLs, valid for 1 hour

Transcription runs on a machine that starts on demand, so it is never synchronous: Transcribe returns StatusProcessing as soon as the upload is done and you poll, while TranscribeFile does the polling for you (two to three minutes before recognition starts on a cold queue, then roughly one minute of work per five minutes of audio).

t, err := vx.Transcribe(ctx, voixa.TranscribeRequest{Audio: f, FileName: "talk.m4a"})
done, err := vx.WaitForTranscript(ctx, t.TranscriptID)
// progress while processing: done.DoneSec / done.DurationSec, done.JobStatus

page, err := vx.ListTranscripts(ctx, voixa.ListTranscriptsOptions{Project: "episode-12"}) // no text in lists
full, err := vx.GetTranscript(ctx, page.Items[0].TranscriptID)                             // text + segments + URLs
n, err := vx.DeleteTranscripts(ctx, []string{t.TranscriptID})                              // also deletes the upload

The container format is inferred from FileName (or the name of an *os.File); set Format when the name does not say. A reader whose length cannot be known (not a file, bytes.Reader, bytes.Buffer or strings.Reader, and no Size) is read into memory before upload. Languages you may pin, the accepted formats and the limits come from STTLanguages. Transcription is charged per second of audio when it finishes.

Library

projects, _ := vx.Projects(ctx)                                                       // groups with clip counts
page, _ := vx.ListClips(ctx, voixa.ListClipsOptions{Project: "episode-12"})           // newest first
fromAPI, _ := vx.ListClips(ctx, voixa.ListClipsOptions{Source: voixa.SourceAPI})      // made with an API key (vs. Studio)
podcasts, _ := vx.ListClips(ctx, voixa.ListClipsOptions{Kind: voixa.ClipKindPodcast})
next, _ := vx.ListClips(ctx, voixa.ListClipsOptions{Project: "episode-12", Cursor: page.Cursor})

shareURL, _ := vx.ShareClip(ctx, clipID) // public page, no account needed
_ = vx.UnshareClip(ctx, clipID)          // the link stops working
n, _ := vx.DeleteClips(ctx, []string{clipID})

Errors and limits

Every API failure is a *voixa.Error with Status, Message, Code (NOT_APPROVED, INSUFFICIENT_TOKENS, CLONE_LIMIT, VOICE_NOT_READY, INVALID, AUDIO_TOO_LARGE, NO_AUDIO; QUOTA and AUDIO_QUOTA are legacy) and RetryAfter. The client adds TIMEOUT (a wait gave up while the work was still processing — the work continues), SYNTH_FAILED, TRANSCRIBE_FAILED and DESIGN_FAILED.

_, err := vx.Say(ctx, req)
var ve *voixa.Error
switch {
case voixa.IsInsufficientTokens(err):
	// top up, or wait for ve.RefillsAt
case voixa.IsRateLimited(err):
	// wait ve.RetryAfter
case errors.As(err, &ve):
	log.Printf("%d %s: %s", ve.Status, ve.Code, ve.Message)
case errors.Is(err, context.Canceled):
	// you cancelled
case err != nil:
	// network error
}

Helpers: IsInsufficientTokens, IsNotFound, IsRateLimited, IsTimeout, ErrorCode, ErrorStatus.

  • One Speak call takes up to 5 000 characters (1 200 for Chinese, Japanese and Korean). Whitespace and line breaks collapse to one space before counting; the client checks the length before sending (CollapseSpace, TextLength).
  • One podcast takes up to 100 000 characters (NormalizePodcastText).
  • One transcription takes up to 5 GB and eight hours of audio.
  • API keys are rate-limited to 1 request/second and 1 000 calls/day; new keys activate after two to three minutes.
  • Speech, streaming, transcription, designs and clones are paid in tokens (see Tokens above); read the wallet with Account.

Default waits

MethodPoll intervalGives up after
WaitForClip, Say, WaitForVoice, CreateVoice, SaveDesign2 s3 minutes
WaitForDesign4 s10 minutes
MakePodcast, WaitForTranscript, TranscribeFile5 s2 hours

Pass voixa.WaitOptions{PollInterval: …, Timeout: …} as the last argument to change them; zero fields keep the default.

Was this page helpful?