Speak

speak turns text into a clip: a WAV in your library with a signed URL. Short texts come back ready in the same call; longer ones come back processing and finish a few seconds later.

The clip model

Properties

  • Name
    clipId
    Type
    string
    Description

    Unique id of the clip.

  • Name
    voiceId
    Type
    string
    Description

    The voice it was read with; voiceName is its display name.

  • Name
    language
    Type
    string
    Description

    The voice's language.

  • Name
    text
    Type
    string
    Description

    The text as it was read, after cleaning (see Text cleaning).

  • Name
    chars
    Type
    integer
    Description

    Characters counted against your allowance.

  • Name
    project
    Type
    string
    Description

    The content group you passed, if any.

  • Name
    source
    Type
    string
    Description

    api for calls with an API key, studio for clips made in Studio.

  • Name
    status
    Type
    string
    Description

    processing, ready or failed (with failReason).

  • Name
    url
    Type
    string
    Description

    Signed WAV URL, valid for one hour. Present when ready.

  • Name
    durationSec
    Type
    number
    Description

    Length of the audio, once ready. sampleRate is its sample rate.

  • Name
    parts
    Type
    integer
    Description

    For long texts: how many parts were produced in parallel. partsDone reports progress while processing.

  • Name
    shareUrl
    Type
    string
    Description

    Public listening page, after sharing.


POST/v1/speak

Create a clip

Reads text with the voice and stores the clip in your library. The API waits up to about 20 seconds for the audio, which covers most texts; otherwise the clip comes back processing and you poll GET /clips/{clipId} every second or two. Texts longer than about 250 characters are split into parts that run in parallel and are joined into one WAV — to you it is still one clip.

Required attributes

  • Name
    voiceId
    Type
    string
    Description

    A voice from GET /voices.

  • Name
    text
    Type
    string
    Description

    Up to 5 000 characters (1 200 for Chinese, Japanese and Korean). Longer? Use a podcast.

Optional attributes

  • Name
    project
    Type
    string
    Description

    A content group such as episode-12. Letters, digits, spaces and . _ -, up to 60.

  • Name
    wait
    Type
    boolean
    Description

    true (default) waits up to ~20 s for the audio; false answers at once with processing.

Request

POST
/v1/speak
curl -X POST https://api.voixa.vovix.io/v1/speak \
  -H "x-api-key: $VOIXA_API_KEY" \
  -H "content-type: application/json" \
  -d '{
    "voiceId": "vi-truc-ly",
    "text": "Xin chào, cuộc gọi của bạn đang được kết nối.",
    "project": "ivr"
  }'

Response

{
  "clip": {
    "clipId": "e0f43e3b-e343-430d-be01-6b6d4309f8d8",
    "voiceId": "vi-truc-ly",
    "voiceName": "Trúc Ly",
    "language": "vi",
    "text": "Xin chào, cuộc gọi của bạn đang được kết nối.",
    "chars": 45,
    "project": "ivr",
    "source": "api",
    "status": "ready",
    "durationSec": 3.04,
    "sampleRate": 48000,
    "url": "https://…"
  }
}

Text cleaning

Before counting and reading, Voixa removes what would sound wrong read aloud: markdown symbols (*, #, >, list bullets, code fences), URLs, emoji, bracketed notes, and repeated punctuation (!!! becomes !). Line breaks and runs of spaces collapse to one space. Numbers, dates, times, percentages and money stay — they are read naturally (8h30, 1.250.000 đồng, 68,5%).

Was this page helpful?