Voice design

Describe a voice in plain words and Voixa performs it: up to three takes of that voice, each a different performer. Keep the take you like and it becomes a voice of your account, usable everywhere a voice is: speak, podcasts, streaming.

Made for dubbing and storytelling: design one voice per character — the old storyteller, the child, the villain — then read every line with it at normal speech prices. Designing is a one-off; reading is as cheap as any other voice.

How it works

  1. Describe. Age, gender, timbre, pace, mood. Any language; English, Vietnamese and Chinese descriptions work best.
  2. Listen. POST /designs performs 1–3 takes (about a minute each; the first design after a quiet spell waits for the engine to warm up). Poll GET /designs/{designId}.
  3. Keep. POST /designs/{designId}/save turns one take into a voice. From then on it is an ordinary cloned voice with a voiceId.

Takes are drafts kept for 23 hours (expiresAt). Voice design works for every language with voice cloning: Vietnamese, English, French, German, Italian, Spanish and Portuguese (designable in GET /languages).

Tokens. Each take costs rates.design tokens (2 000 by default); keeping one costs rates.clone like any clone. A take that fails is refunded. See Tokens.

Writing a good description

  • Say who, then how. "A seven-year-old girl, clear and high voice, cheerful" beats "cute voice".
  • Give three to five traits. Age, timbre (deep, husky, bright, breathy), pace (slow, fast), mood (calm, menacing, warm).
  • Three takes, not one. The same description gives three different performers; pick the one that fits.
  • Change the line to fit the character. text is what the takes read. A line in character tells you more than the default sentence.

The design model

  • Name
    designId
    Type
    string
    Description

    Unique identifier.

  • Name
    description
    Type
    string
    Description

    The description the takes were performed from.

  • Name
    language
    Type
    string
    Description

    Language of the voice.

  • Name
    text
    Type
    string
    Description

    The line every take reads.

  • Name
    count
    Type
    integer
    Description

    Number of takes, 1–3.

  • Name
    status
    Type
    string
    Description

    processing until every take is done, then ready (at least one take worked) or failed.

  • Name
    candidates
    Type
    array
    Description

    One entry per take: index, status, durationSec, and url (signed, one hour) once ready.

  • Name
    savedVoiceIds
    Type
    array
    Description

    Voices kept from this design.

  • Name
    expiresAt
    Type
    string
    Description

    When the takes are deleted. expired: true after that.

  • Name
    tokens
    Type
    integer
    Description

    Tokens charged (refunds for failed takes are in your usage).


POST/v1/designs

Design a voice

Always answers at once with "status": "processing". Poll GET /designs/{designId} every few seconds, or let the SDK's waitForDesign() do it.

Required attributes

  • Name
    description
    Type
    string
    Description

    The voice in plain words, 3–500 characters.

  • Name
    language
    Type
    string
    Description

    vi, en, fr, de, it, es or pt.

Optional attributes

  • Name
    count
    Type
    integer
    Description

    Takes to perform, 1–3. Default 3.

  • Name
    text
    Type
    string
    Description

    The line to perform, up to 160 characters (about 8–10 seconds — enough to clone from). Defaults to the language's sample sentence.

Request

POST
/v1/designs
curl -X POST https://api.voixa.vovix.io/v1/designs \
  -H "x-api-key: $VOIXA_API_KEY" \
  -H "content-type: application/json" \
  -d '{
    "description": "An old man in his seventies, hoarse and deep, telling a story slowly",
    "language": "vi",
    "count": 3
  }'

Response

{
  "design": {
    "designId": "5b2f0c1e-8f7a-4d0b-9a51-2f0f3d3f6c11",
    "description": "An old man in his seventies, hoarse and deep, telling a story slowly",
    "language": "vi",
    "text": "Xin chào, đây là giọng đọc mẫu của Voixa. Nội dung của bạn sẽ được đọc bằng chất giọng này.",
    "count": 3,
    "status": "processing",
    "candidates": [
      { "index": 0, "status": "processing" },
      { "index": 1, "status": "processing" },
      { "index": 2, "status": "processing" }
    ],
    "savedVoiceIds": [],
    "tokens": 6000,
    "expiresAt": "2026-09-30T10:12:00.000Z",
    "expired": false
  }
}

GET/v1/designs/:designId

Retrieve a design

Returns the design with a signed url for every take that is ready. GET /v1/designs lists your designs, newest first (limit, cursor).

Response

{
  "design": {
    "designId": "5b2f0c1e-8f7a-4d0b-9a51-2f0f3d3f6c11",
    "status": "ready",
    "candidates": [
      { "index": 0, "status": "ready", "durationSec": 6.72, "url": "https://…/0.wav?…" },
      { "index": 1, "status": "ready", "durationSec": 7.36, "url": "https://…/1.wav?…" },
      { "index": 2, "status": "failed" }
    ]
  }
}

POST/v1/designs/:designId/save

Keep a take

Saves one take as a voice of your account and returns it, like POST /voices. Its sample is generated right away. The voice carries designId. You can keep several takes from the same design.

  • Name
    candidate
    Type
    integer
    Description

    Index of the take (0-based). Required.

  • Name
    name
    Type
    string
    Description

    Voice name, 1–60 characters. Required.

  • Name
    gender
    Type
    string
    Description

    female, male or neutral.

  • Name
    description
    Type
    string
    Description

    Defaults to the design's description.

Request

POST
/v1/designs/:designId/save
curl -X POST https://api.voixa.vovix.io/v1/designs/$DESIGN_ID/save \
  -H "x-api-key: $VOIXA_API_KEY" \
  -H "content-type: application/json" \
  -d '{ "candidate": 1, "name": "Grandpa Ba", "gender": "male" }'

Was this page helpful?