Transcripts

The other direction: an audio or video file becomes text with timed segments, plus SRT, VTT and plain-text files. Up to 5 GB and eight hours per file, in far more languages than Voixa speaks.

How it works

The bytes never pass through the API, which is what lets files be this large:

  1. POST /transcribe registers the transcript and returns an uploadUrl (valid 15 minutes).
  2. You PUT the file to uploadUrl. The job starts by itself the moment the upload lands.
  3. You poll GET /transcripts/{transcriptId} until it is ready.

Recognition runs on machines that start on demand, so it is never synchronous. Expect two to three minutes before work starts when nothing has run for a while, then roughly one minute per five to ten minutes of audio — long files are cut into parts that run side by side. doneSec, durationSec and partsDone report progress.

Accepted formats: WAV, MP3, M4A, MP4, OGG, WebM, FLAC and AAC. Video is fine; only the audio track is read.

The transcript model

Properties

  • Name
    transcriptId
    Type
    string
    Description

    Unique id.

  • Name
    status
    Type
    string
    Description

    processing, ready or failed (with failReason).

  • Name
    language
    Type
    string
    Description

    The language detected (or the one you pinned), once ready.

  • Name
    durationSec
    Type
    number
    Description

    Length of the audio.

  • Name
    text
    Type
    string
    Description

    The whole transcript. Only on GET /transcripts/{transcriptId}, not in lists.

  • Name
    segments
    Type
    array
    Description

    { id, start, end, text } in seconds, plus words when you asked for them, and speaker (with speakerScore, and turns when several people talk in one segment) when you asked for diarize. Only on the single transcript.

  • Name
    speakers
    Type
    array
    Description

    With diarize: the voices found, most talkative first — { id, seconds, lines, pitchHz, known, similarity, samples }. samples are segment ids where the voice is clearest; known means the voice was already in the project. Only on the single transcript. speakersError is set instead when labelling failed; the transcript itself is still complete.

  • Name
    srtUrl
    Type
    string
    Description

    Signed one-hour URLs for the subtitle and text files: srtUrl, vttUrl, txtUrl, jsonUrl, and audioUrl for your upload.

  • Name
    project
    Type
    string
    Description

    Content group — the same names as clips.

  • Name
    source
    Type
    string
    Description

    api or studio.


POST/v1/transcribe

Create a transcript

Registers the transcript and returns where to upload. The response is always processing; the job waits for your upload.

Optional attributes

  • Name
    format
    Type
    string
    Description

    wav, mp3 (default), m4a, mp4, ogg, webm, flac or aac. It sets the content-type your PUT must use.

  • Name
    language
    Type
    string
    Description

    Pin the language (see GET /stt/languages). Leave it out and it is detected — usually the right choice.

  • Name
    task
    Type
    string
    Description

    transcribe (default), or translate to recognise any language and return English.

  • Name
    words
    Type
    boolean
    Description

    Per-word timestamps (about 10% slower).

  • Name
    prompt
    Type
    string
    Description

    Spellings of names and terms that appear in the audio, up to 400 characters.

  • Name
    title
    Type
    string
    Description

    Shown in Studio; fileName is kept for display too.

  • Name
    project
    Type
    string
    Description

    A content group. With diarize, transcripts of one project share their speakers.

  • Name
    diarize
    Type
    boolean
    Description

    Label who says each segment (default false). See Speaker labels.

Request

POST
/v1/transcribe
# 1. register, get the upload URL
curl -X POST https://api.voixa.vovix.io/v1/transcribe \
  -H "x-api-key: $VOIXA_API_KEY" -H "content-type: application/json" \
  -d '{ "format": "mp3", "fileName": "interview.mp3", "project": "episode-12" }'

# 2. upload — this starts the job
curl -X PUT "$UPLOAD_URL" -H "content-type: audio/mpeg" --data-binary @interview.mp3

# 3. poll
curl https://api.voixa.vovix.io/v1/transcripts/$TRANSCRIPT_ID \
  -H "x-api-key: $VOIXA_API_KEY"

Response

{
  "transcript": { "transcriptId": "…", "status": "processing", "format": "mp3" },
  "uploadUrl": "https://…",
  "expiresIn": 900,
  "maxBytes": 5368709120
}

GET/v1/transcripts

List transcripts

Your transcripts, newest first — summaries without text or segments. Filter with project, source and status; page with limit and cursor.

Request

GET
/v1/transcripts
curl "https://api.voixa.vovix.io/v1/transcripts?status=ready" \
  -H "x-api-key: $VOIXA_API_KEY"

Speaker labels

Send diarize: true and every segment says who speaks it: speaker: "S1", "S2", … The transcript also carries a speakers list — how long each voice talks, its median pitch (below about 165 Hz is usually a man), and the ids of the segments where it is clearest, so you can play a few seconds and name the person.

{ "id": 41, "start": 240.4, "end": 244.8, "text": "You just put the book in that woman's hand.", "speaker": "S3", "speakerScore": 0.82 }
  • A series keeps its cast. Transcripts with the same project share their speakers: the voice that was S3 in the first episode is S3 in the second (known: true, with a similarity). New voices get new ids. Run the episodes of one project one after another, not at the same time.
  • It is a best guess. Lines of a second or more in clean speech are usually right. Interjections under half a second, people talking over each other, and speech under loud music are where it goes wrong — speakerScore is low on the lines worth a second look.
  • Timing. While it runs the transcript stays processing with jobStatus: "LABELLING_SPEAKERS": about a minute per ten minutes of audio; recordings over 90 minutes run on a machine that takes two to three minutes to start. No extra charge.
  • The SRT, VTT and TXT files do not carry speaker names; build them from segments if you need them.

DELETE/v1/transcripts

Delete transcripts

Deletes up to 100 transcripts with their files and the uploaded audio, and cancels a job that is still running.

Request

DELETE
/v1/transcripts
curl -X DELETE https://api.voixa.vovix.io/v1/transcripts \
  -H "x-api-key: $VOIXA_API_KEY" -H "content-type: application/json" \
  -d '{ "transcriptIds": ["'$TRANSCRIPT_ID'"] }'

Was this page helpful?