Developers

API reference

Transcribe audio programmatically. All endpoints are under the API base URL and require an API key.

01

Base URL

https://api-sotaspeech.notexapp.com

Authentication

Send your API key as a Bearer token, or as the x-api-key header on the queued endpoints. /v1/audio/transcriptions reads only the Authorization header, matching the OpenAI SDK.

Authorization: Bearer YOUR_API_KEY
# or
x-api-key: YOUR_API_KEY

Getting an API key

This is a closed evaluation deployment: keys are issued by the team, and there is no self-service signup. The service is not for sale and has no paid plans. Ask your project contact for a key.

02

Data ownership

Every job is tied to the API key that created it. GET /history and GET /transcribe/{task_id} only return that key's jobs; a different valid key gets 404. The key used by the web UI is public (embedded in the page bundle), so jobs created through the UI share one bucket.

Limits

file size2000 MB
response_formatjson · verbose_json · text · srt · vtt
rate limit30 / min / key
concurrency1

Maximum 2 GB per file (413 above that). No separate audio-duration cap. Job-creating endpoints allow 30 requests per minute per API key by default, then answer 429 with a Retry-After header. One GPU session runs at a time, so jobs queue.

03

Recommended flow (large files)

Upload the bytes straight to storage with a presigned URL, enqueue by object key, then poll — so multi-MB audio never transits the API.

1 · presign
POST/uploads/presign
curl -X POST https://api-sotaspeech.notexapp.com/uploads/presign \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"filename":"meeting.wav","content_type":"audio/wav"}'
# -> { "object_key": "...", "upload_url": "https://.../..." }
2 · upload bytes
PUT{upload_url}
curl -X PUT "{upload_url}" \
  -H "Content-Type: audio/wav" \
  --data-binary @meeting.wav
3 · enqueue
POST/transcribe/object

Language hint: vi, en, ja, ko, or omit for auto-detect.

curl -X POST https://api-sotaspeech.notexapp.com/transcribe/object \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"object_key":"...","language":"vi","context":"SotaSpeech, Hà Nội"}'
# -> { "task_id": "abc123", "status": "queued" }

Hotwords / context

Pass a context string to bias proper nouns — people, products, acronyms. /v1/audio/transcriptions takes the same hint through its prompt field.

4 · Poll the job
GET/transcribe/{task_id}

Poll until status is done; the result carries the transcript, timestamped segments, and speaker labels.

Job states

queuedaccepted, waiting for the GPU
processingdecoding now
doneresult is present
failederror explains why

queued then processing, then done or failed. While a job is queued or processing the poll response also carries a queue block with the number waiting and your position.

curl https://api-sotaspeech.notexapp.com/transcribe/abc123 \
  -H "Authorization: Bearer YOUR_API_KEY"
# -> { "status": "done", "result": {
#      "text": "...", "language": "Vietnamese",
#      "duration_sec": 300.9, "segments": [
#        { "start": 1.7, "end": 10.1, "start_fmt": "00:01.720",
#          "text": "...", "speaker": "SPEAKER_00" } ],
#      "speaker_turns": [ { "start": 1.7, "end": 10.1,
#                           "speaker": "SPEAKER_00" } ],
#      "segment_count": 19 } }
04

Speaker labels

The queued result includes segments[].speaker and speaker_turns. The OpenAI-compatible response does NOT: its segments carry only id, start, end and text, matching the SDK's typed models.

POST/transcribe

Language hint: vi, en, ja, ko, or omit for auto-detect.

curl -X POST https://api-sotaspeech.notexapp.com/transcribe \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"file_url":"https://example.com/audio.mp3","language":"vi"}'
# -> { "task_id": "...", "status": "queued" }  (then poll as above)
# Public http(s) URLs only: private, loopback and cloud-metadata
# addresses are refused, on the original URL and on every redirect.
05

One-shot (OpenAI-compatible)

A synchronous multipart endpoint compatible with OpenAI’s audio transcription API.

POST/v1/audio/transcriptions
curl -X POST https://api-sotaspeech.notexapp.com/v1/audio/transcriptions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -F "file=@meeting.wav" \
  -F "model=sotaspeech" \
  -F "language=vi" \
  -F "prompt=SotaSpeech, Hà Nội" \
  -F "response_format=srt"

Response formats

json{ "text": "..." }
verbose_jsontext + language + duration + segments
textplain text, no JSON envelope
srtSubRip subtitles
vttWebVTT subtitles

Only /v1/audio/transcriptions accepts response_format: json, verbose_json, text, srt, vtt. The queued endpoints always return the full JSON result, which carries more than the OpenAI shape (per-segment timings, diarization turns, audio_url).

06

Errors

400bad language, unusable file_url, bad response_format
401missing or wrong API key
404unknown task id, expired link, or another key's job
413file larger than 2 GB
422a query/path parameter out of range
429rate limit exceeded — retry after Retry-After
503model not configured, storage down, or the queue is reconnecting

The queued endpoints use FastAPI's {"detail": "..."} envelope. /v1/audio/transcriptions uses OpenAI's {"error": {message, type, param, code}}, because the SDK parses that shape.

# queued endpoints
{ "detail": "Missing or invalid API key. ..." }

# /v1/audio/transcriptions
{ "error": { "message": "...", "type": "invalid_request_error",
             "param": null, "code": null } }
07

Polling only — no webhooks

Results are collected by polling GET /transcribe/{task_id}. There is no webhook or push notification yet; the web UI polls every 1.2 s.

OpenAPI and SDKs

The full spec and an interactive console are served by the API itself (links below). There is no first-party SDK yet — use any HTTP client, or the OpenAI SDK against the compatible endpoint.

https://api-sotaspeech.notexapp.com/openapi.jsonhttps://api-sotaspeech.notexapp.com/docs