Developers
API reference
Transcribe audio programmatically. All endpoints are under the API base URL and require an API key.
Base URL
https://api-sotaspeech.notexapp.comAuthentication
Send your API key as a Bearer token, or as the x-api-key header on the queued endpoints. /v1/audio/transcriptions reads only the Authorization header, matching the OpenAI SDK.
Authorization: Bearer YOUR_API_KEY
# or
x-api-key: YOUR_API_KEYGetting an API key
This is a closed evaluation deployment: keys are issued by the team, and there is no self-service signup. The service is not for sale and has no paid plans. Ask your project contact for a key.
Data ownership
Every job is tied to the API key that created it. GET /history and GET /transcribe/{task_id} only return that key's jobs; a different valid key gets 404. The key used by the web UI is public (embedded in the page bundle), so jobs created through the UI share one bucket.
Limits
file size2000 MBresponse_formatjson · verbose_json · text · srt · vttrate limit30 / min / keyconcurrency1Maximum 2 GB per file (413 above that). No separate audio-duration cap. Job-creating endpoints allow 30 requests per minute per API key by default, then answer 429 with a Retry-After header. One GPU session runs at a time, so jobs queue.
Recommended flow (large files)
Upload the bytes straight to storage with a presigned URL, enqueue by object key, then poll — so multi-MB audio never transits the API.
/uploads/presigncurl -X POST https://api-sotaspeech.notexapp.com/uploads/presign \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"filename":"meeting.wav","content_type":"audio/wav"}'
# -> { "object_key": "...", "upload_url": "https://.../..." }{upload_url}curl -X PUT "{upload_url}" \
-H "Content-Type: audio/wav" \
--data-binary @meeting.wav/transcribe/objectLanguage hint: vi, en, ja, ko, or omit for auto-detect.
curl -X POST https://api-sotaspeech.notexapp.com/transcribe/object \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"object_key":"...","language":"vi","context":"SotaSpeech, Hà Nội"}'
# -> { "task_id": "abc123", "status": "queued" }Hotwords / context
Pass a context string to bias proper nouns — people, products, acronyms. /v1/audio/transcriptions takes the same hint through its prompt field.
/transcribe/{task_id}Poll until status is done; the result carries the transcript, timestamped segments, and speaker labels.
Job states
queuedaccepted, waiting for the GPUprocessingdecoding nowdoneresult is presentfailederror explains whyqueued then processing, then done or failed. While a job is queued or processing the poll response also carries a queue block with the number waiting and your position.
curl https://api-sotaspeech.notexapp.com/transcribe/abc123 \
-H "Authorization: Bearer YOUR_API_KEY"
# -> { "status": "done", "result": {
# "text": "...", "language": "Vietnamese",
# "duration_sec": 300.9, "segments": [
# { "start": 1.7, "end": 10.1, "start_fmt": "00:01.720",
# "text": "...", "speaker": "SPEAKER_00" } ],
# "speaker_turns": [ { "start": 1.7, "end": 10.1,
# "speaker": "SPEAKER_00" } ],
# "segment_count": 19 } }Speaker labels
The queued result includes segments[].speaker and speaker_turns. The OpenAI-compatible response does NOT: its segments carry only id, start, end and text, matching the SDK's typed models.
/transcribeLanguage hint: vi, en, ja, ko, or omit for auto-detect.
curl -X POST https://api-sotaspeech.notexapp.com/transcribe \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"file_url":"https://example.com/audio.mp3","language":"vi"}'
# -> { "task_id": "...", "status": "queued" } (then poll as above)
# Public http(s) URLs only: private, loopback and cloud-metadata
# addresses are refused, on the original URL and on every redirect.One-shot (OpenAI-compatible)
A synchronous multipart endpoint compatible with OpenAI’s audio transcription API.
/v1/audio/transcriptionscurl -X POST https://api-sotaspeech.notexapp.com/v1/audio/transcriptions \
-H "Authorization: Bearer YOUR_API_KEY" \
-F "file=@meeting.wav" \
-F "model=sotaspeech" \
-F "language=vi" \
-F "prompt=SotaSpeech, Hà Nội" \
-F "response_format=srt"Response formats
json{ "text": "..." }verbose_jsontext + language + duration + segmentstextplain text, no JSON envelopesrtSubRip subtitlesvttWebVTT subtitlesOnly /v1/audio/transcriptions accepts response_format: json, verbose_json, text, srt, vtt. The queued endpoints always return the full JSON result, which carries more than the OpenAI shape (per-segment timings, diarization turns, audio_url).
Errors
400bad language, unusable file_url, bad response_format401missing or wrong API key404unknown task id, expired link, or another key's job413file larger than 2 GB422a query/path parameter out of range429rate limit exceeded — retry after Retry-After503model not configured, storage down, or the queue is reconnectingThe queued endpoints use FastAPI's {"detail": "..."} envelope. /v1/audio/transcriptions uses OpenAI's {"error": {message, type, param, code}}, because the SDK parses that shape.
# queued endpoints
{ "detail": "Missing or invalid API key. ..." }
# /v1/audio/transcriptions
{ "error": { "message": "...", "type": "invalid_request_error",
"param": null, "code": null } }Polling only — no webhooks
Results are collected by polling GET /transcribe/{task_id}. There is no webhook or push notification yet; the web UI polls every 1.2 s.
OpenAPI and SDKs
The full spec and an interactive console are served by the API itself (links below). There is no first-party SDK yet — use any HTTP client, or the OpenAI SDK against the compatible endpoint.
https://api-sotaspeech.notexapp.com/openapi.jsonhttps://api-sotaspeech.notexapp.com/docs