API

English only. Turn audio into transcripts from your own code: send a short clip and read the response, or submit a durable job for long audio.

Quickstart

Sign in and create a key

Sign in or create an account and verify your email. Open the Account page’s API keys tab to create a key, then copy the secret and store it securely. It is shown only once.

Your signed-in account includes free minutes for evaluation in the web app. API-key usage is currently separate from that balance. To try the API, create a key and make the request below.

Make your first request

Set MACHINERA_API_KEY in your server environment to your saved key. Replace audio.mp3 with your file path, then run:

curl https://api.machinera.com/v1/audio/transcriptions \
  -H "Authorization: Bearer $MACHINERA_API_KEY" \
  -F file=@audio.mp3 \
  -F model=transcribe-v1

Read the result

A successful request returns JSON with a text field:

{"text":"Your transcript appears here.","usage":{"type":"duration","seconds":12.345}}

usage is omitted when measured audio duration is unavailable. Use the transcript from the text field directly, or request word timings and quality warnings.

This deployment’s API base URL:

https://api.machinera.com/v1

Authentication and keys

Send your key in the Authorization: Bearer header on every API request.

“Shown once” means the secret cannot be displayed again: only its digest is stored. Afterwards, the list identifies it by label, key ID and masked hint. If you lose a key, create a replacement and revoke the old key by its ID; a lost key stays valid until revoked.

Keys start with ma_, followed by a mode, an underscore and 32 letters or digits. Masked examples: ma_live_0123…STUV (live); ma_test_0123…STUV (test). Keys do not cross environments: use a live-mode key for production and a test-mode key for non-production.

Manage your keys and usage on your Account page.

Production API base URL: https://api.machinera.com/v1. Use this deployment's base URL below with a key created in this environment.

Key lifecycle

You can hold multiple keys and label each for its service or use. Revoke a key by its key ID on the Account page without revoking your other keys. Revocation retains a revoked record; it does not delete the record. Keys do not expire. To rotate a key, create its replacement, move your calls to it, then revoke the old key.

Keys carry no per-key permission scopes: all keys on an account have equivalent access.403 2004: API key is forbidden.

/v1 is server-to-server: OPTIONS carries no CORS headers. Browser clients use the signed-in app path, /transcribe/api/uploads, keeping bearer keys out of browser code.

Transcribe audio: synchronous request

The call blocks until the transcript is ready and returns it in the response — the simplest path for short clips. Multipart uploads go up to 25 MiB by default; a larger body is rejected with 413 1021 (use the asynchronous mode below, or send a URL by reference). For upload advice, see the cap entries linked here and in the asynchronous section below.

Upload a file

Multipart POST /v1/audio/transcriptions with the audio as a file part:

curl https://api.machinera.com/v1/audio/transcriptions \
  -H "Authorization: Bearer $MACHINERA_API_KEY" \
  -F file=@audio.mp3 \
  -F model=transcribe-v1

By reference (URL)

Send a JSON body with a url instead of uploading bytes — same endpoint, Content-Type: application/json (exactly one of file or url):

curl https://api.machinera.com/v1/audio/transcriptions \
  -H "Authorization: Bearer $MACHINERA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/audio.mp3", "model": "transcribe-v1"}'

The urlmust be publicly reachable, including any required access signature. Because the body is just the URL, the multipart upload limit doesn't apply. 5003: url intake is not configured.

Use HTTP in your application

These examples send the same multipart request. Let the client set the multipart boundary.

Python · requests

import os
import requests

with open("audio.mp3", "rb") as audio:
    response = requests.post(
        "https://api.machinera.com/v1/audio/transcriptions",
        headers={"Authorization": "Bearer " + os.environ["MACHINERA_API_KEY"]},
        files={"file": audio},
        data={"model": "transcribe-v1"},
    )
response.raise_for_status()
print(response.json()["text"])

JavaScript · built-in fetch

import { readFile } from "fs/promises";

const form = new FormData();
form.set("file", new Blob([await readFile("audio.mp3")]), "audio.mp3");
form.set("model", "transcribe-v1");
const response = await fetch("https://api.machinera.com/v1/audio/transcriptions", {
  method: "POST",
  headers: { Authorization: `Bearer ${process.env.MACHINERA_API_KEY}` },
  body: form,
});
if (!response.ok) throw new Error(await response.text());
console.log((await response.json()).text);

Send a descriptive User-Agent. Browsers and curl already send one; a raw HTTP client should set one that names your app, e.g. my-app/1.0, so a support request can be traced back to the caller that made it.

Long audio: async jobs

Submit the audio now and get back 202 with a job id and a status of queued; then poll the job (or receive a webhook) until it finishes. Use this mode for long-form audio (full episodes) and whenever you want a durable job that survives a dropped connection. For audio above the async multipart cap below, submit it by reference (a url); a by-reference body carries no audio bytes.

For direct file uploads, follow the Upload grants and recovery contract.

Upload a file

Multipart POST /v1/transcription_jobs with the audio as a file part. Its separate default whole-body cap is 95 MiB, configurable downward per deployment. A body over it receives 413 1022; use the by-reference example below instead. request body exceeds the upload cap; submit a url via POST /v1/transcription_jobs instead. Submit a file like this:

curl https://api.machinera.com/v1/transcription_jobs \
  -H "Authorization: Bearer $MACHINERA_API_KEY" \
  -F file=@audio.mp3 \
  -F model=transcribe-v1

By reference (URL)

Send a JSON body with a url (same URL rules as above: publicly reachable, SSRF-guarded, no multipart size limit) and an optional webhook_url for a best-effort terminal notification:

curl https://api.machinera.com/v1/transcription_jobs \
  -H "Authorization: Bearer $MACHINERA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/audio.mp3", "model": "transcribe-v1", "webhook_url": "https://example.com/webhook"}'

Poll and read the result

Poll the job by its id with GET /v1/transcription_jobs/{id} (replace JOB_ID with the id from the submit response):

curl https://api.machinera.com/v1/transcription_jobs/JOB_ID \
  -H "Authorization: Bearer $MACHINERA_API_KEY"

Keep polling until status is completed (the transcript is then in result) or error (which carries an error object). A completed job's result is always the full verbose_json body: language (the served language), duration, usage, warnings[], and words (per-word timings, always returned), plus inference_seconds (total processing time spent producing this transcript. null when unavailable).timestamp_granularities[]=word is optional and is not required to receive words — whatever the submit asked for. That matches a synchronous response only when the sync request sets response_format=verbose_json; a sync call that omits it returns the default JSON body: {"text":"Your transcript appears here.","usage":{"type":"duration","seconds":12.345}}. usage is omitted when measured audio duration is unavailable.

Webhooks

The webhook_url callback is a best-effort hint: it carriesid and status, with optional warning notices or a numeric error, but no result. Create a webhook signing secret on the Account page and the delivery is signed, so your server can tell a real notification from a forged one; until then it is unsigned and forgeable. Either way, treat it as a nudge to poll, never as proof of job state — always re-fetch GET /v1/transcription_jobs/{id} over the authenticated API for the actual result. The signature headers and a verification recipe are in the API reference.

Public asynchronous submission validates the same options before accepting a job. Poll accepted jobs for terminal status. Completed jobs return full verbose JSON regardless of response_format; extract result.text for text.

Batch

POST /v1/transcription_jobs/batch submits up to 64 by-reference jobs in a JSON jobs array. Give each job its own idempotency_key; replace the example key for each new logical job.

curl https://api.machinera.com/v1/transcription_jobs/batch \
  -H "Authorization: Bearer $MACHINERA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"jobs":[{"url":"https://example.com/audio.mp3","model":"transcribe-v1","idempotency_key":"UNIQUE_KEY_PER_JOB"}]}'

GET /v1/transcription_jobs?ids=JOB_ID,ANOTHER_JOB_ID polls up to 300 known job IDs; it is not a job listing. Replace the placeholders with IDs returned by submission.

curl --get https://api.machinera.com/v1/transcription_jobs \
  -H "Authorization: Bearer $MACHINERA_API_KEY" \
  --data-urlencode "ids=JOB_ID,ANOTHER_JOB_ID"

Both routes return ordered per-item results with each item's status, headers, and body. The call itself returns 200whenever the batch was accepted for processing, including when some items failed, so branch on each item's status_code, never on the outer status. Submissions can partially succeed; inspect each result before retrying. Honor each item's Retry-After header. For successful 200 single-job status responses, Retry-After is present for queued or processing and absent for completed or error. Batch polls also expose a top-level Retry-After when any item carries it.

The two ceilings differ deliberately and neither implies the other: the submit ceiling is sized against the request body those job descriptors add up to and the work one call fans out to, the poll ceiling against the size of the status response it returns. Chunk a larger set against whichever route you are calling.

Request options and response formats

Both modes require model=transcribe-v1. Accepted upload formats: mp3, mp4, mpeg, mpga, m4a, wav, webm, flac, ogg.

Options for synchronous and asynchronous requests
FieldUse
file / urlSend exactly one audio source: a multipart file or a publicly reachable URL in JSON.
modelRequired: transcribe-v1.
response_formatjson (default), verbose_json, or text.
timestamp_granularitiesThe word value is optional; verbose_json always includes word timings.
languageOptional English hint; see the accepted values below.

For synchronous requests, response_format=json returns {"text":"Your transcript appears here.","usage":{"type":"duration","seconds":12.345}}. usage is omitted when measured audio duration is unavailable. Pass response_format=verbose_json for the full result: language (the served language), the audio duration, usage, a warnings[] quality channel, and words (per-word timings, always returned), plus inference_seconds (total processing time spent producing this transcript. null when unavailable).timestamp_granularities[]=word is optional and is not required to receive words. Both sync and async multipart accept the bracketed form and the unbracketed field name. Use response_format=text to get the transcript back as plain text.

Inspect warnings[] in verbose results as a quality channel, even when the request succeeds. A warning is not an HTTP error.

response_format is unsupported

timestamp_granularities is unsupported

language is unsupported (English only; omit 'language' or send 'en')

Errors and retries

Error codes
CodeStatusDescription
1001404No such upload.
1002409Upload is not complete.
1003410Upload has expired; recover any accepted transcription before starting a new upload with a new initialization key.
1004400Uploaded audio does not match the declared file.
1005409Upload is already assigned to a transcription.
3001429Request rate exceeded for this source and tier.
5016503File upload storage is full; contact support before retrying.
5001503File uploads are unavailable on this endpoint.
1006414request path exceeds its byte limit
1007431request headers exceed their aggregate byte limit
1008400provide a non-empty model field
1009400unknown model alias (send a published alias: 'transcribe-v1')
1010400audio url must use http or https
1011400audio url could not be fetched
1012400audio url could not be fetched
1013400webhook url could not be reached
4001503input capacity is busy
1014413request input exceeds its byte limit
1015400audio exceeds the permitted duration for the selected model
1016400measured audio duration exceeds max_duration_seconds
1017400max_duration_seconds must be a finite positive number
4018400request body is shorter than Content-Length
4019408request body did not arrive before its deadline
1018499request was aborted
1019400request body, descriptor or parameter was not accepted
1020400audio is corrupt or unsupported
4002502synchronous processing is unavailable right now; submit the job with POST /v1/transcription_jobs
4003504synchronous processing timed out; processing may have started; contact support before resubmitting
4004502upload transfer failed; retry the request, honoring Retry-After when present
4005503synchronous processing is unavailable right now; submit the job with POST /v1/transcription_jobs
4006503synchronous processing could not start in time; submit the job with POST /v1/transcription_jobs
4007503synchronous processing is unavailable right now; processing may have started; contact support before resubmitting
4008503capacity is temporarily unavailable; retry later
4009503the service is temporarily unable to process this request; please retry
4010503the service is temporarily unavailable; retry later
5002502the request could not be processed; contact support before retrying
1021413request body exceeds the synchronous input cap; submit via POST /v1/transcription_jobs instead
1022413request body exceeds the upload cap; submit a url via POST /v1/transcription_jobs instead
1023411multipart request bodies must carry a Content-Length
1024400batch request contains too many items; split it into smaller batches
3002429batch status response byte limit reached; request fewer job IDs
1025400uploaded media type is unsupported
1026400response_format is unsupported
1027400timestamp_granularities is unsupported
1028400language is unsupported (English only; omit 'language' or send 'en')
1029422The audio appears to be predominantly non-English. Only English audio is supported.
4020400The received bytes do not match X-Content-MD5; verify the checksum and file before retrying.
4021400request body is shorter than Content-Length
3003429this request exceeds the capacity available to your key; submit a shorter clip
4011503capacity is temporarily unavailable; retry later
1030409the original idempotent response is no longer available
1031422this Idempotency-Key was first used with a different request
1032400strong result encryption requires a registered result public key
1033400the registered strong-mode result public key is unusable
1034400the audio could not be read
1035400the job's audio reference expired before it ran; resubmit
5003501url intake is not configured
5004502Failed to transcribe
4012502Failed to transcribe
4013500This service is not configured correctly; contact support before retrying.
5005200the job result is unavailable; resubmit
5006200the job could not be processed due to demand; it may be resubmitted
5007200the job could not start in time; it may be resubmitted
5008200processing was interrupted and no result is available; it may be resubmitted
5009200the job ran out of execution attempts before producing a result; it may be resubmitted
5010200the job was cancelled before running and has no result body; it may be resubmitted
5011200the terminal job result is unreadable; do not resubmit, contact support
1036404no such job
3004429source authentication failure budget exhausted
1037405method not allowed
2003401invalid API key
2004403API key is forbidden
1038404unknown request URL
4014503This service is not configured correctly; contact support before retrying.
4015503This service is not configured correctly; contact support before retrying.
4016503This service is not configured correctly; contact support before retrying.
4017503This service is not configured correctly; contact support before retrying.
3005429request capacity is temporarily exhausted; retry later
5012500Failed to transcribe
5013500Failed to transcribe
1042400The request could not be completed.
1043409The request could not be completed.
5014502Failed to transcribe
5015500Failed to transcribe
Warning codes
CodeDescription
8001Transcription quality for this audio is degraded
8002Transcription quality for this audio is degraded
8003Part of the audio could not be transcribed.
8004Some word timings may be less precise.
8005Speech boundaries may be less precise.
8006Some moments couldn't be heard clearly; they're marked [inaudible] in the transcript.
8007Repeated text that could not be verified was removed.
8008Text that could not be verified against the audio was removed.
8009Automatic punctuation and capitalization could not be applied.
8010The notification could not be delivered; check the transcription status.

Errors use an error object. Required fields are message (a readable explanation), type (the error category), code (a numeric identifier), and doc_url (the code’s one-line meaning). The fields param (the offending input), retryable (retry advice), and details (additional context) are optional.

Rate limits and backoff

Absence of retryable means not specified, not false. Retry only when retryable is true, or when it is absent and the status is 429 or 503 with Retry-After. When retryable is false, do not retry the request unchanged. Wait the Retry-After interval when present; zero means no known timed wait, not immediate recovery. When retryable is true and no wait is specified, use bounded client backoff. Use exponential backoff with jitter and bound the total retry time. For asynchronous submits, reuse the same idempotency key when retrying the same request.

For successful 200 single-job status responses, Retry-After is present for queued or processing and absent for completed or error.

Check Content-Type before parsing: some refusals have a non-JSON body without an error object or request identifier. Record the HTTP status, a short response summary, and the request time when reporting these failures.

Limits

Upload limits
RequestDefault body capIf exceeded
Synchronous upload25 MiBUse an async job or a URL.
Async multipart upload95 MiBUse a URL; deployments may set a lower cap.

Audio duration is constrained by serving capacity and key tier. A terminal 429 3003 : this request exceeds the capacity available to your key; submit a shorter clip. Submitting a URL avoids uploading audio bytes in the request body, but does not bypass duration or capacity checks.

Upload grants and recovery

Initialize with POST /v1/uploads, send the exact file bytes and required headers to put_url, then submit upload_id to POST /v1/transcription_jobs. Read the returned limits. The PUT capability expires at expires_at; plan to finish the transfer before it. Before starting or retrying a transfer after grant expiry, replay with the same initialization key and file descriptor for a fresh grant while the file is incomplete and before upload_deadline. Do not interrupt an active PUT.

Submit immediately after PUT success. submit_expires_at is the inclusive first server observation time of the complete file plus limits.submit_grace_seconds, capped by the absolute retention bound upload_deadline. It is null until that observation. Submit itself records the observation; initialization replay can report the deadline without a write grant. Later inspections and retries never extend either deadline.

Code 1002 means finish or retry the PUT, then submit again. A lost PUT response may cause 412 on retry; submit to confirm. Code 1003 means the grace or absolute deadline passed. First recover any ambiguous submission using the same upload ID, submission key and options; an accepted replay returns the same job. Only after definite non-acceptance, use a new initialization key, re-upload and submit with a new submission key. If recovery history has expired, stop automatic replacement and ask the caller how to proceed.

A pending initialization response without a PUT URL can mean submission recovery is in progress; replay the original submission before requesting another grant. Retryable API refusals carry Retry-After; zero allows client backoff without a timed wait and does not promise immediate recovery. A storage-capacity refusal directs you to support.

Rate limits and authentication budget

Upload initialization is limited by tier and source request rate. Its rate headers report the active allowance; Retry-After reports seconds until reset. Processing limits also apply to submitted jobs. Follow the Rate limits and backoff guidance.

Repeated authentication failures share a source budget, including callers behind the same source address. Once exhausted, even valid keys receive 429 3004 until the next UTC minute. The x-ratelimit-limit-requests, x-ratelimit-remaining-requests and x-ratelimit-reset-requests headers describe this failure budget, not a tenant request allowance. Reset is seconds remaining, not an epoch timestamp. Read the limit header for the failure allowance and honor Retry-After before retrying. Correct invalid credentials before sending more requests.

Audio URL policy

Audio URLs must use http or https and resolve only to publicly routable addresses. Private, loopback and link-local addresses are refused. Redirects are not followed; provide the final audio URL. Check public reachability and the audio format before resubmitting.

Idempotency

Use an Idempotency-Key to make asynchronous submits safe to retry. Generate a unique key per logical submit and reuse it only when retrying that exact request: while retained, the repeat replays the original job instead of re-running or double-billing it. 409 1030: the original idempotent response is no longer available. Send it on the async submit route as a header or an idempotency_key body field. Public synchronous requests have no idempotency replay guarantee; a retry may run again.

Each key is bound to a request fingerprint: the audio you submitted plus every option that shapes the transcript — model, response_format, language, and the rest. For URL input, the fingerprint uses the reference, not newly fetched bytes. Use an immutable URL or content version when deriving a key. Fields that do not change what is transcribed (webhook_url, prompt) are outside it, and omitting an option is the same request as sending its default. Reuse a key with a different fingerprint — one static key across different files, say — and the stored response answers the first request. Use a new key for a changed request. To deduplicate identical submissions rather than just network retries, derive the key from the content plus every request-affecting option — e.g. a sha256 of the file bytes with model, response_format, and language. Keys are scoped to your account, so such a content hash can never collide into another customer's result.

Result and idempotency retention

Save results before they expire. Result access and idempotency protection are time-limited; retaining a key does not guarantee that the original response remains retrievable. A retained key whose response is unavailable returns 409 1030. Submitting with a new key may incur a new charge. After retention expires, do not assume deduplication.

OpenAI-compatible clients

Optional: if you already use an OpenAI-style client, the synchronous endpoint accepts the OpenAI transcription request shape. Point the client at this deployment’s base URL and use your API key. The stock OpenAI SDK covers synchronous file transcription. Use raw HTTP for asynchronous jobs, by-URL input and batch routes.

import os
from openai import OpenAI

client = OpenAI(api_key=os.environ["MACHINERA_API_KEY"], base_url="https://api.machinera.com/v1")
with open("clip.mp3", "rb") as audio:
    transcript = client.audio.transcriptions.create(
        model="transcribe-v1",
        file=audio,
        response_format="verbose_json",
        timestamp_granularities=["word"],
    )
print(transcript.text)

Support and contact

Include the x-request-idresponse header value and the affected calls' timestamps when you report an API problem. The x-request-id shape is ^[A-Za-z0-9_.:-]{1,128}$.

Production volume and commitments

If you are evaluating production volume or need commitments, contact sales.