Arabic Transcription API with Word Timestamps and Speaker Labels

Kalemio turns Arabic audio or video into transcript text, word-level timestamps, speaker labels and SRT subtitles. It is built for applications that need to keep spoken Arabic wording, dialect included, while linking the transcript back to the recording.

Use the JSON result for a searchable recording or a custom playback interface. Use SRT when the next step needs timed subtitle cues. If you only need a subtitle file, not an integration, the separate Arabic Subtitles tool makes one without code.

You can evaluate the API on the Free plan: 60 audio minutes a month with no card, in a slower queue, one job at a time. Paid plans get priority processing; compare them in the pricing table before you choose a production setup.

Read the API guide Get an API key

What to test in an Arabic speech-to-text API

Start with recordings that resemble your application’s real inputs: the dialects, microphones, speaker counts and vocabulary you expect. A clean single-speaker demo can’t tell you how a system handles a noisy discussion.

Check transcription and timing separately. A transcript can contain the right word while placing its boundary too early for word highlighting, and a plausible-looking transcript can still get a name or a number wrong.

Keep a short reviewed reference for each test recording and note the corrections you make. Look at dialect wording, English terms inside Arabic speech, speaker changes and overlapping voices. Decide which errors your application can tolerate, and which need a review step before people rely on the result.

What Kalemio returns

A completed job’s JSON result has the transcript text, the recording’s duration and a list of words. Each word has its text, its start and end time in seconds, and a speaker label such as speaker_0.

You can use these fields to build:

  • Search results that seek to the right moment in a recording
  • Word highlighting synchronized with playback
  • Transcript views grouped by speaker label
  • Custom subtitle layouts with your own line-breaking rules

Your application supplies the search, the player and the interface. A speaker label tells voices apart within one recording; it doesn’t name a person or show that the same voice appears in another file. The SRT download groups words into cues of up to eight words or five seconds, one speaker each. The API guide describes every field.

Make a first request with the asynchronous API

The native API works in three steps: upload the recording, check the job, then download the result. The upload answers 202 Accepted with a job ID, not the finished transcript.

Keep the API key on your server, in an environment variable or a secret store, and out of browser code and source control. These commands expect it in KALEMIO_API_KEY. It needs two permissions: Submit transcriptions and Read transcripts.

1. Upload one recording

Use a short Arabic recording you have permission to process. This example sends an M4A file and asks for automatic speaker detection:

Shell
curl 'https://kalemio.app/v1/transcriptions' \
  -H "Authorization: Bearer $KALEMIO_API_KEY" \
  -H 'Idempotency-Key: evaluation-recording-001' \
  -F 'file=@conversation.m4a' \
  -F 'language=ar' \
  -F 'speakers=auto'

A successful upload returns HTTP 202 and the job. The values below are illustrative:

JSON
{ "id": "tr_5d0c3f6a8e2b4c71a9f0e6d2b8c4a1f3",
  "status": "queued", "duration_seconds": 1.92,
  "billed_seconds": 0, "billed_minutes": 0,
  "reserved_seconds": 2, "billing_method": "credits",
  "processing": "priority",
  "expires_at": "2026-10-01T09:30:00.000Z" }

Save the id. Give each recording its own Idempotency-Key, 8 to 128 letters, digits, dots, colons, underscores or hyphens. Sending the same key again returns the same job rather than starting, and billing, another; use a new key for a different recording or speaker setting.

2. Check the job

Set JOB_ID to the ID the upload returned, then request the job:

Shell
JOB_ID='tr_your_job_id'  # the "id" from the upload response
curl "https://kalemio.app/v1/transcriptions/$JOB_ID" \
  -H "Authorization: Bearer $KALEMIO_API_KEY"

A job is queued, running, completed or failed. Check every few seconds, and less often while a job waits in the queue. A queued Free plan job has been accepted but can wait a while before it starts; don’t upload the recording again while it waits.

3. Download the results

Once the job is completed, download the JSON and the SRT:

Shell
curl --fail "https://kalemio.app/v1/transcriptions/$JOB_ID/result" \
  -H "Authorization: Bearer $KALEMIO_API_KEY" \
  -o transcript.json
curl --fail "https://kalemio.app/v1/transcriptions/$JOB_ID/srt" \
  -H "Authorization: Bearer $KALEMIO_API_KEY" \
  -o subtitles.srt

With --fail, curl stops on an error instead of saving the error message as the file. In production, check the HTTP status of every response and handle failed jobs, an exhausted allowance and rate limits explicitly, including any Retry-After header. The quickstart has a complete script in Python, Bash and Node.js that runs all three steps and saves both files.

Works with the OpenAI SDKs

Kalemio also has an OpenAI-compatible endpoint, POST /v1/audio/transcriptions, which returns the transcript in the same request. Code written for OpenAI’s speech-to-text API needs only the base URL https://kalemio.app/v1, a Kalemio key and the model kalemio-arabic-1. It differs from the native steps above:

  • Each request waits up to 100 seconds. A job still running then returns 504 with transcription_pending and its job ID; follow that job by its ID instead of uploading again.
  • Only kalemio-arabic-1 is accepted. OpenAI model names such as whisper-1 return 404.
  • It transcribes Arabic only, and refuses options it doesn’t support, such as a prompt, instead of ignoring them.
  • It answers as JSON, text, SRT, WebVTT, verbose JSON with word and segment timestamps, or diarized JSON with speakers labelled A, B and so on.
  • The same file limits, plans and billing apply as on the native API.

The API guide lists every difference, with Python, JavaScript and LiteLLM examples.

Plan for file limits and longer recordings

Both endpoints accept MP3, M4A, PCM WAV, MP4 and MOV files of up to 20 MiB and 10 minutes each; Kalemio extracts the audio from video. A paid account can have two jobs queued or running at once, and a Free plan account one.

For longer recordings, the API guide shows how to split the media, submit each piece as its own job, and join the results by adding each piece’s start time to its timestamps. Keep the original timeline, check the joins, and don’t assume a speaker label in one piece refers to the same person in the next.

Save results and review data handling

Download what you need promptly: results are available for 24 hours after submission. Decide where your application stores transcripts, and who can read them, before you accept customer recordings.

Kalemio doesn’t use submitted recordings to train models, and working audio is deleted after processing. The API privacy policy covers service providers, retention and cleanup. Your application still needs its own policy for the media and results it keeps.

Evaluate with a complete workflow

Choose a few representative clips, submit them, save both formats and note the corrections each one needs. Test how your application behaves when a job waits, fails or reaches an allowance limit. Measure processing time on the queue you plan to use: a Free plan evaluation doesn’t show paid production speed.

When you’re ready, follow the API guide to build your first integration. Start with a recording you know well, and judge the transcript, timing and speaker labels against what your application needs.

Read the API guide Get an API key

Need a subtitle file rather than an integration? Arabic Subtitles is a separate no-code product, with guides for dialect subtitles, YouTube and Arabic SRT files.