Arabic dialects.
Every word, timed.

An Arabic speech-to-text API that doesn’t rewrite the dialect into formal Arabic. Get word-level timestamps, speaker labels, JSON, and SRT for captions and searchable transcripts.

Starting at 1,250 minutes for USD 5.00 /month. Word timestamps, speaker labels, JSON, and SRT included.

Request
export KALEMIO_API_BASE='https://kalemio.app'
export KALEMIO_API_KEY='YOUR_API_KEY'

curl "$KALEMIO_API_BASE/v1/transcriptions" \
  -H "Authorization: Bearer $KALEMIO_API_KEY" \
  -H "Idempotency-Key: conversation-001" \
  -F 'file=@conversation.m4a' \
  -F 'speakers=auto'

Response202 Accepted

{ "id": "tr_5d0c3f6a8e2b4c71a9f0e6d2b8c4a1f3",
  "status": "queued", "duration_seconds": 1.92,
  "billed_seconds": 0, "billed_minutes": 0,
  "reserved_seconds": 2, "billing_method": "credits",
  "processing": "priority",
  "expires_at": "2026-10-01T09:30:00.000Z" }

Asynchronous: the job ID comes back at once. Poll GET /v1/transcriptions/{id} until "status" is "completed", then download /result as JSON or /srt. Complete quickstart script

Completed result

A little Levantine Arabic

أهلاً، كيف فيني ساعدك؟

“Hello, how can I help you?”

  • أهلاً،0.24–0.68
  • كيف0.72–0.94
  • فيني0.98–1.22
  • ساعدك؟1.26–1.78
{ "text": "أهلاً، كيف فيني ساعدك؟",
  "words": [
    {"text":"أهلاً،","start":0.24,"end":0.68,"speaker_id":"speaker_0"},
    {"text":"كيف","start":0.72,"end":0.94,"speaker_id":"speaker_0"},
    {"text":"فيني","start":0.98,"end":1.22,"speaker_id":"speaker_0"},
    {"text":"ساعدك؟","start":1.26,"end":1.78,"speaker_id":"speaker_0"}
  ],
  "duration_seconds": 1.92 }
1
00:00:00,240 --> 00:00:01,780
[speaker_0] أهلاً، كيف فيني ساعدك؟
Illustrative transcript. Times are in seconds.

The words matter. So does the dialect.

Interviews, podcasts and everyday conversations don’t sound like a textbook. Kalemio is built for Modern Standard Arabic and spoken dialects, keeping the words as they were said.

In our September 2026 benchmark on 2.06 hours of Levantine Arabic, Kalemio had the lowest dialect-tolerant word error rate of the 6 systems tested: 4.6%. Every system got the same audio and the same scoring.

Explore word-level timestamps

Fewer transcription errors.

Dialect-tolerant word error rateLower is better

  1. Kalemio4.6%
  2. Meta Muse Voice 1.06.3%
  3. Hamsa General V27.4%
  4. ElevenLabs Scribe v27.4%
  5. Groq Whisper large-v317.0%
  6. Groq Whisper turbo25.1%

Kalemio benchmark · September 2026 · 2.06 hours of Levantine audio. Same audio and scoring for every model, with spelling and equivalent word forms normalized. Results vary by recording.

PlanMinutesPrice
FreeSlower, queued processing60 per monthUSD 0.00 / monthStart free
Starter1,250 per monthUSD 5.00 / monthChoose Starter
StarterBilled yearly1,250 per monthUSD 50.00 / yearChoose Starter annual
Studio5,000 per monthUSD 13.93 / monthChoose Studio
StudioBilled yearly5,000 per monthUSD 159.00 / yearChoose Studio annual
Top-up500 extra minutesUSD 5.00 one timeFor subscribers. These minutes never expire, even if you cancel your plan.

Free plan: 60 minutes a month, no card needed, with slower, queued processing on a best-effort basis. Jobs are typically several times slower than on paid plans and may wait in a queue. Meant for evaluation, one job at a time. Paid jobs are processed separately with priority and never wait behind the free queue.

Plans renew monthly or yearly until cancelled. Minutes reset every month, on annual plans too, and don’t roll over. Upgrade or cancel from your account.

Prefer no monthly allowance? Optional usage-based billing costs US$0.015 per minute, more than any plan, with a 2-minute minimum (US$0.03) per job. How plans and usage-based billing compare

Send the audio. Or the video.

Formats
MP3, M4A, PCM WAV, MP4 and MOV. We extract the audio from video.
File size
Up to 20 MiB and 10 minutes per file.
Jobs
Two active jobs per account on paid plans; one on the Free plan.
OpenAI SDKs
An OpenAI-compatible transcription endpoint: set the base URL to https://kalemio.app/v1 and the model to kalemio-arabic-1. It answers in the same request, waiting up to 100 seconds; a job still running then returns 504 transcription_pending with its job ID to resume.

Your recordings stay yours.

Training
We don’t use your recordings to train models.
Audio
Working audio is deleted after processing.
Results
Yours to download for 24 hours after submission.

Just need a subtitle file?

Arabic Subtitles is a separate no-code product for creating SRT files. These guides cover subtitle workflows; the Kalemio API remains the product for software integrations.

Generate an Arabic SRT without writing code