Arabic speech-to-text API guide

Transcribe Arabic dialects, get a timestamp for every word, and download subtitles. Follow these steps to make your first API request.

API base URL: https://kalemio.app. Use a live API key from your account.

Download OpenAPI JSON View the raw document

Quickstart: from a recording to JSON and SRT

This script submits one recording, waits for the job with growing pauses between checks, and saves its JSON result and SRT beside the file. You need:

  • A Kalemio account. The Free plan works: 60 minutes a month and no card, with slower, queued processing, one job at a time.
  • An API key with two permissions: Submit transcriptions (transcriptions:write) and Read transcripts (transcriptions:read). Read usage isn’t needed. Create the key in the API console.
  • Any short Arabic recording you have the right to process: MP3, M4A, PCM WAV, MP4 or MOV, up to 20 MiB and 10 minutes.
"""Transcribe an Arabic recording with Kalemio; save JSON and SRT beside it.

    pip install requests
    export KALEMIO_API_KEY='YOUR_API_KEY'
    python kalemio_quickstart.py conversation.m4a
"""
import hashlib
import os
import sys
import time
from pathlib import Path

import requests

API = os.environ.get("KALEMIO_API_BASE", "https://kalemio.app")
AUTH = {"Authorization": "Bearer " + os.environ["KALEMIO_API_KEY"]}
recording = Path(sys.argv[1])
audio = recording.read_bytes()
# The same file always sends the same key, so running this again after an
# interruption returns the first job instead of starting and billing another.
# A second argument, such as retry1, makes a new key to transcribe it afresh.
key = "quickstart-" + hashlib.sha256(audio).hexdigest()[:32]
if len(sys.argv) > 2:
    key += "-" + sys.argv[2]


def message(response):
    try:
        return response.json().get("error", response.text)
    except ValueError:
        return response.text


def stop(what, response):
    sys.exit(f"{what} ({response.status_code}): {message(response)}")


def retryable(response):
    if response.status_code == 429:  # The Free plan's own limits don't clear soon.
        return not response.json().get("code", "").startswith("free_")
    return response.status_code >= 500


def send(method, path, headers=None, **options):
    """Sends a request; rate limits and server errors are retried with backoff."""
    headers = {**AUTH, **(headers or {})}
    for attempt in range(6):
        response = requests.request(method, API + path, headers=headers,
                                    timeout=120, **options)
        if not retryable(response):
            return response
        time.sleep(float(response.headers.get("Retry-After", 2 ** attempt)))
    return response


response = send("POST", "/v1/transcriptions",
                headers={"Idempotency-Key": key},
                files={"file": (recording.name, audio)},
                data={"speakers": "auto"})
if response.status_code == 402:
    sys.exit(f"Not enough minutes: {message(response)} Nothing was charged. Add "
             f"a plan or minutes at {API}/developers/portal/ and run this again.")
if response.status_code != 202:
    stop("Upload refused", response)
job = response.json()
print(f"Job {job['id']} accepted ({job['processing']} processing)")

delay = 3
while job["status"] in ("queued", "running"):
    time.sleep(delay)
    delay = min(delay * 2, 30)  # Free plan jobs can wait in a queue for a while.
    response = send("GET", f"/v1/transcriptions/{job['id']}")
    if not response.ok:
        stop("Status check failed", response)
    job = response.json()
if job["status"] != "completed":
    sys.exit(f"Job {job['id']} {job['status']}: {job.get('error', '')}")

for name, suffix in (("result", ".json"), ("srt", ".srt")):
    response = send("GET", f"/v1/transcriptions/{job['id']}/{name}")
    if not response.ok:
        stop("Download failed", response)
    recording.with_suffix(suffix).write_bytes(response.content)
    print("Saved", recording.with_suffix(suffix))

Save the script, set KALEMIO_API_KEY, and pass the recording’s path. It prints the job ID and saves two files next to the recording:

conversation.json
The body of GET /v1/transcriptions/{id}/result: text, every word with start, end and speaker_id, and duration_seconds.
conversation.srt
The body of GET /v1/transcriptions/{id}/srt: cues of up to eight words or five seconds, one speaker each, labelled such as [speaker_0].

Delays, minutes and retries

  • Free plan queue. Free jobs wait in their own queue and are typically several times slower than paid ones, so the script keeps checking, every 30 seconds at most. A free job that cannot start within a few hours fails without using your minutes. Free accounts run one job at a time; a second upload meanwhile returns 429 with free_concurrency_limit.
  • Interrupted runs. The Idempotency-Key comes from the file’s contents, so running the script again on the same file follows the same job instead of starting, and paying for, another. To transcribe it afresh, after a failure or once its result has expired, add a word as a second argument, such as retry1.
  • Not enough minutes. The upload returns 402 before a job exists, so nothing is charged: free_quota_exhausted on the Free plan, or a request to add minutes on a prepaid plan. Add a plan or minutes in the API console, then run the script again. The OpenAI-compatible endpoint reports the same condition as 429 with insufficient_quota.
  • Failed jobs end with "status": "failed" and an error. Prepaid minutes are returned, and a failed free job doesn’t count.

Resuming on the OpenAI-compatible endpoint

The compatible endpoint waits up to 100 seconds per request. A job still running then returns 504 with transcription_pending and its ID in the X-Kalemio-Transcription-Id header, after the OpenAI libraries have retried the identical request twice. The job carries on and is charged once. Follow it by its ID with the native status and result endpoints, which also needs the Read transcripts permission, or send the identical request again with the same Idempotency-Key. Uploading with a new key or a changed file starts, and bills, a second job.

Python
import os
import sys
import time

import openai
import requests

API = "https://kalemio.app"
KEY = os.environ["KALEMIO_API_KEY"]
client = openai.OpenAI(base_url=API + "/v1", api_key=KEY)
try:
    with open("conversation.m4a", "rb") as audio:
        transcript = client.audio.transcriptions.create(
            model="kalemio-arabic-1", file=audio,
            extra_headers={"Idempotency-Key": "openai-conversation-001"})
    print(transcript.text)
except openai.APIStatusError as error:
    if error.code != "transcription_pending":
        raise
    # Still running: follow the same job by its ID. Uploading again with a new
    # Idempotency-Key would start, and bill, a second job.
    job_id = error.response.headers["x-kalemio-transcription-id"]
    job = f"{API}/v1/transcriptions/{job_id}"
    auth = {"Authorization": "Bearer " + KEY}
    status = "queued"
    while status in ("queued", "running"):
        time.sleep(15)
        status = requests.get(job, headers=auth).json()["status"]
    if status != "completed":
        sys.exit(f"Transcription {job_id} {status}")
    print(requests.get(job + "/result", headers=auth).json()["text"])

The native result is /result JSON or /srt, not the compatible endpoint’s response formats. Results stay available for 24 hours after submission.

Authentication

Create a key in the API console, choose only the permissions that integration needs, then send it in the Authorization header. Use environment variables on your server.

Set these variables in your terminal before running the examples. Replace the key placeholder with your live key; keep it out of source control, browser code and shared shell history.

Shell
export KALEMIO_API_BASE='https://kalemio.app'
export KALEMIO_API_KEY='YOUR_LIVE_API_KEY'

curl "$KALEMIO_API_BASE/v1/account" \
  -H "Authorization: Bearer $KALEMIO_API_KEY"

Keys are shown once and cannot be recovered. Rotate a lost or exposed key in the portal; rotation revokes the old key immediately. You can see when a key was last used in the console.

Verify your key with GET /v1/account. It returns your mode, available seconds, and whether transcription is enabled. Invalid or revoked keys return 401; a key missing the required permission returns 403.

Upload audio or video

Submit one audio or video file as multipart form data. You can send video directly; we extract the audio for you. MP3, M4A, PCM WAV, MP4 and MOV files are supported up to 20 MiB and 10 minutes; for anything longer, see long recordings. The language is Arabic, including dialectal speech. Choose automatic speaker detection or set the number of speakers if you know it.

Shell
curl "$KALEMIO_API_BASE/v1/transcriptions" \
  -H "Authorization: Bearer $KALEMIO_API_KEY" \
  -H "Idempotency-Key: recording-2026-001" \
  -F 'file=@conversation.m4a' \
  -F 'language=ar' \
  -F 'speakers=auto'

A successful upload returns 202 Accepted with the job, not the transcript. Store the ID to retrieve the result later. Before running the polling example, set export JOB_ID='tr_example', replacing tr_example with the returned ID. Use a new Idempotency-Key for each different recording.

JSON
{ "id": "tr_5d0c3f6a8e2b4c71a9f0e6d2b8c4a1f3",
  "status": "queued", "duration_seconds": 1.92,
  "billed_seconds": 0, "billed_minutes": 0,
  "reserved_seconds": 2, "billing_method": "credits",
  "processing": "priority",
  "expires_at": "2026-10-01T09:30:00.000Z" }
FieldValue
fileOne MP3, M4A, PCM WAV, MP4 or MOV file
languagear (default)
speakersauto (default), or 1–20

Check progress

Shell
curl "$KALEMIO_API_BASE/v1/transcriptions/$JOB_ID" \
  -H "Authorization: Bearer $KALEMIO_API_KEY"

Poll every 3–5 seconds. Jobs move from queued to running, then completed or failed.

Download results

Once complete, retrieve JSON with GET /v1/transcriptions/{id}/result or subtitles with GET /v1/transcriptions/{id}/srt. Authenticate both requests.

JSON
{ "text": "مرحبا", "words": [
  { "text": "مرحبا", "start": 1.24,
    "end": 1.78, "speaker_id": "speaker_0" }
], "duration_seconds": 2.1 }

Work with word-level timestamps

Each word has a start and end time in seconds from the beginning of the recording. In the example above, مرحبا starts at 1.24 seconds and ends at 1.78 seconds.

text
The transcribed word, with the dialect wording preserved.
start / end
The word’s time range in seconds. Use it to sync captions or seek to a passage.
speaker_id
A label for the voice within this recording, such as speaker_0. It does not identify the same person across recordings.

For word highlighting, show the word while playback is between its start and end times. For custom subtitles, group adjacent words into readable lines and use the first word’s start and last word’s end. Use the SRT download if you want ready-made subtitle cues.

These examples illustrate the response format. The published benchmark measures word error rate on Levantine Arabic; it does not measure word timing.

Long recordings

Each upload is limited to 20 MiB and 10 minutes. To transcribe a longer recording today, split it into consecutive audio pieces of under 10 minutes, submit each piece as its own job, then join the results.

  1. Split the audio. Keeping only the audio in mono also keeps each piece well under 20 MiB. For example, with FFmpeg:
    Shell
    ffmpeg -i interview.mp4 -vn -ac 1 -c:a aac -b:a 96k \
      -f segment -segment_time 590 -reset_timestamps 1 \
      -segment_list pieces.csv part_%03d.m4a
  2. Submit each piece. Give every piece its own Idempotency-Key, such as interview-part-000, so a retry never creates a second job. Paid plans run two jobs at a time; send the next piece when one finishes.
  3. Join the results. Word times start at zero in each piece. Add the piece’s start time (pieces.csv lists each piece’s start time in seconds) to every start and end, then concatenate the words in order.

Two things change at the joins. A word spoken exactly at a cut can be split or missed, so cut at a pause if your tool allows it. Speaker labels such as speaker_0 are assigned per job, so the same voice can carry a different label in the next piece. Each piece is billed as its own job. If you only need a subtitle file, Arabic Subtitles accepts recordings of up to six hours without splitting.

OpenAI-compatible API

Code written for OpenAI’s speech-to-text endpoint can call Kalemio by changing two settings: the base URL becomes https://kalemio.app/v1 and the key becomes your Kalemio API key. POST /v1/audio/transcriptions takes the same multipart fields and returns the transcript in the same response, so the official openai libraries for Python and JavaScript, LiteLLM and OpenAI’s own curl examples work with only the model name changed. The key needs the transcriptions:write permission.

Set model to kalemio-arabic-1. OpenAI model names such as whisper-1 and gpt-4o-transcribe aren’t accepted: they return 404 with "code": "model_not_found". GET /v1/models lists the valid ids and needs no key.

Shell
curl "$KALEMIO_API_BASE/v1/audio/transcriptions" \
  -H "Authorization: Bearer $KALEMIO_API_KEY" \
  -F file=@conversation.m4a \
  -F model=kalemio-arabic-1 \
  -F response_format=verbose_json \
  -F 'timestamp_granularities[]=word'
Python
import os
from openai import OpenAI

client = OpenAI(base_url="https://kalemio.app/v1",
                api_key=os.environ["KALEMIO_API_KEY"])
with open("conversation.m4a", "rb") as audio:
    srt = client.audio.transcriptions.create(
        model="kalemio-arabic-1", file=audio, response_format="srt")
print(srt)
JavaScript
import fs from "node:fs";
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://kalemio.app/v1",
  apiKey: process.env.KALEMIO_API_KEY });
const transcript = await client.audio.transcriptions.create({
  model: "kalemio-arabic-1",
  file: fs.createReadStream("conversation.m4a"),
});
console.log(transcript.text);
LiteLLM
import os
import litellm

with open("conversation.m4a", "rb") as audio:
    transcript = litellm.transcription(
        model="openai/kalemio-arabic-1", file=audio,
        api_base="https://kalemio.app/v1",
        api_key=os.environ["KALEMIO_API_KEY"])
print(transcript.text)

Response formats

response_formatWhat you get
json (default)text, and usage with the recording’s duration in seconds
textThe transcript as plain text
srt, vttSubtitle cues of up to eight words or five seconds, one speaker per cue
verbose_jsonlanguage ("arabic"), duration, text and segments; add timestamp_granularities[]=word for words with start and end times
diarized_jsonSegments labelled with speakers A, B, C… in order of first appearance

Differences from OpenAI’s API

  • Arabic only. Omit language or set it to ar. Dialects need no hint, and codes such as ar-EG or en are refused.
  • Files. MP3, M4A, PCM WAV, MP4 or MOV, up to 20 MiB and 10 minutes, the same limits as the native API. FLAC, OGG and WebM aren’t accepted. Larger files return 413 with file_too_large; longer ones return 400 with audio_too_long.
  • Refused, not ignored. A non-empty prompt, keywords[], include[]=logprobs, known_speaker_names[], known_speaker_references[], a temperature other than 0, a chunking_strategy other than auto and unknown fields each return 400 naming the parameter.
  • Placeholder fields. verbose_json segments include seek, tokens, temperature, avg_logprob, compression_ratio and no_speech_prob so OpenAI-typed clients can read them, but they are fixed values (0 or []), not model scores.
  • Streaming. stream=true works with json, text and diarized_json, but the events arrive together once the whole file is transcribed. They aren’t progressive.
  • Speakers. An extra speakers field (auto or 1–20) sets the speaker count, as in the native API. With the Python library, pass extra_body={"speakers": "2"}.
  • No translation. /v1/audio/translations returns 404.
  • Billing. Requests use your Kalemio plan, minutes and limits exactly like POST /v1/transcriptions; see plans and billing. usage.seconds is the recording’s duration rounded up.

Waiting, retries and errors

Each request waits up to 100 seconds for its job. Free plan jobs wait in a slower queue and often take longer. A job still processing at that point returns 504 with "code": "transcription_pending", its ID in error.transcription_id and in the X-Kalemio-Transcription-Id header, and a Retry-After header. Send the same request again to keep waiting; the OpenAI libraries retry it automatically, twice by default. The recording isn’t charged twice: with an Idempotency-Key header the key identifies the job, and without one an identical upload (same file and speakers setting) that is still processing or finished within the last hour is reused. You can also poll GET /v1/transcriptions/{id} and download /result or /srt as described above.

Errors use OpenAI’s shape, {"error": {"message", "type", "param", "code"}}. Running out of minutes returns 429 with "code": "insufficient_quota" (the native API returns 402). Rate limits return 429 with rate_limit_exceeded, and a failed job returns 500 with transcription_failed; prepaid minutes are returned. Responses that retrying can’t fix carry x-should-retry: false, so the OpenAI libraries don’t repeat them.

Plans and usage-based billing

An account pays in one of two ways. Plans are prepaid, monthly or yearly: each includes a set number of minutes every month, such as 1,250 minutes a month for US$5.00 on Starter, and one-time top-ups add more. There is no overage charge; when the minutes run out, uploads return 402 until you add more. Usage-based billing is an optional alternative with no included minutes and no allowance to run out: each successful job costs US$0.015 per processed audio minute, rounded up to whole minutes, with a minimum of 2 minutes (US$0.03) per job, invoiced in USD at the end of each month. Per minute, usage-based billing costs more than every monthly plan: 1,250 minutes would cost US$18.75 with it, against US$5.00 on Starter. Choose it only when a workload must never stop at an allowance.

Prepaid minutes are reserved when a job is accepted, measured from the uploaded media duration rounded up to the next second, and returned if the job fails. GET /v1/account reports the method in billing_method: credits for a monthly plan, metered_api_v1 for usage-based billing, or free. Compare the plans in the API console.

Free plan and processing speed

Accounts without a paid plan are on the Free plan: 60 audio minutes per calendar month (UTC), with no card required. Each job counts its duration in whole minutes, rounded up, with a minimum of one minute. Failed jobs don’t count.

The Free plan is an evaluation queue, not a service level. Free jobs wait in their own best-effort queue: they are typically several times slower than paid jobs, a free account runs one job at a time, and a free job that cannot start within a few hours fails without counting against your minutes. Use it to check quality and the response format, not to measure speed or reliability.

Paid plans do not use the free queue. Paid jobs are admitted and processed separately with priority, and free jobs never hold them up, however busy the free queue is.

Every job reports how it is processed in its processing field: "standard" for Free plan jobs and "priority" for paid ones. GET /v1/account returns free_plan with the minutes used and remaining this month and when they reset, or null on a paid plan.

JSON
{ "id": "tr_example", "status": "queued",
  "processing": "standard", "billing_method": "free" }

When this month’s free minutes run out, uploads return 402 with "code": "free_quota_exhausted". Upgrade in the API console to keep going with priority processing.

Errors and limits

StatusWhat to do
401Check your API key.
402Add minutes before submitting another file, unless the account uses metered billing. On the Free plan, free_quota_exhausted means this month’s free minutes are used up; upgrade to continue.
404Check the job ID and the account that created it.
409The result is not ready. Check job status.
413 / 422Check file size and request fields.
429Wait before retrying; respect Retry-After. Free plan codes: free_concurrency_limit (one job at a time), free_rate_limited (too many free uploads) and free_queue_full (the free queue is busy).
503The API is temporarily unavailable.

Uploads are limited to 20 MiB and 10 minutes, with two active jobs per account on paid plans and one on the Free plan. Prepaid credit accounts are billed by uploaded media duration rounded up to the next second, reserve minutes on acceptance, and receive them back if a job fails. Usage-based accounts are billed after each successful job at US$0.015 per processed audio minute, rounded up to whole minutes, with a minimum of 2 minutes (US$0.03) per job; see plans and billing. No automatic overages on prepaid accounts. Results are available for 24 hours from submission. Reuse an Idempotency-Key only for the same file and speaker settings; it remains associated with that job. Test keys do not process recordings.

Billing and data handling

Plans and minutes

Plans are billed monthly or yearly in USD or CAD. Included minutes are granted every month, on annual plans too, and do not roll over. Top-up minutes never expire and are used after monthly minutes. Every charge is at least US$5 (CA$6.95). Upgrades keep the billing interval and charge and grant minutes proportionally for the rest of the billing period after successful payment; an upgrade whose prorated charge would be under the minimum is not offered. The renewal date stays the same. Cancel from the API console before renewal to stop the next charge.

Failed jobs and refunds

Failed transcription jobs return minutes to their original allowance; expired monthly minutes remain expired; a network interruption during submission can also cause a job to fail without a charge. Refunded or disputed payments remove the corresponding credits and may leave a negative balance if those credits have already been used. Contact Support email for purchase or refund help.

Your recordings and results

Audio is processed on our transcription infrastructure and hosting providers. Uploaded media is not used to train models. Working audio and upstream job files are deleted after processing; downloadable results expire 24 hours after submission. Cleanup retries during service outages. Account, payment and usage metadata remain for service/accounting. Only upload files you have permission to process. Contact support to request account deletion.