Sarj Speech APIs provide production-ready speech services for Arabic and multilingual applications.
The platform currently offers:
- Speech-to-Text (STT) — Convert audio into accurate text with low latency.
- Text-to-Speech (TTS) — Generate natural, expressive speech from text, including support for voice cloning.
These APIs are designed for conversational AI, voice assistants, telephony, accessibility, media generation, and other speech-enabled applications.
Authentication
All Speech API requests require an API key.
Include your API key in the Authorization header of every request.
Authorization: Bearer YOUR_API_KEY
Base URLs
Speech-to-Text
https://stt-rnnt-ar.sarj.ai
Text-to-Speech
https://sarj-omni-tts.sarj.ai
Available APIs
Speech-to-Text
Endpoint:
POST /openai/v1/audio/transcriptions
Learn how to transcribe audio into text, supported request parameters, response formats, and usage examples in the Speech-to-Text guide.
Text-to-Speech
Endpoint:
POST /v1/audio/speech
Generate high-quality speech from text using one of the available voices. Voice cloning is also supported through the Text-to-Speech API. Last modified on July 7, 2026