Skip to content
ffiniti byte
Home/Real-Time Speech Translation
Speech AI · Deployed in mosques across the Gulf

Real-time sermon translation into 80+ languages, running in production

A Friday congregation in the Gulf may speak a dozen languages while the sermon is delivered in one. Qudwah translates the imam as he speaks, into each listener’s own language, on their own phone, with published latency figures that remain unresolved.

Inspect the evidence

Published latency figures remain unresolved: Qudwah’s product FAQ says 2–3 seconds, while the earlier engineering account below reports under 500 ms. The published conditions do not reconcile the two.

Read the measurement status and open questions alongside the engineering account below. Internal product metrics are not independently verified results.

80+Languages translated in real time
<500ms (disputed)Speech-to-audio translation latency
96%Terminology accuracy on domain terms
iOS + AndroidPublished in both app stores

The problem

Friday congregations in the Gulf are among the most linguistically mixed audiences anywhere in the world. The sermon is delivered in Arabic, and a large share of the room does not follow it closely enough to take anything away.

The usual remedies do not scale. A printed summary arrives after the fact. A second sermon in a second language serves one more group and no others. A human interpreter serves one language at a time, and mosques do not have interpreter booths. What was needed was something that works for every listener at once, in the room, while the imam is still speaking.

What we built

Qudwah listens to the sermon, translates it as it is spoken, and delivers it to each listener as text and audio in the language they chose. A listener joins by scanning a QR code at the mosque entrance, picks a language, and puts in their earphones. Nothing is installed on the mosque side beyond the audio feed.

1. CaptureThe imam’s audio is transcribed live by a speech model fine-tuned on Arabic religious speech, including regional accent and recitation cadence.
2. TranslateTranscribed speech is translated by models trained on Islamic vocabulary rather than general-purpose text, then spoken back with text-to-speech.
3. DeliverAudio and text are streamed over WebRTC to every connected listener in their own language, thousands at a time, under half a second behind the speaker.

What made it difficult

  • Religious terminology tolerates no driftA general translation model renders Islamic terms loosely, and a loose rendering can change the meaning of a passage entirely. The models were trained on Islamic vocabulary specifically, and the output was reviewed by religious scholars before the system went into any mosque.
  • Latency is the product, not a metricSpeech translation that lands two seconds late stops being translation and becomes subtitling: the listener is always behind, and the room’s reactions give away what they missed. Holding the full path from spoken word to audio in the ear under half a second is what makes it usable at all.
  • One speaker, thousands of listeners, many languagesThe system fans a single audio source out to thousands of concurrent listeners, each on a different language track, over consumer mobile networks inside buildings that were not designed for connectivity.
  • It has to work for people who will not read a manualWorshippers arrive minutes before the sermon and are not going to configure anything. The entire path from arriving to listening is a QR code and one tap on a language.

Beyond translation

Once congregations were using the app weekly, it grew into the surrounding needs: a mosque locator, access to Quran and Hadith, recordings of past sermons, a daily prayer tracker, and a halal reference. The translation engine remains the core, but the product people actually keep on their phone is the whole set.

Technology

Selected for this system, not from a house standard.

OpenAI Whisper (fine-tuned)Custom transformer modelsAzure TTSWebRTCRedisKubernetes

Check it yourself

Qudwah is published in the Apple App Store and on Google Play, and the product site is open to anyone. You do not have to take our word for the latency or the language coverage. Download it and listen.

Open qudwah.ai →

Planning something in this direction?

Tell us what the system has to do and where the data sits. You get a 30-minute technical call within 24 hours, and we can sign an NDA before you share anything commercial.

Book a technical discovery call

Start a conversation

What needs to
work better?

Tell us about the problem, the people and the data.
We will help you work out the next step.

Discuss your project