The numbers behind our own systems, and what they actually mean
A figure without its measuring conditions is not a claim, it is decoration. This page covers the figures for our two public systems and the group figures that appear alongside them: what was measured, under which conditions, and where the honest answer is that it has not been independently verified. Numbers on other pages that are not listed here have not yet been through this treatment, and saying so is part of the point.
Product evidence and measurement status
We checked the public sources below on 10 September 2026. You can inspect Sylos’s product description and account entry point, and Qudwah’s app-store listings. These sources establish the published products and their described workflows. They do not establish output quality, deployment scale or customer results.
Qudwah: a timing difference we need to resolve
The Qudwah product FAQ describes a delay of 2–3 seconds. Our engineering account below reports under 500 ms from speech to translated audio. The published test conditions do not explain the difference. Treat the earlier number as an internal report awaiting reconciliation, not a verified end-to-end benchmark.
Language coverage also varies across the product and store descriptions. Check the supported language, platform and output mode for your intended deployment.
Sylos: page delivery and answer generation
The 180 ms median below measures delivery of the web page. It does not measure a legal answer. The document count, query-language coverage and assistant-response figures are internal reports; we have not published a dated corpus export or a reproducible answer-quality evaluation.
Evidence to agree before a pilot
- Define the task, product version, dataset and sample count.
- Agree what counts as a correct answer, including source fidelity and appropriate abstention.
- For latency, define the start and end events, network and concurrent load; report median and tail results.
- Record errors as well as successes and agree acceptance criteria before deployment.
The reading-assistant and meeting-translation figures elsewhere on this site also remain company-reported. Their published descriptions do not include a reproducible evaluation set. Group staffing and deployment figures describe the group, not the headcount or independent results of the Dubai entity.
Public source links
Sylos product and account access · Qudwah product and FAQ · Qudwah on the App Store · Qudwah on Google Play
Response times you can reproduce yourself
Both of our public systems can be timed by anyone. Method: 50 sequential full page requests per target on 6 September 2026, from a machine in Europe, measuring total wall-clock time including DNS resolution, TCP connect and the TLS handshake, using a standard command-line HTTP client. Percentiles are computed by nearest rank over those 50 samples: the 95th percentile is the 48th fastest run, which is the third slowest of the fifty. At this sample size a 99th percentile would simply be the slowest run, so we report the maximum separately rather than dressing it up as a tail estimate.
- sylos.aip50 180 ms · p90 225 ms · p95 244 ms. Fastest run 139 ms, slowest 285 ms, across 50 samples.
- qudwah.aip50 251 ms · p90 354 ms · p95 371 ms. Fastest run 196 ms, slowest 618 ms, across 50 samples. That slowest run is a single cold start, and we report it rather than trimming the tail.
What this does not measure: it is the time to serve the product’s web front end, not the inference latency inside the system. Those are different numbers and we keep them apart on purpose. Run the same test yourself and you should get figures in the same range; if you do not, we would like to know.
The system figures, and where each one comes from
- Previously reported under 500 ms speech-to-audio latency (Qudwah; disputed)End to end, from spoken input to translated audio reaching the listener, under production conditions in deployed installations. This is an internal operational measurement from the running system, not an independently audited benchmark. It has not been verified by a third party.
- 96 per cent terminology accuracy on domain terms (Qudwah)Measured on religious and domain terminology, with output reviewed by religious scholars before deployment. Internal measurement, not third-party audited.
- 3M+ legal documents indexed (Sylos)Count of documents in the production index across the covered jurisdictions. Coverage is not uniform across all four, and the index is only as current as the last ingestion run.
- 110+ languages for queries, under 2 s average query response (Sylos)Average across production queries. Average, not a percentile: a median and a 95th percentile would tell you more, and publishing those is on our list rather than done.
- 200+ languages supported by our AI modelsThis counts the language coverage across the group’s speech and translation models taken together, not the coverage of any single system on this page. Sylos answers queries in 110+ languages; Qudwah translates into 80+. Treat the 200+ figure as a capability total, not as a property of either system.
- 12 AI systems in GCC government production, 8+ years, 600+ engineers, 5 delivery centersThese belong to the group this company is part of, not to the Dubai entity, and are labelled that way everywhere on this site. They are not independently audited figures.
What we have not published, and why
We do not publish accuracy benchmarks against competing systems. Running such a comparison properly means agreeing a task, a dataset and a scoring method that both sides accept, and anything short of that is a marketing exercise dressed as measurement.
We do not publish business outcomes for client projects, because those numbers belong to the clients and most of that work sits under confidentiality. Where you need one before you commit, ask for a reference call and we will arrange it under NDA.
There is no independent penetration test report on this site yet. When there is one, it will be linked here, including whatever it found.
How we would measure yours
On your project the same discipline applies, before anything is built. The evaluation set is drawn from your material rather than a public benchmark, your team owns it, and the threshold that decides whether we ship is agreed in writing at the start. The harness is written before the pipeline and re-runs on every change, which is what makes a later model swap a measurement rather than an argument.
Where a task has no clean automatic metric, we define a sampled human review with a written rubric and a target agreement rate, and we say so, rather than inventing a score that looks precise.
A number here you want to test?
Both systems are public. Time them, push them, and tell us what you got. If your measurement disagrees with ours, that is worth a conversation.
Tell us what you measuredStart a conversation
What needs to
work better?
Tell us about the problem, the people and the data.
We will help you work out the next step.