Home/Evaluation
Know how your model performs in the real MENA market
Benchmarking and evaluation across the region's languages, dialects, and use cases — reviewed by people, not scored by a single automated metric.
What we evaluate
Correctness, fluency, and register-appropriateness of Arabic model output, reviewed against a rubric by native reviewers rather than scored by an automated metric alone.
How well a model produces and understands specific dialects, not just Modern Standard Arabic — the gap between the two is often where products fail in market.
Word-error rate and intelligibility benchmarks run against real dialectal speech samples, not studio-clean read speech.
Whether a model's output is contextually appropriate for the region it's being evaluated in — not flagged as unsafe by an overly broad global filter, and not missing regional context a generic benchmark wouldn't catch.
Region-specific harm and safety review, run by native-language reviewers trained on a safety rubric.
Model performance on mixed-language input and output, which is how a large share of the region's users actually communicate.
Accuracy on domain-specific content (medical, legal, financial Arabic) where general-purpose benchmarks don't reach.
Recognition accuracy against regionally authentic imagery — signage, packaging, environments — rather than generic vision benchmarks.
Head-to-head comparison between model versions or competitors, scored by human reviewers rather than an automated proxy metric.
An illustrative example
The table below shows the shape of a benchmark report — the dimensions we score and how results are presented — using placeholder numbers. It is not a real result from any model we've evaluated; we don't have a published benchmark suite live yet.
Benchmark suite in development — first evaluation cohort targeted for the quarter following our initial buyer pilots. Tell us your target model and market on the project form and we'll scope a custom evaluation now, ahead of the general suite.
| Saudi Arabic | 87% |
| Egyptian Arabic | 91% |
| Jordanian Arabic | 84% |
| Code-switching | 73% |
| Cultural alignment | 88% |