Home/Evaluation

Know how your model performs in the real MENA market

Benchmarking and evaluation across the region's languages, dialects, and use cases — reviewed by people, not scored by a single automated metric.

What we evaluate

Correctness, fluency, and register-appropriateness of Arabic model output, reviewed against a rubric by native reviewers rather than scored by an automated metric alone.

How well a model produces and understands specific dialects, not just Modern Standard Arabic — the gap between the two is often where products fail in market.

Word-error rate and intelligibility benchmarks run against real dialectal speech samples, not studio-clean read speech.

Whether a model's output is contextually appropriate for the region it's being evaluated in — not flagged as unsafe by an overly broad global filter, and not missing regional context a generic benchmark wouldn't catch.

Region-specific harm and safety review, run by native-language reviewers trained on a safety rubric.

Model performance on mixed-language input and output, which is how a large share of the region's users actually communicate.

Accuracy on domain-specific content (medical, legal, financial Arabic) where general-purpose benchmarks don't reach.

Recognition accuracy against regionally authentic imagery — signage, packaging, environments — rather than generic vision benchmarks.

Head-to-head comparison between model versions or competitors, scored by human reviewers rather than an automated proxy metric.

An illustrative example

The table below shows the shape of a benchmark report — the dimensions we score and how results are presented — using placeholder numbers. It is not a real result from any model we've evaluated; we don't have a published benchmark suite live yet.

Benchmark suite in development — first evaluation cohort targeted for the quarter following our initial buyer pilots. Tell us your target model and market on the project form and we'll scope a custom evaluation now, ahead of the general suite.

Illustrative example — not a real result
Saudi Arabic87%
Egyptian Arabic91%
Jordanian Arabic84%
Code-switching73%
Cultural alignment88%

Evaluate Your Model