Dialect-level results
See performance broken down the way your users speak.
We build held-out test sets from verified native speakers for each dialect and condition you care about, then score your system against them.
ASR & speech-to-text
Word error rate by dialect, speaker group, and environment.
Voice agents
Whether the agent understood the request, across dialects.
Conversational AI
End-to-end behavior on realistic Arabic conversations.
| Segment | Clean audio WER | Noisy audio WER |
|---|---|---|
| Saudi Arabic | ||
| Najdi | — | — |
| Hijazi | — | — |
| Eastern | — | — |
| Egyptian | — | — |
| Jordanian | — | — |
| Arabic–English code-switching | — | — |
Illustrative layout only. Scores are produced for your system, on your engagement — we don’t publish benchmark results.
A continuous cycle
Evaluation is where improvement starts.
Every weakness we find maps to data that can fix it. Measure, collect what’s missing, retrain, and measure again.
- Evaluate
- Find model weaknesses
- Build a targeted dataset
- Improve the model
- Evaluate again
Find out where your model struggles.
Tell us your system and target dialects. We’ll scope an evaluation pilot.