Home/Data

What data do you need?

MenaData covers six modalities across the region. Each one is collected and reviewed by people who speak the language, live the dialect, or hold the professional credential the task requires — not a generic crowd panel.

Every project starts the same way regardless of modality: we define the task with you, recruit and screen the specific contributor profile it needs, and run a review layer before anything ships. What differs by modality is the tooling, the review criteria, and who does the reviewing.

Text
LLM training, fine-tuning, reasoning traces, instruction data, translation, domain-specific corpora.
Audio & speech
ASR/TTS pairs, dialect recordings, conversational speech, code-switching, call-center audio.
Image
Real-world MENA imagery for computer vision, OCR, retail, mapping, document AI.
Video
Human actions and environments for multimodal training and robotics.
Human feedback
Preference ranking, RLHF, safety review, cultural-alignment judgments.
Expert data
Physician, lawyer, engineer, and other professional review or authored content.

Each type below has its own page covering exactly how it's collected, what a reviewer checks for, and what you receive.