Dialect is the variable that matters most here, and it's also the one generic crowd platforms handle worst: a script read by an Egyptian speaker and one read by a Gulf speaker are not interchangeable training data, even though both are "Arabic." We recruit by dialect region, not just by language, and we screen for natural delivery — read-aloud scripts sound different from spontaneous speech, and a model trained on one performs poorly on the other.
What a reviewer checks
Audio quality (background noise, clipping, mic distance), transcription accuracy against the actual recording, and — the check that's easy to skip and shouldn't be — whether the speech is actually representative of natural dialect use, not a script read in an unnatural register.
Worked example. A speech-AI team needed conversational Gulf Arabic with natural code-switching into English for customer-service terms. We recruited bilingual Gulf-based contributors and collected unscripted conversational pairs rather than script reads, which is what code-switching data actually requires to be usable.