A reviewer reads a prompt and one or more candidate responses and records a judgment: which response they'd accept, whether either is unsafe, whether the dialect used is authentic, whether a cultural reference lands correctly or misses. The judgment plus the reviewer's stated reason is the training signal.
Task types we run
- Pairwise preference ranking (A vs. B, with rationale)
- Absolute quality scoring against a rubric
- Safety and harm review, including region-specific sensitivity that a global safety rubric misses
- Dialect-authenticity checks — does this actually sound like natural Levantine Arabic, or does it read as MSA wearing a dialect's vocabulary
- Hallucination review against a source document or known facts
How disagreement is handled
Reviewers disagree — that's expected, not an error state. Tasks with low inter-annotator agreement are routed to a senior reviewer rather than resolved by majority vote alone, and disagreement rates are reported back to you as part of the QA package, not hidden. See quality & methodology for how agreement is measured.