Text is the modality with the widest range of task shapes: instruction-response pairs, multi-turn conversations, reasoning traces, translation pairs, and domain corpora (legal, medical, financial Arabic text with its own vocabulary and register). We staff each project with contributors matched to the register the task actually needs — Modern Standard Arabic for formal or legal content, a specific spoken dialect for conversational data, code-switched text where that's how the target users actually write.
What a reviewer checks
A text task is not accepted on submission — it goes through a review pass against a written rubric specific to the project: factual accuracy, register consistency (no MSA leaking into a dialect conversation), instruction-following fidelity, and — for reasoning data — whether the intermediate steps actually support the conclusion, not just whether the final answer is right.
Worked example. A buyer needed 20,000 Arabic instruction-response pairs for a customer-support fine-tune, in Egyptian Arabic specifically. Contributors were screened for native Egyptian Arabic and support-domain familiarity; a second reviewer tier checked each response against a tone-and-accuracy rubric before delivery.
Formats we deliver
JSONL or CSV with the fields your training pipeline expects, plus a metadata file (contributor dialect/region, review pass results, task version) and a QA summary. See quality & methodology for what that QA report actually contains.