“Collect 300 hours of spontaneous Saudi Najdi conversations from 500 verified native speakers.”
Targeting
Specify every dimension that matters to your model.
Geography & dialect
Country, dialect, and region — Najdi rather than just “Saudi”.
Speakers
Gender, age range, and number of unique speakers.
Recording environment
Quiet rooms, cars, streets, offices, or call-center conditions.
Device
Mobile phones, headsets, far-field microphones, or telephony audio.
Speech type
Read scripts, prompted responses, or spontaneous conversation.
Domain
Banking, healthcare, telecom, retail, automotive, or your own topics.
Code-switching
Natural Arabic–English mixing, where your users speak that way.
Volume
Hours of speech, split however your training plan needs.
Managed end to end
You set the requirements. We run the collection.
-
01
Speaker recruitment
We source native speakers who match your profile and verify their dialect.
-
02
Consent
Each speaker agrees to the specific recording and its commercial use.
-
03
Recording
Speakers record under the conditions and scripts the project defines.
-
04
Validation
Recordings are checked for audio quality, dialect, and completeness.
-
05
QA
A separate reviewer approves each item or sends it back.
-
06
Delivery
Audio, transcripts, metadata, and consent records in your format.
Deliverables
What you receive.
- Audio files in the sample rate and format you specify
- Transcripts if the project includes annotation
- Speaker & session metadata dialect, demographics, device, environment
- Dataset documentation collection method, guidelines, and QA summary
- Consent records documenting the rights granted for your use
Start with a representative sample.
A pilot lets you validate the data against your model before committing to full volume.