Arabic is not one language to a model — it is a morphologically rich core plus living dialects that diverge from the textbook. Models trained on translated or scraped text speak a faded copy of it. We build the missing layer: knowledge captured from inside its own context, dialect-stamped and source-documented, with a clean chain of rights — consented, owned, and quality-controlled.
What we provide
Delivered as clean JSON with full metadata — dialect, country, device, timestamp, and consent status on every record.
Native speakers pick the better of two responses in one click (RLHF-style). Pairwise, gold-controlled, dialect-tagged.
Rating and error-flagging of model answers in Arabic — fluency, accuracy, cultural fit — by real Arabic speakers.
Authentic dialect prompts, responses, and labeled text spanning the Arab world's real linguistic diversity.
The same sentence in Modern Standard Arabic and in real Levantine, Egyptian, Gulf, or Maghrebi — aligned pairs that teach models the divergence textbooks hide.
Root, pattern, and part-of-speech labels by native speakers — structured signal for the morphology that makes Arabic hard for tokenizers and models alike.
Original prose written by Arabic speakers about their own markets, customs, and daily life — a growing owned archive, not scraped or translated content.
Why Riziq
Not textbook Arabic and not translation — real Levantine, Egyptian, Gulf, and Maghrebi written by in-region native speakers about their own world. The wide authentic archive is the decisive factor for Arabic AI, and it can't be scraped into existence.
Embedded gold items and multiple annotators per item, with inter-annotator agreement (Cohen/Fleiss κ) and below-threshold exclusion.
Every contributor grants explicit, dated, versioned consent to commercialize. We license data we own — no scraped, no third-party content.
Per-record country tagging and a full chain of custody, so your compliance team can trace every data point.
A/B balance checks and per-annotator accuracy scoring filter out low-effort work before delivery.
Personal identity is separated before export — you receive a pseudonymous id, never a name or email.
How it works
You share the spec (type, dialect, size, schema, acceptance metrics). We agree an SOW + NDA.
We deliver a small labeled sample (~200 records) with a real QC report so you can verify quality first.
Clean JSON + metadata, IAA report, and consent status per record. Formats to match your pipeline.
Once approved, we scale the workforce and cadence — recurring datasets, one relationship.
التسعير · SLA
خمس باقات بعمق تحقق متصاعد — لكل باقة اتفاقية مستوى خدمة واضحة: عدد المقيّمين لكل عنصر، الدقة المتوقعة، والسعر. اختر العمق الذي يحتاجه مشروعك.
| الباقة | عمق التحقق | الدقة المتوقعة | السعر | الأنسب لـ |
|---|---|---|---|---|
| Bronze | 1× + أسئلة تحقق | ~85% | ×1 | نماذج أولية، طلاب |
| Silver | 3× إجماع | ~94% | ×2 | RLHF قياسي، شركات ناشئة |
| Gold الافتراضي | 5× + أسئلة تحقق + تقرير IAA | ~97% | ×3.5 | مختبرات AI جادة |
| Platinum | 7× + حَكَم خبير + تدقيق 5% + سجل حيازة كامل | ~99% | ×7 | طبي/قانوني/مالي |
| Benchmark | 25×+ + إجماع خبراء + توثيق فردي | 99.9%+ | تسعير خاص | مجموعات تقييم النماذج |
قل لنا كم خطأ تتحمل في المليون — نبيعك العمق الذي يضمنه، مع تقرير جودة يثبته لكل عنصر.
Tell us the dialect and task you need. We'll send a labeled pilot sample and a QC report — no commitment.
Request a sample →