ARABIC DATA · CONSENTED · DEMAND-FIRST

Riziq Data — spec in, then we collect.

Arabic is not one language to a model. We do not sell a ready-made dialect catalog. You create a project — name, task type, dialect, requested rows, criteria, budget, deadline, JSONL or CSV, and quality — and we collect from native speakers with dated consent. The public labeled sample is whatever rows exist; today that count is zero.

Riziq Data is a consented Arabic data partner, not a task app and not a stocked warehouse. Preference pairs, dialect text, MSA↔dialect, or evaluation sets are capabilities we open after a spec — not inventory on this page.

After a written spec, the work language can be Arabic or English (labeling, transcription, or evaluation). That is a capability we open when you write the spec — not inventory on this page.

What we can collect — after a spec

Capabilities, not a shelf.

These are types we can open once you write the project — task type, dialect, rows, criteria, delivery as JSONL or CSV, and quality. They are not listed as available stock. Delivery follows the format you choose; we do not invent row counts. When a buyer writes that project, we can also collect first-person video with Arabic narration (the contributor's own work; identifiable faces and voices only with consent). That is a capability after a spec — not a catalog SKU and not stock on this page.

⚖️

Human-preference pairs

Native speakers pick the better of two Arabic responses (RLHF/DPO-style). Dialect tags and gold items only if your spec asks for them and work is done.

🔍

Model-output evaluation

Rating and error-flagging of model answers in Arabic — if that is the spec — by Arabic speakers who consent.

🗂️

Dialect text

Prompts and labeled text in a dialect you name. We do not fill dialect doors on this page; representation is whatever the live sample and stats show.

MSA ↔ dialect pairs

The same sentence in Modern Standard Arabic and a named dialect — collected when a spec asks for aligned pairs.

📜

Original long-form text

Prose written by Arabic speakers about their own world, if requested. Not a growing archive we claim today — the public sample is empty until rows exist.

🕌

Religious Q&A (if specified)

رياض المتقين — sourced Islamic items only, kept unmixed with general trivia. Opened only if a spec asks for that domain.

الحلقة المغلقة

مساهمة → توسيم → جودة → منتج مشتق → مهام جديدة من المنتج.

كل طبقة يقيّمها المستخدمون أنفسهم. المساهمة الخام (جملة لهجية، صوت، سؤال ديني) تُوسَم وتُحكَّم، ثم تصبح أزواجاً تفضيلية أو كوربوساً مجالياً أو مجموعة تقييم محجوزة — ثم قد تعود مهام حكم جديدة. لا أرقام حجم مخترعة هنا.

أسطح المساهم: قيّم واربح · ساحة التحكيم · ساهم بالبيانات (لهجتك في جملة) · فريق التقييم · رياض المتقين (ديني / فتاوى فقط). موافقة صريحة قبل أي ترخيص، وتدريب ذهبي قبل الطابور الحيّ إن وُجد ذهب.

العائلة (عنصر واحد يتضاعف): مساهمة خام → وسم رخيص (تصنيف / نسخ / كيانات إن وُجدت) → اختبار بشري (ذهب + تدريب + تحكيم) → تفضيل أ/ب → طبقات تُباع: DPO · SFT · صوت+تفريغ+لهجة · كوربوس ديني مجالي · تقييم محجوز لا يُتدرَّب عليه. بلا جدول أسعار مخترع وبلا أرقام حجم.

Why a spec first

Quality is a method we apply to work that exists — not a badge on an empty shelf.

🌍

From inside the language

Not textbook Arabic and not translation. Dialects are collected when a spec names them. We do not advertise a filled Levantine / Egyptian / Gulf / Maghrebi warehouse.

🎯

Gold and IAA — if the spec asks

Embedded gold and more than one annotator are tools we can run on a job. We do not claim a current IAA report or gold-controlled inventory while the public sample is empty.

🎙️

Labeled and spoken QC — when work runs

For labeled text — and for spoken or video work if a spec orders it — we require complete connected sentences. If a three-part narration spec is used: task, then scene, then action and effect. We reject object-lists and action-before-scene. This is a method, not a live SKU.

📝

Consented before license

A contribution enters a licensable corpus only after explicit, dated, versioned consent. No scrape, no third-party dump.

🛡️

Provenance on delivered records

Country and consent status travel with a record when we have one to deliver. There is no chain of custody to inspect on zero rows.

🚫

Low-effort filters when work runs

Balance checks and accuracy scoring exist in the pipeline. They do not imply a live annotator bench you can buy today.

🤖

Flags go to a human

Anomalous timing or repetition can be flagged. Penalties are not automated. We do not sell an ISO or SLA package.

🔒

Identity stays off the export

A buyer receives a pseudonymous id, never a name or email — on records that actually exist.

Honest about compliance: consent, access control, and audit logging are in the product. We do not claim a badge, ISO, or SOC report we do not hold.

How it works

Spec in. Then tasks. Then rows.

Spec

You write the project on this page: name, task type, dialect, rows, criteria, budget, deadline, JSONL or CSV, and quality. That is the sale path — not a dialect catalog and not a partner «اطلب هذه المجموعة» button.

Collect

We open work against that project, with consent. No checkout in this step.

Deliver what exists

Delivery is JSONL or CSV as you chose. An IAA or gold report is produced only if the project asked for it and enough labeled rows exist. The public sample stays empty until then.

Buyer project

Create a data project.

Required: project name, task type, work language, and delivery format. We save the project on our server first, then email hello@riziqapp.com when mail is configured. A person sends a price and timeline estimate after that — we do not invent a unit rate.

This is the language of the dataset and the work — not the language of this page. English labeling or transcription is opened only after this spec is written. It is not listed as stock.

Or email hello@