The Arabic reads perfectly. The obligation is gone.
A model renders consideration as الاعتبار and the sentence scans. BLEU sees high overlap. A bilingual reviewer without legal training sees nothing wrong. The promise is no longer enforceable. A 3,009-item corpus of exactly those moments — each paired with the correct rendering, the plausible-but-wrong one, and the reason, written out by the attorney who caught it.
Which rendering would a lawyer sign off?
Four real items from the public sample. Choose, then read the reason.
Render for use in a binding contract:consideration
Standard metrics grade fluency. The law sits in the word choice.
Three things make Arabic legal translation quietly unscoreable.
The errors are invisible to metrics
These failures are fluent by construction. Surface overlap is high, perplexity is low, and an LLM judge without doctrinal grounding agrees with the wrong answer. Detection needs the rationale — which is exactly what this corpus supplies.
The public data is thin
Arabic legal parallel text is scraped statute and broadcast. It teaches register, not the boundary between باطل and قابل للفسخ, or between ربح and فائدة. Nothing public exercises those distinctions deliberately.
The stakes are asymmetric
A dropped modal turns an obligation into a forecast. “Immunity from execution” rendered as الحصانة من الإعدام moves the reader into a different body of law. The cost of the failure has no relation to the size of the edit.
Every item is a trap, a fix, and a reason.
One JSON object per line. Two field sets keyed by direction. IDs run alfe-0001–alfe-3009, continuous, canonical key order, with no Latin script in the Arabic fields and no Arabic in the English fields.
The schema
{
"id": "alfe-0010",
"domain": "islamic_finance",
"register": "MSA",
"task": "en_to_ar_term",
"source_en": "constructive possession",
"gold_ar": "القبض الحُكمي",
"chosen_ar": "القبض الحُكمي",
"rejected_ar": "الحيازة البنائية",
"error_type": "domain_term_of_art",
"rationale": "the established fiqh term, distinguished
from actual possession; 'constructive'
calqued reads as building-related.",
"difficulty": "hard"
}
What each field carries
gold_* | Reference rendering, usually with a bracketed gloss of the concept. |
chosen_* | Preferred rendering — shorter, no gloss. The positive side of a preference pair. |
rejected_* | The trap — wrong, but the wrong a model or a non-lawyer actually produces. Clause items stack two to seven faults. |
error_type | One of fifteen named failure modes. |
rationale | The core IP: why chosen is right and exactly what rejected does to the meaning — written to serve as a grading rubric or chain-of-thought. |
register | MSA for drafting terms; Egyptian, Levantine and Gulf for client-facing adaptation. |
Terms, clauses, homographs, and the client in the room.
en_to_ar_term
The drafting term, the tempting calque, and the sense a dictionary hands you first. The bulk of the corpus.
clause_faithfulness
Real contract sentences where the rejected version loses a “shall”, inverts a party, or drops a limiter. 480+ clauses.
ar_to_en_term
Where one unpointed Arabic word decides which body of law applies: النقض as cassation or as treaty denunciation. 240+ items.
register_adaptation
Explain the rule to a client in Egyptian, Levantine or Gulf Arabic without losing the legal content. Each scenario written three ways.
Fifteen named ways it fails.
Every item carries exactly one error_type. Group by it and you get a failure profile, not a score.
| error_type | What goes wrong |
|---|---|
polysemy_false_sense | Right word, wrong semantic branch |
legal_meaning_drift | Plausible Arabic that loses or inverts the legal effect |
near_synonym_drift | Adjacent term that shifts doctrine or register |
concept_calque | Structural transfer that strips the legal concept |
false_cognate | Diacritical or near-form confusion |
domain_term_of_art | Wrong choice between two doctrinal bodies |
literal_calque | Word-by-word where a term of art exists |
under_translation | Right root, too weak — strips an element |
register_and_compliance | Register failure and compliance failure stacked |
compliance_critical | Sharia-compliance violation (riba framing and kin) |
term_of_art | Non-standard, academic, or wrong-tradition term |
register_mismatch | Right concept, wrong register for the audience |
unnecessary_transliteration | Latin importation where the Arabic is settled |
modality_loss | Obligation downgraded to prediction |
void_vs_voidable | Conflates void ab initio with rescindable |
What a trap actually looks like.
“The financing carries a profit rate of 5% per annum.”
The return must be framed as profit (ربح), never interest (فائدة). Rendering it as معدّل فائدة does not merely mistranslate — it recharacterises the product as riba-bearing and is disqualifying. The single highest-stakes trap in the set.
drag-along rights
Drag-along lets a majority compel the minority to join a sale. حقوق السحب المصاحب literally calques “drag along” and conveys nothing to a practitioner reading the SPA.
“Explain to a client, plainly, that the contract is null and void (void ab initio).”
ملغى (“cancelled”) wrongly implies a once-valid contract later undone — closer to voidable. The distinction changes remedies and restitution. The gold pairs the precise MSA term with a natural Levantine gloss.
Seventy-four chunks, each a designed set.
Every domain area ships as a set: term-level traps, the operative clauses that re-embed them, the AR→EN homographs a reader meets going the other way, client-facing dialect scenarios, and compound clauses that stack several errors for the rationale to name.
Core commercial
Islamic finance
Regulatory & specialist
New economy & institutional
Four pipelines, one file.
Evaluation with a gold reference
Score EN↔AR output against gold_* and grade with the rationale as rubric instead of surface overlap.
Preference tuning
The (source, chosen, rejected) triples are drop-in RLHF or DPO pairs. Clause items supply multi-error negatives.
Supervised fine-tuning
source → gold with the rationale as chain-of-thought, so the model learns the reason, not just the string.
Error-mode diagnostics
Group by error_type, domain or register to find where a model systematically breaks.
How it is built
Authored top-down from known failure modes by a dual-barred (NY/MA) corporate attorney practising M&A, PE/VC and Islamic finance, native across MSA and Egyptian, Levantine and Gulf Arabic. Renderings are checked for legal correctness against a GCC and Egyptian civil-law baseline, for register in the stated dialect, and for compliance framing in Islamic-finance items. Every item is reviewed and adopted by the author before it enters the licensed corpus.
No scraped text. No personal data. No third-party corpora. No crowd labour. Original authored work — which is why it can be warranted and licensed cleanly.
Machine checks on every build: schema and key order, taxonomy values, contiguous IDs, script integrity in every field, chosen ≠ rejected, and no duplicate sources within or across builds.
Disclosed up front
Nine source headwords appear twice, all within alfe-0001–1001 and listed with IDs in the data card. Six pairs test materially different failure modes; three are near-redundant. Deduplicating on source loses three items and nothing of substance.
gold and chosen are not interchangeable. gold is the full reference, usually carrying a bracketed gloss or an accepted variant; chosen is the tight form a practitioner writes. Use gold for scored evaluation and chosen for preference pairs.
Some items describe regimes that move — passenger-rights compensation, virtual-asset licensing, apostille membership, the BNPL perimeter. These are re-verified at each annual refresh.
Chunk notes travel with each block of 42 items, listing the domain, the distribution, and the doctrinal points your own reviewers will want to test first.
Sample first. Then the tier that fits the job.
A short data-licence agreement governs permitted use — training, fine-tuning, evaluation, benchmarking and alignment — plus exclusivity and non-redistribution. The models you train are yours, and so are their outputs.
| Option | Scope | Indicative |
|---|---|---|
| Sample | 17 items on Hugging Face, gated. Evaluation only; no redistribution or derivative datasets. | Free |
| Pilot set | About 500 items, non-exclusive, one project. | $4k – 8k |
| Scaled corpus | 2,000 – 5,000 items, non-exclusive. Shaped to your taxonomy, difficulty mix, dialect mix and format. | $15k – 60k |
| Full / exclusive | All 3,009 items, exclusivity by field or language pair, or a bespoke build to your specification. | On application |
| Kept current | Annual refresh as statutes, AAOIFI standards and case law move. | 20 – 30% / yr |
Indicative only; final scope and price by agreement.
Before you ask.
Can we evaluate it before committing?
Yes — start with the free 17-item sample on Hugging Face. It carries the full schema and rationale style, so you can wire up the judge prompt and see what your model does before any money moves. The pilot set then gives you roughly 500 items across the taxonomy for a single project.
Is 3,009 items enough to train on?
It is not a pretraining corpus and does not pretend to be. It is a diagnostic and preference set: dense, adversarial, and labelled with the reason each wrong answer is wrong. Used as DPO pairs or as an eval harness, that density is the point. If you need volume built to your own taxonomy, that is the scaled tier.
Who does the Arabic hold up for?
MSA carries the black-letter drafting vocabulary and is weighted accordingly. The 220+ dialect items are the hardest slice and the one most datasets omit entirely: explaining a legal point colloquially in Egyptian, Levantine or Gulf Arabic while holding the technical term stable. Each scenario is written three ways with identical legal substance, so register is isolated from content.
What exactly are we buying?
JSONL, one object per line, with a pretty-printed JSON review copy, per-chunk files, and chunk notes listing the domain, distribution and review priorities for each block of 42 items. Plus a commercial licence under which you keep the models and the outputs.
Test it on your own model before you talk to us.
Request the gated sample, run it through your evaluation harness, and see what the rejected renderings do to your scores. Then tell us the taxonomy and volume you need.