How We Translate
Last updated: 2026-07-23
Every page on ShamelaTranslate is produced by a defined, repeatable process — and we believe you should be able to see exactly what that process is. This page describes how a text travels from its source to the translation you read: how we acquire the Arabic, how the AI translation is produced and constrained, which automated checks stand between the machine and publication, and where human eyes come in.
Step 1 — The Arabic source
Most works come from the public Shamela library, imported page by page and volume by volume so that nothing is dropped and nothing is reordered. (Our page numbering follows the digital library's sequence, so it can differ from the printed edition's page numbers.) We try to store the Arabic text as it reaches us: we do not rewrite or abridge it — at most we correct an obvious, unmistakable spelling slip or a clearly erroneous vowel mark (ḥarakah), never the substance of the text. Where the source marks speaker changes, Qur'anic citations, or editorial insertions, we do our best to carry that structure over so the page reads the way it was set — though such a marking can occasionally be lost in transfer.
Works that exist only as scanned images go through our double-pass extraction: two separate AI readings of the same page are compared word by word, and text is only ever accepted where the two agree. Where they differ, the page is read a third time against the original photograph — never guessed at. That third reading can settle two things: a vowel mark (ḥarakah) when both readings already agree on the letters, and a single word when it can identify with high confidence what is actually printed. Anything it cannot settle is flagged and decided by a person before the page enters the library — with one exception we make deliberately: a lone vowel mark that stays uncertain does not hold back an otherwise agreed page. That page is published, and the uncertainty is recorded for our reviewers rather than shown to you — there is no mark on the page you read. With volumes of this size it can be a while before a reviewer reaches it, so we would rather say so here than imply a check has already happened. A page that cannot be confidently read is not published half-guessed.
Qurʾānic verses — verified letter-by-letter against the muṣḥaf
Scripture is the one place we never allow a model to be the authority. After the AI reads a page, every Qurʾānic quotation it marks is matched — letter by letter, disregarding diacritics and spelling variants — against the verified muṣḥaf text (the Ḥafṣ ʿan ʿĀṣim reading, from Tanzil). When a quotation matches with certainty, we replace it with the authentic verse, rendered in full ʿUthmānī script: the exact spelling and diacritics, the ornate ﴿ ﴾ brackets, and the end-of-āyah signs with their numbers — so the wording never depends on what the model happened to type.
A quotation that cannot be matched with certainty is left exactly as read and flagged for a human — never guessed at. This is a deterministic step of our own, not the AI's: the model only tells us where a verse is; the wording comes from the muṣḥaf.
Step 2 — A glossary before a single page is translated
Classical Islamic scholarship has a precise technical vocabulary — terms of jurisprudence, hadith criticism, and theology whose established renderings matter. Before translation begins we maintain a curated glossary: each Arabic term is paired with the exact wording it must receive in German and English, together with notes on its intended sense. The glossary is reviewed by hand, and every translation request carries it, so that a term is rendered the same way on page 3,000 as on page 3.
The glossary fixes wording, never meaning: when a word is used in a different, everyday sense in a particular passage, it is translated according to its context — the glossary entry applies only where the technical sense does.
Step 3 — Translation with context, not in fragments
Classical texts do not break neatly at page boundaries — a sentence often begins on one page and ends on the next. Our pipeline therefore never translates a page in isolation. Each translation window receives the surrounding Arabic context and the already-published translation bordering it, and is required to continue that text seamlessly: same terminology, same register, no re-translated overlaps, no dropped clauses at the seams.
The main text (matn) and the editor's footnotes (taḥqīq) are translated as separate streams, because they are different kinds of text: running scholarly prose on the one hand, dense citation apparatus on the other. Footnote numbers, bracketed editorial remarks, and the source's own paragraph and line structure are preserved.
Step 4 — Automated checks before anything is published
Every translated window must pass validation before it is stored. Among other things we verify that:
- every requested page came back — none missing, none duplicated, none reordered;
- no Arabic was left untranslated and no passage was simply echoed back in the source language;
- the translation's length is plausible for the source — a strong signal of dropped or invented text;
- the seams match: text continuing across a page boundary must join the already-published wording without repeating or contradicting it.
A window that fails these checks is retried with stricter instructions; if it still fails, it is withheld and flagged for human review. We would rather show you an honest "not yet translated" notice than a page that failed its own quality gate.
Step 5 — Human review and the status you see
AI translation, however carefully constrained, is a draft — and we label it as such. Every translated page carries a status you can see:
- AI translation — produced by the pipeline and validated automatically, not yet checked by a person.
- Human-reviewed — a person has read this page's translation against the Arabic and corrected it where needed.
- Verified — the translation has passed a final human confirmation.
Reviewed pages are protected: an automated re-run can never silently overwrite a translation a person has corrected. Review proceeds continuously, and reader reports (see below) move pages up the queue.
Limitations — read this before relying on a translation
We are transparent about what this library is: a reading aid produced with AI assistance under human oversight, not a substitute for the Arabic original or for qualified scholarship. Machine translation can miss nuance, mishandle rare constructions, or flatten a deliberately ambiguous phrase. That is why the Arabic original is always available beside the translation — one tap away, never hidden — and why every page states its review status.
For religious practice, legal rulings, or academic citation, consult the Arabic original and qualified scholars. Our translations are a bridge into the text — they are not a fatwa, and they are not the text itself.
Found a mistake? Tell us — it genuinely helps
Signed-in readers have a feedback control on every reading page, and the contact page is open to everyone. Reported passages are checked against the Arabic by a person, corrected where needed, and marked as human-reviewed. Reader corrections are among the most valuable contributions this project receives: they improve the very page you were reading, for everyone who comes after you.
