Skip to content

feat(pdfrx_engine): restore logical order to Hebrew text stored in visual order - #732

Open
palmoni5 wants to merge 1 commit into
espresso3389:masterfrom
palmoni5:feat/hebrew-text-order
Open

palmoni5 wants to merge 1 commit into
espresso3389:masterfrom
palmoni5:feat/hebrew-text-order

Conversation

@palmoni5

Copy link
Copy Markdown
Contributor

Many scanned Hebrew books come with an OCR text layer that stores each line's words left to right. Some older fonts go further and store every word's letters left to right too. Either way, loadText/loadStructuredText return the text backwards, so phrase search never matches and copied text comes out reversed.

This adds PdfHebrewTextNormalizer, opt-in through Pdfrx.normalizeHebrewText (default false).

How it decides. It uses the character geometry, not guesswork:

  • A word whose Hebrew letters advance left to right is reversed. Combining points stay after their letter.
  • A line whose Hebrew words advance left to right is reordered right to left. Whitespace keeps its slots.
  • A word or line with too little evidence (a single pair) follows the page majority, so short lines on a logical page are never flipped.
  • Text already in logical order is returned untouched (the same instance).

Legacy encodings. It also decodes Windows-1255 Hebrew that a font without a ToUnicode map exposes as Latin-1, or as Mac Roman. This only happens on a page where those characters outnumber the Latin letters, so French or German text is never touched.

Character boxes are permuted together with the text, so indices into the result keep addressing the right glyphs; highlighting and selection keep working. The PDFium backend runs the normalizer on the worker.

Measured on real books

  • 15 random HebrewBooks scans with a text layer: 7 had reversed word order, with 0–13% of adjacent word pairs right-to-left. After the fix they measure 87–100%.
  • The 8 scans that were already correct are left byte-for-byte unchanged.
  • In a 39-tractate Talmud set, only the two broken files change. One of them has its Windows-1255 / Mac Roman runs decoded, about 76k letters over 8 pages. The other 37 are untouched.

Unit tests cover logical text, word-order reversal, full visual order, points, Latin-1 decoding, accented Latin text, and the page-majority rule.

…sual order

Scanned books with an OCR text layer often write each line's words left to right, and some older fonts write every word's letters left to right as well, so loadText/loadStructuredText return the text backwards and phrase search never matches. PdfHebrewTextNormalizer uses the character geometry to tell the two apart: only a word whose letters advance left to right, or a line whose words do, is reordered; a short word or line with too little evidence follows the page majority. Text already in logical order is returned untouched.

It also decodes Windows-1255 Hebrew that a font without a ToUnicode map exposes as Latin-1 or Mac Roman characters, only on a page where those characters outnumber the Latin letters.

Character boxes are permuted together with the text. Opt-in through Pdfrx.normalizeHebrewText (default false); the PDFium backend runs it on the worker.
@espresso3389

Copy link
Copy Markdown
Owner

@palmoni5 Very impressive and interesting work. I'll look into it.

palmoni5 added a commit to palmoni5/otzaria that referenced this pull request Sep 28, 2026
הפעלת Pdfrx.normalizeHebrewText, בחלון הראשי ובחלונות המשניים. pdfrx_engine מחזיר עכשיו טקסט בסדר לוגי, והתיקון חל על חיפוש, העתקה, 'חפש מקבילות', בחירה ואינדוקס. שורה שהמילים בה נשמרו משמאל לימין מסודרת מחדש, ומילה שהאותיות בה נשמרו משמאל לימין מתהפכת. ההכרעה נעשית לפי מיקום התווים על הדף. עברית בקידוד Windows-1255 שגופן בלי ToUnicode חושף כתווים לטיניים או כ-Mac Roman מפוענחת, כמו במסכת מעילה.

במדידה: 7 מתוך 15 ספרים סרוקים מהיברובוקס היו בסדר מילים הפוך. אחרי התיקון 87%–100% מזוגות המילים הסמוכות מימין לשמאל. ספרים תקינים לא משתנים בכלל.

pdfrx_engine מאותו ענף בפורק, עד שיתמזג espresso3389/pdfrx#732.
Y-PLONI pushed a commit to palmoni5/otzaria that referenced this pull request Sep 28, 2026
הפעלת Pdfrx.normalizeHebrewText, בחלון הראשי ובחלונות המשניים. pdfrx_engine מחזיר עכשיו טקסט בסדר לוגי, והתיקון חל על חיפוש, העתקה, 'חפש מקבילות', בחירה ואינדוקס. שורה שהמילים בה נשמרו משמאל לימין מסודרת מחדש, ומילה שהאותיות בה נשמרו משמאל לימין מתהפכת. ההכרעה נעשית לפי מיקום התווים על הדף. עברית בקידוד Windows-1255 שגופן בלי ToUnicode חושף כתווים לטיניים או כ-Mac Roman מפוענחת, כמו במסכת מעילה.

במדידה: 7 מתוך 15 ספרים סרוקים מהיברובוקס היו בסדר מילים הפוך. אחרי התיקון 87%–100% מזוגות המילים הסמוכות מימין לשמאל. ספרים תקינים לא משתנים בכלל.

pdfrx_engine מאותו ענף בפורק, עד שיתמזג espresso3389/pdfrx#732.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants