Conversation
…sual order Scanned books with an OCR text layer often write each line's words left to right, and some older fonts write every word's letters left to right as well, so loadText/loadStructuredText return the text backwards and phrase search never matches. PdfHebrewTextNormalizer uses the character geometry to tell the two apart: only a word whose letters advance left to right, or a line whose words do, is reordered; a short word or line with too little evidence follows the page majority. Text already in logical order is returned untouched. It also decodes Windows-1255 Hebrew that a font without a ToUnicode map exposes as Latin-1 or Mac Roman characters, only on a page where those characters outnumber the Latin letters. Character boxes are permuted together with the text. Opt-in through Pdfrx.normalizeHebrewText (default false); the PDFium backend runs it on the worker.
Owner
|
@palmoni5 Very impressive and interesting work. I'll look into it. |
palmoni5
added a commit
to palmoni5/otzaria
that referenced
this pull request
Sep 28, 2026
הפעלת Pdfrx.normalizeHebrewText, בחלון הראשי ובחלונות המשניים. pdfrx_engine מחזיר עכשיו טקסט בסדר לוגי, והתיקון חל על חיפוש, העתקה, 'חפש מקבילות', בחירה ואינדוקס. שורה שהמילים בה נשמרו משמאל לימין מסודרת מחדש, ומילה שהאותיות בה נשמרו משמאל לימין מתהפכת. ההכרעה נעשית לפי מיקום התווים על הדף. עברית בקידוד Windows-1255 שגופן בלי ToUnicode חושף כתווים לטיניים או כ-Mac Roman מפוענחת, כמו במסכת מעילה. במדידה: 7 מתוך 15 ספרים סרוקים מהיברובוקס היו בסדר מילים הפוך. אחרי התיקון 87%–100% מזוגות המילים הסמוכות מימין לשמאל. ספרים תקינים לא משתנים בכלל. pdfrx_engine מאותו ענף בפורק, עד שיתמזג espresso3389/pdfrx#732.
Y-PLONI
pushed a commit
to palmoni5/otzaria
that referenced
this pull request
Sep 28, 2026
הפעלת Pdfrx.normalizeHebrewText, בחלון הראשי ובחלונות המשניים. pdfrx_engine מחזיר עכשיו טקסט בסדר לוגי, והתיקון חל על חיפוש, העתקה, 'חפש מקבילות', בחירה ואינדוקס. שורה שהמילים בה נשמרו משמאל לימין מסודרת מחדש, ומילה שהאותיות בה נשמרו משמאל לימין מתהפכת. ההכרעה נעשית לפי מיקום התווים על הדף. עברית בקידוד Windows-1255 שגופן בלי ToUnicode חושף כתווים לטיניים או כ-Mac Roman מפוענחת, כמו במסכת מעילה. במדידה: 7 מתוך 15 ספרים סרוקים מהיברובוקס היו בסדר מילים הפוך. אחרי התיקון 87%–100% מזוגות המילים הסמוכות מימין לשמאל. ספרים תקינים לא משתנים בכלל. pdfrx_engine מאותו ענף בפורק, עד שיתמזג espresso3389/pdfrx#732.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Many scanned Hebrew books come with an OCR text layer that stores each line's words left to right. Some older fonts go further and store every word's letters left to right too. Either way,
loadText/loadStructuredTextreturn the text backwards, so phrase search never matches and copied text comes out reversed.This adds
PdfHebrewTextNormalizer, opt-in throughPdfrx.normalizeHebrewText(defaultfalse).How it decides. It uses the character geometry, not guesswork:
Legacy encodings. It also decodes Windows-1255 Hebrew that a font without a ToUnicode map exposes as Latin-1, or as Mac Roman. This only happens on a page where those characters outnumber the Latin letters, so French or German text is never touched.
Character boxes are permuted together with the text, so indices into the result keep addressing the right glyphs; highlighting and selection keep working. The PDFium backend runs the normalizer on the worker.
Measured on real books
Unit tests cover logical text, word-order reversal, full visual order, points, Latin-1 decoding, accented Latin text, and the page-majority rule.