Skip to content

feat(pdf): expose bounded PDFium page inventory and source-layout rasters #4

Description

@ZhiXiao-Lin

Problem

A3S Parser cannot truthfully admit PDF pages into its Visual, Balanced, Planned, or Deep profiles because the released Office Rust boundary has no page-level PDF contract.

Reproduced on Office main at ac094d58bf04b6710cebc58da86c9e1eca680a7b and the Parser-pinned revision 7903df9b58bf979cfeccdf4f66d9041434542a91:

  • DocumentKind has only Word, Spreadsheet, and Presentation.
  • NativeOfficeUnitLocator has Document, Worksheet, and Slide, but no Page.
  • NativeOfficeLayoutRenderer has no concrete PDF provider.
  • NativeOfficePptxImageLayoutRenderer correctly returns typed unsupported results outside its exact PNG-slide contract.
  • Consequently A3sOfficeBackend::support(DocumentFormat::Pdf) is empty and Parser must reject PDF before creating run state.

This is the PDF follow-up to #1. The semantic screenshot path is not acceptable source-layout evidence.

Required contract

Expose an object-safe, Send + Sync, browser-neutral Office API backed by PDFium that:

  • inventories every PDF page under an explicit hard unit bound;
  • uses a one-based page locator with stable page identity and canonical path;
  • binds immutable source size and SHA-256 before and after inventory/render;
  • reports media box/crop box, effective rotation, physical surface, and page pixel dimensions;
  • renders exactly one selected page under explicit deadline and output-byte limits;
  • returns a source-layout receipt with PNG SHA-256, size, dimensions, rotation, PDFium engine version/binary identity, DPI, viewport, locale/timezone, and deterministic render-profile SHA-256;
  • performs no external fetch and grants no Browser or filesystem authority to the Parser agent;
  • distinguishes corrupt, encrypted/password-required, unsupported-feature, source-mutated, timeout, page-limit, and output-limit failures with stable typed codes.

The core contract must not require a3s-use-browser or chromiumoxide in Parser. A PDFium implementation may live behind an injected Office provider or an Office-owned optional runtime boundary.

Acceptance criteria

  • A generated two-page fixture inventories pages 1 and 2 in stable order and rejects page 0, page 3, duplicate/reordered/foreign locators, and truncated inventories.
  • Rendering page 1 never includes page 2 pixels and vice versa; repeated renders have identical profile and pixel identities.
  • Normal, scanned/image-only, non-default crop box, and 0/90/180/270-degree page fixtures retain exact canvas geometry.
  • Encrypted, corrupt, zero-page, over-unit-limit, over-output-limit, timeout, and source-mutation cases fail with the required stable codes before publication.
  • Output is staged privately, atomically published without clobbering, rehashed from disk, and removed or left unpublished after failure.
  • Parser integration tests prove governed render → OCR → canonical canvas/region output, frontend overlay projection, transfer-policy enforcement, and full-process backend/model-free resume.
  • Parser dependency CI continues to reject a3s-use-browser and chromiumoxide.
  • Office and Parser documentation advertise PDF only after the full fixture and cross-platform gates pass.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions