Skip to content

feat(perfs): add VLM-based picture/chart extraction to RAG eval pipeline - #57

Merged
ceberam merged 18 commits into
mainfrom
dev/rag-tree-opt
Sep 21, 2026
Merged

ceberam merged 18 commits into
mainfrom
dev/rag-tree-opt

Conversation

@ceberam

@ceberam ceberam commented Sep 21, 2026

Copy link
Copy Markdown
Member

Summary

This PR refactors the agentic RAG evaluation pipeline (perfs/agentic_rag_eval.py) to streamline the conversion step, drop a redundant heading-fix step, and introduce new VLM-powered enrichments during PDF conversion. Output directories are now namespaced by dataset name to make multi-dataset runs self-contained and easier to manage.

What changed

Pipeline restructuring (4 steps → 3 steps)

The previous four-step pipeline included a dedicated Step 2 that used DoclingEditingAgent to correct flat heading hierarchies in converted documents. This step has been removed — heading hierarchy detection is now handled natively by Docling's conversion pipeline via HeadingHierarchyOptions(enabled=True). The steps are renumbered accordingly:

Before After
Step 1 – Convert PDFs Step 1 – Convert PDFs (richer, with hierarchy)
Step 2 – Fix heading levels (agent) Step 2 – Enrich with summaries
Step 3 – Enrich with summaries Step 3 – Evaluate RAG
Step 4 – Evaluate RAG (removed)

Richer, configurable PDF conversion (Step 1)

The converter is no longer a bare DocumentConverter(). A dedicated _create_document_converter() factory now builds a fully configured converter with:

  • Native heading hierarchy detection via HeadingHierarchyOptions.
  • Platform-aware OCR selection: Nemotron OCR on Linux with the correct language (English or French, inferred from the dataset directory name); auto-selection on macOS/other platforms.
  • CUDA acceleration: when a GPU is available, the threaded pipeline is used with increased batch sizes for OCR, layout, and table detection.
  • Optional picture description (--picture-description / picture_description: true in config): captions every image using ibm-granite/granite-vision-4.1-4b.
  • Optional chart extraction (--chart-extraction / chart_extraction: true in config): produces a natural-language summary for each detected chart using the Granite Vision v4 model in chart2summary mode.

Conversion errors are now handled gracefully: soft failures (PARTIAL_SUCCESS) are logged as warnings instead of raising, and a per-run summary counter (succeeded / partial / failed) is printed at the end of the step.

Dual output format in Step 1

Each converted document is now saved in two forms:

  1. A .dclx archive that bundles the full DoclingDocument together with page images — useful for visual inspection and downstream VLM tasks.
  2. A lean .json file with page images stripped out — the format consumed by Steps 2 and 3.

Dataset-namespaced output directories

All step output directories are now scoped under a subdirectory named after the dataset:

output/
  step1_converted/<dataset_name>/
  step2_enriched/<dataset_name>/
  step3_evaluation/<dataset_name>/

This makes it safe to run the pipeline on multiple datasets pointing at the same output root without outputs colliding.

Dependency additions

The evaluation extras group gains:

  • ocrmac (macOS only) for OCR engine selection on Apple Silicon.
  • nemotron-ocr (Linux x86-64, Python 3.12 only) for GPU-accelerated OCR.
  • A pinned transformers range as a temporary workaround for a Docling compatibility issue.

@github-actions

Copy link
Copy Markdown
Contributor

✅ DCO Check Passed

Thanks @ceberam, all your commits are properly signed off. 🎉

@mergify

mergify Bot commented Sep 21, 2026

Copy link
Copy Markdown
Contributor

Merge Protections

🟢 Merge protection satisfied — ready to merge.

Show 1 satisfied protection

🟢 Enforce conventional commit

Make sure that we follow https://www.conventionalcommits.org/en/v1.0.0/

  • title ~= ^(fix|feat|docs|style|refactor|perf|test|build|ci|chore|revert)(?:\(.+\))?(!)?:

PeterStaar-IBM
PeterStaar-IBM previously approved these changes Sep 21, 2026

@PeterStaar-IBM PeterStaar-IBM left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm!

Upgrade 'uv.lock' dependencies

Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Adds a picture_description flag (default: false) that enables per-image
captioning in Step 1 using the Granite vision model
(ibm-granite/granite-vision-3.3-2b) via Docling's
PictureDescriptionVlmEngineOptions.from_preset('granite_vision').

- AgenticRAGEvaluator.__init__ gains a picture_description: bool = False
  parameter, stored as self.picture_description
- step1_convert_pdfs_to_docling conditionally sets
  pdf_options.do_picture_description and
  pdf_options.picture_description_options on the shared PdfPipelineOptions
  before the converter is built
- --picture-description CLI flag wired up in main() following the same
  pattern as --page-level
- picture_description: false added to agentic_rag_eval_config.yaml with
  a cost / VLM availability warning

Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
The PDF pipeline already handles heading-level detection via
HeadingHierarchyOptions (enabled in step 1) and the _hierarchize()
call that follows conversion, making the LLM-based DoclingEditingAgent
heading-fix step redundant.

- Remove step2_fix_heading_levels and the DoclingEditingAgent /
  SectionHeaderItem imports it depended on
- Renumber: old step 3 (enrich) → step 2, old step 4 (evaluate) → step 3
- Rename directory slots accordingly:
    step2_enriched  (was step2_hierarchical / step3_enriched)
    step3_evaluation (was step4_evaluation)
- Both enrich helpers (_enrich_element_level, _enrich_page_level) now
  read from step1_dir and write to step2_dir
- step3_evaluate_rag reads enriched docs from step2_dir and writes
  results to step3_dir
- CLI --step now accepts {1,2,3}; all help strings, log messages, and
  docstring cross-references updated to match

Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
…anguage from dataset name

- Extract PDF pipeline setup into module-level _create_document_converter(),
  with CUDA detection and ThreadedStandardPdfPipeline support
- Add _FRENCH_DATASETS frozenset and _ocr_lang_for_dataset() helper that maps
  vidore_v3_energy / vidore_v3_finance_fr / vidore_v3_physics to 'iso:fr'
  and all other datasets to 'english'
- step1_convert_pdfs_to_docling now calls _ocr_lang_for_dataset(self.dataset_path)
  and logs the resolved language before building the converter

Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
…N in Step 1

- generate_page_images = True in the PDF pipeline so page PNGs are
  available on the document object after conversion
- Save .dclx archive first; save_as_doclang_archive packs the page
  PNGs into pages/<n>.png inside the OPC zip
- Clear page.image on all pages before save_as_json so the JSON stays
  lightweight — page images are not needed for the RAG pipeline
  (PageItem.image is Optional[ImageRef], so None is always valid)

Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
…ation

Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
…Options with summary-only mode

Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
…entences

Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
…tion

Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
@codecov

codecov Bot commented Sep 21, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@PeterStaar-IBM PeterStaar-IBM left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm!

@ceberam
ceberam merged commit 9c754cf into main Sep 21, 2026
11 checks passed
@ceberam
ceberam deleted the dev/rag-tree-opt branch September 21, 2026 12:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants