Repository navigation
feat(perfs): add VLM-based picture/chart extraction to RAG eval pipeline - #57
Merged
Merged
Conversation
Contributor
|
✅ DCO Check Passed Thanks @ceberam, all your commits are properly signed off. 🎉 |
Contributor
Merge Protections🟢 Merge protection satisfied — ready to merge. Show 1 satisfied protection🟢 Enforce conventional commitMake sure that we follow https://www.conventionalcommits.org/en/v1.0.0/
|
Upgrade 'uv.lock' dependencies Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Adds a picture_description flag (default: false) that enables per-image
captioning in Step 1 using the Granite vision model
(ibm-granite/granite-vision-3.3-2b) via Docling's
PictureDescriptionVlmEngineOptions.from_preset('granite_vision').
- AgenticRAGEvaluator.__init__ gains a picture_description: bool = False
parameter, stored as self.picture_description
- step1_convert_pdfs_to_docling conditionally sets
pdf_options.do_picture_description and
pdf_options.picture_description_options on the shared PdfPipelineOptions
before the converter is built
- --picture-description CLI flag wired up in main() following the same
pattern as --page-level
- picture_description: false added to agentic_rag_eval_config.yaml with
a cost / VLM availability warning
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
The PDF pipeline already handles heading-level detection via
HeadingHierarchyOptions (enabled in step 1) and the _hierarchize()
call that follows conversion, making the LLM-based DoclingEditingAgent
heading-fix step redundant.
- Remove step2_fix_heading_levels and the DoclingEditingAgent /
SectionHeaderItem imports it depended on
- Renumber: old step 3 (enrich) → step 2, old step 4 (evaluate) → step 3
- Rename directory slots accordingly:
step2_enriched (was step2_hierarchical / step3_enriched)
step3_evaluation (was step4_evaluation)
- Both enrich helpers (_enrich_element_level, _enrich_page_level) now
read from step1_dir and write to step2_dir
- step3_evaluate_rag reads enriched docs from step2_dir and writes
results to step3_dir
- CLI --step now accepts {1,2,3}; all help strings, log messages, and
docstring cross-references updated to match
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
…anguage from dataset name - Extract PDF pipeline setup into module-level _create_document_converter(), with CUDA detection and ThreadedStandardPdfPipeline support - Add _FRENCH_DATASETS frozenset and _ocr_lang_for_dataset() helper that maps vidore_v3_energy / vidore_v3_finance_fr / vidore_v3_physics to 'iso:fr' and all other datasets to 'english' - step1_convert_pdfs_to_docling now calls _ocr_lang_for_dataset(self.dataset_path) and logs the resolved language before building the converter Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
…N in Step 1 - generate_page_images = True in the PDF pipeline so page PNGs are available on the document object after conversion - Save .dclx archive first; save_as_doclang_archive packs the page PNGs into pages/<n>.png inside the OPC zip - Clear page.image on all pages before save_as_json so the JSON stays lightweight — page images are not needed for the RAG pipeline (PageItem.image is Optional[ImageRef], so None is always valid) Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
…ation Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
…Options with summary-only mode Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
…entences Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
…tion Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
ceberam
force-pushed
the
dev/rag-tree-opt
branch
from
September 21, 2026 11:04
2b60314 to
b0530f5
Compare
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR refactors the agentic RAG evaluation pipeline (
perfs/agentic_rag_eval.py) to streamline the conversion step, drop a redundant heading-fix step, and introduce new VLM-powered enrichments during PDF conversion. Output directories are now namespaced by dataset name to make multi-dataset runs self-contained and easier to manage.What changed
Pipeline restructuring (4 steps → 3 steps)
The previous four-step pipeline included a dedicated Step 2 that used
DoclingEditingAgentto correct flat heading hierarchies in converted documents. This step has been removed — heading hierarchy detection is now handled natively by Docling's conversion pipeline viaHeadingHierarchyOptions(enabled=True). The steps are renumbered accordingly:Richer, configurable PDF conversion (Step 1)
The converter is no longer a bare
DocumentConverter(). A dedicated_create_document_converter()factory now builds a fully configured converter with:HeadingHierarchyOptions.--picture-description/picture_description: truein config): captions every image usingibm-granite/granite-vision-4.1-4b.--chart-extraction/chart_extraction: truein config): produces a natural-language summary for each detected chart using the Granite Vision v4 model inchart2summarymode.Conversion errors are now handled gracefully: soft failures (
PARTIAL_SUCCESS) are logged as warnings instead of raising, and a per-run summary counter (succeeded / partial / failed) is printed at the end of the step.Dual output format in Step 1
Each converted document is now saved in two forms:
.dclxarchive that bundles the fullDoclingDocumenttogether with page images — useful for visual inspection and downstream VLM tasks..jsonfile with page images stripped out — the format consumed by Steps 2 and 3.Dataset-namespaced output directories
All step output directories are now scoped under a subdirectory named after the dataset:
This makes it safe to run the pipeline on multiple datasets pointing at the same
outputroot without outputs colliding.Dependency additions
The
evaluationextras group gains:ocrmac(macOS only) for OCR engine selection on Apple Silicon.nemotron-ocr(Linux x86-64, Python 3.12 only) for GPU-accelerated OCR.transformersrange as a temporary workaround for a Docling compatibility issue.