Skip to content

feat(web-fetch): extract text from PDF documents - #12

Merged
ZhiXiao-Lin merged 1 commit into
mainfrom
recovery/web-fetch-pdf
Jul 19, 2026
Merged

ZhiXiao-Lin merged 1 commit into
mainfrom
recovery/web-fetch-pdf

Conversation

@ZhiXiao-Lin

Copy link
Copy Markdown

Summary

  • detect PDF responses by media type or file signature and extract their text in web_fetch
  • run PDF parsing off Tokio worker threads and return clear errors for malformed or image-only documents
  • expose normalized content type and document kind in tool metadata while preserving HTML redirect handling
  • add in-memory PDF coverage for detection, extraction, malformed files, and empty documents

Validation

  • cargo fmt --all -- --check
  • cargo clippy --workspace --lib --bins -- -D warnings
  • cargo test --workspace
  • cargo test --workspace --all-features --lib
  • git diff --check

@ZhiXiao-Lin
ZhiXiao-Lin merged commit 6f43697 into main Jul 19, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants