Skip to content

DEMO (do not merge): read documents with AI_COMPLETE on claude-opus-5, unparsed - #4409

Draft
sfc-gh-hjayakumar wants to merge 2 commits into
mainfrom
hjayakumar-qinyi-demo
Draft

sfc-gh-hjayakumar wants to merge 2 commits into
mainfrom
hjayakumar-qinyi-demo

Conversation

@sfc-gh-hjayakumar

@sfc-gh-hjayakumar sfc-gh-hjayakumar commented Oct 1, 2026 •

Copy link
Copy Markdown

Draft, and not proposed for merge as-is. This is a demo branch: the changes are
hardcoded defaults marked DEMO BRANCH in the source, and it deliberately carries no
tests. It exists so the document reader can be run end-to-end on real customer documents
with the configuration that reads them most accurately. Opening it as a PR for the diff and
the discussion, not to land.

  1. Which Jira issue is this PR addressing?

    None yet. Follow-on work to the document reader added in SNOW-4057099 (epic SNOW-4057090).
    A Jira issue needs filing before any of this is proposed as a real default.

  2. Pre-review checklist:

    • I am adding a new automated test(s) to verify correctness of my new code
    • I am adding new logging messages
    • I am adding a new telemetry message
    • I am adding new credentials
    • I am adding a new dependency
    • If this is a new feature/behavior, I'm adding the Local Testing parity changes.
    • I acknowledge that I have ensured my changes to be thread-safe.
    • If adding any arguments to public Snowpark APIs or creating new public Snowpark APIs, I acknowledge that I have ensured my changes include AST support.

    Unchecked on purpose — see "Not done" below.

  3. What the change does

Supplying a schema to DataFrameReader._documents() currently means: parse the document to
text with AI_PARSE_DOCUMENT, then hand that text to AI_EXTRACT with no model and no
model_parameters. This branch changes that to: hand the document itself to AI_COMPLETE on
claude-opus-5, with no parse step and an explicit output cap.

Four default changes, all in two files:

before after
extraction_engine ai_extract ai_complete
_DEFAULT_EXTRACTION_MODEL claude-4-sonnet claude-opus-5
parse step always, parse_mode="layout" skipped when a schema is present and the engine is ai_complete
model_parameters not sent at all temperature: 0 + a per-model max_tokens

Why AI_EXTRACT → AI_COMPLETE

AI_EXTRACT flattens tabular and clause-structured content rather than extracting it, and its
schema path cannot express the verbatim multi-paragraph answers that clause-level review asks
for. It also silently ignores the model option: extract_with_ai_extract never passes one,
so on the old default path there was no way to select a model at all.

AI_EXTRACT remains reachable with .option("extraction_engine", "ai_extract"). Note that
its per-field confidence scores (EXTRACTION_SCORES_COLUMN) have no AI_COMPLETE equivalent,
so that column disappears from the default path.

Why skip the parse

Parsing flattens the document, and layout is information: table columns, and the printed
section numbers clause extraction is usually asked to quote, survive in the page image and not
in flattened text. Skipping it also removes a whole Cortex call per document along with its own
failure and latency modes.

Conditional on there being something to extract. With no schema, AI_COMPLETE has nothing to
do and the parsed text is the read's only output — skipping the parse there would hand back a
DataFrame with nothing in it. An explicit parse_mode always wins, and the default is filled
before validate() so validation sees the parse_mode the read will actually use.

Why model_parameters, and why per-model

No model_parameters were sent, so AI_COMPLETE applied its own max_tokens default of 4096.
That is low enough to truncate a document answering more than a handful of fields, and
truncation discards the whole response rather than the overflow — surfacing as a bare
internal error indistinguishable from a server fault, or occasionally as success with a
silently partial answer, so status alone is not a quality signal.

The ceiling is per-model, which makes a single shared constant unsafe. Asking each model
directly:

model maximum accepted evidence
claude-4-sonnet 32000 24576 accepted; 32768 → 400 'max_tokens parameter exceeds the maximum possible value (32000)'
claude-opus-5 128000 128000 accepted; 200000 rejected

So a value one model accepts makes every call to the other fail outright. Only models whose
limit has been confirmed are listed; anything else sends no max_tokens and gets the server
default, which is always accepted — adding a model later cannot break it.

claude-opus-5 is capped at 65536 rather than its 128000 maximum, because a larger ceiling
does not rescue a document whose answer cannot fit in one response. It only lets the call run
longer before failing.

temperature: 0 for every model: the same document and schema should not yield different
answers on consecutive reads, which is a correctness property for an extraction API rather
than a preference.

Verified

The planned call is now:

model => 'claude-opus-5',
predicate => 'Extract the requested fields from this document.',
file => "SOURCE_FILE",
model_parameters => {'temperature': 0, 'max_tokens': 65536},
response_format => {'type': 'json', 'schema': {...}}

No AI_PARSE_DOCUMENT, no AI_EXTRACT. A live extraction returns a typed DOUBLE for a
currency field and a real nested ARRAY for line items. The 6 existing
tests/unit/test_document_reader.py tests pass unchanged.

Flat question maps

A response_format like {"vendor": "Who issued this invoice?"} is AI_EXTRACT's own shape.
AI_COMPLETE requires {"type": "json", "schema": {...}} and rejects the bare map with
invalid response format object, so passing it through unchanged was only safe while
AI_EXTRACT was the default engine. The second commit rewrites it for AI_COMPLETE: each
field becomes a string property whose description is the question, which is how a model
consumes it under either engine. AI_EXTRACT still receives the caller's shape exactly as
given, so its behaviour is unchanged. The array forms (["name", "question"] pairs and
"name: question" strings) are handled too, and keep their questions rather than dropping
them.

What the conversion cannot do: every converted field is typed string, because a question
implies no type. A field whose answer should be a list comes back as a JSON-encoded string
rather than an ARRAY, and nullability is never declared. Both need a real JSON Schema -- the
conversion makes the call succeed, it cannot add a contract the input never expressed.
field_types is deliberately left unset rather than populated with StringType, since a
question is not grounds to start casting output columns, least of all on the AI_EXTRACT path
which shares that field.

Not done, deliberately

  • No tests for the new behaviour. This is a demo branch.
  • Thread-safety, Local Testing parity and AST support not assessed. _documents() is
    private and the change alters the payload sent to Cortex rather than adding public API
    surface, but someone who knows those requirements should confirm before this is taken
    seriously.
  • Only claude-opus-5 and claude-4-sonnet ceilings are known. Every other model falls
    back to the server default.
  • No document splitting. A document whose answer exceeds one response still fails, and
    raising the cap further does not fix it.

Demo branch. Three default changes so that supplying a schema gets the document
read as a file by AI_COMPLETE on claude-opus-5, rather than parsed to text and
handed to AI_EXTRACT.

Engine default ai_extract -> ai_complete. AI_EXTRACT flattens tabular and
clause-structured content instead of extracting it, and its schema path cannot
express the verbatim multi-paragraph answers clause-level review asks for. The
AI_EXTRACT path stays reachable by explicit option, where its per-field confidence
scores still have no AI_COMPLETE equivalent.

Model default claude-4-sonnet -> claude-opus-5.

Skip AI_PARSE_DOCUMENT when a schema is present and the engine is AI_COMPLETE, and
pass the document as a FILE instead. Parsing flattens the document, and layout is
information: table columns, and the printed section numbers clause extraction is
asked to quote, survive in the page image and not in flattened text. It also
removes a Cortex call per document along with its own failure and latency modes.
Conditional on there being something to extract -- with no schema the parsed text
is the read's only output, so skipping the parse there would return a DataFrame
with nothing in it. An explicit parse_mode always wins, and the default is filled
before validate() so validation sees the parse_mode the read will use.

Send model_parameters, which were previously absent entirely. AI_COMPLETE then
applied its own max_tokens default of 4096, low enough to truncate a document
answering more than a handful of fields -- and truncation discards the whole
response rather than the overflow, surfacing as a bare internal error that is
indistinguishable from a server fault, or occasionally as success with a silently
partial answer.

The ceiling is per-model, so one shared constant is unsafe: claude-opus-5's
ceiling sent to claude-4-sonnet fails every call. Only models whose limit has been
confirmed are listed; anything else sends no max_tokens and gets the server
default, which is always accepted. claude-opus-5 is capped at 65536 rather than
its 128000 maximum, because a larger ceiling does not rescue a document whose
answer cannot fit in one response -- it only lets the call run longer before
failing. temperature is 0 for every model: the same document and schema should not
yield different answers on consecutive reads.

Known gap: a flat {field: question} map is an AI_EXTRACT-only shape and is passed
through unchanged, so AI_COMPLETE rejects it with "invalid response format
object". A JSON Schema is required until schema preparation can convert one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@codecov-commenter

codecov-commenter commented Oct 1, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 18.75000% with 26 lines in your changes missing coverage. Please review.
✅ Project coverage is 80.60%. Comparing base (b1173b1) to head (53b0c6a).

Files with missing lines Patch % Lines
...lake/snowpark/_internal/document_reader_options.py 12.50% 21 Missing ⚠️
...rc/snowflake/snowpark/_internal/document_reader.py 37.50% 5 Missing ⚠️
Additional details and impacted files
@@             Coverage Diff             @@
##             main    #4409       +/-   ##
===========================================
- Coverage   95.47%   80.60%   -14.87%     
===========================================
  Files         176      175        -1     
  Lines       45271    45205       -66     
  Branches     7759     7767        +8     
===========================================
- Hits        43222    36439     -6783     
- Misses       1269     7251     +5982     
- Partials      780     1515      +735     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

A response_format like {"vendor": "Who issued this invoice?"} is AI_EXTRACT's own
shape. AI_COMPLETE requires {"type": "json", "schema": {...}} and rejects the bare
map with "invalid response format object", so passing it through unchanged was only
safe while AI_EXTRACT was the default engine. Now that it is not, a question map --
the form a caller reaches for when they have questions rather than a type
contract -- fails outright.

Rewrite it for AI_COMPLETE instead: each field becomes a string property whose
description is the question, which is how a model consumes it under either engine.
AI_EXTRACT continues to receive the caller's shape exactly as given, so its
behaviour is unchanged.

The array form is handled too, and keeps its questions. Both ["name", "question"]
pairs and "name: question" strings carry the question inside the item, and taking
only the field name would discard the only instruction the model gets for that
field.

Every converted field is typed string, because a question implies no type. A field
whose answer should be a list or a number still needs a real JSON Schema -- the
conversion makes the call succeed, it cannot add a contract the input never
expressed. field_types is therefore left unset rather than populated with
StringType: a question is not grounds to start casting output columns, least of all
on the AI_EXTRACT path, which shares field_types.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants