DEMO (do not merge): read documents with AI_COMPLETE on claude-opus-5, unparsed - #4409
Draft
sfc-gh-hjayakumar wants to merge 2 commits into
Draft
sfc-gh-hjayakumar wants to merge 2 commits into
sfc-gh-hjayakumar wants to merge 2 commits into
Conversation
Demo branch. Three default changes so that supplying a schema gets the document
read as a file by AI_COMPLETE on claude-opus-5, rather than parsed to text and
handed to AI_EXTRACT.
Engine default ai_extract -> ai_complete. AI_EXTRACT flattens tabular and
clause-structured content instead of extracting it, and its schema path cannot
express the verbatim multi-paragraph answers clause-level review asks for. The
AI_EXTRACT path stays reachable by explicit option, where its per-field confidence
scores still have no AI_COMPLETE equivalent.
Model default claude-4-sonnet -> claude-opus-5.
Skip AI_PARSE_DOCUMENT when a schema is present and the engine is AI_COMPLETE, and
pass the document as a FILE instead. Parsing flattens the document, and layout is
information: table columns, and the printed section numbers clause extraction is
asked to quote, survive in the page image and not in flattened text. It also
removes a Cortex call per document along with its own failure and latency modes.
Conditional on there being something to extract -- with no schema the parsed text
is the read's only output, so skipping the parse there would return a DataFrame
with nothing in it. An explicit parse_mode always wins, and the default is filled
before validate() so validation sees the parse_mode the read will use.
Send model_parameters, which were previously absent entirely. AI_COMPLETE then
applied its own max_tokens default of 4096, low enough to truncate a document
answering more than a handful of fields -- and truncation discards the whole
response rather than the overflow, surfacing as a bare internal error that is
indistinguishable from a server fault, or occasionally as success with a silently
partial answer.
The ceiling is per-model, so one shared constant is unsafe: claude-opus-5's
ceiling sent to claude-4-sonnet fails every call. Only models whose limit has been
confirmed are listed; anything else sends no max_tokens and gets the server
default, which is always accepted. claude-opus-5 is capped at 65536 rather than
its 128000 maximum, because a larger ceiling does not rescue a document whose
answer cannot fit in one response -- it only lets the call run longer before
failing. temperature is 0 for every model: the same document and schema should not
yield different answers on consecutive reads.
Known gap: a flat {field: question} map is an AI_EXTRACT-only shape and is passed
through unchanged, so AI_COMPLETE rejects it with "invalid response format
object". A JSON Schema is required until schema preparation can convert one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #4409 +/- ##
===========================================
- Coverage 95.47% 80.60% -14.87%
===========================================
Files 176 175 -1
Lines 45271 45205 -66
Branches 7759 7767 +8
===========================================
- Hits 43222 36439 -6783
- Misses 1269 7251 +5982
- Partials 780 1515 +735 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
A response_format like {"vendor": "Who issued this invoice?"} is AI_EXTRACT's own
shape. AI_COMPLETE requires {"type": "json", "schema": {...}} and rejects the bare
map with "invalid response format object", so passing it through unchanged was only
safe while AI_EXTRACT was the default engine. Now that it is not, a question map --
the form a caller reaches for when they have questions rather than a type
contract -- fails outright.
Rewrite it for AI_COMPLETE instead: each field becomes a string property whose
description is the question, which is how a model consumes it under either engine.
AI_EXTRACT continues to receive the caller's shape exactly as given, so its
behaviour is unchanged.
The array form is handled too, and keeps its questions. Both ["name", "question"]
pairs and "name: question" strings carry the question inside the item, and taking
only the field name would discard the only instruction the model gets for that
field.
Every converted field is typed string, because a question implies no type. A field
whose answer should be a list or a number still needs a real JSON Schema -- the
conversion makes the call succeed, it cannot add a contract the input never
expressed. field_types is therefore left unset rather than populated with
StringType: a question is not grounds to start casting output columns, least of all
on the AI_EXTRACT path, which shares field_types.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Which Jira issue is this PR addressing?
None yet. Follow-on work to the document reader added in SNOW-4057099 (epic SNOW-4057090).
A Jira issue needs filing before any of this is proposed as a real default.
Pre-review checklist:
Unchecked on purpose — see "Not done" below.
What the change does
Supplying a schema to
DataFrameReader._documents()currently means: parse the document totext with
AI_PARSE_DOCUMENT, then hand that text toAI_EXTRACTwith no model and nomodel_parameters. This branch changes that to: hand the document itself toAI_COMPLETEonclaude-opus-5, with no parse step and an explicit output cap.Four default changes, all in two files:
extraction_engineai_extractai_complete_DEFAULT_EXTRACTION_MODELclaude-4-sonnetclaude-opus-5parse_mode="layout"ai_completemodel_parameterstemperature: 0+ a per-modelmax_tokensWhy
AI_EXTRACT→AI_COMPLETEAI_EXTRACTflattens tabular and clause-structured content rather than extracting it, and itsschema path cannot express the verbatim multi-paragraph answers that clause-level review asks
for. It also silently ignores the
modeloption:extract_with_ai_extractnever passes one,so on the old default path there was no way to select a model at all.
AI_EXTRACTremains reachable with.option("extraction_engine", "ai_extract"). Note thatits per-field confidence scores (
EXTRACTION_SCORES_COLUMN) have noAI_COMPLETEequivalent,so that column disappears from the default path.
Why skip the parse
Parsing flattens the document, and layout is information: table columns, and the printed
section numbers clause extraction is usually asked to quote, survive in the page image and not
in flattened text. Skipping it also removes a whole Cortex call per document along with its own
failure and latency modes.
Conditional on there being something to extract. With no schema,
AI_COMPLETEhas nothing todo and the parsed text is the read's only output — skipping the parse there would hand back a
DataFrame with nothing in it. An explicit
parse_modealways wins, and the default is filledbefore
validate()so validation sees theparse_modethe read will actually use.Why
model_parameters, and why per-modelNo
model_parameterswere sent, soAI_COMPLETEapplied its ownmax_tokensdefault of 4096.That is low enough to truncate a document answering more than a handful of fields, and
truncation discards the whole response rather than the overflow — surfacing as a bare
internal errorindistinguishable from a server fault, or occasionally as success with asilently partial answer, so status alone is not a quality signal.
The ceiling is per-model, which makes a single shared constant unsafe. Asking each model
directly:
claude-4-sonnet400 'max_tokens parameter exceeds the maximum possible value (32000)'claude-opus-5So a value one model accepts makes every call to the other fail outright. Only models whose
limit has been confirmed are listed; anything else sends no
max_tokensand gets the serverdefault, which is always accepted — adding a model later cannot break it.
claude-opus-5is capped at 65536 rather than its 128000 maximum, because a larger ceilingdoes not rescue a document whose answer cannot fit in one response. It only lets the call run
longer before failing.
temperature: 0for every model: the same document and schema should not yield differentanswers on consecutive reads, which is a correctness property for an extraction API rather
than a preference.
Verified
The planned call is now:
No
AI_PARSE_DOCUMENT, noAI_EXTRACT. A live extraction returns a typedDOUBLEfor acurrency field and a real nested
ARRAYfor line items. The 6 existingtests/unit/test_document_reader.pytests pass unchanged.Flat question maps
A
response_formatlike{"vendor": "Who issued this invoice?"}isAI_EXTRACT's own shape.AI_COMPLETErequires{"type": "json", "schema": {...}}and rejects the bare map withinvalid response format object, so passing it through unchanged was only safe whileAI_EXTRACTwas the default engine. The second commit rewrites it forAI_COMPLETE: eachfield becomes a string property whose
descriptionis the question, which is how a modelconsumes it under either engine.
AI_EXTRACTstill receives the caller's shape exactly asgiven, so its behaviour is unchanged. The array forms (
["name", "question"]pairs and"name: question"strings) are handled too, and keep their questions rather than droppingthem.
What the conversion cannot do: every converted field is typed
string, because a questionimplies no type. A field whose answer should be a list comes back as a JSON-encoded string
rather than an
ARRAY, and nullability is never declared. Both need a real JSON Schema -- theconversion makes the call succeed, it cannot add a contract the input never expressed.
field_typesis deliberately left unset rather than populated withStringType, since aquestion is not grounds to start casting output columns, least of all on the
AI_EXTRACTpathwhich shares that field.
Not done, deliberately
_documents()isprivate and the change alters the payload sent to Cortex rather than adding public API
surface, but someone who knows those requirements should confirm before this is taken
seriously.
claude-opus-5andclaude-4-sonnetceilings are known. Every other model fallsback to the server default.
raising the cap further does not fix it.