Problem
detect_drift() derives the monitored feature list from the reference dataset and verifies only that those features are present in production.
If the production dataset gains one or more unexpected columns, they are silently ignored.
Why it matters
Unexpected columns can indicate an upstream schema/version change. Ignoring them makes schema drift invisible even though the monitoring system should surface contract changes.
Suggested fix
Compute both sides of the schema diff:
- missing expected features
- unexpected production features
Either fail fast on unexpected features or record them explicitly in the drift report under a schema-drift section.
Acceptance criteria
- missing expected feature remains an explicit error
- unexpected production feature is surfaced rather than ignored
- target/known non-feature columns are handled deliberately
- regression test covers an added production-only column
Problem
detect_drift()derives the monitored feature list from the reference dataset and verifies only that those features are present in production.If the production dataset gains one or more unexpected columns, they are silently ignored.
Why it matters
Unexpected columns can indicate an upstream schema/version change. Ignoring them makes schema drift invisible even though the monitoring system should surface contract changes.
Suggested fix
Compute both sides of the schema diff:
Either fail fast on unexpected features or record them explicitly in the drift report under a schema-drift section.
Acceptance criteria