Skip to content

Drift detection silently ignores unexpected production features #4

Description

@Ayushdevo

Problem

detect_drift() derives the monitored feature list from the reference dataset and verifies only that those features are present in production.

If the production dataset gains one or more unexpected columns, they are silently ignored.

Why it matters

Unexpected columns can indicate an upstream schema/version change. Ignoring them makes schema drift invisible even though the monitoring system should surface contract changes.

Suggested fix

Compute both sides of the schema diff:

  • missing expected features
  • unexpected production features

Either fail fast on unexpected features or record them explicitly in the drift report under a schema-drift section.

Acceptance criteria

  • missing expected feature remains an explicit error
  • unexpected production feature is surfaced rather than ignored
  • target/known non-feature columns are handled deliberately
  • regression test covers an added production-only column

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions