Skip to content

Support native variant_explode and variant_explode_outer #5432

Description

@peterxcli

What is the problem the feature request solves?

Spark 4.x provides variant_explode and variant_explode_outer to turn a Variant object or array into rows. Comet does not support these generators, so queries must leave native execution:

SELECT id, e.* FROM t LATERAL VIEW variant_explode(v) e AS pos, key, value;
SELECT id, e.* FROM t LATERAL VIEW variant_explode_outer(v) e AS pos, key, value;

Spark defines the generated columns as pos INT NOT NULL, key STRING, and value VARIANT NOT NULL, and implements object/array iteration in VariantExplode. The exact SQL names are registered as generators in Spark's 4.1.3 function registry.

Describe the potential solution

Add a Spark-4-specific generator path that:

  • iterates arrays by index and objects by their stored field order;
  • emits Spark-compatible pos, nullable key, and Variant value columns;
  • preserves the Variant result Field marker and canonical [value, metadata] storage for each emitted value;
  • matches Spark behavior for scalar values, Variant null, SQL NULL, empty containers, and the outer form; and
  • keeps unrelated Variant-bearing generators/operators on explicit fallback.

Add focused parity and native-plan tests for objects, arrays, nested values, empty containers, scalars, Variant/SQL nulls, Unicode keys, aliases, lateral-view syntax, and outer versus non-outer cardinality.

Additional context

Related work: #4295, #5407, #5425, and the general generator execution path.

Non-goals: Variant mutation, subfield pruning or pushdown, shuffle/spill, C2R, writing, Python transport, and Iceberg-specific work.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions