From 65a0ebbb51ef7e3e0296313bd14d14f87347d4f5 Mon Sep 17 00:00:00 2001 From: Speculator55005 <50082482+fas89@users.noreply.github.com> Date: Mon, 5 Oct 2026 11:25:52 +0200 Subject: [PATCH] docs: state 0.7.5 stable and 0.7.6 preview, rewrite the specification against 0.7.5, and make every example validate The version and status model (0.7.5 latest stable, 0.7.6 preview); new 0.7.6 preview notes derived from the schema, with the authoritative bucket-policy warning; the 0.7.5 release note and changelog corrected; specification.md rewritten against 0.7.5, with an accurate, non-normative description of multi-file composition in place of the .fluid/ recommendation; anatomy, cheatsheet and minimal-contract examples that validate, and the complete Lake Formation admins warning; licence mentions that match LICENSE. Signed-off-by: Speculator55005 <50082482+fas89@users.noreply.github.com> --- README.md | 19 +- docs/.vuepress/config.ts | 20 +- docs/README.md | 14 +- docs/concepts/README.md | 2 +- docs/concepts/agentic-native.md | 2 +- docs/concepts/comparisons.md | 12 +- docs/concepts/forge-cli.md | 65 ++- docs/contributing/README.md | 102 +++-- docs/deck/README.md | 12 +- docs/examples/README.md | 176 +++++++- docs/guide/quickstart.md | 25 +- docs/how-to/README.md | 4 + docs/how-to/advanced.md | 163 ++++---- docs/how-to/airflow.md | 7 +- docs/how-to/build-patterns.md | 6 +- docs/how-to/datavault.md | 10 +- docs/how-to/dbt.md | 44 +- docs/how-to/mcp.md | 31 +- docs/how-to/source-aligned-data-product.md | 125 +++--- docs/how-to/source-aligned-kafka.md | 101 ++--- docs/releases/0.7.1.md | 225 +++++----- docs/releases/0.7.2.md | 29 +- docs/releases/0.7.3.md | 43 +- docs/releases/0.7.4.md | 16 +- docs/releases/0.7.5.md | 174 ++++++-- docs/releases/0.7.6.md | 293 +++++++++++++ docs/releases/README.md | 25 +- docs/schema/README.md | 15 +- docs/schema/anatomy.md | 198 +++++---- docs/schema/changelog.md | 137 +++++-- docs/schema/cheatsheet.md | 375 +++++++++-------- docs/schema/minimal-contract.md | 122 +++--- docs/schema/specification.md | 452 ++++++++------------- docs/schema/versions.md | 85 +++- 34 files changed, 1958 insertions(+), 1171 deletions(-) create mode 100644 docs/releases/0.7.6.md diff --git a/README.md b/README.md index 03f45aa..e7cc494 100644 --- a/README.md +++ b/README.md @@ -17,7 +17,7 @@ --- -FLUID is one YAML file that describes a data product end to end — **schema, build, orchestration, agentic governance, sovereignty, and semantics**. Write it once; validate it, compile it, and deploy it anywhere. It compiles to Bitol ODPS + ODCS for catalog interop via the reference compiler, [`forge-cli`](https://github.com/Agenticstiger/forge-cli). +FLUID is one YAML file that describes a data product end to end — **schema, build, orchestration, agentic governance, sovereignty, and semantics**. Write it once; validate it, compile it, and deploy it anywhere. It compiles to Bitol ODPS + ODCS for catalog interop via the reference implementation, [`forge-cli`](https://github.com/Agenticstiger/forge-cli). ## How the pieces fit @@ -61,14 +61,25 @@ exposes: ## JSON Schema -- **Latest:** `https://open-data-protocol.github.io/fluid/schema/fluid-schema-0.7.5.json` -- Use it in your editor — add this line to any `.fluid.yml`: +- **Latest stable: 0.7.5** — `https://open-data-protocol.github.io/fluid/schema/fluid-schema-0.7.5.json` +- **Preview: 0.7.6** — `https://open-data-protocol.github.io/fluid/schema/fluid-schema-0.7.6.json`. A contract uses it only by declaring `fluidVersion: "0.7.6"`, and it can still change before it is promoted to stable. See **[Stable and preview versions](https://open-data-protocol.github.io/fluid/schema/versions#stable-and-preview-versions)** and **[What's new in 0.7.6](https://open-data-protocol.github.io/fluid/releases/0.7.6)**. +- A document is validated against the schema of the `fluidVersion` it declares. To use a field, declare the version that added it — see **[Choosing `fluidVersion`](https://open-data-protocol.github.io/fluid/schema/versions#choosing-fluidversion)**. +- Editor support — add this line to a file that declares `fluidVersion: "0.7.5"` (use the URL matching the version the file declares): ```yaml # yaml-language-server: $schema=https://open-data-protocol.github.io/fluid/schema/fluid-schema-0.7.5.json ``` - All versions, diffs, and the generated HTML reference: **[Schema → Versions](https://open-data-protocol.github.io/fluid/schema/versions)**. -> **Note:** 0.7.5 ("Streaming Kafka → Iceberg Sink & Confluent Tableflow") is **additive and fully backward-compatible** with 0.7.4 — every valid 0.7.4 contract still validates. See **[What's New in 0.7.5](https://open-data-protocol.github.io/fluid/releases/0.7.5)**. +0.7.5 is additive over 0.7.4: every valid 0.7.4 contract stays valid when it declares `"0.7.5"`. See **[What's New in 0.7.5](https://open-data-protocol.github.io/fluid/releases/0.7.5)**. + +## Validate a contract + +With a JSON Schema Draft 2020-12 validator, against the schema of the declared version (validators that compile `pattern` as ECMA-262, as JavaScript ones do, should read the [known interoperability issue](https://open-data-protocol.github.io/fluid/schema/specification#known-interoperability-issue-one-pattern-is-not-an-ecma-262-regular-expression) first). Or with the reference implementation, [`data-product-forge`](https://open-data-protocol.github.io/fluid/concepts/forge-cli): + +```bash +pip install data-product-forge +fluid validate contract.fluid.yaml +``` ## Build the docs locally diff --git a/docs/.vuepress/config.ts b/docs/.vuepress/config.ts index 499edca..a6b0134 100644 --- a/docs/.vuepress/config.ts +++ b/docs/.vuepress/config.ts @@ -75,6 +75,7 @@ export default defineUserConfig({ { text: 'Core Principles', link: '/concepts/principles' }, { text: 'Agentic-Native Layer', link: '/concepts/agentic-native' }, { text: 'FLUID vs ODCS / ODPS', link: '/concepts/comparisons' }, + { text: 'Reference Implementation', link: '/concepts/forge-cli' }, ], }, { @@ -82,16 +83,28 @@ export default defineUserConfig({ children: [ { text: 'Anatomy', link: '/schema/anatomy' }, { text: 'Cheatsheet', link: '/schema/cheatsheet' }, - { text: 'Full Specification', link: '/schema/specification' }, + { text: 'Minimal Contract', link: '/schema/minimal-contract' }, + { text: 'Specification', link: '/schema/specification' }, { text: 'Versions', link: '/schema/versions' }, - { text: 'JSON Schema 0.7.5 ↗', link: 'https://open-data-protocol.github.io/fluid/schema/fluid-schema-0.7.5.json', target: '_blank' }, - { text: 'Reference (HTML) ↗', link: 'https://open-data-protocol.github.io/fluid/specs/0.7.5/fluid-spec.html', target: '_blank' }, + { text: 'Changelog', link: '/schema/changelog' }, + { text: 'JSON Schema 0.7.5 (stable) ↗', link: 'https://open-data-protocol.github.io/fluid/schema/fluid-schema-0.7.5.json', target: '_blank' }, + { text: 'Reference 0.7.5 (HTML) ↗', link: 'https://open-data-protocol.github.io/fluid/specs/0.7.5/fluid-spec.html', target: '_blank' }, + { text: 'Preview: 0.7.6', link: '/releases/0.7.6' }, ], }, { text: 'Examples', link: '/examples/' }, { text: 'How-to', link: '/how-to/' }, { text: "What's New", link: '/releases/' }, { text: 'Deck', link: '/deck/' }, + { + text: 'Project', + children: [ + { text: 'Contributing', link: '/contributing/' }, + { text: 'Governance ↗', link: 'https://github.com/open-data-protocol/fluid/blob/main/GOVERNANCE.md', target: '_blank' }, + { text: 'Conformance Corpus ↗', link: 'https://github.com/open-data-protocol/fluid/blob/main/tests/README.md', target: '_blank' }, + { text: 'Vision', link: '/vision/' }, + ], + }, { text: 'GitHub', link: 'https://github.com/open-data-protocol/fluid' }, ], @@ -162,6 +175,7 @@ export default defineUserConfig({ text: "What's New", children: [ '/releases/README.md', + '/releases/0.7.6.md', '/releases/0.7.5.md', '/releases/0.7.4.md', '/releases/0.7.3.md', diff --git a/docs/README.md b/docs/README.md index 1850c71..f680609 100644 --- a/docs/README.md +++ b/docs/README.md @@ -17,16 +17,16 @@ features: - title: Contract-First details: One .fluid.yml is the source of truth — schema, quality, build, lineage, governance. Version-controlled and schema-validated. - title: Agentic-Native - details: agentPolicy, sovereignty, and semantics give LLM agents deterministic answers to PII, allowed-use, metric, and residency questions — enforced at the MCP gateway. + details: agentPolicy, sovereignty, and semantics give LLM agents machine-readable answers to PII, allowed-use, metric, and residency questions — and an MCP gateway can enforce them on every read. - title: Operational Superset - details: The only open spec covering build + orchestration + source-aligned acquisition alongside contract and governance. + details: Covers build, orchestration and source-aligned acquisition alongside the data contract and its governance. - title: Interoperable - details: Compiles to Bitol ODPS + ODCS via the forge-cli reference compiler for DataHub / OpenMetadata / Datamesh Manager catalog interop. + details: Compiles to Bitol ODPS + ODCS via the forge-cli reference implementation for DataHub / OpenMetadata / Datamesh Manager catalog interop. - title: Federated by Design details: Built for Data Mesh — decentralized ownership, globally unique product ids, one unified fabric. - title: Open & Apache 2.0-Licensed - details: A community-led protocol. Good ideas backed by real use cases get in. -footer: Apache 2.0 Licensed | © open-data-protocol — Federated Layered Unified Interchange Definition + details: An open specification under the Apache License 2.0, with a public conformance corpus and no CLA. +footer: Licensed under the Apache License 2.0 | Copyright 2025 The FLUID Authors --- > **Your agents are only as trustworthy as the data products they consume.** @@ -65,5 +65,7 @@ exposes: - **[Guide](/fluid/guide/)** — what FLUID is, the quickstart, and the FAQ. - **[Concepts](/fluid/concepts/)** — the agentic-native layer and how FLUID compares to ODCS / ODPS. - **[Schema Reference](/fluid/schema/anatomy)** — every top-level block, a cheatsheet, and the full specification. -- **[What's New in 0.7.5](/fluid/releases/0.7.5)** — streaming Kafka → Iceberg sink + Confluent Cloud Tableflow. +- **[What's New in 0.7.5](/fluid/releases/0.7.5)** — the latest stable version: streaming Kafka → Iceberg sink, Confluent Cloud Tableflow, and a pgvector output port. +- **[0.7.6 (preview)](/fluid/releases/0.7.6)** — packaging modes, declared consumers, cross-mesh pins; opt-in, and still subject to change. - **[See the deck](/fluid/deck/)** — the FLUID story in slides. +- **[Contributing & governance](/fluid/contributing/)** — how the standard is maintained, and how to take part. diff --git a/docs/concepts/README.md b/docs/concepts/README.md index bebd066..4a4f107 100644 --- a/docs/concepts/README.md +++ b/docs/concepts/README.md @@ -35,7 +35,7 @@ A FLUID contract describes a data product end to end: its schema, build logic, d This structure separates **interface** (what you get) from **implementation** (how it's built), enabling reliable data ecosystems ready for both humans and AI agents. -> The latest schema version is **0.7.4**, which adds runtime `agentPolicy` enforcement at the MCP gateway. Like the 0.7.1 → 0.7.3 line, **0.7.4 is additive and fully backward-compatible** — every valid 0.7.3 contract still validates. See the release notes for details. +> The latest stable schema version is **0.7.5**; **0.7.6** is a preview that a contract uses only by declaring it. Every stable release after 0.7.2 has been additive over its predecessor; the one recorded narrowing is the step into 0.7.2 (0.7.1 → 0.7.2). See [What's New](/fluid/releases/) and [Versions](/fluid/schema/versions). --- diff --git a/docs/concepts/agentic-native.md b/docs/concepts/agentic-native.md index 1afa38e..0ea5e80 100644 --- a/docs/concepts/agentic-native.md +++ b/docs/concepts/agentic-native.md @@ -18,7 +18,7 @@ Each spec was evaluated against the **canonical example file its maintainers pub | Bitol ODCS v3.1 | [`full-example.odcs.yaml`](https://github.com/bitol-io/open-data-contract-standard/blob/main/docs/examples/all/full-example.odcs.yaml) (seller/payments contract) | | Bitol ODPS v1.0 | [`customer-data-product.odps.yaml`](https://github.com/bitol-io/open-data-product-standard/blob/main/docs/examples/customer-data-product.odps.yaml) | | ODPS v4 | [`urbanpulse_final.yml`](https://github.com/Open-Data-Product-Initiative/v4.0/blob/main/source/examples/Refs/urbanpulse_final.yml) (UrbanPulse Events) | -| FLUID v0.7.3 | [Example 10 (Customers CDC)](/fluid/examples/#10-source-aligned-acquisition) | +| FLUID v0.7.3 | [Example 10 (Customers CDC)](/fluid/examples/#_10-source-aligned-acquisition) | ## The four agent failure modes diff --git a/docs/concepts/comparisons.md b/docs/concepts/comparisons.md index 6ef893f..66ccb15 100644 --- a/docs/concepts/comparisons.md +++ b/docs/concepts/comparisons.md @@ -23,7 +23,7 @@ Four active "open data product" specs share overlapping names and adjacent scope | **ODCS** (Open Data Contract Standard) | LF AI & Data · Bitol | v3.1.0 — Dec 2025 | Column-level technical contract; producer↔consumer agreement for one dataset | | **Bitol ODPS** (Open Data Product Standard) | LF AI & Data · Bitol | v1.0.0 — Sep 2025 | Thin product manifest; bundles ODCS contracts via `contractId` on input/output ports | | **ODPS v4** (Open Data Product Specification) | LF · [Open-Data-Product-Initiative](https://github.com/Open-Data-Product-Initiative/v4.0) | v4.0 — Jul 2025 · v4.1 — Oct 2025 | Business + commercial wrapper: pricing · license · multi-language · marketplace | -| **FLUID** | open-data-protocol | v0.7.3 — this repo | End-to-end operational contract: schema + build + orchestration + agentic governance + sovereignty + semantics | +| **FLUID** | open-data-protocol | v0.7.5 (stable) — this repo | End-to-end operational contract: schema + build + orchestration + agentic governance + sovereignty + semantics | ## How they actually fit together @@ -34,9 +34,9 @@ flowchart LR classDef core fill:#5B8DEF,color:#fff,stroke:#1E3A8A,stroke-width:4px,font-weight:bold classDef opt fill:#94A3B8,color:#fff,stroke:#475569,stroke-width:1px,stroke-dasharray:6 3 - F["FLUID v0.7.3
your .fluid.yml
standalone — complete on its own"]:::core + F["FLUID v0.7.5
your .fluid.yml
standalone — complete on its own"]:::core - FC["forge-cli
reference compiler
(emits IaC + Airflow DAGs)"]:::opt + FC["forge-cli
reference implementation
(emits IaC + Airflow DAGs)"]:::opt BIT["Bitol ODPS + ODCS
(catalog interop)"]:::opt V4["ODPS v4 wrapper
(commercial publishing)"]:::opt @@ -62,7 +62,7 @@ flowchart LR ## 📊 Capability matrix -Legend: ✅ deterministic in spec · ⚠️ partial · ❌ silent. Headers abbreviated for width: **F** = FLUID v0.7.3 · **ODCS** = Bitol ODCS v3.1 · **ODPS** = Bitol ODPS v1.0 · **v4** = ODPS v4.0. Field-level detail lives in the [**Schema Cheatsheet**](/fluid/schema/cheatsheet) — this matrix is the at-a-glance scoreboard. +Legend: ✅ deterministic in spec · ⚠️ partial · ❌ silent. Headers abbreviated for width: **F** = FLUID v0.7.5 · **ODCS** = Bitol ODCS v3.1 · **ODPS** = Bitol ODPS v1.0 · **v4** = ODPS v4.0. Field-level detail lives in the [**Schema Cheatsheet**](/fluid/schema/cheatsheet) — this matrix is the at-a-glance scoreboard. ### 📐 Data shape & quality @@ -111,7 +111,7 @@ Legend: ✅ deterministic in spec · ⚠️ partial · ❌ silent. Headers abbre > Each capability links to its field-level reference in the [**Schema Cheatsheet**](/fluid/schema/cheatsheet). MetricFlow round-trip is on the FLUID roadmap. -> ⚙️ The matrix above shows what each spec *covers*. **[`forge-cli`](/fluid/concepts/forge-cli)** is the reference compiler that turns a FLUID contract into deployed reality and Bitol-compatible outputs. +> ⚙️ The matrix above shows what each spec *covers*. **[`forge-cli`](/fluid/concepts/forge-cli)** is the reference implementation that turns a FLUID contract into deployed reality and Bitol-compatible outputs. --- @@ -138,7 +138,7 @@ The cleanest production stack uses all four where each is strongest: - [Bitol ODCS — open-data-contract-standard](https://github.com/bitol-io/open-data-contract-standard) (v3.1.0) - [Bitol ODPS — open-data-product-standard](https://github.com/bitol-io/open-data-product-standard) (v1.0.0) - [opendataproducts.org ODPS v4](https://opendataproducts.org/v4.0/) · [v4.0 repo](https://github.com/Open-Data-Product-Initiative/v4.0) · [v4.1 release](https://github.com/Open-Data-Product-Initiative/v4.1) -- [forge-cli — the FLUID reference compiler](https://github.com/Agenticstiger/forge-cli) · [forge-docs](https://agenticstiger.github.io/forge_docs/) (emits Bitol ODPS + ODCS) +- [forge-cli — the FLUID reference implementation](https://github.com/Agenticstiger/forge-cli) · [forge-docs](https://agenticstiger.github.io/forge_docs/) (emits Bitol ODPS + ODCS) - [Linux Foundation AI & Data — Bitol project](https://lfaidata.foundation/projects/bitol/) --- diff --git a/docs/concepts/forge-cli.md b/docs/concepts/forge-cli.md index ed0644a..aa22a4d 100644 --- a/docs/concepts/forge-cli.md +++ b/docs/concepts/forge-cli.md @@ -1,34 +1,65 @@ -# The Reference Compiler — forge-cli +# The Reference Implementation — forge-cli -A capability matrix shows what each spec *covers*. **[`forge-cli`](https://github.com/Agenticstiger/forge-cli)** is what turns a FLUID contract into deployed reality and Bitol-compatible outputs. +A capability matrix shows what each spec *covers*. **[`forge-cli`](https://github.com/Agenticstiger/forge-cli)** is the reference implementation of FLUID: it validates FLUID contracts and turns them into deployed infrastructure, generated orchestration and catalog-interop documents. [![Repo](https://img.shields.io/badge/Agenticstiger%2Fforge--cli-FF6B35?logo=github&logoColor=white&style=for-the-badge)](https://github.com/Agenticstiger/forge-cli) [![License Apache 2.0](https://img.shields.io/badge/license-Apache%202.0-1E3A8A?style=for-the-badge)](https://github.com/Agenticstiger/forge-cli/blob/main/LICENSE) [![Docs](https://img.shields.io/badge/docs-forge--docs-A78BFA?style=for-the-badge&logo=readthedocs&logoColor=white)](https://agenticstiger.github.io/forge_docs/) -> **FLUID stands on its own.** A FLUID contract is complete and useful with no compiler at all — you can author, validate, and version it as the single source of truth. `forge-cli` is an **optional adapter**: reach for it once you want compiled Infrastructure-as-Code, generated orchestration DAGs, and catalog-interop outputs. +> **FLUID stands on its own.** A FLUID contract is complete and useful with no compiler at all — you can author, validate and version it with any JSON Schema validator. forge-cli is **one implementation**, not the referee: what makes a document valid FLUID is the published schema and the [conformance corpus](https://github.com/open-data-protocol/fluid/blob/main/tests/README.md), and forge-cli is implementation #1 under test against them ([GOVERNANCE.md](https://github.com/open-data-protocol/fluid/blob/main/GOVERNANCE.md)). -## What forge-cli does +## Install -`forge-cli` consumes a `.fluid.yml` and compiles it to: +```bash +pip install data-product-forge +fluid --version +``` -- **Infrastructure-as-Code** — OpenTofu / Terraform for BigQuery, Snowflake, AWS, and GCP. -- **Airflow DAGs** (and Dagster / Prefect) generated from the contract's `orchestration` block. -- **Bitol ODPS + ODCS** — for catalog and data-mesh registry interop (DataHub, OpenMetadata, Datamesh Manager). +- The package is **`data-product-forge`** on PyPI; the command it installs is `fluid`, and the Python import module is `fluid_build`. +- The older distribution name **`fluid-forge`** is frozen at 0.7.9; `pip install fluid-forge` installs that old line, not the current one. +- forge-cli is licensed under Apache 2.0, like this specification. -## Compile stages +## Which FLUID versions it supports -It emits the following, stage by stage: +The JSON Schemas are authored in forge-cli and published here (see [CONTRIBUTING.md](https://github.com/open-data-protocol/fluid/blob/main/CONTRIBUTING.md#the-schemas-are-not-authored-here)). As of data-product-forge 0.18.1: -| Stage | Output | +- It **bundles** the schemas for 0.7.1 to 0.7.6 and validates a contract against the schema of the `fluidVersion` it declares (`fluid validate --list-versions` lists them). +- **0.7.5 is its default** — the newest stable version — and `fluid init --quickstart` writes `fluidVersion: 0.7.5`. +- **0.7.6 is a preview**: validated when a contract declares it, never chosen by default. See [Stable and preview versions](/fluid/schema/versions#stable-and-preview-versions). +- For an older version it does not bundle, it fetches the schema from this repository and caches it (`--offline` restricts it to bundled and cached schemas). +- Its bundled 0.7.1 is a different document from the published 0.7.1 under the same `$id`; see [Versions](/fluid/schema/versions#_0-7-1-two-documents-share-one-id). + +`fluid validate` also applies rules the schema cannot express — for example, a `confluent` binding must use `format: iceberg` and name its environment, cluster, bucket and role. Those are the implementation's own checks: a document can be schema-valid and still be refused by `fluid validate`. + +## What it does with a contract + +| Area | What forge-cli produces | |---|---| -| 🟢 **Bitol export** | `1 ODPS doc + N .odcs.yaml` files — Bitol v1.0.0 as the default, center-stage ODPS. | -| 🟠 **Orchestration** | Native **Airflow / Dagster / Prefect** DAGs from the `orchestration` block. | -| 🟣 **Infrastructure** | **OpenTofu / Terraform** IaC for BigQuery / Snowflake / AWS / GCP. | -| 🔵 **Governance** | **IAM bindings** from `accessPolicy.grants[]` + **AI gateway** enforcement of `agentPolicy`. | -| 🔴 **Supply chain** | **Cosign-verified** connector images + **SLSA provenance** checks on ingest. | +| **Validation** | `fluid validate` — the declared version's schema plus provider rules. | +| **Infrastructure** | OpenTofu / Terraform for the binding's platform (`fluid plan`, `fluid apply`), and `fluid verify` to check the deployed result. | +| **Orchestration** | Airflow, Dagster and Prefect workflows (`fluid generate schedule`). | +| **Standards** | Bitol ODPS and ODCS documents, and Linux Foundation ODPS v4.1 on request (`fluid generate standard`), for catalog interop. | +| **Governance** | Cloud access grants and governance resources derived from the contract (which ones depend on the platform), and an MCP output-port gateway that enforces `agentPolicy` on exposes that carry an `mcp` block. | + +## Composing a contract from several files + +FLUID defines one document; how a contract may be split across files is implementation-defined (see [Specification → Composing a document from several files](/fluid/schema/specification#composing-a-document-from-several-files)). forge-cli's mechanism: + +```bash +fluid split contract.fluid.yaml # write fragments/ and replace blocks with $ref nodes +fluid bundle contract.fluid.yaml # resolve every $ref back into one document +``` + +`fluid validate`, `plan` and `apply` resolve `$ref` nodes before anything else. Since 0.18, refs are confined to the root contract's directory tree — URLs, absolute paths and escapes through `..` or symlinks are refused — and `FLUID_REF_ROOT` widens that root for monorepos. See [Composing a contract from fragments](https://agenticstiger.github.io/forge_docs/concepts/contract-refs.html), [`fluid split`](https://agenticstiger.github.io/forge_docs/cli/split.html) and [`fluid bundle`](https://agenticstiger.github.io/forge_docs/cli/bundle.html). + +## Checking forge-cli against the corpus + +This repository's [`conformance/check_reference.py`](https://github.com/open-data-protocol/fluid/blob/main/conformance/check_reference.py) runs every corpus case through forge-cli and reports any case where its verdict differs from the corpus: -> *"What Terraform did for infrastructure, FLUID Forge does for data products."* — [forge-docs](https://agenticstiger.github.io/forge_docs/) +```bash +pip install data-product-forge +python3 conformance/check_reference.py +``` ## Learn more diff --git a/docs/contributing/README.md b/docs/contributing/README.md index 30f8582..dab9e31 100644 --- a/docs/contributing/README.md +++ b/docs/contributing/README.md @@ -1,77 +1,93 @@ -# Contributing to FLUID: Shape the Standard +# Contributing to FLUID -Welcome! FLUID is more than a specification — it's a community-driven standard for the future of data. We thrive on collaboration, innovation, and practical solutions. If you want to help define a great standard and solve real problems, you're in the right place. +FLUID is an open specification, licensed under the [Apache License 2.0](https://github.com/open-data-protocol/fluid/blob/main/LICENSE) and held by **The FLUID Authors** — everyone whose work has been merged into the repository. This page summarises how to take part. The authoritative versions are [CONTRIBUTING.md](https://github.com/open-data-protocol/fluid/blob/main/CONTRIBUTING.md) and [GOVERNANCE.md](https://github.com/open-data-protocol/fluid/blob/main/GOVERNANCE.md) in the repository; where they and this page differ, they win. --- -## 🚀 Our philosophy: build with us +## The schemas are not authored here -We believe in a simple, meritocratic process: **good ideas, backed by clear use cases, get in**. Your expertise and creativity can help shape the future of data engineering. +The JSON Schemas under `schema/` are **vendored**: they are authored in [forge-cli](https://github.com/Agenticstiger/forge-cli), the reference implementation, and this repository is their public distribution point. Nothing copies them here automatically: the `drift` job in `.github/workflows/conformance.yml` runs on every pull request and fails when a synced schema differs by even one byte from the copy in the forge-cli release **pinned** in [`scripts/schema-versions.json`](https://github.com/open-data-protocol/fluid/blob/main/scripts/schema-versions.json) (`upstream.ref`, and `upstream.commit`, which is what the check fetches), when that release bundles a schema version not published here, or when its preview versions or latest stable version differ from that file. A pull request that edits a file under `schema/` by hand fails the `drift` job — not because the change is wrong, but because the published schema must be the one the reference implementation ships. The job fails that pull request's checks; whether a failing check also blocks the merge is a branch-protection setting, not something this page promises. Propose schema changes upstream in forge-cli. ---- - -## 🛠️ How to contribute - -We keep things lean and effective — no unnecessary bureaucracy. +The comparison is with the pinned release, not with whatever forge-cli released last, so a new forge-cli release does not turn unrelated pull requests red; moving to one is a change to the pin, reviewed in a pull request. A separate, advisory `drift-latest` job warns when forge-cli has released after the pinned release, and cannot fail a run. The `generated` job also fails when a file under `schema/` or `specs/` is neither generated for a synced version nor listed, with its sha256, in the `frozen` map of `scripts/schema-versions.json`, or when a frozen file changes. -### 1. Have an idea? Start the conversation +`scripts/schema-versions.json` is the record of which versions are synced, which is the latest **stable** version and which are **previews**; `python3 generate-docs.py --print-latest-stable` prints the stable one. To re-vendor after a drift failure or a new forge-cli release (`python3 scripts/check-schema-drift.py --latest` says whether there is one): choose a release tag (never `main`) and find its commit with `git ls-remote`, set `upstream.ref` and `upstream.commit` in `scripts/schema-versions.json` (and `synced`, `preview` and `latestStable` if a version was added or its status changed), copy each changed schema byte for byte from that commit, regenerate `specs/` and `schema-diffs/` with `generate-docs.py` and `generate-schema-diffs.py` (the renderer is installed with `pip install --require-hashes -r scripts/requirements-docs.lock`), run the gates — `python3 scripts/check-schema-drift.py` now compares with the commit you recorded — and name the release and commit in the pull request. The full steps are in [CONTRIBUTING.md](https://github.com/open-data-protocol/fluid/blob/main/CONTRIBUTING.md#re-vendoring-a-schema). -- **Big ideas, new features, or breaking changes:** - Open a [GitHub Issue](https://github.com/open-data-protocol/fluid/issues) to start a discussion. - *Tip: A great proposal explains the "why" as much as the "what."* +What you can change here: -- **Bug fixes, typos, or small improvements:** - Skip the issue — just open a Pull Request! +| Area | What it is | +|---|---| +| `tests/**` | The **conformance corpus** — where "FLUID-conformant" is defined. See [tests/README.md](https://github.com/open-data-protocol/fluid/blob/main/tests/README.md). | +| `conformance/**` | The corpus runner, the reference-implementation check, and the mutation-coverage tool. | +| `scripts/check-compat.py`, `scripts/compat-waivers.txt` | The backward-compatibility gate and its decision record. | +| `docs/**` | This documentation site. | +| `examples/**` | Example contracts. | --- -### 2. Fork, code, and submit a Pull Request +## The most valuable contribution: conformance cases -1. **Fork** the repository. -2. **Make your changes** directly to the specification files. -3. **Add a real-world example.** - *Features without clear, practical examples will not be accepted.* -4. **Submit your PR.** - Link to your issue if you created one. +The corpus is what makes conformance a fact rather than a claim, and it lets anyone check an implementation without access to the reference one. The rules, in short (full version in [tests/README.md](https://github.com/open-data-protocol/fluid/blob/main/tests/README.md)): ---- +- Assert on **JSON Schema keyword + RFC 6901 pointer**, never on message text. +- An invalid case must declare **why** it is invalid. +- Keep cases **minimal**: an invalid document differs from a valid one in exactly the way under test. +- The bar is **falsifiability**: if deleting the constraint you meant to pin leaves the corpus green, the case is not doing its job (`python3 conformance/mutation_coverage.py --group `). -### 3. Community review: fast, fair, focused +Run the gates locally before opening a pull request: -Your PR will be reviewed by core contributors and the community. We look for: +```bash +pip install "jsonschema[format]>=4.22" +python3 conformance/run.py # the corpus must be green +python3 tests/meta_test.py # the gates must be able to fail +python3 scripts/check-compat.py # no release may break its promise +``` -- **Alignment with FLUID principles:** - Federated, Layered, Unified, Interchange Definition. -- **Clarity and simplicity:** - Is it easy to understand and elegant? -- **Real-world impact:** - Does it solve an actual problem? +Optionally, check the corpus against the reference implementation: -If it's good, it's in. We value thoughtful debate and feedback, but we have a strong bias for action. +```bash +pip install data-product-forge +python3 conformance/check_reference.py +``` --- -## 🌟 Becoming a core contributor +## Changing what the schema means -Consistently submit high-quality PRs, provide thoughtful reviews, and help shape discussions. That's the path to becoming a core contributor and a recognized member of the FLUID project. +Anything that changes which documents validate is **normative**. Open an issue before writing code, and expect to be asked for the corpus cases that pin the new behaviour. Editorial changes — typos, prose, descriptions that do not affect validation — can go straight to a pull request. -Your name, ideas, and expertise will be embedded in a standard powering the next generation of data engineering. +**Compatibility.** Every release promises that valid contracts keep validating, and `scripts/check-compat.py` checks it on every pull request. A deliberate narrowing goes in `scripts/compat-waivers.txt` *with the evidence that settled it*; the waiver file is a decision record, not a mute button. --- -## 🤝 Code of Conduct +## Pull requests + +- One concern per pull request. +- Say what you verified and how — the command you ran and what it printed. +- New behaviour needs a case that would fail without it. +- Be honest about what you did not do. -We expect everyone to follow our Code of Conduct: +**Licence and provenance.** Contributions are accepted under the repository's Apache License 2.0. By opening a pull request you certify that you wrote the contribution or otherwise have the right to submit it under that licence — the [Developer Certificate of Origin](https://developercertificate.org/). Sign your commits with `git commit -s` if you want that recorded explicitly. There is **no contributor licence agreement**. -- Be respectful. -- Challenge ideas, not people. +--- + +## Governance in brief + +- FLUID is pre-1.0 and is stewarded by its originating maintainers in the [open-data-protocol](https://github.com/open-data-protocol) organization; decisions are made in the open, in issues and pull requests. +- **Two homes:** the JSON Schemas are authored in forge-cli; the conformance corpus lives here, so that conformance can be checked by someone who has neither the engine nor an account with the maintainers. forge-cli is implementation #1 under test — never the referee. +- Anything the specification does not decide is written down as an open question rather than settled by implementation behaviour (see "Known ambiguities" in [tests/README.md](https://github.com/open-data-protocol/fluid/blob/main/tests/README.md)). -Let's build something great — together. +Read [GOVERNANCE.md](https://github.com/open-data-protocol/fluid/blob/main/GOVERNANCE.md) for the full model, including vendor neutrality and trademarks. --- +## Reporting problems + +- **Bugs and spec questions** — open an [issue](https://github.com/open-data-protocol/fluid/issues). +- **Security vulnerabilities** — do not open an issue; see [SECURITY.md](https://github.com/open-data-protocol/fluid/blob/main/SECURITY.md). +- **Conduct** — see [CODE_OF_CONDUCT.md](https://github.com/open-data-protocol/fluid/blob/main/CODE_OF_CONDUCT.md). + ## Where to go next -- **[Core Principles](/fluid/concepts/principles)** — the five principles every contribution should align with. -- **[Schema Reference](/fluid/schema/)** — the specification files you'll be editing. -- **[Examples](/fluid/examples/)** — where your real-world example will live. +- **[Core Principles](/fluid/concepts/principles)** — the principles every contribution should align with. +- **[Schema Versions](/fluid/schema/versions)** — stable and preview versions, and what each status means. +- **[Vision](/fluid/vision/)** — where the standard is heading. diff --git a/docs/deck/README.md b/docs/deck/README.md index dbe393e..6d63d70 100644 --- a/docs/deck/README.md +++ b/docs/deck/README.md @@ -111,13 +111,13 @@ A visual tour of FLUID. Use ← / → (or the controls) to move between slides.

Structuring an Enterprise Data Product

-

The "gold standard" for a complex product: a multi-file layout that lets different teams own different parts of one contract.

+

One contract, owned by several teams. FLUID itself defines a single document; splitting it across files is left to tools. With the reference implementation, fluid split writes a root contract plus a fragments/ directory that maps onto team ownership:

    -
  • Data Modeler — owns exposes.yml and the schema_*.yml files.
  • -
  • Data Engineer — owns build.yml and consumes.yml.
  • -
  • Governance Lead — owns quality.yml and accessPolicy.yml.
  • +
  • Data Modeler — owns fragments/exposes/<exposeId>.yaml: each output port with its schema and quality rules.
  • +
  • Data Engineer — owns fragments/builds/<id>.yaml.
  • +
  • Governance Lead — owns fragments/sovereignty.yaml and fragments/access-policy.yaml.
-

A root fluid.yml composes the detailed implementation files (via $ref) that live in a .fluid/ directory — separating high-level identity from its parts while keeping a single source of truth.

+

The root file keeps identity and ownership and points at each fragment with $ref; fluid bundle resolves it back into the one document that is validated. See Composing a document from several files.

@@ -165,7 +165,7 @@ A visual tour of FLUID. Use ← / → (or the controls) to move between slides.
  • Vendors — build the next generation of FLUID-aware tools.
  • Developers — shape the spec and grow the ecosystem.
  • -

    Help us build the missing protocol for the agentic era. Start with the guide, explore the concepts, or dive into the schema (current version 0.7.4).

    +

    Help us build the missing protocol for the agentic era. Start with the guide, explore the concepts, or dive into the schema (latest stable version 0.7.5).

    diff --git a/docs/examples/README.md b/docs/examples/README.md index 40ab14e..3b13222 100644 --- a/docs/examples/README.md +++ b/docs/examples/README.md @@ -1,7 +1,9 @@ -# FLUID by Example — Eleven Steps from "Hello World" to Production +# FLUID by Example — From "Hello World" to Production -> All examples target **`fluidVersion: "0.7.4"`** (the [latest schema](/fluid/schema/fluid-schema-0.7.4.json)). +> Examples 1–13 target **`fluidVersion: "0.7.5"`**, the [latest stable schema](/fluid/schema/fluid-schema-0.7.5.json). Example 14 uses the **0.7.6 preview**. > Each example **adds one block** to the previous one. Read in order to learn the schema by accretion. +> Every complete contract on this page validates against the JSON Schema of the version it declares (Draft 2020-12) and with the reference implementation's `fluid validate`; examples that show only the new block validate when merged into the example before them. +> File names such as `01-minimal.fluid.yml` are illustrative: FLUID does not prescribe a file name, and the reference implementation scaffolds `contract.fluid.yaml`. > For the schema's mental model first, see [**Anatomy**](/fluid/schema/anatomy). > For a one-row-per-field lookup, see [**Cheatsheet**](/fluid/schema/cheatsheet). @@ -22,17 +24,20 @@ | 9 | [Define business semantics](#_9-define-business-semantics) | `exposes[].semantics` | | 10 | [Source-aligned acquisition](#_10-source-aligned-acquisition) | `build.pattern: acquisition` | | 11 | [Agent-consumable output port (MCP)](#_11-agent-consumable-output-port-mcp) | `exposes[].mcp` (⭐ 0.7.4) | +| 12 | [Stream Kafka into Iceberg](#_12-stream-kafka-into-iceberg) | `kafka-connect` Iceberg streaming sink + Iceberg catalog location (⭐ 0.7.5) | +| 13 | [Vector output port (pgvector)](#_13-vector-output-port-pgvector) | `platform: pgvector` + `vectorConfig` (⭐ 0.7.5) | +| 14 | [Preview: packaging and consumers](#_14-preview-packaging-and-declared-consumers-0-7-6) | `packaging`, `consumers` (🧪 0.7.6 preview) | --- ## 1. Minimal valid contract -> **Goal:** get a `.fluid.yml` that validates against v0.7.4 with the fewest possible lines. +> **Goal:** get a contract that validates against 0.7.5 with the fewest possible lines. > **New in this step:** the six required top-level keys (`fluidVersion`, `kind`, `id`, `name`, `metadata`, `exposes`) and the four required fields inside each expose (`exposeId`, `kind`, `contract`, `binding`). ```yaml # 01-minimal.fluid.yml -fluidVersion: "0.7.4" +fluidVersion: "0.7.5" kind: DataProduct id: demo.bronze.hello_world name: "Hello World" @@ -63,7 +68,7 @@ That's it. Everything else in this guide is opt-in. ```yaml # 02-with-schema.fluid.yml -fluidVersion: "0.7.4" +fluidVersion: "0.7.5" kind: DataProduct id: demo.bronze.payments name: "Raw Payments" @@ -94,6 +99,7 @@ Why bother: every consumer (humans, BI tools, AI agents) now has a machine-reada > **Goal:** assert that certain quality conditions must hold for the data to be considered valid. > **New in this step:** `exposes[].contract.dq.rules[]`. + ```yaml # 03-with-dq.fluid.yml # ... (everything from example 2, plus:) @@ -135,6 +141,7 @@ DQ rules run as part of every build. `severity: error` fails the pipeline; `warn > **Goal:** describe how the data actually gets produced. > **New in this step:** the `build` block with `pattern: embedded-logic`. + ```yaml # 04-with-build.fluid.yml # ... (everything from example 3, plus:) @@ -171,7 +178,7 @@ build: ```yaml # 05-customer-ltv.fluid.yml -fluidVersion: "0.7.4" +fluidVersion: "0.7.5" kind: DataProduct id: analytics.silver.customer_ltv name: "Customer Lifetime Value" @@ -198,7 +205,7 @@ exposes: binding: platform: gcp format: bigquery_table - location: { project: company-data, dataset: silver_analytics, table: customer_ltv } + location: { project: company-data, dataset: silver_analytics, table: customer_ltv, region: europe-west1 } build: pattern: hybrid-reference @@ -217,6 +224,7 @@ Now the orchestrator can auto-build the DAG: `demo.bronze.payments` → `analyti > **Goal:** declare exactly who can read this product. The platform generates cloud IAM bindings from this block. > **New in this step:** root-level `accessPolicy`. + ```yaml # 06-with-access.fluid.yml # ... (everything from example 5, plus:) @@ -243,6 +251,7 @@ The `resources` field uses JSONPath to target subsets of `exposes` — useful wh > **New in this step:** `agentPolicy` under `exposes[].policy.agentPolicy`. > ⚠️ **Important shape note:** `agentPolicy` is **per-expose**, not top-level — it lives inside `exposes[].policy.agentPolicy`. (Some prior release notes show it at the root; the schema has never accepted it there.) + ```yaml # 07-with-agent-policy.fluid.yml # ... (everything from example 6, plus inside the relevant expose:) @@ -285,6 +294,7 @@ This block is enforced by FLUID-aware AI gateways. An LLM request that doesn't m > **Goal:** enforce data residency — apply-time blocks any `binding` that would land data outside the allowed region. > **New in this step:** root-level `sovereignty`. + ```yaml # 08-with-sovereignty.fluid.yml # ... (everything from example 7, plus:) @@ -310,7 +320,7 @@ With `enforcementMode: strict` and `validationRequired: true`, contract-apply di ```yaml # 09-with-semantics.fluid.yml -fluidVersion: "0.7.4" +fluidVersion: "0.7.5" kind: DataProduct id: analytics.gold.orders_revenue name: "Orders Revenue Semantic Model" @@ -380,7 +390,7 @@ The shape mirrors **dbt MetricFlow** and **Snowflake Semantic Views** — portab ```yaml # 10-acquisition.fluid.yml -fluidVersion: "0.7.4" +fluidVersion: "0.7.5" kind: DataProduct id: crm.bronze.customers_cdc name: "Customers CDC from Production Postgres" @@ -422,7 +432,7 @@ exposes: binding: platform: aws format: iceberg - location: { bucket: acme-bronze, path: "crm/customers/" } + location: { bucket: acme-bronze, path: "crm/customers/", region: eu-west-1 } icebergConfig: # NB: icebergConfig.partitionSpec uses OBJECT form; writeVersion: 2 # acquisitionSink.partitionBy below uses STRING form fileFormat: parquet @@ -439,7 +449,7 @@ build: mode: cdc cursor_field: updated_at connection: - secretRef: "vault://pg-prod-readonly" # URI form required (vault:// aws:// gcp:// azure:// env://) + secretRef: "vault://pg-prod-readonly" # a URI (://…), e.g. vault://, env:// streams: [public.customers] watermark: { strategy: lsn } # Postgres LSN-based @@ -480,7 +490,7 @@ build: debezium: connector_class: io.debezium.connector.postgresql.PostgresConnector - deployment: { mode: managed } # Forge provisions via Helm + deployment: { mode: managed } # the platform provisions the connector image_signature: verifier: cosign # cosign is the only verifier today publicKey: "k8s://acme/cosign-pub" @@ -534,12 +544,12 @@ A FLUID-aware platform takes this file, provisions the connector, registers it i > **Goal:** publish an output port that an AI agent can describe, sample, and query directly — with governance enforced from the contract, not bolted on at the gateway. > **New in this step:** `exposes[].mcp` (⭐ 0.7.4). -Adding an `mcp` block to an expose opts that output port into the **Fluid MCP gateway**. A tool such as `fluid mcp output-port serve` then surfaces the port to Claude Code, Cursor, or any MCP client, exposing `describe` / `sample` / `query` operations against the contract. +Adding an `mcp` block to an expose opts that output port into an MCP output-port gateway. The reference implementation's `fluid mcp output-port serve` then surfaces the port to Claude Code, Cursor, or any MCP client, exposing `describe` / `sample` / `query` operations against the contract. -The contract below is `0.7.4`-valid and runs without cloud credentials (DuckDB over a local CSV): +The contract below is valid against 0.7.5 and runs without cloud credentials (DuckDB over a local CSV next to the contract): ```yaml -fluidVersion: "0.7.4" +fluidVersion: "0.7.5" kind: DataProduct id: silver.demo.customer_segments_v1 name: Customer Segments (MCP demo) @@ -635,15 +645,145 @@ The `mcp` block opts the port in; the actual *who/what/how* rules come from the - **`policy.agentPolicy`** (Example 7) is enforced **at the gateway on every read**. `allowedModels` / `deniedModels` gate which LLM may call the port, and `allowedUseCases` / `deniedUseCases` gate why. A request that doesn't satisfy `allowedModels` ∧ `allowedUseCases` is rejected before any data is read. - **Column-level `sensitivity`** drives **value redaction**. A column marked `sensitivity: pii` (like `email` above) has its *values* replaced with a redaction marker (e.g. `[REDACTED-PII]`) in every `sample` / `query` result, while the column itself stays visible — the agent learns the field exists but never sees a real address. `sensitivity: phi` is treated the same way for protected health information. -The result: an agent connected through the Fluid MCP gateway can explore the port's shape and semantics, run governed queries, and stay inside the contract's model/use-case/redaction guardrails — without a single flag or proxy outside the contract. +The result: an agent connected through such a gateway can explore the port's shape and semantics, run governed queries, and stay inside the contract's model/use-case/redaction guardrails — without a single flag or proxy outside the contract. > See the [MCP how-to](/fluid/how-to/mcp) for the end-to-end agentic-access story, and the [0.7.4 release notes](/fluid/releases/0.7.4) for the full `exposes[].mcp` reference. --- +## 12. Stream Kafka into Iceberg + +> **Goal:** land Kafka topics continuously in an Iceberg table through Kafka Connect, with the table's catalog named in the contract. +> **New in this step:** the `kafka-connect` Iceberg streaming sink and the Iceberg catalog location fields (⭐ 0.7.5). + +```yaml +# 12-kafka-to-iceberg.fluid.yml +fluidVersion: "0.7.5" +kind: DataProduct +id: sales.bronze.orders_stream +name: "Orders (streamed)" +metadata: + owner: { team: sales-data } + productType: SDP +builds: + - id: ingest_orders + pattern: acquisition + engine: kafka-connect + capabilities: [streaming] + outputs: [orders] + properties: + source: + kind: kafka + mode: streaming + connection: { secretRef: "env://KAFKA_BOOTSTRAP" } + streams: [orders, payments] + kafka-connect: + iceberg_sink_enabled: true # derive the Iceberg sink connector + sink_topics: [orders, payments] + streamingSink: { commitIntervalMs: 60000, autoCreate: true, evolveSchema: true } +exposes: + - exposeId: orders + kind: table + contract: + schema: + - { name: order_id, type: STRING, required: true } + - { name: amount, type: NUMERIC } + binding: + platform: aws + format: iceberg + location: + bucket: acme-lake + database: sales + table: orders + region: eu-west-1 + catalog: glue # ⭐ 0.7.5 — Iceberg catalog kind + warehouse: "s3://acme-lake/warehouse" # ⭐ 0.7.5 +``` + +The sink settings live under the build's `properties`, in the key named after the engine. For a managed alternative, the `confluent` platform (Confluent Cloud Tableflow) materialises a topic as an Iceberg table — see [What's New in 0.7.5](/fluid/releases/0.7.5#confluent-confluent-cloud-tableflow-binding). + +--- + +## 13. Vector output port (pgvector) + +> **Goal:** publish embeddings of a text column as a vector output port that a retrieval (RAG) application can query. +> **New in this step:** `binding.platform: pgvector`, `format: pgvector_table` and `binding.vectorConfig` (⭐ 0.7.5). + +```yaml +# 13-vector-port.fluid.yml +fluidVersion: "0.7.5" +kind: DataProduct +id: support.gold.ticket_embeddings +name: "Support ticket embeddings" +metadata: + owner: { team: support-ai } +exposes: + - exposeId: tickets + kind: vector + contract: + schema: + - { name: ticket_id, type: STRING, required: true } + - { name: body, type: TEXT, labels: { ai-embeddable: "true" } } + binding: + platform: pgvector + format: pgvector_table + location: { database: support, schema: public, table: tickets } + vectorConfig: + dimensions: 1536 # required, 1–16000 + embeddingModel: text-embedding-3-small + distanceMetric: cosine # cosine | l2 | inner_product | l1 + indexType: hnsw # hnsw | ivfflat | none + hnsw: { m: 16, efConstruction: 64 } + sourceKeyColumn: ticket_id # ties each embedding back to its source row +``` + +Which columns are embedded is not part of the schema; the reference implementation embeds the columns labelled `ai-embeddable: "true"`. + +--- + +## 14. Preview: packaging and declared consumers (0.7.6) + +> 🧪 **0.7.6 is a preview.** It is opt-in — declare `fluidVersion: "0.7.6"` — and can still change before it is promoted. See [0.7.6 (preview)](/fluid/releases/0.7.6). +> **New in this step:** top-level `packaging` and `consumers`. + +```yaml +# 14-preview-packaging.fluid.yml +fluidVersion: "0.7.6" +kind: DataProduct +id: finance.gold.revenue +name: "Revenue" +metadata: + owner: { team: finance-analytics } +packaging: + mode: shared # write into platform-owned containers… + pool: analytics-eu + containers: { warehouse: isolated } # …but own the warehouse +consumers: + - name: weekly_revenue_dashboard + type: dashboard # dashboard | notebook | analysis | ml | application + label: "Weekly Revenue Dashboard" + maturity: high + exposeIds: [revenue] +exposes: + - exposeId: revenue + kind: table + contract: + schema: + - { name: day, type: DATE, required: true } + - { name: revenue, type: NUMERIC, required: true } + binding: + platform: snowflake + format: snowflake_table + location: { account: acme-eu, database: FINANCE, schema: GOLD, table: REVENUE } +``` + +The 0.7.6 release note covers the rest of the preview: cross-mesh pins on `consumes[]`, encryption keys and principal mapping on bindings, and expiring data with `lifecycle.expire`. + +--- + ## Where to go from here - [**Anatomy**](/fluid/schema/anatomy) — guided tour of every top-level block. - [**Cheatsheet**](/fluid/schema/cheatsheet) — one-row-per-field lookup table. -- [**What's New**](/fluid/releases/) — auto-generated diffs between each version. -- [**Full Specification**](/fluid/schema/specification) — exhaustive field-by-field reference. +- [**What's New**](/fluid/releases/) — release notes for each version. +- [**Specification**](/fluid/schema/specification) — validation semantics and the document structure. diff --git a/docs/guide/quickstart.md b/docs/guide/quickstart.md index a7053b1..4a81bee 100644 --- a/docs/guide/quickstart.md +++ b/docs/guide/quickstart.md @@ -1,10 +1,10 @@ # Quickstart -A FLUID contract is **one YAML file**. This page shows the smallest file that validates against FLUID **0.7.5**, explains each required block, and shows how to validate it. +A FLUID contract is **one YAML file**. This page shows the smallest file that validates against FLUID **0.7.5** — the latest stable version — explains each required block, and shows how to validate it. ## The smallest valid contract -Every block below is required. Every other top-level block — `description`, `domain`, `tags`, `labels`, `consumes`, `build`, `orchestration`, governance (`agentPolicy`, `sovereignty`, `accessPolicy`, `retention`), `lineage`, `lifecycle`, `environments`, `docs` — is opt-in and layered on as you need it. +Every block below is required. Every other top-level block — `description`, `domain`, `tags`, `labels`, `consumes`, `build` / `builds`, `orchestration`, `sovereignty`, `accessPolicy`, `governance`, `retention`, `lineage`, `lifecycle`, `environments`, `docs`, `extensions` — is opt-in and layered on as you need it. (`agentPolicy` is opt-in too, but it lives inside an expose, under `exposes[].policy`.) ```yaml fluidVersion: "0.7.5" @@ -29,11 +29,11 @@ exposes: | Block | Required | Meaning | |---|---|---| -| `fluidVersion` | ✅ | Which contract version this file targets. Use `"0.7.5"` for the latest schema. | +| `fluidVersion` | ✅ | The schema version this file declares; it is validated against that version's schema. Use `"0.7.5"`, the latest stable version. | | `kind` | ✅ | `DataProduct` or `MLPipeline`. | | `id` | ✅ | Globally unique product id. Convention: `domain.layer.name`. | | `name` | ✅ | Human-readable display name. | -| `metadata` | ✅ | Only `metadata.owner` is required; by convention `owner.team` is the one truly required field. | +| `metadata` | ✅ | Only `metadata.owner` is required. The schema requires none of `owner`'s members, but name a `team`: it is how tools route ownership and alerts. | | `exposes` | ✅ | The ports you publish. Each entry requires `exposeId`, `kind`, `contract`, and `binding`. | Inside each `exposes[]` entry: @@ -41,21 +41,30 @@ Inside each `exposes[]` entry: - **`contract`** — the schema columns (or an `openapiRef`) plus optional data-quality rules. - **`binding`** — where the data physically lives: `platform` + `format` + `location`. -> **0.7.5 note.** The `kafka-connect` acquisition pattern gains an opt-in Iceberg streaming sink (`iceberg_sink_enabled`, `streamingSink`, `sink_topics`), and `binding.platform` adds `confluent` (Confluent Cloud Tableflow). 0.7.5 is **additive and fully backward-compatible** — every valid 0.7.4 contract still validates, so bumping `fluidVersion` is all it takes. +> **0.7.5 note.** 0.7.5 adds an opt-in Iceberg streaming sink on the `kafka-connect` acquisition engine, the `confluent` (Tableflow) and `pgvector` platforms, and new location fields — see [What's New in 0.7.5](/fluid/releases/0.7.5). It is additive: a valid 0.7.4 contract stays valid when it declares `"0.7.5"`, and declaring `"0.7.5"` is what lets a contract use the new fields. +> +> **0.7.6 is a preview.** Use it only if you need one of its fields, by declaring `fluidVersion: "0.7.6"`; it can still change. See [0.7.6 (preview)](/fluid/releases/0.7.6). ## Validate it -Validate any `.fluid.yml` against the published JSON Schema for 0.7.5: +Validate a contract against the published JSON Schema **of the version it declares** — for this file, 0.7.5: ``` https://open-data-protocol.github.io/fluid/schema/fluid-schema-0.7.5.json ``` -Point any JSON-Schema validator (for example `ajv`, `check-jsonschema`, or your CI's schema step) at that URL. +Use a validator that supports JSON Schema Draft 2020-12, which every FLUID schema declares. Validators that compile `pattern` as an ECMA-262 regular expression, as JavaScript-based ones typically do, may refuse one pattern in the 0.7.2–0.7.6 schemas; see the [known interoperability issue](/fluid/schema/specification#known-interoperability-issue-one-pattern-is-not-an-ecma-262-regular-expression). Python's `jsonschema`, which this repository's conformance tooling uses, validates them. + +Or use the reference implementation, which validates against the declared version and adds its own provider checks: + +```bash +pip install data-product-forge +fluid validate contract.fluid.yaml # ✅ Valid FLUID contract (schema v0.7.5) +``` ### Editor autocomplete & inline validation -Add this line as the **first line** of your `.fluid.yml` so the [YAML Language Server](https://github.com/redhat-developer/yaml-language-server) (bundled with the VS Code YAML extension) gives you live completion and validation: +Add this line as the **first line** of your contract file so the [YAML Language Server](https://github.com/redhat-developer/yaml-language-server) (bundled with the VS Code YAML extension) gives you live completion and validation. Use the URL of the version the file declares; pointing a 0.7.4 file at the 0.7.5 schema would hide 0.7.5-only fields that its own schema rejects. ```yaml # yaml-language-server: $schema=https://open-data-protocol.github.io/fluid/schema/fluid-schema-0.7.5.json diff --git a/docs/how-to/README.md b/docs/how-to/README.md index 5fd203e..cbf36b8 100644 --- a/docs/how-to/README.md +++ b/docs/how-to/README.md @@ -2,6 +2,10 @@ Task-focused walkthroughs that show FLUID working alongside the tools you already run — ingestion sources, transformation engines, orchestrators, and AI agents. Each guide is self-contained; start with the one closest to your problem. +::: warning Most of these guides use a legacy manifest shape +The YAML in these guides predates the current schema — most of it declares `fluidVersion: "1.0"`, a version that was never published, and the build-patterns guide uses 0.3.0 — and it does **not** validate against the latest stable schema, 0.7.5. Each guide says so at the top. Read them for the integration patterns; for YAML that validates, use [**Examples**](/fluid/examples/), the [**Anatomy**](/fluid/schema/anatomy) and the [**release notes**](/fluid/releases/). +::: + For a step-by-step tour of the schema itself, see [**Examples**](/fluid/examples/). For field-level reference, see the [**Cheatsheet**](/fluid/schema/cheatsheet). | Guide | What it covers | diff --git a/docs/how-to/advanced.md b/docs/how-to/advanced.md index fe26273..f8f1a5f 100644 --- a/docs/how-to/advanced.md +++ b/docs/how-to/advanced.md @@ -2,7 +2,7 @@ Explore the full power and extensibility of the FLUID specification with these advanced, real-world use cases. -> ℹ️ These examples use an earlier, illustrative manifest shape (`fluidVersion: "1.0"`, `metadata.dataProduct`, `exposes[].location`) to sketch the *art of the possible*. They are conceptual, not `0.7.4`-valid templates. For current, schema-valid syntax see [**Examples**](/fluid/examples/). +> ℹ️ These examples use an earlier, illustrative manifest shape (`fluidVersion: "1.0"`, `metadata.dataProduct`, `exposes[].location`) to sketch the *art of the possible*. They are conceptual and do **not** validate against any published FLUID schema (`1.0` was never published; the latest stable schema is `0.7.5`). For current, schema-valid syntax see [**Examples**](/fluid/examples/). --- @@ -14,6 +14,7 @@ Serve different data to different partners, enforcing dynamic, context-aware acc
    YAML: 11-dynamic-policies.fluid.yml + ```yaml fluidVersion: "1.0" kind: VirtualDataProduct @@ -25,24 +26,24 @@ metadata: consumes: - type: fluid-product - name: inventory.gold.live_stock_by_warehouse + name: inventory.gold.live_stock_by_warehouse exposes: - location: { type: 'virtual' } - contract: - schema: - columns: - - { name: product_sku, type: STRING } - - { name: quantity_on_hand, type: INT64 } + contract: + schema: + columns: + - { name: product_sku, type: STRING } + - { name: quantity_on_hand, type: INT64 } dynamicPolicies: rules: - name: "Allow authorized partners based on JWT claim" - condition: "agent.jwt.claims.can_access_stock_api == true" - grant: - permissions: [readData] - scope: - rowFilter: "partner_id = '{{ agent.jwt.claims.partner_id }}'" + condition: "agent.jwt.claims.can_access_stock_api == true" + grant: + permissions: [readData] + scope: + rowFilter: "partner_id = '{{ agent.jwt.claims.partner_id }}'" build: execution: { trigger: { type: 'manual' } } @@ -60,6 +61,7 @@ Bridge data engineering and MLOps by delivering features directly to a feature s
    YAML: 12-feature-store.fluid.yml + ```yaml fluidVersion: "1.0" kind: DataProduct @@ -72,21 +74,21 @@ metadata: consumes: - type: fluid-product - name: customers.silver.trusted_customers + name: customers.silver.trusted_customers exposes: - location: - type: redis - connection: secret:ml-feature-store-redis-creds - properties: - keyPrefix: 'customer_churn_features' - contract: - schema: - columns: - - { name: 'customer_id', type: 'STRING' } - - { name: 'recency_days', type: 'INT64' } - - { name: 'frequency_30d', type: 'INT64' } - - { name: 'last_updated_ts', type: 'TIMESTAMP' } + type: redis + connection: secret:ml-feature-store-redis-creds + properties: + keyPrefix: 'customer_churn_features' + contract: + schema: + columns: + - { name: 'customer_id', type: 'STRING' } + - { name: 'recency_days', type: 'INT64' } + - { name: 'frequency_30d', type: 'INT64' } + - { name: 'last_updated_ts', type: 'TIMESTAMP' } build: transformation: @@ -111,6 +113,7 @@ Centralize observability by consuming execution logs from all FLUID products, po
    YAML: 13-active-metadata.fluid.yml + ```yaml fluidVersion: "1.0" kind: DataProduct @@ -123,34 +126,34 @@ metadata: consumes: - type: gcs - connection: secret:gcp-prod-sa-key - format: { type: 'jsonl' } - properties: - bucket: 'fluid-execution-logs-prod' - path: 'runs/' + connection: secret:gcp-prod-sa-key + format: { type: 'jsonl' } + properties: + bucket: 'fluid-execution-logs-prod' + path: 'runs/' exposes: - location: - type: bigquery - properties: { project: 'acme-prod-dwh', dataset: 'observability', table: 'fluid_runs' } - contract: - schema: - columns: - - { name: 'run_id', type: 'STRING' } - - { name: 'data_product_name', type: 'STRING' } - - { name: 'status', type: 'STRING' } - - { name: 'start_time', type: 'TIMESTAMP' } - - { name: 'duration_ms', type: 'INT64' } - - { name: 'rows_written', type: 'INT64' } - - { name: 'error_message', type: 'STRING' } - quality: - - rule: in_set - columns: [status] - set: ['SUCCESS', 'FAILED', 'QUARANTINED'] - onFailure: - action: 'alert' - notifications: - - { channel: 'slack', target: '#platform-alerts' } + type: bigquery + properties: { project: 'acme-prod-dwh', dataset: 'observability', table: 'fluid_runs' } + contract: + schema: + columns: + - { name: 'run_id', type: 'STRING' } + - { name: 'data_product_name', type: 'STRING' } + - { name: 'status', type: 'STRING' } + - { name: 'start_time', type: 'TIMESTAMP' } + - { name: 'duration_ms', type: 'INT64' } + - { name: 'rows_written', type: 'INT64' } + - { name: 'error_message', type: 'STRING' } + quality: + - rule: in_set + columns: [status] + set: ['SUCCESS', 'FAILED', 'QUARANTINED'] + onFailure: + action: 'alert' + notifications: + - { channel: 'slack', target: '#platform-alerts' } build: execution: { trigger: { type: 'streaming' }, runtime: { type: 'gcp-cloud-run' } } @@ -168,6 +171,7 @@ Enrich a Gold-layer product with formal semantic meaning from an external ontolo
    YAML: 14-semantic-product.fluid.yml + ```yaml fluidVersion: "1.0" kind: DataProduct @@ -180,30 +184,30 @@ metadata: consumes: - type: fluid-product - name: products.silver.cleaned_catalog + name: products.silver.cleaned_catalog exposes: - location: - type: bigquery - properties: { project: 'acme-prod-dwh', dataset: 'gold', table: 'product_catalog' } - contract: - schema: - columns: - - { name: 'product_id', type: 'STRING' } - - { name: 'name', type: 'STRING' } - - { name: 'description', type: 'STRING' } - - { name: 'price', type: 'NUMERIC' } - semantics: - ontology: "https://schema.org/docs/schema_org_rdfa.html" - classifications: - - column: product_id - term: "schema:sku" - - column: name - term: "schema:name" - - column: description - term: "schema:description" - - column: price - term: "schema:price" + type: bigquery + properties: { project: 'acme-prod-dwh', dataset: 'gold', table: 'product_catalog' } + contract: + schema: + columns: + - { name: 'product_id', type: 'STRING' } + - { name: 'name', type: 'STRING' } + - { name: 'description', type: 'STRING' } + - { name: 'price', type: 'NUMERIC' } + semantics: + ontology: "https://schema.org/docs/schema_org_rdfa.html" + classifications: + - column: product_id + term: "schema:sku" + - column: name + term: "schema:name" + - column: description + term: "schema:description" + - column: price + term: "schema:price" build: # ... build definition ... ``` @@ -219,6 +223,7 @@ Create a temporary, virtual data product for a single user conversation. Joins m
    YAML: 15-ephemeral-product.fluid.yml + ```yaml fluidVersion: "1.0" kind: VirtualDataProduct @@ -230,19 +235,19 @@ metadata: consumes: - type: fluid-product - name: sales.silver.clean_orders - alias: orders + name: sales.silver.clean_orders + alias: orders - type: fluid-product - name: customers.silver.trusted_customers - alias: customers + name: customers.silver.trusted_customers + alias: customers exposes: - location: { type: 'virtual' } - contract: - schema: - columns: - - { name: order_id, type: STRING } - - { name: order_total, type: NUMERIC } + contract: + schema: + columns: + - { name: order_id, type: STRING } + - { name: order_total, type: NUMERIC } build: transformation: diff --git a/docs/how-to/airflow.md b/docs/how-to/airflow.md index da6538e..3e651f5 100644 --- a/docs/how-to/airflow.md +++ b/docs/how-to/airflow.md @@ -1,6 +1,8 @@ # Integrating dbt and Airflow with FLUID: A Practical Guide -This guide presents a hands-on approach to integrating the [FLUID specification](https://github.com/open-data-protocol/fluid/blob/main/specification.md) with core data stack tools like **dbt** and **Airflow**. It explains how FLUID bridges the gap between these tools, clarifies their responsibilities, and demonstrates how to orchestrate robust, contract-aware data pipelines. +> ℹ️ **Legacy manifest shape.** The YAML in this guide uses an earlier, illustrative manifest shape (`fluidVersion: "1.0"`, a version that was never published) and does **not** validate against any published FLUID schema. Read it as a description of the integration pattern. On the current schema (latest stable `0.7.5`), orchestration is expressed with the `orchestration` block — `engine: airflow`, with `airflow.dagId` and `airflow.tasks[]` — see [**Anatomy §6**](/fluid/schema/anatomy#_6-orchestration-who-actually-runs-the-build). + +This guide presents a hands-on approach to integrating the [FLUID specification](/fluid/schema/specification) with core data stack tools like **dbt** and **Airflow**. It explains how FLUID bridges the gap between these tools, clarifies their responsibilities, and demonstrates how to orchestrate robust, contract-aware data pipelines. --- @@ -75,6 +77,7 @@ Organize by data domain. FLUID files live with the dbt project they orchestrate. **`silver_stg_customers/product.fluid.yml`**: + ```yaml fluidVersion: 1.0 kind: DataProduct @@ -188,5 +191,3 @@ for spec_file in glob.glob(f"{FLUID_PRODUCT_PATH}/**/*.fluid.yml", recursive=Tru ## Conclusion By adopting FLUID, you move from implicit, brittle glue code to explicit, version-controlled, contract-aware data products. This is the foundation for a scalable, trustworthy data platform. - -> ℹ️ The manifest above uses the legacy `fluidVersion: 1.0` shape. On the current schema (latest `0.7.4`), orchestration is expressed via the top-level `orchestration` block (e.g. `engine: airflow`) — see [**Examples → Source-aligned acquisition**](/fluid/examples/#_10-source-aligned-acquisition). diff --git a/docs/how-to/build-patterns.md b/docs/how-to/build-patterns.md index 447038c..b9f7531 100644 --- a/docs/how-to/build-patterns.md +++ b/docs/how-to/build-patterns.md @@ -2,7 +2,7 @@ *A Practical Guide to Build Patterns for FLUID v0.3.0* -> ℹ️ **This guide reflects the v0.3.0 build-pattern model.** The pattern *names* and the `build` block have evolved since: the current schema (latest `0.7.4`) uses `build.pattern` values such as `embedded-logic`, `hybrid-reference`, and `acquisition`, with the transformation engine and properties as direct children of `build`. Read this as a conceptual tour of the available *styles* of build; for current, schema-valid syntax see [**Examples**](/fluid/examples/). +> ℹ️ **This guide reflects the v0.3.0 build-pattern model.** The pattern *names* and the `build` block have evolved since: the current schema (latest stable `0.7.5`) uses `build.pattern` values such as `embedded-logic`, `hybrid-reference`, and `acquisition`, with the transformation engine and properties as direct children of `build`. Its YAML does not validate even against the 0.3.0 schema (for example, its ids carry a `:1.0.0` suffix that the 0.3.0 id pattern rejects). Read this as a conceptual tour of the available *styles* of build; for current, schema-valid syntax see [**Examples**](/fluid/examples/). --- @@ -32,6 +32,7 @@ This guide explores **four canonical Build Patterns**, complete with real-world ### Example: High-Value Customers in BigQuery + ```yaml fluidVersion: "0.3.0" kind: DataProduct @@ -127,6 +128,7 @@ build: ### Example: Quarterly Financial Report with dbt + Snowflake + ```yaml fluidVersion: "0.3.0" kind: DataProduct @@ -207,6 +209,7 @@ build: ### Example: Spark Structured Streaming (5G Latency Alerts) + ```yaml fluidVersion: "0.3.0" kind: DataProduct @@ -288,6 +291,7 @@ build: ### Example: Data Vault 2.0 Model with Iceberg + ```yaml fluidVersion: "0.3.0" kind: DataProduct diff --git a/docs/how-to/datavault.md b/docs/how-to/datavault.md index 3ddf414..e2bdbb8 100644 --- a/docs/how-to/datavault.md +++ b/docs/how-to/datavault.md @@ -1,5 +1,7 @@ # A Practical Guide to Integrating Data Vault 2.0 +> ℹ️ **Legacy manifest shape.** The YAML in this guide uses an earlier, illustrative manifest shape (`fluidVersion: "1.0"`, a version that was never published) to illustrate the Data Vault pattern, and does **not** validate against any published FLUID schema. On the current schema (latest stable `0.7.5`), express dbt builds with `build.pattern: hybrid-reference` and upstream links with `consumes[]` — see the [**dbt how-to**](/fluid/how-to/dbt) and [**Examples**](/fluid/examples/). + This guide presents a hands-on proposal for integrating the [FLUID specification](https://github.com/open-data-protocol/fluid) with core data stack tools like **dbt** and **datavault 2.0**. It explains how this integration streamlines data engineering workflows, clarifies tool responsibilities, and solves common pain points in modern data platforms. ## Example: Building a Data Vault 2.0 Product @@ -63,6 +65,7 @@ FROM source_data ### FLUID Files **1. Hub Data Product (`customer_hub/product.fluid.yml`):** + ```yaml fluidVersion: 1.0 kind: DataProduct @@ -71,9 +74,9 @@ metadata: owner: { team: 'data-architecture' } consumes: - type: dbt-model - name: stg_crm_customers + name: stg_crm_customers - type: dbt-model - name: stg_ecommerce_users + name: stg_ecommerce_users exposes: location: type: bigquery @@ -99,6 +102,7 @@ build: ``` **2. Satellite Data Product (`crm_satellite/product.fluid.yml`):** + ```yaml fluidVersion: 1.0 kind: DataProduct @@ -141,5 +145,3 @@ Engineers declare dependencies in FLUID files; the orchestrator builds and maint --- > **FLUID + dbt + Airflow = Declarative, contract-aware, and maintainable data pipelines.** - -> ℹ️ The manifests above use the legacy `fluidVersion: 1.0` shape to illustrate the Data Vault pattern. On the current schema (latest `0.7.4`), express dbt builds via `build.pattern: hybrid-reference` and upstream links via `consumes[]` — see the [**dbt how-to**](/fluid/how-to/dbt) and [**Examples**](/fluid/examples/). diff --git a/docs/how-to/dbt.md b/docs/how-to/dbt.md index 0252efb..a68a181 100644 --- a/docs/how-to/dbt.md +++ b/docs/how-to/dbt.md @@ -1,5 +1,7 @@ # A Practical Guide to Integrating dbt with FLUID +> ℹ️ **Legacy manifest shape.** The FLUID YAML in this guide uses an earlier, illustrative manifest shape (`fluidVersion: "1.0"`, a version that was never published) to illustrate the dbt-orchestration pattern, and does **not** validate against any published FLUID schema. On the current schema (latest stable `0.7.5`), the equivalent is `build.pattern: hybrid-reference` with `engine: dbt` — see [**Examples → Consume another product**](/fluid/examples/#_5-consume-another-product). + > **Unlock seamless, contract-aware data engineering with FLUID and dbt.** --- @@ -100,6 +102,7 @@ Ingest raw customer data from GCS → Load to "bronze" table → Use dbt to tran #### File 1: `bronze_raw_customers/product.fluid.yml` *Ingest raw data from GCS into BigQuery.* + ```yaml fluidVersion: 1.0 kind: DataProduct @@ -143,6 +146,7 @@ build: #### File 2: `silver_stg_customers/product.fluid.yml` *Orchestrate dbt to build the silver product.* + ```yaml fluidVersion: 1.0 kind: DataProduct @@ -201,10 +205,10 @@ build: version: 2 sources: - name: staging - database: bq-prod-lakehouse - schema: bronze - tables: - - name: raw_customers + database: bq-prod-lakehouse + schema: bronze + tables: + - name: raw_customers ``` --- @@ -234,21 +238,21 @@ from source version: 2 models: - name: stg_customers - description: "Cleansed customer records from raw source." - columns: - - name: customer_id - description: "The unique customer identifier." - tests: - - not_null - - unique - - name: first_name - description: "Customer's first name." - - name: last_name - description: "Customer's last name." - - name: created_at - description: "Timestamp when the record was ingested." - tests: - - not_null + description: "Cleansed customer records from raw source." + columns: + - name: customer_id + description: "The unique customer identifier." + tests: + - not_null + - unique + - name: first_name + description: "Customer's first name." + - name: last_name + description: "Customer's last name." + - name: created_at + description: "Timestamp when the record was ingested." + tests: + - not_null ``` --- @@ -270,5 +274,3 @@ models: --- > **By adopting FLUID, you move from a world of implicit, brittle glue code to explicit, version-controlled, contract-aware data products. This is the foundation for a scalable, trustworthy data fabric.** - -> ℹ️ The manifests above use the legacy `fluidVersion: 1.0` shape to illustrate the dbt-orchestration pattern. On the current schema (latest `0.7.4`), the equivalent is `build.pattern: hybrid-reference` with `engine: dbt` — see [**Examples → Consume another product**](/fluid/examples/#_5-consume-another-product). diff --git a/docs/how-to/mcp.md b/docs/how-to/mcp.md index 65dd595..9fd7f59 100644 --- a/docs/how-to/mcp.md +++ b/docs/how-to/mcp.md @@ -1,8 +1,8 @@ # Unlocking Governable AI: Agentic Data Access with MCP & FLUID -> ⭐ **As of schema `0.7.4`, the recommended path is the `exposes[].mcp` block.** Add `mcp.sampling` / `mcp.classification` to an output port to opt it into the **Fluid MCP gateway**, where `policy.agentPolicy` (allowed/denied models and use cases) is enforced at runtime on every read and column-level `sensitivity: pii` / `phi` redacts values. See the [**0.7.4 release notes**](/fluid/releases/0.7.4) for the full reference and [**Examples → Agent-consumable output port (MCP)**](/fluid/examples/#_11-agent-consumable-output-port-mcp) for a complete, valid contract. +> ⭐ **As of schema `0.7.4`, the recommended path is the `exposes[].mcp` block.** Add `mcp.sampling` / `mcp.classification` to an output port to opt it into an MCP output-port gateway — the reference implementation provides one — where `policy.agentPolicy` (allowed/denied models and use cases) is enforced at runtime on every read and column-level `sensitivity: pii` / `phi` redacts values. See the [**0.7.4 release notes**](/fluid/releases/0.7.4) for the full reference and [**Examples → Agent-consumable output port (MCP)**](/fluid/examples/#_11-agent-consumable-output-port-mcp) for a complete, valid contract. > -> The walkthrough below explains the *conceptual* request/authorization flow that underpins this gateway. It uses an earlier illustrative manifest shape; treat the YAML as a narrative aid, not a `0.7.4` template. +> The walkthrough below explains the *conceptual* request/authorization flow that underpins this gateway. It uses an earlier illustrative manifest shape; treat the YAML as a narrative aid: it declares `fluidVersion: 1.0`, a version that was never published, and does not validate against any published FLUID schema. For a valid contract, use the example linked above. **Imagine this:** A sales executive simply asks their AI assistant: @@ -25,6 +25,7 @@ Together, they enable **agentic data access**: AI agents can act on behalf of us A data engineering team publishes a FLUID data product: + ```yaml fluidVersion: 1.0 kind: DataProduct @@ -54,23 +55,23 @@ exposes: grants: # Sales team: see treated PII for their region - principal: group:sales-de@company.com - permissions: [readData] - scope: - privacyView: treated - columns: [customer_id, full_name, email, country, total_lifetime_value] - rowFilter: "country = 'DE'" + permissions: [readData] + scope: + privacyView: treated + columns: [customer_id, full_name, email, country, total_lifetime_value] + rowFilter: "country = 'DE'" # Fraud agent: see cleartext PII - principal: agent:fraud_investigation_agent_v1 - permissions: [readData] - scope: - privacyView: cleartext - columns: [customer_id, full_name, email, last_purchase_date] + permissions: [readData] + scope: + privacyView: cleartext + columns: [customer_id, full_name, email, last_purchase_date] # AI Assistant: limited, read-only access (no PII) - principal: agent:ai_assistant_prod - permissions: [readData] - scope: - privacyView: treated - columns: [customer_id, country, total_lifetime_value, last_purchase_date] + permissions: [readData] + scope: + privacyView: treated + columns: [customer_id, country, total_lifetime_value, last_purchase_date] ``` **Key Takeaway:** diff --git a/docs/how-to/source-aligned-data-product.md b/docs/how-to/source-aligned-data-product.md index ff331fa..afaab7f 100644 --- a/docs/how-to/source-aligned-data-product.md +++ b/docs/how-to/source-aligned-data-product.md @@ -1,5 +1,7 @@ # Source-Aligned Ingestion from On-Prem Oracle to Cloud +> ℹ️ **Legacy manifest shape.** This guide predates the current schema: its YAML uses an earlier manifest shape (`fluidVersion: "1.0"`, `metadata.dataProduct`, `exposes[].location`) and does **not** validate against any published FLUID schema. For the current source-aligned ingestion model (latest stable `0.7.5`), see the `build.pattern: acquisition` example in [**Examples → Source-aligned acquisition**](/fluid/examples/#_10-source-aligned-acquisition). + This example showcases a core FLUID ecosystem pattern: **Source-Aligned Data Product**. The objective is to create a reliable, governable, and secure mirror of a source system in the cloud's bronze layer—without altering the data's meaning. --- @@ -16,6 +18,7 @@ This example showcases a core FLUID ecosystem pattern: **Source-Aligned Data Pro A single `customer.bronze.raw_oracle_customers.fluid.yml` file drives the entire process—no extra configuration or code required. + ```yaml fluidVersion: "1.0" kind: DataProduct @@ -34,70 +37,70 @@ metadata: # 2. CONSUMES: Source system definition. consumes: - type: oracle-db - connection: secret:onprem-oracle-erp-readonly-creds - properties: - query: | - SELECT - CUST_ID, - F_NAME, - L_NAME, - CUST_EMAIL_ADDR, - PHONE_INTL, - COUNTRY_CODE, - CREATED_TS, - LAST_UPDATED_TS - FROM ERP.CUSTOMERS - WHERE LAST_UPDATED_TS > '{{ watermark.last_updated_ts }}' + connection: secret:onprem-oracle-erp-readonly-creds + properties: + query: | + SELECT + CUST_ID, + F_NAME, + L_NAME, + CUST_EMAIL_ADDR, + PHONE_INTL, + COUNTRY_CODE, + CREATED_TS, + LAST_UPDATED_TS + FROM ERP.CUSTOMERS + WHERE LAST_UPDATED_TS > '{{ watermark.last_updated_ts }}' # 3. EXPOSES: Output interface. exposes: - location: - type: gcs - connection: secret:gcp-prod-sa-key - format: { type: 'parquet' } - properties: - bucket: 'prod-customer-landing-zone' - path: 'raw_oracle_customers/' - partitionBy: ['load_date'] - - # 4. CONTRACT: Governance and enforcement. - contract: - schema: - columns: - - { name: 'customer_id', type: 'INT64', nullable: false } - - { name: 'first_name_pii', type: 'STRING' } - - { name: 'last_name_pii', type: 'STRING' } - - { name: 'email_hash', type: 'STRING' } - - { name: 'phone_token', type: 'STRING' } - - { name: 'country_code', type: 'STRING' } - - { name: 'created_ts', type: 'TIMESTAMP' } - - { name: 'last_updated_ts', type: 'TIMESTAMP' } - - { name: 'load_date', type: 'DATE' } - - quality: - - rule: not_null - columns: [customer_id] - onFailure: { action: 'reject_row' } - - rule: regex_match - columns: [CUST_EMAIL_ADDR] - pattern: '^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$' - onFailure: { action: 'quarantine_row', location: 'gs://prod-customer-quarantine/invalid_emails/' } - - privacy: - - classification: PII - columns: [CUST_EMAIL_ADDR] - treatment: - type: hashing - properties: { algorithm: 'SHA256' } - newColumn: 'email_hash' - - classification: SPI - columns: [PHONE_INTL] - treatment: - type: tokenization - properties: - vault: 'gcp-dlp-service' - keyId: 'customer-phone-key' - newColumn: 'phone_token' + type: gcs + connection: secret:gcp-prod-sa-key + format: { type: 'parquet' } + properties: + bucket: 'prod-customer-landing-zone' + path: 'raw_oracle_customers/' + partitionBy: ['load_date'] + + # 4. CONTRACT: Governance and enforcement. + contract: + schema: + columns: + - { name: 'customer_id', type: 'INT64', nullable: false } + - { name: 'first_name_pii', type: 'STRING' } + - { name: 'last_name_pii', type: 'STRING' } + - { name: 'email_hash', type: 'STRING' } + - { name: 'phone_token', type: 'STRING' } + - { name: 'country_code', type: 'STRING' } + - { name: 'created_ts', type: 'TIMESTAMP' } + - { name: 'last_updated_ts', type: 'TIMESTAMP' } + - { name: 'load_date', type: 'DATE' } + + quality: + - rule: not_null + columns: [customer_id] + onFailure: { action: 'reject_row' } + - rule: regex_match + columns: [CUST_EMAIL_ADDR] + pattern: '^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$' + onFailure: { action: 'quarantine_row', location: 'gs://prod-customer-quarantine/invalid_emails/' } + + privacy: + - classification: PII + columns: [CUST_EMAIL_ADDR] + treatment: + type: hashing + properties: { algorithm: 'SHA256' } + newColumn: 'email_hash' + - classification: SPI + columns: [PHONE_INTL] + treatment: + type: tokenization + properties: + vault: 'gcp-dlp-service' + keyId: 'customer-phone-key' + newColumn: 'phone_token' # 5. BUILD: Implementation logic. build: @@ -147,5 +150,3 @@ build: > - **Governed:** Data contracts and privacy rules are enforced in-flight. > - **Cloud-Native:** Runs on ephemeral, scalable GCP DataProc clusters. > - **Extensible:** Easily adapted for other sources, targets, or domains. - -> ℹ️ This guide predates the current schema and uses the legacy `fluidVersion: "1.0"` manifest shape (`metadata.dataProduct`, `exposes[].location`, `contract.quality`). For the current source-aligned ingestion model, see the `build.pattern: acquisition` example in [**Examples → Source-aligned acquisition**](/fluid/examples/#_10-source-aligned-acquisition) (the latest schema is `0.7.4`). diff --git a/docs/how-to/source-aligned-kafka.md b/docs/how-to/source-aligned-kafka.md index a329a36..27ebf0f 100644 --- a/docs/how-to/source-aligned-kafka.md +++ b/docs/how-to/source-aligned-kafka.md @@ -1,5 +1,7 @@ # Source-Aligned Ingestion from On-Prem Kafka to Cloud +> ℹ️ **Legacy manifest shape.** This guide predates the current schema: its YAML uses an earlier manifest shape (`fluidVersion: "1.0"`) and does **not** validate against any published FLUID schema. For streaming ingestion on the current schema (latest stable `0.7.5`), see the `build.pattern: acquisition` example in [**Examples → Source-aligned acquisition**](/fluid/examples/#_10-source-aligned-acquisition) and the Kafka → Iceberg streaming sink in [**0.7.5**](/fluid/releases/0.7.5). + This example showcases a powerful FLUID pattern: **Source-Aligned Data Product**. The objective is to mirror a source system in the cloud's bronze layer—reliably, securely, and with strong governance—without altering the data's meaning. --- @@ -16,6 +18,7 @@ This example showcases a powerful FLUID pattern: **Source-Aligned Data Product** A single `finance.bronze.raw_kafka_payments.fluid.yml` file drives the entire process—no extra code or config needed. + ```yaml fluidVersion: "1.0" kind: DataProduct @@ -34,58 +37,58 @@ metadata: # 2. CONSUMES: Source System Definition consumes: - type: kafka - connection: secret:onprem-kafka-cluster-creds - format: { type: 'json' } - properties: - topic: 'prod.financial.payments' - consumerGroup: 'fluid-gcs-sink-v1' - # Optional: startingOffsets: 'earliest' + connection: secret:onprem-kafka-cluster-creds + format: { type: 'json' } + properties: + topic: 'prod.financial.payments' + consumerGroup: 'fluid-gcs-sink-v1' + # Optional: startingOffsets: 'earliest' # 3. EXPOSES: Output Interface exposes: - location: - type: gcs - connection: secret:gcp-prod-sa-key - format: { type: 'parquet' } - properties: - bucket: 'prod-finance-landing-zone' - path: 'raw_payments/' - partitionBy: ['load_date'] - - # 4. CONTRACT: Governance & Enforcement - contract: - schema: - columns: - - { name: 'payment_id', type: 'STRING', nullable: false } - - { name: 'amount', type: 'NUMERIC' } - - { name: 'currency', type: 'STRING' } - - { name: 'payment_method_token', type: 'STRING' } - - { name: 'customer_email_hash', type: 'STRING' } - - { name: 'event_timestamp', type: 'TIMESTAMP' } - - { name: 'load_date', type: 'DATE' } - quality: - - rule: not_null - columns: [payment_id, amount] - onFailure: { action: 'reject_row' } - - rule: in_set - columns: [currency] - set: ['USD', 'EUR', 'GBP', 'JPY'] - onFailure: { action: 'quarantine_row', location: 'gs://prod-finance-quarantine/invalid_currency/' } - privacy: - - classification: PII - columns: [user_email] - treatment: - type: hashing - properties: { algorithm: 'SHA256' } - newColumn: 'customer_email_hash' - - classification: SPI - columns: [payment_method_details] - treatment: - type: tokenization - properties: - vault: 'gcp-dlp-service' - keyId: 'payment-method-key' - newColumn: 'payment_method_token' + type: gcs + connection: secret:gcp-prod-sa-key + format: { type: 'parquet' } + properties: + bucket: 'prod-finance-landing-zone' + path: 'raw_payments/' + partitionBy: ['load_date'] + + # 4. CONTRACT: Governance & Enforcement + contract: + schema: + columns: + - { name: 'payment_id', type: 'STRING', nullable: false } + - { name: 'amount', type: 'NUMERIC' } + - { name: 'currency', type: 'STRING' } + - { name: 'payment_method_token', type: 'STRING' } + - { name: 'customer_email_hash', type: 'STRING' } + - { name: 'event_timestamp', type: 'TIMESTAMP' } + - { name: 'load_date', type: 'DATE' } + quality: + - rule: not_null + columns: [payment_id, amount] + onFailure: { action: 'reject_row' } + - rule: in_set + columns: [currency] + set: ['USD', 'EUR', 'GBP', 'JPY'] + onFailure: { action: 'quarantine_row', location: 'gs://prod-finance-quarantine/invalid_currency/' } + privacy: + - classification: PII + columns: [user_email] + treatment: + type: hashing + properties: { algorithm: 'SHA256' } + newColumn: 'customer_email_hash' + - classification: SPI + columns: [payment_method_details] + treatment: + type: tokenization + properties: + vault: 'gcp-dlp-service' + keyId: 'payment-method-key' + newColumn: 'payment_method_token' # 5. BUILD: Implementation Logic build: @@ -136,5 +139,3 @@ build: --- > **Tip:** Use this pattern to bootstrap trusted, analytics-ready data products from any source system—no custom code required! - -> ℹ️ This guide predates the current schema and uses the legacy `fluidVersion: "1.0"` manifest shape. For streaming/CDC ingestion on the current schema, see the `build.pattern: acquisition` example in [**Examples → Source-aligned acquisition**](/fluid/examples/#_10-source-aligned-acquisition) (the latest schema is `0.7.4`). diff --git a/docs/releases/0.7.1.md b/docs/releases/0.7.1.md index 124dcae..ad1967e 100644 --- a/docs/releases/0.7.1.md +++ b/docs/releases/0.7.1.md @@ -1,175 +1,164 @@ # What's New in FLUID 0.7.1 — Agentic Governance + Provider-First Orchestration -**FLUID 0.7.1** represents a significant evolution focused on **Agentic Governance** and **Provider-First Orchestration**. Built with **100% backward compatibility** with v0.5.7, it adds powerful new capabilities for the AI-driven enterprise. +**FLUID 0.7.1** adds **agentic governance** — which AI models may read an expose, and for what — together with data **sovereignty**, a root-level **access policy**, and **provider actions** as orchestration tasks. It is **additive** over 0.5.7: `scripts/check-compat.py --from 0.5.7 --to 0.7.1` finds no narrowing change. -> 💡 **`agentPolicy` location.** AI/LLM consumption policy lives **per-expose** under `exposes[].policy.agentPolicy` — not at the root, despite the top-level snippets in this historical release note. The schema has never accepted it at the top level. See the [Schema Anatomy](/fluid/schema/anatomy) for the correct shape. +Every example on this page validates against the published [0.7.1 schema](/fluid/schema/fluid-schema-0.7.1.json) (JSON Schema Draft 2020-12) when placed in a minimal contract. + +::: tip This page was corrected +Earlier versions of this page showed `agentPolicy` at the top level, `sovereignty` fields that do not exist, free-text use cases, and a top-level `orchestration` block. None of those validate against 0.7.1. The examples below use the shapes the 0.7.1 schema defines. Note also that the 0.7.1 schema bundled by the reference implementation differs from the published one — see [Versions](/fluid/schema/versions#_0-7-1-two-documents-share-one-id). +::: --- -## 🤖 Agentic Governance (NEW) +## 🤖 Agentic governance: `exposes[].policy.agentPolicy` (NEW) -Control **which AI models** can access your data and **how they can use it**: +Control **which AI models** can read an expose and **how they may use it**. The policy is per-expose, under `exposes[].policy.agentPolicy`: + ```yaml -# NEW in v0.7.1: AI/LLM usage policies -agentPolicy: - allowedModels: - - "gpt-4" - - "claude-3-opus" - - "gemini-1.5-pro" - maxTokensPerRequest: 8192 - maxTokensPerDay: 100000 - allowedUseCases: - - "customer-insights" - - "market-analysis" - deniedUseCases: - - "political-profiling" - - "credit-scoring" - requiresHumanReview: true - auditLog: - enabled: true - includePrompts: true +exposes: + - exposeId: customer_insights + kind: table + # ...contract, binding... + policy: + agentPolicy: # NEW in v0.7.1 + allowedModels: [gpt-4, claude-3-opus, gemini-1.5-pro] + deniedModels: [gpt-3.5-turbo] + maxTokensPerRequest: 8192 + maxTokensPerDay: 100000 + allowedUseCases: [analysis, summarization, qa] # controlled vocabulary + deniedUseCases: [training, fine_tuning] + canReason: false + canStore: false + auditRequired: true + purposeLimitation: "Customer-insights analysis only; no credit scoring or political profiling." ``` -**Why this matters:** As AI agents become primary data consumers, organizations need granular control over: - -- ✅ **Model-specific access** — whitelist/blacklist AI models -- ✅ **Usage boundaries** — define permitted and prohibited use cases -- ✅ **Rate limiting** — token quotas per request and per day -- ✅ **Audit compliance** — full logging of AI interactions with data -- ✅ **Human oversight** — require review for sensitive operations +- **Model access** — `allowedModels` / `deniedModels`. +- **Use cases** — `allowedUseCases` / `deniedUseCases` take values from a fixed vocabulary: `inference`, `reasoning`, `analysis`, `summarization`, `classification`, `embedding`, `search`, `qa`, `code_generation`, `fine_tuning`, `training`, `rag`. Anything more specific goes in the free-text `purposeLimitation`. +- **Rate limits** — `maxTokensPerRequest`, `maxTokensPerDay`. +- **Audit and retention** — `auditRequired`, `canStore`, `retentionPolicy` (`maxRetentionDays`, `requireDeletion`). --- -## 🌍 Sovereignty Constraints (NEW) +## 🌍 Sovereignty constraints (NEW) -Enforce **data residency** and **jurisdictional compliance** at the contract level: +Declare **where data may live** at the contract level, with the top-level `sovereignty` block: + ```yaml -# NEW in v0.7.1: Top-level sovereignty requirements -sovereignty: - jurisdiction: "EU" - dataResidency: - allowedRegions: - - "europe-west1" - - "europe-west3" - deniedRegions: - - "us-central1" - complianceFrameworks: - - "GDPR" - - "HIPAA" - crossBorderTransfer: - allowed: false +sovereignty: # NEW in v0.7.1 + jurisdiction: EU + allowedRegions: [europe-west1, europe-west3] + deniedRegions: [us-central1] + dataResidency: true # data must stay within the jurisdiction + crossBorderTransfer: false + regulatoryFramework: [GDPR, HIPAA] + enforcementMode: strict # strict | advisory | audit ``` -**Why this matters:** Global compliance requires infrastructure-level enforcement: - -- ✅ **Jurisdictional boundaries** — enforce EU, US, APAC data laws -- ✅ **Regional constraints** — specify allowed/denied cloud regions -- ✅ **Compliance frameworks** — declare GDPR, HIPAA, SOC2 requirements -- ✅ **Transfer controls** — block cross-border data movement +- **Jurisdiction** — `EU`, `US`, `UK`, `CA`, `AU`, `JP`, `CN`, `IN`, `BR`, `Global` or `Multi-Region`. +- **Regions** — `allowedRegions` / `deniedRegions`. +- **Frameworks** — `regulatoryFramework` from a fixed list (`GDPR`, `CCPA`, `CPRA`, `HIPAA`, `PIPEDA`, `LGPD`, `PDPA`, `POPIA`, `DPA`, `APPI`). +- **Transfers** — `crossBorderTransfer` (boolean) and, when transfers are allowed, `transferMechanisms` (`SCCs`, `BCRs`, `Adequacy`, `DPF`, `Consent`, `Derogation`). --- -## ⚙️ Provider-First Orchestration (NEW) +## ⚙️ Provider-first orchestration (NEW) -Direct invocation of **provider actions** as first-class orchestration tasks: +A build's `execution.orchestration` can describe its Airflow DAG, including **provider actions** as tasks: + ```yaml -# NEW in v0.7.1: Provider actions without wrapper operators -orchestration: - engine: "airflow" - tasks: - - taskId: "ensure_s3_bucket" - type: "provider_action" - provider: "aws.s3" - action: "ensure_bucket" - parameters: - bucket_name: "customer-data-lake" - region: "us-west-2" - - - taskId: "load_to_snowflake" - type: "provider_action" - provider: "snowflake.table" - action: "ensure" - parameters: - database: "ANALYTICS" - schema: "GOLD" - table: "CUSTOMER_360" - dependsOn: ["ensure_s3_bucket"] +builds: + - id: customer_360 + pattern: embedded-logic + engine: sql + properties: + sql: "SELECT * FROM staging.customers" + execution: + orchestration: # NEW in v0.7.1 + engine: airflow + airflow: + dagId: customer_360 + tasks: + - taskId: ensure_s3_bucket + type: provider_action + operator: ProviderActionOperator # see the note below + provider: aws # aws | gcp | azure | snowflake | databricks | kafka | kubernetes | local | custom + action: s3.ensure_bucket # . + params: + bucket_name: customer-data-lake + region: us-west-2 + - taskId: load_to_snowflake + type: provider_action + operator: ProviderActionOperator + provider: snowflake + action: table.ensure + params: + database: ANALYTICS + schema: GOLD + table: CUSTOMER_360 + dependencies: [ensure_s3_bucket] ``` -**Why this matters:** Simplifies multi-cloud orchestration: +::: warning Name each task's `operator` +In the 0.7.1 to 0.7.6 schemas, the conditional rules for `fluid_execute`, `bash` and `python` tasks also match any task that has **no** `operator` member, so such a task must satisfy all three at once (`params.contract_path`, a bash command and a Python callable) and is in practice rejected. Setting `operator` avoids it; any value that does not contain `FluidExecute`, `Bash` or `Python` leaves a `provider_action` task to its own rule (`provider`, `action` and `params` required). +::: + +Tasks can also declare the data products they wait on (`dataProductDependencies`), what they produce (`produces`), a cost estimate (`costEstimate`), and access control (`accessControl`); `taskDefaults.errorHandling` sets a retry strategy and which errors are retryable. -- ✅ **Native provider actions** — AWS, GCP, Azure, Snowflake primitives -- ✅ **No wrapper complexity** — direct action invocation -- ✅ **Cross-provider workflows** — multi-cloud pipelines without vendor lock-in -- ✅ **Strong typing** — provider-specific validation +A **top-level** `orchestration` key arrives in [0.7.2](/fluid/releases/0.7.2). Be aware that its `tasks` list is not defined by any FLUID schema up to 0.7.6: `orchestration` is an open object there, so a validator accepts any `orchestration.tasks` content without checking it. The task shape the schema does check is the one above, under `airflow.tasks`. --- -## 📊 Enhanced Access Control (NEW) +## 📊 Root-level access control (NEW) -Root-level **accessPolicy** for automated IAM binding generation: +The top-level `accessPolicy` declares grants for the whole product: + ```yaml -# NEW in v0.7.1: Declarative access grants -accessPolicy: +accessPolicy: # NEW in v0.7.1 grants: - principal: "group:data-analytics@company.com" - permissions: ["read", "select", "query"] + permissions: [read, select, query] resources: - "$.exposes[?(@.kind=='table')]" - - principal: "serviceAccount:pipeline@project.iam.gserviceaccount.com" - permissions: ["write", "insert", "update"] + permissions: [write, insert, update] conditions: ipRanges: ["10.0.0.0/8"] ``` -**Why this matters:** Infrastructure-as-code for data access: - -- ✅ **Automated IAM** — generate cloud IAM bindings from the FLUID spec -- ✅ **Resource targeting** — JSONPath expressions for fine-grained access -- ✅ **Conditional access** — IP restrictions, time windows -- ✅ **Audit-ready** — version-controlled access policies +Each grant requires only `principal`. `permissions` come from a fixed list (`read`, `select`, `query`, `write`, `insert`, `update`, `delete`, `admin`, `manage`; `create` is added in 0.7.2), `resources` holds JSONPath expressions selecting the exposes in scope, and `conditions` is an open object. --- -## 📈 Key improvements over v0.5.7 +## 📈 What 0.7.1 adds over 0.5.7 -| Feature | v0.5.7 | v0.7.1 | Impact | -|---|---|---|---| -| **AI Model Control** | ❌ None | ✅ agentPolicy | Govern AI/LLM data access | -| **Data Sovereignty** | ❌ Manual | ✅ sovereignty | Automated compliance enforcement | -| **Orchestration** | ⚠️ Abstract | ✅ Provider-first | Direct cloud provider actions | -| **Access Control** | ⚠️ Expose-level | ✅ Root-level accessPolicy | Centralized IAM automation | -| **Cross-Provider** | ⚠️ Complex | ✅ Native | Simplified multi-cloud workflows | -| **Task Dependencies** | ⚠️ Build-only | ✅ Data product deps | Richer dependency graphs | -| **Error Handling** | ⚠️ Basic | ✅ Categorized | Intelligent retry strategies | -| **Cost Tracking** | ⚠️ Estimated | ✅ Actual vs estimated | Budget enforcement | +| Capability | 0.5.7 | 0.7.1 | +|---|---|---| +| AI model control | — | `exposes[].policy.agentPolicy` | +| Data sovereignty | — | top-level `sovereignty` | +| Orchestration | — | `build(s).execution.orchestration` with Airflow DAG, tasks and provider actions | +| Access control | — | top-level `accessPolicy` | +| Task dependencies on data products | — | `airflow.tasks[].dataProductDependencies` | +| Error handling | — | `airflow.taskDefaults.errorHandling` | +| Cost tracking | — | `airflow.tasks[].costEstimate` | +| Triggers | `schedule`, `event`, `manual`, `dependency` | adds `dataset`, `schedule_and_dataset`, `timetable` | --- -## 🔄 100% backward compatible +## 🔄 Backward compatibility -**All v0.5.7 contracts work unchanged in 0.7.1:** +- ✅ No narrowing change from 0.5.7 (`scripts/check-compat.py`). +- ✅ Every addition is optional. -- ✅ No breaking changes -- ✅ New features are opt-in -- ✅ Existing patterns fully preserved -- ✅ Gradual adoption path - -**Migration is simple:** +**Migration:** + ```yaml -# Change version number — that's it! -fluidVersion: "0.7.1" # was "0.5.7" - -# Optionally add new features -agentPolicy: { ... } -sovereignty: { ... } -accessPolicy: { ... } +fluidVersion: "0.7.1" # was "0.5.7"; the 0.7.1 schema accepts only "0.7.1" ``` See the [**Schema Changelog**](/fluid/schema/changelog) for the full version history, or the [0.5.7 → 0.7.1 diff](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.5.7-to-0.7.1.md) for the auto-generated change list. diff --git a/docs/releases/0.7.2.md b/docs/releases/0.7.2.md index 5a8ac36..04f2d49 100644 --- a/docs/releases/0.7.2.md +++ b/docs/releases/0.7.2.md @@ -1,6 +1,6 @@ # What's New in FLUID 0.7.2 — Semantic Truth Engine -**FLUID 0.7.2** is the **Semantic Truth Engine** release. It is **additive** over 0.7.1 — every 0.7.1 contract remains valid unchanged — and completes the agentic contract: *agentPolicy* decides **whether** an AI agent may act on a data product, and *semantics* tells it **how to act correctly**. +**FLUID 0.7.2** is the **Semantic Truth Engine** release. It adds to 0.7.1 with **one narrowing change** — `notification` became a closed object, so a 0.7.1 contract with an extra member on a `notifications[]` entry is rejected by 0.7.2 ([details below](#backward-compatibility-with-0-7-1-one-narrowing-change)). It completes the agentic contract: *agentPolicy* decides **whether** an AI agent may act on a data product, and *semantics* tells it **how to act correctly**. --- @@ -8,6 +8,7 @@ Each `exposes[]` entry now accepts an optional `semantics` block that maps physical columns to business concepts — entities, measures, dimensions, and metrics — in a form agents and BI tools can reason about without re-deriving KPIs. + ```yaml # NEW in v0.7.2: machine-readable business logic on an exposed table exposes: @@ -17,19 +18,24 @@ exposes: semantics: entities: - name: customer - primaryKey: customer_id + type: primary # primary | foreign | unique | natural + expr: customer_id measures: - name: revenue - expr: "SUM(order_total)" agg: sum + expr: order_total + - name: active_customers + agg: count_distinct + expr: customer_id dimensions: - name: signup_month + type: time # categorical | time expr: "DATE_TRUNC('month', created_at)" metrics: - name: monthly_active_customers - type: simple - measure: customer_id - filters: ["last_active_at >= CURRENT_DATE - INTERVAL 30 DAY"] + type: simple # simple | derived | ratio + measure: active_customers + filter: "last_active_at >= CURRENT_DATE - INTERVAL 30 DAY" ``` **Why this matters:** eliminates the *semantic hallucination* failure mode where an LLM asked "what's our MRR?" invents an SQL expression because the contract never told it how MRR is defined. The shape is aligned with dbt MetricFlow, Snowflake Semantic Views, and OSI-format metric definitions — portable across engines. @@ -40,6 +46,7 @@ exposes: `binding.icebergConfig` lets contracts declare Iceberg table format specifics — write version, file format, partition spec, sort order — directly in the FLUID contract, so provisioners don't need an out-of-band config file. + ```yaml binding: platform: aws @@ -54,7 +61,7 @@ binding: --- -## 🔄 Backward compatibility with 0.7.1 — one narrowing change +## Backward compatibility with 0.7.1: one narrowing change - ✅ `semantics` and `icebergConfig` are both opt-in - ✅ All 0.7.1 features (agentPolicy, sovereignty, provider-first orchestration, root accessPolicy) fully preserved @@ -68,6 +75,7 @@ gained `"additionalProperties": false` in 0.7.2, so a contract carrying **any** member on a `build.execution.notifications[]` entry validated under 0.7.1 and is rejected by 0.7.2: + ```yaml build: execution: @@ -90,10 +98,11 @@ field the schema does declare is unchanged. **Migration:** + ```yaml -# Change version number — that's it. -fluidVersion: "0.7.2" # was "0.7.1" -# Optionally add a semantics block on any expose. +fluidVersion: "0.7.2" # was "0.7.1"; the 0.7.2 schema accepts only "0.7.2" ``` +Change the version, and drop any member the 0.7.2 `notification` object does not declare. Then add a `semantics` block on any expose if you want one. + See the [**Schema Changelog**](/fluid/schema/changelog) for the full version history, or the [0.7.1 → 0.7.2 diff](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.7.1-to-0.7.2.md) for the auto-generated change list. diff --git a/docs/releases/0.7.3.md b/docs/releases/0.7.3.md index 8233e6c..5f7715c 100644 --- a/docs/releases/0.7.3.md +++ b/docs/releases/0.7.3.md @@ -1,6 +1,6 @@ # What's New in FLUID 0.7.3 — Source-Aligned Acquisition -**FLUID 0.7.3** is the **Source-Aligned Acquisition** release. It is **additive** over 0.7.2 — every existing contract remains valid — and finally makes *ingestion from external systems* a first-class FLUID concept rather than something you bolt on with Airbyte/Meltano configs outside the contract. +**FLUID 0.7.3** is the **Source-Aligned Acquisition** release. It is **additive** over 0.7.2 — a valid 0.7.2 contract stays valid when it declares `"0.7.3"` — and makes *ingestion from external systems* a first-class FLUID concept rather than something you bolt on with Airbyte/Meltano configs outside the contract. --- @@ -8,6 +8,7 @@ The `build` block gains a fourth pattern, `acquisition`, designed for source-aligned data products: contracts that *ingest* an external system (Salesforce, Postgres, Kafka, files, …) into your mesh, with no transformation logic. + ```yaml build: pattern: acquisition # tells the schema that build.properties below is an acquisitionPattern @@ -19,7 +20,7 @@ build: mode: incremental_dedup # 6 modes: full_refresh | incremental_* (3) | cdc | streaming cursor_field: updated_at connection: - secretRef: "vault://pg-prod-readonly" # must be a URI: vault:// aws:// gcp:// azure:// env:// + secretRef: "vault://pg-prod-readonly" # a URI (://…), e.g. vault://, aws://, gcp://, azure://, env:// streams: [public.customers, public.orders] sink: format: iceberg # iceberg | delta | parquet | snowflake_table | bigquery_table | … @@ -102,6 +103,7 @@ Runners publish what they can do (`full_refresh`, `incremental_dedup`, `cdc`, `s For production-grade ingestion, the connector image you run is part of the data supply chain: + ```yaml airbyte: image_signature: @@ -114,6 +116,7 @@ airbyte: ## 🧹 Top-level `retention` block + ```yaml retention: runState: P30D # ISO-8601 durations @@ -126,16 +129,42 @@ A single sweeper job honors these — no more per-tool TTL knobs scattered acros --- -## 🔄 100% backward compatible with 0.7.2 +## 🧩 Also new in 0.7.3 -- ✅ No breaking changes -- ✅ `acquisition`, `retention`, `build.capabilities`, and the six new engines are all opt-in -- ✅ All 0.7.2 semantics + 0.7.1 agentPolicy/sovereignty/accessPolicy fully preserved +The sections above cover acquisition. The same release added these, all optional: + +| Addition | What it declares | +|---|---| +| Top-level `governance.lakeFormation` | Account-wide AWS Lake Formation settings: `admins` and LF-tag `tagDefinitions`. **`admins` is authoritative** — see the warning below. | +| `exposes[].binding.governance.lakeFormation` | Per-resource Lake Formation settings: `registerLocation`, principal `grants`, LF-tag associations (`tags`) and a `rowFilter`. | +| Top-level `extensions` | An open object for vendor- or plugin-namespaced configuration. | +| `metadata.productType` | `SDP` (source-aligned) \| `ADP` (aggregated) \| `CDP` (consumption-aligned) — the data-mesh product type. | +| `metadata.classification`, `metadata.experimental` | A default classification (`public` \| `internal` \| `confidential` \| `restricted`) and a list of experimental features the contract opts into. | +| `exposes[].contract.schemaPolicy` | How the output schema may change: `strict` \| `discover_and_freeze` \| `evolve_safe` \| `evolve_all`. | +| `exposes[].observability.alert` | Alert channels (`log`, `file`, `webhook`) for acquisition alerts such as DLQ overflow or a changed source schema. | +| `binding.format` values | `snowflake_view`, `redshift_table`, `redshift_serverless`, `redshift_external_schema`. | +| `runtime.platform` values | `athena`, `glue`, `redshift`. | +| Identifiers | `$defs/identifier` now also admits upper-case letters, so existing catalog ids round-trip unchanged. | + +::: warning Lake Formation `admins` replaces the admin list +`governance.lakeFormation.admins` is **authoritative, not additive**: applying it replaces the account's Lake Formation admin list, so an admin you do not list is removed — including the role or user that runs the apply. The 0.7.3, 0.7.4 and 0.7.5 schemas first published here described it as additive; their description has since been corrected (see the [Changelog](/fluid/schema/changelog#corrections-to-published-schemas)). + +The corrected description adds: "The terraform-provider-aws resource also clears the data lake settings its configuration omits, and this block sets only the admins, so the account's create-database and create-table default permissions (the IAMAllowedPrincipals default for new databases and tables), trusted resource owners and parameters are cleared too. Destroying the resource empties the admin list and resets CROSS_ACCOUNT_VERSION to 1." +::: + +--- + +## 🔄 Backward compatibility with 0.7.2 + +- ✅ No narrowing change: `scripts/check-compat.py --from 0.7.2 --to 0.7.3` reports none. (It flags the identifier pattern change for review, because pattern containment cannot be decided in general; the change only widens the pattern, and the waiver in `scripts/compat-waivers.txt` records the exhaustive check that showed it.) +- ✅ Every addition in this release is optional. **Migration:** ```yaml -fluidVersion: "0.7.3" # was "0.7.2" — that's it +fluidVersion: "0.7.3" # was "0.7.2" ``` +The 0.7.3 schema accepts only `"0.7.3"` as `fluidVersion`, so this line is the one change a valid 0.7.2 contract needs. + See the [**Schema Changelog**](/fluid/schema/changelog) for the full version history, or the [0.7.2 → 0.7.3 diff](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.7.2-to-0.7.3.md) for the auto-generated change list. diff --git a/docs/releases/0.7.4.md b/docs/releases/0.7.4.md index 0326e99..b09f811 100644 --- a/docs/releases/0.7.4.md +++ b/docs/releases/0.7.4.md @@ -12,6 +12,7 @@ The headline change is behavioral, not structural: `policy.agentPolicy`'s shape A new `exposeMcp` block declares an expose as agent-consumable over MCP and carries per-expose overrides for the output-port server: + ```yaml exposes: - exposeId: customer_profiles @@ -44,7 +45,7 @@ Requests that violate the policy are rejected at the gateway rather than served. When a column is classified as PII/PHI, the gateway **redacts the values, not the column**: the column stays visible in the response shape, but each value is replaced with `[REDACTED-PII]`. Agents still see the schema and can reason about structure without ever receiving the sensitive payload. -> `phi` is **not new** in 0.7.4 — it has been a value in the column `sensitivity` enum since 0.7.3. +> `phi` is **not new** in 0.7.4 — it has been a value in the column `sensitivity` enum since 0.5.7. --- @@ -53,18 +54,19 @@ When a column is classified as PII/PHI, the gateway **redacts the values, not th - **`binding.platform`** gains **`postgres`** — for the MCP output-port PostgreSQL driver. Existing platforms (`gcp`, `aws`, `azure`, `snowflake`, `databricks`, `kafka`, `local`, `kubernetes`, `other`) are unchanged. - **`binding.format`** gains **`postgres_table`**, **`athena_table`**, and **`glue_table`** — added alongside (not replacing) the existing formats including `snowflake_view` and the `redshift_*` values. + ```yaml binding: platform: postgres # NEW in v0.7.4 format: postgres_table # NEW in v0.7.4 (also: athena_table, glue_table) - location: { ... } + location: { database: analytics, schema: public, table: customer_profiles } ``` `fluidVersion` also moves from `const: "0.7.3"` to `enum: ["0.7.3", "0.7.4"]`, so existing `0.7.3` values still validate against the 0.7.4 schema. --- -## Backward compatibility — fully additive +## Backward compatibility: fully additive **0.7.4 is fully backward-compatible with 0.7.3.** Every valid 0.7.3 contract validates as 0.7.4 unchanged — the release only adds optional fields and enum values. Nothing was removed: @@ -77,11 +79,15 @@ The only schema-level change beyond the additions is that `fluidVersion` moves f ### Upgrading + ```yaml -fluidVersion: "0.7.4" # optional — "0.7.3" still validates against the 0.7.4 schema +fluidVersion: "0.7.4" ``` -Bumping the version is all it takes — there are no removed fields to migrate off. The new `exposes[].mcp` block, the `postgres` platform, and the `postgres_table` / `athena_table` / `glue_table` formats are all optional and additive. +- **To use any field added in 0.7.4** (`exposes[].mcp`, `platform: postgres`, the new formats), declare `"0.7.4"`. A document is validated against the schema of the version it declares, so a contract that declares `"0.7.3"` and uses `exposes[].mcp` is invalid, even though the 0.7.4 schema's `fluidVersion` enum also accepts `"0.7.3"`. +- **If you use no new field, bumping is optional** — the contract stays valid either way. + +Nothing was removed, so there is nothing to migrate off. See [Choosing `fluidVersion`](/fluid/schema/versions#choosing-fluidversion). --- diff --git a/docs/releases/0.7.5.md b/docs/releases/0.7.5.md index 3fa5979..0abb337 100644 --- a/docs/releases/0.7.5.md +++ b/docs/releases/0.7.5.md @@ -1,83 +1,185 @@ # What's New in FLUID 0.7.5 — Streaming Kafka → Iceberg Sink & Confluent Tableflow -**FLUID 0.7.5** extends the acquisition layer with a **streaming Kafka → Iceberg sink** and a managed **Confluent Cloud Tableflow** binding, and generalizes `build.engine` so dbt **adapter-qualified** engines (e.g. `dbt-databricks`) validate. It is the contract side of the streaming-ingestion RFC shipped in the forge CLI 0.9.0 line. +**FLUID 0.7.5 is the latest stable version.** It adds an opt-in Iceberg streaming sink to the `kafka-connect` acquisition engine, two binding platforms — `confluent` (Confluent Cloud Tableflow) and `pgvector` (a vector / embeddings output port) — and location fields for Iceberg catalogs, Redshift Serverless and Kinesis. -> ✅ **0.7.5 is fully backward-compatible with 0.7.4.** This release is purely **additive** — every valid 0.7.4 contract validates as 0.7.5 unchanged. See [Backward compatibility — fully additive](#backward-compatibility-fully-additive) for details. +> ✅ **Additive over 0.7.4.** `scripts/check-compat.py --from 0.7.4 --to 0.7.5` finds no narrowing change: every valid 0.7.4 contract stays valid when it declares `"0.7.5"`. To *use* a 0.7.5 field, the contract must declare `fluidVersion: "0.7.5"` — see [Upgrading](#upgrading). + +- **Schema:** [`fluid-schema-0.7.5.json`](/fluid/schema/fluid-schema-0.7.5.json) · **Reference:** [`specs/0.7.5/fluid-spec.html`](/fluid/specs/0.7.5/fluid-spec.html) · **Diff:** [`diff-0.7.4-to-0.7.5.md`](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.7.4-to-0.7.5.md) --- -## 🧊 Streaming Iceberg sink on `kafka-connect` acquisition (NEW) +## Streaming Iceberg sink on `kafka-connect` acquisition -The `kafka-connect` acquisition pattern gains an opt-in Iceberg streaming sink (RFC §6.2). All fields are optional and **off by default** — a hand-written `sink_connector_config` is untouched unless you opt in. +The `kafka-connect` engine's settings gain an opt-in Iceberg streaming sink. They live where every acquisition engine's settings live: under `builds[].properties`, which the schema validates as an acquisition pattern when `pattern: acquisition`, in a key named after the engine. ```yaml +fluidVersion: "0.7.5" +kind: DataProduct +id: sales.bronze.orders_stream +name: Orders (streamed) +metadata: + owner: { team: sales-data } builds: - - name: ingest_orders + - id: ingest_orders + pattern: acquisition engine: kafka-connect - acquisition: + capabilities: [streaming] + outputs: [orders] + properties: + source: # required for pattern: acquisition + kind: kafka + mode: streaming + connection: { secretRef: "env://KAFKA_BOOTSTRAP" } + streams: [orders, payments] kafka-connect: - iceberg_sink_enabled: true # NEW — opt-in; derive the Iceberg sink connector config - sink_topics: [orders, payments] # NEW — explicit topics the sink consumes (else derived from source streams) - streamingSink: # NEW — optional tuning (all keys optional) + iceberg_sink_enabled: true # NEW — opt in to a derived Iceberg sink connector + sink_topics: [orders, payments] # NEW — topics the sink consumes (else derived from source.streams) + streamingSink: # NEW — optional tuning, every key optional commitIntervalMs: 60000 dynamicEnabled: true routeField: _topic autoCreate: true evolveSchema: true - iceberg_catalog_overrides: # NEW — operator escape hatch, merged LAST over the derived config - iceberg.catalog.warehouse: s3://lake/warehouse + iceberg_catalog_overrides: # NEW — string → string, merged last over the derived config + iceberg.catalog.warehouse: "s3://acme-lake/warehouse" +exposes: + - exposeId: orders + kind: table + contract: + schema: + - { name: order_id, type: STRING, required: true } + - { name: amount, type: NUMERIC } + binding: + platform: aws + format: iceberg + location: + bucket: acme-lake + database: sales + table: orders + region: eu-west-1 + catalog: glue # NEW — Iceberg catalog kind + warehouse: "s3://acme-lake/warehouse" # NEW — Iceberg warehouse ``` -- **`iceberg_sink_enabled`** — boolean opt-in. Defaults OFF when a hand-written `sink_connector_config` is present. -- **`sink_topics`** — explicit topic list the derived sink consumes; otherwise derived from the source streams. -- **`streamingSink`** — optional sink tuning (`commitIntervalMs`, `routeField`, `dynamicEnabled`, `autoCreate`, `evolveSchema`, …). All keys optional. -- **`iceberg_catalog_overrides`** — a `string → string` map merged **last** over the derived Iceberg sink config, for operator escape-hatch keys. +- **`iceberg_sink_enabled`** — boolean opt-in. The schema description adds that it defaults to off when a hand-written `sink_connector_config` is present. +- **`sink_topics`** — the topics the derived sink consumes; when absent, they are derived from the source streams. +- **`streamingSink`** — a closed object of optional tuning keys: `commitIntervalMs`, `routeField`, `dynamicEnabled`, `autoCreate`, `evolveSchema`, `upsertMode`, `controlTopic`. +- **`iceberg_catalog_overrides`** — a `string → string` map merged last over the derived sink configuration. --- -## ☁️ Confluent Cloud Tableflow binding (NEW, additive) +## `confluent`: Confluent Cloud Tableflow binding + +`binding.platform` gains `confluent`, and `binding.location` gains the fields Tableflow needs: -- **`binding.platform`** gains **`confluent`** — the Confluent Cloud Tableflow managed Kafka → Iceberg emitter. Existing platforms (`gcp`, `aws`, `azure`, `snowflake`, `databricks`, `kafka`, `local`, `kubernetes`, `postgres`, `other`) are unchanged. -- **`binding.location`** gains three Tableflow fields: - - **`environment_id`** — Confluent Cloud environment id (`env-xxxxx`). - - **`kafka_cluster_id`** — the Kafka cluster id (`lkc-xxxxx`) the Tableflow topic belongs to. - - **`confluent_role_arn`** — ARN of the pre-created AWS IAM role Confluent Tableflow assumes (BYOB AWS). +| Field | Meaning | +|---|---| +| `environment_id` | Confluent Cloud environment id (`env-…`). | +| `kafka_cluster_id` | The Kafka cluster (`lkc-…`) the topic belongs to. | +| `confluent_role_arn` | ARN of the pre-created AWS IAM role Tableflow assumes when it writes to your own bucket. | + ```yaml binding: - platform: confluent # NEW in v0.7.5 + platform: confluent # NEW in 0.7.5 + format: iceberg location: - environment_id: env-abc12 - kafka_cluster_id: lkc-7x8y9 - confluent_role_arn: arn:aws:iam::123456789012:role/tableflow-byob + environment_id: env-abc12 # NEW + kafka_cluster_id: lkc-7x8y9 # NEW + confluent_role_arn: "arn:aws:iam::123456789012:role/tableflow-byob" # NEW + bucket: acme-tableflow + database: sales + topic: orders + region: eu-west-1 ``` +The schema only adds the fields. The reference implementation also requires, for a `confluent` binding, `format: iceberg` and all of `environment_id`, `kafka_cluster_id`, `bucket` and `confluent_role_arn`; `fluid validate` reports each one that is missing. + --- -## `fluidVersion` widened (additive) +## `pgvector`: vector / embeddings output port -`fluidVersion` widens from `enum: ["0.7.3", "0.7.4"]` to `enum: ["0.7.3", "0.7.4", "0.7.5"]`, so existing `0.7.3` / `0.7.4` values still validate against the 0.7.5 schema. +`binding.platform` gains `pgvector`, `binding.format` gains `pgvector_table`, and a binding can carry `vectorConfig`: + + +```yaml +kind: vector +contract: + schema: + - { name: ticket_id, type: STRING, required: true } + - { name: body, type: TEXT, labels: { ai-embeddable: "true" } } +binding: + platform: pgvector # NEW in 0.7.5 + format: pgvector_table # NEW in 0.7.5 + location: { database: support, schema: public, table: tickets } + vectorConfig: # NEW in 0.7.5 + dimensions: 1536 # required — 1 to 16000 + embeddingModel: text-embedding-3-small + vectorType: vector # vector | halfvec + distanceMetric: cosine # cosine | l2 | inner_product | l1 + indexType: hnsw # hnsw | ivfflat | none + hnsw: { m: 16, efConstruction: 64 } + sourceKeyColumn: ticket_id +``` + +`vectorConfig` is a closed object; only `dimensions` is required. `hnsw` (`m`, `efConstruction`) and `ivfflat` (`lists`) carry index tuning, and `table` names the embeddings table. Which columns get embedded is not part of the schema: the reference implementation embeds the columns labelled `ai-embeddable: "true"`, and `fluid validate` warns when there are none. --- -## Backward compatibility — fully additive +## New `binding.location` fields -**0.7.5 is fully backward-compatible with 0.7.4.** Every valid 0.7.4 contract validates as 0.7.5 unchanged — the release only adds optional fields and enum values. Nothing was removed: +| Fields | For | +|---|---| +| `catalog`, `warehouse`, `uri`, `partitionBy` | Iceberg catalogs: the catalog kind (for example `glue` or `rest`), the warehouse, a REST catalog endpoint, and partition columns. | +| `namespace`, `workgroup`, `iam_role_arn`, `external_schema`, `glue_database` | Amazon Redshift Serverless, and a Redshift external schema over a Glue database. | +| `stream` | An Amazon Kinesis Data Stream. | +| `environment_id`, `kafka_cluster_id`, `confluent_role_arn` | Confluent Tableflow (above). | -- The streaming-sink keys (`iceberg_sink_enabled`, `sink_topics`, `streamingSink`, `iceberg_catalog_overrides`) and the Confluent Tableflow binding fields are all **optional**. -- All 0.7.4 surfaces — `exposes[].mcp`, the `postgres` platform, the Lake Formation `governance` blocks, the `postgres_table` / `athena_table` / `glue_table` formats — are unchanged. + +```yaml +binding: + platform: aws + format: redshift_external_schema + location: + region: eu-west-1 + namespace: analytics # NEW + workgroup: analytics-wg # NEW + iam_role_arn: "arn:aws:iam::123456789012:role/redshift-spectrum" # NEW + external_schema: sales_ext # NEW + glue_database: sales # NEW + table: orders +``` + +--- + +## `fluidVersion` + +The `fluidVersion` enum is `["0.7.3", "0.7.4", "0.7.5"]`. That lets the 0.7.5 schema accept documents that declare an older version; it does not let them use 0.7.5 fields, because a document is validated against the schema of the version it declares. -### Upgrading +--- + +## Upgrading ```yaml -fluidVersion: "0.7.5" # optional — "0.7.3" / "0.7.4" still validate against the 0.7.5 schema +fluidVersion: "0.7.5" ``` -Bumping the version is all it takes — there are no removed fields to migrate off. The streaming Iceberg sink, the `confluent` binding, and the adapter-qualified dbt engines are all optional and additive. +- **To use any field above, declare `"0.7.5"`.** A contract that declares `"0.7.4"` and uses, say, `platform: confluent` is invalid: its schema is 0.7.4, which does not have that value. The reference implementation reports exactly that against schema v0.7.4. +- **If you use no new field, bumping is optional** — the contract stays valid either way. +- Point your editor's `# yaml-language-server: $schema=…` line at `fluid-schema-0.7.5.json` when the file declares 0.7.5, and at the matching older schema when it does not. See [Choosing `fluidVersion`](/fluid/schema/versions#choosing-fluidversion). + +Nothing was removed, so there is nothing to migrate off. --- +## Corrections to this page + +- An earlier version of this page said that 0.7.5 widened `build.engine` to accept adapter-qualified dbt engines such as `dbt-databricks`. That pattern (`^dbt-[a-z0-9]+([_-][a-z0-9]+)*$`) has been in the schema since **0.7.2**; `build.engine` is identical in 0.7.4 and 0.7.5. +- Its headline example put the sink under `builds[].acquisition` with a `name` key. A build has neither key; the example above shows where the settings go, inside a complete contract that validates. +- The 0.7.5 schema first published on this site was a snapshot taken while 0.7.5 was still a preview in the reference implementation. The stable 0.7.5 also has the `pgvector` platform, `pgvector_table`, `vectorConfig` and the Iceberg-catalog, Redshift Serverless and Kinesis location fields; the published file has been re-synchronised with it and this page now describes it. See the [Changelog](/fluid/schema/changelog#_0-7-4-to-0-7-5-streaming-kafka-to-iceberg-sink-confluent-tableflow). + ## References -- [**Schema Changelog**](/fluid/schema/changelog) — the field-by-field machine diff for every version. -- [Full 0.7.4 → 0.7.5 diff](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.7.4-to-0.7.5.md) — the complete auto-generated change list (additive: the streaming Iceberg sink on `kafka-connect`, the `confluent` Tableflow binding, and `fluidVersion` widened to include `0.7.5`). +- [**Schema Changelog**](/fluid/schema/changelog) — the history of every version. +- [Full 0.7.4 → 0.7.5 diff](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.7.4-to-0.7.5.md) — the auto-generated change list. +- What comes next: [0.7.6 (preview)](/fluid/releases/0.7.6). diff --git a/docs/releases/0.7.6.md b/docs/releases/0.7.6.md new file mode 100644 index 0000000..fd6187e --- /dev/null +++ b/docs/releases/0.7.6.md @@ -0,0 +1,293 @@ +# FLUID 0.7.6 (preview) — Declarative Packaging Modes + +::: warning Preview — not yet stable +0.7.6 is a **preview**. A contract uses it only by declaring `fluidVersion: "0.7.6"`, and its fields and their meaning can still change before it is promoted to stable. The latest stable version is [0.7.5](/fluid/releases/0.7.5). See [Stable and preview versions](/fluid/schema/versions#stable-and-preview-versions). +::: + +0.7.6 adds declarative ownership of infrastructure containers (`packaging`), a declared list of downstream consumers, pins for upstreams in another mesh, and binding-level fields that let one cloud-neutral contract carry encryption, identity mapping and data expiry to each platform it is deployed on. + +- **Schema:** [`fluid-schema-0.7.6.json`](/fluid/schema/fluid-schema-0.7.6.json) · **Reference:** [`specs/0.7.6/fluid-spec.html`](/fluid/specs/0.7.6/fluid-spec.html) · **Diff:** [`diff-0.7.5-to-0.7.6.md`](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.7.5-to-0.7.6.md) +- **Compatibility:** additive over 0.7.5. `scripts/check-compat.py --from 0.7.5 --to 0.7.6` reports no narrowing change, so a valid 0.7.5 contract stays valid when it declares `"0.7.6"`. + +--- + +## What is new + +Every addition below is optional. The list is derived from the schemas: these are the property paths present in 0.7.6 and absent in 0.7.5. + +| Path | What it declares | +|---|---| +| `packaging` | Contract-wide ownership mode for the containers a product writes into. | +| `exposes[].binding.packaging` | A per-binding override of `packaging`. | +| `consumers[]` | The downstream artifacts (dashboards, notebooks, applications…) built on this product. | +| `consumes[].upstreamWorkspace`, `consumes[].upstreamDigest` | A pin on an upstream that lives in another mesh. | +| `exposes[].semantics.measures[].aggParams` | Parameters for an aggregation, used by `agg: percentile`. | +| `exposes[].binding.encryption.kms` | The key that data at rest is encrypted with. | +| `exposes[].binding.principals` | What each logical principal named in the contract is on this binding's platform. | +| `exposes[].binding.governance.lakeFormation.bucketPolicy` | Which grantees may read a Lake Formation location directly from object storage. | +| `exposes[].lifecycle.expire` | Whether data older than the expose's `retention` is deleted. | +| `exposes[].policy.privacy.masking[].params.{keepFirst, keepLast, saltEnv, keyEnv}` | Typed parameters for masking rules. | + +`fluidVersion` accepts `"0.7.6"` in addition to `"0.7.3"`, `"0.7.4"` and `"0.7.5"`. + +0.7.6 also writes down the meaning of a field that already existed: `exposes[].policy.authz.columnRestrictions` ([below](#column-restrictions-semantics-written-down)). + +::: tip Reading the 0.7.6 schema descriptions +Many 0.7.6 field descriptions explain what the reference implementation does with the field (`fluid apply`, the OpenTofu resources it emits, `fluid verify`). Those passages are **non-normative**: they describe one implementation. A FLUID document is valid or invalid by its schema; an implementation that validates the same documents conforms whether or not it provisions the same resources. +::: + +--- + +## `packaging` — who owns the containers + + +```yaml +packaging: + mode: shared # isolated | shared — the default for every container kind + pool: analytics-eu # id of the platform-owned pool the product writes into + containers: # per-kind overrides: bucket | database | dataset | schema | warehouse | cluster + warehouse: isolated # hybrid tier: shared database, owned warehouse +``` + +- `isolated` — the product creates and owns the container. +- `shared` — the container already exists and belongs to the platform; the product writes into it but does not create or own it. +- `exposes[].binding.packaging` has the same shape and overrides `packaging` key by key for that binding. An absent `packaging` block declares nothing, and the 0.7.6 schema describes that case as the behaviour from before the block existed. + +The schema constrains the shape; it does not require `pool`. The reference implementation refuses a contract in which any container resolves to `shared` and no `pool` is set (`packaging (pool-required)`), and refuses `isolated` for the `cluster` container kind on a Confluent or Kafka binding (`cluster-isolated-unsupported`). Which platform resource each container kind stands for is implementation-defined; the 0.7.6 schema's description of `containers` lists the reference implementation's mapping. + +## `consumers` — what is built on this product + + +```yaml +consumers: + - name: weekly_revenue_dashboard # required — identifier, unique within the contract + type: dashboard # required — dashboard | notebook | analysis | ml | application + label: Weekly Revenue Dashboard + maturity: high # high | medium | low + owner: { team: finance-analytics, email: finance-analytics@example.com } + url: https://bi.example.com/dashboards/42 + exposeIds: [orders] # which output ports it reads; absent = all +``` + +The shape follows dbt exposures. The block is purely declarative: the 0.7.6 schema says consumers never affect plan or apply, and marks the block as **reserved** — the reference implementation carries it through but does not yet act on it. + +## Cross-mesh pins on `consumes[]` + + +```yaml +consumes: + - productId: crm.bronze.accounts + exposeId: accounts + upstreamWorkspace: crm-mesh # the other mesh that owns this upstream + upstreamDigest: "sha256:4f1c9a0d6e2b7c8a1f3e5d9b0c2a4e6f8d1b3c5e7a9f0b2d4c6e8a1b3d5f7091" +``` + +- `upstreamWorkspace` names the federated workspace that owns the upstream. Omit it for an upstream in the same workspace. +- `upstreamDigest` is the digest of the upstream contract this product was composed against, written `sha256:` followed by 64 lowercase hex characters. +- **Setting `upstreamWorkspace` requires `upstreamDigest`.** The schema states this with `dependentRequired`, a JSON Schema Draft 2020-12 keyword. 0.7.6 is the first FLUID schema to use a keyword that Draft 7 does not have, so a validator that applies Draft 7 rules ignores this requirement. Every FLUID schema declares Draft 2020-12 (see [Validation semantics](/fluid/schema/specification#validation-semantics)). + +What the digest is computed over is **not defined by the standard**; it is implementation-defined. In the reference implementation it is a SHA-256 over the parsed, canonicalised upstream contract, printed by `fluid contract digest `, and `fluid apply` compares it with the upstream's current digest to detect that the upstream changed since composition. + +## `aggParams` on semantic measures + + +```yaml +semantics: + measures: + - name: p95_amount + agg: percentile + expr: amount + aggParams: + percentile: 0.95 # in [0, 1]; 0.5 is the median + useDiscretePercentile: false # discrete instead of interpolated percentile +``` + +`aggParams` mirrors `agg_params` in dbt-semantic-interfaces. In 0.7.6 it is used by `agg: percentile`. + +## Binding fields: encryption, principals, Lake Formation bucket policy + + +```yaml +binding: + platform: gcp + format: bigquery_table + location: { project: acme-analytics, dataset: sales, table: orders, region: europe-west1 } + encryption: + kms: product # product (default) | none | a key reference + principals: # logical principal -> identity on this platform + "group:analysts": "group:analysts@example.com" + "group:stewards": ["group:data-stewards@example.com"] +``` + +**`encryption.kms`** accepts `product`, `none`, an AWS KMS alias (`alias/`), an AWS KMS key or alias ARN, or a Google Cloud KMS key name (`projects/

    /locations//keyRings//cryptoKeys/`). The schema refuses an `alias/…` or `arn:…` value on a `gcp` binding, and a `projects/…` value on an `aws` binding. In the 0.7.6 schema's description, `product` means a customer-managed key created for this product, and `none` means the platform's default encryption. + +**`principals`** maps the principals the contract names — in `accessPolicy.grants[].principal` and `exposes[].policy.authz.columnRestrictions[].principal` — to identities on this binding's platform, so the base contract stays cloud-neutral and each environment binds it. A value is one identity or a list of them; `[]` says the principal has no identity on this platform. On a `gcp` binding each identity must be an IAM member (`user:`, `group:`, `serviceAccount:` or `domain:`), and on an `aws` binding an IAM principal ARN; the schema enforces both. The schema's description adds that, when the block is present, every principal the contract names for the expose must be mapped; the reference implementation refuses an unmapped one. + +**`governance.lakeFormation.bucketPolicy`** (`cross-account` (default) | `none` | `all-grantees`) decides which Lake Formation grantees also get a statement in the S3 bucket policy emitted beside the grants. A bucket-policy statement lets its principal read the objects directly, which bypasses Lake Formation's column, row and cell filters; `none` emits no bucket policy at all. + +::: danger An emitted bucket policy is authoritative +The 0.7.6 schema's description of `bucketPolicy` ends: "An emitted aws_s3_bucket_policy is authoritative: it replaces every other statement on the bucket. When no grantee is left none is emitted." Applying a contract that emits a bucket policy therefore replaces every other statement on that bucket. + +Whether one is emitted depends on the value. `cross-account`, **the default**, gives a statement only to principals "whose ARN names an AWS account other than the one running the apply, decided at plan time against the caller identity"; same-account grantees get none, "since on a registered location Lake Formation vends them credentials". So a contract that never sets `bucketPolicy` still has an emitted policy replace every other statement on the bucket as soon as one of its grantees is in another AWS account. `all-grantees` gives every grantee a statement, same-account included, "which is the output before this field existed and reopens the direct S3 read path". `none` emits no bucket policy, so nothing is replaced. + +A statement (`s3:ListBucket` and `s3:GetBucketLocation` on the bucket, `s3:GetObject` under `location.path`) "lets its principal read the objects straight from S3, which skips Lake Formation's column, row and cell filters and its revocations": Lake Formation's filters do not apply to that read, and revoking the grant in Lake Formation does not close it. + +The schema describes `none` as "the tightest setting when the location is registered (registerLocation) and every reader, cross-account ones included, queries through a Lake Formation integrated engine, or when the bucket policy is managed elsewhere". Like the other 0.7.6 descriptions of what `fluid apply` emits, this is a non-normative description of the reference implementation (see the tip above). +::: + +## `exposes[].lifecycle.expire` + + +```yaml +lifecycle: + retention: P400D # ISO-8601 duration + expire: true # default false: delete data older than retention +``` + +An expose's `lifecycle` has the same fields as the contract-root `lifecycle`, plus `expire`. With `expire: false` — the default — `retention` is a declaration only, so a contract that already declares a retention period does not start deleting data. With `expire: true`, data older than `retention` is deleted. How a platform expires data is implementation-defined. The reference implementation honours `expire` on an AWS S3 binding and on a BigQuery table, and refuses it on other GCP bindings rather than ignoring it. + +## Typed masking parameters + + +```yaml +policy: + privacy: + masking: + - { column: phone, strategy: mask, params: { keepLast: 4 } } # ********4567 + - { column: email, strategy: hash, params: { saltEnv: FLUID_PII_HASH_SECRET } } +``` + +| Param | Strategy | Meaning | +|---|---|---| +| `keepFirst`, `keepLast` | `mask` | Leading / trailing characters left unmasked (0–64; defaults 0 and 4). | +| `saltEnv` | `hash` | Name of the environment variable that holds the salt (default `FLUID_PII_HASH_SECRET`). | +| `keyEnv` | `tokenize`, `encrypt` | Name of the environment variable that holds the key. | + +`params` stays open (`additionalProperties: true`). The typed keys name secrets by environment variable rather than carrying them; the schema's description adds that the reference implementation's runner refuses any other key, including a literal salt or key. The 0.7.6 description of `strategy` also gives each strategy's output form — for example, `hash` is the lowercase hex SHA-256 of salt followed by value — and notes that `k_anonymity` is accepted by the schema but is a property of a whole table, not a per-value transform. + +## Column restrictions: semantics written down + +`exposes[].policy.authz.columnRestrictions` exists in 0.7.5 with the one-line description "Column-level access control". The 0.7.6 description defines it: + +- A column named in any restriction is restricted. +- `access: deny` — the principal may not read the columns. `access: allow` — the columns are readable only by the principals an `allow` names. +- A deny beats an allow, and a restriction never grants access: the readers are still the expose's readers. +- Principals are logical, resolved per platform through `binding.principals`. + +--- + +## Two complete contracts + +Both validate against the 0.7.6 JSON Schema (Draft 2020-12) and with the reference implementation, `data-product-forge` 0.18.1 (`fluid validate` prints `✅ Valid FLUID contract (schema v0.7.6)`). + +**BigQuery:** column restriction, mapped principals, a product-owned key, expiring partitions, a percentile measure, packaging and a consumer. + +```yaml +fluidVersion: "0.7.6" +kind: DataProduct +id: sales.gold.orders +name: Orders +metadata: + owner: { team: sales-analytics, email: sales-analytics@example.com } +accessPolicy: + grants: + - principal: "group:analysts" + permissions: [read] + - principal: "group:stewards" + permissions: [read] +exposes: + - exposeId: orders + kind: table + contract: + schema: + - { name: order_id, type: STRING, required: true } + - { name: order_date, type: DATE, required: true } + - { name: customer_email, type: STRING } + - { name: amount, type: NUMERIC } + policy: + authz: + columnRestrictions: + - principal: "group:analysts" + columns: [customer_email] + access: deny + semantics: + measures: + - name: p95_amount + agg: percentile + expr: amount + aggParams: { percentile: 0.95 } + lifecycle: + retention: P400D + expire: true + binding: + platform: gcp + format: bigquery_table + location: { project: acme-analytics, dataset: sales, table: orders, region: europe-west1 } + encryption: { kms: product } + principals: + "group:analysts": "group:analysts@example.com" + "group:stewards": "group:data-stewards@example.com" +packaging: + mode: isolated +consumers: + - name: weekly_revenue_dashboard + label: Weekly Revenue Dashboard + type: dashboard + maturity: high + owner: { team: finance-analytics } + exposeIds: [orders] +``` + +**S3 + Lake Formation:** a cross-mesh upstream pin, masking parameters, an existing key, a shared bucket and no bucket policy. + +```yaml +fluidVersion: "0.7.6" +kind: DataProduct +id: sales.silver.customers +name: Customers +metadata: + owner: { team: sales-data } +consumes: + - productId: crm.bronze.accounts + exposeId: accounts + upstreamWorkspace: crm-mesh + upstreamDigest: "sha256:4f1c9a0d6e2b7c8a1f3e5d9b0c2a4e6f8d1b3c5e7a9f0b2d4c6e8a1b3d5f7091" +exposes: + - exposeId: customers + kind: table + contract: + schema: + - { name: customer_id, type: STRING, required: true } + - { name: phone, type: STRING } + - { name: email, type: STRING } + policy: + privacy: + masking: + - { column: phone, strategy: mask, params: { keepLast: 4 } } + - { column: email, strategy: hash, params: { saltEnv: FLUID_PII_HASH_SECRET } } + binding: + platform: aws + format: parquet + location: { bucket: acme-lake, path: sales/customers/, region: eu-west-1, database: sales, table: customers } + encryption: { kms: "alias/acme-lake" } + packaging: { mode: shared, pool: lake-pool-eu, containers: { bucket: shared } } + governance: + lakeFormation: + registerLocation: true + bucketPolicy: none + grants: + - principal: "arn:aws:iam::123456789012:role/analyst" + permissions: [SELECT] +``` + +The example sets `bucketPolicy: none` on purpose. Its bucket is `shared`, so by the `packaging` section above it already exists and belongs to the platform, and the product writes into it without owning it; the location is registered (`registerLocation: true`), so Lake Formation governs the `analyst` grant. An emitted bucket policy would replace every other statement on that bucket and would give a grantee a path to the objects that skips Lake Formation's filters (see [the box above](#binding-fields-encryption-principals-lake-formation-bucket-policy)). With the default, `cross-account`, a grantee in another AWS account would get exactly such a statement. `none` is the value for a registered location whose readers query through a Lake Formation integrated engine, or whose bucket policy is managed elsewhere. + +Remove `upstreamDigest` from the second contract and it becomes invalid: `'upstreamDigest' is a dependency of 'upstreamWorkspace'` (keyword `dependentRequired`, at `/consumes/0`). + +--- + +## References + +- [Schema Versions](/fluid/schema/versions) — status of every version, and how to choose `fluidVersion`. +- [Changelog: 0.7.5 → 0.7.6](/fluid/schema/changelog#_0-7-5-to-0-7-6-preview-declarative-packaging-modes). +- Reference implementation: the [governance parity notes](https://github.com/Agenticstiger/forge-cli/blob/main/docs/governance-parity.md) describe what `data-product-forge` emits and verifies for retention, encryption, column restrictions and principals on AWS and GCP. diff --git a/docs/releases/README.md b/docs/releases/README.md index 05cb8bb..0908f15 100644 --- a/docs/releases/README.md +++ b/docs/releases/README.md @@ -1,21 +1,24 @@ # What's New -Release notes for the FLUID specification, newest first. Each entry links to the full page and the auto-generated schema diff. +Release notes for the FLUID specification, newest first. Each entry links to its full page and to the auto-generated schema diff. -| Version | Codename | Compatibility | Summary | -|---|---|---|---| -| **[0.7.5](/fluid/releases/0.7.5)** *(latest)* | Streaming Kafka → Iceberg Sink & Confluent Tableflow | ✅ Additive over 0.7.4 | Adds an opt-in Iceberg streaming sink on the `kafka-connect` acquisition pattern (`iceberg_sink_enabled`, `streamingSink`, `sink_topics`, `iceberg_catalog_overrides`) and the `confluent` Cloud Tableflow binding platform (`environment_id` / `kafka_cluster_id` / `confluent_role_arn`). | -| **[0.7.4](/fluid/releases/0.7.4)** | Runtime agentPolicy Enforcement at the MCP Gateway | ✅ Additive over 0.7.3 | Makes an expose agent-consumable over MCP and turns `policy.agentPolicy` into a gate the MCP output-port gateway enforces on every read. Adds the optional `exposes[].mcp` block, the `postgres` binding platform, and the `postgres_table` / `athena_table` / `glue_table` formats. | -| **[0.7.3](/fluid/releases/0.7.3)** | Source-Aligned Acquisition | ✅ Additive over 0.7.2 | Ingestion becomes a first-class FLUID concept: the `acquisition` build pattern, six ingestion engines, delivery guarantees + DLQ, schema-evolution policy, supply-chain image signing, and a top-level `retention` block. | -| **[0.7.2](/fluid/releases/0.7.2)** | Semantic Truth Engine | ✅ Additive over 0.7.1 | A per-expose `semantics` block (entities, measures, dimensions, metrics) gives agents and BI tools a verbatim metric definition, plus `binding.icebergConfig` for Iceberg table specifics. | -| **[0.7.1](/fluid/releases/0.7.1)** | Agentic Governance + Provider-First Orchestration | ✅ Additive over 0.5.7 | `agentPolicy`, `sovereignty`, root-level `accessPolicy`, and provider-first orchestration tasks. | +**0.7.5 is the latest stable version. 0.7.6 is a preview**: a contract uses it only by declaring `fluidVersion: "0.7.6"`, and it can still change before it is promoted. See [Stable and preview versions](/fluid/schema/versions#stable-and-preview-versions). + +| Version | Status | Codename | Compatibility | Summary | +|---|---|---|---|---| +| **[0.7.6](/fluid/releases/0.7.6)** | **Preview** | Declarative Packaging Modes | ✅ Additive over 0.7.5 | Adds top-level `packaging` and `consumers`, cross-mesh pins on `consumes[]` (`upstreamWorkspace` + `upstreamDigest`), `aggParams` on semantic measures, and on bindings `encryption.kms`, `principals` and Lake Formation `bucketPolicy`; an expose's `lifecycle` gains `expire`, and masking rules gain typed `params`. | +| **[0.7.5](/fluid/releases/0.7.5)** | **Stable (latest)** | Streaming Kafka → Iceberg Sink & Confluent Tableflow | ✅ Additive over 0.7.4 | Adds an opt-in Iceberg streaming sink on the `kafka-connect` acquisition engine, the `confluent` (Tableflow) and `pgvector` binding platforms, the `pgvector_table` format and `binding.vectorConfig`, and location fields for Tableflow, Iceberg catalogs, Redshift Serverless and Kinesis. | +| **[0.7.4](/fluid/releases/0.7.4)** | Stable | Runtime agentPolicy Enforcement at the MCP Gateway | ✅ Additive over 0.7.3 | Adds the optional `exposes[].mcp` block, the `postgres` binding platform, and the `postgres_table` / `athena_table` / `glue_table` formats. | +| **[0.7.3](/fluid/releases/0.7.3)** | Stable | Source-Aligned Acquisition | ✅ Additive over 0.7.2 | The `acquisition` build pattern and six ingestion engines, top-level `retention`, `governance` (AWS Lake Formation) and `extensions`, `binding.governance`, `metadata.productType`, and `exposes[].contract.schemaPolicy`. | +| **[0.7.2](/fluid/releases/0.7.2)** | Stable | Semantic Truth Engine | ⚠️ One narrowing: `notification` became a closed object | A per-expose `semantics` block, `binding.icebergConfig`, the top-level `orchestration` key, and adapter-qualified dbt engines (`build.engine: dbt-`). | +| **[0.7.1](/fluid/releases/0.7.1)** | Stable | Agentic Governance + Provider-First Orchestration | ✅ Additive over 0.5.7 | `exposes[].policy.agentPolicy`, top-level `sovereignty` and `accessPolicy`, and provider-action tasks under `builds[].execution.orchestration`. | --- ## Reading the compatibility column -The 0.7.1 → 0.7.5 line is strictly **additive**: bump `fluidVersion` and every prior contract still validates. +"Additive" means what [`scripts/check-compat.py`](https://github.com/open-data-protocol/fluid/blob/main/scripts/check-compat.py) checks on every pull request: no document that validates under one version stops validating under the next. Every transition in the table is additive except **0.7.1 → 0.7.2**, where `$defs/notification` gained `additionalProperties: false`. That break is recorded in [`scripts/compat-waivers.txt`](https://github.com/open-data-protocol/fluid/blob/main/scripts/compat-waivers.txt) and explained in the [0.7.2 note](/fluid/releases/0.7.2#backward-compatibility-with-0-7-1-one-narrowing-change). -**0.7.5 continues that streak.** It only adds optional fields, enum values, and a wider `engine` shape, so every valid 0.7.4 contract validates as 0.7.5 unchanged. See [What's New in 0.7.5](/fluid/releases/0.7.5#backward-compatibility-fully-additive) for the details. +Additive does **not** mean that an older `fluidVersion` may use newer fields. A document is validated against the schema of the version it declares, so to use a field you declare the version that added it. See [Choosing `fluidVersion`](/fluid/schema/versions#choosing-fluidversion). -For the field-by-field machine diff of every version, see the [**Schema Changelog**](/fluid/schema/changelog). +For the field-by-field history, see the [**Schema Changelog**](/fluid/schema/changelog). diff --git a/docs/schema/README.md b/docs/schema/README.md index e7f469d..3e885f7 100644 --- a/docs/schema/README.md +++ b/docs/schema/README.md @@ -1,19 +1,20 @@ # Schema Reference -Everything you need to read, write, and validate a FLUID contract. The latest schema version is **0.7.4**. +Everything you need to read, write, and validate a FLUID contract. **The latest stable schema is 0.7.5; 0.7.6 is a preview** (see [Versions](/fluid/schema/versions#stable-and-preview-versions)). | Page | Use it for | |---|---| | [**Anatomy**](/fluid/schema/anatomy) | A guided tour of every top-level block, in authoring order, with a "what / when / why" per block. Start here if you've never written a contract. | | [**Cheatsheet**](/fluid/schema/cheatsheet) | One row per field — meaning, required flag, and version added. The fast lookup table while you author. | | [**Minimal Valid Contract**](/fluid/schema/minimal-contract) | The smallest file that validates, plus a glance at every top-level block. | -| [**Full Specification**](/fluid/schema/specification) | The narrative protocol specification. | -| [**Versions**](/fluid/schema/versions) | Every published schema version, linked to its JSON Schema and generated HTML reference. | -| [**Changelog**](/fluid/schema/changelog) | Version-to-version diffs, including the breaking changes in 0.7.4. | +| [**Specification**](/fluid/schema/specification) | Validation semantics, document structure, and how multi-file composition relates to validity. | +| [**Versions**](/fluid/schema/versions) | Every published schema version and its status, how to choose `fluidVersion`, and known caveats. | +| [**Changelog**](/fluid/schema/changelog) | What changed between consecutive versions, including the pre-0.7.1 breaking changes and the one narrowing in 0.7.2. | ## Machine-readable artifacts -- **Raw JSON Schema (0.7.4):** [`/fluid/schema/fluid-schema-0.7.4.json`](/fluid/schema/fluid-schema-0.7.4.json) — point your validator here. -- **Generated HTML reference (0.7.4):** [`/fluid/specs/0.7.4/fluid-spec.html`](/fluid/specs/0.7.4/fluid-spec.html) — the authoritative field-by-field rendering of every type, enum, and validation rule. +- **JSON Schema (0.7.5, stable):** [`/fluid/schema/fluid-schema-0.7.5.json`](/fluid/schema/fluid-schema-0.7.5.json) — point your validator here. +- **Generated HTML reference (0.7.5):** [`/fluid/specs/0.7.5/fluid-spec.html`](/fluid/specs/0.7.5/fluid-spec.html) — every type, enum and validation rule, rendered from the schema. +- **Preview (0.7.6):** [`fluid-schema-0.7.6.json`](/fluid/schema/fluid-schema-0.7.6.json) · [`specs/0.7.6/fluid-spec.html`](/fluid/specs/0.7.6/fluid-spec.html). -For older versions, see [**Versions**](/fluid/schema/versions). +Validate a document against the schema of the `fluidVersion` it declares. For older versions, see [**Versions**](/fluid/schema/versions). diff --git a/docs/schema/anatomy.md b/docs/schema/anatomy.md index 4b52a41..f5b1787 100644 --- a/docs/schema/anatomy.md +++ b/docs/schema/anatomy.md @@ -1,10 +1,13 @@ # Schema Anatomy -A FLUID contract is one YAML file. This page walks **every top-level block** in v0.7.4, in the order you'd typically author them, with a one-paragraph "what / when / why" for each and a link into the generated reference doc for the field-by-field details. +A FLUID contract is one YAML file. This page walks **every top-level block** of the latest stable schema, **0.7.5**, in the order you'd typically author them, with a one-paragraph "what / when / why" for each and a link into the generated reference doc for the field-by-field details. > 📋 If you just want a one-line-per-field table, see the [**Cheatsheet**](/fluid/schema/cheatsheet). > 🟢 If you want a copy-pasteable minimal example, see the [**Minimal Valid Contract**](/fluid/schema/minimal-contract). -> 📚 Full field-by-field reference: [`specs/0.7.4/fluid-spec.html`](/fluid/specs/0.7.4/fluid-spec.html). +> 📚 Full field-by-field reference: [`specs/0.7.5/fluid-spec.html`](/fluid/specs/0.7.5/fluid-spec.html). +> 🧪 Blocks marked **0.7.6 preview** need `fluidVersion: "0.7.6"` and can still change — see [0.7.6 (preview)](/fluid/releases/0.7.6). + +The YAML blocks on this page fit together: merged in order into the [minimal contract](/fluid/schema/minimal-contract), they form one contract that validates against the 0.7.5 schema. --- @@ -13,13 +16,14 @@ A FLUID contract is one YAML file. This page walks **every top-level block** in ``` ┌────────────────────────── IDENTITY ──────────────────────────┐ │ fluidVersion kind id name description domain │ -│ metadata { owner, layer, businessContext } │ +│ metadata { owner, layer, productType, businessContext } │ │ tags labels │ └──────────────────────────────────────────────────────────────┘ ┌────── INTERFACE (what you publish) ───────┐ ┌── INTAKE ────┐ │ exposes[] │ │ consumes[] │ -│ ├─ contract (schema, dq, policy) │ │ │ +│ ├─ contract (schema, dq, schemaPolicy)│ │ │ +│ ├─ policy (authz, privacy, agent…) │ │ │ │ ├─ semantics (entities, measures, …) │ │ │ │ ├─ mcp (MCP gateway opt-in) ⭐0.7.4│ │ │ │ └─ binding (platform, location, …) │ │ │ @@ -28,14 +32,14 @@ A FLUID contract is one YAML file. This page walks **every top-level block** in ┌────── IMPLEMENTATION (how it's built) ───────────────────────┐ │ build { pattern, engine, properties, capabilities, … } │ │ └─ acquisition (source-aligned ingestion) ⭐ 0.7.3 │ -│ orchestration { engine, tasks } ⭐ 0.7.0+ │ +│ orchestration { engine, airflow, … } top-level ⭐ 0.7.2 │ └──────────────────────────────────────────────────────────────┘ ┌────── GOVERNANCE (rules around the product) ─────────────────┐ │ sovereignty (jurisdiction) top-level ⭐ 0.7.1 │ │ accessPolicy (IAM grants) top-level ⭐ 0.7.1 │ │ retention (data/log TTLs) top-level ⭐ 0.7.3 │ -│ governance (AWS Lake Formation) top-level ⭐ 0.4.0 │ +│ governance (AWS Lake Formation) top-level ⭐ 0.7.3 │ │ exposes[].policy.agentPolicy (AI/LLM) per-expose ⭐ 0.7.1 │ │ lineage schemaEvolution observability │ └──────────────────────────────────────────────────────────────┘ @@ -44,14 +48,18 @@ A FLUID contract is one YAML file. This page walks **every top-level block** in │ lifecycle { state } environments { dev, staging, prod } │ │ machineLearning { ... } docs { ... } │ └──────────────────────────────────────────────────────────────┘ + + 0.7.6 preview adds: packaging, consumers[], consumes[].upstream*, + binding.{encryption, principals, packaging}, lifecycle.expire ``` --- -## 1. Identity — *who is this product?* +## 1. Identity: *who is this product?* + ```yaml -fluidVersion: "0.7.4" # required +fluidVersion: "0.7.5" # required kind: DataProduct # required — DataProduct | MLPipeline id: finance.gold.customer_360 # required — globally unique, dot-separated name: "Customer 360" # required — display name @@ -67,30 +75,33 @@ labels: { team: customer-analytics, criticality: high, cost-center: engineering --- -## 2. `metadata` — *who owns it, where does it sit in the medallion?* +## 2. `metadata`: *who owns it, where does it sit in the medallion?* + ```yaml metadata: # required (only metadata.owner is required inside) - owner: # required - team: customer-analytics # required + owner: # required — no member of owner is required by the schema + team: customer-analytics # name a team: tools route alerts and ownership by it email: customer-analytics@company.com slack: "#customer-analytics-eng" layer: Gold # free-form string; convention: Bronze | Silver | Gold + productType: CDP # SDP | ADP | CDP (source-aligned, aggregated, consumption-aligned) businessContext: domain: "Customer Experience" subdomain: "Customer Intelligence" ``` **What:** ownership and business context that doesn't fit in `id`. -**When:** always — `metadata` is required, and its only required sub-field is `owner.team`. +**When:** always — `metadata` is required, and inside it only `owner` is. The schema requires no member of `owner`, so `owner: {}` validates; name at least a `team`. **Why:** every alert, catalog page, and ownership audit traces back here. Without an `owner.email`, a broken product has no human to wake up. > 💡 `metadata.layer` is a **free-form string** (not an enum). `Bronze`/`Silver`/`Gold` is the conventional medallion vocabulary, but the schema doesn't enforce it. --- -## 3. `consumes` — *what upstream products do I need?* +## 3. `consumes`: *what upstream products do I need?* + ```yaml consumes: - productId: finance.bronze.raw_payments @@ -102,14 +113,17 @@ consumes: **What:** explicit, version-constrained dependencies on other FLUID products. **When:** any product that isn't strictly source-aligned (i.e. not `build.pattern: acquisition`). -**Why:** the orchestrator builds the DAG from `consumes`. No `consumes` → no automatic upstream waiting, no automatic lineage. +**Why:** the orchestrator builds the DAG from `consumes`. No `consumes` → no automatic upstream waiting, no automatic lineage. Each entry can also state `qosExpectations` (`freshnessMax`, `maxStaleness`, `minCompleteness`) and `requiredPolicies`. + +> 🧪 **0.7.6 preview:** `upstreamWorkspace` and `upstreamDigest` pin an upstream that lives in another mesh; setting the first requires the second. See [0.7.6](/fluid/releases/0.7.6#cross-mesh-pins-on-consumes). --- -## 4. `exposes` — *what does this product publish?* +## 4. `exposes`: *what does this product publish?* This is the heart of the contract. A product can expose multiple ports (e.g. a Snowflake table **and** a Kafka stream **and** an API). + ```yaml exposes: - exposeId: customer_profiles # required @@ -130,12 +144,21 @@ exposes: type: valid_values # type ∈ freshness | completeness | uniqueness | valid_values | accuracy | schema | anomaly_detection | drift_detection selector: "ltv >= 0" severity: error # severity ∈ info | warn | error | critical - policy: - privacy: - masking: - - { column: email, strategy: tokenize } - authn: iam - authz: { readers: [analytics-team, ml-agents] } + schemaPolicy: evolve_safe # strict | discover_and_freeze | evolve_safe | evolve_all + + # ── POLICY — a sibling of contract, not inside it ───────── + policy: + authn: iam + authz: { readers: [analytics-team, ml-agents] } + privacy: + masking: + - { column: email, strategy: tokenize } + + # ── QoS ─────────────────────────────────────────────────── + qos: + availability: "99.9%" # a percentage string + freshnessSLO: PT1H # ISO-8601 durations + latencyP95: PT2S # ── 4b. SEMANTICS (⭐ 0.7.2) ───────────────────────────── semantics: @@ -155,42 +178,42 @@ exposes: # ── 4d. BINDING ─────────────────────────────────────────── binding: - platform: gcp # gcp | aws | azure | snowflake | databricks | kafka | local | kubernetes | postgres ⭐0.7.4 | other - format: bigquery_table # bigquery_table | snowflake_table | iceberg | delta_table | parquet | http_api | kafka_topic | postgres_table ⭐0.7.4 | athena_table ⭐0.7.4 | glue_table ⭐0.7.4 | … - location: { project: company-data, dataset: gold_customer, table: profiles_v1 } - icebergConfig: # ⭐ 0.7.2 — only when format = iceberg - writeVersion: 2 - fileFormat: parquet - partitionSpec: [ { sourceColumn: created_at, transform: month } ] + platform: gcp # gcp | aws | azure | snowflake | databricks | kafka | confluent ⭐0.7.5 | local | kubernetes | postgres ⭐0.7.4 | pgvector ⭐0.7.5 | other + format: bigquery_table # bigquery_table | snowflake_table | iceberg | delta_table | parquet | http_api | kafka_topic | postgres_table ⭐0.7.4 | athena_table ⭐0.7.4 | glue_table ⭐0.7.4 | pgvector_table ⭐0.7.5 | … + location: { project: company-data, dataset: gold_customer, table: profiles_v1, region: europe-west1 } ``` -### `contract` — *the data shape and the rules around it* +### `contract`: *the data shape and the rules around it* -- **`schema`**: columns, types, nullability, sensitivity (`cleartext` | `tokenized` | `hashed` | `encrypted` | …), tags. +- **`schema`**: columns, types, `required`, sensitivity (`cleartext` | `tokenized` | `pseudonymized` | `encrypted` | `pii` | …), tags. - **`dq.rules`**: declarative data-quality assertions (`valid_values`, `freshness`, `completeness`, `uniqueness`, `accuracy`, `anomaly_detection`, `drift_detection`, `schema`). -- **`policy`**: privacy treatments, RBAC readers/writers, column-level redaction, and the per-expose `agentPolicy` (see §7). +- **`schemaPolicy`** (0.7.3+): how the output schema may change. + +`contract` is closed: `policy` is **not** a member of it. Privacy treatments, readers/writers, column restrictions and the per-expose `agentPolicy` go in **`exposes[].policy`**, next to `contract` (see §7). `qos` states service levels for the port: `availability`, `freshnessSLO`, `dataLossSLO`, `latencyP95`, `completenessTarget`, `errorBudget`. -### `semantics` ⭐ 0.7.2 — *the business meaning layer* +### `semantics` (0.7.2): *the business meaning layer* `entities`, `measures`, `dimensions`, `metrics` — the same primitives as dbt MetricFlow and Snowflake Semantic Views. An LLM or BI tool reading this block can answer *"what is our MRR?"* without inventing SQL. -### `mcp` ⭐ 0.7.4 — *opt this expose into the MCP output-port gateway* +### `mcp` (0.7.4): *opt this expose into the MCP output-port gateway* **What:** the `exposes[].mcp` block declares this expose as agent-consumable over MCP. Its two sub-blocks are `sampling.maxRows` (integer ≥1, default 100 — a hard cap on the `sample` tool so an over-curious agent can't pull petabyte-scale rows) and `classification.dataClass` (`public` | `internal` | `confidential` | `restricted` — an advisory hint surfaced on the `describe` tool so consumer agents can declare downstream handling). **When:** any expose you want served through `fluid mcp output-port serve` and reachable by AI agents over MCP. -**Why:** the *presence* of a (non-empty) `mcp` block is what flips on runtime enforcement. Once present, the expose's `policy.agentPolicy` (`allowedModels` / `deniedModels`, `allowedUseCases` / `deniedUseCases`, token caps) is enforced by the Fluid MCP gateway **on every read** — closing the gap where those fields were previously declarative metadata only. The `agentPolicy` shape itself is unchanged from 0.7.3; 0.7.4 makes it a runtime gate. +**Why:** the *presence* of a (non-empty) `mcp` block is what flips on runtime enforcement. Once present, the expose's `policy.agentPolicy` (`allowedModels` / `deniedModels`, `allowedUseCases` / `deniedUseCases`, token caps) is enforced **on every read** by an MCP gateway that implements 0.7.4 (the reference implementation's does) — where before those fields were declarative metadata only. The `agentPolicy` shape itself is unchanged from 0.7.3; 0.7.4 makes it a runtime gate. > 🔗 See the [**MCP how-to**](/fluid/how-to/mcp) for an end-to-end walkthrough of serving an expose over MCP with governed agent access. -### `binding` — *where the data physically lives* +### `binding`: *where the data physically lives* + +Platform + format + location. Plus `icebergConfig` (0.7.2+) when format is `iceberg` (write version, file format, `partitionSpec`, `sortOrder`), so the contract owns table-format details too. **0.7.4** adds the `postgres` platform and the `postgres_table` / `athena_table` / `glue_table` formats. **0.7.5** adds the `confluent` (Tableflow) and `pgvector` platforms, `pgvector_table`, `vectorConfig`, and location fields for Iceberg catalogs, Redshift Serverless and Kinesis — see [0.7.5](/fluid/releases/0.7.5). Optional `binding.governance` carries per-resource AWS Lake Formation settings. -Platform + format + location. Plus `icebergConfig` (0.7.2+) when format is `iceberg`, so the contract owns table-format details too. **0.7.4** adds the `postgres` platform and the `postgres_table` / `athena_table` / `glue_table` formats (additive — `snowflake_view` and the `redshift_*` formats remain valid). Optional `binding.governance` carries per-resource AWS Lake Formation grants. +> 🧪 **0.7.6 preview** adds `binding.encryption.kms`, `binding.principals` (logical principal → platform identity), `binding.packaging`, and Lake Formation `bucketPolicy`; a bucket policy it emits is **authoritative** and replaces every other statement on the bucket. See [0.7.6](/fluid/releases/0.7.6#binding-fields-encryption-principals-lake-formation-bucket-policy). --- -## 5. `build` — *how the product is produced* +## 5. `build`: *how the product is produced* Four patterns, validated conditionally: @@ -201,10 +224,11 @@ Four patterns, validated conditionally: | `multi-stage` | Multi-step pipeline (bronze → silver → gold) in one product. | | `acquisition` ⭐ 0.7.3 | Source-aligned ingestion from external systems (Postgres, Salesforce, Kafka, files, …). | + ```yaml build: pattern: acquisition # ⭐ 0.7.3 — picks acquisitionPattern shape for build.properties below - engine: airbyte # dbt | sql | python | spark | glue | custom | duckdb | airbyte | meltano | dlt | kafka-connect | debezium + engine: airbyte # dbt | dbt- | sql | python | spark | glue | custom | duckdb | airbyte | meltano | dlt | kafka-connect | debezium capabilities: [incremental_dedup, schema_evolution, dlp_scan] properties: # ← schema validates this as acquisitionPattern when pattern=acquisition source: @@ -212,11 +236,11 @@ build: mode: incremental_dedup # full_refresh | incremental_append | incremental_dedup | incremental_merge | cdc | streaming cursor_field: updated_at connection: - secretRef: "vault://pg-prod-readonly" # URI form required: vault:// aws:// gcp:// azure:// env:// + secretRef: "vault://pg-prod-readonly" # must be a URI (://…), e.g. vault://, env:// streams: [public.customers, public.orders] sink: - format: iceberg - partitionBy: ["day(ingested_at)"] # array of strings (function-form). NOT the object-form used by binding.icebergConfig.partitionSpec + format: bigquery_table # iceberg | delta | parquet | csv | json | snowflake_table | bigquery_table | redshift_table | duckdb_table + partitionBy: ["day(ingested_at)"] # array of strings (function-form) — not the object form of binding.icebergConfig.partitionSpec delivery: guarantee: at_least_once # at_most_once | at_least_once | exactly_once idempotencyKey: "${stream}|${batch_id}" @@ -241,39 +265,52 @@ build: > 🧭 **Where `acquisitionPattern` actually lives:** in the schema, `build.properties` is validated against a different shape depending on `build.pattern` (`hybrid-reference` → `hybridReferencePattern`, `embedded-logic` → `embeddedLogicPattern`, `multi-stage` → `multiStagePattern`, `acquisition` → `acquisitionPattern`). The pattern name is the discriminator; `properties` is always the payload key. -**When:** any product that is built rather than virtual. +**When:** any product that is built rather than virtual. Use `build` for one build, or `builds[]` (each with an `id`) for several; both have the same shape. --- -## 6. `orchestration` — *who actually runs the build?* +## 6. `orchestration`: *who actually runs the build?* + ```yaml orchestration: - engine: airflow # airflow | dagster | prefect | kubeflow | custom | none + engine: airflow # required — airflow | dagster | prefect | kubeflow | custom | none + mode: generated # generated | manual | hybrid generateOnChange: true - tasks: - - taskId: ensure_warehouse - type: provider_action # provider_action | fluid_validate | fluid_plan | fluid_apply | fluid_execute | fluid_verify | fluid_dq_check | bash | python | sensor | … - provider: snowflake.warehouse - action: ensure - parameters: { name: ANALYTICS_WH, size: medium } - - taskId: load_customers - type: provider_action - provider: snowflake.table - action: ensure - buildStepRef: customers_ingest - dependsOn: [ensure_warehouse] + airflow: + dagId: customer_360 # required when airflow is present + dagConfig: { schedule: "0 2 * * *", catchup: false } + tasks: + - taskId: ensure_warehouse + type: provider_action # provider_action | fluid_validate | fluid_plan | fluid_apply | fluid_execute | fluid_verify | fluid_dq_check | bash | python | sensor | … + operator: ProviderActionOperator + provider: snowflake # aws | gcp | azure | snowflake | databricks | kafka | kubernetes | local | custom + action: warehouse.ensure # . + params: { name: ANALYTICS_WH, size: medium } + - taskId: load_customers + type: provider_action + operator: ProviderActionOperator + provider: snowflake + action: table.ensure + params: { name: CUSTOMERS } + buildStepRef: customers_ingest + dependencies: [ensure_warehouse] ``` -**What:** how the build is scheduled and runs. Provider actions (0.7.0+) let you target cloud primitives directly without wrapper operators. -**Why:** the orchestrator config used to be the part most likely to drift from the contract. Embedding it here keeps "what runs" and "what gets produced" in one file. +**What:** how the build is scheduled and run. The top-level `orchestration` key exists from 0.7.2; the same block can sit under `build.execution.orchestration` (from 0.7.1). +**Why:** the orchestrator config used to be the part most likely to drift from the contract. Embedding it keeps "what runs" and "what gets produced" in one file. + +> ⚠️ **Only `orchestration.airflow.tasks[]` is checked.** `orchestration` is an open object, so a top-level `orchestration.tasks` list — which some FLUID tooling reads — is accepted with **no validation at all**: a typo in it is not reported. The key names in such a list are implementation-defined. In the reference implementation, the AWS and GCP code generators read `params` and `dependsOn` from `orchestration.tasks[]`, while its Snowflake Airflow generator reads `parameters`. +> +> Name each task's **`operator`**, as above. In the 0.7.1 to 0.7.6 schemas the conditional rules for `fluid_execute`, `bash` and `python` tasks also match a task that has no `operator`, so such a task must satisfy all three at once and is in practice rejected. --- -## 7. Governance — `sovereignty`, `accessPolicy`, `governance`, and `exposes[].policy.agentPolicy` +## 7. Governance: `sovereignty`, `accessPolicy`, `governance`, and `exposes[].policy.agentPolicy` -The agentic-era governance layer is mostly 0.7.1 (`sovereignty`, `accessPolicy`, `agentPolicy`), plus the older top-level `governance` block (AWS Lake Formation, since 0.4.0). **Three blocks are top-level (`sovereignty`, `accessPolicy`, `governance`); `agentPolicy` is per-expose** — it lives inside `exposes[].policy.agentPolicy`, not at the root, even though some prior release notes show it at the top level. +The agentic-era governance layer is mostly 0.7.1 (`sovereignty`, `accessPolicy`, `agentPolicy`), plus the top-level `governance` block (AWS Lake Formation, since 0.7.3; an unrelated `governance` block existed from 0.0.1 to 0.4.0 and was removed in 0.5.7). **Three blocks are top-level (`sovereignty`, `accessPolicy`, `governance`); `agentPolicy` is per-expose** — it lives inside `exposes[].policy.agentPolicy`, not at the root (earlier versions of the 0.7.1 release note showed it at the top level). + ```yaml sovereignty: # ⭐ 0.7.1 — where data may live (top-level) jurisdiction: EU @@ -294,9 +331,11 @@ accessPolicy: # ⭐ 0.7.1 — IAM grants genera permissions: [write, insert, update] conditions: { ipRanges: ["10.0.0.0/8"] } -governance: # ⭐ 0.4.0 — account/project-wide AWS Lake Formation (top-level) +governance: # ⭐ 0.7.3 — account-wide AWS Lake Formation (top-level) lakeFormation: # admins + LF-tag definitions (per-resource grants live under binding.governance) - admins: ["arn:aws:iam::123456789012:role/LakeAdmins"] + admins: # AUTHORITATIVE: replaces the account's admin list — see the warning below + - "arn:aws:iam::123456789012:role/LakeAdmins" + - "arn:aws:iam::123456789012:role/fluid-deployer" # the role that runs the apply tagDefinitions: # tag key -> allowed values (emitted as aws_lakeformation_lf_tag) classification: [public, internal, confidential, restricted] @@ -326,14 +365,21 @@ exposes: purposeLimitation: "Customer support chatbot only — not for marketing." ``` -**Why:** the contract becomes the source of truth for compliance reviewers. Drift between "what the contract says" and "what IAM bindings exist in prod" disappears. +**Why:** the contract becomes the source of truth for compliance reviewers, and drift between "what the contract says" and "what IAM bindings exist" becomes checkable. + +::: danger `governance.lakeFormation.admins` replaces the admin list +`admins` is **authoritative, not additive**. Applying it sets the account's Lake Formation admins to exactly this list, so any admin you do not list is **removed** — including the role or user that runs the apply, after which Lake Formation grants and registrations that need an admin can fail. List every principal that must stay an admin, the deploying identity included. (The 0.7.3–0.7.5 schemas were first published with a description saying the opposite; see the [Changelog](/fluid/schema/changelog#corrections-to-published-schemas).) In the reference implementation, applying this block also clears the account's other data-lake settings that it does not set: the create-database and create-table default permissions, trusted resource owners and parameters. Destroying the emitted `aws_lakeformation_data_lake_settings` resource empties the admin list and resets `CROSS_ACCOUNT_VERSION` to 1. See the field's description in the [schema](/fluid/specs/0.7.5/fluid-spec.html). +::: + +> 🧪 **0.7.6 preview** adds `binding.principals`, which maps the logical principals named in `accessPolicy` and `columnRestrictions` to each platform's identities, writes down what `columnRestrictions` means (deny beats allow; a restriction never grants access), and adds typed masking `params`. See [0.7.6](/fluid/releases/0.7.6). -> ⭐ **0.7.4 — runtime enforcement.** When an expose also carries an [`mcp` block](#mcp-0-7-4-opt-this-expose-into-the-mcp-output-port-gateway), the `agentPolicy` above stops being passive metadata: the Fluid MCP gateway enforces `allowedModels` / `deniedModels` and `allowedUseCases` / `deniedUseCases` on every read. The `agentPolicy` shape is unchanged — 0.7.4 only changes *when* it bites. +> ⭐ **0.7.4 — runtime enforcement.** When an expose also carries an [`mcp` block](#mcp-0-7-4-opt-this-expose-into-the-mcp-output-port-gateway), the `agentPolicy` above stops being passive metadata: an MCP gateway that implements 0.7.4, such as the reference implementation's, enforces `allowedModels` / `deniedModels` and `allowedUseCases` / `deniedUseCases` on every read. The `agentPolicy` shape is unchanged — 0.7.4 only changes *when* it bites. --- -## 8. `retention` ⭐ 0.7.3 — *TTLs for everything around the product* +## 8. `retention` (0.7.3): *TTLs for everything around the product* + ```yaml retention: # ISO-8601 durations; defaults shown runState: P30D @@ -342,8 +388,8 @@ retention: # ISO-8601 durations; defaults sh dlq: P180D ``` -**What:** how long the platform retains operational artifacts. -**Why:** before 0.7.3 these knobs were scattered across Airflow, S3 lifecycle rules, Datadog, etc. The retention sweeper now reads this block. +**What:** how long the platform retains operational records — run state, run logs, lineage events, dead-letter records. It does not govern the product's data; that is `lifecycle.retention`. +**Why:** before 0.7.3 these knobs were scattered across tools. One block in the contract states them once. --- @@ -351,15 +397,15 @@ retention: # ISO-8601 durations; defaults sh These are the deeper-cut blocks — every one is opt-in: -- **`lineage`** — upstream/downstream graph, optional field-level mappings. -- **`schemaEvolution`** — compatibility strategy (`backward`, `forward`, `full`), version pinning. -- **`observability`** — SLIs and dashboard refs (default SLIs auto-generated). -- **`machineLearning`** — model spec for ML-product kinds (features, training, inference). -- **`environments`** — per-env (`dev` / `staging` / `prod`) override blocks. -- **`lifecycle`** — `state: preview | active | deprecated | retired` + deprecation date. -- **`docs`** — external doc URLs (runbook, design doc, dashboard). +- **`lineage`** — `granularity` (`table_level` | `field_level`), `upstream[]` (with optional `fieldMappings`) and `downstream[]`. +- **`schemaEvolution`** — `strategy` (`semantic_versioning` | `date_based` | `sequential`), `compatibility` (`backward_compatible` | `forward_compatible` | `full_compatible` | `breaking`) and `changePolicy`. +- **`observability`** (per expose) — `metrics`, `onBreach` actions, `defaultSLIs`, and (0.7.3+) `alert` channels. +- **`machineLearning`** — `enabled`, `framework` (`mlflow` | `kubeflow` | `sagemaker` | `vertex_ai` | `custom`), `models[]`. +- **`environments`** — per-environment overrides of `metadata` and `exposes`. +- **`lifecycle`** — `state: preview | active | deprecated | retired`, `retention` (an ISO-8601 duration) and `deprecationPolicy` (`noticePeriod`, `contact`, `replacement`). Note that `preview` here is a *product's* state; it has nothing to do with preview *schema versions*. 🧪 In the 0.7.6 preview, an expose's `lifecycle` also takes `expire: true`, which deletes data older than `retention`. +- **`docs`** — `homepage`, `runbook`, `dictionary`, `changeLog`. -For the field-by-field detail on each, see the generated reference: [`specs/0.7.4/fluid-spec.html`](/fluid/specs/0.7.4/fluid-spec.html). +For the field-by-field detail on each, see the generated reference: [`specs/0.7.5/fluid-spec.html`](/fluid/specs/0.7.5/fluid-spec.html). --- @@ -371,4 +417,4 @@ If you've never written a FLUID contract: 2. Copy the [Minimal Valid Contract](/fluid/schema/minimal-contract) and run it. 3. Walk through the [**Examples**](/fluid/examples/) — each example adds one block from this anatomy until you reach a production-grade source-aligned acquisition product. 4. Use the [**Cheatsheet**](/fluid/schema/cheatsheet) as a lookup table when you're authoring. -5. Drop into the generated [`fluid-spec.html`](/fluid/specs/0.7.4/fluid-spec.html) for exact types and enums. +5. Drop into the generated [`fluid-spec.html`](/fluid/specs/0.7.5/fluid-spec.html) for exact types and enums. diff --git a/docs/schema/changelog.md b/docs/schema/changelog.md index a862623..75ee722 100644 --- a/docs/schema/changelog.md +++ b/docs/schema/changelog.md @@ -1,70 +1,149 @@ # Changelog -Human-readable diffs between consecutive versions of the FLUID schema. Each entry links to the full auto-generated change list on GitHub. The latest version is **0.7.4**. +Human-readable summaries of what changed between consecutive versions of the FLUID schema. Each entry links to the full auto-generated change list on GitHub. **0.7.5 is the latest stable version; 0.7.6 is a preview** (see [Versions](/fluid/schema/versions#stable-and-preview-versions)). -## 0.7.3 → 0.7.4 — Runtime agentPolicy Enforcement at the MCP Gateway +"No narrowing change" below means that `python3 scripts/check-compat.py --from --to ` reports none: every document valid under the earlier version stays valid under the later one (once it declares the later `fluidVersion`). Each summary lists the property paths present in the later schema and absent in the earlier one, plus new enum values. -[Full diff →](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.7.3-to-0.7.4.md) +## 0.7.5 to 0.7.6 (preview): Declarative Packaging Modes -This release makes an expose *agent-consumable over MCP* and turns `policy.agentPolicy` into a gate the Fluid MCP output-port gateway enforces on every read. The `policy.agentPolicy` **shape is unchanged** from 0.7.3 — the enforcement is behavioral, not a new field. **Fully backward-compatible: this release is purely additive, so every valid 0.7.3 contract validates as 0.7.4 unchanged.** +[Full diff →](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.7.5-to-0.7.6.md) · [Release note →](/fluid/releases/0.7.6) + +No narrowing change. **Preview:** the 0.7.6 schema can still change before it is promoted. + +**Added** + +- Top-level `packaging` (`mode`, `pool`, `poolManifest`, `containers`) and its per-binding override `exposes[].binding.packaging`. +- Top-level `consumers[]` (`name` and `type` required; `label`, `owner`, `url`, `maturity`, `description`, `exposeIds`). +- `consumes[].upstreamWorkspace` and `consumes[].upstreamDigest` (`sha256:` + 64 lowercase hex), with `dependentRequired: {upstreamWorkspace: [upstreamDigest]}` — the first JSON Schema keyword in a FLUID schema that Draft 7 does not have. +- `exposes[].semantics.measures[].aggParams` (`percentile`, `useDiscretePercentile`). +- `exposes[].binding.encryption.kms`, with platform rules: no AWS key reference on a `gcp` binding, no Cloud KMS name on an `aws` binding. +- `exposes[].binding.principals`, with identity patterns for `gcp` (IAM members) and `aws` (IAM ARNs). +- `exposes[].binding.governance.lakeFormation.bucketPolicy` (`cross-account` | `none` | `all-grantees`; `cross-account` is the default). Its description says that an emitted `aws_s3_bucket_policy` is authoritative and replaces every other statement on the bucket; see the [0.7.6 notes](/fluid/releases/0.7.6#binding-fields-encryption-principals-lake-formation-bucket-policy). +- `exposes[].lifecycle` now refers to a new `exposeLifecycle` definition: the root `lifecycle` fields plus `expire` (default `false`). +- Typed `exposes[].policy.privacy.masking[].params`: `keepFirst`, `keepLast`, `saltEnv`, `keyEnv`. +- `fluidVersion` accepts `"0.7.6"`. + +**Described:** `exposes[].policy.authz.columnRestrictions` gains a description of its semantics (deny beats allow; a restriction never grants access; principals are logical). + +## 0.7.4 to 0.7.5: Streaming Kafka to Iceberg Sink & Confluent Tableflow + +[Full diff →](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.7.4-to-0.7.5.md) · [Release note →](/fluid/releases/0.7.5) + +No narrowing change. **Added** -- **`exposes[].mcp`** (new `exposeMcp` block) opts an expose into the MCP gateway. `mcp.sampling.maxRows` (integer ≥ 1, default 100) hard-caps the `sample` tool; `mcp.classification.dataClass` (`public` / `internal` / `confidential` / `restricted`) is an advisory hint surfaced on `describe`. -- **`binding.platform`** gains `postgres`. -- **`binding.format`** gains `postgres_table`, `athena_table`, `glue_table` (added alongside the existing formats, including `snowflake_view` and the `redshift_*` values). -- **`fluidVersion`** moves from `const: "0.7.3"` to `enum: ["0.7.3", "0.7.4"]`, so existing `0.7.3` values still validate. +- On the `kafka-connect` acquisition engine (`build(s).properties.kafka-connect`): `iceberg_sink_enabled`, `sink_topics`, `streamingSink`, `iceberg_catalog_overrides`. +- `binding.platform` values `confluent` and `pgvector`; `binding.format` value `pgvector_table`; `binding.vectorConfig`. +- `binding.location` fields: `environment_id`, `kafka_cluster_id`, `confluent_role_arn` (Confluent Tableflow); `catalog`, `warehouse`, `uri`, `partitionBy` (Iceberg catalogs); `namespace`, `workgroup`, `iam_role_arn`, `external_schema`, `glue_database` (Redshift Serverless and external schemas); `stream` (Kinesis). +- `fluidVersion` accepts `"0.7.5"`. -**Unchanged & still valid.** Nothing was removed — the top-level `governance` block and `binding.governance` (AWS Lake Formation), the `snowflake_view` / `redshift_table` / `redshift_serverless` / `redshift_external_schema` formats, the `athena` / `glue` / `redshift` runtimes, `datamesh_manager` catalog registration, and the `opentofu` deployment target all remain valid in 0.7.4. +**Re-published.** The 0.7.5 schema first published here was a snapshot taken while 0.7.5 was still a preview in the reference implementation. It lacked `pgvector`, `pgvector_table`, `vectorConfig` and the Iceberg-catalog, Redshift Serverless and Kinesis location fields, which the stable 0.7.5 has. The published file has been re-vendored from forge-cli v0.18.1, which carries the stable 0.7.5; every document the snapshot accepted is still accepted. -**Upgrading.** Set `fluidVersion: "0.7.4"` (or keep `"0.7.3"` — it still validates). There are no removed fields to migrate off. +## 0.7.3 to 0.7.4: Runtime agentPolicy Enforcement at the MCP Gateway -## 0.7.2 → 0.7.3 — Source-Aligned Acquisition +[Full diff →](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.7.3-to-0.7.4.md) · [Release note →](/fluid/releases/0.7.4) -[Full diff →](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.7.2-to-0.7.3.md) +No narrowing change. The gate flags `fluidVersion`, which became an enum; the waiver in `scripts/compat-waivers.txt` records why that is intended. -Adds the `acquisition` build pattern, six ingestion engines, `build.capabilities`, delivery guarantees + DLQ, source-side `schemaEvolution`, supply-chain image signatures, and the top-level `retention` block. Additive over 0.7.2. +**Added** + +- `exposes[].mcp` (`sampling.maxRows`, `classification.dataClass`). +- `binding.platform` value `postgres`; `binding.format` values `postgres_table`, `athena_table`, `glue_table`. +- Acquisition `catalog.register` values `glue`, `snowflake_horizon`, `unity`. +- `fluidVersion` becomes `enum: ["0.7.3", "0.7.4"]`. + +`policy.agentPolicy` is unchanged in shape; 0.7.4's headline change is how the reference implementation's MCP gateway enforces it. + +## 0.7.2 to 0.7.3: Source-Aligned Acquisition -## 0.7.1 → 0.7.2 — Semantic Truth Engine +[Full diff →](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.7.2-to-0.7.3.md) · [Release note →](/fluid/releases/0.7.3) -[Full diff →](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.7.1-to-0.7.2.md) +No narrowing change. The gate flags the identifier pattern (upper-case letters became allowed); the waiver records the exhaustive check that shows the new pattern only widens. -Adds the per-expose `semantics` block (entities, measures, dimensions, metrics) and `binding.icebergConfig`. Additive over 0.7.1. +**Added** + +- `build.pattern` value `acquisition`, the acquisition properties (`source`, `sink`, `delivery`, `schemaEvolution`, `preLand`, `quality`, `cost`, `catalog`, `concurrency`, `lineage`, and one key per engine: `duckdb`, `airbyte`, `meltano`, `dlt`, `kafka-connect`, `debezium`), the matching `build.engine` values, and `build.capabilities`. +- Top-level `retention`, `governance` (AWS Lake Formation `admins` and `tagDefinitions`) and `extensions`. +- `exposes[].binding.governance` (per-resource Lake Formation settings). +- `metadata.productType`, `metadata.classification`, `metadata.experimental`. +- `exposes[].contract.schemaPolicy`, `exposes[].observability.alert`. +- `binding.format` values `snowflake_view`, `redshift_table`, `redshift_serverless`, `redshift_external_schema`; `runtime.platform` values `athena`, `glue`, `redshift`. -## 0.5.7 → 0.7.1 — Agentic Governance & Provider-First Orchestration +## 0.7.1 to 0.7.2: Semantic Truth Engine -[Full diff →](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.5.7-to-0.7.1.md) +[Full diff →](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.7.1-to-0.7.2.md) · [Release note →](/fluid/releases/0.7.2) -Adds `exposes[].policy.agentPolicy`, top-level `sovereignty`, top-level `accessPolicy`, and provider-first `orchestration` tasks. +**One narrowing change:** `$defs/notification` gained `additionalProperties: false`, so an extra member on a `notifications[]` entry that 0.7.1 accepted is rejected by 0.7.2. It shipped before the gate existed and is recorded in `scripts/compat-waivers.txt`. + +**Added** -## 0.4.0 → 0.5.7 +- `exposes[].semantics` (entities, measures, dimensions, metrics). +- `exposes[].binding.icebergConfig`, `exposes[].binding.properties`, and `binding.format` value `iceberg`. +- Top-level `orchestration` (an open object; see the note under [Cheatsheet → orchestration](/fluid/schema/cheatsheet#orchestration-fields)). +- `exposes[].description`, `exposes[].crawler`, `exposes[].iceberg`, `exposes[].contract.quality`. +- `metadata.provenance`, `metadata.tags`. +- `hybrid-reference` build properties `models`, `select`, `target`. +- `build.engine` accepts adapter-qualified dbt engines matching `^dbt-[a-z0-9]+([_-][a-z0-9]+)*$` (for example `dbt-databricks`). +- `accessPolicy.grants[].permissions` value `create`. + +## 0.5.7 to 0.7.1: Agentic Governance & Provider-First Orchestration + +[Full diff →](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.5.7-to-0.7.1.md) · [Release note →](/fluid/releases/0.7.1) + +No narrowing change. 0.7.1 is the first version the compatibility promise covers; the gate runs from here on. + +**Added** + +- `exposes[].policy.agentPolicy`, top-level `sovereignty`, top-level `accessPolicy`. +- `build(s).execution.orchestration`, including Airflow DAG settings and `provider_action` tasks under `airflow.tasks[]`. +- Runtime fields (`platform`, `image`, `executor`, `serviceAccount`) and trigger fields (`cron`, `timezone`, `datasets`, `datasetsOperator`, `timetable`); trigger types `dataset`, `schedule_and_dataset`, `timetable`. + +--- + +## Before 0.7.1: history before the compatibility promise + +These versions predate the promise, and several transitions removed or narrowed fields. The gate is not run over them in CI; the result of running it by hand is given for each. + +### 0.4.0 to 0.5.7 (breaking) [Full diff →](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.4.0-to-0.5.7.md) -Adds `builds`, `machineLearning`, `environments`, `lifecycle`, `docs`, and `extensions`. +`check-compat` reports breaking changes. The contract was restructured around today's shape: `exposes[]` gains `exposeId`, `kind`, `contract` and `binding`; `consumes[]` gains `productId`, `exposeId` and `versionConstraint`; `build` gains `id`, `pattern`, `engine` and `properties`. Added at the top level: `builds`, `tags`, `labels`, `lineage`, `schemaEvolution`, `machineLearning`, `environments`, `lifecycle`, `docs`. **Removed** at the top level: `accessPolicy`, `governance`, `operations`, `security`, `slo` (`accessPolicy` returns in 0.7.1 and a different `governance` in 0.7.3). -## 0.3.0 → 0.4.0 +### 0.3.0 to 0.4.0 [Full diff →](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.3.0-to-0.4.0.md) -Adds `lineage`, `governance`, and `schemaEvolution`. +Apart from the `$id`, the only change is in `build.transformation`: its `properties` are now chosen by `pattern` through `if`/`then` rules instead of a `oneOf`. `check-compat` reports no narrowing change. -## 0.2.0 → 0.3.0 +### 0.2.0 to 0.3.0 (breaking) [Full diff →](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.2.0-to-0.3.0.md) -Adds the `build` block. +`check-compat` reports breaking changes. `build` is restructured into `build.transformation` (with build patterns) and `build.execution`, replacing `build.engine`, `runtime`, `trigger`, `retries` and `notifications` at the `build` level. `governance` gains `lineage`, `regulatory` and `stewardship` in place of `rules`. -## 0.1.1 → 0.2.0 +### 0.1.1 to 0.2.0 (breaking) [Full diff →](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.1.1-to-0.2.0.md) -Adds the `consumes` block. +`check-compat` reports breaking changes: patterns are added to ids, `domain`, column names and access principals, bounds to SLA and lifecycle numbers, and a pattern to cron triggers. `exposes[].mappings` is added. -## 0.1.0 → 0.1.1 +### 0.1.0 to 0.1.1 [Full diff →](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.1.0-to-0.1.1.md) -## 0.0.1 → 0.1.0 +No narrowing change. Note that 0.1.0 and 0.1.1 declare the same `$id` (see [Versions](/fluid/schema/versions#_0-4-0-and-earlier-ids-that-do-not-resolve)). + +### 0.0.1 to 0.1.0 (breaking) [Full diff →](https://github.com/open-data-protocol/fluid/blob/main/schema-diffs/diff-0.0.1-to-0.1.0.md) + +`check-compat` reports breaking changes. Top-level `conformance`, `dynamicPolicies` and `extensions` are removed (`extensions` returns in 0.7.3), and `slo` is added. `consumes` and `build` exist from 0.0.1 on. + +--- + +## Corrections to published schemas + +- **Lake Formation `admins` (0.7.3, 0.7.4, 0.7.5).** These schemas were first published with a description of `governance.lakeFormation.admins` saying that principals not listed are not removed. That is the opposite of what applying it does: the list is **authoritative**, and an admin not listed — including the identity running the apply — loses Lake Formation admin. The description was corrected in the reference implementation and the published files carry the corrected text. Only the description changed; the same documents validate. +- **0.7.5** was re-published with the stable content (above). diff --git a/docs/schema/cheatsheet.md b/docs/schema/cheatsheet.md index c4f0be0..b3047a5 100644 --- a/docs/schema/cheatsheet.md +++ b/docs/schema/cheatsheet.md @@ -1,276 +1,299 @@ # Schema Cheatsheet -One row per top-level field in the latest schema (**0.7.4**). Use this as a lookup table when you're reading or writing a `.fluid.yml` — find the field on the left, get a one-line meaning, find out if it's required, and jump to the deep-dive. +One row per field of the latest stable schema, **0.7.5**: a one-line meaning, whether it is required, and the version it first appeared in. Rows marked 🧪 are in the **0.7.6 preview** only; a contract needs `fluidVersion: "0.7.6"` to use them, and they can still change. > 🧭 For a guided tour with context per block, read the [**Schema Anatomy**](/fluid/schema/anatomy). -> 📚 For the exhaustive field-by-field reference, see [`specs/0.7.4/fluid-spec.html`](/fluid/specs/0.7.4/fluid-spec.html). +> 📚 For the exhaustive field-by-field reference, see [`specs/0.7.5/fluid-spec.html`](/fluid/specs/0.7.5/fluid-spec.html) (stable) or [`specs/0.7.6/fluid-spec.html`](/fluid/specs/0.7.6/fluid-spec.html) (preview). + +"Since" is the version from which the field has been in the schema without interruption; where a member of the same name existed earlier and was removed, the row says so. Contracts before 0.5.7 had a different overall shape (see the [Changelog](/fluid/schema/changelog#before-0-7-1-history-before-the-compatibility-promise)), so for those blocks the current *shape* is newer than the name. --- ## Top-level fields -| Field | Purpose | Required | Added in | Type | +The root object is closed: a top-level key not listed here makes the document invalid. + +| Field | Purpose | Required | Since | Type | |---|---|:-:|:-:|---| -| `fluidVersion` | Which contract version this file targets. Accepts `"0.7.4"` (or `"0.7.3"`). | ✅ | 0.1.0 | `enum` | -| `kind` | Product kind: `DataProduct` \| `MLPipeline`. | ✅ | 0.1.0 | `enum` | -| `id` | Globally unique, dot-separated product id. The public address of this product. | ✅ | 0.1.0 | `string` | -| `name` | Human-readable display name. | ✅ | 0.1.0 | `string` | -| `description` | Business-facing summary. | | 0.1.0 | `string` | -| `domain` | Owning business domain (e.g. `Finance`, `Customer Experience`). | | 0.1.0 | `string` | -| `tags` | Free-text categorization labels. | | 0.1.0 | `string[]` | -| `labels` | Key/value categorization for catalog and IAM. | | 0.1.0 | `map` | -| `metadata` | Owner, layer, business context. | ✅ | 0.1.0 | `object` | -| `consumes` | Upstream FLUID products this product depends on (with version constraints). | | 0.2.0 | `object[]` | -| `exposes` | Ports this product publishes (table / api / stream / feature_store / …). | ✅ | 0.1.0 | `object[]` | -| `build` | How the product is produced (pattern + engine + properties). | | 0.3.0 | `object` | -| `builds` | Multi-build orchestration (when one product has several builds). | | 0.5.7 | `object[]` | -| `orchestration` | Schedule + tasks for the build. Provider-first since 0.7.0. | | 0.7.0 | `object` | -| `sovereignty` ⭐ | Data residency & jurisdictional compliance. | | **0.7.1** | `object` | -| `accessPolicy` ⭐ | Root-level IAM grants — generates cloud IAM bindings. | | **0.7.1** | `object` | -| `retention` ⭐ | TTLs for runState / runLogs / lineage / dlq (ISO-8601). | | **0.7.3** | `object` | -| `lineage` | Upstream/downstream graph + optional field-level mappings. | | 0.4.0 | `object` | -| `governance` | Contract-level (account/project-wide) governance — AWS Lake Formation admins + LF-tag definitions. Per-resource grants live under `binding.governance`. | | 0.4.0 | `object` | -| `schemaEvolution` | Compatibility strategy (`backward` / `forward` / `full`). | | 0.4.0 | `object` | -| `machineLearning` | Model spec for ML-product kinds (features, training, inference). | | 0.5.7 | `object` | -| `environments` | Per-environment overrides (`dev` / `staging` / `prod`). | | 0.5.7 | `object` | -| `lifecycle` | `state: preview \| active \| deprecated \| retired` + deprecation date. | | 0.5.7 | `object` | -| `docs` | External documentation URLs (runbook, design doc, dashboard). | | 0.5.7 | `object` | -| `extensions` | Escape hatch for vendor-specific additions. | | 0.5.7 | `object` | +| `fluidVersion` | The schema version this document declares; it selects the schema the document is validated against. 0.7.5 accepts `"0.7.3"`, `"0.7.4"`, `"0.7.5"` — but see [Choosing `fluidVersion`](/fluid/schema/versions#choosing-fluidversion). | ✅ | 0.0.1 | `enum` | +| `kind` | `DataProduct` \| `MLPipeline`. | ✅ | 0.0.1 | `enum` | +| `id` | Globally unique product id — the product's public address. | ✅ | 0.0.1 | identifier | +| `name` | Human-readable display name. | ✅ | 0.0.1 | `string` | +| `metadata` | Owner, layer, product type, business context. | ✅ | 0.0.1 | `object` | +| `exposes` | Ports this product publishes. | ✅ | 0.0.1 | `object[]` | +| `description` | Business-facing summary. | | 0.0.1 | `string` | +| `domain` | Owning business domain. | | 0.0.1 | `string` | +| `consumes` | Upstream FLUID products this product reads. | | 0.0.1 | `object[]` | +| `build` | How the product is produced (pattern + engine + properties). | | 0.0.1 | `object` | +| `builds` | Several builds for one product, same shape as `build`. | | 0.5.7 | `object[]` | +| `tags` | Lower-case categorization tags (`^[a-z0-9][a-z0-9-]*[a-z0-9]$`). | | 0.5.7 | `string[]` | +| `labels` | Key/value labels. | | 0.5.7 | `map` | +| `lineage` | `granularity`, `upstream[]` (with `fieldMappings`), `downstream[]`. | | 0.5.7 | `object` | +| `schemaEvolution` | `strategy` and `compatibility` of the product's schema changes. | | 0.5.7 | `object` | +| `machineLearning` | `enabled`, `framework`, `models[]`. | | 0.5.7 | `object` | +| `environments` | Per-environment overrides of `metadata` and `exposes`. | | 0.5.7 | `map` | +| `lifecycle` | `state`, `retention`, `deprecationPolicy`. | | 0.5.7 | `object` | +| `docs` | `homepage`, `runbook`, `dictionary`, `changeLog`. | | 0.5.7 | `object` | +| `sovereignty` | Jurisdiction and data residency. | | **0.7.1** | `object` | +| `accessPolicy` | Access grants for the product. (A different `accessPolicy` existed from 0.0.1 to 0.4.0.) | | **0.7.1** | `object` | +| `orchestration` | Scheduling and workflow-engine settings. Open object — see [below](#orchestration-fields). | | **0.7.2** | `object` | +| `retention` | ISO-8601 TTLs for run state, run logs, lineage events and DLQ records. | | **0.7.3** | `object` | +| `governance` | Account-wide AWS Lake Formation settings: `admins`, `tagDefinitions`. (A different `governance` existed from 0.0.1 to 0.4.0.) | | **0.7.3** | `object` | +| `extensions` | Vendor- or plugin-namespaced configuration. Open object. (Also in 0.0.1; absent from 0.1.0 to 0.7.2.) | | **0.7.3** | `object` | +| 🧪 `packaging` | Container ownership: `mode` (`isolated` \| `shared`), `pool`, `poolManifest`, `containers`. | | 0.7.6 | `object` | +| 🧪 `consumers` | Declared downstream artifacts: `name` and `type` required; `label`, `owner`, `url`, `maturity`, `exposeIds`. | | 0.7.6 | `object[]` | --- -## `metadata` (required) — fields +## `metadata` (required) fields -`metadata` itself is required at the top level, but only `metadata.owner` is required inside it. +`metadata` is required, and inside it only `owner` is. The schema requires **no** member of `owner`, so `owner: {}` validates; name at least a `team`. -| Field | Purpose | Required | -|---|---|:-:| -| `metadata.owner` | Owning team object. | ✅ | -| `metadata.owner.team` | Owning team handle. | ✅ | -| `metadata.owner.email` | Owner email for alerts and audit. | | -| `metadata.owner.slack` | Owner Slack channel. | | -| `metadata.layer` | Medallion layer label. **Free-form string** — `Bronze`/`Silver`/`Gold` is convention, not enum. | | -| `metadata.businessContext.domain` | Business domain. | | -| `metadata.businessContext.subdomain` | Business subdomain. | | -| `metadata.provenance` | Auto-injected generation envelope (tool + version + timestamp). Machine-written. | | +| Field | Purpose | Required | Since | +|---|---|:-:|:-:| +| `metadata.owner` | Owner object: `team`, `email`, `slack`, `oncall`. | ✅ | 0.0.1 | +| `metadata.layer` | Layer label. **Free-form string** — `Bronze`/`Silver`/`Gold` is convention, not enum. | | 0.0.1 | +| `metadata.productType` | `SDP` (source-aligned) \| `ADP` (aggregated) \| `CDP` (consumption-aligned). | | 0.7.3 | +| `metadata.classification` | `public` \| `internal` \| `confidential` \| `restricted`. (A `classification` also existed in 0.0.1.) | | 0.7.3 | +| `metadata.experimental` | Experimental features the contract opts into. | | 0.7.3 | +| `metadata.businessContext` | `domain`, `subdomain`, `businessCapability`, `valueStream`. | | 0.5.7 | +| `metadata.tags` | Tags on the metadata. (Also existed from 0.0.1 to 0.4.0.) | | 0.7.2 | +| `metadata.createdAt` | Creation time (`format: date-time`, an annotation). | | 0.5.7 | +| `metadata.provenance` | Generation envelope written by tooling. | | 0.7.2 | --- -## `exposes[]` — fields +## `exposes[]` fields -Each entry in `exposes[]` requires `exposeId`, `kind`, `contract`, and `binding`. +Each entry requires `exposeId`, `kind`, `contract` and `binding`, and is closed. | Field | Purpose | Required | |---|---|:-:| | `exposes[].exposeId` | Stable id of this port within the product. | ✅ | | `exposes[].kind` | `table` \| `view` \| `api` \| `file` \| `stream` \| `topic` \| `feature_store` \| `model` \| `vector` \| `graph` \| `time_series` \| `other`. | ✅ | -| `exposes[].contract` | Schema + dq + policy. Must define at least `schema` or `openapiRef`. | ✅ | -| `exposes[].binding` | Where it lives — `platform`, `format`, `location`. | ✅ | -| `exposes[].description` | What this port is. *(NEW in 0.7.2)* | | +| `exposes[].contract` | The data's shape and promises. Must contain `schema` or `openapiRef`. | ✅ | +| `exposes[].binding` | Where it lives — `platform`, `format`, `location` (all three required). | ✅ | +| `exposes[].policy` | **Sibling of `contract`**, not inside it: `authn`, `authz`, `privacy`, `classification`, `agentPolicy`. | | +| `exposes[].title`, `exposes[].description` | Display name and description of the port (`description` since 0.7.2; it also existed from 0.1.0 to 0.4.0). | | | `exposes[].version` | Semver of this port. | | -| `exposes[].contract.schema[]` | Column list with type, required, sensitivity, tags. | (one of) | -| `exposes[].contract.openapiRef` | Reference to an external OpenAPI spec (for `kind: api`). | (one of) | -| `exposes[].contract.dq.rules[]` | Declarative quality assertions. | | -| `exposes[].contract.policy` | Privacy + RBAC + column-level access. | | -| `exposes[].semantics` ⭐ 0.7.2 | Entities, measures, dimensions, metrics. | | -| `exposes[].mcp` ⭐ 0.7.4 | Opts this expose into the Fluid MCP output-port gateway. Presence enables runtime `agentPolicy` enforcement on every read. | | -| `exposes[].mcp.sampling.maxRows` ⭐ 0.7.4 | Hard cap for the `sample` MCP tool. Integer ≥ 1, default `100`. | | -| `exposes[].mcp.classification.dataClass` ⭐ 0.7.4 | Advisory data class surfaced on `describe`: `public` \| `internal` \| `confidential` \| `restricted`. | | -| `exposes[].binding.platform` | `gcp` \| `aws` \| `azure` \| `snowflake` \| `databricks` \| `kafka` \| `local` \| `kubernetes` \| `postgres` ⭐ 0.7.4 \| `other`. | | -| `exposes[].binding.format` | `bigquery_table` \| `snowflake_table` \| `snowflake_view` \| `iceberg` \| `delta_table` \| `parquet` \| `csv` \| `json` \| `http_api` \| `grpc_api` \| `kafka_topic` \| `pubsub_topic` \| `gcs_file` \| `s3_file` \| `redshift_table` \| `redshift_serverless` \| `redshift_external_schema` \| `postgres_table` ⭐ 0.7.4 \| `athena_table` ⭐ 0.7.4 \| `glue_table` ⭐ 0.7.4 \| `other`. | | -| `exposes[].binding.icebergConfig` ⭐ 0.7.2 | Iceberg specifics when `format: iceberg`. | | -| `exposes[].binding.governance` | Per-resource AWS Lake Formation governance — `registerLocation`, principal grants, tag associations, row/column filters. Account-wide settings live in the top-level `governance` block. | | -| `exposes[].qos` | Availability, freshness, latency targets. | | +| `exposes[].contract.schema[]` | Columns: `name` and `type` required; `required`, `description`, `sensitivity`, `semanticType`, `businessName`, `businessDefinition`, `validationRules`, `tags`, `labels`. | (one of) | +| `exposes[].contract.openapiRef` | Reference to an OpenAPI document (for `kind: api`). | (one of) | +| `exposes[].contract.dq.rules[]` | Declarative quality assertions (below). | | +| `exposes[].contract.schemaPolicy` | `strict` \| `discover_and_freeze` \| `evolve_safe` \| `evolve_all` (since 0.7.3). | | +| `exposes[].contract.schemaSignature` | `sha256:` + 64 hex — a digest of the schema. | | +| `exposes[].qos` | `availability` (a percentage string such as `"99.9%"`), `freshnessSLO`, `latencyP95` (ISO-8601 durations), `dataLossSLO`, `completenessTarget`, `errorBudget`. | | +| `exposes[].semantics` | Entities, measures, dimensions, metrics (since 0.7.2; a `semantics` member with a different shape existed from 0.1.0 to 0.4.0). | | +| `exposes[].mcp` | Serve the port to AI agents over MCP (since 0.7.4): `sampling.maxRows` (integer ≥ 1), `classification.dataClass` (`public` \| `internal` \| `confidential` \| `restricted`). | | +| `exposes[].lifecycle`, `exposes[].observability`, `exposes[].docs` | Per-port lifecycle, observability (`metrics`, `onBreach`, `defaultSLIs`, `alert`) and docs. | | +| `exposes[].binding.platform` | `gcp` \| `aws` \| `azure` \| `snowflake` \| `databricks` \| `kafka` \| `confluent` (0.7.5) \| `local` \| `kubernetes` \| `postgres` (0.7.4) \| `pgvector` (0.7.5) \| `other`. | ✅ | +| `exposes[].binding.format` | `bigquery_table` \| `snowflake_table` \| `snowflake_view` \| `iceberg` \| `delta_table` \| `parquet` \| `csv` \| `json` \| `http_api` \| `grpc_api` \| `kafka_topic` \| `pubsub_topic` \| `gcs_file` \| `s3_file` \| `redshift_table` \| `redshift_serverless` \| `redshift_external_schema` \| `postgres_table` (0.7.4) \| `athena_table` (0.7.4) \| `glue_table` (0.7.4) \| `pgvector_table` (0.7.5) \| `other`. | ✅ | +| `exposes[].binding.location` | Platform address; no member is required by the schema. 0.7.5 adds `environment_id`, `kafka_cluster_id`, `confluent_role_arn` (Tableflow), `catalog`, `warehouse`, `uri`, `partitionBy` (Iceberg catalogs), `namespace`, `workgroup`, `iam_role_arn`, `external_schema`, `glue_database` (Redshift Serverless), `stream` (Kinesis). | ✅ | +| `exposes[].binding.icebergConfig` | Iceberg table settings when `format: iceberg` (since 0.7.2). | | +| `exposes[].binding.vectorConfig` | Vector output port for `pgvector` (since 0.7.5): `dimensions` (required, 1–16000), `embeddingModel`, `vectorType`, `indexType`, `distanceMetric`, `hnsw`, `ivfflat`, `table`, `sourceKeyColumn`. | | +| `exposes[].binding.governance` | Per-resource AWS Lake Formation settings: `registerLocation`, `grants[]`, `tags`, `rowFilter` (since 0.7.3). Account-wide settings live in the top-level `governance` block. | | +| 🧪 `exposes[].binding.encryption.kms` | `product` \| `none` \| an AWS KMS alias or ARN \| a Cloud KMS key name. | | +| 🧪 `exposes[].binding.principals` | Logical principal → identity (or list) on this platform. | | +| 🧪 `exposes[].binding.packaging` | Per-binding override of `packaging`. | | +| 🧪 `exposes[].binding.governance.lakeFormation.bucketPolicy` | `cross-account` (default) \| `none` \| `all-grantees`. **Authoritative** — an emitted bucket policy replaces every other statement on the bucket; see the [0.7.6 notes](/fluid/releases/0.7.6#binding-fields-encryption-principals-lake-formation-bucket-policy). | | +| 🧪 `exposes[].lifecycle.expire` | `true` deletes data older than `retention` (default `false`). | | | `exposes[].tags`, `exposes[].labels` | Port-level categorization. | | -### `contract.dq.rules[]` — fields +### `contract.dq.rules[]` fields | Field | Purpose | Required | |---|---|:-:| -| `dq.rules[].id` | Rule id (unique within the contract). | ✅ | +| `dq.rules[].id` | Rule id. | ✅ | | `dq.rules[].type` | `freshness` \| `completeness` \| `uniqueness` \| `valid_values` \| `accuracy` \| `schema` \| `anomaly_detection` \| `drift_detection`. | ✅ | | `dq.rules[].severity` | `info` \| `warn` \| `error` \| `critical`. | ✅ | -| `dq.rules[].selector` | SQL-like predicate (e.g. `"amount > 0"`). | | +| `dq.rules[].selector` | Predicate or column the rule applies to (e.g. `"amount > 0"`). | | | `dq.rules[].threshold` | Numeric threshold for the operator. | | | `dq.rules[].operator` | `>=` \| `>` \| `<=` \| `<` \| `==` \| `!=`. | | | `dq.rules[].window` | ISO-8601 duration window (e.g. `PT15M`). | | -| `dq.rules[].description` | Human-readable purpose. | | --- -## `build` — fields +## `build` / `builds[]` fields + +A build is closed and has no required member. | Field | Purpose | Required | |---|---|:-:| -| `build.pattern` | `hybrid-reference` \| `embedded-logic` \| `multi-stage` \| `acquisition` ⭐ 0.7.3. | | -| `build.engine` | `dbt` \| `sql` \| `python` \| `spark` \| `glue` \| `custom` \| `duckdb` ⭐ \| `airbyte` ⭐ \| `meltano` ⭐ \| `dlt` ⭐ \| `kafka-connect` ⭐ \| `debezium` ⭐ (⭐ added 0.7.3). | | -| `build.capabilities` ⭐ 0.7.3 | What the build asks for: `full_refresh`, `incremental_dedup`, `cdc`, `schema_discovery`, `dlp_scan`, `exactly_once`, … | | -| `build.properties` | Pattern-specific properties (validated conditionally). | | -| `build.execution.trigger` | `schedule` / `event` / `manual`. | | -| `build.execution.runtime` | Where it runs (platform + resources). | | -| `build.execution.retries` | Retry policy. | | -| `build.outputs` | List of `exposeId`s this build produces. | | -| `build.dependencies` | Other builds this one waits on. | | -| `build.properties` | Pattern-specific payload — schema validates this against `acquisitionPattern` when `pattern: acquisition` (see below). | (when `pattern: acquisition`) | +| `build.id` | Build id; give one to each entry of `builds[]`. | | +| `build.pattern` | `hybrid-reference` \| `embedded-logic` \| `multi-stage` \| `acquisition` (0.7.3). Selects the shape of `properties`. | | +| `build.engine` | `dbt` \| `sql` \| `python` \| `spark` \| `glue` \| `custom` \| `duckdb` \| `airbyte` \| `meltano` \| `dlt` \| `kafka-connect` \| `debezium` (the last six since 0.7.3), or `dbt-` such as `dbt-databricks` (since 0.7.2). | | +| `build.capabilities` | What the build asks of its runner: `full_refresh`, `incremental_append`, `incremental_dedup`, `incremental_merge`, `cdc`, `streaming`, `schema_discovery`, `schema_evolution`, `dlp_scan`, `at_most_once`, `at_least_once`, `exactly_once` (since 0.7.3). | | +| `build.properties` | Pattern-specific settings. `hybrid-reference` requires `model`; `embedded-logic` requires `sql`; `acquisition` requires `source` (below). | | +| `build.execution.trigger` | `type`: `schedule` \| `event` \| `manual` \| `dependency` \| `dataset` \| `schedule_and_dataset` \| `timetable`, plus `schedule`, `cron`, `timezone`, … | | +| `build.execution.runtime` | Where it runs: `platform`, `resources`, `image`, `executor`, `serviceAccount`, `timeout`. | | +| `build.execution.retries` | `maxAttempts`, `backoffStrategy` (`fixed` \| `exponential` \| `linear`), `initialDelay`, `maxDelay`. | | +| `build.execution.notifications[]` | `type` (`email` \| `slack` \| `webhook` \| `pagerduty`), `target`, `condition` (`success` \| `failure` \| `always`). Closed since 0.7.2. | | +| `build.execution.orchestration` | The same shape as top-level `orchestration`. | | +| `build.outputs` | The `exposeId`s this build produces. | | +| `build.repository` | Where referenced code lives. | | --- -## `build.properties` when `pattern: acquisition` ⭐ 0.7.3 — fields +## `build.properties` when `pattern: acquisition` fields -When `build.pattern: acquisition`, the schema validates `build.properties` against the `acquisitionPattern` shape below. (For other patterns, `properties` conforms to `hybridReferencePattern` / `embeddedLogicPattern` / `multiStagePattern` respectively.) +When `pattern: acquisition` (0.7.3+), the schema validates `properties` against the acquisition shape below. (For the other patterns it uses the hybrid-reference, embedded-logic and multi-stage shapes.) | Field (under `build.properties`) | Purpose | Required | |---|---|:-:| -| `source.kind` | `filesystem` \| `postgres` \| `mysql` \| `sqlite` \| `http` \| `salesforce` \| `stripe` \| … (free string with example values, no enum). | ✅ | +| `source.kind` | Source system — a free string, e.g. `postgres`, `kafka`, `salesforce`. | ✅ | | `source.mode` | `full_refresh` \| `incremental_append` \| `incremental_dedup` \| `incremental_merge` \| `cdc` \| `streaming`. | ✅ | -| `source.cursor_field` | Column used as incremental cursor. | | -| `source.connection` | Connection details. `secretRef` must be a URI: `vault://...`, `aws://...`, `gcp://...`, `azure://...`, or `env://...`. Inline secrets are rejected by the validator. | | -| `source.streams` | List of streams/tables/objects to ingest. | | +| `source.cursor_field` | Column used as the incremental cursor. | | +| `source.connection` | Connection details (open object). `secretRef` must be a URI matching `://…`; the schema's description names `vault://`, `aws://`, `gcp://`, `azure://` and `env://`. | | +| `source.streams` | Streams / tables / objects to ingest. | | | `sink.format` | `iceberg` \| `delta` \| `parquet` \| `csv` \| `json` \| `snowflake_table` \| `bigquery_table` \| `redshift_table` \| `duckdb_table`. | | -| `sink.partitionBy` | Array of **strings** (function-form, e.g. `["day(ingested_at)"]`). Note: `binding.icebergConfig.partitionSpec` uses an object form — different shape. | | -| `sink.catalog` | Catalog identifier (e.g. `glue`, `snowflake_horizon`). | | -| `delivery.guarantee` | `at_most_once` \| `at_least_once` \| `exactly_once`. | | -| `delivery.idempotencyKey` | Template, e.g. `"${stream}\|${batch_id}"`. | | -| `delivery.dlq.enabled` | Boolean — DLQ on/off (default `true`). | | -| `delivery.dlq.sink.format` | `parquet` \| `json` \| `ndjson`. | | -| `delivery.dlq.sink.location` | DLQ destination URI. | | -| `delivery.dlq.maxRecordsBeforeAbort` | Integer (default 10000). | | -| `delivery.dlq.alertOn` | List: `pii_classification_failed` \| `schema_violation` \| `destination_write_failed` \| `quality_gate_failed`. | | -| `schemaEvolution.policy` | `strict` \| `discover_and_freeze` \| `evolve_safe` \| `evolve_all`. | | -| `schemaEvolution.onAddedColumn` | `include` \| `warn` \| `fail`. | | -| `schemaEvolution.onRemovedColumn` | `drop` \| `warn` \| `fail`. | | -| `schemaEvolution.onTypeChange` | `cast` \| `warn` \| `fail`. | | -| `schemaEvolution.sourceFingerprint` | `required` \| `optional` \| `disabled` (default: `required`). | | +| `sink.catalog` | `rest` \| `glue` \| `nessie` \| `unity` \| `snowflake-managed` \| `hive`. | | +| `sink.partitionBy` | Array of **strings** (function form, e.g. `["day(ingested_at)"]`) — not the object form of `binding.icebergConfig.partitionSpec`. | | +| `delivery.guarantee` | `at_most_once` \| `at_least_once` (default) \| `exactly_once`. | | +| `delivery.idempotencyKey` | Key template (default `{run_id}:{stream}:{record_pk}`). | | +| `delivery.dlq` | `enabled` (default `true`), `sink.format` (`parquet` \| `json` \| `ndjson`), `sink.location`, `maxRecordsBeforeAbort` (default 10000), `alertOn` (`pii_classification_failed` \| `schema_violation` \| `destination_write_failed` \| `quality_gate_failed`). | | +| `schemaEvolution.policy` | `strict` (default) \| `discover_and_freeze` \| `evolve_safe` \| `evolve_all`. | | +| `schemaEvolution.onAddedColumn` / `onRemovedColumn` / `onTypeChange` | `include` \| `warn` \| `fail` / `drop` \| `warn` \| `fail` / `cast` \| `warn` \| `fail`. | | +| `schemaEvolution.sourceFingerprint` | `required` (default) \| `optional` \| `disabled`. | | | `preLand` | Hook chain: `dlp_scan`, `tokenize_pii`, `quality_gate`, `emit_lineage_input`. | | -| `quality` | Pre-land quality gates (gate-or-quarantine semantics). | | -| `cost.budget.monthly` | `{ rows, bytes (e.g. "50GB"), computeMinutes }`. | | -| `cost.budget.onExceed` | `warn` (default) \| `abort`. | | -| `cost.chargeback` | `{ team, project, costCenter }`. | | -| `catalog.register` | Catalog auto-registration targets: `datahub`, `openmetadata`, `datamesh_manager`, `unity`, `glue`, `snowflake_horizon`. | | -| `concurrency.lock` | Single-flight lock scope (`product` or `build`). | | -| `lineage.emit` | Whether to emit OpenLineage events (default: `true`). | | -| `` (`airbyte` \| `meltano` \| `dlt` \| `duckdb` \| `kafka-connect` \| `debezium`) | Engine-specific config block. | | -| `.deployment.mode` | `embedded` \| `bring-your-own` \| `managed`. | | -| `.deployment.managed.target` | When `managed`: `docker` \| `kubernetes` \| `terraform` \| `opentofu`. | | -| `.image_signature.verifier` | `cosign` (only value today, default `cosign`). | (prod) | -| `.image_signature.publicKey` | Public key reference for signature verification. | (prod) | -| `.image_signature.slsaProvenance` | `required` \| `optional` (default) \| `disabled`. | (prod) | +| `quality` | Pre-land quality `gates[]` (`rule`, `severity` required), `onError`, `anomalies[]`. | | +| `cost.budget` | `monthly` (`rows`, `bytes`, `computeMinutes`) and `onExceed` (`warn` (default) \| `abort`). | | +| `cost.chargeback` | `team`, `project`, `costCenter`. | | +| `catalog.register` | `datahub`, `openmetadata`, `datamesh_manager`, `unity`, `glue`, `snowflake_horizon` (the last three since 0.7.4). | | +| `concurrency.lock` | `scope` (`product` (default) \| `build`), `timeout`, `onContended` (`abort` \| `queue` \| `replace`). | | +| `lineage.emit` | Emit lineage events (default `true`). | | +| `` (`duckdb` \| `airbyte` \| `meltano` \| `dlt` \| `kafka-connect` \| `debezium`) | Engine-specific settings. | | +| `.deployment.mode` | `embedded` (default) \| `bring-your-own` \| `managed`; with `managed`, `managed.target` is `docker` \| `kubernetes` \| `terraform` \| `opentofu`. | | +| `.image_signature` | `verifier` (`cosign`), `publicKey`, `slsaProvenance` (`required` \| `optional` (default) \| `disabled`). | | +| `kafka-connect.iceberg_sink_enabled`, `sink_topics`, `streamingSink`, `iceberg_catalog_overrides` | Opt-in Iceberg streaming sink (since 0.7.5). | | --- -## `exposes[].mcp` ⭐ 0.7.4 — fields - -The presence of a (non-empty) `mcp` block opts the expose into the Fluid MCP output-port gateway, where `policy.agentPolicy` is enforced at runtime on every read. The `agentPolicy` shape is unchanged from 0.7.3. +## `exposes[].policy` fields | Field | Purpose | Required | |---|---|:-:| -| `exposes[].mcp.sampling.maxRows` | Hard cap for the `sample` MCP tool (defends against an over-curious agent). Integer ≥ 1, default `100`. | | -| `exposes[].mcp.classification.dataClass` | Advisory data class surfaced on the `describe` tool: `public` \| `internal` \| `confidential` \| `restricted`. The real gate is `policy.agentPolicy`. | | - ---- +| `exposes[].policy.authn` | `oidc` \| `oauth2` \| `api_key` \| `none` \| `custom` \| `iam` \| `jwt`. | | +| `exposes[].policy.authz.readers` / `writers` | Principals with read / write access. | | +| `exposes[].policy.authz.columnRestrictions[]` | `principal`, `columns`, `access` (`allow` \| `deny`). 🧪 The 0.7.6 schema defines its meaning: a column named in any restriction is restricted; `deny` removes access; `allow` limits the columns to the principals an `allow` names; deny beats allow; a restriction never grants access. | | +| `exposes[].policy.privacy.masking[]` | `{column, strategy}` required; `strategy` is `mask` \| `hash` \| `tokenize` \| `encrypt` \| `k_anonymity`; `params` is open. 🧪 0.7.6 types `params`: `keepFirst`, `keepLast` (mask), `saltEnv` (hash), `keyEnv` (tokenize, encrypt). | | +| `exposes[].policy.privacy.rowLevelPolicy.expression` | Provider-specific row filter predicate (required inside `rowLevelPolicy`). | | +| `exposes[].policy.classification` | `Public` \| `Internal` \| `Confidential` \| `Restricted`. | | +| `exposes[].policy.agentPolicy` | AI-consumer policy (below). | | -## `exposes[].policy.agentPolicy` ⭐ 0.7.1 — fields +### `exposes[].policy.agentPolicy` (0.7.1) fields -> ⚠️ **Location.** `agentPolicy` is **per-expose**, nested under `exposes[].policy.agentPolicy`. It is not a top-level property in the schema (despite the 0.7.1 release notes showing it that way). +> ⚠️ **Location.** `agentPolicy` is **per-expose**, under `exposes[].policy.agentPolicy`. It is not a top-level property. > -> ⭐ **0.7.4 — runtime enforcement.** When the same expose carries an `mcp` block, these `allowedModels` / `deniedModels` and `allowedUseCases` / `deniedUseCases` fields are enforced by the Fluid MCP gateway on every read (no shape change). +> **0.7.4 — runtime enforcement.** When the same expose carries an `mcp` block, an MCP gateway that implements 0.7.4 enforces the model and use-case lists on every read; the reference implementation's gateway does. The shape is unchanged. -| Field | Purpose | Required | +| Field | Purpose | Default | |---|---|:-:| -| `…policy.agentPolicy.allowedModels` | Whitelist of AI model ids (e.g. `gpt-4`, `claude-3-opus`). Empty array = no AI access. | | -| `…policy.agentPolicy.deniedModels` | Blacklist (takes precedence over `allowedModels`). | | -| `…policy.agentPolicy.maxTokensPerRequest` | Per-request token cap (positive integer). | | -| `…policy.agentPolicy.maxTokensPerDay` | Per-day token cap. | | -| `…policy.agentPolicy.allowedUseCases` | Permitted use cases from the schema's controlled vocabulary: `inference` \| `reasoning` \| `analysis` \| `summarization` \| `classification` \| `embedding` \| `search` \| `qa` \| `code_generation` \| `fine_tuning` \| `training` \| `rag`. **Custom strings are rejected.** | | -| `…policy.agentPolicy.deniedUseCases` | Same controlled vocabulary as `allowedUseCases`. Useful pattern: allow `inference, qa, rag`; deny `training, fine_tuning`. | | -| `…policy.agentPolicy.canReason` | Boolean — allow chain-of-thought / multi-step reasoning? (default `false`) | | -| `…policy.agentPolicy.canStore` | Boolean — allow caching / persistence? (default `false`) | | -| `…policy.agentPolicy.retentionPolicy.maxRetentionDays` | Max days AI may retain data (`0` = no retention). | | -| `…policy.agentPolicy.retentionPolicy.requireDeletion` | Boolean — must AI delete after use? (default `true`) | | -| `…policy.agentPolicy.auditRequired` | Boolean — log all AI consumption? (default `true`) | | -| `…policy.agentPolicy.purposeLimitation` | Free-text purpose statement, e.g. `"Customer support chatbot only"`. | | - -## `exposes[].policy` — sibling fields +| `…agentPolicy.allowedModels` | Model ids allowed to read the expose (lower-case ids such as `gpt-4`, `claude-3-opus`). | | +| `…agentPolicy.deniedModels` | Model ids denied. | | +| `…agentPolicy.maxTokensPerRequest` / `maxTokensPerDay` | Token caps (integers ≥ 1). | | +| `…agentPolicy.allowedUseCases` / `deniedUseCases` | From a fixed vocabulary: `inference` \| `reasoning` \| `analysis` \| `summarization` \| `classification` \| `embedding` \| `search` \| `qa` \| `code_generation` \| `fine_tuning` \| `training` \| `rag`. **Other strings are invalid.** | | +| `…agentPolicy.canReason` | Allow multi-step reasoning over the data. | `false` | +| `…agentPolicy.canStore` | Allow caching or persisting the data. | `false` | +| `…agentPolicy.retentionPolicy` | `maxRetentionDays` (≥ 0), `requireDeletion`. | `requireDeletion: true` | +| `…agentPolicy.auditRequired` | Log AI consumption. | `true` | +| `…agentPolicy.purposeLimitation` | Free-text purpose statement. | | + +--- + +## `consumes[]` fields | Field | Purpose | Required | |---|---|:-:| -| `exposes[].policy.authn` | `oidc` \| `oauth2` \| `api_key` \| `none` \| `custom` \| `iam` \| `jwt`. | | -| `exposes[].policy.authz.readers` | List of principals with read access. | | -| `exposes[].policy.authz.writers` | List of principals with write access. | | -| `exposes[].policy.authz.columnRestrictions[]` | Per-principal column allow/deny rules. | | -| `exposes[].policy.privacy.masking[]` | Column-level masking rules: `{column, strategy: mask\|hash\|tokenize\|encrypt\|k_anonymity}`. | | -| `exposes[].policy.privacy.rowLevelPolicy.expression` | Provider-specific row-level filter predicate. | | -| `exposes[].policy.classification` | Data classification label. | | +| `consumes[].productId` | Upstream product id. | ✅ | +| `consumes[].exposeId` | Upstream port. | ✅ | +| `consumes[].versionConstraint` | Semver range, e.g. `^2.0.0`. | | +| `consumes[].qosExpectations` | `freshnessMax`, `maxStaleness` (durations), `minCompleteness` (0–1). | | +| `consumes[].requiredPolicies` | Policies the upstream must carry. | | +| `consumes[].purpose` | Why this product reads the upstream. | | +| 🧪 `consumes[].upstreamWorkspace` | The other mesh that owns the upstream. Requires `upstreamDigest`. | | +| 🧪 `consumes[].upstreamDigest` | `sha256:` + 64 lowercase hex — the upstream contract this product was composed against. | | --- -## `sovereignty` ⭐ 0.7.1 — fields +## `sovereignty` (0.7.1) fields -| Field | Purpose | Required | -|---|---|:-:| -| `sovereignty.jurisdiction` | Required legal jurisdiction (e.g. `EU`, `US`, `Multi-Region`). | | -| `sovereignty.allowedRegions` | Explicit allowed cloud regions. | | -| `sovereignty.deniedRegions` | Explicit denied cloud regions. | | -| `sovereignty.dataResidency` | Boolean — must data stay within jurisdiction? | | -| `sovereignty.crossBorderTransfer` | Boolean — is cross-border transfer permitted? | | -| `sovereignty.transferMechanisms` | If permitted, legal mechanisms (`SCCs`, `BCRs`, …). | | -| `sovereignty.regulatoryFramework` | Frameworks: `GDPR`, `HIPAA`, `CCPA`, `LGPD`. | | -| `sovereignty.enforcementMode` | `strict` (block apply) \| `advisory` (warn only) \| `audit`. | | -| `sovereignty.validationRequired` | Whether `binding` regions are validated against this block at apply-time. | | +| Field | Purpose | +|---|---| +| `sovereignty.jurisdiction` | `EU` \| `US` \| `UK` \| `CA` \| `AU` \| `JP` \| `CN` \| `IN` \| `BR` \| `Global` \| `Multi-Region`. | +| `sovereignty.allowedRegions` / `deniedRegions` | Cloud regions allowed / denied. | +| `sovereignty.dataResidency` | Boolean — must data stay within the jurisdiction? | +| `sovereignty.crossBorderTransfer` | Boolean — is cross-border transfer permitted? | +| `sovereignty.transferMechanisms` | `SCCs` \| `BCRs` \| `Adequacy` \| `DPF` \| `Consent` \| `Derogation`. | +| `sovereignty.regulatoryFramework` | `GDPR` \| `CCPA` \| `CPRA` \| `HIPAA` \| `PIPEDA` \| `LGPD` \| `PDPA` \| `POPIA` \| `DPA` \| `APPI`. | +| `sovereignty.enforcementMode` | `strict` \| `advisory` \| `audit`. | +| `sovereignty.validationRequired` | Whether bindings are checked against this block. | --- -## `accessPolicy` ⭐ 0.7.1 — fields +## `accessPolicy` (0.7.1) fields -Each grant requires only `principal`; everything else is opt-in. +Each grant requires only `principal`, and is closed. | Field | Purpose | Required | |---|---|:-:| -| `accessPolicy.grants[].principal` | `user:`, `group:`, `serviceAccount:` identifier. | ✅ | -| `accessPolicy.grants[].permissions` | List of permissions: `read` \| `select` \| `query` \| `write` \| `insert` \| `update` \| `delete` \| `create` \| `admin` \| `manage`. | | -| `accessPolicy.grants[].resources` | JSONPath expressions selecting which exposes are in scope. | | -| `accessPolicy.grants[].conditions` | Conditional access (IP ranges, time windows). | | +| `accessPolicy.grants[].principal` | The principal, e.g. `group:analysts@example.com`. | ✅ | +| `accessPolicy.grants[].permissions` | `read` \| `select` \| `query` \| `write` \| `insert` \| `update` \| `delete` \| `create` (0.7.2) \| `admin` \| `manage`. | | +| `accessPolicy.grants[].resources` | Strings selecting the exposes in scope, e.g. JSONPath. | | +| `accessPolicy.grants[].conditions` | Open object for conditional access (IP ranges, time windows). | | --- -## `retention` ⭐ 0.7.3 — fields +## `governance` (0.7.3) fields -| Field | Purpose | Default | +| Field | Purpose | +|---|---| +| `governance.lakeFormation.admins` | IAM ARNs. **Authoritative:** applying it replaces the account's Lake Formation admin list, so list every admin, including the identity that applies it. The schema's description adds that applying it also clears the account's create-database and create-table default permissions, trusted resource owners and parameters, and that destroying the emitted `aws_lakeformation_data_lake_settings` resource empties the admin list and resets `CROSS_ACCOUNT_VERSION` to 1. | +| `governance.lakeFormation.tagDefinitions` | LF-tag key → allowed values. | + +--- + +## `retention` (0.7.3) fields + +| Field | Purpose | Default (schema description) | |---|---|:-:| -| `retention.runState` | How long the orchestrator keeps run state. | `P30D` | +| `retention.runState` | How long run state is kept. | `P30D` | | `retention.runLogs` | How long run logs are kept. | `P90D` | | `retention.lineage` | How long emitted lineage events are kept. | `P365D` | -| `retention.dlq` | How long DLQ records are kept before purge. | `P180D` | +| `retention.dlq` | How long DLQ records are kept. | `P180D` | -All values are [ISO-8601 durations](https://en.wikipedia.org/wiki/ISO_8601#Durations). +All values are [ISO-8601 durations](https://en.wikipedia.org/wiki/ISO_8601#Durations). `retention` covers operational records, not the product's data; for data, see `lifecycle.retention` (and, 🧪 in 0.7.6, `exposes[].lifecycle.expire`). --- -## `orchestration` — fields +## `orchestration` fields | Field | Purpose | Required | |---|---|:-:| | `orchestration.engine` | `airflow` \| `dagster` \| `prefect` \| `kubeflow` \| `custom` \| `none`. | ✅ | -| `orchestration.mode` | Pull or push generation. | | -| `orchestration.generateOnChange` | Auto-regenerate DAGs when contract changes. | | -| `orchestration.tasks[].taskId` | Task id within the DAG. | | -| `orchestration.tasks[].type` | `provider_action` \| `fluid_validate` \| `fluid_plan` \| `fluid_apply` \| `fluid_execute` \| `fluid_verify` \| `fluid_dq_check` \| `bash` \| `python` \| `branch_python` \| `sensor` \| `email` \| `http` \| `snowflake_query` \| `bigquery_query` \| `bigquery_job` \| `glue_job` \| `databricks_job` \| `custom`. | | -| `orchestration.tasks[].provider` | For `provider_action`: e.g. `aws.s3`, `snowflake.table`. | | -| `orchestration.tasks[].action` | For `provider_action`: e.g. `ensure_bucket`, `ensure`. | | -| `orchestration.tasks[].parameters` | Action parameters. | | -| `orchestration.tasks[].dependsOn` | Upstream task ids. | | -| `orchestration.tasks[].buildStepRef` | Reference into `build` step. | | +| `orchestration.mode` | `generated` \| `manual` \| `hybrid`. | | +| `orchestration.generateOnChange` | Regenerate the workflow when the contract changes. | | +| `orchestration.airflow.dagId` | DAG id; required when `airflow` is present. | | +| `orchestration.airflow.dagConfig` | `schedule`, `startDate`, `catchup`, `maxActiveRuns`, … | | +| `orchestration.airflow.tasks[].taskId` | Task id (lower-case). | ✅ | +| `orchestration.airflow.tasks[].type` | `provider_action` \| `fluid_validate` \| `fluid_plan` \| `fluid_apply` \| `fluid_execute` \| `fluid_verify` \| `fluid_dq_check` \| `bash` \| `python` \| `branch_python` \| `sensor` \| `email` \| `http` \| `snowflake_query` \| `bigquery_query` \| `bigquery_job` \| `glue_job` \| `databricks_job` \| `custom`. | | +| `orchestration.airflow.tasks[].operator` | Operator name. **Set it** — see the note below. | | +| `orchestration.airflow.tasks[].provider` | For `provider_action`: `aws` \| `gcp` \| `azure` \| `snowflake` \| `databricks` \| `kafka` \| `kubernetes` \| `local` \| `custom`. | (provider_action) | +| `orchestration.airflow.tasks[].action` | For `provider_action`: `.`, e.g. `s3.ensure_bucket`. | (provider_action) | +| `orchestration.airflow.tasks[].params` | Action parameters (open object). | (provider_action) | +| `orchestration.airflow.tasks[].dependencies` | Upstream task ids. | | +| `orchestration.airflow.tasks[].buildStepRef` | Reference into a build step. | | + +> ⚠️ **`orchestration.tasks` is not validated.** `orchestration` is an open object up to 0.7.6, so a top-level `orchestration.tasks` list is accepted without any check. Its keys are implementation-defined: the reference implementation's AWS and GCP code generators read `params` and `dependsOn` there, and its Snowflake Airflow generator reads `parameters`. +> +> **Name each Airflow task's `operator`.** In the 0.7.1 to 0.7.6 schemas the conditional rules for `fluid_execute`, `bash` and `python` tasks also match a task with no `operator`, so a task that has no `operator` and carries `params` must satisfy all three at once and is in practice rejected. A `provider_action` task always carries `params`, so it is always affected. A task with neither `operator` nor `params` validates. --- -## `lifecycle` — values - -`lifecycle.state` ∈ `preview` \| `active` \| `deprecated` \| `retired`. +## `lifecycle` values -A `deprecated` product still serves reads but new `consumes:` references should fail validation in CI. +`lifecycle.state` ∈ `preview` \| `active` \| `deprecated` \| `retired` — the state of the *product*, unrelated to preview *schema versions*. Other members: `retention` (ISO-8601 duration) and `deprecationPolicy` (`noticePeriod`, `contact`, `replacement`). --- ## Where each field is exhaustively documented -The auto-generated reference at [`specs/0.7.4/fluid-spec.html`](/fluid/specs/0.7.4/fluid-spec.html) is the authoritative source for every field's exact type, enum values, validation rules, and examples. +The auto-generated reference at [`specs/0.7.5/fluid-spec.html`](/fluid/specs/0.7.5/fluid-spec.html) renders every field's exact type, enum values and validation rules from the schema. diff --git a/docs/schema/minimal-contract.md b/docs/schema/minimal-contract.md index 619a7d1..505fbe9 100644 --- a/docs/schema/minimal-contract.md +++ b/docs/schema/minimal-contract.md @@ -1,65 +1,13 @@ # Minimal Contract & FLUID at a Glance -Two views of the contract for v0.7.4: the smallest file that validates, and a one-screen map of every top-level block. - -## FLUID at a Glance - -A FLUID contract is one YAML file. The shape below shows **every top-level block** in v0.7.4, with one-line meaning and the version each was introduced. Required blocks are marked `[req]`; everything else is opt-in. - -```yaml -fluidVersion: "0.7.4" # [req] which contract version this file targets -kind: DataProduct # [req] DataProduct | MLPipeline -id: domain.layer.name # [req] globally unique product id -name: "Human-readable name" # [req] display name -description: "..." # business-facing summary -domain: "Finance" # owning business domain -tags: [pii, gold-layer] # free-text categorization -labels: { team: analytics } # key/value categorization - -metadata: # [req] only metadata.owner is required - owner: { team, email } # [req] team is the one truly required field - layer: Gold # free-form; convention: Bronze | Silver | Gold - -consumes: [ ... ] # upstream FLUID products you depend on -exposes: [ ... ] # [req] ports you publish (each requires exposeId+kind+contract+binding) - └── contract # [req] schema columns (or openapiRef), dq rules - └── semantics # ⭐ 0.7.2 — entities, measures, dimensions, metrics - └── policy # authn, authz, privacy, classification, agentPolicy (⭐ 0.7.1) - └── mcp # ⭐ 0.7.4 — opt this expose into the MCP output-port gateway - └── binding # [req] where it lives (platform + format + location, + icebergConfig*) - -build: # how the product is produced - pattern: hybrid-reference # | embedded-logic | multi-stage | acquisition (⭐ 0.7.3) - engine: dbt | sql | python | spark | glue | custom # transformation engines - | duckdb | airbyte | meltano | dlt | kafka-connect | debezium # ⭐ 0.7.3 acquisition - properties: { ... } # pattern-specific (⭐ 0.7.3: acquisitionPattern) - -orchestration: { engine: airflow | dagster | prefect | kubeflow | custom | none, tasks: [...] } # ⭐ 0.7.0+ -sovereignty: { jurisdiction, allowedRegions, deniedRegions, enforcementMode, … } # ⭐ 0.7.1 -accessPolicy: { grants: [{ principal, permissions, resources, conditions }] } # ⭐ 0.7.1 -retention: { runState, runLogs, lineage, dlq } # ⭐ 0.7.3 -lineage: { upstream, downstream, fieldLevel } -schemaEvolution: { strategy, compatibility } -machineLearning: { ... } -environments: { dev: {...}, staging: {...}, prod: {...} } -lifecycle: { state: preview | active | deprecated | retired } -docs: { ... } -``` - -::: tip `agentPolicy` location -AI/LLM consumption policy lives **per-expose** under `exposes[].policy.agentPolicy` — not at the root. (Earlier release notes show it at the top level; the schema has never accepted it there.) As of 0.7.4, when an expose also carries an `mcp` block, that `agentPolicy` is enforced at runtime by the Fluid MCP gateway on every read. See [Anatomy §7](/fluid/schema/anatomy#_7-governance-sovereignty-accesspolicy-and-exposes-policy-agentpolicy) for the correct shape. -::: - -For a one-line-per-field reference with required/optional flags, see the [**Cheatsheet**](/fluid/schema/cheatsheet). For a tour of each block with deep-dive links, see the [**Anatomy**](/fluid/schema/anatomy). - ---- +Two views of a contract on the latest stable schema, **0.7.5**: the smallest file that validates, and a one-screen map of every top-level block. ## Minimal Valid Contract -The smallest file that passes JSON Schema validation against v0.7.4 — every other top-level block is opt-in: +The smallest file that passes JSON Schema validation against 0.7.5 — every other block is opt-in: ```yaml -fluidVersion: "0.7.4" +fluidVersion: "0.7.5" kind: DataProduct id: demo.bronze.hello_world name: "Hello World" @@ -77,6 +25,68 @@ exposes: location: { path: "./hello.parquet" } ``` -**What you can drop:** `description`, `domain`, `tags`, `labels`, `consumes`, `build`, `orchestration`, all governance blocks (`agentPolicy`, `sovereignty`, `accessPolicy`, `retention`), `exposes[].mcp`, `lineage`, `lifecycle`, `environments`, `docs`. Everything else is layered on as you need it. +It validates against `fluid-schema-0.7.5.json` with any JSON Schema Draft 2020-12 validator, and with the reference implementation (`fluid validate` prints `✅ Valid FLUID contract (schema v0.7.5)`). + +**What is required:** the six top-level keys above; inside `metadata`, only `owner` (the schema requires no member of `owner`, but name a `team`); and inside each expose, `exposeId`, `kind`, `contract` (with `schema` or `openapiRef`) and `binding` (with `platform`, `format` and `location`). See the [**Examples**](/fluid/examples/) for the step-by-step progression from this minimal file to a production source-aligned acquisition product. + +--- + +## FLUID at a Glance + +Every top-level block in 0.7.5, with the version each first appeared in (in its current form). `[req]` marks required blocks. This is a **map, not a contract** — the placeholders are not valid values. + +```text +fluidVersion: "0.7.5" [req] the schema version this file declares +kind: DataProduct [req] DataProduct | MLPipeline +id: domain.layer.name [req] globally unique product id +name: "Human-readable name" [req] display name +description, domain business-facing summary, owning domain +tags: [pii, gold-layer] lower-case categorization tags +labels: { team: analytics } key/value labels + +metadata: [req] only metadata.owner is required + owner: { team, email, slack, oncall } [req] owner itself; none of its members + layer, productType, classification, businessContext, … + +exposes: [ ... ] [req] ports you publish; each requires exposeId, kind, contract, binding + ├── contract [req] schema columns (or openapiRef), dq rules, schemaPolicy + ├── policy authn, authz, privacy, classification, agentPolicy (0.7.1) + ├── semantics entities, measures, dimensions, metrics (0.7.2) + ├── qos availability, freshnessSLO, latencyP95, … + ├── mcp serve the port to AI agents over MCP (0.7.4) + └── binding [req] platform + format + location (+ icebergConfig, vectorConfig, governance) + +consumes: [ ... ] upstream FLUID products you read +build / builds[] how the product is produced + pattern: hybrid-reference | embedded-logic | multi-stage | acquisition (0.7.3) + engine: dbt | dbt- | sql | python | spark | glue | custom + | duckdb | airbyte | meltano | dlt | kafka-connect | debezium (0.7.3) + properties: { ... } shape chosen by pattern + +orchestration { engine, mode, airflow, … } (0.7.2, top level) +sovereignty { jurisdiction, allowedRegions, deniedRegions, … } (0.7.1) +accessPolicy { grants: [{ principal, permissions, resources, conditions }] } (0.7.1) +governance { lakeFormation: { admins, tagDefinitions } } (0.7.3) +retention { runState, runLogs, lineage, dlq } (0.7.3) +extensions { ... } (0.7.3) +lineage { granularity, upstream, downstream } +schemaEvolution { strategy, compatibility, changePolicy } +machineLearning { enabled, framework, models } +environments { dev: {...}, staging: {...}, prod: {...} } +lifecycle { state, retention, deprecationPolicy } +docs { homepage, runbook, dictionary, changeLog } + +0.7.6 preview only (fluidVersion "0.7.6"): +packaging { mode, pool, containers } +consumers [ { name, type, ... } ] +``` + +Field-level lineage is `lineage.granularity: field_level` with `fieldMappings` on each `upstream[]` entry; there is no `fieldLevel` member. + +::: tip `agentPolicy` location +AI/LLM consumption policy lives **per-expose** under `exposes[].policy.agentPolicy` — not at the root; the schema has never accepted it there. See [Anatomy §7](/fluid/schema/anatomy#_7-governance-sovereignty-accesspolicy-governance-and-exposes-policy-agentpolicy) for the shape. +::: + +For a one-line-per-field reference with required flags, see the [**Cheatsheet**](/fluid/schema/cheatsheet). For a tour of each block, see the [**Anatomy**](/fluid/schema/anatomy). diff --git a/docs/schema/specification.md b/docs/schema/specification.md index 85d62c4..54d7835 100644 --- a/docs/schema/specification.md +++ b/docs/schema/specification.md @@ -1,88 +1,29 @@ -# Fluid Protocol Specification +# FLUID Specification -::: tip Latest schema version -The latest published JSON Schema is **0.7.5**. This page is the narrative protocol specification; for the version-specific, field-by-field reference see the [**Cheatsheet**](/fluid/schema/cheatsheet), the [**Anatomy**](/fluid/schema/anatomy), and the [**Versions**](/fluid/schema/versions) index (raw JSON Schema + generated HTML per version). +::: tip Versions +**Latest stable: 0.7.5. Preview: 0.7.6.** This page specifies the structure of a FLUID document as defined by the **0.7.5** JSON Schema. For every version's schema and generated field-by-field reference, see [**Versions**](/fluid/schema/versions); for what is new in the preview, see [0.7.6 (preview)](/fluid/releases/0.7.6). ::: -This document provides the complete, official specification for the FLUID (Federated Layered Unified Interchange Definition) protocol. It is intended for data architects, platform engineers, and developers who are building the next generation of data infrastructure, as well as for vendors seeking to make their tools compliant with this open standard. +FLUID (Federated Layered Unified Interchange Definition) is an open, declarative specification for **data products**, written in YAML or JSON and kept in version control. It is not a platform or a single tool: it is a shared language that tools read to build, deploy, govern and serve a data product. -### The Strategic Imperative: A Protocol for the Agentic Era -The contemporary enterprise is shifting from process automation to an Agentic Ecosystem, where autonomous AI agents drive operations with unprecedented speed and intelligence. This paradigm shift, enabled by communication standards like the Model Context Protocol (MCP), exposes a foundational vulnerability in modern data architecture: the lack of a common language for defining, governing, and interacting with data assets. +A FLUID document describes one data product: what it **consumes** (its inputs), what it **exposes** (its output ports, each with a contract and a binding to where the data lives), how it is **built**, and the policies around it — access, sovereignty, AI-agent use, retention. This page is for anyone implementing a FLUID-aware tool or checking that one conforms. -Today's data landscape is a fragmented collection of imperative pipelines, siloed tool configurations, and implicit knowledge. This static, brittle foundation cannot support the dynamic, real-time demands of an agentic workforce. Agents require a data fabric that is not only accessible but also discoverable, trustworthy, and context-aware. - -FLUID is the standard designed to create this fabric. It addresses this challenge by providing a declarative, universal protocol for defining Data Products. It is the missing piece of the puzzle, serving as the foundational layer that makes an organization truly MCP-ready. While MCP standardizes how agents communicate, FLUID standardizes the trustworthy Data Products they communicate with. - -### What is FLUID? -FLUID is an open, declarative specification, written in YAML and managed in version control. It is not a platform or a single tool, but a shared language that enables a decentralized ecosystem of compliant tools to work in concert. - -It re-frames the data lifecycle around the concept of a Data Product: a versioned, autonomous asset with a clearly defined interface, contract, and implementation. By unifying the definition of what a data product consumes (its dependencies), what it exposes (its public interface), and how it is built (its implementation logic), FLUID provides a holistic, auditable, and machine-readable blueprint for every data asset in the enterprise. - -This document details the full specification for this protocol, providing the technical foundation required to build the governable, scalable, and agent-ready data ecosystems of the future. - -> 📖 *FLUID Data Products* is [available on Amazon](https://amzn.eu/d/ikMlWNV). - -🌊 FLUID: Federated Layered Unified Interchange Definition - ---- - -## 🧭 Core Principles - -- **Data as a Product** - Data is a first-class asset with a clear owner, a versioned interface, and a machine-readable contract. - -- **Declarative, Not Imperative** - Contracts define the desired end state of a data product. The FLUID-aware framework is responsible for the implementation. - -- **Contracts as Code** - Governance (schema, quality, build, privacy) is embedded directly into version-controlled files, enabling automated, proactive enforcement. - -- **Federated Ownership** - Data products are owned and managed by the domain teams who know the data best, enabling a true, scalable Data Mesh. +**What defines conformance.** The published JSON Schemas define which documents are valid FLUID documents, and the [conformance corpus](https://github.com/open-data-protocol/fluid/blob/main/tests/README.md) pins that behaviour case by case, so that "FLUID-conformant" can be checked without any particular implementation ([GOVERNANCE.md](https://github.com/open-data-protocol/fluid/blob/main/GOVERNANCE.md)). [`data-product-forge`](/fluid/concepts/forge-cli) is the reference implementation: implementation #1 under test, not the referee. Where this page and a schema disagree, the schema is authoritative and this page is wrong. --- -## 🏗️ Contract Structure: Monolithic vs. Modular - -The FLUID specification is designed for flexibility. -A data product contract can be defined in a single, **monolithic file** or composed from multiple, specialized files. - -### Monolithic Structure (For Simplicity) - -For simple data products owned by a single team, all definitions can be contained within a single root `fluid.yml` file. +## Core principles -``` -/dp-simple-product/ -└── 📄 fluid.yml # All definitions are inline. -``` +- **Data as a product** — data is a first-class asset with an owner, a versioned interface and a machine-readable contract. +- **Declarative, not imperative** — a document states the desired end state; FLUID-aware tools decide how to reach it. +- **Contracts as code** — schema, quality, build and policy live in version-controlled files, so they can be checked automatically. +- **Federated ownership** — data products are owned by the domain teams that know the data. --- -### Modular Structure (For Complexity & Federation) +## Documents and files -For complex, enterprise-grade products with multiple stakeholders, the contract can be broken into logical, linked files. -This is the recommended best practice. - -The root `fluid.yml` acts as a "table of contents," referencing detailed configuration files stored in a dedicated `.fluid/` directory using the `$ref` keyword. - -``` -/dp-complex-product/ -│ -├── 📄 fluid.yml # The main entrypoint, contains high-level identity. -│ -└── 📁 .fluid/ # A dedicated folder for all contract details. - ├── 📄 consumes.yml - ├── 📄 build.yml - ├── 📄 exposes.yml - ├── 📄 schema.yml - └── 📄 quality.yml -``` - -## Preamble - -This document provides the complete, official specification for the FLUID (Federated Layered Unified Interchange Definition) protocol. It is intended for data architects, platform engineers, and developers who are building the next generation of data infrastructure, as well as for vendors seeking to make their tools compliant with this open standard. - ---- +A FLUID document is a single YAML or JSON object. **The file name is not normative.** Pages on this site name files `*.fluid.yml`; the reference implementation scaffolds and looks for `contract.fluid.yaml`. Either is fine. ## Validation semantics @@ -98,7 +39,7 @@ This is the Draft 2020-12 default, and FLUID keeps it for two reasons of its own The first is that conformance must not depend on which validator you run. `format` vocabularies are optional in Draft 2020-12 and implementations differ widely in which formats they recognise and what they pull in to check them. Were FLUID to make `format` assertive, the same document could be conformant in one language and non-conformant in another, which would defeat the purpose of publishing a conformance corpus at all. -The second is that the choice is not symmetric. Declaring `format` assertive would invalidate documents that are valid today — a narrowing, which [GOVERNANCE.md](https://github.com/open-data-protocol/fluid/blob/main/GOVERNANCE.md) forbids between versions and `scripts/check-compat.py` enforces on every pull request. Annotation-only is therefore the only reading available to a pre-1.0 specification that has already published twelve schema versions. A future version may add assertive checking behind a new, opt-in keyword; it may not retroactively sharpen this one. +The second is that the choice is not symmetric. Declaring `format` assertive would invalidate documents that are valid today — a narrowing, which [GOVERNANCE.md](https://github.com/open-data-protocol/fluid/blob/main/GOVERNANCE.md) forbids between versions and `scripts/check-compat.py` enforces on every pull request. Annotation-only is therefore the only reading available to a pre-1.0 specification that has already published many schema versions. A future version may add assertive checking behind a new, opt-in keyword; it may not retroactively sharpen this one. Implementations that do want to assert formats are served by the conformance corpus rather than left to guess: the cases that depend on assertion live in [`tests/optional/`](https://github.com/open-data-protocol/fluid/tree/main/tests/optional), separated from the core tier for exactly this reason, and `conformance/run.py` runs them under a format-asserting validator. @@ -108,225 +49,188 @@ FLUID defines a **field** named `format` in several places — `exposes[].bindin The collision of names is unfortunate and is called out here because it is easy to read "`format` is an annotation" as applying to them. It does not. The sentence above is about the JSON Schema *keyword*; this paragraph is about a FLUID *field* that happens to share its spelling. ---- - -## 1. Specification Root - -The FLUID definition is a YAML or JSON file (`.fluid.yml` OR `.fluid.json`) with the following root-level objects. - -| Key | Type | Required | Description | -|----------------|----------------|----------|-----------------------------------------------------------------------------| -| fluidVersion | String | Yes | The version of the FLUID specification this file adheres to (e.g., 0.7.4). | -| kind | String | Yes | The type of data product definition. See section 1.1. | -| id | String | Yes | A globally unique, versioned id for the data product (customer360_v1). | -| name | String | Yes | A human-readable name for the data product (e.g., "Customer 360"). | -| description | String | Yes | A brief description of the product's purpose. | -| domain | String | Yes | The business domain that owns this product (e.g., "Marketing"). | -| metadata | Object | Yes | Identification, ownership, and classification information. See section 1.2. | -| consumes | Object / List | Yes | Defines the input data sources needed to build the product. See section 1.4.| -| build | Object | Yes | Contains the implementation logic for how the product is built. See 1.5. | -| exposes | Object / List | Yes | Defines the public output interface(s) of the product. See section 1.3. | -| accessPolicy | Object | No | Defines the static access control policies for the data product. See 1.6. | -| dynamicPolicies| Object | No | Defines context-aware access policies that adapt at runtime. See 1.7. | -| operations | Object | No | Defines SLAs, lifecycle, and observability characteristics. See 1.8. | -| extensions | Object | No | Registers required external plugins. See section 1.9. | - ---- - -### 1.1 kind Enumeration - -The `kind` key specifies the nature of the data product. - -- **DataProduct**: A standard, materialized data asset. -- **VirtualDataProduct**: A product that exists only as a logical view or query, without its own physical storage. -- **EgressFlow**: A product specifically designed to export data securely to an external system. - ---- - -### 1.2 metadata Block - -| Key | Type | Required | Description | -|-------------------|----------------|----------|-----------------------------------------------------------------------------| -| owner | Object | Yes | Ownership details (e.g., { team: 'finance', email: 'finance@company.com' }).| -| layer | String | No | The architectural layer (e.g., "Gold", "Silver", "Bronze") | -| status | String | No | The lifecycle status (e.g., "Published", "Development", "Deprecated"). | -| sensitivity_level | String | No | The overall data classification (e.g., "Confidential", "Public"). | -| cost_center | String | No | An identifier for automated FinOps cost attribution. | -| tags | Map[String] | No | Key-value pairs for categorization (e.g., layer: gold, domain: finance). | -| classification | String | Yes | Default privacy level: public, internal, confidential, restricted. | -| purpose | Object | No | Detailed, machine-readable description of the business purpose. | -| version | String | No | (Recommended) Semantic version of this data product definition (e.g., 1.0.0)| - -#### 1.2.1 purpose Block - -| Key | Type | Required | Description | -|--------------------|----------------|----------|-----------------------------------------------------------------------------| -| business_purpose | String | Yes | A clear statement of why this data product exists. | -| use_cases | List | No | A list of specific business use cases it supports. | -| target_group | List | No | The intended audience for this product. | -| limitations | String | No | Any known limitations or constraints of the data. | +### The declared `fluidVersion` selects the schema ---- - -### 1.3 exposes Block (The Output Port) - -Defines the public interface of the data product. This is what consumers interact with. - -| Key | Type | Required | Description | -|----------|--------|----------|--------------------------------------------------------------------| -| name | String | If list | A unique name for this output port within the product. | -| location | Object | Yes | The physical or virtual location where the data product is materialized. See 1.3.1. | -| contract | Object | Yes | The schema, quality, and privacy promises for this output. See 1.3.2. | - -#### 1.3.1 location Object - -| Key | Type | Required | Description | -|------------|--------|--------------|--------------------------------------------------------------------| -| type | String | Yes | bigquery, s3, snowflake, gcs, iceberg, api, kafka, or virtual. | -| connection | String | Yes | Reference to a secret in a vault (e.g., secret:gcp-prod-dwh-key). | -| format | Object | If not virtual| Describes the data format (e.g., { type: 'parquet' }). | -| properties | Object | Yes | Technology-specific properties (e.g., project, dataset, table). | - -#### 1.3.2 contract Object - -| Key | Type | Required | Description | -|-------------|----------|--------------|--------------------------------------------------------------------| -| inheritFrom | String | No | dbt, fluid-product, or openApi. Populates the contract from a source.| -| model / spec| String | If inheriting| The dbt model name or path to the OpenAPI spec file/URL. | -| schema | Object | Yes | Defines the columns and data types. See 1.3.2.1. | -| quality | List | No | List of data quality rules to enforce. See 1.3.2.2. | -| privacy | List | No | List of privacy classifications and treatments. See 1.3.2.3. | -| semantics | Object | No | Adds machine-readable meaning to the data. See 1.3.2.4. | - -##### 1.3.2.1 schema.columns Array - -| Key | Type | Required | Description | -|---------|---------|----------|----------------------------------------------| -| name | String | Yes | Column name. | -| type | String | Yes | Data type (STRING, INT64, NUMERIC, TIMESTAMP, JSON, BOOLEAN, DATE). | -| nullable| Boolean | No | true by default. | - -##### 1.3.2.2 quality Array Item - -| Key | Type | Required | Description | -|----------|--------|--------------|--------------------------------------------------------------------| -| rule | String | Yes | not_null, unique, regex_match, in_set, or a custom SQL expression. | -| columns | List | If applicable| Column(s) to apply the rule to. | -| pattern/set| String/List|If applicable| Parameters for regex_match or in_set. | -| onFailure| Object | Yes | action (reject_row, quarantine_row, fail_pipeline, alert) and optional notifications. | - -##### 1.3.2.3 privacy Array Item - -| Key | Type | Required | Description | -|---------------|--------|--------------|--------------------------------------------------------------------| -| classification| String | No | PII, SPI, Confidential. Overrides metadata.classification. | -| columns | List | Yes | Column(s) to apply the treatment to. Can be ['*']. | -| treatment | Object | Yes | type (hashing, masking, encryption, tokenization) and properties. | - -##### 1.3.2.4 semantics Object - -| Key | Type | Description | -|---------------|--------|--------------------------------------------------------------------| -| ontology | String | Reference to an external ontology (e.g., URL to an OWL or RDF file).| -| classifications| List | Maps columns to terms in a business glossary or formal ontology. | +Because the schema is chosen by the document's own `fluidVersion`, a later schema accepting an earlier `fluidVersion` value is not a verdict on documents that declare that earlier version. The 0.7.5 schema's `fluidVersion` enum lists `"0.7.3"`, `"0.7.4"` and `"0.7.5"`; a document declaring `"0.7.4"` is still validated against the 0.7.4 schema, so it cannot use a field that only 0.7.5 defines. See [Choosing `fluidVersion`](/fluid/schema/versions#choosing-fluidversion). ---- - -### 1.4 consumes Block (The Input Port) - -| Key | Type | Required | Description | -|-----------------|--------|--------------|--------------------------------------------------------------------| -| type | String | Yes | gcs, kafka, s3, api, postgres-cdc, sftp, or fluid-product. | -| name | String | If type: fluid-product | The dataProduct name of the upstream FLUID definition. | -| alias | String | If list | A local alias to refer to this source in the build block. | -| onUpstreamChange| String | No | Action on upstream contract change: fail, alert, triggerRebuild. | -| connection | String | If physical type | Reference to a secret in a vault. | -| format | Object | If physical type | Describes the data format (e.g., { type: 'json' }). | -| properties | Object | If physical type | Technology-specific properties (e.g., Kafka topic, API endpoint). | - ---- - -### 1.5 build Block (The Implementation) +### Known interoperability issue: one `pattern` is not an ECMA-262 regular expression -| Key | Type | Required | Description | -|-----------------|--------|--------------|---------------------------------------------------------------------------------| -| engine | String | Yes | The engine to use: sql, python, dbt, dbt-cloud, spark-sql. | -| script | String | Yes | The specific asset to execute (e.g., a script path or dbt model selector). | -| trigger | Object | Yes | Defines how the build is initiated (schedule, event, manual). | -| runtime | Object | Yes | The underlying compute platform where the build will run (airflow, gcp-cloud-run). | -| dependencies | Object | No | Defines the execution dependencies on other data products. | -| retries | Object | No | Configuration for handling transient failures: count, delay. | -| notifications | Object | No | Defines how to send alerts (channel, target) on onSuccess or onFailure. | +Draft 2020-12 says a `pattern` SHOULD be a valid ECMA-262 regular expression. In the 0.7.2 to 0.7.6 schemas, the second alternative of the column `type` (`$defs/column/properties/type`) begins with the inline flag `(?i)`, which ECMA-262 does not have: JavaScript's `RegExp` rejects it as an invalid group. Python's `re`, which the conformance tooling in this repository uses, accepts it and matches case-insensitively. A validator that compiles patterns as ECMA-262 may therefore refuse that pattern, or the schema, where a Python-based validator does not. This is recorded here rather than resolved: changing the pattern would be a schema change, and schemas are changed upstream (see [CONTRIBUTING.md](https://github.com/open-data-protocol/fluid/blob/main/CONTRIBUTING.md)). --- -### 1.6 accessPolicy (Static Access) - -| Key | Type | Required | Description | -|------------|--------|----------|--------------------------------------------------------------------| -| visibility | String | No | Default discoverability: private, internal. | -| grants | List | Yes | A list of static, explicit access grants. | - -#### 1.6.1 grants Array Item -| Key | Type | Required | Description | -|------------|--------|----------|--------------------------------------------------------------------| -| principal | String | Yes | The actor receiving the grant. Format: `user:`, `group:`, `agent:`. | -| permissions| List | Yes | Rights granted: readData, readMetadata, manage. | -| scope | Object | No | Fine-grained access. See 1.6.2. | +## Composing a document from several files -#### 1.6.2 scope Object +FLUID defines **one document**. The standard does not define how a document may be assembled from several files; any such mechanism is **implementation-defined**, and nothing in this section is normative. -| Key | Type | Required | Description | -|-------------|--------|----------|--------------------------------------------------------------------| -| columns | List | No | Allow-list of columns the principal can view. | -| rowFilter | String | No | SQL WHERE clause for secure view. | -| privacyView | String | No | treated (default, views post-privacy data), cleartext (views pre-privacy data). | +What follows from the rest of this specification: ---- +1. **A composed root is not a FLUID document.** The published schemas have no `$ref` property in a document, and the root object and most nested objects are closed (`additionalProperties: false`), so a file that stands in `{ $ref: ... }` for a block fails validation. Splitting a valid contract into fragments with the reference implementation's `fluid split` and validating the root file directly against the 0.7.5 schema gives errors such as `Additional properties are not allowed ('$ref' was unexpected)`. +2. **Validate the resolved document.** Conformance is a property of the single document a composition mechanism produces. Resolve first, then validate that result against the schema of its `fluidVersion`. -### 1.7 dynamicPolicies (Adaptive Access) +### How the reference implementation composes contracts (non-normative) -| Key | Type | Required | Description | -|-------|------|----------|--------------------------------------------------------------------| -| rules | List | Yes | A list of contextual access rules, evaluated in order. | +`data-product-forge` resolves `$ref` nodes before it validates, plans or applies a contract, so the rest of its pipeline sees one document. As of data-product-forge 0.18: -#### 1.7.1 rules Array Item +- A `$ref` node is an object whose only key is `$ref`; its value is a path to a YAML or JSON file, resolved relative to the file that contains it, optionally followed by `#` and an [RFC 6901](https://www.rfc-editor.org/rfc/rfc6901) JSON Pointer. A referenced YAML file must hold an object, not a list. Same-document refs (`#/...`) are left as written; cycles and nesting deeper than 20 levels are errors. +- **Refs are confined to the root contract's directory tree.** A ref that resolves outside it — through `..` or a symlink — is refused, as are absolute paths and URLs (`https://`, `file://`, `s3://`, …). Setting `FLUID_REF_ROOT` (or passing `ref_root=` to the loader) widens the root to a directory that contains the contract, for monorepos that share fragments; a `FLUID_REF_ROOT` that does not contain the contract is ignored with a `ref_root_env_ignored` warning. +- `fluid split` turns a single-file contract into a root plus a `fragments/` directory (`fragments/exposes/.yaml`, `fragments/builds/.yaml`, `fragments/sovereignty.yaml`, `fragments/access-policy.yaml`); `fluid bundle` resolves a root back into one document. +- Its `.fluid/` directories hold the tool's own runtime state (receipts, run records, staging data), which its generated `.gitignore` partly excludes from version control. They are not a place for contract fragments. -| Key | Type | Required | Description | -|-----------|--------|----------|--------------------------------------------------------------------| -| name | String | Yes | A descriptive name for the policy rule. | -| condition | String | Yes | An expression evaluated against the agent's context. Supports interpolation. | -| grant | Object | Yes | The permissions and scope to grant if the condition is met. | +Details: [Composing a contract from fragments](https://agenticstiger.github.io/forge_docs/concepts/contract-refs.html), [`fluid split`](https://agenticstiger.github.io/forge_docs/cli/split.html) and [`fluid bundle`](https://agenticstiger.github.io/forge_docs/cli/bundle.html) in the reference implementation's documentation. --- -### 1.8 operations Block - -| Key | Type | Required | Description | -|-------------|--------|----------|--------------------------------------------------------------------| -| sla | Object | No | Defines the Service Level Agreements for this product. See 1.8.1. | -| lifecycle | Object | No | Manages data retention and archival policies. See 1.8.2. | -| observability| Object| No | Configures logging, alerting, and monitoring. See 1.8.3. | - -#### 1.8.1 sla Object - -Defines cost, latency, freshness, accuracy, sustainability, and feedbackSignals. - -#### 1.8.2 lifecycle Object - -Defines retention (period, condition) and archival (trigger, destination). - -#### 1.8.3 observability Object - -Defines logging (level, destination) and alerting (onFailure notifications). +## 1. Document structure (0.7.5) + +The tables summarise the 0.7.5 schema. They name every top-level member and the main nested blocks; the [generated reference](/fluid/specs/0.7.5/fluid-spec.html) has every field. **Closed** means `additionalProperties: false`: a member the table does not list makes the document invalid. + +### 1.1 Root + +The root object is **closed**. + +| Member | Type | Required | Meaning | +|---|---|:-:|---| +| `fluidVersion` | string, one of `"0.7.3"`, `"0.7.4"`, `"0.7.5"` | ✅ | The schema version the document declares. It selects the schema the document is validated against. | +| `kind` | `DataProduct` \| `MLPipeline` | ✅ | The kind of product. | +| `id` | identifier | ✅ | Globally unique product id, e.g. `finance.gold.customer_360`. | +| `name` | string | ✅ | Display name. | +| `metadata` | object (§1.2) | ✅ | Ownership and classification. | +| `exposes` | array of expose (§1.3) | ✅ | The product's output ports. | +| `description`, `domain` | string | | Business description and owning domain. | +| `tags` | array of unique tags | | Each tag lower-case: `^[a-z0-9][a-z0-9-]*[a-z0-9]$` or a single character. | +| `labels` | map of string → string | | Key/value labels. | +| `consumes` | array of consume (§1.4) | | Upstream data products this product reads. | +| `build` | build (§1.5) | | How the product is built. | +| `builds` | array of build (§1.5) | | Several builds for one product. | +| `orchestration` | object (§1.6) | | Scheduling and workflow-engine settings. | +| `accessPolicy` | object (§1.7) | | Access grants for the product. | +| `sovereignty` | object (§1.7) | | Jurisdiction and data-residency constraints. | +| `governance` | object (§1.7) | | Account-wide AWS Lake Formation settings. | +| `retention` | object | | ISO-8601 durations for operational records: `runState`, `runLogs`, `lineage`, `dlq`. | +| `lifecycle` | object | | `state` (`preview` \| `active` \| `deprecated` \| `retired`), `retention`, `deprecationPolicy`. | +| `lineage` | object | | `granularity` (`table_level` \| `field_level`), `upstream[]`, `downstream[]`. | +| `schemaEvolution` | object | | `strategy` (`semantic_versioning` \| `date_based` \| `sequential`), `compatibility` (`backward_compatible` \| `forward_compatible` \| `full_compatible` \| `breaking`), `changePolicy`. | +| `machineLearning` | object | | `enabled`, `framework`, `models[]`. | +| `environments` | map of name → environment | | Per-environment overrides of `metadata` and `exposes`. | +| `docs` | object | | `homepage`, `runbook`, `dictionary`, `changeLog`. | +| `extensions` | object (open) | | Vendor- or plugin-namespaced configuration. | + +An **identifier** matches `^[A-Za-z0-9_][A-Za-z0-9_.-]*[A-Za-z0-9_]$` (or is a single letter, digit or underscore). A **duration** is an ISO-8601 duration such as `P30D` or `PT15M`. + +### 1.2 `metadata` + +**Closed.** `owner` is required; no member of `owner` is. + +| Member | Type | Meaning | +|---|---|---| +| `owner` (required) | object, closed: `team`, `email`, `slack`, `oncall` | Who owns the product. `email` carries `format: email`, which is an annotation (see [Validation semantics](#validation-semantics)). | +| `layer` | string | Free-form layer label; `Bronze` / `Silver` / `Gold` is a convention, not an enum. | +| `productType` | `SDP` \| `ADP` \| `CDP` | Source-aligned, aggregated or consumption-aligned data product. | +| `classification` | `public` \| `internal` \| `confidential` \| `restricted` | Default classification. Optional. | +| `businessContext` | object: `domain`, `subdomain`, `businessCapability`, `valueStream` | Business context. | +| `experimental` | array of unique strings | Experimental features the document opts into. | +| `createdAt` | string, `format: date-time` | Creation time. | +| `provenance` | object | Generation envelope written by tooling (tool, version, command, time). | +| `tags` | array of tags | | + +### 1.3 `exposes[]` — output ports + +Each expose is **closed** and requires `exposeId`, `kind`, `contract` and `binding`. + +| Member | Type | Meaning | +|---|---|---| +| `exposeId` (required) | identifier | Id of the port, unique within the product. | +| `kind` (required) | `table` \| `view` \| `api` \| `file` \| `stream` \| `topic` \| `feature_store` \| `model` \| `vector` \| `graph` \| `time_series` \| `other` | What the port is. | +| `contract` (required) | object, closed | The data's shape and promises. Must contain `schema` or `openapiRef` (or both). | +| `binding` (required) | object, closed | Where the data lives (§1.3.2). | +| `policy` | object, closed | `authn`, `authz` (`readers`, `writers`, `columnRestrictions`), `privacy` (`masking[]`, `rowLevelPolicy`), `classification`, `agentPolicy`. | +| `semantics` | object, closed | Business meaning: `entities`, `measures`, `dimensions`, `metrics`. | +| `qos` | object, closed | `availability`, `freshnessSLO`, `dataLossSLO`, `latencyP95`, `completenessTarget`, `errorBudget`. | +| `mcp` | object, closed | `sampling.maxRows`, `classification.dataClass` for serving the port to AI agents over MCP. | +| `lifecycle`, `observability`, `docs` | object | Per-port lifecycle, observability and documentation. | +| `title`, `description`, `version` | string (`version` is semver) | | +| `crawler`, `iceberg` | object | AWS Glue crawler and Iceberg table-maintenance settings. | +| `tags`, `labels` | | | + +#### 1.3.1 `contract` + +| Member | Type | Meaning | +|---|---|---| +| `schema` | array of column | The columns. | +| `openapiRef` | string | Reference to an OpenAPI document, for `kind: api`. | +| `dq` | object: `rules[]`, `monitoring` | Data-quality rules. Each rule is closed and requires `id`, `type` (`freshness` \| `completeness` \| `uniqueness` \| `valid_values` \| `accuracy` \| `schema` \| `anomaly_detection` \| `drift_detection`) and `severity` (`info` \| `warn` \| `error` \| `critical`). | +| `schemaPolicy` | `strict` \| `discover_and_freeze` \| `evolve_safe` \| `evolve_all` | How the output schema may change. | +| `schemaSignature` | `sha256:` + 64 hex | A digest of the schema. | +| `guarantees`, `quality` | object, array | Compatibility promises and additional quality checks. | + +A **column** is closed and requires `name` and `type`. `type` is a type name such as `string`, `int64`, `numeric`, `timestamp` or `json` from a fixed list, matched case-insensitively and optionally with parameters (`VARCHAR(255)`, `DECIMAL(10,2)`). Other members: `required` (boolean), `description`, `sensitivity` (`none` \| `internal` \| `confidential` \| `restricted` \| `pii` \| `phi` \| `cleartext` \| `treated` \| `anonymized` \| `pseudonymized` \| `tokenized` \| `encrypted`), `semanticType`, `businessName`, `businessDefinition`, `validationRules`, `tags`, `labels`. + +#### 1.3.2 `binding` + +Closed; requires `platform`, `format` and `location`. + +| Member | Type | Meaning | +|---|---|---| +| `platform` (required) | `gcp` \| `aws` \| `azure` \| `snowflake` \| `databricks` \| `kafka` \| `confluent` \| `local` \| `kubernetes` \| `postgres` \| `pgvector` \| `other` | The platform. | +| `format` (required) | `bigquery_table`, `snowflake_table`, `snowflake_view`, `gcs_file`, `s3_file`, `http_api`, `grpc_api`, `pubsub_topic`, `kafka_topic`, `delta_table`, `iceberg`, `parquet`, `csv`, `json`, `redshift_table`, `redshift_serverless`, `redshift_external_schema`, `postgres_table`, `athena_table`, `glue_table`, `pgvector_table`, `other` | The physical format. | +| `location` (required) | object, closed | Platform-specific address. Members: `account`, `project`, `dataset`, `database`, `schema`, `table`, `bucket`, `path`, `gateway`, `baseUrl`, `topic`, `subscription`, `region`, `zone`, `environment_id`, `kafka_cluster_id`, `confluent_role_arn`, `stream`, `namespace`, `workgroup`, `iam_role_arn`, `external_schema`, `glue_database`, `catalog`, `warehouse`, `uri`, `partitionBy`. None is required by the schema. | +| `icebergConfig` | object | Iceberg table settings (write version, file format, partition spec, sort order). | +| `vectorConfig` | object, closed; `dimensions` required | Vector / embeddings output-port settings. | +| `governance` | object: `lakeFormation` | Per-resource AWS Lake Formation settings: `registerLocation`, `grants[]`, `tags`, `rowFilter`. | +| `properties` | object (open) | Platform-specific extra properties. | +| `tags`, `labels` | | | + +### 1.4 `consumes[]` — input ports + +Each entry is **closed** and requires `productId` and `exposeId`: the upstream product and the port it reads. Optional: `versionConstraint` (a semver range such as `^2.0.0`), `qosExpectations` (`freshnessMax`, `maxStaleness`, `minCompleteness`), `requiredPolicies[]`, `purpose`, `tags`, `labels`. + +### 1.5 `build` and `builds[]` + +A build is **closed**; no member is required. + +| Member | Type | Meaning | +|---|---|---| +| `id` | identifier | Build id (useful when there are several). | +| `pattern` | `hybrid-reference` \| `embedded-logic` \| `multi-stage` \| `acquisition` | Selects the shape of `properties`. | +| `engine` | `dbt`, `sql`, `python`, `spark`, `glue`, `custom`, `duckdb`, `airbyte`, `meltano`, `dlt`, `kafka-connect`, `debezium`, or `dbt-` | The engine. | +| `properties` | object | The pattern's settings (below). | +| `execution` | object, closed | `trigger`, `runtime`, `retries`, `notifications[]`, `orchestration`. | +| `capabilities` | array | What the build asks of its runner, e.g. `incremental_dedup`, `cdc`, `streaming`, `exactly_once`. | +| `repository`, `description` | string | | +| `outputs` | array of identifier | The `exposeId`s the build produces. | +| `dependencies`, `transformations` | array | | + +`properties` is validated according to `pattern`: + +| `pattern` | `properties` shape | Required | +|---|---|---| +| `hybrid-reference` | `model`, `target`, `select`, `models`, `vars`, `materializations` — a reference to a model kept elsewhere, such as a dbt project | `model` | +| `embedded-logic` | `sql`, `language` (`sql` \| `flink_sql` \| `pyspark` \| `scala` \| `python` \| `r`), `parameters` | `sql` | +| `multi-stage` | `stages[]`, `orchestration` | — | +| `acquisition` | `source` (`kind`, `mode` required), `sink`, `delivery`, `schemaEvolution`, `preLand`, `quality`, `cost`, `catalog`, `concurrency`, `lineage`, and one settings key per engine: `duckdb`, `airbyte`, `meltano`, `dlt`, `kafka-connect`, `debezium` | `source` | + +### 1.6 `orchestration` + +Requires `engine` (`airflow` \| `dagster` \| `prefect` \| `kubeflow` \| `custom` \| `none`). Defined members: `mode` (`generated` \| `manual` \| `hybrid`), `generateOnChange`, and per-engine settings `airflow` (which requires `dagId`, and holds `tasks[]`), `dagster`, `prefect`. + +The object is **open**: it does not set `additionalProperties`, so members it does not define — such as an `orchestration.tasks` list — are accepted **without being checked**. The only task shape the schema checks is `orchestration.airflow.tasks[]` (and the same under `build.execution.orchestration`). See the [Cheatsheet](/fluid/schema/cheatsheet#orchestration-fields). + +### 1.7 Access and governance + +| Block | Shape | +|---|---| +| `accessPolicy` | Closed; `grants[]`, each closed with `principal` (required), `permissions` (`read`, `select`, `query`, `write`, `insert`, `update`, `delete`, `create`, `admin`, `manage`), `resources` (array of strings, e.g. JSONPath selecting exposes) and `conditions` (open object). | +| `sovereignty` | Closed: `jurisdiction`, `allowedRegions`, `deniedRegions`, `dataResidency` (boolean), `crossBorderTransfer` (boolean), `transferMechanisms`, `regulatoryFramework`, `enforcementMode` (`strict` \| `advisory` \| `audit`), `validationRequired`. | +| `exposes[].policy.agentPolicy` | Closed: `allowedModels`, `deniedModels`, `maxTokensPerRequest`, `maxTokensPerDay`, `allowedUseCases` / `deniedUseCases` (from `inference`, `reasoning`, `analysis`, `summarization`, `classification`, `embedding`, `search`, `qa`, `code_generation`, `fine_tuning`, `training`, `rag`), `canReason`, `canStore`, `retentionPolicy`, `auditRequired`, `purposeLimitation`. It exists only per expose, not at the root. | +| `governance.lakeFormation` | Closed: `admins` (IAM ARNs) and `tagDefinitions` (tag key → allowed values). `admins` is **authoritative**: applying it replaces the account's Lake Formation admin list. The schema's description adds that applying it also clears the account's create-database and create-table default permissions, trusted resource owners and parameters, and that destroying the emitted `aws_lakeformation_data_lake_settings` resource empties the admin list and resets `CROSS_ACCOUNT_VERSION` to 1. | --- -### 1.9 extensions Block +## Further reading -| Key | Type | Description | -|----------------------|------|--------------------------------------------------------------------| -| customTransformations| List | A list of custom transformation engines the build requires. | -| policyEngines | List | A list of external policy engines (e.g., OPA) needed to evaluate policies. | -| observabilityHooks | List | A list of custom hooks to send metrics and traces to external systems. | +- [Anatomy](/fluid/schema/anatomy) — a guided tour of the blocks, with examples. +- [Cheatsheet](/fluid/schema/cheatsheet) — one row per field. +- [Changelog](/fluid/schema/changelog) — what changed in each version. +- *FLUID Data Products* (book) is [available on Amazon](https://amzn.eu/d/ikMlWNV). diff --git a/docs/schema/versions.md b/docs/schema/versions.md index feeb138..6658993 100644 --- a/docs/schema/versions.md +++ b/docs/schema/versions.md @@ -1,20 +1,69 @@ # Schema Versions -Every published version of the FLUID JSON Schema. Point your validator at a version's JSON Schema; read its generated HTML reference for the field-by-field detail. The latest version is **0.7.5**. - -| Version | JSON Schema | HTML Reference | -|---|---|---| -| **0.7.5** **(latest)** | [`fluid-schema-0.7.5.json`](/fluid/schema/fluid-schema-0.7.5.json) | [`0.7.5/fluid-spec.html`](/fluid/specs/0.7.5/fluid-spec.html) | -| 0.7.4 | [`fluid-schema-0.7.4.json`](/fluid/schema/fluid-schema-0.7.4.json) | [`0.7.4/fluid-spec.html`](/fluid/specs/0.7.4/fluid-spec.html) | -| 0.7.3 | [`fluid-schema-0.7.3.json`](/fluid/schema/fluid-schema-0.7.3.json) | [`0.7.3/fluid-spec.html`](/fluid/specs/0.7.3/fluid-spec.html) | -| 0.7.2 | [`fluid-schema-0.7.2.json`](/fluid/schema/fluid-schema-0.7.2.json) | [`0.7.2/fluid-spec.html`](/fluid/specs/0.7.2/fluid-spec.html) | -| 0.7.1 | [`fluid-schema-0.7.1.json`](/fluid/schema/fluid-schema-0.7.1.json) | [`0.7.1/fluid-spec.html`](/fluid/specs/0.7.1/fluid-spec.html) | -| 0.5.7 | [`fluid-schema-0.5.7.json`](/fluid/schema/fluid-schema-0.5.7.json) | [`0.5.7/fluid-spec.html`](/fluid/specs/0.5.7/fluid-spec.html) | -| 0.4.0 | [`fluid-schema-0.4.0.json`](/fluid/schema/fluid-schema-0.4.0.json) | [`0.4.0/fluid-spec.html`](/fluid/specs/0.4.0/fluid-spec.html) | -| 0.3.0 | [`fluid-schema-0.3.0.json`](/fluid/schema/fluid-schema-0.3.0.json) | [`0.3.0/fluid-spec.html`](/fluid/specs/0.3.0/fluid-spec.html) | -| 0.2.0 | [`fluid-schema-0.2.0.json`](/fluid/schema/fluid-schema-0.2.0.json) | [`0.2.0/fluid-spec.html`](/fluid/specs/0.2.0/fluid-spec.html) | -| 0.1.1 | [`fluid-schema-0.1.1.json`](/fluid/schema/fluid-schema-0.1.1.json) | — | -| 0.1.0 | [`fluid-schema-0.1.0.json`](/fluid/schema/fluid-schema-0.1.0.json) | — | -| 0.0.1 | [`fluid-schema-0.0.1.json`](/fluid/schema/fluid-schema-0.0.1.json) | — | - -For what changed between any two versions, see the [**Changelog**](/fluid/schema/changelog). +Every published version of the FLUID JSON Schema. Point your validator at a version's JSON Schema; read its generated HTML reference for the field-by-field detail. + +- **Latest stable: 0.7.5.** Use it for new contracts. +- **Preview: 0.7.6.** Published so that it can be read and validated against. It can still change before it is promoted. See [Stable and preview versions](#stable-and-preview-versions). + +| Version | Status | JSON Schema | HTML Reference | +|---|---|---|---| +| **0.7.6** | **Preview** | [`fluid-schema-0.7.6.json`](/fluid/schema/fluid-schema-0.7.6.json) | [`0.7.6/fluid-spec.html`](/fluid/specs/0.7.6/fluid-spec.html) | +| **0.7.5** | **Stable (latest)** | [`fluid-schema-0.7.5.json`](/fluid/schema/fluid-schema-0.7.5.json) | [`0.7.5/fluid-spec.html`](/fluid/specs/0.7.5/fluid-spec.html) | +| 0.7.4 | Stable | [`fluid-schema-0.7.4.json`](/fluid/schema/fluid-schema-0.7.4.json) | [`0.7.4/fluid-spec.html`](/fluid/specs/0.7.4/fluid-spec.html) | +| 0.7.3 | Stable | [`fluid-schema-0.7.3.json`](/fluid/schema/fluid-schema-0.7.3.json) | [`0.7.3/fluid-spec.html`](/fluid/specs/0.7.3/fluid-spec.html) | +| 0.7.2 | Stable | [`fluid-schema-0.7.2.json`](/fluid/schema/fluid-schema-0.7.2.json) | [`0.7.2/fluid-spec.html`](/fluid/specs/0.7.2/fluid-spec.html) | +| 0.7.1 | Stable — [see note](#_0-7-1-two-documents-share-one-id) | [`fluid-schema-0.7.1.json`](/fluid/schema/fluid-schema-0.7.1.json) | [`0.7.1/fluid-spec.html`](/fluid/specs/0.7.1/fluid-spec.html) | +| 0.5.7 | Historical | [`fluid-schema-0.5.7.json`](/fluid/schema/fluid-schema-0.5.7.json) | [`0.5.7/fluid-spec.html`](/fluid/specs/0.5.7/fluid-spec.html) | +| 0.4.0 | Historical | [`fluid-schema-0.4.0.json`](/fluid/schema/fluid-schema-0.4.0.json) | [`0.4.0/fluid-spec.html`](/fluid/specs/0.4.0/fluid-spec.html) | +| 0.3.0 | Historical | [`fluid-schema-0.3.0.json`](/fluid/schema/fluid-schema-0.3.0.json) | [`0.3.0/fluid-spec.html`](/fluid/specs/0.3.0/fluid-spec.html) | +| 0.2.0 | Historical | [`fluid-schema-0.2.0.json`](/fluid/schema/fluid-schema-0.2.0.json) | [`0.2.0/fluid-spec.html`](/fluid/specs/0.2.0/fluid-spec.html) | +| 0.1.1 | Historical | [`fluid-schema-0.1.1.json`](/fluid/schema/fluid-schema-0.1.1.json) | — | +| 0.1.0 | Historical | [`fluid-schema-0.1.0.json`](/fluid/schema/fluid-schema-0.1.0.json) | — | +| 0.0.1 | Historical | [`fluid-schema-0.0.1.json`](/fluid/schema/fluid-schema-0.0.1.json) | — | + +For what changed between any two versions, see the [**Changelog**](/fluid/schema/changelog). For the release notes, see [**What's New**](/fluid/releases/). + +--- + +## Stable and preview versions + +| Status | Meaning | +|---|---| +| **Stable** | Published and settled. Each stable version is checked on every pull request by [`scripts/check-compat.py`](https://github.com/open-data-protocol/fluid/blob/main/scripts/check-compat.py): a document valid under one version must stay valid under the next (see [GOVERNANCE.md](https://github.com/open-data-protocol/fluid/blob/main/GOVERNANCE.md#versioning-and-compatibility)). | +| **Preview** | Published so that its `$id` resolves and a document that declares it can be validated by any JSON Schema validator. It is **not settled**: until it is promoted to stable, its fields and their meaning can still change, and the file at its URL is replaced when they do. | +| **Historical** | 0.5.7 and earlier. These predate the compatibility promise, which the gate enforces from 0.7.1 on. Several of those transitions removed or narrowed fields; the [Changelog](/fluid/schema/changelog) marks which. | + +### What "preview" means for you + +- **A contract opts in explicitly**, by declaring `fluidVersion: "0.7.6"`. Nothing else selects a preview. +- **Expect to re-validate.** The compatibility gate compares two *versions*. It cannot compare one revision of the 0.7.6 file with an earlier revision of the same file, so a contract written against an earlier revision of 0.7.6 may need changes after a later one. The schemas are authored in the reference implementation, [forge-cli](https://github.com/Agenticstiger/forge-cli), where `fluid_build/schemas/fluid-schema-0.7.6.json` has changed several times since 0.7.6 opened; its history is the record of those changes. +- **Promotion** makes a preview the latest stable version and is announced in its release notes. 0.7.5 went through the same path: it was an opt-in preview in forge-cli from 0.9.0 and was promoted to stable in forge-cli 0.12.0, when 0.7.6 opened as the next preview. + +### How the reference implementation treats a preview (non-normative) + +[`data-product-forge`](/fluid/concepts/forge-cli) bundles 0.7.6 and validates a contract that declares it (`✅ Valid FLUID contract (schema v0.7.6)`), but never chooses it on its own: its newest *stable* bundled version is its default, and `fluid init --quickstart` writes `fluidVersion: 0.7.5`. This describes one implementation. The standard's rule is only the one in [Validation semantics](/fluid/schema/specification#validation-semantics): a document is validated against the schema of the version it declares. + +--- + +## Choosing `fluidVersion` + +A document is validated against the schema whose version matches its `fluidVersion` ([Validation semantics](/fluid/schema/specification#validation-semantics)). Three consequences are easy to miss: + +1. **To use a field, declare the version that defines it.** The 0.7.5 schema's own `fluidVersion` enum also accepts `"0.7.3"` and `"0.7.4"`. That does not make a document declaring `"0.7.4"` valid when it uses a 0.7.5 field: its matching schema is 0.7.4, which rejects the field. For example, `binding.platform: confluent` (added in 0.7.5) in a contract declaring `"0.7.4"` is invalid, although the 0.7.5 schema alone would accept it. The reference implementation agrees: it reports `'confluent' is not one of [...]` against schema v0.7.4. +2. **Point your editor at the matching schema.** In a `# yaml-language-server: $schema=…` line, use the URL of the version the file declares. Pointing an older file at a newer schema hides exactly the mistake in (1). +3. **Upgrading is changing the number.** Every stable transition from 0.7.2 on keeps every valid document valid, so moving a valid contract to a later stable version is a one-line change. The one recorded exception is 0.7.1 → 0.7.2, where `notification` became a closed object (see the [0.7.2 release note](/fluid/releases/0.7.2#backward-compatibility-with-0-7-1-one-narrowing-change)). + +--- + +## 0.7.1: two documents share one `$id` + +The 0.7.1 file published here and the 0.7.1 file the reference implementation bundles carry the same `$id` (`…/schema/fluid-schema-0.7.1.json`) but are different documents, and they disagree in both directions: + +- The reference implementation's copy accepts fields the published file rejects — for example top-level `orchestration`, `binding.icebergConfig` and `exposes[].contract.quality`. +- It also rejects documents the published file accepts — for example an extra member on a `notifications[]` entry, or a column `type` outside its list of type names. + +This repository does not re-synchronise 0.7.1 (see [CONTRIBUTING.md](https://github.com/open-data-protocol/fluid/blob/main/CONTRIBUTING.md#the-schemas-are-not-authored-here)), so the file above is the 0.7.1 that this site documents. If a 0.7.1 result has to be reproducible across tools, validate against this file, or move the contract to 0.7.2 or later, where the published files and the reference implementation's copies agree on which documents are valid. + +## 0.4.0 and earlier: `$id`s that do not resolve + +The schemas from 0.0.1 to 0.4.0 declare `$id`s on `open-data-protocol.org`, a host that does not resolve, and their `$id`s do not follow the FLUID version (0.1.0 and 0.1.1 both declare `…/fluid.schema.v2.0.json`, 0.2.0 declares `…/fluid.schema.v2.3.json`). Two different documents sharing an `$id` cannot both be registered in a validator that keys schemas by `$id`. Load these files by their URL on this site instead. They are kept unchanged as history; from 0.5.7 on, every `$id` is the file's own URL here.