The open, declarative standard for Data Products.
FLUID provides the foundational protocol for building trustworthy, governable, and scalable data ecosystems—ready for the agentic era.
Quick Start:
- 📖 FLUID v0.7.1 Specification
- 🔗 JSON Schema v0.7.1
- 🚀 Examples in Action
- 🤝 Contributing Guide
- 🆕 What's New in v0.7.1
While both FLUID and ODPS aim to standardize data product specifications, they represent fundamentally different paradigms for data management:
- FLUID: DataOps-native approach emphasizing compliance-as-code, automated governance, and infrastructure-first data engineering
- ODPS: Business-first approach emphasizing data marketplace operations and commercial exchange
- FLUID: Enable end-to-end data product lifecycle automation with embedded compliance, quality, and governance from inception
- ODPS: Facilitate data product discovery, pricing, and commercial exchange between organizations
- Single Source of Truth: All governance, quality, lineage, and access policies embedded in version-controlled
.fluid.yml - Proactive Compliance: Governance enforced at build-time, not bolt-on post-deployment
- Infrastructure Automation: Native CI/CD integration with GitOps workflows
- Developer-Centric: Engineers define compliance rules alongside code, ensuring alignment
- Separation of Concerns: Business metadata separate from technical implementation
- Reactive Governance: Quality and SLA monitoring applied after deployment
- Commercial Operations: Built for data monetization and external sales
- Business-Centric: Product managers define commercial terms separately from technical teams
| DataOps Capability | FLUID v0.5.7 | ODPS v4.0 | FLUID Advantage |
|---|---|---|---|
| Compliance-as-Code | ✅ Native: Quality rules, policies, lineage embedded in specification | Unified compliance reduces tool sprawl and config drift | |
| GitOps Integration | ✅ Native: Version-controlled .fluid.yml drives entire lifecycle |
Automated deployments with compliance validation | |
| Developer Experience | ✅ Streamlined: Single file defines data product + governance | Faster development with embedded governance | |
| Environment Promotion | ✅ Automated: Same .fluid.yml works across dev/staging/prod |
Consistent governance across environments | |
| Change Management | ✅ Integrated: Schema evolution + quality rules versioned together | Atomic updates prevent configuration skew | |
| Audit Trail | ✅ Complete: Full lineage from source to governance in git history | Comprehensive audit for compliance teams | |
| Testing Strategy | ✅ Holistic: Data quality + business logic tested together | Higher confidence in production deployments | |
| Rollback Capability | ✅ Atomic: Entire data product + governance rolled back as unit | Safer operations with unified rollback |
| Feature | FLUID v0.5.7 | ODPS v4.0 | Why FLUID Wins |
|---|---|---|---|
| Build Automation | ✅ Comprehensive: dbt, Airflow, Python, multi-stage orchestration | ❌ None: No pipeline automation capabilities | End-to-end automation reduces operational overhead |
| Compliance-as-Code | ✅ Native: Quality, lineage, policies embedded in spec | Unified governance prevents compliance drift | |
| DataOps Workflows | ✅ Native: GitOps, CI/CD, environment promotion built-in | ❌ Manual: No workflow automation | Faster, safer deployments with automated validation |
| Schema Evolution | ✅ Managed: Built-in schema versioning and compatibility rules | Reduced breaking changes with automated compatibility checks | |
| Dependency Management | ✅ Explicit: Formal consumes relationships with version constraints |
Reliable data lineage prevents upstream breakage | |
| AI/ML Integration | ✅ Native: ML pipelines, feature stores, model deployment patterns | Complete ML lifecycle support for modern data teams | |
| Developer Velocity | ✅ High: Single file defines entire data product lifecycle | Faster iteration with unified development experience | |
| Operational Excellence | ✅ Proactive: Issues prevented through design-time validation | Higher reliability with shift-left quality approach |
- Single specification eliminates context switching between business and technical tools
- Embedded governance removes compliance bottlenecks from development cycle
- Automated deployments with built-in quality gates reduce manual toil
- Compliance-as-code makes governance requirements explicit and testable
- Version-controlled policies provide complete audit trails for regulatory requirements
- Proactive validation prevents non-compliant data products from reaching production
- Unified monitoring of technical and business metrics from single specification
- Atomic updates eliminate configuration drift between environments
- Comprehensive lineage enables rapid impact analysis for changes
- Native ML support for modern data teams building intelligent products
- Contract-driven development enables reliable AI agent integration
- Feature store patterns built into the specification
- ✅ DataOps transformation initiatives
- ✅ Compliance-heavy industries (finance, healthcare, government)
- ✅ Engineering-led data teams prioritizing automation
- ✅ AI/ML-centric organizations building intelligent products
- ✅ Internal data products requiring tight governance
- ✅ Data marketplace operations
- ✅ Commercial data sales with complex pricing models
- ✅ Business-led data product organizations
- ✅ External data distribution requiring legal frameworks
- ✅ Multi-vendor ecosystems needing business standardization
FLUID's focused scope is a deliberate design decision. Rather than trying to be everything to everyone, FLUID concentrates on what it does best—DataOps and technical governance—while acknowledging where ODPS provides superior capabilities:
Commercial Data Operations:
- Sophisticated pricing models: 12 standardized pricing patterns with payment gateway integration
- Legal framework management: Comprehensive licensing, IPR, and contract governance
- Multi-stakeholder governance: Business process workflows with detailed lifecycle states
- Marketplace operations: Product catalogs, payment processing, and customer relationship management
Business-Oriented Data Products:
- Rich business metadata: Value propositions, use cases, brand management, and marketing content
- Multi-language support: ISO 639-1 compliant internationalization for global data products
- Access diversity: Multiple consumption patterns (API, file, SQL, AI agents) per single product
- SLA sophistication: 11 monitoring dimensions with enterprise tool integrations (SodaCL, Montecarlo, DQOps)
What FLUID Deliberately Doesn't Do:
- ❌ Commercial operations: No pricing, billing, or payment processing
- ❌ Legal frameworks: No licensing or IPR management
- ❌ Marketing metadata: No brand slogans, value propositions, or sales content
- ❌ Multi-language UIs: English-first specification for technical teams
Why These Are Design Choices, Not Limitations:
-
Focus Drives Excellence: By concentrating on DataOps and technical governance, FLUID delivers deeper automation and better developer experience in its domain
-
Tool Ecosystem Integration: FLUID is designed to work with existing business systems, not replace them. Your data products can use FLUID for technical implementation while leveraging other tools for commercial operations
-
Separation of Concerns: Technical teams need different abstractions than business teams. FLUID optimizes for engineering workflows while remaining compatible with business-oriented specifications
-
Evolutionary Architecture: Organizations can start with FLUID for technical governance and later add ODPS for commercial operations as they mature their data product strategy
FLUID's design explicitly enables complementary coexistence with business-focused specifications:
# FLUID: Technical implementation and governance
fluidVersion: "0.7.1"
kind: "DataProduct"
id: "analytics.gold.customer_segments"
# NEW in v0.7.1: AI model governance
agentPolicy:
allowedModels: ["gpt-4", "claude-3-opus"]
maxTokensPerRequest: 4096
# Technical contract and automation
exposes:
- exposeId: "segments_api"
kind: "api"
contract:
# Reference to ODPS business specification
businessMetadata: "./customer-segments-odps.yaml"
# FLUID handles technical contract
schema: [...]
dq: [...]
binding:
platform: "kubernetes"
format: "http_api"
# FLUID handles build automation
build:
engine: "python"
pattern: "embedded-logic"
# ... technical implementation details# ODPS: Business packaging and commercialization
# File: customer-segments-odps.yaml
schema: https://opendataproducts.org/v4.0/schema/odps.yaml
version: 4.0
product:
details:
en:
name: "Customer Segmentation Analytics"
valueProposition: "AI-powered customer segments for personalized marketing"
# ... business metadata
pricingPlans:
declarative:
en:
- name: "Professional API Access"
price: 299
# ... commercial details
# Reference back to FLUID technical implementation
dataAccess:
api:
accessURL: "https://api.company.com/segments" # ← Deployed by FLUID
specsURL: "./fluid-generated-openapi.yaml" # ← Generated by FLUIDFLUID's "Do One Thing Well" Approach:
- Technical Excellence: Deep automation capabilities for data engineering teams
- Ecosystem Friendly: Designed to integrate with, not replace, existing business tools
- Evolutionary Path: Start with FLUID for technical governance, add business layers as needed
The Result: Best of Both Worlds
- Use FLUID for rapid development, automated compliance, and technical governance
- Use ODPS for commercial operations, legal frameworks, and business metadata
- Combine them for enterprises needing both technical excellence and business operations
This architectural approach allows organizations to:
✅ Start fast with FLUID's engineering-focused approach
✅ Scale commercially by adding ODPS business layers
✅ Avoid vendor lock-in through specification compatibility
✅ Optimize teams by matching tools to team responsibilities
FLUID represents the evolution of data engineering from reactive, tool-specific configurations to proactive, unified specifications. By embedding compliance, quality, and governance directly into the data product definition, FLUID enables organizations to achieve:
- Higher velocity through automated compliance validation
- Better reliability through design-time quality enforcement
- Reduced complexity through unified specifications
- Enhanced auditability through version-controlled governance
In an era where data governance is becoming a competitive advantage, FLUID provides the foundation for building trustworthy, scalable, and compliant data ecosystems ready for both human and AI consumption.
FLUID 0.7.1 represents a significant evolution focused on Agentic Governance and Provider-First Orchestration. Built with 100% backward compatibility with v0.5.7, it adds powerful new capabilities for the AI-driven enterprise:
Control which AI models can access your data and how they can use it:
# NEW in v0.7.1: AI/LLM usage policies
agentPolicy:
allowedModels:
- "gpt-4"
- "claude-3-opus"
- "gemini-1.5-pro"
maxTokensPerRequest: 8192
maxTokensPerDay: 100000
allowedUseCases:
- "customer-insights"
- "market-analysis"
deniedUseCases:
- "political-profiling"
- "credit-scoring"
requiresHumanReview: true
auditLog:
enabled: true
includePrompts: trueWhy this matters: As AI agents become primary data consumers, organizations need granular control over:
- ✅ Model-specific access - Whitelist/blacklist AI models
- ✅ Usage boundaries - Define permitted and prohibited use cases
- ✅ Rate limiting - Token quotas per request and per day
- ✅ Audit compliance - Full logging of AI interactions with data
- ✅ Human oversight - Require review for sensitive operations
Enforce data residency and jurisdictional compliance at the contract level:
# NEW in v0.7.1: Top-level sovereignty requirements
sovereignty:
jurisdiction: "EU"
dataResidency:
allowedRegions:
- "europe-west1"
- "europe-west3"
deniedRegions:
- "us-central1"
complianceFrameworks:
- "GDPR"
- "HIPAA"
crossBorderTransfer:
allowed: falseWhy this matters: Global compliance requires infrastructure-level enforcement:
- ✅ Jurisdictional boundaries - Enforce EU, US, APAC data laws
- ✅ Regional constraints - Specify allowed/denied cloud regions
- ✅ Compliance frameworks - Declare GDPR, HIPAA, SOC2 requirements
- ✅ Transfer controls - Block cross-border data movement
Direct invocation of provider actions as first-class orchestration tasks:
# NEW in v0.7.1: Provider actions without wrapper operators
orchestration:
engine: "airflow"
tasks:
- taskId: "ensure_s3_bucket"
type: "provider_action"
provider: "aws.s3"
action: "ensure_bucket"
parameters:
bucket_name: "customer-data-lake"
region: "us-west-2"
- taskId: "load_to_snowflake"
type: "provider_action"
provider: "snowflake.table"
action: "ensure"
parameters:
database: "ANALYTICS"
schema: "GOLD"
table: "CUSTOMER_360"
dependsOn: ["ensure_s3_bucket"]Why this matters: Simplifies multi-cloud orchestration:
- ✅ Native provider actions - AWS, GCP, Azure, Snowflake primitives
- ✅ No wrapper complexity - Direct action invocation
- ✅ Cross-provider workflows - Multi-cloud pipelines without vendor lock-in
- ✅ Strong typing - Provider-specific validation
Root-level accessPolicy for automated IAM binding generation:
# NEW in v0.7.1: Declarative access grants
accessPolicy:
grants:
- principal: "group:data-analytics@company.com"
permissions: ["read", "select", "query"]
resources:
- "$.exposes[?(@.kind=='table')]"
- principal: "serviceAccount:pipeline@project.iam.gserviceaccount.com"
permissions: ["write", "insert", "update"]
conditions:
ipRanges: ["10.0.0.0/8"]Why this matters: Infrastructure-as-code for data access:
- ✅ Automated IAM - Generate cloud IAM bindings from FLUID spec
- ✅ Resource targeting - JSONPath expressions for fine-grained access
- ✅ Conditional access - IP restrictions, time windows
- ✅ Audit-ready - Version-controlled access policies
| Feature | v0.5.7 | v0.7.1 | Impact |
|---|---|---|---|
| AI Model Control | ❌ None | ✅ agentPolicy | Govern AI/LLM data access |
| Data Sovereignty | ❌ Manual | ✅ sovereignty | Automated compliance enforcement |
| Orchestration | ✅ Provider-first | Direct cloud provider actions | |
| Access Control | ✅ Root-level accessPolicy | Centralized IAM automation | |
| Cross-Provider | ✅ Native | Simplified multi-cloud workflows | |
| Task Dependencies | ✅ Data product deps | Richer dependency graphs | |
| Error Handling | ✅ Categorized | Intelligent retry strategies | |
| Cost Tracking | ✅ Actual vs estimated | Budget enforcement |
All v0.5.7 contracts work unchanged in v0.7.1:
- ✅ No breaking changes
- ✅ New features are opt-in
- ✅ Existing patterns fully preserved
- ✅ Gradual adoption path
Migration is simple:
# Change version number - that's it!
fluidVersion: "0.7.1" # was "0.5.7"
# Optionally add new features
agentPolicy: { ... }
sovereignty: { ... }
accessPolicy: { ... }The "modern data stack"—a disaggregated ecosystem of best-in-class tools—has enabled rapid progress, but is held together by fragile scripts, proprietary configs, and tribal knowledge. This complexity, manageable by humans, becomes a liability in the Agentic Revolution.
Agentic AI—capable of complex reasoning and autonomous tool use—will soon be the primary consumer of enterprise data. Their potential, however, is capped by the quality and reliability of accessible data.
Key Questions:
- How can an agent trust the data it consumes?
- How does an agent discover the correct data product?
- How can we govern and audit thousands of autonomous agents accessing sensitive data?
The current landscape, built on disconnected pipelines, offers no scalable answers. Deploying agents atop this foundation is like building a skyscraper on sand. What’s needed is a paradigm shift: from data as pipeline output to data as a product with a contract.
FLUID is that foundational, declarative protocol.
FLUID is a declarative specification (YAML/JSON, version-controlled) that defines a data product's complete lifecycle. It's not an execution engine, but a universal contract language for the data ecosystem.
Core Philosophy (F.L.U.I.D):
- Federated: Distributed ownership and governance, enabling domain teams to own their data products while participating in a unified ecosystem. No central bottlenecks—each team controls their data destiny.
- Labeled: Rich metadata and semantic tagging throughout the specification, making data products discoverable, categorizable, and governable at scale. Every asset carries its context.
- Unifying: Single declarative contract that consolidates interface definitions, dependencies, build logic, quality rules, and access policies. One source of truth eliminates scattered configurations.
- Instructional: Clear, executable specifications that tell tools exactly how to build, deploy, and manage data products. The contract becomes the implementation blueprint.
- Declaration: Declarative-first approach where you specify what you want, not how to achieve it. Tools interpret the specification to determine optimal execution strategies.
Key Components in v0.7.1:
exposes: What data this product provides (schema, location, quality guarantees)consumes: What data this product depends on (other FLUID products or external sources)build: How the data gets created (dbt, SQL, Python, multi-stage pipelines)metadata: Ownership, business context, and governance informationagentPolicy⭐NEW: AI/LLM usage governance and controlsovereignty⭐NEW: Data residency and jurisdictional complianceaccessPolicy⭐NEW: Root-level access control with automated IAMorchestration⭐NEW: Provider-first task orchestration- Enhanced Features: Multi-modal builds, improved lineage, ML pipeline support, agentic governance
This structure separates interface (what you get) from implementation (how it's built), enabling reliable data ecosystems ready for both humans and AI agents.
FLUID is not a new central tool or platform. It does not replace Airflow, dbt, or Snowflake. It does not require a monolithic "Agentic Executor."
Instead, FLUID fosters a decentralized, compliant ecosystem. Tools become "FLUID-aware"—for example, Airflow dynamically generates DAGs from FLUID files, and data catalogs ingest lineage from FLUID repositories. FLUID is the shared language, not the central brain.
Understanding when to use FLUID versus OPDS v4 is crucial for making the right architectural decisions for your data ecosystem.
| Use FLUID When | Use OPDS v4 When |
|---|---|
| Building data products with complex transformations | Exposing data APIs with simple CRUD operations |
| Need end-to-end governance (build → deploy → consume) | Need API contract definition and documentation |
| Multi-modal pipelines (batch, streaming, ML) | Request/response data access patterns |
| Domain-driven data mesh architecture | Service-oriented or microservices architecture |
| Declarative infrastructure as code | Imperative API development workflows |
| Aspect | FLUID 0.7.1 | OPDS v4 |
|---|---|---|
| Primary Purpose | End-to-end data product lifecycle management | API specification and documentation |
| Scope | Data ingestion → transformation → consumption | HTTP API endpoints and schemas |
| Governance Model | Built-in data quality, lineage, and access policies | API versioning and compatibility |
| Build Patterns | hybrid-reference, embedded-logic, multi-stage |
Code generation from OpenAPI specs |
| Data Paradigms | Batch, streaming, ML pipelines, feature stores | Request/response, real-time queries |
| Metadata Richness | Business context, ownership, SLAs, observability | API documentation, examples, parameters |
| Execution Model | Tool-agnostic specification (dbt, Airflow, etc.) | HTTP server implementations |
| Consumer Experience | Data contracts with quality guarantees | API contracts with response schemas |
| Versioning | Semantic versioning with schema evolution | API version paths and deprecation |
| Discovery | Federated catalogs, lineage graphs | API registries, service mesh |
Data Mesh / Domain-Driven Architecture:
# FLUID: Complete data product specification
fluidVersion: "0.7.1"
kind: "DataProduct"
id: "finance.gold.risk_metrics"
# NEW in v0.7.1: Agentic governance
agentPolicy:
allowedModels: ["gpt-4", "claude-3"]
maxTokensPerDay: 50000
# Includes: sources, transformations, quality, access, observability
consumes: [...]
exposes: [...]
build: [...]
metadata: [...]Multi-Stage Data Pipelines:
- Bronze → Silver → Gold transformations
- ML training → inference → monitoring
- Streaming + batch processing hybrid
Enterprise Governance:
- Data quality as code
- Automated lineage tracking
- Policy-driven access control
- SLA monitoring and alerting
API-First Data Access:
# OPDS: API specification focus
openapi: 3.1.0
info:
title: Customer Data API
version: 4.0.0
paths:
/customers/{id}:
get:
responses:
'200':
content:
application/json:
schema:
$ref: '#/components/schemas/Customer'Microservices Data Layer:
- Service-to-service data exchange
- Real-time query interfaces
- API gateway integration
- Developer portal documentation
Request/Response Patterns:
- Interactive dashboards
- Mobile applications
- Third-party integrations
- Real-time analytics APIs
Many organizations benefit from using both specifications together:
# FLUID: Data product that exposes an API
fluidVersion: "0.7.1"
kind: "DataProduct"
id: "customer.api.profiles_v1"
# NEW: AI model restrictions
agentPolicy:
allowedModels: ["gpt-4-turbo"]
maxTokensPerRequest: 4096
exposes:
- exposeId: "customer_api"
kind: "api"
contract:
openapiRef: "./customer-profiles-api-v4.yaml" # ← OPDS v4 spec
binding:
platform: "kubernetes"
format: "http_api"
location:
baseUrl: "https://api.company.com/customers"
build:
engine: "python"
pattern: "embedded-logic"
# API server implementation details- Wrap existing APIs in FLUID specifications
- Add governance layers (quality, lineage, policies)
- Extend to full pipelines beyond just API endpoints
- Implement data mesh patterns gradually
- Extract API specifications from FLUID
exposes.contract.openapiRef - Focus on service boundaries rather than data pipelines
- Simplify to request/response patterns
- Optimize for developer experience
Let me provide a more accurate comparison based on careful analysis of both specifications:
| Concept Domain | FLUID 0.7.1 | OPDS v4.0 | Analysis |
|---|---|---|---|
| Product Definition | id, name, description, domain |
productID, name, description, valueProposition, productSeries |
OPDS stronger: Richer business context with value propositions and product series grouping |
| Lifecycle Management | lifecycle.state (4 states: preview→active→deprecated→retired) |
status (8 states: announcement→draft→development→testing→acceptance→production→sunset→retired) |
OPDS stronger: More granular lifecycle tracking for business processes |
| Data Contracts | Embedded contract.schema[] with inline field definitions |
External contract with contractURL or inline spec, supports ODCS/DCS standards |
Different approaches: FLUID=embedded simplicity, OPDS=external contract management standards |
| Quality Management | Built-in dq.rules[] with anomaly detection |
Comprehensive dataQuality with both declarative targets AND executable monitoring (SodaCL, Montecarlo, DQOps, Custom) |
OPDS stronger: Industry-standard DQ tool integration + declarative/executable pattern |
| SLA Framework | Basic qos (availability, freshness, latency) |
Comprehensive SLA with declarative objectives AND executable monitoring, support contacts, detailed dimensions |
OPDS significantly stronger: Production-grade SLA management |
| Access Methods | Single binding per expose |
Multiple dataAccess[] items with different outputPortType (file, API, SQL, AI, gRPC, sFTP) and formats |
OPDS stronger: Multiple access patterns per product |
| Business Operations | No commercial support | Complete pricingPlans[], paymentGateways[], license with legal frameworks |
OPDS exclusive: Full commercial data product support |
| Data Governance | Technical governance via policy, observability |
Business governance via license.governance, dataHolder legal entities |
Different focus: FLUID=technical, OPDS=business/legal |
| Pipeline Orchestration | Complete build patterns (hybrid-reference, embedded-logic, multi-stage) |
No transformation/pipeline logic | FLUID exclusive: Data engineering and pipeline management |
| Dependency Management | Formal consumes[] with version constraints |
Informal recommendedDataProducts[] |
FLUID stronger: Explicit dependency management |
| Metadata Richness | Technical metadata (tags, labels, businessContext) |
Business metadata (categories, standards, useCases[], brandSlogan) |
Different purposes: FLUID=technical discovery, OPDS=business discovery |
| Versioning Strategy | Semantic versioning with schemaEvolution |
Product versioning with versionNotes and issues tracking |
FLUID stronger: Technical schema evolution, OPDS stronger for business version communication |
| AI/LLM Governance | ✅ NEW: agentPolicy with model whitelisting, usage quotas, audit logging |
❌ No AI-specific governance | FLUID exclusive: Granular control over AI model access and usage boundaries |
| Data Sovereignty | ✅ NEW: sovereignty with jurisdiction, residency, cross-border controls |
FLUID stronger: Automated compliance enforcement at infrastructure level | |
| Orchestration | ✅ NEW: Provider-first tasks with direct cloud action invocation | ❌ No orchestration capabilities | FLUID exclusive: Native multi-cloud workflow management |
- Production SLA Management: Comprehensive monitoring-as-code with industry tools
- Business Product Management: Value propositions, use cases, product series
- Commercial Operations: Complete pricing, billing, legal, and payment frameworks
- Multi-Access Patterns: Supporting diverse consumption methods per product
- Quality Tooling: Integration with enterprise DQ tools (SodaCL, Montecarlo, DQOps)
- Lifecycle Granularity: Detailed business process states
- Data Engineering: Complete pipeline orchestration and transformation logic
- Technical Governance: Embedded contracts, lineage tracking, schema evolution
- AI/ML Workflows: Native support for ML pipelines and agentic consumption
- Agentic Governance ⭐NEW: AI model whitelisting, usage quotas, audit logging
- Data Sovereignty ⭐NEW: Jurisdiction enforcement, regional constraints, cross-border controls
- Provider-First Orchestration ⭐NEW: Direct cloud provider action invocation
- Access Automation ⭐NEW: Root-level IAM policy generation
- Dependency Management: Formal inter-product relationships with version constraints
- Multi-Environment: Environment-specific configurations (dev/staging/prod)
- Developer Experience: Unified specification for technical teams
- Underestimated OPDS quality management - It's actually more comprehensive with tool integrations
- Missed OPDS SLA sophistication - It's production-grade with monitoring-as-code
- Overlooked OPDS access diversity - Multiple access methods vs FLUID's single binding
- Didn't appreciate business vs technical focus - They serve different organizational needs
Choose FLUID 0.7.1 if you need:
- ✅ End-to-end data pipeline governance
- ✅ AI/ML pipeline orchestration
- ✅ Automated quality & lineage tracking
- ✅ Multi-environment data mesh architecture
- ✅ Agentic AI consumption with contracts
- ✅ AI model governance (NEW: agentPolicy)
- ✅ Data sovereignty enforcement (NEW: jurisdiction control)
- ✅ Provider-first orchestration (NEW: direct cloud actions)
Choose OPDS v4.0 if you need:
- ✅ Commercial data marketplace
- ✅ Legal compliance & licensing frameworks
- ✅ Business-oriented data catalogs
- ✅ Payment processing & billing integration
- ✅ Multi-access method data products
Use both together when:
- ✅ Building commercial data platforms with technical governance
- ✅ Need marketplace capabilities + pipeline orchestration
- ✅ Hybrid internal/external data product distribution
- ✅ Enterprise governance + ecosystem monetization
Can’t a smart AI just “get the data”? Why bother with data products?
No matter how advanced, an AI agent cannot operate on data it does not understand or trust. Connecting to raw databases is a liability, not an asset. FLUID closes three critical gaps:
- Semantic Gap: Without a contract, data is just bits. FLUID’s contract and semantics provide essential context—schema, descriptions, business ontology links.
- Trust Gap: How does an agent know data is correct or fresh? FLUID’s quality and SLA blocks provide enforceable guarantees.
- Governance Gap: How do we control and audit agent access? FLUID’s accessPolicy and dynamicPolicies create a programmatic access control layer.
Conclusion: AI cannot “just get the data.” FLUID provides the machine-readable contracts and policies that transform raw data into safe, trustworthy, and understandable Data Products.
- Replace glue code with declarative
.fluid.yml - Built-in governance, compliance & versioning
- Treat data as products
- Discoverable, composable, contract-driven data
- Machine-readable
- Secure
- Ready for AI-first enterprise infrastructure
| Principle | Description |
|---|---|
| Data as a Product | The core mental model of FLUID is to shift from thinking about "pipelines" to thinking about "products." A pipeline is an imperative process. A product is a versioned asset with a defined interface, quality guarantees, and a clear owner. FLUID files are the specification for these products. |
| Declarative, Not Imperative | You define the desired end state of your data product—what it consumes, what it exposes, and the contract it must adhere to. You do not define the step-by-step "how." This is the job of a FLUID-compliant tool, which reads your definition and figures out the best way to implement it. |
| Contracts as Code | The contract block is the heart of every data product. It embeds schema, quality rules, and privacy treatments directly into a version-controlled file. This makes governance an automated, proactive part of the development lifecycle, not a reactive, manual process. |
| Federated Ownership | FLUID is designed for a Data Mesh. .fluid.yml files are intended to be decentralized and co-located with the domain teams that own them. The standard's use of globally unique dataProduct names allows a central orchestrator or catalog to discover these distributed files and weave them into a single, unified data fabric. |
| Compliant Ecosystem | FLUID is not a monolithic platform. It is a standard that delegates execution to the tools you already use. An orchestrator, a catalog, or an ingestion service becomes "FLUID-aware" by learning to read .fluid.yml files to configure itself. This fosters an open, composable ecosystem rather than creating a new silo. |
A FLUID data product that ingests payment events with built-in quality controls:
# payments.fluid.yml
fluidVersion: "0.7.1"
kind: "DataProduct"
id: "finance.bronze.raw_payments"
name: "Raw Payment Events"
# NEW in v0.7.1: Data sovereignty
sovereignty:
jurisdiction: "US"
dataResidency:
allowedRegions: ["us-central1", "us-east1"]
metadata:
layer: "Bronze"
owner:
team: "data-platform"
email: "data-platform@company.com"
# What this data product creates
exposes:
- exposeId: "payment_events"
kind: "table"
contract:
schema:
- name: "payment_id"
type: "STRING"
required: true
description: "Unique payment identifier"
- name: "amount"
type: "NUMERIC"
required: true
description: "Payment amount"
- name: "currency"
type: "STRING"
required: true
description: "Currency code (USD, EUR, etc.)"
dq:
rules:
- id: "positive_amount"
type: "valid_values"
selector: "amount > 0"
severity: "error"
binding:
platform: "gcp"
format: "bigquery_table"
location:
project: "company-data"
dataset: "bronze_finance"
table: "payments"
# How it gets built
build:
engine: "sql"
pattern: "embedded-logic"
properties:
sql: |
SELECT
payment_id,
amount,
currency,
created_at
FROM raw_source.payments
WHERE amount > 0A FLUID data product that transforms raw data into business-ready insights:
# customer_metrics.fluid.yml
fluidVersion: "0.7.1"
kind: "DataProduct"
id: "analytics.silver.customer_metrics"
name: "Customer Metrics"
# NEW in v0.7.1: Access control automation
accessPolicy:
grants:
- principal: "group:analytics-team@company.com"
permissions: ["read", "select"]
metadata:
layer: "Silver"
owner:
team: "analytics"
email: "analytics@company.com"
# What data this consumes
consumes:
- productId: "finance.bronze.raw_payments"
exposeId: "payment_events"
- productId: "crm.bronze.raw_customers"
exposeId: "customer_data"
# What this creates
exposes:
- exposeId: "customer_ltv"
kind: "table"
contract:
schema:
- name: "customer_id"
type: "STRING"
required: true
- name: "total_spent"
type: "NUMERIC"
required: true
- name: "order_count"
type: "INTEGER"
required: true
- name: "avg_order_value"
type: "NUMERIC"
required: true
binding:
platform: "gcp"
format: "bigquery_table"
location:
project: "company-data"
dataset: "silver_analytics"
table: "customer_ltv"
build:
engine: "dbt"
pattern: "hybrid-reference"
properties:
model: "customer_ltv"A FLUID data product optimized for machine learning consumption:
# ml_features.fluid.yml
fluidVersion: "0.7.1"
kind: "DataProduct"
id: "ml.gold.churn_features"
name: "Churn Prediction Features"
# NEW in v0.7.1: AI model governance
agentPolicy:
allowedModels: ["gpt-4", "claude-3-opus"]
maxTokensPerRequest: 8192
allowedUseCases: ["churn-prediction", "customer-analytics"]
requiresHumanReview: false
metadata:
layer: "Gold"
owner:
team: "ml-engineering"
email: "ml@company.com"
consumes:
- productId: "analytics.silver.customer_metrics"
exposeId: "customer_ltv"
exposes:
- exposeId: "churn_features"
kind: "feature_store"
contract:
schema:
- name: "customer_id"
type: "STRING"
required: true
tags: ["identifier"]
- name: "recency_days"
type: "INTEGER"
description: "Days since last purchase"
- name: "frequency_score"
type: "FLOAT"
description: "Purchase frequency score"
- name: "monetary_score"
type: "FLOAT"
description: "Monetary value score"
policy:
authn: "iam"
authz:
readers: ["ml-agents", "data-scientists"]
binding:
platform: "gcp"
format: "bigquery_table"
location:
project: "company-ml"
dataset: "features"
table: "churn_v1"
build:
engine: "python"
pattern: "embedded-logic"
properties:
language: "python"
sql: |
SELECT
customer_id,
DATE_DIFF(CURRENT_DATE(), last_order_date, DAY) as recency_days,
LOG(1 + order_count) as frequency_score,
LOG(1 + total_spent) as monetary_score
FROM {{ ref('customer_ltv') }}A specification is only as strong as its ability to withstand scrutiny. Here, we address the toughest questions head-on.
A: FLUID eliminates complexity by unifying scattered configurations. Instead of separate dbt models, Airflow DAGs, data quality scripts, and access policies, you get one declarative file. Less moving parts = less complexity.
A: No. FLUID makes your tools work better together. dbt, Airflow, Snowflake, and other tools become "FLUID-aware" by reading the .fluid.yml specification to auto-configure themselves. It's the shared language, not a replacement platform.
A: Start small:
- Pick one critical data pipeline
- Write a
.fluid.ymlfile describing it (see examples above) - Use FLUID-compliant tools or build adapters for your existing stack
- Gradually expand to more data products
A: FLUID 0.7.1 supports multiple build patterns:
hybrid-reference: For dbt-style transformationsembedded-logic: For custom SQL/Python codemulti-stage: For complex multi-step orchestration
The lineage block maintains full traceability even with custom code.
A: AI agents need contracts, not chaos. FLUID 0.7.1 provides:
- Discoverable data: Agents can find the right data products
- Trustworthy contracts: Schema, quality, and freshness guarantees
- Secure access: Policy-driven permissions for autonomous systems
- Rich context: Business semantics and lineage for better decision-making
- ⭐ AI governance (NEW): Model whitelisting, usage quotas, and audit trails
- ⭐ Sovereignty controls (NEW): Automated jurisdictional compliance
- ⭐ Fine-grained permissions (NEW): Root-level access policies with automated IAM
📖 FLUID v0.7.1 Full Specification
🔗 JSON Schema v0.7.1
📚 Generated Schema Documentation
🆕 Version Diff: 0.5.7 → 0.7.1
🧑💻 FLUID Contribution Guide
📜 License (MIT)
FLUID is an open-source standard, and we welcome contributions from the community! Whether you are interested in refining the specification, building compliant tools, or creating new examples, there are many ways to get involved.
- Help build the agentic data future
- Contribute examples, tooling, or feedback
- Be part of an open, community-led protocol
Your agents are only as trustworthy as the data products they consume. Make FLUID your foundation.
📄 License This project is licensed under the MIT License - see the LICENSE.md file for details.
