From 2a0f83ec4e791827560a85a51684aa858043d80f Mon Sep 17 00:00:00 2001 From: Speculator55005 <50082482+fas89@users.noreply.github.com> Date: Fri, 9 Oct 2026 02:09:07 +0200 Subject: [PATCH 1/2] docs(iceberg): document the forge-cli defect-scan fix wave (forge-cli #710) - DynamoDB and JDBC sinks need a warehouse (location.warehouse, one derived from an explicit bucket on aws or gcp, or an override); BigQuery sinks need location.project and map region to gcp.bigquery.location - a {{ env.* }} bucket whose variable is unset warns at validate and is refused by the run preflight - a derived Kafka Connect sink writes the one Iceberg expose its outputs pick; a derived Debezium Server sink writes several only within one database and one catalog - platform: confluent with no catalog reads as glue; policy compile grants the Glue table Tableflow publishes - policy compile prints its warnings and fails on a compiler crash, with a new Errors section the CLI's policy_compiler_crashed link points to; policy apply echoes the bindings warnings - the Snowflake EXTERNAL VOLUME guard message names both possible causes --- docs/advanced/error-codes.md | 10 +- docs/advanced/governance.md | 2 +- docs/advanced/production-troubleshooting.md | 2 +- docs/advanced/source-aligned-acquisition.md | 208 ++++++++++++++++++-- docs/advanced/typed-cli-errors.md | 2 + docs/cli/apply.md | 12 +- docs/cli/generate-artifacts.md | 2 +- docs/cli/policy-apply.md | 1 + docs/cli/policy-compile.md | 43 +++- docs/cli/validate.md | 7 +- docs/concepts/builds-exposes-bindings.md | 1 + docs/providers/aws.md | 2 +- docs/providers/gcp.md | 16 +- docs/providers/snowflake.md | 14 +- 14 files changed, 290 insertions(+), 32 deletions(-) diff --git a/docs/advanced/error-codes.md b/docs/advanced/error-codes.md index fa8409a..5dfff88 100644 --- a/docs/advanced/error-codes.md +++ b/docs/advanced/error-codes.md @@ -38,7 +38,8 @@ The route table sends an event to one of these pages, and the entries below name |---|---| | [`fluid providers`](../cli/providers.md) | The provider events | | [`fluid secrets`](../cli/secrets.md) | `copilot_missing_llm_api_key` | -| [Sovereignty](../concepts/sovereignty.md) | The policy and sovereignty events | +| [Sovereignty](../concepts/sovereignty.md) | The policy and sovereignty events, except `policy_compiler_crashed` | +| [`fluid policy compile`, Errors](../cli/policy-compile.md#errors) | `policy_compiler_crashed` *(unreleased, [forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710))* | | [`fluid verify-signature`](../cli/verify-signature.md) | The signing events | | [Getting started](../getting-started/README.md) | `opentofu_engine_install_failed` | | [Typed CLI errors](./typed-cli-errors.md) | The schema-version events and the connectivity events | @@ -360,6 +361,13 @@ The state refusals (`state_shared_with_another_provider`, `state_migration_ambig - Check the agent-policy block in the contract; run 'fluid policy check <contract>' +### policy_compiler_crashed + +*(unreleased, [forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710))* `ERR_POLICY_COMPILER_CRASHED`. The documentation link lands on [`fluid policy compile`, Errors](../cli/policy-compile.md#errors). + +- Run 'fluid validate <contract>': policy compile reads accessPolicy and exposes without validating them against the contract schema +- If the contract validates, re-run with 'fluid --log-level DEBUG policy compile <contract>' to see the compiler's traceback + ### policy_apply_failed `ERR_POLICY_APPLY_FAILED`. The documentation link lands on [Sovereignty](../concepts/sovereignty.md). diff --git a/docs/advanced/governance.md b/docs/advanced/governance.md index 4068593..9de75d5 100644 --- a/docs/advanced/governance.md +++ b/docs/advanced/governance.md @@ -170,7 +170,7 @@ fluid policy-compile contract.fluid.yaml --out runtime/policy/bindings.json } ``` -`read`-style permissions map to a viewer role and `write`, `insert`, `update` or `delete` to an owner role. A contract with no grants compiles to an empty list and a `No grants found in accessPolicy` warning. +`read`-style permissions map to a viewer role and `write`, `insert`, `update` or `delete` to an owner role. A contract with no grants compiles to an empty list and a `No grants found in accessPolicy` warning. *([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* Other warnings, such as one for a grant that compiled to no binding, are also printed at WARNING level, and a crash inside the compiler exits `1` with `policy_compiler_crashed` and writes no file. See [Warnings](../cli/policy-compile.md#warnings) and [Errors](../cli/policy-compile.md#errors). | Option | Description | Default | |--------|-------------|---------| diff --git a/docs/advanced/production-troubleshooting.md b/docs/advanced/production-troubleshooting.md index 8115fe8..7f9a3a0 100644 --- a/docs/advanced/production-troubleshooting.md +++ b/docs/advanced/production-troubleshooting.md @@ -107,7 +107,7 @@ Remote OpenTofu state is keyed per contract and provider. See [Environment varia | `state_migration_unverified` | After the copy, the new key holds different resources than the old one. The old object is untouched | Compare the two states before re-running | | `state_migration_probe_failed` | The old key could not be read | Check access to the bucket and the key named in the message | | `opentofu_region_moved` | State holds this contract's resources in a region other than the one the bindings now name. Applying would create them again and leave the originals unmanaged | If they should stay, set the binding's `location.region` to the region the error names. If they should move, empty and remove them there first (`tofu destroy` in the state directory named, with `AWS_REGION` set to the old region), then apply again | -| `iceberg_catalog_move_blocked` | *(unreleased, [forge-cli #707](https://github.com/Agenticstiger/forge-cli/pull/707) and [forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709))* An Iceberg expose's `location.catalog` names a catalog an earlier release did not honour, and state still holds what that release created for it: a Glue database on AWS, and the Glue table when the expose names one, or an EXTERNAL VOLUME on Snowflake. The plan would destroy them, and destroying a Glue database deletes every table in it | Run the `tofu -chdir= state rm
` commands the error prints, which change nothing in the cloud, then apply again. Do not pass `--allow-data-loss`. If the table belongs in Glue or in Snowflake's own catalog, remove `location.catalog` instead. Steps: [Upgrading an AWS contract that names another catalog](./source-aligned-acquisition.md#upgrading-an-aws-contract-that-names-another-catalog), [Upgrading a Snowflake contract that names another catalog](./source-aligned-acquisition.md#upgrading-a-snowflake-contract-that-names-another-catalog) | +| `iceberg_catalog_move_blocked` | *(unreleased, [forge-cli #707](https://github.com/Agenticstiger/forge-cli/pull/707), [forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709) and [forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710))* An Iceberg expose's `location.catalog` names a catalog an earlier release did not honour, and state still holds what that release created for it: a Glue database on AWS, and the Glue table when the expose names one, or an EXTERNAL VOLUME on Snowflake. The Snowflake volume is named per contract, so it may instead have served a Snowflake-managed Iceberg expose that this change removed or moved to another catalog; the message names both causes. The plan would destroy them, and destroying a Glue database deletes every table in it | Run the `tofu -chdir= state rm
` commands the error prints, which change nothing in the cloud, then apply again. Do not pass `--allow-data-loss`. If the table belongs in Glue or in Snowflake's own catalog, remove `location.catalog` instead. Steps: [Upgrading an AWS contract that names another catalog](./source-aligned-acquisition.md#upgrading-an-aws-contract-that-names-another-catalog), [Upgrading a Snowflake contract that names another catalog](./source-aligned-acquisition.md#upgrading-a-snowflake-contract-that-names-another-catalog) | | `iceberg_catalog_move_probe_skipped` (WARNING) | *(unreleased, [forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709))* The catalog-move guard could not read the state, or its check failed, so the apply went on without it. The `reason` field says why | If the plan then destroys Glue resources or an EXTERNAL VOLUME of the exposes the warning names, stop: release them with `tofu state rm` and apply again, rather than passing `--allow-data-loss` to the data-loss gate | ## Contract load and overlay errors diff --git a/docs/advanced/source-aligned-acquisition.md b/docs/advanced/source-aligned-acquisition.md index d153c35..b0cb58d 100644 --- a/docs/advanced/source-aligned-acquisition.md +++ b/docs/advanced/source-aligned-acquisition.md @@ -218,7 +218,7 @@ For the DuckDB engine, each connection runs in DuckDB's own sandbox ([DuckDB san ## Iceberg catalogs (`location.catalog`) ::: warning Not in a release yet -This section describes forge-cli [PR #707](https://github.com/Agenticstiger/forge-cli/pull/707) and the follow-up fixes in [forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709), marked *([forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709), unreleased)*. No release includes either yet. forge-cli 0.19.0 and earlier behave as described in [On 0.19.0 and earlier](#on-0-19-0-and-earlier), at the end of this section. +This section describes forge-cli [PR #707](https://github.com/Agenticstiger/forge-cli/pull/707) and the follow-up fixes in [forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709), marked *([forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709), unreleased)*, and in [forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), marked *([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)*. No release includes any of them yet. forge-cli 0.19.0 and earlier behave as described in [On 0.19.0 and earlier](#on-0-19-0-and-earlier), at the end of this section. ::: An Iceberg expose names the catalog that owns its table in `binding.location.catalog`. The streaming sinks (Kafka Connect and Debezium Server), dbt's `catalogs.yml`, the AWS and Confluent modules, `fluid policy compile`, `fluid diff`, `fluid test` and `fluid validate` read the value through one table in forge-cli (`fluid_build/providers/_iceberg_catalog.py`), so they agree on which catalog holds the table. The Snowflake module reads it for the Iceberg prerequisites only: the EXTERNAL VOLUME and the Glue catalog integration follow the table. It still emits a `snowflake_database`, `snowflake_schema` and `snowflake_table` for the expose's `location.database`, `location.schema` and `location.table`, whatever the catalog. @@ -316,7 +316,7 @@ The one resource is `aws_s3_bucket.bronze_orders_stream_acme_lake`. Without `buc ### How a value is read - **Spelling.** Case, surrounding whitespace, and `-` against `_` are folded, so `Lakekeeper` is `lakekeeper` and `SNOWFLAKE_MANAGED` is `snowflake-managed`. Two spellings are aliases: `iceberg-rest` (or `iceberg_rest`) is `rest`, and `snowflake` is `snowflake-managed`. -- **No value.** An Iceberg expose with no `location.catalog` gets its platform's default: `glue` on `platform: aws`, `snowflake-managed` on `platform: snowflake`, and `rest` on any other platform. On `platform: gcp` the default means two things: a streaming sink writes through a REST catalog, while dbt-bigquery and the GCP module create a BigLake table. *([forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709), unreleased)* So a `platform: gcp` Iceberg expose that a streaming sink writes must name its catalog. Without one, `fluid validate` fails, and so does the Kafka Connect or embedded Debezium Server run before it creates anything. Set `catalog: bigquery`, or the REST kind your catalog is. #707 alone accepted it. An expose that no streaming sink writes may still leave the catalog out. The error, for a Kafka Connect build `stream_events` writing a GCP expose with no catalog: +- **No value.** An Iceberg expose with no `location.catalog` gets its platform's default: `glue` on `platform: aws` and `platform: confluent`, `snowflake-managed` on `platform: snowflake`, and `rest` on any other platform. *([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* `platform: confluent` reads as `glue` because the Tableflow module publishes such a table to AWS Glue, so `fluid policy compile` and dbt read the catalog the table is in. With #707 and #709 alone it read as `rest`, and `fluid policy compile` dropped the expose's grants with a warning that the table was cataloged in `rest`; see [Other commands](#other-commands). On `platform: gcp` the default means two things: a streaming sink writes through a REST catalog, while dbt-bigquery and the GCP module create a BigLake table. *([forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709), unreleased)* So a `platform: gcp` Iceberg expose that a streaming sink writes must name its catalog. Without one, `fluid validate` fails, and so does the Kafka Connect or embedded Debezium Server run before it creates anything. Set `catalog: bigquery`, or the REST kind your catalog is. #707 alone accepted it. An expose that no streaming sink writes may still leave the catalog out. The error, for a Kafka Connect build `stream_events` writing a GCP expose with no catalog: ```text 1. iceberg sink (build 'stream_events'): the GCP Iceberg expose sets no binding.location.catalog, so it is read two ways: the sink would write through a REST catalog (the 'gcp' platform default) while dbt-bigquery and the GCP IaC, which read only binding.location.catalog, create a BigLake metastore table. Set binding.location.catalog: bigquery, or the REST kind your catalog is (e.g. rest, lakekeeper) @@ -366,11 +366,11 @@ Declare the catalog the sink writes to in `binding.location.catalog`. A REST end | `polaris` | `type=rest` | `iceberg_rest` | nothing; validate warns | the bucket only | `uri`, `warehouse` | | `unity` | `type=rest` | `iceberg_rest` | nothing; validate warns | the bucket only | `uri`, `warehouse` | | `nessie` | `type=nessie` | `iceberg_rest` | nothing; validate warns | the bucket only | `uri`, `warehouse` | -| `bigquery` | `type=bigquery` | `iceberg_rest` | nothing; validate warns | the bucket only | nothing (a Kafka Connect build warns) | +| `bigquery` | `type=bigquery` | `iceberg_rest` | nothing; validate warns | the bucket only | `project` (warns with no `gs://` warehouse to derive, and on a Kafka Connect build) | | `hive` | `type=hive` | left out, with a warning | nothing; validate errors | the bucket only | nothing | -| `jdbc` | `type=jdbc` | left out, with a warning | nothing; validate errors | the bucket only | `uri` | +| `jdbc` | `type=jdbc` | left out, with a warning | nothing; validate errors | the bucket only | `uri`, and `warehouse` or a `bucket` to derive it from | | `hadoop` | `type=hadoop` | left out, with a warning | nothing; validate errors | the bucket only | `warehouse` | -| `dynamodb` | `catalog-impl=org.apache.iceberg.aws.dynamodb.DynamoDbCatalog` | left out, with a warning | nothing; validate errors | the bucket only | nothing | +| `dynamodb` | `catalog-impl=org.apache.iceberg.aws.dynamodb.DynamoDbCatalog` | left out, with a warning | nothing; validate errors | the bucket only | `warehouse`, or a `bucket` to derive it from | | `snowflake-managed` | `type=rest` | `built_in` | `EXTERNAL VOLUME` | the bucket only | `uri`, `warehouse` | How to read the columns: @@ -378,18 +378,160 @@ How to read the columns: - **Sink catalog selector.** Kafka Connect takes the key with the `iceberg.catalog.` prefix (`iceberg.catalog.type`, `iceberg.catalog.catalog-impl`), and Debezium Server with `debezium.sink.iceberg.`. The `type` values are catalog types Apache Iceberg defines. Iceberg has no `lakekeeper`, `polaris` or `unity` type, so those catalogs are reached over Iceberg REST, and no `dynamodb` type, so DynamoDB is selected by class. `type=bigquery` needs an Iceberg runtime of 1.10 or later. For `nessie` the sink uses Iceberg's Nessie client, which the stock Apache Iceberg Kafka Connect runtime does not bundle, so `fluid validate` warns on a `kafka-connect` build. *([forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709), unreleased)* A `kafka-connect` build that sends `type=bigquery` draws a warning too, because the published Apache Iceberg Kafka Connect sink predates it: ```text - 1. iceberg sink (build 'stream_events'): the sink config sets iceberg.catalog.type=bigquery, which the published Apache Iceberg Kafka Connect sink (1.9.2 on Confluent Hub) cannot load: Iceberg's CatalogUtil gains the bigquery type in 1.10. Run a sink built from Iceberg >= 1.10, or the connector fails at start + 1. iceberg sink (build 'stream_events'): the sink config sets iceberg.catalog.type=bigquery, which the published Apache Iceberg Kafka Connect sink (1.9.2 on Confluent Hub) cannot load: Iceberg's CatalogUtil gains the bigquery type in 1.10, so on that sink the connector fails at start. Run a sink built from Iceberg >= 1.10 ``` Both warnings follow the catalog that reaches the worker, selected by `type` or by a `catalog-impl` class (`org.apache.iceberg.nessie.NessieCatalog`, `org.apache.iceberg.gcp.bigquery.BigQueryMetastoreCatalog`), after `iceberg_catalog_overrides` and a hand-written `sink_connector_config` are merged in. A hand-written config that sets `iceberg.catalog.type: rest` for a `bigquery` or `nessie` expose draws neither, and is an error instead, because it selects another catalog than the expose's ([How a value is read](#how-a-value-is-read)). - **dbt `catalogs.yml` on Snowflake** and **Snowflake module** apply to `platform: snowflake`. The **Snowflake module** column lists the Iceberg prerequisite the module creates; for every kind it also emits the database, schema and `snowflake_table` the binding's `location` names. Apart from `glue`, Snowflake reaches the `iceberg_rest` kinds through a catalog integration that authenticates with a secret. The module is credential-free, so it creates none, and `fluid validate` warns (an error under `--strict`). Snowflake has no catalog integration for `hive`, `jdbc`, `hadoop` or `dynamodb`, so an expose naming one fails `fluid validate`. The prerequisites each emitted object needs are in [Iceberg tables via dbt](../providers/snowflake.md#iceberg-tables-via-dbt-since-0-13-1). - **AWS module** applies to `platform: aws`: the bucket is the one the binding names. See [On AWS](#on-aws-a-table-in-another-catalog). -- **A streaming sink needs** the listed `binding.location` keys when a Kafka Connect build, or a Debezium Server build in `embedded` mode, writes the expose. Debezium in `bring-your-own` or `managed` mode creates only the source connector, so these checks do not apply to it. The Kafka Connect runner and the embedded Debezium Server runner run the same checks before they create anything, so a contract `fluid validate` refuses also fails its run, for these builds: +- **A streaming sink needs** the listed `binding.location` keys when a Kafka Connect build, or a Debezium Server build in `embedded` mode, writes the expose. *([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* A `dynamodb` or `jdbc` warehouse may instead derive from a `bucket` or come from an override, and a `bigquery` project from an override; see [What a DynamoDB, JDBC or BigQuery sink needs](#what-a-dynamodb-jdbc-or-bigquery-sink-needs). Debezium in `bring-your-own` or `managed` mode creates only the source connector, so these checks do not apply to it. The Kafka Connect runner and the embedded Debezium Server runner run the same checks before they create anything, so a contract `fluid validate` refuses also fails its run, for these builds: - *([forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709), unreleased)* a build that declares `sink.format: iceberg`, whether the runner derives the sink config or pushes a hand-written one: `sink_connector_config` on Kafka Connect, `server.sink.config` on an embedded Debezium Server build whose `server.sink.type` is `iceberg` (the default). On #707 alone the runners checked only a config they derived, so a hand-written config that `fluid validate` refuses was still deployed; - an embedded Debezium Server build that derives its sink: `server.sink.type: iceberg` (the default) with no hand-written `server.sink.config`, or with `server.sink.iceberg_sink_enabled: true`. *([forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709), unreleased)* In a contract with several builds, each of these two runners reads its properties from the build it runs, the build the checks read. They used to read the first build's. +### What a DynamoDB, JDBC or BigQuery sink needs + +*([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* Apache Iceberg's `DynamoDbCatalog` and `JdbcCatalog` refuse to start without a warehouse, and its `BigQueryMetastoreCatalog` refuses to start without `gcp.bigquery.project-id`. A sink config forge-cli derives now carries them, and `fluid validate` and the run preflight refuse a build whose sink would start without them. With #707 and #709 alone forge-cli derived neither and checked for neither, so the connector failed when it started. + +**DynamoDB and JDBC.** The warehouse is `location.warehouse`. Without one, forge-cli derives it from an explicit `location.bucket` on `platform: aws` (`s3://`) or `platform: gcp` (`gs://`): `:///`, where `path` defaults to `//`. A bucket on any other platform derives nothing, and the account-derived bucket a Glue table falls back to is never used, because no module creates it for a table in another catalog. A DynamoDB expose that a Kafka Connect build writes: + +```yaml +exposes: + - exposeId: orders + kind: table + binding: + platform: aws + format: iceberg + location: + catalog: dynamodb + bucket: acme-lake # the warehouse derives from it + database: streaming + table: orders + region: eu-west-1 +``` + +The table and catalog keys of the sink config the Kafka Connect runner derives from it: + +```json +{ + "iceberg.tables": "streaming.orders", + "iceberg.catalog.catalog-impl": "org.apache.iceberg.aws.dynamodb.DynamoDbCatalog", + "iceberg.catalog.warehouse": "s3://acme-lake/streaming/orders/", + "iceberg.catalog.io-impl": "org.apache.iceberg.aws.s3.S3FileIO", + "iceberg.catalog.client.region": "eu-west-1" +} +``` + +The sink's own `warehouse` property counts too: `iceberg.catalog.warehouse` in `iceberg_catalog_overrides` or `sink_connector_config` on Kafka Connect, `warehouse` in `server.sink.config` on embedded Debezium Server. With no `location.warehouse`, no bucket to derive one from and no override, `fluid validate` fails, and so does the run, before it creates anything: + +```text + 1. iceberg sink (build 'stream_orders'): dynamodb catalog requires binding.location.warehouse (an object-store location), or a binding.location.bucket on platform aws or gcp to derive it from, or the sink's warehouse property in an override; the dynamodb catalog refuses to start without a warehouse +``` + +`jdbc` needs `location.uri` as well. A `location.warehouse` of only whitespace counts as unset. + +**BigQuery.** A `catalog: bigquery` expose maps to the catalog's own properties: + +| `binding.location` | Derived sink config | +|---|---| +| `project` | `gcp.bigquery.project-id`. Required: without it, or with only whitespace, `fluid validate` and the run refuse the build, unless an override sets that property | +| `region` | `gcp.bigquery.location`. The config sets no `client.region` for this catalog | +| a `gs://` `warehouse`, else `bucket` and `path` | `warehouse`: the `gs://` storage dbt-bigquery and the GCP module use, `gs://` plus `path` when it is set | + +The Kafka Connect runner passes them with the `iceberg.catalog.` prefix, and the embedded Debezium Server runner with `debezium.sink.iceberg.`. [Iceberg on BigQuery via dbt](../providers/gcp.md#iceberg-on-bigquery-via-dbt-since-0-14-0) shows the derived config for its example binding. Without `project`: + +```text + 1. iceberg sink (build 'stream_events'): bigquery catalog requires binding.location.project (the sink's gcp.bigquery.project-id), or that property in an override; the bigquery catalog refuses to start without it +``` + +A `warehouse` with another scheme, or no `gs://` warehouse and no bucket, derives no warehouse: the config sets none, and `fluid validate` warns without refusing. On Kafka Connect, table auto-create then fails: with `iceberg.tables.auto-create-enabled` the sink calls `createNamespace` for each table it creates, which `BigQueryMetastoreCatalog` refuses without a warehouse, even in a dataset that exists. Tables that exist need no warehouse. Embedded Debezium Server does not boot without `debezium.sink.iceberg.warehouse`, which has no default. On `platform: gcp` the [Iceberg prerequisite checks](../cli/validate.md#iceberg-prerequisite-checks-since-0-14-0) already require a `bucket` or a `gs://` warehouse. + +**A bucket written as a `{{ env.* }}` template.** `fluid validate` reads such a bucket as written. When a variable it names, such as `LAKE_ENV` in `acme-{{ env.LAKE_ENV }}-lake`, is unset or empty where validate runs, the warehouse cannot be derived there, and validate warns and names the variable instead of asking for a bucket: + +```text + 1. iceberg sink (build 'stream_orders'): the dynamodb catalog's warehouse derives from binding.location.bucket 'acme-{{ env.LAKE_ENV }}-lake', and LAKE_ENV is unset or empty here, so it cannot be derived at validate time. Set it where the sink runs: the runner refuses the build when the bucket does not resolve there, because the dynamodb catalog refuses to start without a warehouse +``` + +The run preflight checks the variable again where the sink config is derived. When it is unset or empty there, the preflight refuses the build instead of pushing a warehouse in a bucket the contract does not name, such as `acme--lake`. That includes a run from a plan, whose embedded contract it reads as written: + +```text +iceberg sink preflight failed: iceberg sink (build 'stream_orders'): the dynamodb catalog's warehouse derives from binding.location.bucket 'acme-{{ env.LAKE_ENV }}-lake', and LAKE_ENV is unset or empty in the runner's environment, so the bucket does not resolve: the sink would get a warehouse in a bucket the contract does not name, or none, and the dynamodb catalog refuses to start without a warehouse. Set it here, or set binding.location.warehouse +``` + +The preflight refuses a DynamoDB or JDBC sink, a BigQuery sink on embedded Debezium Server, and a BigQuery sink on Kafka Connect with auto-create on (`streamingSink.autoCreate: true`, or `iceberg.tables.auto-create-enabled` in an override). A Kafka Connect BigQuery sink reads the warehouse only to create tables, so without auto-create the run warns and pushes no warehouse. A warehouse an override sets is kept, and the bucket's variables are not checked then. They are not checked for a hand-written sink config forge-cli does not derive either. + +### Which exposes a streaming sink writes + +*([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* A Kafka Connect or embedded Debezium Server build whose Iceberg sink config forge-cli derives writes the Iceberg exposes its `outputs` name. An Iceberg expose here is one with an Iceberg format whose `binding.platform` is not `confluent`: a Tableflow expose is published by its own module. `fluid validate`, the run preflight and both runners resolve the exposes the same way, so they name the same tables. On 0.19.0 and earlier, and with #707 and #709 alone, a derived sink wrote the contract's first Iceberg expose, whatever the build's `outputs` named. + +Two Kafka Connect builds, one per expose (each expose's `contract` and each build's `source` are left out): + +```yaml +exposes: + - exposeId: orders + kind: table + binding: + platform: aws + format: iceberg + location: {bucket: acme-lake, database: sales, table: orders, region: eu-west-1} + - exposeId: refunds + kind: table + binding: + platform: aws + format: iceberg + location: {bucket: acme-lake, database: sales, table: refunds, region: eu-west-1} +builds: + - id: stream_orders + pattern: acquisition + engine: kafka-connect + outputs: [orders] + properties: + sink: {format: iceberg} + - id: stream_refunds + pattern: acquisition + engine: kafka-connect + outputs: [refunds] + properties: + sink: {format: iceberg} +``` + +`stream_orders` pushes `iceberg.tables=sales.orders` and `stream_refunds` pushes `iceberg.tables=sales.refunds`. With #707 and #709 alone, `stream_refunds` pushed `iceberg.tables=sales.orders` too, and `fluid validate`, the preflight and the run all passed. + +- **Kafka Connect** writes one expose, because the derived config carries one `iceberg.tables` entry. The build's `outputs` must name exactly one Iceberg expose; a build with no `outputs` is accepted only when the contract has exactly one. Otherwise `fluid validate` and the run preflight refuse it. With `outputs: [orders, refunds]` on `stream_refunds`: + + ```text + 1. iceberg sink (build 'stream_refunds'): its outputs ['orders', 'refunds'] name 2 of the Iceberg sink exposes ['orders (sales.orders)', 'refunds (sales.refunds)']; a derived Kafka Connect sink writes one expose (one iceberg.tables entry). List exactly one of them in the build's outputs, split the build into one build per expose, or hand-write the sink config (properties.kafka-connect.sink_connector_config) + ``` + + Outputs that name no Iceberg expose are refused the same way. With #707 and #709 alone they drew only a warning that the join is implicit. +- **Embedded Debezium Server** writes every captured table under one `table-namespace`, through one catalog. It writes the Iceberg exposes its `outputs` name, or all of the contract's with no `outputs`, and is refused when: + - its outputs name none of them; + - they sit in more than one `binding.location.database`; + - they resolve to different catalogs. The error names the settings that differ, such as `catalog`, `uri` or `warehouse`. A Glue warehouse is a per-table prefix, so it is not compared, and the sink takes the first expose's; + - they are DynamoDB or JDBC exposes whose warehouses, derived from `location.bucket`, differ. Such a catalog creates every missing table under the one warehouse the sink is given, and each derived warehouse is that expose's own table prefix. Set one `location.warehouse` on those exposes, or split the build. + + One embedded Debezium Server build over both exposes above (its `source` left out), with `refunds` moved to `database: finance`: + + ```yaml + - id: cdc_sales + pattern: acquisition + engine: debezium + outputs: [orders, refunds] + properties: + sink: {format: iceberg} + debezium: + deployment: {mode: embedded} + ``` + + `fluid validate` refuses it: + + ```text + 1. iceberg sink (build 'cdc_sales'): a derived Debezium Server sink writes every captured table under one table-namespace, but the exposes its outputs ['orders', 'refunds'] name, ['orders (sales.orders)', 'refunds (finance.refunds)'], sit in the databases ['finance', 'sales']. Give them one binding.location.database, split the build into one build per database, or hand-write the sink config (properties.debezium.server.sink.config) + ``` + +- **A config that names its own tables.** A hand-written `sink_connector_config` or `server.sink.config` that forge-cli does not derive from, and an override that sets `iceberg.tables` (Kafka Connect) or `table-namespace` (Debezium Server), behave as before: the config is checked against the first Iceberg expose, and outputs that name no Iceberg expose draw the warning that the join is implicit. + ### Kafka Connect: `catalog-impl` or `type`, never both Apache Iceberg refuses a catalog configured with both `type` and `catalog-impl`. On 0.19.0 and earlier the Kafka Connect sink config for a Glue table carried both, and the sink failed at startup with: @@ -495,33 +637,67 @@ The [OpenTofu data-loss gate](../cli/apply.md#opentofu-data-loss-gate) still sto ```text ❌ iceberg_catalog_move_blocked [ERR_ICEBERG_CATALOG_MOVE_BLOCKED] kind: iceberg-catalog-move - error: iceberg catalog move blocked — this contract's OpenTofu state holds 1 Snowflake EXTERNAL VOLUME(s) for Iceberg table(s) that now live in another catalog: - - exposes[orders_iceberg]: location.catalog lakekeeper + error: iceberg catalog move blocked — this contract's OpenTofu state holds 1 Snowflake EXTERNAL VOLUME(s) that this contract's configuration no longer declares: snowflake_external_volume.sales_orders_lake_vol_FLUID_SALES_ORDERS_LAKE_VOL -forge-cli no longer creates an EXTERNAL VOLUME for an Iceberg table in a catalog Snowflake does not manage (dbt now writes it as an externally cataloged table, not a Snowflake-managed one on a volume), so applying now would plan to DROP these volumes, and any Snowflake-managed Iceberg table an earlier dbt run wrote onto one still uses it. +The volume is named for the contract, not for an expose, so the state does not say which expose it was created for. Possible causes: this contract was applied by a forge-cli release that gave an Iceberg table in a catalog Snowflake does not manage an EXTERNAL VOLUME (this release gives it none, and dbt writes it as an externally cataloged table); or this change removed a Snowflake-managed Iceberg expose, or moved one to another catalog. Applying now would plan to DROP the volume, and any Snowflake-managed Iceberg table written onto it still uses it. + +Iceberg exposes whose catalog earlier releases gave an EXTERNAL VOLUME: + + exposes[orders_iceberg]: location.catalog lakekeeper Drop each from this contract's state, then re-run apply: tofu -chdir=.fluid/iac/snowflake/sales_orders_lake state rm snowflake_external_volume.sales_orders_lake_vol_FLUID_SALES_ORDERS_LAKE_VOL -`tofu state rm` touches ZERO bytes of infrastructure: the resources stay in Snowflake, and only this contract's claim on them is released. Drop a volume by hand (DROP EXTERNAL VOLUME) only once no Iceberg table uses it. If the table belongs in Snowflake's own catalog, remove location.catalog (or set it to 'snowflake') instead. +`tofu state rm` touches ZERO bytes of infrastructure: the resources stay in Snowflake, and only this contract's claim on them is released. Drop a volume by hand (DROP EXTERNAL VOLUME) only once no Iceberg table uses it. If an Iceberg table belongs in Snowflake's own catalog, remove its location.catalog (or set it to 'snowflake') instead. remediation: [...] ``` +*([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* The message names both possible causes because the state cannot tell them apart: an earlier release that gave the volume to an expose in a catalog Snowflake does not manage, or a Snowflake-managed Iceberg expose that this change removed or moved to another catalog. The guard blocks the same applies as before. With #707 and #709 alone the message said forge-cli no longer creates such a volume, and listed the moved exposes as the tables it was held for. + 1. Run the printed command, `tofu -chdir=.fluid/iac/snowflake/ state rm snowflake_external_volume.`. It changes nothing in Snowflake. 2. Run `fluid apply` again. -3. Drop the volume by hand (`DROP EXTERNAL VOLUME`) only once no Iceberg table uses it. A Snowflake-managed table an earlier dbt run wrote onto it still does. +3. Drop the volume by hand (`DROP EXTERNAL VOLUME`) only once no Iceberg table uses it. A Snowflake-managed table written onto it still does. If the table belongs in Snowflake's own catalog, remove `location.catalog` (or set it to `snowflake`) instead. The volume is named per contract (`FLUID__VOL`), not per expose, so the Snowflake guard has no per-expose resource to check: the volume is flagged when it is in state, an expose moved, and the module no longer declares it. The guard finds the volume by the name derived from the contract id, not from the expose's `location`, so it also stops an upgrade whose edit changed `location.warehouse` to the catalog's warehouse name (`warehouse: analytics`) or removed `location.iam_role_arn`. ### Other commands - **dbt-bigquery.** An expose that names a catalog other than `bigquery` is left out of `catalogs.yml`, with a warning, instead of becoming a BigLake table. On `platform: gcp`, `fluid validate` accepts such a catalog's warehouse name and refuses a warehouse in another object store (`s3://`, `abfss://`). -- **Confluent Tableflow.** A `platform: confluent` expose publishes only to AWS Glue. A `location.catalog` other than `glue` is a `fluid validate` error, and the module creates no catalog integration for it. -- **`fluid policy compile`.** A Snowflake-managed Iceberg table compiles to Snowflake grants. A table in another catalog gets no Glue grant, and a warning to enforce access in that catalog. +- **Confluent Tableflow.** A `platform: confluent` expose publishes only to AWS Glue. A `location.catalog` other than `glue` is a `fluid validate` error, and the module creates no catalog integration for it. *([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* With no `location.catalog` the expose reads as `glue`, so dbt-snowflake's `catalogs.yml` writes `catalog_linked_database_type: glue` for it. +- **`fluid policy compile`.** A Snowflake-managed Iceberg table compiles to Snowflake grants. A table in another catalog gets no Glue grant, and a warning to enforce access in that catalog. *([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* A Confluent Tableflow expose compiles to its S3 bucket and the Glue table Tableflow publishes. The Tableflow module names that table for the topic (`location.topic`, else `location.table`, else the expose id) and publishes it into `location.database`, or, with none, into the database Tableflow names after the Kafka cluster id. For this binding, which sets no `database` (`fluid validate` warns about that): + + ```yaml + binding: + platform: confluent + format: iceberg + location: + environment_id: env-abc123 + kafka_cluster_id: lkc-xyz789 + topic: orders + bucket: acme-tableflow + region: eu-west-1 + confluent_role_arn: arn:aws:iam::123456789012:role/tableflow + ``` + + a `read` grant compiles to an `s3.bucket` binding on `acme-tableflow` and this Glue binding: + + ```json + { + "provider": "aws", + "resource_type": "glue.table", + "resource_id": "lkc-xyz789.orders", + "database": "lkc-xyz789", + "table": "orders", + "region": "eu-west-1", + "principal": "group:analysts@acme.com", + "actions": ["glue:GetTable", "glue:GetDatabase", "athena:StartQueryExecution", "athena:GetQueryResults"] + } + ``` + + With neither `database` nor `kafka_cluster_id`, the database is unknown: the expose gets the bucket binding only, and a warning. [`fluid policy compile`](../cli/policy-compile.md#warnings) prints each warning. - **`fluid validate`.** A crash inside the Iceberg sink, Confluent or Iceberg prerequisite check is an error ("the contract was NOT checked for ..."), not a note printed only with `--verbose`. Two `catalog: snowflake` exposes that derive one EXTERNAL VOLUME on different storage are refused at validate instead of failing `fluid apply` mid-emit. The full list is in [Iceberg catalog checks](../cli/validate.md#iceberg-catalog-checks). ### On 0.19.0 and earlier @@ -535,6 +711,8 @@ Each emitter classified `location.catalog` by hand, and they disagreed: - A Kafka Connect sink on Glue failed at startup, as shown above. - An unknown value fell back differently in each emitter, so a typo split one table across catalogs. - `fluid policy compile` compiled every Iceberg expose off GCP to AWS S3 and Glue grants, whatever its platform. +- A Kafka Connect or Debezium Server sink config forge-cli derived wrote the contract's first Iceberg expose, whatever the build's `outputs` named. +- A derived DynamoDB, JDBC or BigQuery sink config carried no warehouse unless `location.warehouse` set one, and a BigQuery one no `gcp.bigquery.project-id`. - A crash inside an Iceberg or Confluent check of `fluid validate` was printed only with `--verbose`, and the contract passed unchecked. - Neither streaming runner checked its Iceberg sink before it created anything. diff --git a/docs/advanced/typed-cli-errors.md b/docs/advanced/typed-cli-errors.md index a0797d8..f179bf2 100644 --- a/docs/advanced/typed-cli-errors.md +++ b/docs/advanced/typed-cli-errors.md @@ -184,6 +184,8 @@ The CLI builds each `doc` link from a fixed route table. A topic that is not in | `sovereignty`, `sovereignty#residency` | [Sovereignty](../concepts/sovereignty.md) | | `supply-chain` | [`fluid verify-signature`](../cli/verify-signature.md) | | `installation` | [Getting started](../getting-started/README.md) | +| `iceberg-catalog-move` | [Iceberg catalog-move guard](../cli/apply.md#iceberg-catalog-move-guard) *(unreleased, [forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709))* | +| `policy-compile#errors` | [`fluid policy compile`, Errors](../cli/policy-compile.md#errors) *(unreleased, [forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710))* | | anything else | [Production troubleshooting](./production-troubleshooting.md) | A few catalogued events land on a page that does not explain them. The CLI chooses these routes, so the explanations are here: diff --git a/docs/cli/apply.md b/docs/cli/apply.md index 76cc0e3..acce81f 100644 --- a/docs/cli/apply.md +++ b/docs/cli/apply.md @@ -286,19 +286,23 @@ The message lists the addresses and a `tofu -chdir=.fluid/iac/aws/ stat ```text ❌ iceberg_catalog_move_blocked [ERR_ICEBERG_CATALOG_MOVE_BLOCKED] kind: iceberg-catalog-move - error: iceberg catalog move blocked — this contract's OpenTofu state holds 1 Snowflake EXTERNAL VOLUME(s) for Iceberg table(s) that now live in another catalog: + error: iceberg catalog move blocked — this contract's OpenTofu state holds 1 Snowflake EXTERNAL VOLUME(s) that this contract's configuration no longer declares: + + snowflake_external_volume.sales_orders_lake_vol_FLUID_SALES_ORDERS_LAKE_VOL + +The volume is named for the contract, not for an expose, so the state does not say which expose it was created for. Possible causes: this contract was applied by a forge-cli release that gave an Iceberg table in a catalog Snowflake does not manage an EXTERNAL VOLUME (this release gives it none, and dbt writes it as an externally cataloged table); or this change removed a Snowflake-managed Iceberg expose, or moved one to another catalog. ... + +Iceberg exposes whose catalog earlier releases gave an EXTERNAL VOLUME: exposes[orders_iceberg]: location.catalog lakekeeper - snowflake_external_volume.sales_orders_lake_vol_FLUID_SALES_ORDERS_LAKE_VOL -... Drop each from this contract's state, then re-run apply: tofu -chdir=.fluid/iac/snowflake/sales_orders_lake state rm snowflake_external_volume.sales_orders_lake_vol_FLUID_SALES_ORDERS_LAKE_VOL ... ``` -The guard finds the volume by the name an earlier release derived from the contract id, so it also stops an upgrade whose edit changed `location.warehouse` to the catalog's warehouse name or removed `location.iam_role_arn`. Run the printed `tofu -chdir=.fluid/iac/snowflake/ state rm snowflake_external_volume.`, which changes nothing in Snowflake, then apply again. Drop the volume by hand (`DROP EXTERNAL VOLUME`) only once no Iceberg table uses it: a Snowflake-managed table an earlier dbt run wrote onto it still does. See [Upgrading a Snowflake contract that names another catalog](../advanced/source-aligned-acquisition.md#upgrading-a-snowflake-contract-that-names-another-catalog). +*([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* The volume is named per contract, so state cannot say which expose it was created for, and the message names both possible causes: an upgrade from a release that gave a volume to an expose in a catalog Snowflake does not manage, or a Snowflake-managed Iceberg expose this change removed or moved to another catalog. It blocks the same applies as before; with #707 and #709 alone it said forge-cli no longer creates the volume and listed the moved exposes as holding it. The guard finds the volume by the name derived from the contract id, so it also stops an upgrade whose edit changed `location.warehouse` to the catalog's warehouse name or removed `location.iam_role_arn`. Run the printed `tofu -chdir=.fluid/iac/snowflake/ state rm snowflake_external_volume.`, which changes nothing in Snowflake, then apply again. Drop the volume by hand (`DROP EXTERNAL VOLUME`) only once no Iceberg table uses it: a Snowflake-managed table written onto it still does. See [Upgrading a Snowflake contract that names another catalog](../advanced/source-aligned-acquisition.md#upgrading-a-snowflake-contract-that-names-another-catalog). *([forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709), unreleased)* When the guard cannot read the state, or its check fails, the apply goes on to plan, and logs an `iceberg_catalog_move_probe_skipped` WARNING with the reason and the exposes. The data-loss gate still stops a plan that destroys the resources; release them with `tofu state rm` rather than passing `--allow-data-loss`. On #707 alone a failed check was logged at debug level. diff --git a/docs/cli/generate-artifacts.md b/docs/cli/generate-artifacts.md index b7b534f..87d1943 100644 --- a/docs/cli/generate-artifacts.md +++ b/docs/cli/generate-artifacts.md @@ -70,7 +70,7 @@ fluid generate artifacts CONTRACT [--out PATH] [--emit KEYS] [--manifest PATH] [ | `odps-bitol` | ODPS-Bitol v1.0.0 product file under `odps-bitol/`, with an ODCS file per exposed port beside it | Schema vendored from `bitol-io/open-data-product-standard`. | | `opds` | OPDS v4.1 (LF/ODPI) product file under `opds/` (`.opds.json`) | `odps` is accepted as a deprecated alias of `opds` and warns. | | `schedule` | Airflow DAG files under `schedule/` | Emitted when `orchestration.engine` is set, or when a build declares a schedule trigger. See [Scheduled builds](#scheduled-builds). | -| `policies` | `policy/bindings.json`, the compiled IAM / GRANT bindings | | +| `policies` | `policy/bindings.json`, the compiled IAM / GRANT bindings | *([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* A crash inside the policy compiler fails the stage with `policy_compiler_crashed` (exit 1). See [`fluid policy compile`](./policy-compile.md#errors). | `builds[].pattern` (for example `hybrid-reference`) decides how the transformation runs and gates no emit key: a reference-only contract gets the same set as any other. diff --git a/docs/cli/policy-apply.md b/docs/cli/policy-apply.md index 09ffb50..9394f5c 100644 --- a/docs/cli/policy-apply.md +++ b/docs/cli/policy-apply.md @@ -87,5 +87,6 @@ fluid policy-apply runtime/policy/bindings.json --mode enforce - Provider and project are read from the first binding that sets them. `fluid policy compile` writes both from the contract's `binding.platform` and `binding.location`. A provider given by flag or environment overrides the file's. - Returns `0` for `ok` or `noop` results, `1` otherwise. +- *([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* Before it hands the bindings to the provider, the command prints each warning in the file's `warnings` array other than `No grants found in accessPolicy` to stderr, as a `policy_bindings_warning` log line at WARNING level, so a grant that [compiled to no binding](./policy-compile.md#warnings) shows in a job that runs apply apart from compile. The warnings do not change the exit code. - Compile bindings first with [`fluid policy compile`](./policy-compile.md); see [`fluid policy check`](./policy-check.md) for static linting of the access policy. - Both spellings share one argument set. Prefer `fluid policy apply` in new code; the hyphenated form is still registered in 0.18.1. diff --git a/docs/cli/policy-compile.md b/docs/cli/policy-compile.md index e43d188..5b6c22d 100644 --- a/docs/cli/policy-compile.md +++ b/docs/cli/policy-compile.md @@ -36,9 +36,50 @@ fluid policy compile contract.fluid.yaml --out build/bindings.json fluid policy-compile contract.fluid.yaml --env prod ``` +## Warnings + +*([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* Each compiler warning goes into the `warnings` array of the output file, and each one other than `No grants found in accessPolicy` is also printed to stderr, as a `policy_compile_warning` log line at WARNING level. Warnings leave the exit code at `0`. A warning can be the only sign that a grant compiled to no binding. A `read` grant on an AWS Iceberg expose that Lakekeeper catalogs compiles to the S3 bucket binding only, and prints: + +```bash +fluid policy compile contract.fluid.yaml +``` + +```text +{"time": "...", "level": "WARNING", "name": "fluid.cli", "message": "policy_compile_warning", "warning": "Iceberg expose 'orders' is cataloged in 'lakekeeper', not AWS Glue, so no table grant was compiled for group:analysts@acme.com ['read']. Enforce it in the 'lakekeeper' catalog's own access control.", "out": "runtime/policy/bindings.json"} +``` + +A contract with no `accessPolicy` grants gets `No grants found in accessPolicy` in the file only: no grant is left unenforced. [`fluid policy apply`](./policy-apply.md#notes) prints the file's warnings again. Before #710 the warnings were written only to the file. + +## Errors + +| Code | When | Exit | +| --- | --- | --- | +| `ERR_POLICY_COMPILER_CRASHED` | *([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* The policy compiler raised an exception. No bindings file is written; one from an earlier run is left as it was. | 1 | +| `ERR_POLICY_COMPILE_FAILED` | The command failed outside the compiler, for example because the contract or overlay is missing or does not parse, or the bindings file cannot be written. | 1 | + +`policy compile` does not validate the contract against its schema, so a value of the wrong type where the compiler reads (`accessPolicy`, its `grants`, a grant's `permissions`, an expose's `binding` and `binding.location`) reaches the compiler. When it crashes there, the error names each such value by path and expected type, never by its content. A grant written as a bare string: + +```yaml +accessPolicy: + grants: + - group:analysts@acme.com +``` + +```text +❌ policy_compiler_crashed [ERR_POLICY_COMPILER_CRASHED] + error: policy compile failed on contract values of the wrong type: accessPolicy.grants[0] is not of type 'object' + +💡 Suggestions: + • Run 'fluid validate ': policy compile reads accessPolicy and exposes without validating them against the contract schema + • If the contract validates, re-run with 'fluid --log-level DEBUG policy compile ' to see the compiler's traceback + +📖 Documentation: https://agenticstiger.github.io/forge_docs/cli/policy-compile.html#errors +``` + +When the schema finds no such error, the message is the compiler's own exception: `policy compiler failed: : `. The policy step of [`fluid generate artifacts`](./generate-artifacts.md) fails with the same error. Before #710 a crash wrote a bindings file with no bindings and the crash as a warning, and exited `0`, so the next stage read it as no grants to enforce. + ## Notes - Loads the contract with the requested env overlay and emits a JSON document with `bindings` and `warnings` arrays at the `--out` path. - The compiler embeds `provider` and `project` on each binding so [`fluid policy apply`](./policy-apply.md) can target the right account without extra flags. -- Compiler failures are caught and surfaced as warnings inside the output file — the command still exits `0` so downstream automation can inspect the warnings. - Pair with [`fluid policy check`](./policy-check.md) for static linting and [`fluid policy apply`](./policy-apply.md) for the stage-8 hand-off to the provider (it changes no permissions in 0.18.1). diff --git a/docs/cli/validate.md b/docs/cli/validate.md index 8b7afec..aa51ad2 100644 --- a/docs/cli/validate.md +++ b/docs/cli/validate.md @@ -292,7 +292,7 @@ With [forge-cli #707](https://github.com/Agenticstiger/forge-cli/pull/707) (unre ## Iceberg catalog checks ::: warning Not in a release yet -These checks come with [forge-cli #707](https://github.com/Agenticstiger/forge-cli/pull/707), which no release includes yet, and with the follow-up fixes in [forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709), marked *([forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709), unreleased)*. +These checks come with [forge-cli #707](https://github.com/Agenticstiger/forge-cli/pull/707), which no release includes yet, and with the follow-up fixes in [forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709), marked *([forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709), unreleased)*, and in [forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), marked *([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)*. ::: `binding.location.catalog` names the Iceberg catalog that owns an expose's table, and every emitter now reads it through one table. [Iceberg catalogs](../advanced/source-aligned-acquisition.md#iceberg-catalogs-location-catalog) has that table, how spellings fold, and a worked Lakekeeper example. `fluid validate` refuses a contract whose catalog the emitters would disagree about: @@ -300,6 +300,9 @@ These checks come with [forge-cli #707](https://github.com/Agenticstiger/forge-c - **An unknown catalog.** A `location.catalog` outside the table, on an Iceberg expose, is an error that lists the accepted spellings. On `platform: confluent` every value other than `glue` is an error instead, because the Tableflow module publishes only to AWS Glue. - **A streaming sink that cannot reach its catalog.** For a Kafka Connect build, or an embedded Debezium Server build, that writes an Iceberg expose: - every REST catalog (`rest`, `lakekeeper`, `polaris`, `unity`), `nessie` and `snowflake-managed` needs `location.uri` and `location.warehouse`; `jdbc` needs `uri`, and `hadoop` needs `warehouse`. On 0.19.0 and earlier only the literal `rest` was checked. + - *([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* `dynamodb` and `jdbc` need a warehouse: `location.warehouse`, one derived from `location.bucket` on `platform: aws` or `gcp`, or the sink's `warehouse` property in an override. `bigquery` needs `location.project`, or `gcp.bigquery.project-id` in an override. A value of only whitespace counts as missing. A `bigquery` expose with no `gs://` warehouse to derive is a warning. See [What a DynamoDB, JDBC or BigQuery sink needs](../advanced/source-aligned-acquisition.md#what-a-dynamodb-jdbc-or-bigquery-sink-needs). + - *([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* a warehouse that derives from a `location.bucket` written as a `{{ env.* }}` template, whose variable is unset or empty where `fluid validate` runs, is a warning that names the variable. The run preflight refuses the build when the variable is unset or empty where the sink config is derived. + - *([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* a build whose sink config forge-cli derives must name the Iceberg exposes it writes in its `outputs`. A Kafka Connect build writes exactly one, so outputs that name none or several of them are an error, and so is a build with no outputs in a contract with several. An embedded Debezium Server build is an error when its outputs name none of them, or when the exposes it writes sit in more than one `location.database`, resolve to different catalogs, or are DynamoDB or JDBC exposes with different warehouses derived from `location.bucket`. See [Which exposes a streaming sink writes](../advanced/source-aligned-acquisition.md#which-exposes-a-streaming-sink-writes). - a `sink.catalog` that names another catalog than the expose is an error, because dbt and the modules read only the expose. - *([forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709), unreleased)* a `sink_connector_config`, an `iceberg_catalog_overrides` entry or an embedded Debezium Server `server.sink.config` whose `type` or `catalog-impl` selects another catalog than the expose's is an error, on every platform: the sink would write one catalog while dbt and the modules read the other. They are compared as the catalog the worker builds, so `type=rest` matches `rest`, `lakekeeper`, `polaris`, `unity` and `snowflake-managed`. Declare the catalog the sink writes to in `location.catalog`; a REST endpoint that fronts Glue (Glue's Iceberg REST endpoint) is `catalog: rest`. A `catalog: bigquery` expose on `platform: gcp` whose hand-written config sets `iceberg.catalog.type: rest`: ```text @@ -315,7 +318,7 @@ These checks come with [forge-cli #707](https://github.com/Agenticstiger/forge-c A crash inside the Iceberg sink, Confluent or Iceberg prerequisite check now fails validation, with an error such as `Iceberg sink check could not run (...); the contract was NOT checked for streaming-sink defects`. On 0.19.0 and earlier the crash was printed only with `--verbose`, and the contract passed unchecked. -The Kafka Connect runner and the embedded Debezium Server runner run the same streaming-sink checks before they create anything, so a contract that skipped `fluid validate` still fails before any Connect REST call or `application.properties` write. *([forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709), unreleased)* That holds for a build that declares `sink.format: iceberg`, with a derived or a hand-written sink config, and for an embedded Debezium Server build that derives its sink. On #707 alone a hand-written config was not checked by the runners. [What each kind produces](../advanced/source-aligned-acquisition.md#what-each-kind-produces) lists the builds. +The Kafka Connect runner and the embedded Debezium Server runner run the same streaming-sink checks before they create anything, so a contract that skipped `fluid validate` still fails before any Connect REST call or `application.properties` write. *([forge-cli #709](https://github.com/Agenticstiger/forge-cli/pull/709), unreleased)* That holds for a build that declares `sink.format: iceberg`, with a derived or a hand-written sink config, and for an embedded Debezium Server build that derives its sink. On #707 alone a hand-written config was not checked by the runners. *([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* The one difference is a `location.bucket` template whose variable is unset or empty: a warning here, since the variable may be set where the sink runs, and an error in the runner when it is unset or empty there and the sink would use the warehouse. [What each kind produces](../advanced/source-aligned-acquisition.md#what-each-kind-produces) lists the builds. ## GCP binding checks (since 0.15.0) diff --git a/docs/concepts/builds-exposes-bindings.md b/docs/concepts/builds-exposes-bindings.md index 076f188..beaaf9b 100644 --- a/docs/concepts/builds-exposes-bindings.md +++ b/docs/concepts/builds-exposes-bindings.md @@ -95,6 +95,7 @@ Four rules, each measured on 0.18.1: - **Acquisition builds** write to the first expose whose `exposeId` appears in the build's `outputs`, and to `exposes[0]` when it names none. As of 0.18.1 their schema-drift check still compares the source with `exposes[0]`'s declared schema, so a second build whose source differs fails with `source schema drift detected` when `exposes[0]` has `schemaPolicy: discover_and_freeze`, the policy `fluid init --discover` writes. With `evolve_safe` on `exposes[0]` the second build lands in its own expose. - **Embedded-SQL builds on the local DuckDB engine** write one result, to `exposes[0]`, whatever `outputs` says. Naming a second expose prints a warning and writes nothing there. +- **Streaming Iceberg sinks.** *([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* A Kafka Connect or embedded Debezium Server build whose Iceberg sink config forge-cli derives writes the Iceberg exposes its `outputs` name: exactly one for Kafka Connect, and for Debezium Server one or more that share a database and a catalog. `fluid validate` refuses a build whose outputs its sink cannot write that way. See [Which exposes a streaming sink writes](../advanced/source-aligned-acquisition.md#which-exposes-a-streaming-sink-writes). As of 0.18.1, two embedded-SQL builds in one contract therefore share a destination. Each names its own expose in `outputs`; both run; and the second overwrites the first's file: diff --git a/docs/providers/aws.md b/docs/providers/aws.md index 5dd4f9f..16290f6 100644 --- a/docs/providers/aws.md +++ b/docs/providers/aws.md @@ -788,7 +788,7 @@ fluid policy-compile contract.fluid.yaml --out runtime/policy/bindings.json } ``` -Each principal gets one entry for the bucket and one for the Glue table. *(forge-cli [#707](https://github.com/Agenticstiger/forge-cli/pull/707), unreleased)* An Iceberg expose in a catalog other than Glue gets the bucket entry only, and a warning (`Iceberg expose 'orders' is cataloged in 'lakekeeper', not AWS Glue, so no table grant was compiled ...`). The permissions map to two action sets: any of `write`, `insert`, `update` or `delete` selects the write set (`s3:PutObject`, `s3:DeleteObject`, `s3:GetObject`, `s3:ListBucket`, and `glue:CreateTable`, `glue:UpdateTable`, `glue:DeleteTable`); every other permission selects the read set shown above. Placeholders in the bucket stay literal, and the principals are written as the contract names them. +Each principal gets one entry for the bucket and one for the Glue table. *(forge-cli [#707](https://github.com/Agenticstiger/forge-cli/pull/707), unreleased)* An Iceberg expose in a catalog other than Glue gets the bucket entry only, and a warning (`Iceberg expose 'orders' is cataloged in 'lakekeeper', not AWS Glue, so no table grant was compiled ...`). *([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* The command also prints the warning, at WARNING level; see [Warnings](../cli/policy-compile.md#warnings). The permissions map to two action sets: any of `write`, `insert`, `update` or `delete` selects the write set (`s3:PutObject`, `s3:DeleteObject`, `s3:GetObject`, `s3:ListBucket`, and `glue:CreateTable`, `glue:UpdateTable`, `glue:DeleteTable`); every other permission selects the read set shown above. Placeholders in the bucket stay literal, and the principals are written as the contract names them. `fluid policy-apply` enforces nothing on AWS. As of 0.18.1 the aws provider has no standalone policy applier, and the command exits 0: diff --git a/docs/providers/gcp.md b/docs/providers/gcp.md index 805fc2d..fe319de 100644 --- a/docs/providers/gcp.md +++ b/docs/providers/gcp.md @@ -338,9 +338,23 @@ Once the expose names its catalog, a hand-written sink config must select that s An expose that no streaming sink writes may still leave `location.catalog` out; it is a BigLake table. With `catalog: bigquery`, a Kafka Connect build whose sink config sends `iceberg.catalog.type=bigquery` (the derived config does) draws a warning, since the published Apache Iceberg Kafka Connect sink (1.9.2) cannot load that type, which Iceberg adds in 1.10: ```text - 1. iceberg sink (build 'stream_events'): the sink config sets iceberg.catalog.type=bigquery, which the published Apache Iceberg Kafka Connect sink (1.9.2 on Confluent Hub) cannot load: Iceberg's CatalogUtil gains the bigquery type in 1.10. Run a sink built from Iceberg >= 1.10, or the connector fails at start + 1. iceberg sink (build 'stream_events'): the sink config sets iceberg.catalog.type=bigquery, which the published Apache Iceberg Kafka Connect sink (1.9.2 on Confluent Hub) cannot load: Iceberg's CatalogUtil gains the bigquery type in 1.10, so on that sink the connector fails at start. Run a sink built from Iceberg >= 1.10 ``` +**What a streaming sink into BigLake needs.** *([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* Iceberg's `BigQueryMetastoreCatalog` refuses to start without `gcp.bigquery.project-id`, so a `catalog: bigquery` expose that a streaming sink writes needs `location.project`, which becomes that property. Without it, or with only whitespace, `fluid validate` and the run refuse the build, unless an override sets the property. `location.region` becomes `gcp.bigquery.location`, and the derived config sets no `client.region`. The warehouse is the `gs://` storage dbt-bigquery and this module use: a `gs://` `location.warehouse`, else `gs://` plus `location.path` when it is set. For the `events` binding above with `catalog: bigquery` and `region: europe-west1` added, the derived Kafka Connect config carries these catalog keys: + +```json +{ + "iceberg.catalog.type": "bigquery", + "iceberg.catalog.warehouse": "gs://my-lake/products/events", + "iceberg.catalog.io-impl": "org.apache.iceberg.gcp.gcs.GCSFileIO", + "iceberg.catalog.gcp.bigquery.project-id": "my-project-id", + "iceberg.catalog.gcp.bigquery.location": "europe-west1" +} +``` + +With no warehouse to derive, the config sets none and `fluid validate` warns: a Kafka Connect sink with auto-create on cannot create tables, and embedded Debezium Server does not boot without `debezium.sink.iceberg.warehouse`. For a bucket written as a `{{ env.* }}` template, a variable that is unset or empty where `fluid validate` runs draws a warning. One that is unset or empty where the sink config is derived makes the run refuse the build, except on a Kafka Connect sink without auto-create, which reads the warehouse only to create tables: it warns and pushes no warehouse. See [What a DynamoDB, JDBC or BigQuery sink needs](../advanced/source-aligned-acquisition.md#what-a-dynamodb-jdbc-or-bigquery-sink-needs). + --- ## Loading data diff --git a/docs/providers/snowflake.md b/docs/providers/snowflake.md index c85095a..c4b7299 100644 --- a/docs/providers/snowflake.md +++ b/docs/providers/snowflake.md @@ -624,12 +624,16 @@ Since `0.14.0`, `fluid validate` errors when an Iceberg expose is missing one of ```text ❌ iceberg_catalog_move_blocked [ERR_ICEBERG_CATALOG_MOVE_BLOCKED] kind: iceberg-catalog-move - error: iceberg catalog move blocked — this contract's OpenTofu state holds 1 Snowflake EXTERNAL VOLUME(s) for Iceberg table(s) that now live in another catalog: + error: iceberg catalog move blocked — this contract's OpenTofu state holds 1 Snowflake EXTERNAL VOLUME(s) that this contract's configuration no longer declares: + + snowflake_external_volume.sales_orders_lake_vol_FLUID_SALES_ORDERS_LAKE_VOL + +The volume is named for the contract, not for an expose, so the state does not say which expose it was created for. Possible causes: this contract was applied by a forge-cli release that gave an Iceberg table in a catalog Snowflake does not manage an EXTERNAL VOLUME (this release gives it none, and dbt writes it as an externally cataloged table); or this change removed a Snowflake-managed Iceberg expose, or moved one to another catalog. ... + +Iceberg exposes whose catalog earlier releases gave an EXTERNAL VOLUME: exposes[orders_iceberg]: location.catalog lakekeeper - snowflake_external_volume.sales_orders_lake_vol_FLUID_SALES_ORDERS_LAKE_VOL -... Drop each from this contract's state, then re-run apply: tofu -chdir=.fluid/iac/snowflake/sales_orders_lake state rm snowflake_external_volume.sales_orders_lake_vol_FLUID_SALES_ORDERS_LAKE_VOL @@ -638,7 +642,9 @@ Drop each from this contract's state, then re-run apply: 1. Run the printed command, in the form `tofu -chdir=.fluid/iac/snowflake/ state rm snowflake_external_volume.`. It changes nothing in Snowflake; it releases this contract's claim on the volume. 2. Run `fluid apply` again. -3. Drop the volume by hand (`DROP EXTERNAL VOLUME`) only once no Iceberg table uses it. A Snowflake-managed table an earlier dbt run wrote onto it still does. +3. Drop the volume by hand (`DROP EXTERNAL VOLUME`) only once no Iceberg table uses it. A Snowflake-managed table written onto it still does. + +*([forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710), unreleased)* The volume is named per contract, so state cannot say which expose it was created for. The message names both possible causes: an upgrade from a release that gave a volume to an expose in a catalog Snowflake does not manage, or a Snowflake-managed Iceberg expose that this change removed or moved to another catalog. It blocks the same applies as before; with #707 and #709 alone it said forge-cli no longer creates the volume and listed the moved exposes as holding it. If the table belongs in Snowflake's own catalog, remove `location.catalog` (or set it to `snowflake`) instead. A volume you name in `binding.icebergConfig.properties.external_volume` is never flagged, because no release created it. The guard finds the volume by the name derived from the contract id, so it also stops an upgrade whose edit changed `location.warehouse` to the catalog's warehouse name or removed `location.iam_role_arn`. The same guard covers Glue resources on AWS; see [Iceberg catalog-move guard](../cli/apply.md#iceberg-catalog-move-guard). From 250c55734d68f27edea4870736fe830fcc884f68 Mon Sep 17 00:00:00 2001 From: Speculator55005 <50082482+fas89@users.noreply.github.com> Date: Fri, 9 Oct 2026 02:10:30 +0200 Subject: [PATCH 2/2] docs(errors): catalogue iceberg_catalog_move_blocked --- docs/advanced/error-codes.md | 9 +++++++++ 1 file changed, 9 insertions(+) diff --git a/docs/advanced/error-codes.md b/docs/advanced/error-codes.md index 5dfff88..9564cbf 100644 --- a/docs/advanced/error-codes.md +++ b/docs/advanced/error-codes.md @@ -39,6 +39,7 @@ The route table sends an event to one of these pages, and the entries below name | [`fluid providers`](../cli/providers.md) | The provider events | | [`fluid secrets`](../cli/secrets.md) | `copilot_missing_llm_api_key` | | [Sovereignty](../concepts/sovereignty.md) | The policy and sovereignty events, except `policy_compiler_crashed` | +| [`fluid apply`, Iceberg catalog move guard](../cli/apply.md#iceberg-catalog-move-guard) | `iceberg_catalog_move_blocked` | | [`fluid policy compile`, Errors](../cli/policy-compile.md#errors) | `policy_compiler_crashed` *(unreleased, [forge-cli #710](https://github.com/Agenticstiger/forge-cli/pull/710))* | | [`fluid verify-signature`](../cli/verify-signature.md) | The signing events | | [Getting started](../getting-started/README.md) | `opentofu_engine_install_failed` | @@ -276,6 +277,14 @@ The state refusals (`state_shared_with_another_provider`, `state_migration_ambig - If the resources should stay where they are, set the binding's location.region to the region the error names - If they should move, empty and remove them there first (tofu destroy in the state directory the error names, with AWS_REGION set to the old region), then apply again +### iceberg_catalog_move_blocked + +`ERR_ICEBERG_CATALOG_MOVE_BLOCKED`. The documentation link lands on [`fluid apply`, Iceberg catalog move guard](../cli/apply.md#iceberg-catalog-move-guard). + +- Run the printed `tofu state rm` commands: they release the resources from this contract's OpenTofu state and touch nothing in the cloud +- Then re-run fluid apply +- The released Glue database/table or Snowflake EXTERNAL VOLUME stays in place; delete it by hand only if nothing else uses it + ## Generate ### generate_iac_failed