From b52c85f676f1d1566ccb36b27787ba7f43ca23c3 Mon Sep 17 00:00:00 2001
From: Talat Uyarer
Date: Tue, 30 Sep 2025 00:46:18 -0700
Subject: [PATCH 01/18] create catalog and iceberg table details
---
.../dsls/sql/extensions/create-catalog.md | 258 ++++++++++++++++++
.../sql/extensions/create-external-table.md | 171 ++++++++++++
2 files changed, 429 insertions(+)
create mode 100644 website/www/site/content/en/documentation/dsls/sql/extensions/create-catalog.md
diff --git a/website/www/site/content/en/documentation/dsls/sql/extensions/create-catalog.md b/website/www/site/content/en/documentation/dsls/sql/extensions/create-catalog.md
new file mode 100644
index 000000000000..0b71330efd8c
--- /dev/null
+++ b/website/www/site/content/en/documentation/dsls/sql/extensions/create-catalog.md
@@ -0,0 +1,258 @@
+---
+type: languages
+title: "Beam SQL extension: CREATE CATALOG Statement"
+aliases:
+ - /documentation/dsls/sql/create-catalog/
+ - /documentation/dsls/sql/statements/create-catalog/
+---
+
+
+# Beam SQL extensions: CREATE CATALOG
+
+Beam SQL's `CREATE CATALOG` statement creates and registers a catalog that manages metadata for external data sources. Catalogs provide a unified interface for accessing different types of data stores and enable features like schema management, table discovery, and cross-catalog queries.
+
+Currently, Beam SQL supports the **Apache Iceberg** catalog type, which provides access to Iceberg tables with full ACID transaction support, schema evolution, and time travel capabilities.
+
+## Syntax
+
+```
+CREATE CATALOG [ IF NOT EXISTS ] catalogName
+TYPE catalogType
+[PROPERTIES (propertyKey = propertyValue [, propertyKey = propertyValue ]*)]
+```
+
+* `IF NOT EXISTS`: Optional. If the catalog is already registered, Beam SQL
+ ignores the statement instead of returning an error.
+* `catalogName`: The case sensitive name of the catalog to create and register,
+ specified as an [Identifier](/documentation/dsls/sql/calcite/lexical#identifiers).
+* `catalogType`: The type of catalog to create. Currently supported values:
+ * `iceberg`: Apache Iceberg catalog
+* `PROPERTIES`: Optional. Key-value pairs for catalog-specific configuration.
+ Each property is specified as `'key' = 'value'` with string literals.
+
+## Apache Iceberg Catalog
+
+The Iceberg catalog provides access to [Apache Iceberg](https://iceberg.apache.org/) tables, which are high-performance table formats for huge analytic datasets.
+
+### Syntax
+
+```
+CREATE CATALOG [ IF NOT EXISTS ] catalogName
+TYPE iceberg
+PROPERTIES (
+ 'catalog-impl' = 'catalogImplementation',
+ 'warehouse' = 'warehouseLocation'
+ [, additionalProperties...]
+)
+```
+
+### Required Properties
+
+* `catalog-impl`: The Iceberg catalog implementation class. Common values:
+ * `org.apache.iceberg.hadoop.HadoopCatalog`: For Hadoop-compatible storage (HDFS, S3, GCS, etc.)
+ * `org.apache.iceberg.gcp.bigquery.BigQueryMetastoreCatalog`: For BigQuery integration
+ * `org.apache.iceberg.jdbc.JdbcCatalog`: For JDBC-based metadata storage
+ * `org.apache.iceberg.rest.RESTCatalog`: For REST-based catalog access
+* `warehouse`: The root location where Iceberg tables and metadata are stored.
+ Format depends on the storage system:
+ * **Local filesystem**: `file:///path/to/warehouse`
+ * **HDFS**: `hdfs://namenode:port/path/to/warehouse`
+ * **S3**: `s3://bucket-name/path/to/warehouse`
+ * **Google Cloud Storage**: `gs://bucket-name/path/to/warehouse`
+
+### Optional Properties
+
+The available optional properties depend on the catalog implementation:
+
+#### Hadoop Catalog Properties
+
+* `io-impl`: The file I/O implementation class. Common values:
+ * `org.apache.iceberg.hadoop.HadoopFileIO`: For Hadoop-compatible storage
+ * `org.apache.iceberg.aws.s3.S3FileIO`: For S3 storage
+ * `org.apache.iceberg.gcp.gcs.GCSFileIO`: For Google Cloud Storage
+* `hadoop.*`: Any Hadoop configuration property (e.g., `hadoop.fs.s3a.access.key`)
+
+#### BigQuery Metastore Catalog Properties
+
+* `io-impl`: Must be `org.apache.iceberg.gcp.gcs.GCSFileIO` for GCS storage
+* `gcp_project`: Google Cloud Project ID
+* `gcp_region`: Google Cloud region (e.g., `us-central1`)
+* `gcp_location`: Alternative to `gcp_region` for specifying location
+
+#### JDBC Catalog Properties
+
+* `uri`: JDBC connection URI
+* `jdbc.user`: Database username
+* `jdbc.password`: Database password
+* `jdbc.driver`: JDBC driver class name
+
+### Examples
+
+#### Hadoop Catalog with Local Storage
+
+```sql
+CREATE CATALOG my_iceberg_catalog
+TYPE iceberg
+PROPERTIES (
+ 'catalog-impl' = 'org.apache.iceberg.hadoop.HadoopCatalog',
+ 'warehouse' = 'file:///tmp/iceberg-warehouse'
+)
+```
+
+#### Hadoop Catalog with S3 Storage
+
+```sql
+CREATE CATALOG s3_iceberg_catalog
+TYPE iceberg
+PROPERTIES (
+ 'catalog-impl' = 'org.apache.iceberg.hadoop.HadoopCatalog',
+ 'warehouse' = 's3://my-bucket/iceberg-warehouse',
+ 'io-impl' = 'org.apache.iceberg.aws.s3.S3FileIO',
+ 'hadoop.fs.s3a.access.key' = 'your-access-key',
+ 'hadoop.fs.s3a.secret.key' = 'your-secret-key'
+)
+```
+
+#### BigQuery Metastore Catalog
+
+```sql
+CREATE CATALOG bigquery_iceberg_catalog
+TYPE iceberg
+PROPERTIES (
+ 'catalog-impl' = 'org.apache.iceberg.gcp.bigquery.BigQueryMetastoreCatalog',
+ 'io-impl' = 'org.apache.iceberg.gcp.gcs.GCSFileIO',
+ 'warehouse' = 'gs://my-bucket/iceberg-warehouse',
+ 'gcp_project' = 'my-gcp-project',
+ 'gcp_region' = 'us-central1'
+)
+```
+
+#### JDBC Catalog
+
+```sql
+CREATE CATALOG jdbc_iceberg_catalog
+TYPE iceberg
+PROPERTIES (
+ 'catalog-impl' = 'org.apache.iceberg.jdbc.JdbcCatalog',
+ 'uri' = 'jdbc:postgresql://localhost:5432/iceberg_metadata',
+ 'jdbc.user' = 'iceberg_user',
+ 'jdbc.password' = 'iceberg_password',
+ 'jdbc.driver' = 'org.postgresql.Driver',
+ 'warehouse' = 's3://my-bucket/iceberg-warehouse'
+)
+```
+
+## Using Catalogs
+
+After creating a catalog, you can use it to manage databases and tables:
+
+### Switch to a Catalog
+
+```sql
+USE CATALOG catalogName
+```
+
+### Create and Use a Database
+
+```sql
+-- Create a database (namespace)
+CREATE DATABASE my_database
+
+-- Use the database
+USE DATABASE my_database
+```
+
+### Create Tables in the Catalog
+
+Once you've switched to a catalog and database, you can create tables:
+
+```sql
+-- Switch to your catalog and database
+USE CATALOG my_iceberg_catalog
+USE DATABASE my_database
+
+-- Create an Iceberg table
+CREATE EXTERNAL TABLE users (
+ id BIGINT,
+ username VARCHAR,
+ email VARCHAR,
+ created_at TIMESTAMP
+)
+TYPE iceberg
+```
+
+## Catalog Management
+
+### List Available Catalogs
+
+```sql
+SHOW CATALOGS
+```
+
+### Drop a Catalog
+
+```sql
+DROP CATALOG [ IF EXISTS ] catalogName
+```
+
+## Best Practices
+
+### Security
+
+* **Credentials**: Store sensitive credentials (access keys, passwords) in secure configuration systems rather than hardcoding them in SQL statements
+* **IAM Roles**: Use IAM roles and service accounts when possible instead of access keys
+* **Network Security**: Ensure proper network access controls for your storage systems
+
+### Performance
+
+* **Warehouse Location**: Choose a warehouse location that's geographically close to your compute resources
+* **Partitioning**: Use appropriate partitioning strategies for your data access patterns
+* **File Formats**: Iceberg automatically manages file formats, but consider compression settings for your use case
+
+### Monitoring
+
+* **Catalog Health**: Monitor catalog connectivity and performance
+* **Storage Usage**: Track warehouse storage usage and implement lifecycle policies
+* **Query Performance**: Monitor query performance and optimize table schemas as needed
+
+## Troubleshooting
+
+### Common Issues
+
+#### Catalog Creation Fails
+
+* **Check Dependencies**: Ensure all required Iceberg dependencies are available in your classpath
+* **Verify Properties**: Double-check that all required properties are provided and correctly formatted
+* **Storage Access**: Ensure your compute environment has access to the specified warehouse location
+
+#### Table Operations Fail
+
+* **Catalog Context**: Make sure you're using the correct catalog with `USE CATALOG`
+* **Database Context**: Ensure you're in the correct database with `USE DATABASE`
+* **Permissions**: Verify that your credentials have the necessary permissions for the storage system
+
+#### Performance Issues
+
+* **Partitioning**: Review your table partitioning strategy
+* **File Size**: Check if files are too large or too small for your use case
+* **Compression**: Consider adjusting compression settings for your data types
+
+### Getting Help
+
+For more information about Apache Iceberg:
+
+* [Apache Iceberg Documentation](https://iceberg.apache.org/docs/)
+* [Iceberg Catalog Implementations](https://iceberg.apache.org/docs/latest/configuration/)
+* [Beam SQL Documentation](/documentation/dsls/sql/)
diff --git a/website/www/site/content/en/documentation/dsls/sql/extensions/create-external-table.md b/website/www/site/content/en/documentation/dsls/sql/extensions/create-external-table.md
index ad6ba66beb20..62e30b515e44 100644
--- a/website/www/site/content/en/documentation/dsls/sql/extensions/create-external-table.md
+++ b/website/www/site/content/en/documentation/dsls/sql/extensions/create-external-table.md
@@ -72,6 +72,8 @@ tableElement: columnName fieldType [ NOT NULL ]
* `kafka`
* `parquet`
* `text`
+ * `iceberg`
+ * `datagen`
* `location`: The I/O specific location of the underlying table, specified as
a [String
Literal](/documentation/dsls/sql/calcite/lexical/#string-literals).
@@ -748,6 +750,175 @@ TYPE text
LOCATION '/home/admin/orders'
```
+## Apache Iceberg
+
+Beam SQL supports reading from and writing to [Apache Iceberg](https://iceberg.apache.org/) tables. Iceberg is a high-performance table format for huge analytic datasets that provides ACID transactions, schema evolution, and time travel capabilities.
+
+**Prerequisites**: Before creating Iceberg tables, you must first create an Iceberg catalog. See the [CREATE CATALOG](/documentation/dsls/sql/extensions/create-catalog/) documentation for details.
+
+### Syntax
+
+```
+CREATE EXTERNAL TABLE [ IF NOT EXISTS ] tableName (tableElement [, tableElement ]*)
+TYPE iceberg
+[PARTITIONED BY (partitionField [, partitionField ]*)]
+[TBLPROPERTIES tblProperties]
+```
+
+* `tableName`: The case sensitive name of the table to create and register.
+* `tableElement`: `columnName` `fieldType` `[ NOT NULL ]`
+ * `columnName`: The case sensitive name of the column.
+ * `fieldType`: The field's type, specified as one of the following types:
+ * `simpleType`: `TINYINT`, `SMALLINT`, `INTEGER`, `BIGINT`, `FLOAT`,
+ `DOUBLE`, `DECIMAL`, `BOOLEAN`, `DATE`, `TIME`, `TIMESTAMP`, `CHAR`,
+ `VARCHAR`, `BINARY`, `VARBINARY`
+ * `MAP`
+ * `ARRAY`
+ * `ROW`
+ * `NOT NULL`: Optional. Indicates that the column is not nullable.
+* `PARTITIONED BY`: Optional. Specifies partition fields for the table. Supports various partition functions:
+ * `identity(columnName)`: Identity partitioning
+ * `bucket(columnName, numBuckets)`: Bucket partitioning
+ * `truncate(columnName, width)`: Truncate partitioning
+ * `year(columnName)`, `month(columnName)`, `day(columnName)`, `hour(columnName)`: Time-based partitioning
+* `TBLPROPERTIES`: Optional. JSON object with table-specific configuration:
+ * `triggering_frequency_seconds`: For streaming pipelines, specifies how often to commit snapshots (in seconds).
+
+### Read Mode
+
+Beam SQL supports reading from existing Iceberg tables. The connector automatically infers the table schema from the Iceberg table metadata and supports:
+
+* **Predicate push-down**: Filters are pushed down to the Iceberg scan level for better performance
+* **Projection push-down**: Only requested columns are read from the table
+* **Schema evolution**: Automatically handles schema changes in the underlying Iceberg table
+
+### Write Mode
+
+Beam SQL supports writing to Iceberg tables with the following features:
+
+* **Automatic table creation**: If the table doesn't exist, it will be created with the specified schema
+* **ACID transactions**: All writes are committed as atomic transactions
+* **Schema validation**: Ensures data matches the table schema
+* **Partitioning**: Supports writing to partitioned tables
+* **Streaming support**: For streaming pipelines, commits are performed at regular intervals
+
+### Schema
+
+Beam SQL types map to Iceberg types as follows:
+
+
+
+ | Beam SQL Type
+ |
+ Iceberg Type
+ |
+
+
+ | TINYINT, SMALLINT, INTEGER, BIGINT
+ |
+ INTEGER, LONG
+ |
+
+
+ | FLOAT, DOUBLE
+ |
+ FLOAT, DOUBLE
+ |
+
+
+ | DECIMAL
+ |
+ DECIMAL
+ |
+
+
+ | BOOLEAN
+ |
+ BOOLEAN
+ |
+
+
+ | DATE, TIME, TIMESTAMP
+ |
+ DATE, TIME, TIMESTAMP
+ |
+
+
+ | CHAR, VARCHAR
+ |
+ STRING
+ |
+
+
+ | BINARY, VARBINARY
+ |
+ BINARY
+ |
+
+
+ | ARRAY
+ |
+ LIST
+ |
+
+
+ | MAP
+ |
+ MAP
+ |
+
+
+ | ROW
+ |
+ STRUCT
+ |
+
+
+
+### Examples
+
+#### Basic Table Creation
+
+```sql
+CREATE EXTERNAL TABLE users (
+ id BIGINT,
+ username VARCHAR,
+ email VARCHAR,
+ created_at TIMESTAMP
+)
+TYPE iceberg
+```
+
+#### Partitioned Table
+
+```sql
+CREATE EXTERNAL TABLE events (
+ event_id BIGINT,
+ user_id BIGINT,
+ event_type VARCHAR,
+ event_timestamp TIMESTAMP,
+ data ROW
+)
+TYPE iceberg
+PARTITIONED BY (
+ 'bucket(user_id, 10)',
+ 'day(event_timestamp)',
+ 'event_type'
+)
+```
+
+#### Streaming Table with Commit Frequency
+
+```sql
+CREATE EXTERNAL TABLE streaming_events (
+ id BIGINT,
+ message VARCHAR,
+ timestamp TIMESTAMP
+)
+TYPE iceberg
+TBLPROPERTIES '{"triggering_frequency_seconds": 60}'
+```
+
## DataGen
The **DataGen** connector allows for creating tables based on in-memory data generation. This is useful for developing and testing queries locally without requiring access to external systems. The DataGen connector is built-in; no additional dependencies are required.It is available for Beam 2.67.0+
From 257cd5ae1d90223fbb9e8c862086d8114cbd4389 Mon Sep 17 00:00:00 2001
From: Ahmed Abualsaud
Date: Fri, 6 Feb 2026 16:47:39 -0500
Subject: [PATCH 02/18] add sql ddl website documentation
---
.../provider/iceberg/IcebergMetastore.java | 8 +-
.../meta/provider/iceberg/IcebergTable.java | 19 +-
.../provider/iceberg/PubsubToIcebergIT.java | 4 +-
.../sdk/io/iceberg/IcebergCatalogConfig.java | 4 +-
.../en/documentation/dsls/sql/ddl/alter.md | 75 +++++
.../en/documentation/dsls/sql/ddl/create.md | 120 ++++++++
.../en/documentation/dsls/sql/ddl/drop.md | 50 ++++
.../en/documentation/dsls/sql/ddl/overview.md | 53 ++++
.../en/documentation/dsls/sql/ddl/show.md | 83 ++++++
.../en/documentation/dsls/sql/ddl/use.md | 48 ++++
.../dsls/sql/extensions/create-catalog.md | 258 ------------------
.../sql/extensions/create-external-table.md | 171 ------------
.../partials/section-menu/en/sdks.html | 11 +
13 files changed, 468 insertions(+), 436 deletions(-)
create mode 100644 website/www/site/content/en/documentation/dsls/sql/ddl/alter.md
create mode 100644 website/www/site/content/en/documentation/dsls/sql/ddl/create.md
create mode 100644 website/www/site/content/en/documentation/dsls/sql/ddl/drop.md
create mode 100644 website/www/site/content/en/documentation/dsls/sql/ddl/overview.md
create mode 100644 website/www/site/content/en/documentation/dsls/sql/ddl/show.md
create mode 100644 website/www/site/content/en/documentation/dsls/sql/ddl/use.md
delete mode 100644 website/www/site/content/en/documentation/dsls/sql/extensions/create-catalog.md
diff --git a/sdks/java/extensions/sql/iceberg/src/main/java/org/apache/beam/sdk/extensions/sql/meta/provider/iceberg/IcebergMetastore.java b/sdks/java/extensions/sql/iceberg/src/main/java/org/apache/beam/sdk/extensions/sql/meta/provider/iceberg/IcebergMetastore.java
index b73aa25c7a2b..72e8988f41e2 100644
--- a/sdks/java/extensions/sql/iceberg/src/main/java/org/apache/beam/sdk/extensions/sql/meta/provider/iceberg/IcebergMetastore.java
+++ b/sdks/java/extensions/sql/iceberg/src/main/java/org/apache/beam/sdk/extensions/sql/meta/provider/iceberg/IcebergMetastore.java
@@ -22,6 +22,9 @@
import java.util.HashMap;
import java.util.Map;
+import java.util.stream.Collectors;
+
+import com.fasterxml.jackson.core.type.TypeReference;
import org.apache.beam.sdk.extensions.sql.TableUtils;
import org.apache.beam.sdk.extensions.sql.impl.TableName;
import org.apache.beam.sdk.extensions.sql.meta.BeamSqlTable;
@@ -59,8 +62,11 @@ public void createTable(Table table) {
getProvider(table.getType()).createTable(table);
} else {
String identifier = getIdentifier(table);
+ Map props = TableUtils.getObjectMapper()
+ .convertValue(table.getProperties(), new TypeReference
Tip: run SHOW CURRENT CATALOG to view the currently active Catalog.
Note: All subsequent DATABASE and TABLE commands will be executed under this Catalog, unless fully qualified.
@@ -164,7 +164,8 @@ to reference Databases directly without their fully-qualified names (e.g.
USE DATABASE sales_data;
{{< /highlight >}}
-Switch to a Database in a specified Catalog. Node: this also switches the default to that Catalog
+Switch to a Database in a specified Catalog.
+Note: this also switches the default to that Catalog
{{< highlight >}}
USE DATABASE other_catalog.sales_data;
{{< /highlight >}}
@@ -209,7 +210,7 @@ DROP DATABASE [ IF EXISTS ] database_name [ RESTRICT | CASCADE ];
{{< /highlight >}}
- - RESTRICT: (Default): Fails if the Database is not empty.
+ - RESTRICT (Default): Fails if the Database is not empty.
- CASCADE: Drops the Database and all tables contained within it. Use with caution.
@@ -217,10 +218,10 @@ DROP DATABASE [ IF EXISTS ] database_name [ RESTRICT | CASCADE ];
## Tables
The actual entity containing data, and is described by a schema. Some
-data sources also support applying a partition spec and attaching table-specific properties.
+data sources also let you apply a partition spec or attach table-specific properties.
{{< tab CREATE >}}
-Creates a new Table within the current Catalog and Database (default), or the specified Catalog and Database.
+Creates a new Table within the current Catalog and Database (default), or the Catalog and Database you specify.
{{< highlight >}}
CREATE EXTERNAL TABLE [ IF NOT EXISTS ] [ catalog. ][ db. ]table_name (
@@ -234,9 +235,9 @@ CREATE EXTERNAL TABLE [ IF NOT EXISTS ] [ catalog. ][ db. ]table_name (
[ TBLPROPERTIES 'properties_json_string' ];
{{< /highlight >}}
- - TYPE: the table type (e.g.
'iceberg', 'text', 'kafka').
+ - TYPE: the table type (for example,
'iceberg', 'text', 'kafka').
- PARTITIONED BY: an ordered list of fields describing the partition spec.
- - LOCATION: explicitly sets the location of the table (overriding the inferred
catalog.db.table_name location)
+ - LOCATION: explicitly sets the location of the table. This overrides the inferred
catalog.db.table_name location.
- TBLPROPERTIES: configuration properties used when creating the table or setting up its IO connection.
@@ -262,8 +263,8 @@ TBLPROPERTIES '{
- This creates an Iceberg table named
orders under the namespace sales_data, within the prod_iceberg catalog.
- The table is partitioned by
region_id, then by the day value of order_date (using Iceberg's hidden partitioning).
- - The table is created with the appropriate properties
"write.format.default" and "read.split.target-size". The Beam property "beam.write.triggering_frequency_seconds"
- - Beam properties (prefixed with
"beam.write." and "beam.read." are intended for the relevant IOs)
+ - The table is created with the appropriate properties
"write.format.default" and "read.split.target-size". The Beam property "beam.write.triggering_frequency_seconds" configures the Iceberg sink.
+ - Beam sink and source configuration properties are prefixed with
"beam.write." and "beam.read.", respectively.
{{< /tab >}}
{{< tab ALTER >}}
@@ -310,7 +311,7 @@ ALTER TABLE orders RESET ( 'write.target-file-size-bytes' );
{{< /tab >}}
{{< tab SHOW >}}
-Lists tables under the currently active database, or a specified database.
+Lists tables under the currently active database, or a database you specify.
{{< highlight >}}
SHOW TABLES [ ( FROM | IN )? [ catalog_name '.' ] database_name ] [ LIKE regex_pattern ]
diff --git a/website/www/site/content/en/documentation/io/built-in/iceberg.md b/website/www/site/content/en/documentation/io/built-in/iceberg.md
index 732d8308c0cf..3f3b15b57867 100644
--- a/website/www/site/content/en/documentation/io/built-in/iceberg.md
+++ b/website/www/site/content/en/documentation/io/built-in/iceberg.md
@@ -102,7 +102,7 @@ When you've met those prerequisites, start by setting up your catalog:
{{< code_sample "sdks/python/apache_beam/examples/snippets/snippets.py" biglake_public_catalog_props >}}
{{< /highlight >}}
{{< highlight yaml >}}
-catalog_props: &catalog_props
+catalog_props: &biglake_catalog_props
type: "rest"
uri: "https://biglake.googleapis.com/iceberg/v1/restcatalog"
warehouse: "gs://biglake-public-nyc-taxi-iceberg"
From 7698036b03a448fc44d0a8c98ad845a89fd28885 Mon Sep 17 00:00:00 2001
From: Ahmed Abualsaud
Date: Thu, 11 Jun 2026 09:31:37 -0400
Subject: [PATCH 12/18] fix tab switch issue
---
website/www/site/.hugo_build.lock | 0
website/www/site/assets/js/language-switch-v2.js | 1 -
2 files changed, 1 deletion(-)
create mode 100644 website/www/site/.hugo_build.lock
diff --git a/website/www/site/.hugo_build.lock b/website/www/site/.hugo_build.lock
new file mode 100644
index 000000000000..e69de29bb2d1
diff --git a/website/www/site/assets/js/language-switch-v2.js b/website/www/site/assets/js/language-switch-v2.js
index e3ee3e41f58a..ff6fdb0253a8 100644
--- a/website/www/site/assets/js/language-switch-v2.js
+++ b/website/www/site/assets/js/language-switch-v2.js
@@ -290,6 +290,5 @@ $(document).ready(function() {
Switcher({"name": "runner", "default": "direct"}).render();
Switcher({"name": "tab"}).render();
Switcher({"name": "shell", "default": "unix"}).render();
- Switcher({"name": "tab"}).render();
Switcher({"name": "version"}).render();
});
From 8310547d8b6e351d6e93d810df27b853428f58e1 Mon Sep 17 00:00:00 2001
From: Ahmed Abualsaud
Date: Thu, 11 Jun 2026 09:32:02 -0400
Subject: [PATCH 13/18] cleanup
---
website/www/site/.hugo_build.lock | 0
1 file changed, 0 insertions(+), 0 deletions(-)
delete mode 100644 website/www/site/.hugo_build.lock
diff --git a/website/www/site/.hugo_build.lock b/website/www/site/.hugo_build.lock
deleted file mode 100644
index e69de29bb2d1..000000000000
From 660da57cf57b16422ff5ebb038b51c08721f1514 Mon Sep 17 00:00:00 2001
From: Ahmed Abualsaud
Date: Thu, 11 Jun 2026 09:37:51 -0400
Subject: [PATCH 14/18] address comments
---
website/www/site/content/en/documentation/dsls/sql/ddl.md | 6 +++---
.../www/site/content/en/documentation/dsls/sql/overview.md | 2 ++
2 files changed, 5 insertions(+), 3 deletions(-)
diff --git a/website/www/site/content/en/documentation/dsls/sql/ddl.md b/website/www/site/content/en/documentation/dsls/sql/ddl.md
index 1b5a355918ff..7f89cfbc8c88 100644
--- a/website/www/site/content/en/documentation/dsls/sql/ddl.md
+++ b/website/www/site/content/en/documentation/dsls/sql/ddl.md
@@ -20,9 +20,9 @@ limitations under the License.
Beam SQL Data Definition Language (DDL) provides a standard three-level hierarchy to manage metadata across external data sources,
making it easy to explore available data structures and query data across different systems.
-1. Catalog: The top-level container representing an external metadata provider. Examples include a Hive Metastore, AWS Glue, or a Lakehouse (formerly BigLake) Catalog.
-2. Database: A logical grouping within a Catalog. This typically maps to a "Schema" in traditional RDBMS or a "Namespace" in systems like Apache Iceberg.
-3. Table: The leaf node containing the schema definition and the underlying data.
+1. **Catalog**: The top-level container representing an external metadata provider. _For example, a Hive Metastore, AWS Glue, or a Lakehouse (formerly BigLake) Catalog._
+2. **Database**: A logical grouping within a Catalog. This typically maps to a "Schema" in traditional RDBMS or a "Namespace" in systems like Apache Iceberg.
+3. **Table**: The leaf node containing the schema definition and the underlying data.
Beam can resolve multiple Catalogs simultaneously. This structure enables Federated Querying, meaning
you can execute complex pipelines that bridge disparate environments within a single SQL statement.
diff --git a/website/www/site/content/en/documentation/dsls/sql/overview.md b/website/www/site/content/en/documentation/dsls/sql/overview.md
index 7aa9d7bab298..a5b9a6e6a686 100644
--- a/website/www/site/content/en/documentation/dsls/sql/overview.md
+++ b/website/www/site/content/en/documentation/dsls/sql/overview.md
@@ -23,6 +23,8 @@ bounded and unbounded `PCollections` with SQL statements. Your SQL query
is translated to a `PTransform`, an encapsulated segment of a Beam pipeline.
You can freely mix SQL `PTransforms` and other `PTransforms` in your pipeline.
+Beam SQL extends [DDL commands](../ddl) for supported catalogs.
+
Beam SQL uses Calcite SQL based on [Apache Calcite](https://calcite.apache.org),
a dialect widespread in big data processing.
From dd6a2f316d357b418f81c315f43e5abb417bfec5 Mon Sep 17 00:00:00 2001
From: Ahmed Abualsaud
Date: Thu, 11 Jun 2026 12:59:12 -0400
Subject: [PATCH 15/18] address comments
---
website/www/site/content/en/documentation/dsls/sql/ddl.md | 4 ++--
.../www/site/content/en/documentation/dsls/sql/overview.md | 2 +-
2 files changed, 3 insertions(+), 3 deletions(-)
diff --git a/website/www/site/content/en/documentation/dsls/sql/ddl.md b/website/www/site/content/en/documentation/dsls/sql/ddl.md
index 7f89cfbc8c88..45ab5d6a7d7b 100644
--- a/website/www/site/content/en/documentation/dsls/sql/ddl.md
+++ b/website/www/site/content/en/documentation/dsls/sql/ddl.md
@@ -156,7 +156,7 @@ CREATE DATABASE other_catalog.sales_data;
{{< /tab >}}
{{< tab USE >}}
Sets the active Database for the current session. This simplifies queries by allowing you
-to reference Databases directly without their fully-qualified names (e.g. my_db instead of my_catalog.my_db)
+to reference Databases directly without their fully-qualified names (for example, my_db instead of my_catalog.my_db)
Note: All subsequent TABLE commands will be executed under this Database, unless fully qualified.
@@ -263,7 +263,7 @@ TBLPROPERTIES '{
- This creates an Iceberg table named
orders under the namespace sales_data, within the prod_iceberg catalog.
- The table is partitioned by
region_id, then by the day value of order_date (using Iceberg's hidden partitioning).
- - The table is created with the appropriate properties
"write.format.default" and "read.split.target-size". The Beam property "beam.write.triggering_frequency_seconds" configures the Iceberg sink.
+ - The table is created with the appropriate properties
"write.format.default" and "read.split.target-size". The Beam property "beam.write.triggering_frequency_seconds" configures the Iceberg sink.
- Beam sink and source configuration properties are prefixed with
"beam.write." and "beam.read.", respectively.
{{< /tab >}}
diff --git a/website/www/site/content/en/documentation/dsls/sql/overview.md b/website/www/site/content/en/documentation/dsls/sql/overview.md
index a5b9a6e6a686..5ff3ef325651 100644
--- a/website/www/site/content/en/documentation/dsls/sql/overview.md
+++ b/website/www/site/content/en/documentation/dsls/sql/overview.md
@@ -23,7 +23,7 @@ bounded and unbounded `PCollections` with SQL statements. Your SQL query
is translated to a `PTransform`, an encapsulated segment of a Beam pipeline.
You can freely mix SQL `PTransforms` and other `PTransforms` in your pipeline.
-Beam SQL extends [DDL commands](../ddl) for supported catalogs.
+Beam SQL extends [DDL commands](../ddl/) for supported catalogs.
Beam SQL uses Calcite SQL based on [Apache Calcite](https://calcite.apache.org),
a dialect widespread in big data processing.
From bf7703856f949d21ee520e3ce2d368159c89fa4a Mon Sep 17 00:00:00 2001
From: Ahmed Abualsaud
Date: Thu, 11 Jun 2026 12:59:41 -0400
Subject: [PATCH 16/18] address comments
---
website/www/site/.hugo_build.lock | 0
website/www/site/content/en/documentation/dsls/sql/ddl.md | 2 +-
2 files changed, 1 insertion(+), 1 deletion(-)
create mode 100644 website/www/site/.hugo_build.lock
diff --git a/website/www/site/.hugo_build.lock b/website/www/site/.hugo_build.lock
new file mode 100644
index 000000000000..e69de29bb2d1
diff --git a/website/www/site/content/en/documentation/dsls/sql/ddl.md b/website/www/site/content/en/documentation/dsls/sql/ddl.md
index 45ab5d6a7d7b..11d588a4c35b 100644
--- a/website/www/site/content/en/documentation/dsls/sql/ddl.md
+++ b/website/www/site/content/en/documentation/dsls/sql/ddl.md
@@ -20,7 +20,7 @@ limitations under the License.
Beam SQL Data Definition Language (DDL) provides a standard three-level hierarchy to manage metadata across external data sources,
making it easy to explore available data structures and query data across different systems.
-1. **Catalog**: The top-level container representing an external metadata provider. _For example, a Hive Metastore, AWS Glue, or a Lakehouse (formerly BigLake) Catalog._
+1. **Catalog**: The top-level container representing an external metadata provider. For example, a Hive Metastore, AWS Glue, or a Lakehouse (formerly BigLake) Catalog.
2. **Database**: A logical grouping within a Catalog. This typically maps to a "Schema" in traditional RDBMS or a "Namespace" in systems like Apache Iceberg.
3. **Table**: The leaf node containing the schema definition and the underlying data.
From a58d6448faee18659e449d1c5642ef722447fc18 Mon Sep 17 00:00:00 2001
From: Ahmed Abualsaud
Date: Thu, 11 Jun 2026 13:38:23 -0400
Subject: [PATCH 17/18] fix link
---
website/www/site/content/en/documentation/dsls/sql/overview.md | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/website/www/site/content/en/documentation/dsls/sql/overview.md b/website/www/site/content/en/documentation/dsls/sql/overview.md
index 5ff3ef325651..12bcec1a077c 100644
--- a/website/www/site/content/en/documentation/dsls/sql/overview.md
+++ b/website/www/site/content/en/documentation/dsls/sql/overview.md
@@ -23,7 +23,7 @@ bounded and unbounded `PCollections` with SQL statements. Your SQL query
is translated to a `PTransform`, an encapsulated segment of a Beam pipeline.
You can freely mix SQL `PTransforms` and other `PTransforms` in your pipeline.
-Beam SQL extends [DDL commands](../ddl/) for supported catalogs.
+Beam SQL extends [DDL commands](/documentation/dsls/sql/ddl) for supported catalogs.
Beam SQL uses Calcite SQL based on [Apache Calcite](https://calcite.apache.org),
a dialect widespread in big data processing.
From 0f29b12f47b2bfcfd8a06b3e555a60bffcc4dfbf Mon Sep 17 00:00:00 2001
From: Ahmed Abualsaud
Date: Thu, 11 Jun 2026 17:21:05 -0400
Subject: [PATCH 18/18] cleanup
---
website/www/site/.hugo_build.lock | 0
1 file changed, 0 insertions(+), 0 deletions(-)
delete mode 100644 website/www/site/.hugo_build.lock
diff --git a/website/www/site/.hugo_build.lock b/website/www/site/.hugo_build.lock
deleted file mode 100644
index e69de29bb2d1..000000000000