Supported Features
The tables below describe which Iceberg features Sail supports. ✅ indicates support, ⚠️ identifies the supported cases in the notes, and 🚧 means not supported.
Format Versions
Sail creates version 2 Iceberg tables by default. Set the format-version table property when creating a table to use another supported version, or change it later to upgrade a table.
| Feature | Supported | Notes |
|---|---|---|
| Format versions 1, 2, and 3 | ✅ | Reads and writes table metadata. Version-specific features are listed below. |
| Format-version upgrades | ✅ | Through the format-version table property. |
| Format-version downgrades | 🚧 | Iceberg permits format-version upgrades only. |
| Parquet data and delete files | ✅ | Delete-file support is listed by version below. |
| Puffin deletion-vector files | ✅ | Version 3 merge-on-read operations. |
| Avro manifests and manifest lists | ✅ | — |
| Avro and ORC data files | 🚧 | — |
Core Table Operations
| Feature | Supported | Notes |
|---|---|---|
| Table creation, CTAS, and replacement | ⚠️ | Through SQL and the DataFrame writer APIs; replacement is available in the Memory catalog. See Table DDL. |
| Current snapshot reads, append, and full overwrite | ✅ | Both partitioned and non-partitioned tables. |
| Metadata-as-data reads | ✅ | Optional manifest scanning during query execution instead of loading the file list on the driver. |
| Predicate pushdown and file pruning | ✅ | Uses partition transforms and available file metrics. |
| Metadata aggregate optimization | ✅ | Eligible COUNT, MIN, and MAX queries use exact metadata. Incomplete metrics or applicable deletes can require scanning data. |
| Schema evolution on write | ✅ | mergeSchema adds fields and applies supported promotions. overwriteSchema replaces the schema during full overwrite. The two options cannot be combined. |
| Predicate overwrite | ⚠️ | DataFrameWriterV2.overwrite(condition) on identity-partition columns. Requires compatible live partition specs and no active delete files. |
| Dynamic partition overwrite | ⚠️ | overwritePartitions() or overwrite-mode = dynamic on existing tables with compatible live partition specs and no active delete files. Replaces partitions present in the input. Empty input is a no-op. |
| Time travel | ✅ | Snapshot ID, timestamp, or an existing branch/tag reference. Requires the referenced metadata and data files. |
| Branch/tag creation and branch writes | 🚧 | — |
| Table property DDL | ⚠️ | SET/UNSET TBLPROPERTIES and format upgrades through Memory, HMS, Glue, and Iceberg REST. See Table DDL for catalog requirements. |
| Column type and default DDL | ⚠️ | Top-level type promotions and SET/DROP DEFAULT; defaults require format version 3 and typed literals. See Table DDL. |
| Commit conflict handling | ⚠️ | Validates metadata requirements and the expected snapshot, with limited metadata publication retries. Row-level conflicts require replanning, including with snapshot isolation. |
Table DDL
DDL support varies by catalog provider. See the Iceberg DDL support matrix for supported operations and limitations.
DML Operations
Copy-on-write rewrites affected data files, while merge-on-read records deletes separately. Sail uses copy-on-write by default for DELETE, UPDATE, and MERGE INTO. The write.delete.mode, write.update.mode, and write.merge.mode table properties let you choose a mode for each operation. For merge-on-read, version 2 MERGE uses Parquet position-delete files, while version 3 DML uses Puffin deletion vectors.
| Operation | Copy-on-write (v1–v3) | Merge-on-read (v2) | Merge-on-read (v3) | Notes |
|---|---|---|---|---|
DELETE | ✅ | ⚠️ | ✅ | V2 uses equality-delete files on unpartitioned tables, with all columns as equality keys. Nested, floating-point, unknown, and variant fields are unsupported in that writer path. |
UPDATE | ✅ | 🚧 | ✅ | V3 merge-on-read supports updates that move rows between partitions. |
MERGE INTO with inserts, updates, and deletes | ✅ | ✅ | ✅ | Matched, insert, and WHEN NOT MATCHED BY SOURCE clauses. Multiple source matches cannot update one target row. |
MERGE WITH SCHEMA EVOLUTION | 🚧 | 🚧 | 🚧 | — |
When metadata shows that a delete removes every row in a file, Sail removes the file reference directly, including for partitioned tables. Version 3 merge-on-read works with partitioned and non-partitioned tables. Merge-on-read UPDATE (version 3) and MERGE (versions 2 and 3) append replacement data for updated rows.
Metadata, Schema, and Layout
Across the supported format versions, Sail reads and writes table metadata, tracks fields by ID, and uses available partition and file metrics when planning queries. The table below lists the details and limits.
| Feature | Supported | Notes |
|---|---|---|
| Table metadata, snapshots, manifest lists, and manifests | ✅ | Reads and writes the supported version-specific fields. |
| Field IDs and schema history | ✅ | Resolves evolved columns by ID, including nested fields. |
| Partition transforms | ✅ | Identity, bucket, truncate, year, month, day, and hour. Existing void transforms are handled. |
| Partition evolution | ⚠️ | Reads existing specs and writes using the current spec. Schema-replacing overwrite can change partitioning. No partition-evolution DDL. |
| Sort orders | ⚠️ | Honors supported existing single-source sort transforms and records sort-order IDs on data files. No sort-order DDL or multi-argument transforms. |
| Column metrics | ✅ | Uses file counts, null counts, bounds, and other available metrics for planning. |
| NaN value counts | ⚠️ | Reads existing counts. The data writer does not populate nan_value_counts. |
| Name mapping | ✅ | Reads imported files without field IDs using an existing schema.name-mapping.default. |
| Statistics-file generation and query use | 🚧 | Preserves existing statistics-file metadata. No statistics-file creation or query consumption. |
| Snapshot/reference history | ✅ | Preserves history and reads existing refs. |
Version 2: Delete Files
| Feature | Read | Write | Notes |
|---|---|---|---|
| Sequence numbers and inheritance | ✅ | ✅ | Used to determine delete applicability. |
| Manifest and data-file content types | ✅ | ✅ | Distinguishes data, equality deletes, and position deletes. |
| Position-delete files | ✅ | ⚠️ | Written by version 2 merge-on-read MERGE. Existing files can still be read after an upgrade to version 3. |
| Equality-delete files | ✅ | ⚠️ | Written by version 2 merge-on-read DELETE. Reads bind keys by field ID and apply partition/sequence rules. See DML Operations. |
| Delete-aware scan planning | ✅ | 🚧 | Applies supported deletes before returning rows, including scans with a limit. |
Version 3: Extended Types and Capabilities
| Feature | Supported | Notes |
|---|---|---|
| Variant | ✅ | Reads and writes logical VARIANT, including Parquet shredding. |
| Unknown type | ✅ | All-null logical fields, including copy-on-write preservation. |
timestamp_ns, timestamptz_ns | ⚠️ | Iceberg/Arrow conversion and defaults are implemented. Spark Connect result-schema conversion does not support nanosecond timestamps. |
| Geometry and geography | 🚧 | Binary storage conversion does not provide their logical type semantics. |
| Initial and write defaults | ⚠️ | Reads existing defaults and supports SQL DEFAULT, CREATE TABLE defaults, and ALTER COLUMN SET/DROP DEFAULT. DDL requires typed literals and a server that supports version 3. |
| Row lineage and first-row-ID inheritance | ✅ | Assigns row IDs on insert. Copy-on-write and merge-on-read updates preserve row IDs and advance update sequence numbers. |
| Multi-argument partition/sort transforms | 🚧 | — |
| Deletion vectors in Puffin files | ✅ | Reads and writes vectors, combining prior positional deletes when replacing the vector for a data file. |
| Encryption keys and AES-GCM stream encryption | 🚧 | — |
Catalogs and Maintenance
Sail works with filesystem-backed tables and tables in Iceberg REST, AWS Glue, or Hive Metastore. For a catalog-backed table, use its catalog name so reads follow the catalog metadata location and commits publish new metadata through the catalog.
| Feature | Supported | Notes |
|---|---|---|
| Catalog-backed reads and commits | ✅ | Uses the catalog metadata pointer. Filesystem discovery and version-hint.text cannot replace it. |
| Iceberg REST views | ⚠️ | Create, load, list, and drop when the server provides those endpoints. Other lifecycle operations are not implemented. |
| Iceberg SQL UDF specification | 🚧 | — |
| Snapshot expiration | 🚧 | — |
| Data-file compaction | 🚧 | rewrite_data_files. |
| Position-delete rewrite procedures | 🚧 | — |
| Multi-table atomic transactions | 🚧 | — |
| Z-order clustering | 🚧 | — |
| Write-audit-publish | 🚧 | — |
| Streaming reads and writes | 🚧 | — |
