Feature Request / Improvement
Streaming upsert writers (e.g. AWS Firehose Iceberg delivery) write equality-delete files on every commit. Compaction applies them into rewritten data files but leaves the entries in the manifests, and expire_snapshots can't touch files the current snapshot still references — so they accumulate without bound. On one of our production tables we measured ~90K dangling delete entries growing ~4.6K/day, and every query planning over recent partitions has to read the ever-growing delete manifests.
Java Iceberg handles this (rewrite_data_files with remove-dangling-deletes, RemoveDanglingDeletesSparkAction), but PyIceberg's MaintenanceTable currently only has expire_snapshots, and engines like Athena expose no statement for it either — so users on Athena/Firehose stacks have no non-Spark way out.
Proposal: table.maintenance.remove_dangling_deletes() — a metadata-only commit that:
- classifies per
(partition_spec_id, partition): an equality delete at sequence s is dangling iff no live data file in that partition has sequence < s (position deletes: <= s); ambiguous cases (unpartitioned specs, unknown content) are kept
- carries data manifests through unchanged, drops fully-dangling delete manifests, rewrites mixed ones to their surviving entries, and commits as a
replace snapshot against the current ref
One enabler is worth a small standalone fix first: ManifestWriterV2 hardcodes content=data, so PyIceberg currently can't write delete-content manifests at all.
We have a working implementation built on PyIceberg 0.12 internals (write_manifest_list, a ManifestWriterV2 subclass, AddSnapshotUpdate/SetSnapshotRefUpdate with AssertRefSnapshotId), validated against production Glue/Athena tables, with a test matrix for the classification rules. Happy to contribute it if there's interest.
Feature Request / Improvement
Streaming upsert writers (e.g. AWS Firehose Iceberg delivery) write equality-delete files on every commit. Compaction applies them into rewritten data files but leaves the entries in the manifests, and
expire_snapshotscan't touch files the current snapshot still references — so they accumulate without bound. On one of our production tables we measured ~90K dangling delete entries growing ~4.6K/day, and every query planning over recent partitions has to read the ever-growing delete manifests.Java Iceberg handles this (
rewrite_data_fileswithremove-dangling-deletes,RemoveDanglingDeletesSparkAction), but PyIceberg'sMaintenanceTablecurrently only hasexpire_snapshots, and engines like Athena expose no statement for it either — so users on Athena/Firehose stacks have no non-Spark way out.Proposal:
table.maintenance.remove_dangling_deletes()— a metadata-only commit that:(partition_spec_id, partition): an equality delete at sequence s is dangling iff no live data file in that partition has sequence < s (position deletes: <= s); ambiguous cases (unpartitioned specs, unknown content) are keptreplacesnapshot against the current refOne enabler is worth a small standalone fix first:
ManifestWriterV2hardcodescontent=data, so PyIceberg currently can't write delete-content manifests at all.We have a working implementation built on PyIceberg 0.12 internals (
write_manifest_list, aManifestWriterV2subclass,AddSnapshotUpdate/SetSnapshotRefUpdatewithAssertRefSnapshotId), validated against production Glue/Athena tables, with a test matrix for the classification rules. Happy to contribute it if there's interest.