Package and access data on Amazon S3 in a squash manner
-
Updated
Aug 6, 2023 - Rust
Package and access data on Amazon S3 in a squash manner
A file uploader specialized for uploading many small files onto HDFS
PySpark script to aggregate small parquet files in a prefix into larger files. Designed to be run on AWS Glue
Read-only "small-file tax" analyzer for Delta Lake. Reads only the _delta_log to reconstruct the per-partition file-size distribution, then prices the wasted compute and S3 requests in dollars using real Spark internals (openCostInBytes, maxPartitionBytes) and emits a compaction plan with ROI and payback plus the OPTIMIZE SQL. Never runs OPTIMIZE.
Command Line Hangman/Word Guessing game.
Read-only Delta Lake table health scanner with safe fix plans for Databricks, Python, and PySpark.
File merge action to merge files in HDFS or local filesystem
To associate your repository with the small-files topic, visit your repo's landing page and select "manage topics."