Package and access data on Amazon S3 in a squash manner
-
Updated
Aug 6, 2023 - Rust
Package and access data on Amazon S3 in a squash manner
A file uploader specialized for uploading many small files onto HDFS
PySpark script to aggregate small parquet files in a prefix into larger files. Designed to be run on AWS Glue
Command Line Hangman/Word Guessing game.
Read-only Delta Lake table health scanner with safe fix plans for Databricks, Python, and PySpark.
Read-only "small-file tax" analyzer for Delta Lake. Reads only the _delta_log to reconstruct the per-partition file-size distribution, then prices the wasted compute and S3 requests in dollars using real Spark internals (openCostInBytes, maxPartitionBytes) and emits a compaction plan with ROI and payback plus the OPTIMIZE SQL. Never runs OPTIMIZE.
File merge action to merge files in HDFS or local filesystem
Add a description, image, and links to the small-files topic page so that developers can more easily learn about it.
To associate your repository with the small-files topic, visit your repo's landing page and select "manage topics."