After reading this blog, I began to wonder which Parquet version and compression methods the everyday tools we rely on actually use, only to find that there's no straightforward way to determine this. That curiosity and the difficulty of quickly discovering such details motivated me to create iparq (Information Parquet). My goal with iparq is to help users easily identify the specifics of the Parquet files generated by different engines, making it clear which features—like newer encodings or certain compression algorithms—the creator of the parquet is using.
- Bloom filters: Detects real Bloom-filter metadata and reports its size. Read more in this great article.
- Encodings and types: Shows physical and logical types plus encodings such as
RLE_DICTIONARY,DELTA_BINARY_PACKED, andBYTE_STREAM_SPLIT. - Indexes and dictionary pages: Reports dictionary pages, column indexes, and offset indexes.
- Statistics: Displays min/max values and available null and distinct counts.
- Row groups and sort order: Shows row-group sizes, row counts, and declared sorting columns.
- Page locations: Reports column-chunk, dictionary-page, data-page, and Bloom-filter offsets.
- Schema details: Includes legacy converted types, decimal precision/scale, nesting levels, and GeoParquet statistics availability.
- Compression: Shows codecs with optional column sizes and compression ratios.
- Machine-readable output: Emits JSON for scripts and agent workflows.
iParq requires Python 3.10 or later.
-
Make sure to have Astral's UV installed by following the steps here:
-
Execute the following command:
uvx --refresh iparq inspect yourparquet.parquet
-
Install the package using pip:
pip install iparq
-
Verify the installation by running:
iparq --help
-
Make sure to have Astral's UV installed by following the steps here:
-
Execute the following command:
uv pip install iparq
-
Verify the installation by running:
iparq --help
-
Run the following:
brew tap MiguelElGallo/tap https://github.com/MiguelElGallo//homebrew-iparq.git brew install MiguelElGallo/tap/iparq iparq --help
iparq supports inspecting single files, multiple files, and glob patterns:
iparq inspect <filename(s)> [OPTIONS]Options include:
--format,-f: Output format, eitherrich(default) orjson--metadata-only,-m: Show only file metadata without column details--column,-c: Filter results to show only a specific column--sizes,-s: Show column sizes and compression ratios--details,-d: Show row groups, sort order, encodings, types, indexes, page locations, Bloom-filter size, and detailed statistics
# Basic inspection .
iparq inspect yourfile.parquet
# Output in JSON format
iparq inspect yourfile.parquet --format json
# Show only metadata
iparq inspect yourfile.parquet --metadata-only
# Filter to show only a specific column
iparq inspect yourfile.parquet --column column_name
# Show column sizes and compression ratios
iparq inspect yourfile.parquet --sizes
# Show storage-level details
iparq inspect yourfile.parquet --details# Inspect multiple specific files
iparq inspect file1.parquet file2.parquet file3.parquet
# Use glob patterns to inspect all parquet files
iparq inspect *.parquet
# Use specific patterns
iparq inspect yellow*.parquet data_*.parquet
# Combine patterns and specific files
iparq inspect important.parquet temp_*.parquetWhen inspecting multiple files, each file's results are displayed with a header showing the filename. The utility will read the metadata of each file and print the compression codecs used in the parquet files.
For scripts and agents, add --format json. A single file produces one JSON
object; multiple files produce one JSON array whose entries include file.
Diagnostics are written to stderr, and any unreadable input makes the command
exit non-zero without corrupting successful JSON output.
iParq publishes CLI and Agent Skill documentation, an Agentic Resource Discovery (ARD) catalog, an agent-readable documentation index, and a portable Parquet inspection skill. The catalog advertises the existing read-only CLI and its JSON output; it does not add a network service or change how iParq accesses files.
Projects that declare iParq as a dependency can install the version-matched skill bundled in the Python package:
uv add iparq
uvx library-skills install --skill iparq-parquet-inspector --yesThis creates a project-local .agents/skills/iparq-parquet-inspector symlink
to the skill in the installed iParq package. It is separate from ephemeral CLI
execution with uvx iparq, which does not add iParq to the project environment.
ParquetMetaModel(
created_by='parquet-cpp-arrow version 14.0.2',
num_columns=3,
num_rows=3,
num_row_groups=1,
format_version='2.6',
serialized_size=2223,
key_value_metadata_keys=['ARROW:schema', 'pandas']
)
Parquet Column Information
┏━━━━━━━━━━━┳━━━━━━━━━━━━━┳━━━━━━━┳━━━━━━━━━━━━━┳━━━━━━━┳━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Row Group ┃ Column Name ┃ Index ┃ Compression ┃ Bloom ┃ Min Value ┃ Max Value ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━╇━━━━━━━╇━━━━━━━━━━━━━╇━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━┩
│ 0 │ one │ 0 │ SNAPPY │ ❌ │ -1.0 │ 2.5 │
│ 0 │ two │ 1 │ SNAPPY │ ❌ │ bar │ foo │
│ 0 │ three │ 2 │ SNAPPY │ ❌ │ False │ True │
└───────────┴─────────────┴───────┴─────────────┴───────┴───────────┴───────────┘
Compression codecs: {'SNAPPY'}
iparq inspect yourfile.parquet --sizes
Parquet Column Information
┏━━━━━━━━┳━━━━━━━━━┳━━━━━━━┳━━━━━━━━┳━━━━━━━┳━━━━━━━━━━━┳━━━━━━━━━━━┳━━━━━━━━┳━━━━━━━━┳━━━━━━━┓
┃ Row ┃ Column ┃ ┃ ┃ ┃ ┃ ┃ ┃ ┃ ┃
┃ Group ┃ Name ┃ Index ┃ Compr… ┃ Bloom ┃ Min Value ┃ Max Value ┃ Values ┃ Compr… ┃ Ratio ┃
┡━━━━━━━━╇━━━━━━━━━╇━━━━━━━╇━━━━━━━━╇━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━╇━━━━━━━━╇━━━━━━━┩
│ 0 │ one │ 0 │ SNAPPY │ ❌ │ -1.0 │ 2.5 │ 3 │ 104.0B │ 1.0x │
│ 0 │ two │ 1 │ SNAPPY │ ❌ │ bar │ foo │ 3 │ 80.0B │ 0.9x │
│ 0 │ three │ 2 │ SNAPPY │ ❌ │ False │ True │ 3 │ 42.0B │ 1.0x │
└────────┴─────────┴───────┴────────┴───────┴───────────┴───────────┴────────┴────────┴───────┘