A command-line tool for listing the contents and metadata of Apache Parquet files and partitioned parquet datasets, modelled on HDF5's h5ls.
curl -fsSL https://github.com/dunnock/pqls/releases/latest/download/install.sh | shcargo install pqlsInspect a single file:
pqls data.parquetDetailed stats (per-column min/max/nulls):
pqls -d data.parquetDump as CSV:
pqls --csv data.parquet
pqls --csv --head 100 data.parquetList a partitioned dataset (shows schema brief per file by default):
pqls /path/to/dataset/
pqls -r /path/to/dataset/Detailed stats (per-row-group column min/max/nulls):
pqls -d -r /path/to/dataset/Machine-readable output:
pqls -q data.parquetpqls [OPTIONS] <PATH> [PATH_B]
ARGS:
<PATH> path to a .parquet file or directory to inspect
[PATH_B] second .parquet file for schema diff (required by --diff)
OPTIONS:
--diff compare schemas of two files; exits 0 if identical, 1 if different
-d, --detail show per-row-group column statistics (min/max/nulls)
-r, --recursive recurse into a directory and list all .parquet files
--csv dump rows as CSV to stdout
--head <N> limit output to the first N rows (applies to --csv and --ndjson)
-q, --quiet suppress human-readable headers; emit tab-separated summary lines
--schema print schema only (column names and types)
--json emit output as JSON (works with --schema, --kv-meta, --check, --partition-stats, --diff)
--ndjson stream rows as newline-delimited JSON (NDJSON)
--sample <N> emit N randomly-sampled rows; requires --ndjson or --csv
--columns <COLS> comma-separated list of column names to project (e.g. id,ts,value)
--kv-meta print Parquet key-value metadata (writer version, custom properties)
--scan-stats scan the full file to compute per-column min/max/nulls/n_distinct; requires -d
--partition-stats aggregate row counts and file sizes across a Hive-partitioned directory; requires -r
--check verify file integrity by reading the footer and all row groups
--deep with --check: read every data page (slower but catches corrupt column data)
-h, --help print help
-V, --version print version
Single binary. No JVM, no Python interpreter, no pip install. Drop the binary on
any Linux box and it runs — sub-100ms startup on the critical path of a data pipeline.
Composable. Stdout is always clean (data only; warnings go to stderr). Pipe anywhere:
pqls --csv file.parquet | xsv stats
pqls --schema file.parquet | diff - expected.schemaAgent-friendly. Machine-readable --schema --json and --ndjson output let code
agents inspect schema and rows without parsing human text. See SKILL.md for patterns.
One-liner install:
curl -fsSL https://github.com/dunnock/pqls/releases/latest/download/install.sh | shFast:
| Tool | Runtime | Startup | Schema dump | Stats | Pipe-composable |
|---|---|---|---|---|---|
| pqls | none | ~50ms | --schema --json |
--scan-stats |
yes |
| parquet-tools | JVM | ~2s | text only | yes | no |
| DuckDB | Go binary | ~200ms | SQL only | SQL | no |
| fastparquet | Python | ~500ms | Python API | Python API | no |
| pqls | parquet-cli (Apache) | pqrs | DuckDB | |
|---|---|---|---|---|
| Single binary, no JVM/Python | yes | no (JAR) | yes | yes |
--schema --json for agents |
yes | no (text only) | no | via SQL |
NDJSON rows (--ndjson) |
yes | no | cat -f json | via SQL |
Column projection (--columns) |
yes | yes | no | via SQL |
Random sampling (--sample N) |
yes | no | yes | ORDER BY random() |
Key-value metadata (--kv-meta) |
yes | footer cmd | no | parquet_kv_metadata() |
| Directory / partition listing | yes | no | no | no |
| SKILL.md for code agents | yes | no | no | no |
| Composable (stdin/stdout clean) | yes | no | partial | no |
pqls is the only single-binary tool in this list that produces JSON schema output and NDJSON rows without requiring SQL. It is designed for shell pipelines and agent tooling where DuckDB's startup time or SQL syntax is overhead.
pqls is designed to be called by code agents (Claude, Codex, Cursor, etc.) without any human at the terminal.
pqls --schema --json /path/to/foo.parquetReturns a JSON object — safe to parse with jq or Python json.loads. One entry per
top-level column:
type— the friendly type, including nested forms:int64,timestamp[us,UTC],list<float64>,map<utf8,int64>,struct.physical_type— the parquet physical type for primitive columns;nullfor nested ones (a group has no single physical type).logical_type—DATE,TIMESTAMP_MICROS,DECIMAL(10,2),LIST, … ornull.
Nested columns are reported under their own name and type, not flattened to parquet
leaves — markout (list<float64>), never an inner element (DOUBLE). --columns
takes those top-level names everywhere.
Row-level output flattens differently per format:
--ndjsonemits lists and structs as native JSON values.--csvis flat, so a fixed-sizeArray[T; n]expands intoname_0 … name_{n-1}columns and a variable-lengthlistrenders as one bracketed cell,"[1.5, 2.0, -0.015]".-dreports statistics per parquet leaf, named by dotted path (markout.list.element).--scan-statscannot compute min/max/n_distinct for nested columns and reports null counts only for them.
pqls --ndjson --sample 50 foo.parquet50 rows, one JSON object per line. Pipe to jq for field inspection.
pqls --ndjson --columns user_id,amount --sample 20 foo.parquetpqls --kv-meta --json foo.parquet | jq '.["pandas"]'# Find which files in a partitioned dataset have more than 1M rows
pqls -q --recursive /data/events/ \
| awk -F'\t' '$2 > 1000000 { print $1 }'Scripts should test $?:
0— success, output on stdout1— file/path error or schema mismatch (with --diff)2— corrupt or invalid parquet, or bad flag combination
Pre-built binaries are attached to every GitHub release:
| Platform | Asset |
|---|---|
| Linux x86_64 | pqls-linux-x86_64.tar.gz |
| Linux aarch64 | pqls-linux-aarch64.tar.gz |
| macOS Intel (x86_64) | pqls-darwin-x86_64.tar.gz |
| macOS Apple Silicon (aarch64) | pqls-darwin-aarch64.tar.gz |
| Windows x86_64 | pqls-windows-x86_64.zip |
One-liner install (Linux and macOS):
curl -fsSL https://github.com/dunnock/pqls/releases/latest/download/install.sh | shBy default this installs to your user bin directory (~/.local/bin) when it is on
PATH, falling back to any other PATH directory under $HOME, and otherwise to
/usr/local/bin (with a confirmation prompt). Override the location with
PQLS_INSTALL=<dir>.
Install from crates.io:
cargo install pqlsThis compiles from source and works on any platform with a Rust toolchain.
System requirements: none beyond a standard Linux/macOS/Windows environment. pqls
has no runtime dependencies — no JVM, no Python — and requires no elevated privileges
or kernel tuning. The Linux binaries are dynamically linked against glibc and are built
on Ubuntu 22.04 (glibc 2.35), so they run on that release and newer; on older
distributions, install with cargo install pqls instead.
pqls accepts s3://bucket/key and s3://bucket/prefix/ paths directly.
# inspect schema of a single S3 object (no full download)
pqls s3://my-bucket/events/2024/data.parquet
# JSON schema for agents / pipelines
pqls --schema --json s3://my-bucket/events/2024/data.parquet
# list all .parquet files under a prefix with brief schema per file
pqls s3://my-bucket/events/2024/
# machine-readable listing
pqls --json s3://my-bucket/events/2024/pqls uses the standard AWS credential provider chain — no new CLI flags.
Set credentials via environment variables, ~/.aws/credentials, an IAM
instance role, or SSO:
# env vars
export AWS_ACCESS_KEY_ID=...
export AWS_SECRET_ACCESS_KEY=...
export AWS_REGION=us-east-1
# named profile
export AWS_PROFILE=my-profile
# IAM role / ECS task role / IRSA — no config needed- Schema inspection uses S3 range gets (last 64 KiB) — no whole-file download.
- No row data access over S3 (
--csv,--ndjsonare not supported for S3 paths). - No local cache — every
pqls s3://...issues fresh range gets.
# default: minor bump (e.g. 0.5.1 → 0.6.0)
make release
# or override the bump type:
make release BUMP=patch
make release BUMP=major
# preview what would happen — no side effects:
make release-dry-runRequirements: run from your host (not a container), on the main branch with a clean
working tree that is in sync with origin/main, with crates.io credentials available
(cargo login, or CARGO_REGISTRY_TOKEN in the environment).
A release has two halves:
make release— bumps the version, commits, tags, and pushes. The tag push triggers.github/workflows/release.yml, which builds the multi-platform binaries and cuts the GitHub release (~20 min). Usemake release BUMP=nonewhen the version inCargo.tomlwas already bumped by hand as part of the change being released.make publish—cargo publish --locked, run by you from this machine once the GitHub release looks right. CI never publishes to crates.io, and no crates.io token is stored in the repository's GitHub secrets.
Recovery helpers (normally not needed):
make release-resume— push if the local commit/tag exist but push previously failed
Licensed under either of MIT or Apache-2.0 at your option.