Skip to content
dunnockPublic

About

CLI for listing parquet tables

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

pqls

Release Latest Release License

A command-line tool for listing the contents and metadata of Apache Parquet files and partitioned parquet datasets, modelled on HDF5's h5ls.

Install

curl -fsSL https://github.com/dunnock/pqls/releases/latest/download/install.sh | sh

Or install with cargo:

cargo install pqls

Examples

Inspect a single file:

pqls data.parquet

Detailed stats (per-column min/max/nulls):

pqls -d data.parquet

Dump as CSV:

pqls --csv data.parquet
pqls --csv --head 100 data.parquet

List a partitioned dataset (shows schema brief per file by default):

pqls /path/to/dataset/
pqls -r /path/to/dataset/

Detailed stats (per-row-group column min/max/nulls):

pqls -d -r /path/to/dataset/

Machine-readable output:

pqls -q data.parquet

CLI

pqls [OPTIONS] <PATH> [PATH_B]

ARGS:
  <PATH>            path to a .parquet file or directory to inspect
  [PATH_B]          second .parquet file for schema diff (required by --diff)

OPTIONS:
      --diff                compare schemas of two files; exits 0 if identical, 1 if different
  -d, --detail              show per-row-group column statistics (min/max/nulls)
  -r, --recursive           recurse into a directory and list all .parquet files
      --csv                 dump rows as CSV to stdout
      --head <N>            limit output to the first N rows (applies to --csv and --ndjson)
  -q, --quiet               suppress human-readable headers; emit tab-separated summary lines
      --schema              print schema only (column names and types)
      --json                emit output as JSON (works with --schema, --kv-meta, --check, --partition-stats, --diff)
      --ndjson              stream rows as newline-delimited JSON (NDJSON)
      --sample <N>          emit N randomly-sampled rows; requires --ndjson or --csv
      --columns <COLS>      comma-separated list of column names to project (e.g. id,ts,value)
      --kv-meta             print Parquet key-value metadata (writer version, custom properties)
      --scan-stats          scan the full file to compute per-column min/max/nulls/n_distinct; requires -d
      --partition-stats     aggregate row counts and file sizes across a Hive-partitioned directory; requires -r
      --check               verify file integrity by reading the footer and all row groups
      --deep                with --check: read every data page (slower but catches corrupt column data)
  -h, --help                print help
  -V, --version             print version

Why pqls?

Single binary. No JVM, no Python interpreter, no pip install. Drop the binary on any Linux box and it runs — sub-100ms startup on the critical path of a data pipeline.

Composable. Stdout is always clean (data only; warnings go to stderr). Pipe anywhere:

pqls --csv file.parquet | xsv stats
pqls --schema file.parquet | diff - expected.schema

Agent-friendly. Machine-readable --schema --json and --ndjson output let code agents inspect schema and rows without parsing human text. See SKILL.md for patterns.

One-liner install:

curl -fsSL https://github.com/dunnock/pqls/releases/latest/download/install.sh | sh

Fast:

Tool Runtime Startup Schema dump Stats Pipe-composable
pqls none ~50ms --schema --json --scan-stats yes
parquet-tools JVM ~2s text only yes no
DuckDB Go binary ~200ms SQL only SQL no
fastparquet Python ~500ms Python API Python API no

How pqls compares

pqls parquet-cli (Apache) pqrs DuckDB
Single binary, no JVM/Python yes no (JAR) yes yes
--schema --json for agents yes no (text only) no via SQL
NDJSON rows (--ndjson) yes no cat -f json via SQL
Column projection (--columns) yes yes no via SQL
Random sampling (--sample N) yes no yes ORDER BY random()
Key-value metadata (--kv-meta) yes footer cmd no parquet_kv_metadata()
Directory / partition listing yes no no no
SKILL.md for code agents yes no no no
Composable (stdin/stdout clean) yes no partial no

pqls is the only single-binary tool in this list that produces JSON schema output and NDJSON rows without requiring SQL. It is designed for shell pipelines and agent tooling where DuckDB's startup time or SQL syntax is overhead.

Agent usage

pqls is designed to be called by code agents (Claude, Codex, Cursor, etc.) without any human at the terminal.

Discover schema

pqls --schema --json /path/to/foo.parquet

Returns a JSON object — safe to parse with jq or Python json.loads. One entry per top-level column:

  • type — the friendly type, including nested forms: int64, timestamp[us,UTC], list<float64>, map<utf8,int64>, struct.
  • physical_type — the parquet physical type for primitive columns; null for nested ones (a group has no single physical type).
  • logical_type — DATE, TIMESTAMP_MICROS, DECIMAL(10,2), LIST, … or null.

Nested columns

Nested columns are reported under their own name and type, not flattened to parquet leaves — markout (list<float64>), never an inner element (DOUBLE). --columns takes those top-level names everywhere.

Row-level output flattens differently per format:

  • --ndjson emits lists and structs as native JSON values.
  • --csv is flat, so a fixed-size Array[T; n] expands into name_0 … name_{n-1} columns and a variable-length list renders as one bracketed cell, "[1.5, 2.0, -0.015]".
  • -d reports statistics per parquet leaf, named by dotted path (markout.list.element).
  • --scan-stats cannot compute min/max/n_distinct for nested columns and reports null counts only for them.

Sample rows to understand data

pqls --ndjson --sample 50 foo.parquet

50 rows, one JSON object per line. Pipe to jq for field inspection.

Project specific columns

pqls --ndjson --columns user_id,amount --sample 20 foo.parquet

Check embedded metadata (Spark / Pandas schema)

pqls --kv-meta --json foo.parquet | jq '.["pandas"]'

Composable pipeline example

# Find which files in a partitioned dataset have more than 1M rows
pqls -q --recursive /data/events/ \
  | awk -F'\t' '$2 > 1000000 { print $1 }'

Exit code contract

Scripts should test $?:

  • 0 — success, output on stdout
  • 1 — file/path error or schema mismatch (with --diff)
  • 2 — corrupt or invalid parquet, or bad flag combination

Releases

Pre-built binaries are attached to every GitHub release:

Platform Asset
Linux x86_64 pqls-linux-x86_64.tar.gz
Linux aarch64 pqls-linux-aarch64.tar.gz
macOS Intel (x86_64) pqls-darwin-x86_64.tar.gz
macOS Apple Silicon (aarch64) pqls-darwin-aarch64.tar.gz
Windows x86_64 pqls-windows-x86_64.zip

One-liner install (Linux and macOS):

curl -fsSL https://github.com/dunnock/pqls/releases/latest/download/install.sh | sh

By default this installs to your user bin directory (~/.local/bin) when it is on PATH, falling back to any other PATH directory under $HOME, and otherwise to /usr/local/bin (with a confirmation prompt). Override the location with PQLS_INSTALL=<dir>.

Install from crates.io:

cargo install pqls

This compiles from source and works on any platform with a Rust toolchain.

System requirements: none beyond a standard Linux/macOS/Windows environment. pqls has no runtime dependencies — no JVM, no Python — and requires no elevated privileges or kernel tuning. The Linux binaries are dynamically linked against glibc and are built on Ubuntu 22.04 (glibc 2.35), so they run on that release and newer; on older distributions, install with cargo install pqls instead.

S3 paths

pqls accepts s3://bucket/key and s3://bucket/prefix/ paths directly.

# inspect schema of a single S3 object (no full download)
pqls s3://my-bucket/events/2024/data.parquet

# JSON schema for agents / pipelines
pqls --schema --json s3://my-bucket/events/2024/data.parquet

# list all .parquet files under a prefix with brief schema per file
pqls s3://my-bucket/events/2024/

# machine-readable listing
pqls --json s3://my-bucket/events/2024/

AWS auth

pqls uses the standard AWS credential provider chain — no new CLI flags. Set credentials via environment variables, ~/.aws/credentials, an IAM instance role, or SSO:

# env vars
export AWS_ACCESS_KEY_ID=...
export AWS_SECRET_ACCESS_KEY=...
export AWS_REGION=us-east-1

# named profile
export AWS_PROFILE=my-profile

# IAM role / ECS task role / IRSA — no config needed

Trade-offs

  • Schema inspection uses S3 range gets (last 64 KiB) — no whole-file download.
  • No row data access over S3 (--csv, --ndjson are not supported for S3 paths).
  • No local cache — every pqls s3://... issues fresh range gets.

Releasing

# default: minor bump (e.g. 0.5.1 → 0.6.0)
make release

# or override the bump type:
make release BUMP=patch
make release BUMP=major

# preview what would happen — no side effects:
make release-dry-run

Requirements: run from your host (not a container), on the main branch with a clean working tree that is in sync with origin/main, with crates.io credentials available (cargo login, or CARGO_REGISTRY_TOKEN in the environment).

A release has two halves:

  1. make release — bumps the version, commits, tags, and pushes. The tag push triggers .github/workflows/release.yml, which builds the multi-platform binaries and cuts the GitHub release (~20 min). Use make release BUMP=none when the version in Cargo.toml was already bumped by hand as part of the change being released.
  2. make publish — cargo publish --locked, run by you from this machine once the GitHub release looks right. CI never publishes to crates.io, and no crates.io token is stored in the repository's GitHub secrets.

Recovery helpers (normally not needed):

  • make release-resume — push if the local commit/tag exist but push previously failed

License

Licensed under either of MIT or Apache-2.0 at your option.

About

CLI for listing parquet tables

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages