s3-turbo-list · Rust CLI

1M objects in ~19s.

Finds the parallel work. Lists and diffs large buckets. Parquet out.

RustApache-2.0 v0.39.0 (Sep 30, 2026) S3-compatibleParquetAgent-safe JSON

Third-party benchmark · Alibaba Cloud OSS

One stream against -c 8

1M objects on Alibaba Cloud OSS, measured by a third party and quoted in the project README. Listing speed is bounded by the requests per second the provider serves.

Elapsed time, single stream and -c 8
~55s ~19s 1 stream ~55s -c 8 ~19s

Approximate times as reported. Results depend on bucket structure and the provider’s request-rate limit.

Know what is here and what changed.

The two questions before a migration or an investigation.

LIST

What is in a bucket

One bucket becomes a Parquet object list, with no scanner tuning first.

DIFF

What changed between two

Source and target list in parallel and merge in key order: one Parquet set flagging equal, missing, or changed.

Not a sync engine It does not copy objects or replace a storage browser. It leaves a reliable basis for the next decision.

The scan plan follows the bucket.

Parallel work opens where the bucket has structure; only long segments split further.

  1. 1

    Read the shape

    Probe real prefix boundaries before recursive work begins.

  2. 2

    Run segments

    Open parallel list work without a prepared hints file.

  3. 3

    Split the long tail

    Fan out long segments only while useful capacity is available.

  4. 4

    Write the list

    An analysis-ready object list, not terminal-only output.

Check the endpoint before scaling up.

Verified: AWS S3Verified: MinIOVerified: Baidu BOSPreset: Cloudflare R2Preset: OSSPreset: B2

AWS S3, MinIO, and Baidu BOS are on the project’s verified compatibility path; R2, OSS, and B2 have presets. Run compat-probe before trusting a long job on any endpoint.

Sep 30, 2026

What’s new in v0.39.0.

  • Startup discovery keeps at most its boundary target
  • Flat runs share the boundary budget by size
  • Startup probes parse pages with the fast Contents parser and reuse what they fetched

Get started

Download the binary for your platform from the GitHub release, verify it against SHA256SUMS, and put it on your PATH.

  1. Local preflight — no S3 access

    s3-turbo-list doctor
    Show real output
    OK    binary_version: 0.39.0
    OK    working_directory: /tmp/inventory
    OK    config_parse: resolved configuration is valid TOML/CLI state
    OK    config_file: no config file (searched ./s3-turbo-list.toml and ~/.s3-turbo-list.toml); built-in defaults apply
    OK    aws_profile: AWS_PROFILE is not set; the AWS SDK uses its default credential chain
    OK    provider: no provider preset; plain S3 settings apply
    OK    endpoint_url: no endpoint URL: requests go to AWS S3 (pass --endpoint-url or --provider for another service)
    SKIP  proxy: the request host depends on the bucket and region (virtual-hosted or AWS endpoint), which doctor does not take; the run log names the proxy, if any, for each resolved endpoint
    SKIP  network: doctor is a local preflight and does not contact S3 endpoints; use compat-probe to validate an endpoint
    Resolved: provider -, endpoint -, addressing auto, concurrency 100, threads 4
    Doctor status: ok
  2. Full recursive inventory → Parquet in out/

    s3-turbo-list list --bucket my-bucket --region us-east-2 --output-dir out

FAQ

Isn’t aws s3 ls enough?

For a quick, small inspection, aws s3 ls is often enough. s3-turbo-list is for large inventories and repeatable comparisons where output and recovery matter.

Do I need a hints file?

No. List mode probes real CommonPrefixes boundaries itself. Hints files remain an optional control for repeated inventories.

What does startup discovery do?

Every run starts with a few delimiter probes that find real CommonPrefixes boundaries; nothing is cached in the working directory. A flat namespace with no prefixes is split by single-key probes instead, and so is a large flat directory under one prefix (like data/part-… under data/), so the run starts parallel either way.

Can an interrupted run pick up where it stopped?

Yes. Every interrupted list run saves a checkpoint on Ctrl-C or SIGTERM. Run it again with --resume and the same --output-dir, and it lists only the key ranges not yet written; that output covers only those ranges, so combine it with the interrupted run’s. A crash leaves the previous checkpoint unchanged, and diff and --start-after runs do not checkpoint.

s3-turbo-list