Skip to content

[DRAFT] Transparent Compression - #133

Open
Erotemic wants to merge 4 commits into
evaleval:mainfrom
Erotemic:dev/compression
Open

Erotemic wants to merge 4 commits into
evaleval:mainfrom
Erotemic:dev/compression

Conversation

@Erotemic

@Erotemic Erotemic commented May 8, 2026

Copy link
Copy Markdown
Collaborator

This PR adds transparent compression support for aggregate and sample result files. Compressed variants of <uuid>.json and <uuid>_samples.jsonl are now treated as equivalents of the plain files for supported codecs: .gz, .zst, .bz2, .xz, and .lz4.

Changes:

  • Add every_eval_ever/io.py with compression suffix detection, transparent openers, result discovery, and duplicate-variant resolution.
  • Route validation and duplicate-entry checks through the new IO helpers.
  • Add a duplicate_variant validation error when both plain and compressed variants of the same logical file are present.
  • Add writer-side --compress, --compress-aggregate, and --compress-samples flags.
  • Compute recorded checksums over the actual on-disk bytes.
  • Document compression behavior and add optional [zst] and [lz4] codec extras.

Tests:

  • Validate compressed files through expand_paths and validate_file.
  • Ensure compressed files with invalid payloads still fail validation.
  • Cover duplicate compressed/plain variants as CI-failing errors.
  • Confirm aggregate and sample files do not false-positive as duplicates.
  • Ensure missing optional codecs produce typed errors instead of crashes.
  • Verify writer-side compression flags and override behavior.

Defaults remain unchanged: outputs are still uncompressed unless explicitly requested.

Out of scope:
The manifest-generation and viewer-Parquet scripts also need this resolver, but they live in EEE_datastore. This PR keeps the package-level compression support isolated and leaves datastore integration for a companion PR.

NOTE:

I would not be comfortable merging this until we have CI running. I have not tested this robustly yet.

Erotemic added 4 commits May 7, 2026 21:58
New module ``every_eval_ever/io.py`` provides a single source of truth
for opening EEE result files (``<uuid>.json`` aggregates and
``<uuid>_samples.jsonl`` per-instance samples) regardless of compression.

Recognized codecs match the HuggingFace Hub's documented set
(https://huggingface.co/docs/hub/en/datasets-adding#file-formats):

  .gz, .zst, .bz2, .xz, .lz4

``.zip`` is intentionally excluded — it is an archive container, not a
stream codec, and would conflict with the duplicate-variant rule.

Helpers:

* ``open_eee_text(path, mode)``        — text I/O with codec auto-detection
* ``is_eee_result(path)``              — 'aggregate' / 'samples' / None
* ``eee_uuid_stem(path)``              — strip kind+codec suffixes
* ``detect_compression(path)``         — name of the trailing codec
* ``add_compression_suffix(path, cs)`` — synthesize the on-disk filename
* ``iter_eee_results(roots)``          — recursive discovery of EEE files
* ``find_duplicate_variants(paths)``   — enforces 'one variant per
                                         (folder, uuid, kind)' rule

``zstandard`` and ``lz4`` are optional runtime dependencies; the open
helper raises ``ImportError`` with an actionable message pointing at
the right ``[zst]`` / ``[lz4]`` extra. The stdlib codecs (.gz, .bz2,
.xz) work out of the box.

This commit only adds the module + unit tests. Subsequent commits wire
the rest of the package (validate, check_duplicate_entries, converters,
CLI flags) through these helpers.

53 passed, 2 skipped (lz4-codec tests skip when ``lz4`` is not
installed; same for ``zstandard`` if absent).
Wire ``validate.py`` and ``check_duplicate_entries.py`` through the
``every_eval_ever.io`` helpers so compressed result files
(``.gz``, ``.zst``, ``.bz2``, ``.xz``, ``.lz4``) are validated
identically to their plain counterparts.

Validator changes
-----------------

* ``validate_aggregate`` and ``validate_instance_file`` open via
  ``io.open_eee_text`` rather than ``Path.open`` / ``Path.read_text``.
* ``validate_file`` dispatches on ``io.is_eee_result(path)`` rather
  than ``path.suffix == '.json' / '.jsonl'``. Compressed forms are
  recognized via the suffix-stripping detector; non-EEE filenames
  (``.zip``, ``.parquet``, plain ``.csv``) report
  ``unsupported_extension`` as before.
* ``expand_paths`` enumerates EEE result files via
  ``io.iter_eee_results`` for directory inputs. Files passed
  explicitly are accepted regardless of suffix and let
  ``validate_file`` produce the unsupported-extension error.
* New ``_duplicate_variant_reports`` produces synthetic
  ``ValidationReport``s with a ``duplicate_variant`` error type for
  every ``(folder, uuid_stem, kind)`` group with more than one
  physical variant. ``validate.main`` runs this check unconditionally
  before per-file validation, so a CI/PR-bot path cannot silently
  green-light a folder containing both ``abc.json`` and ``abc.json.gz``.
* Missing optional codec dependencies surface as ``codec_unavailable``
  validation errors rather than crashing the whole run.

check_duplicate_entries changes
-------------------------------

* ``expand_paths`` recognizes compressed aggregate files; explicit
  file inputs that aren't aggregate-shaped are skipped (preserving
  the historical "ignore non-JSON" behaviour for non-EEE content).
* The aggregate-payload reader uses ``io.open_eee_text``.
  ``_samples.jsonl`` files are intentionally excluded from this
  command's discovery — duplicate detection runs over aggregate
  metadata only.

io.py refinement
----------------

* ``is_eee_result`` and ``eee_uuid_stem`` are lenient about the
  ``_samples`` prefix on per-instance files: bare ``*.jsonl`` is
  still recognized as samples (matching what lm-eval emits when no
  file UUID is supplied). Strict ``_samples.jsonl`` matching is still
  preferred for stem extraction.

Tests
-----

* New ``tests/test_validate_compression.py`` covers: same fixture in
  plain + ``.gz`` form yields identical validation outcomes for both
  aggregate and samples; ``expand_paths`` finds compressed files;
  duplicate-variant reports fire and CI-gate appropriately;
  distinct kinds in the same folder don't false-positive; missing
  optional codec surfaces as a typed error rather than an exception.

213 passed, 5 skipped (4 require optional codec deps, 1 is unrelated
upstream skip).
Threads compression through the converter writer paths. Defaults to
``--compress none`` everywhere for full backwards compatibility — this
commit only enables compression on opt-in.

CLI flags (added to all four ``convert`` subcommands)
-----------------------------------------------------

* ``--compress {none|gz|zst|bz2|xz|lz4}``
    Default codec for both aggregate and per-instance files.
* ``--compress-aggregate {none|gz|zst|bz2|xz|lz4}``
    Override for aggregate ``<uuid>.json`` only. Defaults to
    ``--compress``'s value. Recommended setting: ``none``, since HF and
    GitHub render uncompressed JSON inline for spot-checking.
* ``--compress-samples {none|gz|zst|bz2|xz|lz4}``
    Override for per-instance ``<uuid>_samples.jsonl`` only. Recommended
    setting for any submission going to a public store: ``gz`` (5–15x
    size reduction with no UX cost — the samples files are too large
    for web preview anyway).

Argparse rejects unsupported codecs at parse time. Optional codec deps
(``zstandard``, ``lz4``) trigger ``ImportError`` with an actionable
message at write time when the user picks the codec but hasn't
installed the extra.

Writer-side plumbing
--------------------

* ``cli._write_log`` accepts ``compression``; output filename gets the
  codec suffix appended (``<uuid>.json[.gz]``).
* ``LMEvalInstanceLevelAdapter.transform_and_save`` accepts
  ``compression``; the recorded ``DetailedEvaluationResults.checksum``
  is computed over the on-disk (possibly compressed) bytes so that
  downstream verifiers can validate the file as-stored.
* ``HELMInstanceLevelDataAdapter`` accepts ``compression`` in its
  constructor; the adapter's ``self.path`` reflects the on-disk
  filename including any codec suffix.
* ``cli._cmd_convert_helm`` threads the resolved samples-compression
  via ``metadata_args['samples_compression']`` rather than adding a
  new positional to the long-existing ``transform_from_directory``
  signature.

Tests
-----

* ``tests/test_cli_compression.py``: full coverage of
  ``_resolve_compression``, ``_write_log`` round-trips per codec,
  LM-Eval samples writer compression including checksum-on-disk and
  the recorded ``file_path``, and parser-level wiring (defaults,
  per-kind override, invalid codec rejection).
* ``tests/test_cli_inspect_uuid.py``: existing fake-write-log stubs
  updated to accept the new ``compression=`` keyword.

226 passed, 5 skipped. Backwards compatibility preserved across all
existing tests.
* Add ``[zst]`` and ``[lz4]`` optional-dependency groups for the
  third-party codecs not available in the standard library. ``[all]``
  now pulls them in alongside ``[inspect]`` and ``[helm]``.

* README: new subsection under "Data Validation" documenting the
  recognized compressed forms, the duplicate-variant rule, and the
  ``--compress`` / ``--compress-samples`` writer flags. Notes that
  defaults remain uncompressed for backwards compatibility, and
  recommends ``--compress-samples gz`` (not the global ``--compress``)
  for public submissions so aggregate JSON stays browsable on the HF
  web UI.

226 passed, 5 skipped.
@Erotemic

Copy link
Copy Markdown
Collaborator Author

GPT Review (5m 12s):

Review: Request changes

High — Inspect sample outputs ignore the compression flags

The CLI advertises --compress as controlling both aggregate and per-instance outputs and exposes --compress-samples for every converter. However, _cmd_convert_inspect applies compression only when writing the aggregate, after the Inspect adapter has already generated its sample file. The Inspect instance writer still constructs a plain .jsonl path and uses Path.open() directly.

Consequently, convert inspect --compress gz produces a compressed aggregate but an uncompressed sample file, while --compress-samples is silently ignored. Thread the samples codec through metadata into InspectInstanceLevelDataAdapter, use add_compression_suffix() and open_eee_text(), and add an Inspect end-to-end compression test.

High — Corrupt compressed streams can abort the entire validator

Aggregate validation catches only OSError and ImportError around decompression. Errors from some compression decoders are not necessarily OSError, so malformed compressed content can escape instead of becoming a ValidationReport.

Instance validation is more broadly exposed: its try block covers only open_eee_text(). Actual decompression happens while iterating later, outside the exception handler. A corrupt sample stream can therefore terminate the complete list comprehension in main(), preventing reports for every remaining file.

Wrap the actual reads and iteration in the error boundary, normalize decompression failures into a typed validation error, and test deliberately corrupt bytes—not merely valid compressed streams containing invalid JSON—for every supported codec.

Medium — check-duplicates silently succeeds when an explicit input is unsupported

When an existing file is not recognized as an aggregate result, expand_paths() silently skips it. If every supplied path is skipped, the command checks zero files, prints “No duplicates found,” and exits successfully. That can green-light a mistyped filename or a CI invocation pointed at the wrong artifact.

Reject explicit non-aggregate files with a nonzero result rather than continuing silently, and add a regression test for an existing .txt, .jsonl, or otherwise unsupported path.

Low — Sample duplicate reports use an inconsistent file_type

Normal sample validation reports use file_type='instance', which is also the documented dataclass contract. Duplicate detection instead copies the IO kind directly, producing file_type='samples'. JSON consumers therefore receive two names for the same logical file type.

Map the IO kind to the validation-report vocabulary when constructing synthetic duplicate reports.

Verification status

No GitHub Actions runs are attached to either the current head or the synthetic merge commit, so the reported local test result has not been independently exercised against the merge result.

@mrshu mrshu left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚒️ review-anvil report

Review decision: COMMENT — This is a comment-only review, and one confirmed high-priority issue remains.
Result: The compression design is broad and well tested. A few focused edge cases would make it safer.
Scope: This PR adds transparent compression across CLI conversion, I/O, validation, adapters, tests, and user documentation.

Checks: 3 concerns checked; 2 confirmed and 1 narrowed.
Second check: 2 reviewers checked 4 findings; 2 kept, 2 clarified, and none removed.

Earlier review comments

The earlier review remains relevant at this head:

  • [high] Inspect output — Inspect sample writers still ignore --compress and --compress-samples. Earlier comment
  • [high] validation — Corrupt compressed streams can still escape the validator error boundary. Earlier comment
  • [medium] duplicate CLI — An explicit unsupported file can still produce a successful zero-file check. Earlier comment
  • [low] report schema — Sample duplicate reports still use samples instead of instance. Earlier comment

What I noticed

ID Priority Topic Code location What I noticed
RAVF001 high Zstandard reading every_eval_ever/io.py:223 The shared reader stops after the first frame in a valid multi-frame Zstandard stream. Later JSONL records can escape validation.
RAVF002 medium duplicate variants every_eval_ever/validate.py:437 Explicit-file validation checks only the named path. A compressed sibling can therefore escape the one-variant rule.
RAVF004 medium duplicate CLI errors every_eval_ever/check_duplicate_entries.py:101 Expected codec and decompression failures bypass the file-specific error path. The command exits early with a traceback.
Non-blocking low-priority finding
  • [low] UUID parsing every_eval_ever/converters/common/utils.py:9 — The shared UUID parser does not accept compressed result paths. No in-repository production caller was found, so this looks like a compatibility suggestion. (RAVF003)

Things to try

  • [high] Zstandard reading — The shared reader can read every concatenated frame. A two-frame JSONL test can assert the later row and full line count. (RAVF001)
  • [medium] duplicate variants — Collision discovery can inspect only same-parent, same-stem variants. Payload validation can stay limited to requested files. (RAVF002)
  • [medium] duplicate CLI errors — Known codec and read failures can use the existing file annotation path. Other programming errors can remain visible. (RAVF004)
  • [low] UUID parsing — One central compression suffix can be removed before UUID matching. Parameterized utility tests can cover every supported codec. (RAVF003)
Run details
  • Target: PR #133 (dev/compression, 13 files, +1270/-48)
  • Rounds: 1/1 completed; adaptive off; material findings remain
  • Mix: 3 codex-exec
  • Focus: correctness, maintainability, simplicity, production blast-radius, and constructive suggestions
  • Earlier review comments: 4 still present
  • Finding counts: 0 critical, 1 high, 2 medium, 1 low, 0 nit
  • Checks: concerns=3; confirmed=2, ruled-out=0, set-aside=0, lowered=0, narrowed=1
  • Second check: targeted; reviewers=2; kept=2, clarified=2, set-aside=0, removed=0; approval changed=no
  • Fixes applied: 0 (review-only)

Reviewed with review-anvil.

Comment thread every_eval_ever/io.py
except ImportError as exc: # pragma: no cover - exercised in env w/o dep
raise ImportError(_missing_codec_msg('zst', 'zstandard', 'zst')) from exc
if text_mode == 'rt':
reader = zstandard.ZstdDecompressor().stream_reader(open(p, 'rb'))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[high] Zstandard reading — The shared reader stops after the first frame

stream_reader() omits read_across_frames=True. Valid concatenated Zstandard frames can contain later JSONL records, so validation can finish without seeing them.

The shared reader can include every frame. A two-frame test can assert the later row and complete line count.

# Detect duplicate variants up front. Producing a fail-closed
# report per collision means the CI/PR-bot path cannot silently
# green-light a folder containing both `abc.json` and `abc.json.gz`.
duplicate_reports = _duplicate_variant_reports(file_paths)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[medium] duplicate variants — Explicit-file validation cannot see compressed siblings

expand_paths() returns only the named file, and this call checks that limited list. abc.json.gz can therefore escape the one-variant rule when the caller names only abc.json.

Collision discovery can inspect same-parent, same-stem variants while payload validation stays limited to requested files.

for file_path in file_paths:
try:
with open(file_path, 'r') as f:
with eee_io.open_eee_text(file_path, 'r') as f:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[medium] duplicate CLI errors — Expected compressed-read failures bypass the file annotation

This handler catches only JSONDecodeError. A missing codec raises ImportError, and corrupt gzip data raises BadGzipFile, so the command exits early with a traceback.

Known codec and read failures can use the existing file-specific error path. Other programming errors can remain visible.

@evijit

evijit commented Aug 11, 2026

Copy link
Copy Markdown
Member

@Erotemic -- did you want to resume work on this? The datastore is increasing in size so this is a great idea to finish

@Erotemic

Copy link
Copy Markdown
Collaborator Author

I can probably (90% chance) find time in the next week or two make this unstale. I'll comment if I start working on it. If anyone else wants to bring this to the finish line before that, then feel free.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants