Infrastructure

Building dataset pipelines under free compute

What source-ID deduplication, supervised tokens, and packed windows say about preparation.

On this page

Read the counters literally

A preparation log is not a training benchmark. One supplied Python portion of The Stack v3 records 125,215 JSONL records inspected, 20,638 / 32,767 train windows packed, 170,057,120 supervised tokens, 512 validation windows, and 62,375 identical source-ID duplicates skipped.

These describe a preparation stage, not final corpus size, optimizer-consumed tokens, or model quality.

Tokens and storage differ

A supervised token contributes a prediction target under the packer's accounting. Window size, separators, masks, padding, and prediction shift affect how the count relates to stored token arrays.

Multiplying a counter by a byte width gives file size only when the dtype and format are known. File bytes are measured from files; prediction targets are counted by the preparation or training procedure.

Deduplication has a scope

Skipping an identical source ID avoids processing the same identified upstream record again. It does not prove content or near-duplicate deduplication, nor absence of split contamination.

Identity keys, persistent state, and stable split assignment matter. A restarted export should not silently change what gets counted or where a source lands.

Make failures observable

Resumption under session limits benefits from upstream position, completed shards, counters, and reproducible packing rules. Partially written output should be distinguishable from a finished artifact.

The preparation record belongs beside training. A final validation loss cannot explain a data bug that was never logged.

Keep the comparison controlled

The baseline and SGCA run should share tokenizer and data construction where possible. Order, context, and token budget need the same care as the equation.

The lab note preserves the supplied snapshot as a preparation observation, not a live or final report.

Last updated 01 Oct 2026Discuss this work