Topic index
training
5 entries across research, projects, and writing.
notes
The Stack v3: a packing snapshot
A supplied Python-data preparation record, separate from training progress.
notesPhase I uses 8,240 tokens
The current context decision, separate from earlier long-context curricula.
ArticlePlanning a 1B language-model experiment independently
Architecture targets, constrained compute, and infrastructure for an inspectable run.
researchAlethic-151M
Experimental pretraining and SinGatedAttention exploration on constrained compute.
ArticleBuilding dataset pipelines under free compute
What source-ID deduplication, supervised tokens, and packed windows say about preparation.