On this page
The model is one part of the experiment
Alethic Phase I targets ~1B parameters, combining GQA, SGCA, RMSNorm, SwiGLU, and RoPE in attention. The context target is 8,240 tokens, with a custom tokenizer plan around 48K vocabulary.
Roughly 20B training tokens is a full-run target, not consumed compute. Architecture target, prepared corpus, packed tokens, and tokens consumed by training are different quantities.
Make the run resumable
Free Colab and Kaggle sessions constrain duration and hardware availability. Checkpointing, export state, stable data construction, and a clear run manifest become part of the experiment.
Preparation includes tokenization, export, source-ID deduplication, train/validation packing, and binary data. CPU preparation and GPU training have different bottlenecks. Having a GPU allocated does not mean a tokenizer is using it.
Keep the baseline close
Different tokenizers, mixtures, ordering, optimizers, schedules, contexts, or precision can make an interesting architecture comparison uninformative. A controlled matched baseline therefore belongs in the current direction.
SGCA changes parameters and computation. A fair report states the matching method and remaining differences, without hiding parameter, geometry, and compute matching under one label.
Use history carefully
The 116M predecessor and Alethic-151M are earlier work, not the current specification. Alethic-151M used a 256-token original context and ~20K SentencePiece vocabulary. Importing those values into Phase I would confuse history with current configuration.
Earlier 1.5B and 12K/54K plans are historical too. Longer-context scaling is postponed until the smaller-scale mechanism is evaluated.
What should be published
Architecture, actual parameterization, configuration, tokenizer, data documentation, seeds, checkpoints, logs, and evaluation form a reproducibility package. A model name and parameter target are not a substitute.
The Alethic overview records the current state and available artifacts. A useful progress update is measured state from a particular run with its limits, not a claim that the intended final model already exists.