On this page
Start with the residual update
A decoder block retrieves information and introduces it into a residual stream. In the conceptual baseline used here, normalized input produces a GQA output, and that output is added directly:
Two ideas are contained in the expression: retrieving information and deciding how it changes the representation. SGCA asks whether making the latter a separate learned control path is useful.
The proposed control term
Attention still retrieves with GQA. Its output enters sinusoidal elementwise modulation, alongside a learned linear projection of normalized input. Together they define the residual update.
α is a learned tensor broadcast-compatible with the attention output. Calling it a scalar would narrow the mechanism beyond its definition. Shape and initialization belong in each reproducible run configuration.
A question, not a conclusion
A periodic gate can change sign, scale, and sensitivity as its input varies. That may affect optimization and representation, but could also introduce instability or initialization sensitivity. An algebraically interesting mechanism is not automatically empirically useful.
Earlier SinGatedLM experiments motivated this direction. They used a predecessor mechanism on a small dataset and do not prove current SGCA works at Phase-I scale.
A fair comparison
The projection changes parameterization and computation as well as the gate. Matching the backbone is necessary but insufficient. Tokenizer, corpus, data order, context, optimizer, schedule, precision, batching, and evaluation should match where possible. Parameter and compute matching must be labeled separately. Multiple seeds test whether a difference survives initialization.
Ablations should isolate the sinusoid, modulation, and projection. Comparing sin with identity, tanh, and sigmoid asks a clearer question than comparing two bundles of architecture choices.
What the next experiment should answer
Where does the mechanism help, where does it fail, and what does it cost? A controlled negative result is evidence too. Phase I targets ~1B parameters and an 8,240-token context, but a planned scale is not a benchmark.
The SGCA page keeps the definition, controls, and open questions together. Results belong beside their exact configurations and limitations.