Architecture
We investigate model architectures that selectively allocate expensive computation where it provides the most value. The question is how to use limited resources well while preserving useful language-model capability.
Exploring new Pareto frontiers in language-model capability, compute, and memory.
How much more can a language model do with the resources it already has?
View research statusA useful result must hold up beyond the experiment that first revealed it. We study the balance between model quality, computational cost, and memory requirements, then test how that balance changes across settings.
We are working toward better trade-offs. We are not publishing performance comparisons or numerical results at this stage.
We investigate model architectures that selectively allocate expensive computation where it provides the most value. The question is how to use limited resources well while preserving useful language-model capability.
We evaluate research across increasing model sizes rather than relying on single small-scale wins. Controlled experiments help us decide which ideas warrant a larger investment of compute.
We validate architectural findings beyond toy byte-level settings using modern subword tokenization and broader text distributions. Transfer to real-world language is a research milestone in its own right.
We study how architecture affects the memory required during inference, particularly as sequence length grows. Memory movement and retained state are part of the efficiency question, alongside computational work.
Our evaluation approach is designed to distinguish promising ideas from findings that hold up under more demanding conditions.
Compare candidates with appropriate baselines under comparable experimental conditions.
Use multiple training seeds for important research promotions to test whether findings are repeatable.
Evaluate on data held out from training to examine generalization.
Revisit findings at increasing parameter counts before drawing broader conclusions.
Consider the computational resources behind a result, alongside model quality.
Probe information retention and behavior across longer sequences.
Controlled studies have reached the approximately-100M-parameter class. Real-web transfer is complete.
We are preparing the build of our first approximately-300M-parameter generalist model, intended to become our first model candidate for external evaluation and potential public release.
Preparing our first ~300M generalist model build.
We are interested in conversations with researchers, infrastructure teams, hardware companies, potential design partners, and investors working on the economics of language models.