The research program

Research

Exploring new Pareto frontiers in language-model capability, compute, and memory.

An open question

How much more can a language model do with the resources it already has?

View research status
01 / Research philosophy

Efficiency has
more than one axis.

A useful result must hold up beyond the experiment that first revealed it. We study the balance between model quality, computational cost, and memory requirements, then test how that balance changes across settings.

We are working toward better trade-offs. We are not publishing performance comparisons or numerical results at this stage.

01

Architecture

We investigate model architectures that selectively allocate expensive computation where it provides the most value. The question is how to use limited resources well while preserving useful language-model capability.

02

Scaling

We evaluate research across increasing model sizes rather than relying on single small-scale wins. Controlled experiments help us decide which ideas warrant a larger investment of compute.

03

Representation

We validate architectural findings beyond toy byte-level settings using modern subword tokenization and broader text distributions. Transfer to real-world language is a research milestone in its own right.

04

Memory and inference state

We study how architecture affects the memory required during inference, particularly as sequence length grows. Memory movement and retained state are part of the efficiency question, alongside computational work.

02 / Evaluation methodology

Evidence before
extrapolation.

Our evaluation approach is designed to distinguish promising ideas from findings that hold up under more demanding conditions.

01

Matched controls

Compare candidates with appropriate baselines under comparable experimental conditions.

02

Multiple training seeds

Use multiple training seeds for important research promotions to test whether findings are repeatable.

03

Held-out evaluation

Evaluate on data held out from training to examine generalization.

04

Model-size scaling

Revisit findings at increasing parameter counts before drawing broader conclusions.

05

Compute-aware comparison

Consider the computational resources behind a result, alongside model quality.

06

Recall and long-range diagnostics

Probe information retention and behavior across longer sequences.

03 / Research status

Progress, with
clear boundaries.

Controlled studies have reached the approximately-100M-parameter class. Real-web transfer is complete.

We are preparing the build of our first approximately-300M-parameter generalist model, intended to become our first model candidate for external evaluation and potential public release.

Current focus

Preparing our first ~300M generalist model build.

  1. Completed

    Architecture exploration

  2. Completed

    Controlled small-model studies

  3. Completed

    Subword-token transfer (BPE)

  4. Completed

    ~30M-class validation

  5. Completed

    ~100M-class controlled scaling

  6. Completed

    Real-web transfer at ~100M class

  7. Preparing

    First ~300M generalist model build

  8. Future work

    Production runtime

Let’s compare notes

Interested in efficient
AI infrastructure?

We are interested in conversations with researchers, infrastructure teams, hardware companies, potential design partners, and investors working on the economics of language models.