What it is

Basic question: for a fixed amount of compute, how do you effectively scale model parameters, architecture, and data?

Variables: model params including embeddings (M), dataset size (D), compute (C).

Key sources

To ingest