Tag: Training

Debugging an LLM Training Run

LLM training has more moving parts than classic machine learning, but the debugging method is similar. Engineers check the training pipeline, data, objective, optimization, and evaluation in that order.

The fastest first test is to overfit a tiny dataset.

Take 100 examples and train on them many times. The model should almost memorize them. If it cannot, the problem is probably not data coverage. The problem may be the loss function, masking, tokenizer, optimizer, gradient flow, or training code.

[…265 words]

Train A BPE Tokenizer

I built train_tokenizer: a CLI that trains and evaluates a byte-level BPE tokenizer. It has 100 rows of sample data and a test file.

What a tokenizer does

A language model does not read text. It reads a list of numbers. A tokenizer turns text into that list, and turns the list back into text.

The simplest tokenizer splits text on spaces, into words, and gives each word a number. This breaks fast. Any word the model has not seen has no number. A model that never saw “photosynthesizing” cannot represent it.

[…1152 words]

Marin 535B-A23B Starts Training, in the Open

🚢 Marin 535B-A23B started training this week! As usual, the whole process is open.

Voyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow.

Before kicking off the run, we trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues, and to make a forecast of our hero run. This is by far our biggest run, so definitely expecting the unexpected.

Percy Liang, announcing the run on X.


The Marin project is a great example of being open to AI. The whole process is public: the scaling ladder they ran to debug the pipeline before committing GPU-months to the real thing, the exact token counts, the exact FLOPs, even the admission that they’re “expecting the unexpected” on their biggest run yet.

Most labs treat a run like this as a trade secret until the model ships. It’s great we see another public model build. The last public run at this scale was BLOOM, BigScience’s 176B model. Marin’s 535B total parameters (23B active) puts it well past that, the biggest public pretraining run that I’m aware of.