Tag: Reinforcement-Learning

Jev: State In, Typed Decisions Out

Diogo Almeida posted on X about a new model, Jev, from a company called TypeSafe. TypeSafe calls it its first “System One” model, a term borrowed from Daniel Kahneman’s split between fast, intuitive System 1 thinking and slow, deliberate System 2 reasoning. A normal LLM writes out its reasoning in text. Jev skips straight to a typed answer.

What it actually does

Ask a normal LLM “given this incident, what should we do?” and it answers in text:

[…959 words]

Training Qwen to Paint Watercolors with Pure RL

Surya trained Qwen 3.5 to have taste in watercolor painting. The results are stunning.

It’s purely RL:

The system is a four-step loop, run thousands of times during training.

The model receives a prompt, something like draw a peach hibiscus in watercolour, and writes a complete p5.brush JavaScript sketch. The sketch is rendered in a sandboxed Puppeteer environment, which produces a PNG. The PNG is judged against two random reference paintings sampled from a hand-rated pool, with a separate judge model picking the better watercolour. The judgment is converted into a reward signal, GRPO updates the model, and the loop runs again.

[…297 words]