Teaching AI to draw, four scoring dimensions beat nine — swap "absolute scoring" for "compare against a good example," and the model itself compresses its code from 13,500 tokens to under 2,000. It proves one thing: a reward function isn't better for being more complex. When five rulers measure the same thing, the model just spins in place. The difference is that relative judgment opens up the learning space, while hand-picked anchors nail down what "good" means.

It's a bit like asking nine food critics to score the same dish — five of them are basically saying "the salt level's right," one is counting how many grains of salt were used, one is checking the plating, and the remaining two are earnest but barely carry any weight. Nobody cares whether the dish actually tastes good; the critics just end up nodding at each other. The score stalls at 65 and won't budge, so the chef keeps producing standardized five-petal flowers. Surya's fix was to swap "rate it" for "compare it against the house specialty" — keep one judge that can actually create spread, then bump its say to sixty percent. Strangely, the dish got better, and the chef's recipe (the model's output code) shrank from 13,500 characters to about 2,000, because he finally realized a good dish doesn't need piled-up steps. End of analogy. The real difference: that "judge" is actually a differentiable reward model, and every comparison feeds gradients into the parameters, while the 117 hand-picked "favorites" reference images bake the definition of "good" right into the training data.
Backstory

One prompt couldn't move a single stroke, so they let the model "draw by writing code"

Using reinforcement learning to teach AI to draw, the key isn't likeness — it's how to get a machine to learn aesthetics. A project that got stuck at the reward function.

Anyone who's used mainstream AI image generators knows the feeling: you want to change a flower's color, you have to rewrite the entire prompt and gamble that the model will give you what you want. In a project note published in March 2026, Surya framed this constraint as a starting point — "images aren't editable; the only interaction point is the prompt."

He and Cameron wanted to sidestep that dead end: have the model write JavaScript sketches directly in p5.brush (a watercolor brush library built on p5.js). The rendered PNG becomes a byproduct; the code itself is the product. When users edit the code, they edit the painting.

What really stalled them was a deeper question. Reinforcement learning (a method that lets a model fumble toward an optimal strategy through "try-fail-score" loops) can teach models to play chess and solve math, because right and wrong have objective answers; but "looks good" has no standard answer. Surya laid it out plainly in his notes: if the reward is too rigid, the model converges on a lazy solution; if it's too loose, the model spins in place. The design problem itself became a reward function problem.

The team — Surya, Cameron, and Alex Wang — trained Qwen 3.5. Training ran through a four-step loop for thousands of iterations: the model receives a prompt like "paint a watercolor hibiscus," writes a complete JS sketch; the sketch renders to PNG inside a sandboxed Puppeteer browser; a separate panel of judge models compares the PNG against two randomly drawn images from a reference pool and picks the better one; the win/loss becomes a reward signal, fed back through GRPO (a reinforcement learning algorithm that updates parameters based on relative rankings within a group), and the cycle restarts. Images in the reference pool were individually graded by hand; "favorites"-tier samples entered the comparison pool, so in every round, the model was being measured against works deemed "good."

0.65
Ceiling of the old nine-signal reward function
Source: Hacker News trending (buzzing.cc Chinese translation)
1,664
Total hand-graded reference images
Source: Hacker News trending (buzzing.cc Chinese translation)
117
"Favorites"-tier samples entering the comparison pool
Source: Hacker News trending (buzzing.cc Chinese translation)
13.5K→2K
Compression of per-output code tokens
Source: Hacker News trending (buzzing.cc Chinese translation)

The project page only shows painting screenshots, training flow diagrams, and a thesis defense video — no full reward curves or ablation study (a research method that removes variables one at a time to see their effect) data were released. The subsequent experiment — "collapsing nine scoring dimensions into four, then swapping absolute scoring for pairwise comparison" — is the real meat of this rewrite, and it's what the next section breaks down.

Why It Matters

Nine judges, all measuring the same thing

A reward function isn't better for having more dimensions — when five rulers measure the same length, the model just spins in place.

The first version of the reward function packed in nine signals: whether the code runs, whether p5.brush was used correctly, whether length hits a target tier, what HPSv3 (a scoring model trained on human preference data) says, plus a "committee" of GPT-5.4 and Gemini judging prompt adherence, and finally four independent judges for recognizability, aesthetics, technique, and depth. Looked comprehensive, but training stalled at a reward of 0.65. Every output looked the same: a flat, five-petal paper-cut flower. The reward score kept climbing, but the painting itself never improved.

Tear the nine signals apart and look at them individually, and the problem jumps out. The correlation between the four quality judges and "prompt adherence" sits at 0.85 to 0.95 — put differently, five people wielding five different rulers come up with nearly identical numbers. Those five rulers are really measuring the same thing.

The code-length reward made up about a third of the total score, but it saturated by step 30 (hit the ceiling, can't go any higher), after which it was zero gradient (the directional signal telling the model "improve this way"; near zero means no direction) — the model learned nothing new from it. The only thing still moving was HPSv3, but its weight was just 0.10 — faint as background noise.

The judges turned on each other first. Over and over they told the model the same thing: make a prompt-fitting, recognizable, technically passable flower. The model got the message, so it delivered exactly that — a perpetual five-petal paper-cut flower. By breaking "what's good" into too many pieces, the reward function choked off the model's room to explore.

Behind this lies an underrated engineering truth: a reward function doesn't get more "aesthetic" by being more complex. When multiple signals are highly correlated, complexity just rubber-stamps the same judgment. The most counter-intuitive finding: cut scoring dimensions from nine to four, swap "rate it" for "compare against a good example," and the model actually learned to draw with compositional awareness using simpler code. Next question: how do you actually build that ruler? The next section breaks it down.