I keep running one very stupid-looking test on new models. I give them a banana.
What is the banana test?
Every model gets the same banana image and the same prompt. It has to turn the banana into part of a new drawing without hiding it. Then it has to build the result as a web page and animate the drawing process, including the pencil cursor and hand-drawn color. Publish this post. The prompt calls this an object-integrated drawing. I just call it the banana test.
Why it works
The model has to see more than "banana." It has to understand the curve, color, and proportions well enough to find another use for the object. It has to come up with an idea that keeps the banana recognizable, translate that idea into code, draw it convincingly, and make the animation work.
That gives me a quick look at vision, creativity, frontend execution, and whether the model can hold a visual plan together. It is also hard to fake with polish. A creative result is obvious. A broken animation is obvious. A model that just draws some random stuff around a banana is obvious.
It is more fun than reading another benchmark table, and the outputs are different enough that people can judge them without me inventing a score.
The banana tests so far
Newer runs
Here are more banana tests that have come out since I wrote this article.


