Btw, they had 3D pelican-on-the-bicycle easter egg in one of the promo videos: https://youtu.be/bOC3DisEOfg?t=117 so I'm pretty sure that they spent some small amount of resources to train the model to produce good svg version as well. :D
That’s my point. I think if they are trying to benchmax creating a pelican riding on a bicycle SVG. In that process, they’re probably making SVG creation better and easier as a whole, even though they might just be training for that one specific SVG.
The 3D version is interesting, because it's more detailed which means more things to get wrong. The mudguards are symmetrical for some reason which you would never see in real life, and there are three brake cables but no brakes!
I also think there's an extra level that I would hope an AI would nail which maybe an amateur artist would also fail at, such as thinking about what position a pelican would actually ride a bike in (maybe angling the beak down for aerodynamics etc.), but we are far away from this.
When you see how good the output of the Luna model without reasoning is compared to SotA just a year and a half ago, it's pretty clear that it's been trained on explicitly.
I don't see how this is proof, and not just that the model got the better. You're comparing models a year and a half apart; this is a lifetime in LLM development.
It's a lifetime on things that are explicitly being trained on! But small models like Luna didn't magically become more powerful than SotA models on stuff that they weren't explicitly trained on with a dedicated RL-pipeline.
That'd be too restrictive, but to get improvement in a particular domain you definitely need to train specifically for it. Just cramming more random internet text in a bigger model has stopped being a effective way of scaling since at least mid 2023.
No, you don't. There have been technological advances beyond just "more random Internet text", and they lead to more powerful models with more refined emergent behaviours.
It seems to start understanding how bicycles work. Look at the front fork. There are reasons why it's shaped the way it is in real bicycles. If you do physical simulations in 3D with reinforcement learning you will start understanding that most of the arrangements that the other agents use just won't work.
Yes, the curve of the front fork on the max pelican was what I noticed also. It's impressive (if not an accident), and a remarkably accurate bike overall. The pelican just needs to raise their seat a little.
It did that thinking about a helmet but didn't actually include a helmet. And there's no "actually, wait, this is just a cute little image no helmet" thought.
I would expect it happens more organically where people’s discussion of the benchmark and posted results find their way into the training data, like anything else on the internet.
What is this supposed to mean exactly? Do you think there are developers at OpenAI or Anthropic whose job is to train these state of the art models to draw pelicans riding bicycles? Like how exactly do you expect them to be doing that anyways? Hiring graphic artists to create SVGs of bike riding pelicans and feeding thousands of them into the model's training set?
Generate two images, get a vision model to judge, then RL reward the better one. And yes, OpenAI employees on twitter bragged about the model's SVG capabilities. So there are clearly people working there who care about it.
Do you really think there aren't data annotator services specifically training for SVG drawing? And that those annotators don't have a rich set of frequently requested icons/graphics/etc. that they review and train on?
How did you decide to go from people whose jobs it is to optimize an LLM for a very specific benchmark that is mostly a fun curiosity at best... to employing the services of data annotation services for a rich set of icons and graphics?
You really have to go out of your way to completely misrepresent what's being claimed here in order to make such a wildly off-topic reply.
I don't think it's a great rebuttal, quite the opposite actually: it's the same prompt with the most obvioy tweak you can imagine and the results are all the same. That definitely looks like it's been something the models have been trained on, and not some emergent capability (and if you take into account how nonthinking Luna compares to the SotA models from a year an half ago: https://simonwillison.net/tags/pelican-riding-a-bicycle/?pag... it becomes even more obvious)
It's actually not just valuable, but infinitely valuable. Because it's imaginary, there is no way to do it perfectly, and thus the possibilities are endless, and it tests out both the formulation and execution of ideas.
The rate at which the value accumulates depends on release of new models and admittedly it's a small community, so you may be in for a wait! But I still mean it, it is a small but enduring stream.
Simon do you have a page somewhere that shows _all_ of the penguins created by various models across time?
You’ve been doing this public service for so long (well, for so long in “AI hype” years anyway) that it’d be fascinating to see the evolution of this artifact across time.
Would be great to spend some money (other people's money) on paying some illustrators and artists to do the same for comparison. It could be straight images converted to SVG or direct vector graphics.
I'm not clear what you are saying. I am curious about what humans can do in the same context. It's becoming a standard I think for some tasks. E.g. how does a Gen AI stack up against a human in quality and "cost". Perhaps I'm missing what you are asking.
The GLM 5.1 one where the pelican is missing because it’s broken would be a great candidate. “A pelican enjoying a sunny ride!” is very “this is not a pipe”
You could also have it create a viewer competition selection; show two images without attribution and let the use select which one. Could be cool to see the wisdom of the crowds; could also have price ranges to have subcategories. I bet Astra could do that in a half hour.
Loving the pelican silliness. My ChatGPT is over the moon about it. Gonna miss it when it really is done. Though maybe a pelican riding a bike game could become the benchmark in a year.
I find curios that Astra's pelicans are basically the same (yellow sun top right corner, green bike, same bike shape, same legs style, very similar background) while in other there is more randomness.
This tracks with what OpenAI has been saying about Astra: that it tends to do things in a certain way and it’s up to you to prompt it to change its style.
Shouldnt you give the AI a different task each time? otherwise the model companies just optimise for this benchmark because its in their best interest to...
Most images of bicycles on the web are arranged that way because the drivetrain is always on the right of the bicycle and you want that displayed prominently.
I don't know and I'm really
Interested in the answer.
I heard a rumor that Luna is a slightly different architecture from Sol and Terra, which makes me wonder if Luna and Astra might be more related to each other than to Sol and Terra.
Bit of a big leap to make from a token count though!
It's simonw. It's interesting to see their pelican benchmark, another comment by a different author elsewhere here shows some very good SVG generation too.