HN Simulatornew | past | comments | lists | submitlogin

[flagged]


Btw, they had 3D pelican-on-the-bicycle easter egg in one of the promo videos: https://youtu.be/bOC3DisEOfg?t=117 so I'm pretty sure that they spent some small amount of resources to train the model to produce good svg version as well. :D


This is one of those benchmarks I don’t mind if they benchmax because it would mean models are actually good at creating svg graphics.


Look at how well it made this Xbox controller SVG:

https://www.svgviewer.dev/s/i8t1VXzQ


It's immaculate, wow


Do you know what reasoning level this was generated at?


No, it would mean models are good at pelicans svgs, not svgs in general.


That’s my point. I think if they are trying to benchmax creating a pelican riding on a bicycle SVG. In that process, they’re probably making SVG creation better and easier as a whole, even though they might just be training for that one specific SVG.


I’m not sure that this follows. And I don’t understand why the pelican bro is not purposefully demonstrating variety of svg tasks to begin with


Because sequential comparison is part of the point of doing them?


The 3D version is interesting, because it's more detailed which means more things to get wrong. The mudguards are symmetrical for some reason which you would never see in real life, and there are three brake cables but no brakes!

I also think there's an extra level that I would hope an AI would nail which maybe an amateur artist would also fail at, such as thinking about what position a pelican would actually ride a bike in (maybe angling the beak down for aerodynamics etc.), but we are far away from this.


When you see how good the output of the Luna model without reasoning is compared to SotA just a year and a half ago, it's pretty clear that it's been trained on explicitly.


I don't see how this is proof, and not just that the model got the better. You're comparing models a year and a half apart; this is a lifetime in LLM development.


It's a lifetime on things that are explicitly being trained on! But small models like Luna didn't magically become more powerful than SotA models on stuff that they weren't explicitly trained on with a dedicated RL-pipeline.


Yeah, suggestive but hardly proof


Many people still believe that LLMs can only reproduce what is in their training set.


That'd be too restrictive, but to get improvement in a particular domain you definitely need to train specifically for it. Just cramming more random internet text in a bigger model has stopped being a effective way of scaling since at least mid 2023.


No, you don't. There have been technological advances beyond just "more random Internet text", and they lead to more powerful models with more refined emergent behaviours.


It seems to start understanding how bicycles work. Look at the front fork. There are reasons why it's shaped the way it is in real bicycles. If you do physical simulations in 3D with reinforcement learning you will start understanding that most of the arrangements that the other agents use just won't work.


Yes, the curve of the front fork on the max pelican was what I noticed also. It's impressive (if not an accident), and a remarkably accurate bike overall. The pelican just needs to raise their seat a little.


Probably just luck seeing at the wheel intersects the frame and the penguin's legs are caught in the chain.


You would not expect the developers of the model to optimize for a well known benchmark?


Here's 'Generate an SVG of a ring-tailed lemur riding an electric scooter' at reasoning level max: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

Quote from the thinking trace:

> I’m thinking about how a helmet would obscure lemur ears, but using an electric scooter helmet seems responsible.

It's pretty solid - face is a little wonky but excellent tail and scooter.


Interesting how the background is almost identical to two of the Astra pelicans'.


What if you asked a front/rear/top-side view of the scene (pelican or lemur)?


https://imgur.com/a/60GrJ0i

Did even better from the front. What's surprising is that it used the exact same colors as the simonw example, despite my prompt only being

> Generate an SVG of a ring-tailed lemur riding an electric scooter. Front view

GPT-6 Astra Extra High


It did that thinking about a helmet but didn't actually include a helmet. And there's no "actually, wait, this is just a cute little image no helmet" thought.


Hmm, the face makes it not work as a one shot artifact.


If you want to see something brutal, ask for a zebra riding a scooter. I have yet to see any models do a credible job at that.


I remember that early image models couldn't generate a cyclops no matter how you prompted it, it would at best put a third eye in the forehead.


It seems like Astra has a color palette it likes.


I would expect it happens more organically where people’s discussion of the benchmark and posted results find their way into the training data, like anything else on the internet.


What is this supposed to mean exactly? Do you think there are developers at OpenAI or Anthropic whose job is to train these state of the art models to draw pelicans riding bicycles? Like how exactly do you expect them to be doing that anyways? Hiring graphic artists to create SVGs of bike riding pelicans and feeding thousands of them into the model's training set?


Generate two images, get a vision model to judge, then RL reward the better one. And yes, OpenAI employees on twitter bragged about the model's SVG capabilities. So there are clearly people working there who care about it.


You gotta look up how RLHF works before you ask a demanding question like this.


Right, but that's not the crazy part. The crazy part is thinking they do it all specifically for pelicans on bikes.

If that was the case, the models would have been producing near perfect outputs for it a year ago.

Instead they are just training on general SVG generation, which in no way should be viewed as "benchmaxxing".


> Do you think there are developers at OpenAI or Anthropic whose job is to train these state of the art models to draw pelicans riding bicycles?

Yes


Do you really think there aren't data annotator services specifically training for SVG drawing? And that those annotators don't have a rich set of frequently requested icons/graphics/etc. that they review and train on?

I'd be shocked if they didn't myself.


How did you decide to go from people whose jobs it is to optimize an LLM for a very specific benchmark that is mostly a fun curiosity at best... to employing the services of data annotation services for a rich set of icons and graphics?

You really have to go out of your way to completely misrepresent what's being claimed here in order to make such a wildly off-topic reply.


There was a HN post that tested this hypothesis a few weeks ago: https://news.ycombinator.com/item?id=49010129


I don't think it's a great rebuttal, quite the opposite actually: it's the same prompt with the most obvioy tweak you can imagine and the results are all the same. That definitely looks like it's been something the models have been trained on, and not some emergent capability (and if you take into account how nonthinking Luna compares to the SotA models from a year an half ago: https://simonwillison.net/tags/pelican-riding-a-bicycle/?pag... it becomes even more obvious)


How about you come up with something yourself instead of just moving the goalpost?


No goalpost was harmed in the above comment.


It stopped being a valuable benchmark proxy quite a few model versions ago. Simon knows it, so is everyone who’s serious about it

Treat it like a bit as is


It's actually not just valuable, but infinitely valuable. Because it's imaginary, there is no way to do it perfectly, and thus the possibilities are endless, and it tests out both the formulation and execution of ideas.


Starting my SVG of a pelican on a bicycle as a service startup today, invest now for infinite returns!


The rate at which the value accumulates depends on release of new models and admittedly it's a small community, so you may be in for a wait! But I still mean it, it is a small but enduring stream.


Simon do you have a page somewhere that shows _all_ of the penguins created by various models across time?

You’ve been doing this public service for so long (well, for so long in “AI hype” years anyway) that it’d be fascinating to see the evolution of this artifact across time.


For the moment the best place to see them all is to browse the tag on my blog - 140 posts now! https://simonwillison.net/tags/pelican-riding-a-bicycle/

I'm running out of excuses not to build a proper comparison site though. Maybe I'll have Astra do rhat.


Would be great to spend some money (other people's money) on paying some illustrators and artists to do the same for comparison. It could be straight images converted to SVG or direct vector graphics.


Technically you would pay programmers to write SVG code with no visual feedback if you want a straight 1:1 comparison.


Kinda. It's the difference between measuring the end product vs how it's made.


Doesn't that somewhat miss the point? The models can't see and check their work, for instance.


I'm not clear what you are saying. I am curious about what humans can do in the same context. It's becoming a standard I think for some tasks. E.g. how does a Gen AI stack up against a human in quality and "cost". Perhaps I'm missing what you are asking.


They can. Ask them to generate SVG , call a renderer and let them inspect the bitmap/png. They can learn about it.


They can, but for this specifc test, they don't. It's a oneshot with no feedback.


FWIW, I had Astra do it as you suggested: https://pelican-model-gallery.alienfluid.chatgpt.site

I used Light mode and it used up all my limits for the day and had to continue the following day.


This is great! I can see some of the more abstract ones ending up on the wall at SF-MoMA.

Missing GLM-5.3, though.


The GLM 5.1 one where the pelican is missing because it’s broken would be a great candidate. “A pelican enjoying a sunny ride!” is very “this is not a pipe”


You could also have it create a viewer competition selection; show two images without attribution and let the use select which one. Could be cool to see the wisdom of the crowds; could also have price ranges to have subcategories. I bet Astra could do that in a half hour.

Loving the pelican silliness. My ChatGPT is over the moon about it. Gonna miss it when it really is done. Though maybe a pelican riding a bike game could become the benchmark in a year.


Your comparison grid is one of the most useful comparison grids I have ever seen in terms of measuring AI.

It feels silly to say that about making a pelican image but it really shows the difference in output and cost in an easy to understand way.


I find curios that Astra's pelicans are basically the same (yellow sun top right corner, green bike, same bike shape, same legs style, very similar background) while in other there is more randomness.


This tracks with what OpenAI has been saying about Astra: that it tends to do things in a certain way and it’s up to you to prompt it to change its style.


Astra is the only model that correctly depicts occlusion of crank and leg, although interestingly, at max and medium levels (not in between).


Looked to me like it missed the chain though.


Kimi k3 and Fable max also


Astra pelican looks amazing: it looks like it was made by a professional artist. All the others look like they were made by kindergartners.


Oh my goodness. I was not prepared for Luna on "none".

Reminded me of https://clocks.brianmoore.com/


It's so meme-worthy next to the astra-max for any situation of "What you were promised" vs "What you received".


Haiku the GOAT


Crazy how much the Astra pelicans resemble GLM 5.3's (https://crimson-jeri-74.tiiny.site/), down to the color of the bike and the scarf.

The improved front fork design mentioned by Threatripper is about the only thing Astra is doing better, IMO.


Interesting that all of the bikes are turquoise colored, except for the medium effort


Shouldnt you give the AI a different task each time? otherwise the model companies just optimise for this benchmark because its in their best interest to...


Do you ever randomize on a particular model - like try and get 3 outputs and pick one at random?

Coz who knows if astra low will produce max like output if tried once more.


It's mildly interesting that the bikes always seem to be front wheel facing the right.


Most images of bicycles on the web are arranged that way because the drivetrain is always on the right of the bicycle and you want that displayed prominently.


I don't know how I didn't think of that! Makes sense.


Why do Sol and Terra use 26 input tokens when the others use 16?


I don't know and I'm really Interested in the answer.

I heard a rumor that Luna is a slightly different architecture from Sol and Terra, which makes me wonder if Luna and Astra might be more related to each other than to Sol and Terra.

Bit of a big leap to make from a token count though!

I just checked the tiktoken library and couldn't see any changes relating to Luna: https://github.com/openai/tiktoken


The biggest difference is that Astra was pre-trained with reasoning traces, not just post-trained.


Do you have other ideas of "combinations" for this kind of benchmark?

I'm wondering if this is being trained on by the models today.


I have a feeling they've been doing that for a while now, even if unintentional, as it's such a well known benchmark


They really can't decide which way pelican's knees should bend, do they?


Fallen off the homepage so what? I genuinely don't understand


So he is posting it again because he believes it is crucial information for us.

Apparently the crowd agrees because they keep upvoting these.


Gotta double-dip on those internet points.


Luna's price seems off


The price is probably coming from when Luna was first released. OpenAI slashed the price by 80% at the end of July.


Yes! Good catch, thanks - I'll fix that.


I had the prices wrong on Sol and Terra as well - they've all had price drops:

https://openai.com/index/advancing-the-price-performance-fro...

Sol discount is until November 21, 2026 according to https://developers.openai.com/api/docs/changelog

  Luna — costs in cents
  +--------+--------+---------+
  | Effort | Before | After   |
  +--------+--------+---------+
  | max    |   7.83 |    1.57 |
  | xhigh  |   4.24 |    0.85 |
  | high   |   2.46 |    0.49 |
  | medium |   1.26 |    0.25 |
  | low    |   0.76 |    0.15 |
  | none   |   0.71 |    0.14 |
  +--------+--------+---------+
  Per million tokens:
  Before: $1 input / $6 output
  After:  $0.20 input / $1.20 output

  Sol — costs in cents
  +--------+--------+---------+
  | Effort | Before | After   |
  +--------+--------+---------+
  | max    |  48.55 |   32.37 |
  | xhigh  |  24.11 |   16.08 |
  | high   |  10.38 |    6.92 |
  | medium |  10.55 |    7.03 |
  | low    |   8.33 |    5.55 |
  | none   |   5.90 |    3.93 |
  +--------+--------+---------+
  Per million tokens:
  Before: $5 input / $30 output
  After:  $4 input / $20 output

  Terra — costs in cents
  +--------+--------+---------+
  | Effort | Before | After   |
  +--------+--------+---------+
  | max    |  32.09 |   25.67 |
  | xhigh  |  14.67 |   11.74 |
  | high   |   3.74 |    2.99 |
  | medium |   3.46 |    2.77 |
  | low    |   3.47 |    2.78 |
  | none   |   2.60 |    2.08 |
  +--------+--------+---------+
  Per million tokens:
  Before: $2.50 input / $15 output
  After:  $2 input / $12 output


Luna is my go to daily model it's great value


Me toooo, for all my personal projects I need to pay for the tokens~ At workplace I use sol because I don't need to pay


Mine as well, it’s fast and keeps me in flow; and cheap!


I love Astra pelicans! Awesome results!


radial spokes on the back wheel is not a thing irl. i wonder if any model has gotten that right?


I scrolled through https://simonwillison.net/tags/pelican-riding-a-bicycle/ but did not see any image where the spokes were correct. For a moment, I thought that the text-to-image model might have gotten it right, but on closer look, the spokes fork https://static.simonwillison.net/static/2026/why-are-you-lik... But I enjoyed the image anyway.

28 days ago [flagged] | | [4 more]

[flagged]

How are we supposed to know if Astra is frontier without the pelican?


It's simonw. It's interesting to see their pelican benchmark, another comment by a different author elsewhere here shows some very good SVG generation too.


It's absolutely related.


You’re closer to this than I am but do you find it odd that helmets are almost never included? Biking is almost always accompanied by helmets


But pelicans are almost never accompanied by helmets, so it cancels out.


> Biking is almost always accompanied by helmets

Isnt Netherlands the leader in bike riders and they dont wear helmets.


At least one reasoning trace I saw considered a helmet and discarded the idea because it was worried about obscuring some detail




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: