NHacker Next
login
▲Are AI Labs Pelicanmaxxing?dylancastillo.co
310 points by dcastm 5 hours ago | 126 comments
Loading comments...
simonw 4 hours ago [-]
This is fantastic

I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations.

Catching a lab cheating specifically on my one dumb benchmark would be really funny.

Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than anything I was considering.

His conclusion:

> Nothing jumped out at me. I couldn’t find a case where the pelican-bicycle images looked noticeably better than the rest of that model’s grid.

lukev 3 hours ago [-]
What if they’re not pelicanmaxxing, but svgmaxxxing in general?

Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge.

Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.

qq66 3 hours ago [-]
But that's a genuine worthwhile capability. It's like benchmarkmaxxing on a weightlifting competition by getting really strong.
Balgair 2 hours ago [-]
https://www.youtube.com/watch?v=jgYYOUC10aM

reminds me of this Key and Peele skit

ryukoposting 2 hours ago [-]
How useful actually is this? It generates SVGs of pelicans on bicycles, sure, and some of them are (almost) spatially correct. But, none of them look good.

AI image generation suffers from this more generally. You can generate pictures of pelicans, sure. Newer models clearly generate images with more pelican-ness than before. But all of it is still uglier than sin. Drawing things accurately is one thing, making results that someone might actually want to use (without embarrassing themselves) is something else.

derektank 2 hours ago [-]
Just today I asked Claude Fable to take an SVG map and add a circle with a 150 mile radius around a specific city, cutting off the circle at the edge of certain boundaries. It came back a minute later with more or less exactly what I wanted, saving me maybe 5 to 10 minutes of photoshop time. It’s almost certainly not perfect, but I didn’t need it to be perfect, just a legible representation for a group of 20ish people.
cmckn 53 minutes ago [-]
I used Claude to plan some landscaping, and it gave me SVG diagrams of flower beds showing where to place plants and their approximate mature size. It was pretty useful.
Kamshak 2 hours ago [-]
Great for charts, illustrations etc, can do much more with SVG than pretty pictures. Good spatial understanding in an SVG can be really nice for model to have
zahlman 2 hours ago [-]
Even if the models don't break through any particular "uglier than sin" barrier, with a bit more work, presumably the SVGs could become importable into an editor that would let a human apply taste and discretion. Seems to me like a heck of a head-start.

As for conventional diffusion-model stuff, I happen to think there are some pieces of AI art that still look really good even knowing they're AI.

fiddlerwoaroof 2 hours ago [-]
AI can generate a fairly satisfactory SVG for a favicon now (programmer art quality at least).
amarant 2 hours ago [-]
Vibecoding a SVG based metroidvania as we speak! This is gonna be lit!
OutOfHere 2 hours ago [-]
The fear is that the SVGmaxxing is limited to "X doing Y". If such 'template maxxing' exists, it will break for other templates, e.g. "X not doing Y", "X and Y doing Z", "X doing Y doing Z", etc.
sysguest 2 hours ago [-]
well that holds IF svgmaxxing is 100% "code-writing-maxxing"

...which.. hmm I dunno if they are same or not

tsimionescu 2 hours ago [-]
No, the point is that a general-ish ability to draw good SVGs is a useful ability in itself. People need SVGs for all sorts of purposes, and if AI can generate one for them, that's mostly useful (discussions about art and employment etc notwithstanding).

That said, I think this would correlate relatively little with general programming ability. They're not unrelated, of course, but being able to generate code that paints an accurate + esthetically pleasing image is quite different from generating code that achieves a non-spatial goal.

runarberg 22 minutes ago [-]
If people need to generate good-ish SVG why not simply use a specialized model for a much better result and for far cheaper and quicker?

Why do LLMs need to be able to do this as well, but worse, slower and more expensive?

yorwba 3 minutes ago [-]
Indeed, if you want to generate an aesthetically pleasing SVG, you'd be better off asking a pixel-based image-generation model for "vector art" and then deconstructing it into an equivalent SVG with something like LayerPeeler. https://layerpeeler.github.io/
exhaze 11 minutes ago [-]
Agree directionally - even back in sonnet 3.5 days, I was helping some friends by showing them how to create intermediate representations for SVG building blocks mapping to parametrizable functions that can be used to make interactive SVG-rendered visualizations for various medical needs.
wasabi991011 2 hours ago [-]
I don't see why that's true. LLMs don't have to only be good at code-writing.
exhaze 14 minutes ago [-]
Not an expert on this, so won’t speculate regarding what traits would create robustness specifically attributed to SVG visual representation capability, but felt you might find this paper on reasoning models trained on physical world video data becoming better at general reasoning interesting:

https://arxiv.org/abs/2210.05359

timClicks 1 hours ago [-]
If they're optimizing for SVG generation, then that's an excellent outcome in my opinion. Vector images shouldn't be "pretty niche".
knollimar 47 minutes ago [-]
Please it's almost my primary bench for drafting understanding.

I'm almost about to post that xkcd 810

brikym 2 hours ago [-]
Then it's great. A year ago I couldn't get any AI to draw a simple company logo in SVG or even convert from raster. No doubt the Pelicans put pressure on the labs to fix the awful SVG situation. Now I can even make a decent Peli in 3D.
simonw 53 minutes ago [-]
Gemini have absolutely been SVGmaxxing. They've openly talked about it.
netsec_burn 3 hours ago [-]
Addressed in the article, in case you're curious.
lukev 3 hours ago [-]
Well, it’s mentioned as a limitation of the analysis, very much not ruled out (or in.)

That simonw is causing labs to do extra fine-tuning runs for this seems highly probable :)

kaliqt 2 hours ago [-]
Funnily enough, not that niche, because I have tried many times to do it as part of a wider project.
kelvinjps10 2 hours ago [-]
SVG are really useful you can create images that don't have the AI look
charcircuit 3 hours ago [-]
I agree, other formats, both textual and binary should be tested.
exhaze 20 minutes ago [-]
What if they’re just unsuccessful at it and something about a model grokking how to create a realistic visual representation of a pelican on a bicycle ends being key to the next OOM of capability unhobbling.

And since we have established this silly routine once, I must keep going and ask - yes, you have quite a collection of bona fide pelicans you’ve seen and photographed.

Have you physically seen all the other animals you’ve evaluated as well?

I was born in Soviet Russia, so trust but verify and if still around, maybe there pelican make LLM draw Simon ride bicycle and notice Python code improve.

eob 2 hours ago [-]
Simon I hope from this day hence, your bio always includes:

"Simon Willison, among other things, is an advocate for the inclusion of pelican geometry in LLM training datasets."

exhaze 8 minutes ago [-]
“Today, the NASDAQ fell by 18% moments after Simon Willison published his latest blog evaluating the Legend 7.4 model that rendered an animated SVG of a pelican on a bicycle and the pelican fell off”
gopalv 3 hours ago [-]
> Catching a lab cheating specifically on my one dumb benchmark would be really funny.

Similar thing happened when TPC came up with SQL benchmarks.

If you're not good at TPC, your engineering team is no good.

If you're good at TPC, then (as a customer) we will actually include you in a bake-off benchmark for our specific problem.

Winning on it is the price of admittance into the game, especially in a crowded market.

But how narrowly you benchmarket matters, you can't just hard-code that specific scenario & not fix anything adjacent while you're at it.

For example when it comes to GPUs, the "Quack3" (sic) benchmark on ATI cards comes to mind.

ryandrake 10 minutes ago [-]
Yea, I worked for a competitor to ATI back in the day and we were definitely Quakemaxxing. We were not so unethical as to try to detect the .EXE name so to put the GPU into a "cheating" mode only when that benchmark was run, but we did run Quake 3 pretty much constantly while trying to eke out a few more FPS...
exhaze 4 minutes ago [-]
Wasn’t that almost the norm for both new OS drivers and subsequently GPU SoC firmware of ‘gamer cards’ to literally optimize for the latest AAA titles? I feel in that case, the interests aligned - in a world where I had a chance to actually play Crysis, I wouldn’t care why the frame rate was decent, no?
docheinestages 2 hours ago [-]
I think a more fundamental test is SVG art creation in general. Perhaps a pipeline to take any image, caption it, ask the LLM for an SVG, rasterize to an image, and finally either use a deterministic visual similarity check or ask another LLM to be the judge and score how close the SVG is to the original image.
zahlman 2 hours ago [-]
Fidelity to the original is definitely not how humans would measure "art" in this context.
docheinestages 2 hours ago [-]
True, maybe we can call the generated SVG something else than art.
gilleain 3 hours ago [-]
Perhaps also vary the bird? Wikipedia tells me pelicans are in the order _Pelecaniformes_ so shoebills or herons might do.
cyberax 2 hours ago [-]
> I've been casually spot-checking other animals in other vehicles

Snakes on a plane, weasels on a diesel, spiders on a glider, baboons on a balloon, goats on a boat.

mattertoast 3 hours ago [-]
[dead]
mauvehaus 3 hours ago [-]
> All 21 pelican-bicycle images, across all seven labs, face right. No other animal/vehicle combination does that.

> However, facing right is common: 60% of all 1,008 images do it. How common depends on the animal and the vehicle, and bicycles are one of the two vehicles where it’s strongest

Of course the pelican on the bicycle is facing right. The drivetrain on a bicycle is on the right side. If you want any representation of a bicycle that shows the drivetrain you're going to show the right side of it if you want to do so without the frame occluding it. It's an excellent bet that their training data reflects this.

Citation: https://www.rei.com/c/bikes

Edited to add:

As near as I can tell, all of the bicycles are shown facing right, regardless of the direction the animal is facing (GPT 5.6-Terra, Sample 1/3). Also, in every case where the rider has legs (i.e. not the whale) both of the rider's legs are on the right side of the bicycle. This suggests a pretty serious lack of actual understanding of how a bicycle works.

stusmall 4 hours ago [-]
I'm glad someone ran the numbers on this. Every single Simon Willison post of an SVG is followed with someone dismissing it saying "I'm sure they train on it by now." This is despite a good blog post with sound logic on how easy that is to catch. [1] Glad to see someone took the time for a quantitative analysis of dumb little animals riding dumb little bikes.

1. https://simonwillison.net/2025/Nov/13/training-for-pelicans-...

unholiness 3 hours ago [-]
I don't think this small amount generalization to other animals and vehicles is strong evidence they haven't trained on this, either directly or more generally.
conception 1 hours ago [-]
Honest question how could they possibly train on this as there are no good SVG pelicans to train off right? So they’re just training off a bunch of bad ones which should lead to just bad pelicans, but the pelicans are getting better.
alexthehurst 1 hours ago [-]
It’s not hard for a visual model to score the quality of that output though, which would be a pretty good fitness function.
elicash 2 hours ago [-]
There's that version of the argument version, but there's also the softer version: that there used to be no training material of illustrated pelicans on bicycles, but now you have actual artistically talented individuals drawing it and that could improve the performance even though the AI labs are sucking it up no differently than everything else.

This post proves that hasn't happened yet, either. Although maybe the bad results posted online are being trained on and that explains the UNDER performance.

SyneRyder 33 minutes ago [-]
Huh. They're not "Pelicanmaxxing"... they're Ottermaxxing.

Take a look at the GLM 5.2 and Deepseek V4 "animal on a plane" examples. In every case, the animal is standing on top of the plane, a clear misunderstanding of the concept of "animal on a plane"... with the exception of the Otter. The otters are sitting in a seat on the plane, looking out the window.

That's Ethan Mollick's "Otter On A Plane Using WiFi" image benchmark.

https://www.oneusefulthing.org/p/the-recent-history-of-ai-in...

(Sometimes the Racoon is sitting inside the plane as well, but the racoon is a common backup benchmark. I'm surprised it wasn't also holding a sign saying that it loves trash.)

Also, Grok seemed to really really enjoy "whale on a plane" in that second round, and kudos to GPT Terra for deciding after 3 rounds that the user was terrible at spelling and generated "Antelope On A Plain".

EDIT: I promise I'm a human, but I did just notice my "that's not x... that's y" construction at the start. I am rather Claudepilled — my apologies.

waterproof 1 hours ago [-]
I recently had the following conversation with Claude:

Me: how many P's are in the following text? [Pasted text]

Claude: There are 14 P's, all lowercase (no capital P's)

Me: how many in "strawberry"?

Claude: there are 3 R's in the word "strawberry".

ickyforce 55 minutes ago [-]
I tried this with Gemini 3.6 Flash

Me:how many P's are in the following text? [Pasted text with 10 P's]

Gemini: There are 9 "P"s (1 uppercase P and 8 lowercase ps) in the provided text. [List of words except the one missed]

Me: How many in strawberry?

Gemini: Something went wrong (1096)

bnfcl 3 hours ago [-]
This is funny, I actually did a similar experiment just yesterday.

Looking for evidence of the same, but with another twist: checking if the models would choose to create a pelican on a bicycle, if no specific bird or method of transportation was specified.

My version of it: https://www.modelbias.ai/pelican-on-a-bicycle-test

michaelt 1 hours ago [-]
The outputs of Qwen3.7 Plus have to be seen to be believed:

https://www.modelbias.ai/pelican-on-a-bicycle-test/result/12... https://www.modelbias.ai/pelican-on-a-bicycle-test/result/12...

zahlman 2 hours ago [-]
Simple as they are, there are some really aesthetically pleasing penguins on skateboards in there, including from less capable models. (In fact, I would say the Opus series got progressively worse at it over time.)
wasabi991011 2 hours ago [-]
I find your analysis much more convincing than TFA, since it doesn't require a subjective evaluation and is more robust to animal/transport complexity.
bnfcl 2 hours ago [-]
Thanks! Because I think that models are becoming better at creating SVGs in general. If you look at Claude Fable 5 and Kimi K3 for example.

In my tests it did create bicycles the most, but this is just a general bias I believe, as tested here: https://www.modelbias.ai/prompt/transport

Wowfunhappy 4 hours ago [-]
> The more plausible story is SVGmaxxing

Exactly--and you have to ask yourself at this point what "maxxing" really means, since "get better at drawing SVGs" is a useful skill.

beering 4 hours ago [-]
Really awful how the AI labs are skillmaxxing /s

Pelicans aside, we need to remember that benchmarks are the only good quantitative way we have of comparing models. If someone has complaints about “benchmaxxing”, please ask them to contribute a better benchmark! It is valuable work and very appreciated.

dbt00 3 hours ago [-]
It's a problem because of Goodhart's law.

If you train towards the test, you aren't necessarily improving overall fitness, but you are destroying the value of that test over time because you're decreasing its correlation with overall fitness.

Wowfunhappy 3 hours ago [-]
> If someone has complaints about “benchmaxxing”, please ask them to contribute a better benchmark!

I don't think I'd go that far!

When someone says a model has been benchmaxxed, what they really mean is that it performs better in benchmarks compared to their real world experience. That's a real thing, I've certainly experienced it with some models.

...my take is that some things in life just resist quantitative measurements. Who is the best job candidate? What is the best programming language? Add AI models to the pile.

richardw 1 hours ago [-]
And a whole universe of random tests got baked into the AI’s training data. Websites of antelopes driving trains and hammerhead sharks swinging in a tyre swing were created. It was a short while until AI became so focused on animals that it gave up competing with developers. Life became sane again.
oasisbob 45 minutes ago [-]
As a unicyclist, all the sideways-riding caught my eye for being especially silly.

However, I'm very surprised that most of the models make the same sideways mistake with only some of the animals, and they do it consistently.

With most of the models, cat, raccoons, and otters are almost always riding sideways. Why is that?

simonw 3 hours ago [-]
Underlying data is available on GitHub: https://github.com/dylanjcastillo/blog/tree/main/_extras/pel...
dllu 3 hours ago [-]
I feel like getting LLMs to spit out an SVG is akin to getting a human artist to draw something by just reciting a list of coordinates. It's insanely hard and unnatural.

Image generation models nowadays can easily generate a photorealistic pelican riding a bicycle, where the bicycle has perfect structure. But it is, of course, only a raster image.

It seems that we're missing a kind of step to decompose an image into a list of instructions (say, SVG paths, or even brush strokes with a real brush) to reproduce it properly. Doing so would probably need a true understanding of the structure of the scene, which is something that AI still struggles with to this day.

staticshock 3 hours ago [-]
The pelican on a bicycle test is specifically about generating an SVG, fyi, not a raster.
dllu 3 hours ago [-]
I know. I'm just thinking about how to make AI create SVGs better... in theory, a sufficiently smart AI could "generate an image in its head", think about it, and then output the SVG paths to produce said image. Intuitively that would be somewhat closer to how human artists convert artistic visions into a sequence of arm movements while holding a brush (obviously, humans don't hold a fully formed, photorealistic image in the head while drawing, but rather vague concepts, but still).
cherioo 2 hours ago [-]
I don’t quite agree. Good human artist can visualize in their mind how to draw a picture, i think. Which i think is no different than LLM doing SVG drawing in their “head”. Anthropic’s recent post call this head-space “workspace”.

It just might feel foreign to human who does not have a SVG trained head-space.

dllu 2 hours ago [-]
Sorry, my comment was confusingly worded, but I am precisely thinking about the head-space thing you are talking about. My point is: the fact that image generation models can generate a perfect pelican riding on a topologically correct and highly detailed bike frame [1], suggests that AI does have the capability to correctly understand it internally. But something is disconnected when we just let an LLM directly output the SVG, and all the nuanced understanding is lost.

As such, modern LLMs still kind of suck at generating pelican bike SVGs with obvious errors:

* some omitted the bottom of the diamond which connects from the pedals to the rear wheel

* some added an extra connection from the pedals to the front wheel, making it impossible to steer

* none could align the head tube with the fork

* none added a correct offset to the fork

* none could generate the chain properly in a way that attaches to the two sprockets correctly

... whereas these errors do not appear in the raster image. (To be fair, the raster image has other weirdnesses, like the bird having arms and only one leg)

If we could harness the sort of internal thinking that must have happened when it generated the highly consistent raster image, but make it output SVG instead, we'd get much better SVGs. So this would be indeed "visualize in their mind how to draw a picture" before outputting the SVG.

[1] https://chatgpt.com/s/m_6a611c29c02481918fbb0f165eb83594

0x000xca0xfe 3 hours ago [-]
Image models that support text output like Image2, or general text models that can read images like Claude can vectorize raster images. But they aren't very good at it, doing it manually in Inkscape still produces better quality even when done by non-artists.
apwheele 3 hours ago [-]
So this is not my experience at all for asking about simple SVG icons for web-pages. Here is one of the examples I have tried for in the past, make a simple cartoon SVG knife for a map icon for a crime map.

https://x.com/CrimeDecoder/status/2080008114615537766

Can see the images for ChatGPT/Claude (Sonnet 5), and Gemini are all quite bad.

Jagged edge of LLMs. How do you explain being able to generate very complicated shapes in the Pelican example but cannot make a much simpler icon without just alluding to it is in the training data?

Topfi 1 hours ago [-]
Thanks for sharing another solid data point. I fear you won't get an answer from my experience [0]. Unfortunately, the blog post decided to forgo the very models that I found to be the worst offenders:

> Here is what Qwen3.6-35B-A3B via Openrouter provided for a sloth riding a skateboard: https://imgur.com/a/Dy8fvR5

> Like Grok 4 Fasts attempt at a mushroom in a rowboat, it is barely recognisable as anything despite both Qwen3.6-35B-A3B and Grok 4 Fast having no issue with more popular (i.e. benchmarked) examples. [...]

> And here is Opus 4.7 [which simonw claimed to provide a worse pelican vs Qwen], again via Openrouter: https://imgur.com/a/Qus1Enf

Anyone who hasn't witnessed such deltas either hasn't looked at enough examples, a sufficient variety of models, or both. And they are, unfortunately, not limited to "SVGMaxxing", but a wide range of evals.

[0] https://news.ycombinator.com/item?id=48951229

robocat 3 hours ago [-]
Does asking for a dagger help?
apwheele 2 hours ago [-]
If you look at the raster image ChatGPT generated, that is fine. It is just this example (and other simple SVG icons I have asked for) result in pretty bad SVGs. It just makes me highly suspicious that the LLMs are learning shape primitives and extrapolating to new shapes, vs just having a big dictionary of prior examples and stitching them together.
robocat 56 minutes ago [-]

  Don't judge the dog's technique. The miracle is that it's dancing.
I would guess most programmers struggle to create SVG icons - I don't find it easy. The average person even more so.

Are we best to assume an LLM is a blind programmer? Any HN comments from blind programmers tasked with creating SVG icons? Only relevant comment I could find from ctoth was about accessibility: https://news.ycombinator.com/item?id=7185771

Projecting how you think onto what the LLM is doing or should be doing, is probably a mistake on your part.

I spent a little time trying to understand exactly why Gemini was misexplaining $X. $X = {why the generated LaTeX didn't match what it was asked to do}. It was enlightening.

ertgbnm 3 hours ago [-]
I've had the feeling that labs aren't pelicanmaxxing specifically but that they do have some sort of RL environment for SVGs that they are letting the AIs overcook in. Specifically I'm thinking of the gemini 3.1 pro annoucnement that seemed to have a huge leap in animated SVG performance but not much else impressive about it.

So they aren't pelicanmaxxing but they are benchmaxxing in a way. The benefit of the pelican was originally that uplift on the pelican signaled an overall uplift on intelligence. I don't believe that is the case anymore and it is just another jagged edge of model intelligence.

munk-a 3 hours ago [-]
It's a method to grade LLM output - as such it's something that will receive focus in correcting for. As soon as people who have a say in where funding is going noticed it as a metric the labs started caring about their performance in it. In the best case the labs are focusing on improving SVG capabilities in general and optimizing Pelican production as part of that initiative - but now that it's a known measure it is no longer reliable.
oaxacaoaxaca 2 hours ago [-]
Hilarious question. Imagine someone woke up from a 7 year coma and read this title lol
robviren 2 hours ago [-]
Why let your dreams be dreams? This is a perfect example of following a hypothesis. I love when people dive into an esoteric subject and just go full swing. Reminds me a CGP Grey and the name Tiffany. Sometimes you just need to know.
bluealienpie 2 hours ago [-]
AI rating AI? Am I missing something.
Nnnes 58 minutes ago [-]
I'm surprised (but not really) that you're the only comment I see even mentioning it. The ratings may be even lower quality than the SVGs.

Obviously they're all a bit cartoon-y, what else do you expect from SVGs. But I'm not convinced you could find a single human on Earth over the age of 4 who would seriously give the vehicle in GPT/1/whale/plane a 5/5.

Browse through the options a bit and the rest is not that much better. Grok/2/cat/plane, one of the more accurate planes, got a 2/5. For the most part, vehicles entirely missing do get a 1/5, except for whatever it is in Gemini/1/heron/plane scoring 4. Animals inside planes get completely random vehicle scores I guess.

The cats are all orange, except for a few of the skateboard cats that are black. I'm sure there's nothing to read into there...

Well, I've convinced myself that the next effective test of multimodal models will be whether their judgments of LLM-generated SVG airplanes are anywhere close to reasonable.

zahlman 2 hours ago [-]
That seems to be how we signal "objectivity" nowadays.
recursive 1 hours ago [-]
Missing something? You're missing the boat! Have some AI-prepared koolaid before you get left behind.
johndough 4 hours ago [-]
Another point for consideration: Specialized SVG models create way better looking pelicans riding a bicycle. (E.g. Refract V4: https://jumpshare.com/s/8liB7Aiuoo3yucbWGXjZ mirror: https://postimg.cc/McV70p84 )
solarkraft 4 hours ago [-]
That’s an impressive image, but what a mistake it was to click the second link (on mobile without an ad blocker). I wouldn’t send it to anyone I respect ...
johndough 1 hours ago [-]
Thanks for pointing that out. I haven't noticed any adds in years with Firefox and Ublock Origin extension.

I'll look for a better image host in the future. I guess the economic incentives makes them all turn bad after a while.

ACCount37 4 hours ago [-]
The name is "Recraft V4", and from looking it up: yeah, it sure seems like whatever black magic they use for SVG generation kicks ass.
johndough 1 hours ago [-]
Oops, autocorrect. Sorry about that.
nostrademons 2 hours ago [-]
It's really refreshing to see someone publish a null result.
Gander5739 1 hours ago [-]
Relevant xkcd: https://xkcd.com/2020/
scosman 3 hours ago [-]
join me in building the ideal training set for pelicans riding bicycles: https://github.com/scosman/pelicans_riding_bicycles
BeetleB 3 hours ago [-]
Oh great! You've now made it a lot easier for LLMs to train on this dataset!

Your next iteration will need different animals and different transportation options. You'll run out after a few iterations.

anuramat 3 hours ago [-]
"benchmaxxing by generalizing" is not really benchmaxxing
jonatron 4 hours ago [-]
OK, so we've done animal_vehicle, how about new SVG ideas each time? I just tried "make an SVG of a man sitting in a chair at a computer behind a desk" which gives more interesting results than the animalVehicle test.
ninju 4 hours ago [-]
There probably good set of images of that description already so it does exercise the inference capability of the model
andy99 4 hours ago [-]
If an AI researcher was going to pelicanmaxx, they would almost certainly apply the augmentations mentioned in the article during training, e.g. randomly selecting animals and conveyances. You’d want a model that generalizes well, just sfting in that specific prompt would be pretty bush league for a frontier lab.

I don’t have any reason to believe they are gaming the benchmark, just saying. I do find the idea of a data labeller having to generate thousands of svgs of different animals on different modes of transportation quite funny though.

cute_boi 4 hours ago [-]
At this point, I think there are so many pelican images in the pretraining data that drawing a pelican no longer makes sense as a model evaluation task.
Rooster61 4 hours ago [-]
I find it humorous that the animal + plane combo appears to be such an outlier. I assume this is due to the models assuming the user mean plain and misspelled it in the prompt.
ramses0 3 hours ago [-]
I think it's actually due to "pelican on a plane" isn't the same as "pelican on an airplane" (Sonnet5 @ Flamingo x Plane), some consistent and warranted semantic/linguistic confusion!
flsw 2 hours ago [-]
I noticed this happens especially with herons. My guess is it's because the model links "heron" to Heron's formula and the Cartesian plane
NitpickLawyer 3 hours ago [-]
GLM has 2 combos of "on a plane" literally sitting inside a plane, with a window and a bit of wing showing. That's funny.
zahlman 2 hours ago [-]
... Is that not how it should be interpreted?
comrade1234 3 hours ago [-]
Hilarious. Could you imagine being a programmer at an AI company and this is your assigned task?
stri8ted 3 hours ago [-]
You seem to assume training on pelican would not result in improved performance on other similar tasks. Why?
altcognito 3 hours ago [-]
He didn't. That's why the article exists. You have to do the science to see if it does.

He was asking the question - do we see gains across other tasks? The underlying question was: Is the additional attention given to this specific task creating a false impression of progress?

HarHarVeryFunny 2 hours ago [-]
Either you've memorized the outline (or detailed component shapes) of, say, a horse, or you haven't. Memorizing the outline of a pelican isn't going to help you with the horse.

You could train a model to do something a bit different like a pencil sketch, or vector graphic sketch, of something given a photo of it, and expect that to be a generalized skill, but if you are asking the model to do it "from memory" then memorizing a pelican is no substitute for not having memorized a horse.

tomas789 4 hours ago [-]
Having an objective score is quite difficult. Maybe it would be better to do a pairwise comparison and calculate ELO?
NitpickLawyer 3 hours ago [-]
Just click through the models. At a glance (and highly subjective) I don't see anything jumping out as oom worse than anything else. I only noticed a model placing the animal inside a plane (with seat and small window) but other than that, they all seem similar inside each model to me.
javier123454321 4 hours ago [-]
If you want to, go ahead, but it seems to me the author already exceeded the energy expenditure that this question warranted.
RobRivera 2 hours ago [-]
Chasing metrics Chasing dragons

Tomato, tomato

IshKebab 1 hours ago [-]
Thanks for not using AI to write this. So much more pleasant to read.
dcchambers 4 hours ago [-]
It's incredible that each model has it's own style that remains relatively consistent throughout all of the different generated examples.
busymom0 3 hours ago [-]
How does attempt 2 by Llama 4 Maverick look like a bald eagle??
ck2 3 hours ago [-]
I am not sure if this is how it works but let's say there was a reddit thread talking about the pelican benchmark and in it someone posts mockup examples of what an ideal result would look like

aren't some LLM going to digest that thread at some point and indirectly learn from it?

basically my point is originally this was a good benchmark because it was an absurd never-seen-before thing, but now that it is in content, some models are going to get a benefit in education?

you'd need the "AI" equivalent of an old-school "google whack", something with no previous results

* https://en.wikipedia.org/wiki/Googlewhack

gbalduzzi 3 hours ago [-]
They ingest so much data that a couple of reddit threads do not move the needle.

It is the reinforcement learning that produces more tangible results with less data, but it is something that the AI labs specifically selects and it is not picked up unknowingly

j45 4 hours ago [-]
The models definitely seem to pay attention to the tests.

Since the tests can be generally gamed with directing descriptions at it non-deterministically, there's a greater chance the questions solution can be found.

Of course, hopefully the models are instead adding patterns and types of questions as well and it makes the models more capable, but it may be limited in how it transfers to other types of questions in breadth or depth.

cute_boi 4 hours ago [-]
https://playcode.io/blog/macbook-svg-benchmark

I think we should stop using pelican benchmark.

dllu 3 hours ago [-]
I disagree with this in the blog post:

> Every single one is a pelican, on a bicycle, first try. When every student gets an A, the exam has stopped grading.

Numerous pelicans and their bikes are clearly horribly malformed. In fact none of the bike frames are correct. Fable and Opus come close, but the top of the diamond is disconnected in Fable's case and the head tube is misaligned with the front fork in Opus's case.

And of course, as the parent post shows, labs don't actually seem to be training on the pelican bike case.

ErrantX 3 hours ago [-]
Agreed. And more; the Macbooks are pretty much the same - some are god approximations, some are terrible, all of them are recognisably a MacBook. And if you start using it they can train on it.

The problem isn't the test, its that is a public test.

Simon has previously said he has a list of secret prompts (at least one of which he "burned" as a demonstration a while ago). That's what makes it a good test - his commentary on the public test is something of a proxy for non-public tests. This makes it a good benchmark.

andrewstuart 3 hours ago [-]
The pelican prompt is ridiculous.

Test the LLLM against things you want it to do.

Asking questions that are absurd is like interviewing developers and asking absurd questions on the grounds that it tests creative and critical thinking.

Remember these Microsoft interview questions designed to identify the best developers?

"If you could eliminate one U.S. state, which one would it be?"

"How would you move Mount Fuji?"

Absurd interview questions have an air of legitimacy due to the quasi sophisticated justifications put forward for why they are good tests.

Absurd interview questions are not good tests of people or LLMs.

Relevant questions are good tests.

user- 2 hours ago [-]
The whole point of "AI" is arbitrary task completion. Why isn't a SVG drawing relevant for that?
simonw 3 hours ago [-]
> The pelican prompt is ridiculous

Yes, deliberately so.

It was never intended as a meaningful benchmark. The surprising thing was that for the first ~12 months performance on the stupid pelican benchmark did seem to correspond to the performance of the models on other tasks.

That pattern no longer holds - Fable 5 and GPT-5.6 have both been out-pelicaned by lesser models now.

zahlman 2 hours ago [-]
> Asking questions that are absurd is like interviewing developers and asking absurd questions on the grounds that it tests creative and critical thinking.

Yes; what's wrong with that?

Do you suppose that it doesn't test those qualities?

ErrantX 3 hours ago [-]
As I understand it; the point is to ask for an SVG which would demonstrate a conceptual understanding of what is being asked for and that is an important test IMO.

What sufficiently hard, but useful, problem would you ask the model for?

BigTTYGothGF 3 hours ago [-]
> Test the LLLM against things you want it to do

I agree, it is ridiculous to ask an LLM to replace an artist.

TZubiri 3 hours ago [-]
https://en.wikipedia.org/wiki/Goodhart%27s_law

"Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes."

Or the more pop layman version

"When a measure becomes a metric/KPI, it ceases to be a good measure."

Story time, I live in Argentina, and we don't have Big Macs, the main Mc Donald's brand, here, because during the CFK presidency, one of her tactics was to Goodhart economic metrics. Even the informal obscure ones like the [Big Mac Index](https://en.wikipedia.org/wiki/Big_Mac_Index), I don't know the precise details, but the Big Mac ended up being a very cheap item, like 2 or 3 times cheaper than actual menu items, but it was never on the advertised menu, and it also ended up being very small compared to the other burgers, so it wasn't even like a hack, a shrinkflation type of deal.

But hey, anyone who read the Big Mac Index table would never find Argentina at the bottom of that list along with a couple of other countries with bad brands, so the ploy worked. And now we live with the aftershock, the brand never really turned around, other brands with ridiculous names took over it like the McTasty, which makes me sound like that skit from Tarantino's Pulp Fiction.

zahlman 2 hours ago [-]
How did the president manage to influence McDonalds' local business decisions? And how did that lead to McDonald's pulling out of the country?
Ilya85 2 hours ago [-]
[flagged]
Ilya85 2 hours ago [-]
[flagged]
sbseitz 4 hours ago [-]
I wish I could downvote this for Pelicanmaxxing lmao.
influx 3 hours ago [-]
Would you prefer the term Pelicangate?
theandrewbailey 2 hours ago [-]
We're going to keep maxxmaxxing forever.