7/22/2026 at 7:04:48 PM
This is fantasticI've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations.
Catching a lab cheating specifically on my one dumb benchmark would be really funny.
Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than anything I was considering.
His conclusion:
> Nothing jumped out at me. I couldn’t find a case where the pelican-bicycle images looked noticeably better than the rest of that model’s grid.
by simonw
7/22/2026 at 7:38:28 PM
What if they’re not pelicanmaxxing, but svgmaxxxing in general?Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge.
Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.
by lukev
7/22/2026 at 8:53:07 PM
Then it's great. A year ago I couldn't get any AI to draw a simple company logo in SVG or even convert from raster. No doubt the Pelicans put pressure on the labs to fix the awful SVG situation.by brikym
7/22/2026 at 8:03:07 PM
But that's a genuine worthwhile capability. It's like benchmarkmaxxing on a weightlifting competition by getting really strong.by qq66
7/22/2026 at 8:40:54 PM
https://www.youtube.com/watch?v=jgYYOUC10aMreminds me of this Key and Peele skit
by Balgair
7/22/2026 at 8:11:09 PM
How useful actually is this? It generates SVGs of pelicans on bicycles, sure, and some of them are (almost) spatially correct. But, none of them look good.AI image generation suffers from this more generally. You can generate pictures of pelicans, sure. Newer models clearly generate images with more pelican-ness than before. But all of it is still uglier than sin. Drawing things accurately is one thing, making results that someone might actually want to use (without embarrassing themselves) is something else.
by ryukoposting
7/22/2026 at 8:27:24 PM
Even if the models don't break through any particular "uglier than sin" barrier, with a bit more work, presumably the SVGs could become importable into an editor that would let a human apply taste and discretion. Seems to me like a heck of a head-start.As for conventional diffusion-model stuff, I happen to think there are some pieces of AI art that still look really good even knowing they're AI.
by zahlman
7/22/2026 at 8:32:10 PM
Vibecoding a SVG based metroidvania as we speak! This is gonna be lit!by amarant
7/22/2026 at 8:19:05 PM
AI can generate a fairly satisfactory SVG for a favicon now (programmer art quality at least).by fiddlerwoaroof
7/22/2026 at 8:16:58 PM
well that holds IF svgmaxxing is 100% "code-writing-maxxing"...which.. hmm I dunno if they are same or not
by sysguest
7/22/2026 at 8:41:58 PM
No, the point is that a general-ish ability to draw good SVGs is a useful ability in itself. People need SVGs for all sorts of purposes, and if AI can generate one for them, that's mostly useful (discussions about art and employment etc notwithstanding).That said, I think this would correlate relatively little with general programming ability. They're not unrelated, of course, but being able to generate code that paints an accurate + esthetically pleasing image is quite different from generating code that achieves a non-spatial goal.
by tsimionescu
7/22/2026 at 8:21:14 PM
I don't see why that's true. LLMs don't have to only be good at code-writing.by wasabi991011
7/22/2026 at 8:46:44 PM
Funnily enough, not that niche, because I have tried many times to do it as part of a wider project.by kaliqt
7/22/2026 at 7:46:37 PM
Addressed in the article, in case you're curious.by netsec_burn
7/22/2026 at 7:59:12 PM
Well, it’s mentioned as a limitation of the analysis, very much not ruled out (or in.)That simonw is causing labs to do extra fine-tuning runs for this seems highly probable :)
by lukev
7/22/2026 at 7:45:42 PM
I agree, other formats, both textual and binary should be tested.by charcircuit
7/22/2026 at 8:25:41 PM
Simon I hope from this day hence, your bio always includes:"Simon Willison, among other things, is an advocate for the inclusion of pelican geometry in LLM training datasets."
by eob
7/22/2026 at 8:10:23 PM
I think a more fundamental test is SVG art creation in general. Perhaps a pipeline to take any image, caption it, ask the LLM for an SVG, rasterize to an image, and finally either use a deterministic visual similarity check or ask another LLM to be the judge and score how close the SVG is to the original image.by docheinestages
7/22/2026 at 8:29:01 PM
Fidelity to the original is definitely not how humans would measure "art" in this context.by zahlman
7/22/2026 at 8:40:01 PM
True, maybe we can call the generated SVG something else than art.by docheinestages
7/22/2026 at 7:45:10 PM
> Catching a lab cheating specifically on my one dumb benchmark would be really funny.Similar thing happened when TPC came up with SQL benchmarks.
If you're not good at TPC, your engineering team is no good.
If you're good at TPC, then (as a customer) we will actually include you in a bake-off benchmark for our specific problem.
Winning on it is the price of admittance into the game, especially in a crowded market.
But how narrowly you benchmarket matters, you can't just hard-code that specific scenario & not fix anything adjacent while you're at it.
For example when it comes to GPUs, the "Quack3" (sic) benchmark on ATI cards comes to mind.
by gopalv
7/22/2026 at 7:09:35 PM
Perhaps also vary the bird? Wikipedia tells me pelicans are in the order _Pelecaniformes_ so shoebills or herons might do.by gilleain
7/22/2026 at 8:15:42 PM
> I've been casually spot-checking other animals in other vehiclesSnakes on a plane, weasels on a diesel, spiders on a glider, baboons on a balloon, goats on a boat.
by cyberax
7/22/2026 at 7:26:54 PM
[dead]by mattertoast