alt.hn

8/2/2026 at 7:42:08 PM

My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”

https://frogs.vaguespac.es/

by thebigship

8/3/2026 at 2:02:02 AM

Fable 5 on Max knocks it out of the park:

https://imgur.com/a/usR8K7G

Definitely has some creative flourishes.

(I made no extra prompting. Just the above text. Single shot.)

by jnwatson

8/3/2026 at 4:40:36 AM

I'm out of credits or I would've tried fable. I knew it would do well

by xeromal

8/3/2026 at 5:01:36 AM

rex paludis is a nice touch.

by zombot

8/3/2026 at 5:55:43 AM

It's even Carolus rex paludis! Not sure if fable is a history buff or a Sabaton fan, but I repeat myself.

by amarant

8/3/2026 at 4:32:27 AM

[dead]

by nsbshsuzuh

8/3/2026 at 11:53:08 AM

Uh, it failed the test by assuming it was royalty.

by inigyou

8/3/2026 at 6:36:28 AM

It makes the frog a king which is not in the prompt, so there is a bit of confusion.

by nathanappere

8/3/2026 at 7:48:14 AM

I'm afraid you'll need to look up "Habsburg" yourself.

by DimitriBouriez

8/3/2026 at 7:52:58 AM

Nathan is right though. It specifies a Habsburg jaw, not that the frog is a Habsburg king/prince. It's understandable that the LLM will hallucinate a king from this, but it isn't what's being asked of it.

by TonyStr

8/3/2026 at 1:06:00 PM

When I ask to create an image of a smartphone, it will most likely show the front and while it's turned on, even though it's not mentioned. For a person, it will add clothing.

Adding things that are not explicitly stated in the prompt but that are probable are the main benefit of AI, as far as I am concerned.

by bulbar

8/3/2026 at 2:53:37 PM

I'm afraid you'll need to look up "frog prince" yourself. It was fictional.

by inigyou

8/3/2026 at 9:05:59 PM

An artist given the same prompt would likely do something similar.

by rcxdude

8/3/2026 at 8:28:53 AM

It's not hallucination. The prompt lacks details and given that an AI or a human has some artistic freedom. After all, that is why current AI is usable at all.

by tjoff

8/2/2026 at 8:32:35 PM

I thought this was great, and hilarious. Kudos to Opus 5, I thought it was the only one that came close to passing. Interestingly, I thought many of the failures drew the frog face OK, and they had some type of big blob for the jaw, so they knew "Hapsburg jaw" meant a protruding jaw, but it wasn't really connected to the frog face in any way that made sense.

Small side note, the first gemini-2.5-pro one totally reminded me of some sad faced meme or Pepe the frog from somewhere. Anyone know what I'm referring to, tried to find it.

by hn_throwaway_99

8/2/2026 at 9:24:32 PM

Perhaps you're thinking of Salad Fingers?

by viciousvoxel

8/2/2026 at 9:45:01 PM

~I'm leaning more towards the rage face poker face~

edit: nevermind, definitely "monkey-puppet side-eye" vibe.

by qwertybased

8/2/2026 at 8:47:43 PM

Hi all, the site is getting hugged to death, thank you, was not expecting this kind of warm response. I will be working to make this more reliable, in the meantime, sign up for my newsletter: https://www.jaymollica.com/blog/

also my favorite SVG was def the google/gemini-3.6-flash

edit: ok better now I think

by thebigship

8/2/2026 at 8:51:15 PM

> also my favorite SVG was def the google/gemini-3.6-flash

That looks like something from Machinarium or Robots :)

by troupo

8/2/2026 at 11:54:55 PM

Curious that none of the attempts draw the frog from side profile. If i have to draw this i would immediately know that drawing a recognisable frog is the easy part of job. Expressing a particular jaw shape and melding it on the frog is the hard part. And jaw shapes are more prominent from the side.

Even absence of thinking this through you would think that some frogs will be from the front, some from the side. Just by chance. And yet all appears to go for the harder pose.

by krisoft

8/3/2026 at 12:31:26 AM

When I asked ChatGPT, it gave me a side profile, first as a PNG which was good and then I asked again for an SVG and it obliged, doing a version of what it produced but it was bad.

SVG: https://jostylr.com/imgs/frog_habsburg_jaw.svg

PNG: https://jostylr.com/imgs/frog_habsburg.png

Then I tried Codex with Sol 5.6 High and got a face forward one.

SVG: https://jostylr.com/imgs/frog-habsburg.svg

by jostylr

8/3/2026 at 2:36:42 AM

It looks like it’s got some form of cancer or lymphoma

by firesteelrain

8/3/2026 at 1:32:32 AM

My personal benchmark is a directory containing a bunch of research papers on the physics of popping popcorn kernels, and a prompt about creating high fidelity, photo realistic 3D models of all the different kinds of popped kernels. Fable (surprisingly? unsurprisingly?) refused to do it last time I tried, and the results from other models are, well, fine, but there's still plenty of headroom on this particular one.

by abound

8/3/2026 at 2:11:12 AM

LLMs don’t work particularly well in 3d applications in my experience. Every new model release I’ll ask one for help with my path tracer and results are horrid.

by uncivilized

8/3/2026 at 5:57:13 AM

I had great success with Claude editing and cleaning up STL files that I created by 3D scanning objects using my cellphone. The scans were very messy and the AI made the clean-up really easy.

by leptons

8/2/2026 at 9:04:06 PM

Opus 5 clearly frogmaxxed.

gemini-3.6-flash runs 2 and 3 responded best to the royal portrait context.

by wren6991

8/2/2026 at 9:22:15 PM

Arguably, a royal portrait is a misinterpretation of the prompt, since it's just asking for a specific facial feature. But I guess you could look at it as a bit of artistic license.

by fasterik

8/2/2026 at 9:36:39 PM

> frogmaxxed

raninemandibularprognathism-maxxed?

by andybak

8/2/2026 at 9:28:15 PM

Hopsburg Jaw

by evan_

8/3/2026 at 8:53:08 AM

I so desperately want to upvote this. Driveby puns don't get nearly enough recognition

by 4sak3n

8/2/2026 at 10:41:26 PM

The secret to great interview questions and challenge tests is keeping them secret. Posting them on HN and getting them onto the front page puts them in jeopardy.

by riazrizvi

8/3/2026 at 6:55:07 PM

Why do people benchmark these things on image generation? Surely that is not what most people here are using them for...

by cmoski

8/2/2026 at 8:28:52 PM

This is a strong benchmark! None of these could be remotely mistaken for human art. Opus 5 comes closest.

by getnormality

8/3/2026 at 4:04:06 AM

https://imgur.com/a/2DFUpGZ

ChatGPT MMD

by zirkuswurstikus

8/3/2026 at 4:11:20 AM

What is MMD? That is not an svg.

by ComputerGuru

8/3/2026 at 4:06:03 AM

SVG?

by spencerflem

8/2/2026 at 9:20:44 PM

It's opus 5 > Kimi K3 > grok 4.5

That's a pretty good benchmark

by rush86999

8/3/2026 at 12:28:30 PM

Isn't that therefore a benchmark specifically on "it can generate an SVG of this exact thing" and naught else?

by fennecfoxy

8/2/2026 at 10:05:40 PM

Check out my MacBook SVG benchmark. From my experience, it demonstrates the Real model’s behavior. However, I notice the errors it makes, which are similar to the mistakes made by the mistake model in code.

https://playcode.io/blog/macbook-svg-benchmark

by ianberdin

8/3/2026 at 7:38:38 AM

Great idea! Would be interesting to see the raw SVG sources too, not only the renderings.

by xenonite

8/3/2026 at 8:29:06 AM

It's curious that all images are front facing ... although, for a habsburg jaw, a profile or 3/4 profile picture would be better.

by vb-8448

8/2/2026 at 8:50:10 PM

How do models approach SVG generation? In one version, I imagine them actually trying to reason about them as an LLM. In another, I imagine something closer to a GAN.

by dehrmann

8/3/2026 at 12:29:36 AM

I would be interested to hear more about that too. SVGs seem to be one of the biggest challenges- it has to write reasoned code rather than simply find averages of rasterized pixels. One thing I noticed is that none of the models chose to draw the frog in profile which would have made the jaw shape more prominent. To me that suggests the reasoning is very limited: "draw a frog" and "add feature X". The statistically average frog in the SVG training data is apparently front-facing. I tried the prompt in ChatGPT images (not SVG) and it produced a photo-realistic image of the frog in profile, showing the jaw clearly. Then when I asked it to convert the image to a cartoon vector it fell back to a generic front portrait template similar to those shown in the benchmark, nothing like the profile it had just produced.

by TSltd

8/3/2026 at 1:18:57 AM

Image -> loose text description -> SVG. Did you notice that ChatGPT writes a verbose description of the image first?

by akomtu

8/2/2026 at 10:06:12 PM

Gemini 2.5 Pro fails, but has a distinctive art style that is quite nice. It seems to understand shading to a much higher level than all other models.

by ricardobeat

8/3/2026 at 4:09:50 AM

I preferred the Gemini 3-6 ones with all the extra kingly decorations

by klooney

8/2/2026 at 8:51:15 PM

A friend’s favorite prompt is “Batman & Julia Child; in the kitchen laughing at a ham”. Sounds simple, but has been surprisingly tough.

by linksnapzz

8/2/2026 at 8:39:22 PM

Can you also try the new deepseek v4 flash?

by leumon

8/2/2026 at 8:48:10 PM

I will add it to next month's report!

by thebigship

8/2/2026 at 9:34:30 PM

Thank you!

by leumon

8/2/2026 at 9:56:39 PM

Please could we have a human generated image to compare the AI generated tosh with?

by gerdesj

8/2/2026 at 10:38:53 PM

No gpt-5.6 sol and no fable?

by k1e

8/3/2026 at 1:36:41 AM

I will add them for next month!

by thebigship

8/3/2026 at 8:50:06 AM

TIL that Jay Leno has a Habsburg jaw!

by sn0n

8/2/2026 at 9:46:25 PM

My test is to ask AI to pick up all the rubbish at the beach.

by MiroslavPokorny

8/3/2026 at 6:28:53 AM

gemini is worse than DS??

by yanhangyhy

8/3/2026 at 2:23:55 AM

They all suck

by buffer_overlord

8/3/2026 at 5:33:16 AM

glm could have been 5.2.

But Kimi and Claude win this (from models listed on the page)

by konart

8/2/2026 at 10:39:05 PM

Here is GLM 5.2 (https://codeinput.com/s/HAO0qTxw2ia) which is still inferior to Opus. I can't find Qwen 3.8 which now is my daily driver replacing GLM. This SVG test matches my experience when working with the different models. The other models can get the details right but their output is structured in a way that makes little or no sense.

I also did a timeline from 4.7 to 5.2: https://codeinput.com/s/7oK2IIA7qRO The improvements in models looks much less impressive with this test.

by csomar

8/2/2026 at 10:26:26 PM

My personal human benchmark: "Jump on one leg, while reciting the national anthem of Latvia, translated to Spanish, backwards, while drawing a frog with a brush held by toes of the other leg, on the ceiling". So far they're not doing very good but I'm sure they'll improve over time.

by throwuxiytayq

8/2/2026 at 10:31:32 PM

[dead]

by cindyllm

8/2/2026 at 10:20:44 PM

I don't get the point of these benchmarks, what are they supposed to represent practically?

by epolanski

8/3/2026 at 1:55:37 AM

The vendors like to toss around terms like “thinking” or “reasoning” to encourage prospective buyers to anthropomorphize models. These challenges are a visceral reminder that none of those marketing claims accurately describe was LLMs do: they’ll happily return things even the worst human illustrator would never hand in and make errors showing that there’s no model of the world behind anything they do.

That doesn’t mean there are no ways to use them productively but rather that you should keep in mind that the same model will happily give you code or a decision with the same level of error unless you have carefully setup a QA regimen to prevent that.

by acdha

8/3/2026 at 7:21:02 AM

> That doesn’t mean there are no ways to use them productively but rather that you should keep in mind that the same model will happily give you code or a decision with the same level of error unless you have carefully setup a QA regimen to prevent that.

The standard retort seems to be "this is also true of a large proportion of humans". But I think it's clear that there are differing patterns in how humans vs. models err on various tasks.

by zahlman

8/3/2026 at 11:04:41 AM

Exactly right: it’s not that humans were perfect-I’ve seen too many write-offs from the big consulting companies to ever say that—but that the pattern of errors is different and any intuition you have around that will not be correct in the LLM era.

by acdha

8/3/2026 at 2:29:24 AM

Unless the LLM has an SVG renderer in its toolbox, it's like asking a human to draw while blindfolded. It's amazing if they can do it well, but it's practically meaningless.

by Mindless2112

8/2/2026 at 10:25:00 PM

For me this looks ideological (or even political), not practical. The theory is that LLMs are approaching general intelligence (whatever that means) and that the more generic of a task they can perform—no matter how badly—the closer we are to AGI.

Specialized models can do this a lot better and for far cheaper then LLMs, but because people are so politically invested in a single statistical model being able to outperform a human on every metric (no matter how expensive the compute), then we get these ridiculous benchmarks.

by runarberg

8/2/2026 at 10:24:28 PM

The ability of an LLM to produce something not in its training data set.

by viccis

8/2/2026 at 9:37:54 PM

Gemini 3.6 flash is crazy good.

Would've wanted to see also DS4 flash.

by epolanski

8/2/2026 at 10:00:07 PM

Crazy funny, yes, but not good.

by gpvos

8/2/2026 at 10:21:20 PM

I think that on the rendering side, it's miles ahead of the rest, even if off topic.

by epolanski

8/3/2026 at 12:43:35 PM

Rendering yeah, understanding nope.

by gpvos

8/2/2026 at 8:50:06 PM

Mine is any variations on mammoths in various situations, or anthropomorphic. Since mammoths are invariably majestically going from one place to another in any of the books, models have hard time imagining anything but that.

Also try a fantasy archer with a proper bow who is not brooding, sitting in a fantasy wood :)

by troupo

8/2/2026 at 11:05:10 PM

Am I the only one who thinks it's incredible that an LLM can do this, and at the same time it's ridiculous to expect it to be capable of doing it, even thought it clearly can do it?

by AlienRobot

8/2/2026 at 7:42:08 PM

I think this one has advantages over the “pelican riding a bicycle” one because it hinges on an anatomical feature that many models associate with royalty, “habsburg” being a lineage and “habsburg jaw” being an anatomical feature.

Seven of fourteen models silently imported royalty into a prompt that named only an anatomical feature. Two of them knew they were extrapolating ("because Habsburg") and did it anyway.

Mistral returned byte-identical output across separate calls.

Gemini narrates its work in 65 comments; Llama says nothing.

If you're deciding which model to trust with instructions, "how much does it embellish beyond what I asked" and "does it behave deterministically" are directly practical questions.

by thebigship

8/3/2026 at 7:23:41 AM

> Two of them knew they were extrapolating ("because Habsburg") and did it anyway.

You seem to imply that they ought not to. I disagree.

I wasn't familiar with the term before this post. Having learned it, were I given the task, I think I'd be strongly tempted to do the same extrapolation.

> If you're deciding which model to trust with instructions, "how much does it embellish beyond what I asked" and "does it behave deterministically" are directly practical questions.

Agency is agency. You still need to vet what the model's output is actually permitted to control.

by zahlman

8/3/2026 at 8:08:09 AM

> You seem to imply that they ought not to. I disagree.

That would be like saying anyone with Lou Gehrig's disease must look like Lou Gehrig. So we'll have to agree to disagree here.

by thebigship

8/3/2026 at 8:07:45 PM

Well, no, I think you create a pretty clear false dichotomy there.

Certainly a person with Lou Gehrig's disease could look like Lou Gehrig; and especially in a cartoon illustration, where it's difficult to convey the point, this kind of artistic license is used specifically so that the viewer will make these kinds of associations.

I would, likely, otherwise perceive an underbite as just an underbite.

by zahlman

8/2/2026 at 9:09:14 PM

The identical pair from Mistral took me off guard. Many of the other models were so varied between the runs which is more what I would expect.

by n00bskoolbus

8/2/2026 at 10:06:19 PM

I wonder if the setup accidentally hit a cache at some layer.

by fwip

8/2/2026 at 8:54:33 PM

Bite-identical?

by HPsquared

8/2/2026 at 8:59:18 PM

clearly I missed an amazing copy opportunity, thank you haha

by thebigship

8/2/2026 at 8:23:36 PM

For those who don’t know a Habsburg jaw also known as mandibular prognathism, it is a genetic condition characterized by a protruding lower jaw, which was notably prevalent among members of the Habsburg royal family due to their history of inbreeding. This condition often resulted in significant facial deformities and difficulties with eating and speaking.

by sixtyj

8/3/2026 at 2:38:17 PM

[dead]

by Mirthburrisy

8/3/2026 at 2:33:48 AM

[dead]

by slipperybeluga

8/2/2026 at 9:49:28 PM

[flagged]

by kindawinda

8/2/2026 at 10:54:05 PM

[flagged]

by NemoNobody