Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

it's not even a benchmark, i don't know what's the point of this post. these models aren't pelican painters or anything like that, they're LANGUAGE models (not even language learners, just a lossy model).


Sure, they're the wrong tool for drawing a pelican - but testing their SVG output is a useful way to get a feel for how good they are at step by step reasoning, coordinate systems, spatial awareness and generating valid SVG/XML.

There are genuinely useful applications of SVG-generation from LLMs - outputting simple infographics or charts for example.

I use LLMs to write HTML all the time, of which SVG is a useful optional component.


This benchmark is interesting, because it sidesteps the reasoning and process that humans would excel at.

For example, if I asked you to assemble a bookshelf with some wood, nails, and cement, you might first make a hammer with the cement before trying to assemble the bookshelf.

You can get a much better image by first asking the (multimodal) LLM to draw an image of a pelican on a bicycle, and then generate an SVG using the referenced image.

https://chatgpt.com/share/67609300-9abc-800d-9b26-95074f2149...


> these models aren't pelican painters or anything like that, they're LANGUAGE models

Tools are defined by what people use them for, not by how they were intended—or designed—to be used. (Just ask Nvidia)

adding: so I think someone comparing how various tools perform at a task that's valuable to them—and probably others—is just fine, even if it's different from what the creator of the tool intended?




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: