Karpathy’s Pelican
- YmiYugy - 4432 sekunder sedanI don't think it's a bad way to benchmark new models, I just find it concerning that the author implies that "pelican on a bicycle" has been exhausted. At the risk of making overly broad, unfalsifiable claims I think multi-year exposure to AI content has dramatically raised our expectations for speed and volume but lowered them for quality. We see a very janky pelican and declare the problem solved.
- jmugan - 10878 sekunder sedanA lot of people are posting here about how bad the end product is, but that is kind of the point. Models have moved beyond generating images to a new kind of benchmark that better exposes understanding of the physical world, and we can use benchmarks like this to measure future progress. (Of course, it will have to be a qualitative/subjective measurement.)
- bredren - 14537 sekunder sedanI worked with an LLM to build a ~3D animation of the Back to the Future delorean Time Machine as a way to spice up the hero on a docs page.
That took a fair amount of custom tuning and I had to create a tuning view to get some of the behaviors right.
But it was enough fun that I generalized it to take in ~any scene description from a film. It goes out and gets more detailed descriptions and film stills if available but also takes custom stills if you provide them.
My test scene was the Gauntlet scene from Apocalypto. It is low fidelity but does a pretty amazing sequence with somewhat believable physics of the javelins etc.
Here is the docs page with the vertical takeoff / 88 miles an hour time travel: https://contextify.sh/docs
I can share some of the Apocalypto bit if anyone is interested.
- djhworld - 911 sekunder sedanIt would be interesting to see the models work on a book it hasn't been trained on yet. I guess sadly that means any book released very recently.
Definitely impressive demo, I do wonder though if the countless artwork, films, images etc produced over many decades around Lord of the Rings somewhat influenced the outcome of this though.
- HarHarVeryFunny - 12685 sekunder sedanIt seems pretty clear that Anthropic models have been specifically trained to be good at generating three.js (JavaScript 3-D Graphics) code, so given current state of AI code generation in general, I don't find three.js models/animations as indicative of anything other than the model's ability to write three.js code.
When Fable was first released the day-1 demos of it on Twitter (presumably from people who were given early access, and/or Anthropic employees) were pretty much 100% three.js stuff. Yes, it looks nice, but it doesn't tell me any better than an Erdos proof whether the LLM will be able to run my vending machine.
- try-working - 3488 sekunder sedanthis is not a good benchmark for models, but it's great if you're optimizing for attention on twitter because video content and 3d animations perform best on social media.
a real benchmark is instead running evals on your own traces, and building a cost/quality/speed profile for models based on real workloads. but it doesn't get you a shiny video you can post on twitter.
- qwertox - 15392 sekunder sedanI'd rather have them battle on the topic "Who builds a better Google Wave for LLM chats" to explore the space of how AI studios could be.
"Getting Started with Google Wave": https://www.youtube.com/watch?v=eKUAqNGVwX0
- jcims - 1712 sekunder sedanI’d like to see a human one shot a pelican on a bicycle in raw svg.
- dundarious - 12508 sekunder sedanI can forgive the modeling being godawful jank (windows floating in the air, disconnected from the house). But I expected it to have a better understanding of the text. Instead, we have Bilbo's "disappearance" interpreted as him magically transporting or cloaking, and similarly for his reappearance.
- Lerc - 544 sekunder sedanI think I could tolerate 50 Shades of grey rendered in this style.
- trentor - 13316 sekunder sedanI always thought of the Pelican more of like a gimmicky quick test. There are people who took it as a serious benchmark for overall model performance?
- toolslive - 9823 sekunder sedanReading the title, I was thinking "Karpathy? I don't know this chess player." (The Pelikan is a well known chess opening, and famous chess players often have book titles like "X's Y" where X is the player, and Y is the opening)
- Waterluvian - 4239 sekunder sedanSpeaking of benchmarks has anyone given AIs Where’s Waldo pages and asked it to find Waldo?
I’ve been trying it on them all and can’t find one that does it consistently. The best will tell me they can’t. The worst confidently point out one of countless Waldo-likes.
- baron816 - 14418 sekunder sedanIMO, the area where AI is going to be most useful over the next couple years is in developing manufacturing processes top to bottom. Maybe a million token budget is too small, but something like "design me a sneaker and all the equipment to manufacture it autonomously".
- swe_dima - 1855 sekunder sedanIn my experience SVGs are still too hard for LLMs.
I gave Fable a jpeg and asked to draw an SVG, using a loop that renders the SVG into an image so Fable can inspect it.
Results looked like drawing of a 5 year old.
- informal007 - 1532 sekunder sedanOne difference for human to understand the video is that we only care the changes on a picture compare to LLM
- siliconc0w - 4795 sekunder sedanThere is a tipping point between procedurally generating everything in SVG to maybe giving them tool access to something like 3dsmax (or having them build and then use a tool to do the thing vs doing the thing).
- fzeindl - 14087 sekunder sedanRegarding the argument about LLMs having difficulties auditing their work:
I wonder whether we are entering the era of throwaway software. Just like cheap plastics and improved processes has enabled us to rapidly manufacture anything we want for a very low price, maybe LLMs give us the same for software. Produce it cheaply and if it breaks throws it away and reproduce it.
- hooloovoo_zoo - 3045 sekunder sedanI suspect LotR is a singularly unrepresentative choice here considering how much info exists about it.
- xyzsparetimexyz - 12269 sekunder sedanHow much are the hobbit houses described in the book? The ones here look exactly like the movie
- matsemann - 13075 sekunder sedanI'm pretty tired of the "Y made this game in Z tokens" all over the internet last week. They look impressive, and it's cool that it's even possible, but they're useless as games. None of them are any fun. They're like the most boring variant of basic controllers you can imagine. None have any cool mechanics. None have any tweaks made from hours and hours of testing. All have the same cel-shader.
- eichin - 7365 sekunder sedanIs anyone else getting "mongodb is webscale" vibes? (Except 16 years ago that was a lot smoother, because it used some sort of "render this conversation" engine...)
- sinaatalay - 3662 sekunder sedanOn consumer devices, AI communicates with us through speakers and screens. Screens are the richer medium, so most consumer AI innovation will happen there.
Computer graphics will have enormous applications because they are directly controllable by LLM-generated code. Video models are probabilistic and less suitable when precision matters. In education, for example, we need exact visuals. If an AI wants to plot y = sin(x), it should generate the precise graph through computer graphics rather than approximate it with a video model.
- cocoa19 - 12729 sekunder sedanWe must not be using the same opus 5, because if I tried to generate this it would refuse based on copyright grounds.
- dekhn - 10660 sekunder sedanI'd like to see the Silmarillion, specfically both Ainulindalë and the Fall of Numenor. At this point a visual model would probably produce something better than Amazon (but presumably not Jackson).
- fwlr - 10714 sekunder sedanI really dislike this AI programming thing of “Mr LLM, go slam your face into the problem until there’s no problem left, then call me back”. (Not sure if it’s a recent trend or a fundamental nature.)
It always brings to my mind some words from Rich Hickey:
I don’t think I really have a point to make here, other than it just feels like someone’s released a bunch of carnival bumper cars onto the highways.I think we’re in this world I’d like to call “guardrail programming”. It’s really sad: we’re like, “I can make change because I have tests!”. Who does that? Who drives their car around, banging against the guardrails, saying “whoah, I’m so glad I have these guardrails so I can make it to the show on time!” - barrenko - 9168 sekunder sedanThis has started to feel a bit like the beginning of railroads and then the steampunk fiction of "let's just build railroads to everywhere". We don't need it and there's no use for it.
As with painting, after a while there's nothing really new to paint, we genuinely need 0 new software. We need to fix our broken physical world, our social lives, our kids and what's left of our democracies.
This software crap is done, leave it to the nerds.
- informal007 - 1669 sekunder sedanit shows the possibility that SVG replace PNG/JPG even video.
- mold_aid - 6742 sekunder sedanThe tilde thing remains uniquely obnoxious in a field that seems want to mangle language for fun, so that's innovative I guess
- skybrian - 13915 sekunder sedanStill images seem like a better quick test because we can see them at a glance. Maybe ask it to make a comic?
- hkalbasi - 11571 sekunder sedanThis makes me think about using a game engine and a coding agent instead of current video generation AIs. It will probably cost much more, but it will have almost zero consistency problems. Is this line explored?
- OtherShrezzing - 9912 sekunder sedan> I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom
This is an odd take, given that Karpathy is certainly aware that the LotR films absolutely did create Bag End in digital format; that their creation was outstandingly high quality; and that Claude’s output here very obviously “leans heavily” on their prior art.
- wiradikusuma - 12629 sekunder sedanDo you guys notice that LLM can create fancy viz/animations by coding them instead of leveraging what we humans usually use (e.g Lottie, After Effects)?
I wonder if Flash is still popular... LLM can use that instead...?
- dofm - 10425 sekunder sedanAnthropic spokesman [0] Andrej Karpathy is here to tell you about token-wasting loops, and insists on the weird idea that they are "~free", when in fact, they are fuelled by expensively burning investor money.
[0] Seriously. Get used to mentally prefixing his and Boris Cherny's name like this, every time you see them quoted. These people are speaking while employed; there is no chance they are not aligned with the employers who will make them wealthy. The tech industry does like to pretend that for some reason AI people, uniquely, speak thoughts unbiased and for themselves or even for science or humanity.
- serf - 14077 sekunder sedanyou don't really need screenshots if you have an engine expressive enough for the scene generation while ensuring the visual appearance of the engine output itself is feasible.
that's why these things are actually pretty good at openscad/freecad/F360 mcps , the visual reality is enforced and guaranteed by rigor in the interpretation engine that is anchored to human physical reality.
- mvdtnz - 485 sekunder sedan> I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom but LLMs have all the stamina and patience in the world, so it's an example where we go from "no one would ever do this" to "sure, why not, it's ~free".
Except it's not ~free, it cost ~$10. And no one in their right mind would ever exchange $10 for that crap output except in this brief moment that we're in because it's fun and surprising to see what will happen. The actual result is as close to useless as it's possible to be - it's not interesting in its own right, it's not aesthetically pleasing, nor funny, nor informative. It's just slop.
- croes - 10633 sekunder sedan> I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom
There are people in their right mind who would do that and their are already examples of people who did similar things.
But maybe not in the future if people would confuse all the effort with AI
- xnx - 5810 sekunder sedanAI is now somewhere between "Money for Nothing" (https://www.youtube.com/watch?v=wTP2RUD_cL0) and "Knick Knack" (https://www.youtube.com/watch?v=9uhM_SUhdaw) in capabilities.
- throwaway89864 - 11671 sekunder sedanIt may make sense to switch this to USD/Omniverse.
- mikojan - 4070 sekunder sedanAfter watching this video I am absolutely positive that the issue is not a lack of stamina in humans. It is that humans have the capacity to realize that this is a bad idea long before they complete it.
- xg15 - 11792 sekunder sedan> I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it.
I think it's interesting that the "Bag's End" interpretation in the video clearly looks like the one from the movies, but generated here as a three.js 3D asset.
It makes sense that the movies (or shots/frames from them) were in the training data, and I can also easily imagine an association in concept space between the textual description of Bag's End and the frames from the movie.
But how on earth does the model then go on and convert the latent representation of those images into coordinates for a 3D mesh, without ever even restoring the image? In what kind of representation are the images from the movies stored that it can do that?
- wslh - 8531 sekunder sedanIf you like this check: https://news.ycombinator.com/item?id=47400868 it can be used to generate animations (not games) as well.
- Gooblebrai - 11913 sekunder sedanI can't believe the video demo is $10
- quantumleaper - 14559 sekunder sedanI'm sad that Andrej Karpathy went from being one of the most reasonable, trusted, and credible voices in AI to peddling marketing slop for Anthropic.
8 months ago, he was (very reasonably) claiming that reliable agents are at least a decade away, but this now goes against the interest of his employer, so the narrative has been changed.
- bbstats - 11183 sekunder sedanThis is awful
- blitzar - 13693 sekunder sedanI think the pelican test is better.
- andy99 - 9138 sekunder sedanBenchmarks like the pelican thing are about correlation with “how good the model is”. Better models produce better pelicans.
It’s a useful benchmark (aside from being “cute”) because of its simplicity, both in how many output tokens it takes (though I understand some models think a lot now to do it) and how easily one can subjectively judge. It’s this efficient as a benchmark of performance.
Making a long video takes way more tokens, and presumably is a lot tougher to easily compare. swillison has a presentation that’s pelicans from 2023-present (roughly) showing the progression. Imagine “lord of the rings videos from 2026-2029” or whatever, it would take a long time to watch and be harder to judge, and probably just end up being a comparison of screenshots anyway.
TLDR I feel like the post misunderstands the role of the pelican thing though if find it very hard to believe he really doesn’t understand, so maybe I’m missing something.
- stackedinserter - 9059 sekunder sedanIt would be better to ask model to render segmented 3d, with placeholders, like magenta is water, blue is sky, green is grass, purple is Frodo's face, etc, then pass the result through img2img model to properly "render" it.
- angoragoats - 333 sekunder sedanCan we please link directly to XCancel and leave the fascist garbage site behind?
- c0rruptbytes - 15573 sekunder sedanperfect benchmark to burn more tokens - convenient
- epolanski - 15104 sekunder sedanI wish there was a timeline where I never ever had to see the pelican SVG test ever again.
- forrestthewoods - 13092 sekunder sedanAs a former gamedev watching non-gamedev AI talk about games is so amusing. They really truly do not understand anything about games or consumer entertainment.
There’s a reason AI slop games have literally zero engagement. Last summer that stupid flying game blew up. Maybe a million people “played” the game. Where play means they clicked a link and checked it out not because of what the game was but solely because of how it was made.
In terms of concurrent players that game wouldn’t have cracked the Top 5,000 on Steam.
My metric for AI games is “number of players who spent more than 15 minutes playing”. I’m not aware of any vibeslop that has achieved 1 such player.
Now obviously LLMs are transformative for game dev. But “hyper custom worlds you can drop into” shows an extreme ignorance of what players want imho.
- nozzlegear - 12411 sekunder sedanDon't miss Elon Musk's reply:
> @elonmusk 13h
> Yah
> 158 replies, 74 reposts, 1400 likes
Thanks Elon, you goofy fuck
- miltonlost - 12122 sekunder sedanTech bros continue wasting money to make the absolute worst art
- shapefrog - 13420 sekunder sedan"Check the current situation and make a new Iran Lego (tm) truth bomb video."
- - 996 sekunder sedan
- hansmayer - 2374 sekunder sedan[dead]
- trlhaq - 14229 sekunder sedan[flagged]
- theproblemisyou - 12395 sekunder sedan[flagged]
- yourewrongsorry - 13331 sekunder sedan[flagged]
- andrewstuart - 6825 sekunder sedanThis is equally bad as a pelican test.
LLMs should be tested in the same way people should be tested for a job interview (but often aren’t) - with tasks RELEVANT to usage.
So you don’t just randomly pick some random thing to make the LLM randomly do (like many job interviewers do).
You start with clear statements about real world usage scenarios. THEN you come up with tests that give insight to how well the LLM/hob seeker gets the job done.
Please, stop coming up with random tests like it’s Microsoft in 1990 and you’re asking job seekers how the would move Mount Fuji, as a way of assessing their programming skills.
No stupid irrelevant pelicans on bicycles and no stupid renderings of Lord Of The Rings. Unless those are relevant use cases.
Any test that anyone comes up with must clearly state the context and how the outcome is measured.
- hn22fazjsv - 10127 sekunder sedanScreenshotting for later
- matchagaucho - 14784 sekunder sedanIt's difficult to think in exponentials.
But this demonstrates we're a couple orders of magnitude away from generating 1:1 hyper-personalized entertainment and media for individuals, rather than the masses.
Nördnytt! 🤓