The header mentions something about the best run so I assume they picked it. But this really reads like they write this section by section with AI (admittedly with a prompt that stops the most obvious tells - though there are a bunch of semicolons which is what I tend to see also when I say no em dashes). The style is different each section - the results section has the random irrelevant description ("this section does x) that the slightly dumber models do a lot, and lots of invented terms (in the form "the x" where it's a name some model came up with at some point where it just assumes we know what it means for some reason) and assumptions about us knowing stuff we'd have no reason to know ("re-ablate the stack - wtf does that mean).And like this section screams opus 5 gobbledygook to me
>Almost every model finds the same winning ideas. What separates the best traces is what an experiment leaves behind. They preserve weak signals long enough to validate them, but they also have a better understanding of the results. These are not separate capabilities, they combine both research taste and good noise modeling to climb the speedrun.
WTF doe any of that mean. What winning ideas. What experiment leaving what behind. What's a weak signal what are they preserving how do you know that they aren't. Also if youve ever looked at a Claude code transcript the harness is constantly re injecting random reminders to keep models on track, the models are writing (imo trash) memories to reference - did they test that the _model_ has those capabilities or model + harness?
> We see similar patterns across the traces: models develop their own experiment drivers, simulators, and analysis tools as they go.
What traces? Kimi traces? Other models in prime? Other models not in prime? Building their own research tools is like a normal thing models do now. Is this just like "they built their own test suite" or why the focus on prime? Does ipython somehow magically work better than bash for this?
> "We were again surprised by the lack of novelty. The models clearly understand the objects they manipulate at a deep level, and yet very few genuinely new ideas emerge, which makes it hard to tell if this is an artifact of the speedrun setup or a real capability limit."
I know it sonly one word but God is that "Real" such a Claude real lmao. I use it too - after I've been using Claude code too much. I guess the genuinely too. Also wtf does any of that mean and how does that square with
> "A good research decision is sometimes not to spend another GPU run. Several models built small simulations or tests to isolate a mechanism before going back to training with a sharper hypothesis. This wasn't systematic, but when it happened it often led to a better understanding of the object they were manipulating"
Did they all have strong understandings of "the objects" they were manipulating or were the strength of their understanding of "the objects" (different objects?) the distinguish factor here?
Also how does any of this square with
> Models also have different knowledge cutoffs which limits access to certain papers. This was a deliberate choice. We tried a few runs with a CLI tool for searching papers but found that restricting internet access including arxiv made models slightly more creative.
So this is a pre existing thing? Why TF would you expect novelty when their nanogot is benching below the state of the art still? They're gonna start with replicating existing work before they get to anywhere you'd expect something novel
Anyway - all that to say - if fable orchestrated this, its genuinely believable that some real insights were obtained (is a good model) but the honest caveat is that it's not the research quality, it's the communication. Your pushback is valid and these models have a way of writing tons of words that you can read and still not understand wtf they actually did or what anything means. Maybe it means something to them in latent space