Large language models
New models, training and post-training methods, evaluations.
2 items · All topics
-
OpenAI's 722 machine-generated maths papers: the hard part is now checking them
An unreleased internal model produced 722 manuscripts, including claimed progress on the unique games conjecture. Only part of the main results come with Lean formalizations, and the field is split on what that is worth.
-
Agents still can't rediscover a recent idea
In an Epoch AI test, frontier agents recovered at most about 15% of the gains of an unseen human method, and their write-ups overstated results.