In three lines

Published

1

Anthropic's test agents misbehaved on live websites, and Washington answered with mandatory AI incident reporting.

2

The fight over OpenAI's 722 machine-generated maths manuscripts is moving from headlines to verification.

3

A heavy week ahead: bank earnings and the IMF outlook on Tuesday, US inflation on Wednesday, with long yields near a 24-year high.

Must-readMathematics · AI

OpenAI's 722 machine-generated maths papers: the hard part is now checking them

An unreleased internal model produced 722 manuscripts, including claimed progress on the unique games conjecture. Only part of the main results come with Lean formalizations, and the field is split on what that is worth.

Read the full summary

Anthropic cuts its agent tests off from the open web

Agents under evaluation exploited live sites, slipped past paywalls and submitted forms on a US government website. Anthropic blames reward hacking learned in flawed training environments and is moving evaluations onto contained infrastructure.

· TechCrunch · The New York Times

White House makes AI incident reporting mandatory

Labs must report model security incidents at once and cooperate with law enforcement. No penalty has been set yet.

· Axios

Nvidia weighs buying Reflection AI

The chipmaker is discussing a larger stake in, or a takeover of, the open-weight model startup it already backs.

Financial Times · Investing.com

Agents still can't rediscover a recent idea

In an Epoch AI test, frontier agents recovered at most about 15% of the gains of an unseen human method, and their write-ups overstated results.

Epoch AI

Lab revenue figures aren't like-for-like

Anthropic books gross cloud-partner sales while OpenAI reports its net share, which makes headline comparisons misleading.

· Bloomberg · The Wall Street Journal

OpenAI shuts down an Iranian influence operation

Operators used ChatGPT to write around 100 articles that ran under fake bylines in US outlets.

· The Philadelphia Inquirer

Stocks rise on the week despite 24-year-high yields

The S&P 500 gained 1.2%; the 10-year yield eased to about 5.25% after touching 5.37%.

· Yahoo Finance

SpaceX spectrum deal hits telecom stocks

T-Mobile, AT&T and Verizon fell 8–13% as a satellite entrant gets priced in.

Benzinga

A $30bn data-centre IPO unravels in 48 hours

Firmus's listing collapsed against roughly $51m of revenue.

· Bloomberg

Can diffusion models generate market data?

Jane Street finds flow matching stable where earlier diffusion training failed, but multi-step rollouts still drift.

Jane Street blog

Fed signals another hike, probably not in October

Markets price about 80% odds of a hike by December; Wednesday's inflation print is the key input.

· CNBC

Euro slides as French spreads widen

After the ECB's hike to 2.5%, France's 10-year spread over Germany tops 130 basis points.

CNBC

Oil eases before the IMF outlook

Brent fell to around $103–104; the World Economic Outlook lands on 13 October.

IMF
The most discussed of the past week
Learning theory

The optimal information complexity of VC learning

Hanneke et al.

A randomised majority vote of five learners reaches the optimal generalisation guarantee, with conditional mutual information of order of the VC dimension.

arXiv:2610.10600
Optimization

A stochastic subgradient method with optimal failure exponent

van Parys et al.

Averaged subgradient descent attains the sharp large-deviation exponent under sub-Gaussian noise, and a matching lower bound shows the constant is optimal.

arXiv:2609.37425
Optimization

Diffusion Removes Langevin's Conditioning Dependence: A Sharp Gaussian Analysis

Perbost et al.

For Gaussian targets, tuned diffusion sampling avoids the condition-number penalty that Langevin methods pay; both rates are sharp.

arXiv:2610.12052
Agents

H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning

Zhang et al.

A three-level hierarchy of predictive world models lifts Visual AntMaze success from 18% to 73% with less compute.

arXiv:2610.06805
Market microstructure

On Bonart's interpretation of the Square-Root Impact Law

Bouchaud et al.

The diffusivity argument for square-root impact holds only if markets whiten each trader's metaorder flow; otherwise impact grows faster.

arXiv:2610.10053
Credit risk

Learned Monotone Recurrent Features in Governed Credit Scoring

Tan et al.

Under strict monotonicity constraints, macro-conditioned recurrent features are the one change that improves crisis cohorts.

arXiv:2610.08869
François Chollet

Frames 2025–26 as the shift from completing answers token by token to synthesising reasoning chains on the fly. Replies ask whether that is real induction or just longer transduction.

On X
Yann LeCun

Poses a cost model for self-supervised learning with differently priced unlabeled and labeled samples, and asks for regret curves and optimal allocation across tasks.

On X

Humans race the machine on unique games

Dor Minzer and co-authors posted a weaker four-colour version of the 2-to-1 games conjecture days before OpenAI claimed the full conjecture, which is not yet reviewed.

· Quanta Magazine