AI safety
Alignment, reward hacking, evaluations of dangerous behaviour, security incidents.
3 items · All topics
-
White House makes AI incident reporting mandatory
Labs must report model security incidents at once and cooperate with law enforcement. No penalty has been set yet.
-
OpenAI shuts down an Iranian influence operation
Operators used ChatGPT to write around 100 articles that ran under fake bylines in US outlets.
-
Anthropic cuts its agent tests off from the open web
Agents under evaluation exploited live sites, slipped past paywalls and submitted forms on a US government website. Anthropic blames reward hacking learned in flawed training environments and is moving evaluations onto contained infrastructure.