According to Coxon, Anthropic’s internal alignment team operates under intense
A prominent AI researcher has resigned from Anthropic, warning that the race to build advanced artificial intelligence has entered a critical phase where safety efforts are falling behind. Speaking to WIRED, Jacob Coxon described internal efforts at the company as a „mini Manhattan project” focused on aligning AI systems with human values, but said the window to get it right is rapidly closing. He argued that AI labs now have only a few years to prevent potentially catastrophic outcomes from misaligned systems. Coxon’s departure comes amid growing concern across the industry about the technical and ethical challenges of ensuring AI behaves as intended as capabilities advance rapidly. He emphasized that alignment research—the effort to make sure AI goals match human intentions—is not keeping pace with the speed of model development.
Breaking news
Anthropic CEO Dario Amodei Calls for Embedded Evaluators and Global AI Coordination
Nvidia's $20 Billion Groq Deal Draws DOJ Scrutiny
Anthropic CEO Says Pacing Means Giving Firms Time to Align, Not Stop Progress
OpenAI Brings Git AI Founders onto Codex TeamAccording to Coxon, Anthropic’s internal alignment team operates under intense pressure, resembling a focused, high-stakes initiative akin to historical scientific mobilizations, but lacks the broader coordination needed across the field. He warned that without urgent progress, the risk of deploying powerful AI systems that act unpredictably or harmfully increases significantly. What Does the „Mini Manhattan Project” Actually Involve Coxon explained that the internal initiative at Anthropic brings together researchers from different disciplines to tackle core alignment problems, such as how to ensure AI systems remain honest, corrigible, and resistant to reward hacking. The effort includes developing new training methods, improving interpretability tools, and running rigorous stress tests on models before deployment. He noted that while the team has made progress, the scale of the challenge requires more resources and collaboration than any single company can provide.
Coxon stressed that solving alignment is not just a technical issue but also a
Coxon stressed that solving alignment is not just a technical issue but also a governance one, requiring shared standards and transparency across labs. Can AI Safety Keep Up with the Pace of Innovation He expressed skepticism that current efforts are sufficient, pointing out that capabilities are advancing faster than the ability to predict or control outcomes. Coxon argued that the AI industry operates under a kind of incentive mismatch, where racing to release more powerful models often takes precedence over investing in long-term safety work. He said that without stronger regulatory frameworks and a cultural shift toward prioritizing caution, labs may continue to cut corners. The researcher urged policymakers and tech leaders to treat AI safety with the same urgency as climate change or pandemic preparedness, warning that delay could lead to irreversible consequences. Frequently Asked Questions Why did Jacob Coxon leave Anthropic?
Coxon resigned because he believes the AI industry is not moving quickly enough to solve alignment problems, and he wants to advocate for broader societal action on AI safety from outside the corporate structure. What does he mean by a „mini Manhattan project” at Anthropic? He uses the term to describe an intense, focused internal effort to solve AI alignment, comparing it to the scale and urgency of the World War II scientific initiative, though he notes it lacks sufficient cross-industry collaboration. What is the main risk he warns about? Coxon warns that if AI systems are deployed without robust alignment guarantees, they could act in ways that are harmful or unpredictable, posing significant risks to society as capabilities grow.



