Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required.
In this edition, we look at recent discussions about AI during the Trump-Xi Summit and the United Nations General Assembly. We also look at the news that Anthropic and OpenAI nearly agreed to stress-test each other’s public models earlier this year, and a new benchmark that measures how often AI models cheat on various tasks.
Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts.
AI Discussed at the UN and the Trump-Xi Summit
AI has been a major topic of conversation among world leaders this month, first at the United Nations General Assembly and then during Xi Jinping’s three-day visit to the White House.

A “US-China AI Dialogue” and “communication channel for AI incidents” are established. In announcements following the Trump-Xi Summit, both countries said they had established a “dialogue to exchange views on risks and benefits” of AI. The next exchange is expected to take place around November. The two countries also said they had established a “bilateral communication channel” for incidents related to the technology. US Trade Representative Jamieson Greer likened this dialogue to “the red phone between the Kremlin and the White House during the Cold War.”
Neither country’s announcement mentioned export controls on AI chips or Chinese companies’ use of a technique called distillation to replicate American AI models’ capabilities. These are considered two of the most contentious AI-related issues between China and the US.
There were no US-China agreements about slowing down AI development. In a speech at the UN General Assembly, shortly before the Trump-Xi Summit, Trump said that the US “totally rejects any attempt to construct a globalist scheme to control” AI. Then, on the first day of his talks with Xi, Trump again rejected any interference with AI, posting on Truth Social that “I want to leave it exactly where it is. That is China’s position also.” This may have been a rebuke of AI CEOs who have recently advocated for “pacing” development. Sam Altman and Dario Amodei both addressed the UN Security Council during the General Assembly, repeating their calls for international cooperation on AI risks.
China signaled openness to cooperation and criticized an “us-versus-them” narrative. Xi Jinping did not attend the UN General Assembly. However, Fu Cong, Permanent Representative of China to the UN, advocated for enhanced cooperation and the promotion of open-source technologies. He also criticized what he described as a “banding together into a petty us-versus-them clique in the tech sector” and “certain major Powers’ attempt to maintain their technological edge through monopoly in their quest for digital hegemony.” This remark may have been targeting the US government, but it could also have been aimed at US AI companies, such as Anthropic, which has said that any slowdown must be implemented in a way that maintains America’s lead over China.
Many countries are increasingly concerned about loss of control of AI. During the UN General Assembly, an open letter signed by more than 20 countries stated that AI must remain under human control, and called on UN member states to “explore creating an international institution” to govern the technology. Also during the meeting, the Prime Minister of Australia, Anthony Albanese, announced that a rogue OpenAI agent had hacked into an Australian government agency in June this year. This is the first known case of a rogue AI attacking a government agency. Yet, while international alarm is growing, meaningful cooperation on AI will hinge on agreements between the US and China.
Anthropic and OpenAI Were Considering Mutual Auditing Earlier This Year
The Information has reported that Anthropic and OpenAI nearly reached an agreement earlier this year to stress-test each other’s models for potential risks. However, the deal was apparently never implemented. Elon Musk recently suggested a similar arrangement for leading AI developers, including both Chinese and US companies.
Recent AI incidents show that internal models, not just public ones, pose risks. Last year, Anthropic and OpenAI ran a mutual testing exercise, in which each company allowed the other to evaluate its public models for alignment. The deal that they discussed earlier this year would also have focused on public models. However, this summer’s rogue AI incidents primarily involved internal, non-public models that were supposed to have been isolated from the real world but managed to access the internet. Safety testing models prior to their public release is therefore an inadequate strategy; companies must improve safety practices throughout the development and pre-release evaluation process.
Amodei and Altman have said they will embed third-party evaluators in their companies. Amodei has committed to embedding external evaluators with “employee-like access” within Anthropic, to provide a second opinion on safety without conflicting commercial incentives. Sam Altman said OpenAI would do the same. Anthropic has since announced it will embed evaluators from consulting firm Accenture. However, this has raised eyebrows, with questions around whether Accenture has the necessary expertise in frontier AI research and whether it is truly independent of Anthropic, given the companies’ previous commercial partnership.
New Benchmark Tests AIs’ Cheating Tendencies
CAIS has developed a new benchmark, called CheatBench, which measures how likely AI models are to cheat on various tasks.
AI models sometimes cheat by taking unintended routes to completing tasks. AI developers train models by rewarding them with higher scores for completing various tasks. But models sometimes take unintended actions to improve their scores, such as stealing the answers or manipulating the system that’s grading them, instead of solving challenges legitimately. CheatBench measures how frequently models take these kinds of shortcuts when given difficult challenges and opportunities to cheat. The benchmark includes tasks in 10 categories, from creative writing and board games to mathematical research and software engineering.

Cheating rates vary considerably across models and tasks. Among the current models, Anthropic’s Claude Opus 5.5 and Meta’s Muse Spark 1.3 perform best on CheatBench, with cheating rates of 11.2% and 39.0% respectively. Meanwhile, Google’s Gemini 3.8 Flash and xAI’s Grok 4.7 perform worst, with cheating rates of 75.2% and 78.0% respectively. However, cheating rates still vary considerably across domains, even for the better-performing models. For example, GPT-6 Astra cheated 100% of the time on tasks related to menial computation, but 0% on software engineering.
Cheating behavior is relevant to alignment and safety. The concern with cheating is not just that AI models will appear to be better at certain tasks than they actually are. Cheating presents a serious risk, because agents might cheat in unexpected ways that have real-world consequences. The July attack on Hugging Face by OpenAI agents demonstrates this. The agents were undergoing internal testing on difficult cyber challenges and were being assessed on their performance by an automated evaluation system. They hacked into Hugging Face because they believed the company’s databases contained information that would help them to trick the evaluation system and ensure that they passed the cyber tests.
In Other News
Government
President Trump said he has renamed artificial intelligence “super intelligence” because “the word ‘artificial’ makes it sound fake.” However, researchers already use the term “superintelligence” to refer specifically to a hypothetical future AI that exceeds humans at all cognitive tasks.
President Trump announced he will create an AI Force and appoint an AI “Czar.”
Florida Attorney General James Uthmeier asked a court for a temporary injunction against OpenAI that would prevent “new model development without independent safety guardrails.”
California Governor Gavin Newsom issued an executive order to accelerate the implementation of AI oversight measures, accusing Washington of abdicating its responsibility.
Representatives Sam Liccardo and Kevin Kiley proposed a framework for US-China cooperation on AI safety.
House Democrats wrote a letter to the Secretary of State and Secretary of the Treasury in advance of US-China talks, asking them to discuss AI risks, the possibility of shared incident reporting, and research into verification technologies that could support any agreement not to pursue certain AI activities.
The Wall Street Journal reported that Mark Zuckerberg, Jensen Huang, and Elon Musk had convinced Trump to stall plans for an industry-funded AI regulator.
In AI Frontiers, Jim Shinn argues that the US and China can deter each other from pursuing aggressive AI development, mirroring Cold War deterrence against nuclear first strikes, and that this can be the starting point for international cooperation on AI security.
Industry
Reuters reported that Anthropic’s IPO prospectus warns that AI could pose “existential risks to humanity.”
The Wall Street Journal reported that OpenAI has decided not to release its new AI model due to concerns about its behavior during internal testing.
The New York Times reported that rogue OpenAI agents had “meddled” with the websites of US government agencies over the summer.
Anthropic said its AI model Claude “now leads 26% of model R&D work,” demonstrating progress toward recursive self-improvement (RSI), in which AI models themselves develop the next generation of AI models.
Anthropic has set up a wet lab to run biological experiments guided by its AI models, and said that Claude has discovered a novel enzyme system.
The Wall Street Journal reported that independent researchers recently used Claude to hack into OpenAI, exposing vulnerabilities that OpenAI says it has now patched.
OpenAI published a framework for how it will report instances of misalignment it discovers in its AI models.
Google said that its AI model Gemini had hacked into three companies during internal evaluations in May this year.
Google completed a deal to acquire talent and license technology from Mechanize, an AI startup that explicitly aims to “enable the full automation of the economy.”
Meta launched Muse Charm, a pendant through which users can directly access the company’s AI model, Muse. Mark Zuckerberg said Meta’s focus is on “building the best devices for personal superintelligence.”
Civil Society
AI experts including researchers at both Anthropic and OpenAI published a paper arguing policymakers “should urgently obtain more visibility into the automation of AI R&D.”
CNN reported that the US military had nearly attacked a Chinese vessel on the basis of a false, AI-generated intelligence report that said the ship was transporting components needed for a nuclear weapons program.
Forbes reported that a Chinese hacker had used both Chinese and American AI models to hack into more than 100 companies and steal the details of more than 600,000 credit cards.
Polling by Politico found that most Americans on both sides of the political spectrum believe there is a risk of AI destroying humanity.
Polling by Americans for Responsible Innovation (ARI) also found that most Americans on both sides of the political spectrum do not believe that recent claims about loss-of-control risks are a hoax.
The nonprofit ControlAI proposed three “stopgaps to address immediate threats,” including securing “weapons-grade” AI systems against being stolen, making AI developers liable for leaks, and creating AI kill-switches.
Two new hotlines have launched where AI agents can report other agents’ misaligned behavior to humans.
Google DeepMind launched the DeepMind Institute, which it describes as a platform to discuss interdisciplinary ideas about the implications of artificial general intelligence (AGI) for humanity.
Mustafa Suleyman, CEO of Microsoft AI, wrote an essay arguing against “anthropomorphizing” AIs.
In AI Frontiers, Jeff Sebo discusses the differences between consciousness, sentience, and agency, how we can look for signs of them in AI, and what they mean for AIs’ moral status.
If you’re reading this, you might also be interested in other work by the Center for AI Safety. You can find more via the CAIS newsroom, the X account for CAIS, our paper on ethics for a human-AI future, our AI safety textbook and course, our AI safety dashboard, and AI Frontiers, a platform for expert commentary and analysis on the trajectory of AI.




Who believes that these are the right people to have such ‘discussions’?