Elon Musk wants the companies competing to build the world's most capable artificial intelligence systems to become each other's safety testers. Speaking at the All-In Summit in Los Angeles, Musk proposed that leading U.S. and Chinese AI developers gain limited pre-release access to rival models so they can independently search for dangerous behavior, security weaknesses and other problems before those systems reach the public.
The idea would bring an unusual form of peer review to one of technology's most competitive markets. Musk specifically named xAI, OpenAI, Anthropic, Google and Meta, along with several leading Chinese AI companies, as potential participants. As Yahoo-hosted coverage of the proposal reports, Musk framed the approach as an alternative to companies evaluating only their own work: competitors would have both the expertise and the incentive to uncover weaknesses another laboratory might overlook.
Turning commercial rivalry into a safety mechanism
The basic logic is familiar from other technical disciplines. Internal teams can become accustomed to the assumptions embedded in their own systems, while an outside evaluator approaches the same product with different methods and incentives. In frontier AI, where model behavior can be difficult to predict across thousands of possible tasks, additional independent testing could expose failure modes that an internal evaluation suite does not capture.
Musk described a shared testing arrangement in which rival laboratories could run their own safety harnesses against another company's model. He suggested that pre-release access might last roughly one or two weeks, accompanied by recurring conversations among participating organizations. If one company discovered a serious concern in a competitor's model, it could raise an alarm before deployment.
The proposal comes amid a broader debate over whether frontier AI development is advancing faster than existing safety practices can comfortably manage. Anthropic CEO Dario Amodei recently called for companies to slow the pace of model development and proposed embedded external evaluators with meaningful access to frontier laboratories. OpenAI CEO Sam Altman and Musk subsequently expressed support for stronger safety coordination, creating a rare area of agreement among executives whose companies otherwise compete intensely.
OpenAI, Anthropic and Google are already discussing safety
Cross-company coordination is not entirely hypothetical. TechCrunch reported that OpenAI, Anthropic and Google DeepMind had already been discussing AI safety for weeks. OpenAI global policy chief Chris Lehane said the company had been working with its rivals while U.S. lawmakers considered how to address potentially catastrophic risks associated with increasingly powerful AI systems.
Those discussions are significant because the competitive structure of the industry makes unilateral restraint difficult. If one company delays a model for additional testing while a rival continues advancing, the cautious laboratory may fear losing customers, talent, investment and technological leadership. A common evaluation framework could reduce some of that pressure by creating expectations that apply across multiple frontier developers rather than leaving each company to decide independently how much testing is enough.
Yet cooperation between competitors also creates legal and governance questions. Agreements among dominant companies can attract antitrust scrutiny, particularly if coordination extends beyond narrowly defined safety evaluation into decisions about product releases, capabilities or market behavior. Any workable framework would therefore need clear boundaries around what information is shared, how evaluators gain access and who determines whether a discovered risk is serious enough to delay deployment.
Why Musk wants Chinese AI companies involved
The international dimension is one of the most distinctive parts of Musk's proposal. Rather than limiting reciprocal testing to American laboratories, he suggested involving three or four leading Chinese AI companies. Musk argued that China might agree to such an arrangement because it could provide a practical mechanism for reducing shared risks without immediately requiring a new global regulatory bureaucracy.
That would be difficult in practice. U.S.-China competition over semiconductors, model capabilities and technological leadership has made AI a national-security issue as well as a commercial one. Giving foreign competitors access to unreleased frontier systems could create concerns about intellectual property, security and capability leakage. The amount of access required for meaningful safety testing would therefore have to be balanced against the risk of exposing sensitive technical information.
China has also pushed back against recent calls for an AI slowdown. Its Foreign Ministry characterized those warnings as fearmongering, according to contemporaneous reports. That does not necessarily rule out cooperation on narrowly defined evaluation standards, but it illustrates the diplomatic challenge of convincing countries engaged in an AI race to expose their most advanced systems to foreign competitors before launch.
Independent testing is different from self-regulation
Musk's proposal sits somewhere between conventional corporate self-regulation and formal government oversight. The participating laboratories would remain private organizations, but the company building a model would no longer be the only institution judging whether that model passed its safety tests. Rival evaluators would introduce an adversarial element that could make assessments more rigorous.
That independence would still be limited. Competitors share an interest in avoiding catastrophic failures that could damage the entire industry, but they also have commercial incentives that could influence how they evaluate one another. A company might be motivated to identify legitimate problems in a rival model, yet the same competitive relationship could create disputes over whether a reported weakness represents a genuine safety issue or simply an opportunity to delay a competitor's release.
External academic researchers, dedicated evaluation organizations and public regulators could therefore remain important even if reciprocal testing becomes common. A diverse evaluation ecosystem is more likely to catch different classes of problems than a single mechanism, particularly as frontier models acquire capabilities in cybersecurity, autonomous software operation and long-running agentic tasks.
The hardest question is what happens after a problem is found
Testing itself is only the first part of the governance problem. A cross-laboratory evaluation system would also need procedures for deciding what happens when a serious vulnerability appears. Participants would have to define thresholds for escalation, determine whether findings remain confidential while fixes are developed, and establish what recourse exists if a company decides to launch despite objections from its evaluators.
Musk has pointed to legal liability and reputational pressure as forces that could encourage companies to respond responsibly. Those incentives are real, but they operate most strongly after consequences become visible. Frontier AI safety is increasingly focused on preventing high-impact failures before deployment, which is why clear pre-release standards and credible escalation procedures would matter as much as the tests themselves.
A rare point of agreement in a fiercely competitive industry
The most notable aspect of Musk's proposal may be the willingness of major AI executives to discuss mechanisms that could constrain the race they are simultaneously trying to win. Anthropic, OpenAI, xAI and Google DeepMind differ sharply in strategy, business models and corporate relationships, yet recent statements suggest a growing recognition that increasingly capable systems may require forms of scrutiny that no single company can provide entirely on its own.
Whether reciprocal testing becomes an industry standard will depend on details that remain unresolved: access, confidentiality, antitrust rules, international participation and the authority to act when evaluators disagree. Musk's proposal does not solve those problems, but it identifies a potentially useful principle. In a field where every frontier laboratory has blind spots, letting sophisticated outsiders try to find them before a model is released could make safety evaluation harder to game and easier to scrutinize.