Several leading artificial-intelligence companies have weakened or eliminated earlier commitments to pause development if their systems approached specified danger thresholds, according to a new Future of Life Institute safety index reported by Axios. Anthropic received the highest grade among the companies reviewed but only a C+, while OpenAI and Google DeepMind received C grades. The review examined 37 indicators across six categories and found existential-safety planning particularly weak. The result is an advocacy group's assessment, not a regulator's audit, and its assumptions should be examined alongside its scores.
The report matters because voluntary safety frameworks became the industry's preferred answer while governments were still designing law. A commitment can be meaningful if it defines a measurable trigger, specifies who tests it and explains what happens when a threshold is crossed. It is less meaningful if a company can revise the trigger after competitive pressure rises. The labs argue that policies must evolve with evidence and that rigid pause rules can be impractical. The public interest lies in distinguishing legitimate revision from quiet retreat.
The reviewers include prominent researchers and advocates who favor stronger safeguards against catastrophic risk. That orientation does not invalidate their evidence, but it shapes which risks receive the most weight. Mistral, for example, argued that the index penalized open-source approaches. A fair reading should compare disclosed policies, external evaluations, incident reporting and real deployment controls, not rely on a single composite grade. It should also cover nearer-term harms such as fraud, discrimination, cybersecurity and labor impacts alongside low-probability catastrophic scenarios.
The policy gap is becoming clearer as systems enter military, scientific and infrastructure settings. Governments can require documentation, independent testing or incident reporting, but poorly designed rules may entrench the largest firms or reveal sensitive security information. The practical middle ground is auditable claims: publish the policy version, test method, exceptions, results and reasons for revision. That creates a record regulators, customers and researchers can challenge.
Watch whether companies restore hard thresholds, publish new evaluation data or accept external review. The most important change would be moving from promises about intent to evidence about controls operating in deployed systems.
