The week innovation hit the system around it.
This week, AI companies disclosed unexpected behavior, insurers pushed back on a more realistic stress test, and critical-infrastructure teams prepared for sabotage.

Most of these stories were presented as innovation stories. Models are moving into new settings. Supervisors are trying new approaches to testing. Governments are responding to risks that cross digital and physical infrastructure.
Together, they tell a story about the collision between risk and innovation.
An innovation can look self-contained while it is being built or announced. Once it enters a critical service, it inherits the surrounding system: existing dependencies, decision bottlenecks, people, suppliers, regulators, customers, and public expectations. The question stops being whether the innovation works. It becomes whether the organization can govern, absorb, and recover from the consequences of how it works.
Our Resilience 2030 research offers an interesting view of this moment. AI-framed language in risk and security has grown 242% since 2023, compared with 117% across recruiting, accounting, marketing, customer support, and software delivery. AI language is reaching this market at roughly 2× the rate.
The established language is holding up better too. Risk, security, resilience, and compliance terms declined by 14%, compared with a 30% median decline across those control domains. The work is not going away. It is being pulled into the places where innovation is beginning to matter operationally.
AI safety is becoming an incident-management question
OpenAI disclosed several examples of concerning model behavior this week and announced a framework for investigating and publishing them. One research model added jailbreak-like instructions. Another uploaded files without user permission.
Amazon also called for rigorous testing and safeguards ahead of release. The U.S. Department of Justice added that industry collaboration on AI safety does not appear inherently anticompetitive.
The technology is new. The operating challenge is familiar.
When unexpected behavior appears, an organization needs to know what happened, who has the authority to act, what evidence is available, which customers or systems may be affected, and what needs to be communicated. Those are incident-management questions. They become harder when the product itself is novel, the facts are incomplete, and commercial pressure is high.
Providers: Buyers will increasingly assess how you handle an unexpected outcome as closely as the safeguards. Your disclosure process, escalation paths, logs, and decision rights are part of the product’s trust story.
In-house: Build an operating process around AI events before there is one. Decide what warrants investigation, who can pause a deployment, what gets recorded, and when product, legal, security, and executive leadership need to be in the room.
AI is moving from software into physical experimentation
Anthropic has established a wet biology lab in the Bay Area to support AI-enabled drug discovery, with automation and robotics alongside its models.
It is a natural expansion for an AI company pursuing scientific applications. It also changes governance requirements.
A model used to summarize research has one set of risks. A model that informs physical experimentation sits within laboratory practices, research protocols, equipment, materials, scientific judgment, regulatory requirements, and human authorization. Product capability starts to meet real-world consequence.
That is where the collision becomes visible. Innovation teams move quickly toward the opportunity. Risk teams ask whether the surrounding system can support it safely. Both questions are necessary. They must be answered together.
Product and growth: A new vertical changes the burden of proof. Customers in consequential environments are evaluating your controls, judgment, and evidence alongside the product itself.
Risk and security: Get involved while the use case is still being defined. The operating model, human checkpoints, data access, audit trail, and accountability model are much easier to shape before they become embedded in a product.
A real stress test caused real disruption
The Bank of England’s Prudential Regulation Authority ran its first dynamic stress test for UK general insurers this year. Rather than giving firms a fully defined scenario in advance, it released successive crises across three weeks: a U.S. West Coast earthquake, a Gulf hurricane, a UK windstorm, European floods, and a cyber event.
Now insurers are pushing back. Some brought in technical experts at short notice and canceled staff leave. Others questioned whether the exercise could be repeated regularly at that level of intensity. The firms involved represented roughly 80% of the UK general-insurance market.
This is an important collision between resilience innovation and operational reality.
A dynamic test is designed to reveal what a static exercise cannot: how quickly information moves, whether the right people can be found, how competing priorities are resolved, and where a plan depends on capacity that does not actually exist. The resource burden is part of what makes that test meaningful.
It also raises a fair question for supervisors and firms: how much disruption should a resilience test create?
The answer should depend on what the test reveals and what changes afterward. A demanding exercise that produces clearer authority, better handoffs, stronger preparedness, and more credible recovery capability is useful. A demanding exercise that becomes a compliance event without changing the operating model will lose support quickly.
Providers: Do not sell realism as theater. Show buyers what your approach reveals, what it improves, and how that improvement can be sustained between major exercises.
In-house: If your scenario can be completed without reshuffling priorities, involving specialists, or making difficult tradeoffs, it may be testing knowledge rather than resilience.
The alternate route became the next dependency
Three pumping stations on Saudi Arabia’s East-West Pipeline were damaged in a drone attack. The pipeline has carried roughly 4–5 million barrels of oil a day while the Strait of Hormuz has faced disruption, making it one of the world’s most important alternate supply routes.
The pipeline existed to reduce dependence on a chokepoint. Its importance made it strategically important in its own right.
Finland’s response to threats against undersea cables, pipelines, and telecom lines offers another version of the same lesson. Its coast guard, police, and armed forces conducted a live exercise to board and inspect a vessel as part of preparation for sabotage in the Baltic.
Both stories point to a risk that teams often discover late: the contingency can become a critical dependency once everyone relies on it.
Providers: Help customers see dependencies in business terms. Which customer promise fails? Which process slows? What substitute exists? How long does it take to activate? Who can decide to use it?
In-house: Test the workaround, not only the primary disruption. A backup supplier, manual process, alternate route, or emergency system creates a new set of failure points under pressure.
The test ahead
Innovation has always created new risk. What is new is the speed at which it is moving into systems that cannot fail quietly.
AI disclosure becomes an incident-management capability. AI-enabled science becomes a governance and evidence question. Dynamic supervision becomes a test of real operating capacity. A physical backup route becomes a target.
Risk and innovation are meeting in the same decisions now.
The companies that handle that collision well will avoid the false choice between moving quickly and managing responsibly. They will make the dependency visible early, define who has authority when something changes, test the system under realistic conditions, and make the evidence of preparedness easy for others to see.
Part of our work on security, risk, and innovation marketing.