When the World Runs the Test: Field Events, Functional Safety, and the Living Safety Case
Your safety case isn't finished when you ship. Real-world events will test your autonomous machine whether you plan for them or not. Learn how to turn field events into safety evidence, and why ISO/IEC TS 22440 and the CRA are moving are moving toward safety that continues after deployment..
Sooner or later, the world runs the test on your autonomous machine. Not on your schedule, and not in a lab. That's worth staying with, because here's the part that often goes unsaid: some real-world events are going to happen no matter how thorough your safety case is. In an uncertain and unpredictable world, it is impossible to anticipate every situation a machine will encounter. The real question is what happens when something new does occur, and whether your safety process can take the event, learn from it, adapt, and use it as feedback to update your safety case, and keep the traceability intact.
So the question can't only be how to keep the world from running your safety test cases. It has to also be what you do with the result when it does.
Most teams don't have a good answer. A field event comes in and gets handled as an incident: logged, triaged, patched, closed. The machine gets safer and the paperwork gets longer. But the event never travels the distance that matters, back to the safety case it speaks to, the certification that case underwrites, and the assurance argument everything else stands on.
Because that's what a field event actually is. It's evidence. It either confirms something your safety case already asserts, or it introduces a risk you never foresaw. The assumptions you made at design time, how often a condition would materialize, how a sensor would react, what a person would do at close range, the field is the first place any of them gets checked against reality instead of against a presumption. And if you can't trace the event back to the specific claim it touched, you aren't learning from it. You're just logging it.
Which is the uncomfortable thing the title is pointing at. Your safety case was never finished the day you shipped, and your safety certification isn't the finish line. Each captures what was true at the time, but plenty happens after that moment: you change something in the system, or the field turns up something you didn't expect. Both are new information, and both belong back in the safety case. A safety case isn't a document you close. It's an argument you keep true while the world keeps testing it.
What it means to trace an event back
The reason field events get handled as incidents and not as evidence is mostly mechanical. The safety case lives in one place, the field data lives in another, and the line between a specific thing that happened and the specific claim it bears on has to be drawn by hand, if anyone draws it at all. That work is slow, and it relies on the engineer that outlined the hazards and requirements in the first place. It's exactly the kind of task that quietly slows once an autonomous system is shipped.
The practical version of "a field event impacts the safety case" is narrower, and more useful, than it sounds. It means the event has an address. When something happens, you can follow it to the hazard it relates to, the requirement that was meant to mitigate it, and the claim that requirement was supporting. You can see which part of the argument just got confirmed, and which part just got a question mark next to it.
That distinction is the whole point. Not every field event weakens your safety case. Plenty of them strengthen it: the hazardous scenario you anticipated came up, and the system handled it the way you said it would. That's evidence too, and it's worth keeping, because a safety argument that only ever collects its failures is telling you only part of the story. The events that matter most do one of two things: challenge a claim you were sure of, or reveal a risk you never anticipated, one the open world produced and your case never named.
The recall turns back into a test
There are two ways to meet a hard situation. You can go looking for it on purpose, staging it through physical testing, in a room you control, before the field stages it for you. Or it reaches the field first, which, as argued above, some always will. Catching that event and tracing it to a safety claim gives you something you can act on. You take it back into physical testing and make it a test you run on purpose from then on.
The field found the thing. Physical testing gives you the evidence that the mitigation holds, working alongside other verification methods with added value of running on a real autonomous system. The event that could have been a recall becomes a test, which updates and improves the safety case overall.
That also settles what field learning does to physical testing, because it's a mistake to think that monitoring what happens in the field makes physical testing less necessary. It's the reverse. Without somewhere to stage what the field teaches you, a field event is just a story you tell after the fact. Physical testing is what turns it into standing knowledge. The two are one loop: you test what you can foresee, the field surfaces what you couldn't, and what it surfaces becomes the next thing you test. Or, you can identify edge cases and adversarial conditions and repeatedly test those as well, and that is something you can do prior to shipping. Either way, the safety case becomes a living argument that is constantly being fine-tuned. A program that only tests, or only watches the field, is running half a loop. The safety case gets to be a live argument only when both halves are turning.
ISO/IEC TS 22440 and the Cyber Resilience Act point the same way
If the argument on its own merits doesn't move you, look at where the rules are heading. Look at what's publicly known about ISO/IEC TS 22440, the technical specification on functional safety and AI systems, alongside the EU Cyber Resilience Act (CRA), and a pattern emerges. They cover different ground, one about learned behavior and the other about attackers, but neither treats the product as finished at release.
TS 22440 is still under development, so its final content may change. Based on ISO's published scope and public descriptions of the project, it is expected to cover an AI safety lifecycle with its own data lifecycle, AI-specific fault analysis and mitigations, and testing that supports a justifiable assurance argument. It's designed to build on IEC 61508 instead of replacing it, and its scope includes how security threats can affect the safety of an AI system. The CRA is lifecycle-based too. It requires a cybersecurity risk assessment for each product, vulnerability handling throughout a defined support period, and, since 11 September 2026, reporting of actively exploited vulnerabilities and severe incidents.
They aren't interchangeable. TS 22440 is specific to functional safety of AI systems and isn't expected to be published until early to mid-2027. The CRA covers cybersecurity of any product with digital elements, AI or not, and its reporting duties are already live.
Then there's the stopwatch. The CRA requires manufacturers to send an early warning within 24 hours of becoming aware of an actively exploited vulnerability or severe incident, and a fuller notification within 72 hours. The law already assumes a live channel running from the field back to the people responsible for the product, and it assumes you can act on what comes through it.
That channel covers security events. Safety has reporting duties of its own in places, from accident reporting for consumer products in the EU to crash reporting for automated vehicles in the US, and from January 2027 the EU Machinery Regulation requires manufacturers to inform authorities when a machine presents a risk. What those duties ask for is telling a regulator what happened. They don't ask you to trace it back to the specific claim in your safety case that it touched, and that is the piece still missing. A system, especially a system that incorporates AI or machine learning, can't just be considered one and done with respect to safety certification, and the emerging TS 22440 reflects that. Practitioners commenting on the standard stress continuous field monitoring for AI-enabled systems. No system can be validated against every scenario in advance, but with learned behavior the gaps are harder to predict, which is why monitoring matters more. When a requirement to close that loop arrives, and the standards work suggests it's a question of when rather than if, the programs already running a field-to-case loop will be adapting a habit. Everyone else will be building the capability under a deadline, which is the worst possible time to build anything.
So this isn't only the responsible move. It's the early version of the required one. Doing it now, on your own terms, is less risky than doing it later because an auditor or a regulation made you.
The day something goes wrong in public
Picture the viral clip, the demonstration that went sideways, the single unanticipated instant that becomes the whole story. That day is coming for a range of players in the robotics industry, and as argued above, you can't fully prevent it.
What you can decide ahead of time is what you're holding when it happens. On one side, a team that logged the event, patched something, and now has to reconstruct what it knew and when after the fact. On the other, a team that can show the event, the claim it touched, the analysis that followed, and the change that closed it out, all traced and dated. One of those is scrambling to assemble a story under the worst possible circumstances. The other is closing a feedback loop between its safety narrative and the open world. That's the line between a liability and an event. Exposure grows with every unit in the field, and the defensible team is the one that builds the framework for incorporating what the world throws at it before it's needed, not after.
A case you keep true
None of this makes the safety case self-correcting, and it shouldn't. The loop doesn't decide anything. It puts the right question in front of the right engineer at the moment a field event calls a claim into doubt, and it keeps the record of what happened and what mitigation happened afterward. The judgment stays human. What changes is that the argument stops being something you finished and filed, and becomes something you keep true.
Because the world is going to keep running the test. It always has. The only real choice is whether you're still reading the results.
:format(webp))