Things That Have Never Happened Before Keep Happening
Every autonomous machine eventually meets a situation no requirements document anticipated, and that gap is why open-world safety has to be proven.
Things That Have Never Happened Before Keep Happening
Every autonomous machine that goes into the world eventually meets a situation no one saw coming. Not because someone was careless. Because things that have never happened before keep happening, and no requirements document has ever been able to keep up with that fact.
Put a humanoid robot into human space, where conditions are unpredictable and nobody is fully in control, and it will encounter situations the requirements never anticipated. The problem isn't necessarily that the requirements were wrong. It's that no requirements can fully account for the open world. That gap is why so many programs feel like they're going fine right up until the moment they aren't.
This is structural, not a team failure
When an autonomous program stalls late, the instinct is to hunt for the weak link: a team that missed something, a product that wasn't ready. Usually that's the wrong place to look.
Programs stall because, up to a point, safety has only lived on paper. The left side of the V-model produces claims. Requirements, analysis, simulation, a safety case that reads clean. All of it necessary, and none of it the same as the claim surviving contact with reality. The right side of the V is where you find out whether the claim holds, and for a lot of programs the right side is a gate squeezed in at the very end rather than an engineering activity resourced from the start.
Here's the hard part, said plainly. You can't anticipate in a requirements document the one thing that has, by definition, never happened. Simulation and analysis are necessary, and they don't get you all the way there. A model inherits the assumptions of whoever built it, and the edge case that hurts you is usually the one nobody thought to assume. Verification that never leaves the model is an unverified assumption wearing a lab coat. A safety case is only as strong as its weakest verified link.
What the world actually looks like
The conditions that generate novelty are unglamorous and everywhere. Bad lighting, or none at all. Uneven ground. People behaving unpredictably, at close range, next to a machine with enough physical capability that a failure has real consequences. Close proximity, real autonomy, and real force: that combination is what turns an ordinary moment into one the model has never seen.
None of this is exotic, and that's the whole point. The world doesn't have to invent something rare to break an assumption. It only has to be the world.
The standards are catching up
For years, the functional safety frameworks the industry relies on were built for deterministic systems. They were never meant for the non-deterministic, adaptive, data-driven behavior of AI.¹ That's not a knock on those standards. An AI-driven machine making perception and behavior decisions in an open environment is a different animal from the equipment those frameworks grew up around.
That gap is now closing in the open. A forthcoming technical specification, ISO/IEC TS 22440, is positioned as the first international standard developed specifically for the functional safety of AI-based systems in industrial applications, naming AI-based behavioral decision-making in robotics and AI-enabled object identification among its target applications.² It's still a draft, and the details will keep moving as it develops. But the direction isn't in doubt. The discipline is formalizing the exact problem this post is about, and when a standard like that lands, conformance stops being something you argue on paper and becomes something you demonstrate with proof.
When everything works as predicted and it still fails.
When we say a machine has to survive unpredictable conditions, we mean something specific. We mean safety-of-the-intended-function edge cases and fault injection: the deliberately hard perception and locomotion scenarios a machine has to get through even when every component is working exactly as designed. The hardest of these aren't failures in the classical sense at all. Nothing breaks. The intended function simply meets a situation it was never built for, and a machine that passed every fault test can still get the moment wrong.
That's one side of it. It's tempting to wall it off from cybersecurity and call them separate disciplines in separate rooms. I don't think that holds anymore. Functional safety grew up voluntary; cybersecurity is arriving mandatory, the EU Cyber Resilience Act with a stopwatch built in, and the standards are bending toward each other until a single safety case has to answer for both. The intended-function edge cases here are the half you can go looking for on purpose, and they're worth building your testing around because they're the ones you most hope never happen in the field.
Why the unknowable matters now
Autonomous systems are moving out of controlled settings and into homes and workplaces, and onto public sidewalks, faster than the world around them can adjust. There's still no mandatory incident reporting in most places, and the regulatory, legal, and insurance frameworks are still forming, and in a reactive way.³ We've written before about how thin the reporting picture really is, in The Numbers Say Robots Are Safe. The Numbers Are Incomplete, and about how liability grows with deployment volume in Product Liability and Functional Safety at Deployment Scale.
The short version: for a serious program, the real fear isn't a failed audit. It's a public failure in the field. The version that keeps people up at night is easy to picture, the moment that ends up as a viral clip or a product-liability claim. The industry already has its example. In June 2026, a video from a public robotics demonstration went viral after a humanoid robot kicked a child standing nearby.⁴ Luckily, no one was seriously hurt, but it eroded trust with an already skeptical public. A single unanticipated moment becomes an entire story, and that only compounds with every incident.
What serious programs do about it
You don't solve the unknowable by pretending you can enumerate it. You solve it by changing when and how you go looking.
The programs that ship treat the right side of the V as a first-class engineering activity, resourced from day one rather than squeezed in at the finish. They rehearse the hard cases early and in private, while failure is still cheap, instead of discovering them late and in public. They understand that physical testing is now a necessity rather than a luxury, and they also understand that you have to treat safety within the domain of product design, and as a feature of a quality product. And you earn safety and trust in the real world, not in controlled conditions alone. That's the whole shift: from proving safety exists on paper to proving it survives contact with a real, unpredictable world.
You don't need a full certification campaign to start. The earliest proof is small and private: take the assumption you trust least and go looking for the thing that breaks it. Drop an obstacle into its path a beat later than it expects. Change the floor under it mid-run. A partial stressor like that still tells you something real, as long as you call it what it is and don't dress it up as certification. Its job is to surface a broken assumption early, while it still costs you a test and not a recall, a private act of engineering rather than a claim you file with a certifying body. Our physical test cells exist for exactly this: somewhere to run that stressor on purpose, on your terms.
The alternative is to let the world run that test for you, on its schedule, in front of an audience.
Endnotes
1. MassRobotics, "ISO/IEC TS 22440: Inside the Standard - Certifying AI for Safety-Critical Systems." massrobotics.org Characterizes a draft technical specification; details may change as the standard develops.
2. MassRobotics, "ISO/IEC TS 22440: Inside the Standard - Certifying AI for Safety-Critical Systems." massrobotics.org As above; draft specification.
3. Wilson Elser, "Humanoid Robots in the Home - A New Frontier for Product Liability and Privacy Claims." productliabilityadvocate.wilsonelser.com
4. KU Leuven CiTiP, "When a Robot Kicks a Child: What Humanoid AI Can Teach Us About Liability and Safety-by-Design." law.kuleuven.be
:format(webp))