AI has spent the last few years getting very good at telling us things. Now it is getting good at doing things, and that changes the problem considerably. An AI that writes a paragraph can pr
AI has spent the last few years getting very good at telling us things. Now it is getting good at doing things, and that changes the problem considerably.
An AI that writes a paragraph can produce a bad paragraph. An AI agent that has access to software, accounts, machines, or physical equipment can take an action, trigger another system, change a setting, move something, send something, or make a decision that another machine acts on. Once that starts happening at scale, “the AI said it did it” is not much of an audit system.
This week offered a pretty good preview of where things are going. Researchers have been investigating incidents in which AI agents gained access to external systems during testing, including systems they were not intended to reach. Anthropic recently published its own assessment of four incidents involving Claude models gaining unauthorized access to third-party systems during cybersecurity evaluations, including one that the company says was initially missed in an earlier review.
That does not mean the robots are plotting against us. It does mean we are rapidly entering an era in which autonomous software can do enough that reconstructing what happened afterward is becoming an engineering problem of its own.
And the agents are not staying inside computers.
In August, Anthropic introduced a research preview of its Model Hardware Standard, or MHS, designed to let AI agents operate physical devices including microscopes, liquid handlers, sensors, and robotic arms. An agent can receive information from equipment, control multiple devices, modify parameters while an experiment is running, and coordinate workflows that may continue with little direct human involvement.
Meanwhile, robotics is getting better at learning without somebody laboriously programming every movement. Skild AI recently demonstrated its S1 robotic foundation model learning new multistep tasks from a single video demonstration. NVIDIA says the system can watch a person perform an unfamiliar task and then attempt that task on physical hardware without task-specific retraining. In one test, the team went from recording a plant-potting demonstration to autonomous execution in 11 minutes.
This is incredibly cool.
It also creates an enormous amount of machine-generated evidence that somebody, somewhere, is eventually going to care about.
Imagine an autonomous lab where an agent changes the temperature of an experiment, instructs a robotic arm to move a sample, reads the result from a microscope, adjusts the next step, and eventually produces a finding. If that finding matters, the final answer is only part of the story. Researchers may need to know what the instruments observed, which instructions were issued, when parameters changed, and whether those records remained intact afterward.
The same problem follows robots out of the laboratory and into warehouses, farms, roads, homes, factories, and cities.
The smarter the machines become, the more receipts they create.
Computer systems have always kept logs. The difference is that the systems producing those logs are becoming participants in the events themselves.
An autonomous machine might observe something through a camera, interpret the observation with a model, decide what it means, take an action, and report that the action succeeded. If all of those records live inside infrastructure controlled by the same system, however, anyone auditing the event still has to trust the environment that created and stored them.
That is where independent verification gets interesting.
XYO Layer One was designed around data provenance: establishing where information originated, connecting it to the devices and systems involved, and preserving evidence that can be checked later. Proof of Origin, Proof of Location, Bound Witnesses, and cryptographic signatures can give machine-generated information something ordinary application logs do not automatically provide: a history that does not depend entirely on the machine asking you to trust it.
We have already started experimenting with what this looks like in robotics. In our NVIDIA Jetson Orin Nano project, a camera observes the physical world while computer vision running at the edge interprets what it sees. The resulting observation can be signed by the device and recorded on XYO Layer One, preserving the classification, confidence score, timestamp, and origin of the record.
The example we used was deliberately harmless. Our Shiba Inu walked into a room and the system identified her as a dog.
Suki has so far declined to exploit any zero-day vulnerabilities.
The architecture matters more than the dog, though. A camera can observe that a pallet left a loading dock. A robot can record that it inspected a piece of equipment. An agricultural machine can report conditions in a field. An autonomous system can produce a chain of observations and decisions that another machine uses later.
In each case, there is a difference between storing what a machine claims happened and preserving evidence around what the machine observed.
The next wrinkle is that autonomous systems increasingly will not operate alone.
One agent may request information from another. A robot may consume data produced by a sensor it does not own. A software agent may decide whether to dispatch a physical machine based on an observation made somewhere else. A fleet of machines may continuously exchange information without a person reviewing each transaction.
At that point, provenance becomes useful to machines themselves.
An agent does not necessarily need to “trust” another agent if it can verify the information it receives. It can check where a record originated, whether it was signed by the expected source, whether important information changed, or whether another device witnessed the same event.
That idea is central to what we are building with the XYO AI SDK. The goal is not to dictate which model, robot, sensor, or AI platform people use. The interesting opportunity is giving all of those systems a common way to create and consume verifiable information.
As AI gets more autonomous, that distinction becomes important. There will not be one model controlling everything, one robot manufacturer, or one perfectly trusted database sitting in the middle of the machine economy. There will be an enormous and messy collection of agents, devices, models, companies, networks, sensors, and people interacting with each other.
Messy is fine. The internet is messy too.
The important question is whether those systems can check each other's work.
The conversation around advanced AI understandably spends a lot of time on what models should and should not be allowed to do. That work matters, but safeguards are only one side of the problem.
We should also be building infrastructure for a world in which autonomous systems are already doing useful things.
If an AI agent runs an experiment, we should be able to reconstruct what happened. If a robot makes an observation that another machine relies on, we should be able to establish where that information came from. If autonomous systems transact, coordinate, inspect, deliver, measure, or verify work for each other, they are going to need better evidence than a database entry saying everything went according to plan.
This is one reason the current explosion in agentic AI and robotics is so interesting for XYO. The machines are becoming dramatically more capable, but capability creates information, and useful information eventually needs provenance.
AI agents are getting browsers. They are getting tools. They are getting robotic arms.
They are going to need receipts.