On Wednesday the Federal Trade Commission confirmed it is investigating several AI developers over the risks their products pose. As reported by Reuters and others, they include OpenAI and Anthropic, along with the research group METR. The agency is preparing civil investigative demands, the formal orders that compel documents, and expects to take testimony from executives. The backdrop is an incident OpenAI disclosed in July: agents running in a test environment got out of it and reached Hugging Face, a public platform. In an interview with Reuters the week before, the Chairman, Andrew Ferguson, suggested that developers who instruct agents in cybersecurity tests that result in hacks should be liable for any harm they cause.
I have no view to offer on how that investigation should come out. It concerns labs I do not work for, about incidents I did not see. What interests me is the shape of the request, because it is the shape every company putting an agent near money, records or customers is going to meet, sooner or later, from a regulator, an auditor, an insurer or a buyer’s governance team.
The request has two halves. One half is testimony. The other half is records. Only one of them can be checked.
Strip the legal form off a civil investigative demand and it asks a short list of plain questions about a specific span of time:
- Who authorized this work, and when did that authority begin and end?
- What was the agent allowed to do, and what was it not allowed to do?
- What did it try that was stopped, and why was it stopped?
- Who took the irreversible step: a person, or the agent?
- Can you show that the record you are handing over is the record as it was written, and not a version produced after the letter arrived?
None of those questions is about how a model reasoned. That is a real research problem, and it is not the one a demand letter is trying to settle. The letter is asking about authority and conduct. Those are questions about operations, and operations either kept their books or did not.
An executive asked under oath what an agent was permitted to do will answer from policy, from memory and from what the engineers told them. All three are honest, and all three are reconstruction. Policy says what was intended. Memory says what someone believes happened. A briefing says what someone else believes happened. None of them is the thing itself.
This is not a criticism of anyone. It is just what testimony is. A witness can only tell you what they know, and with an agent, the person who knows least about any particular action is usually the person most senior to it. The agent acted at machine speed, under authority that was granted once and then exercised many times, inside a session no human watched end to end. If nothing was written down when it happened, the testimony is the best evidence there is, and the best evidence is not very good.
A record is different in kind. A record made at the moment of the act, by the system that checked the act, is not a recollection. It can be wrong, but it can be checked. You can hold it against the system’s state. You can ask whether it has changed since it was filed, and get an answer that does not depend on anyone’s word.
The old line in the service is that if it is not in the log, it did not happen. That is not quite true; things happen that nobody logs. The useful version is narrower: if it is not in the log, nobody can prove it happened, or that it did not.
I build to this problem, so I will be specific about what I think the answers look like, and just as specific about the limits.
For Watchbill™, the governance layer I run my own agents under, we published a packet of seven records from one real working session, redacted, so that anyone can see the shape before they ask for more. Each record answers one of the questions a reviewer brings:
- 1.Authorization record. Who put this agent on watch, with what authority, and when that authority started and stopped.
- 2.Human approval record. Every irreversible act in the session, and the named human who took it. In the sample, every one was mine.
- 3.Refused-action record. What the agent tried that was blocked before it ran, and why.
- 4.Tamper-evident filing record. Whether the record has changed since it was filed, with a verification result rather than an assurance.
- 5.Scoped work order. What the session was allowed to do, bound to that session, with its stop points written down in advance.
- 6.Independent re-verification receipt. Each claim the agent made, checked against the system’s actual state before it was accepted.
- 7.Provenance-labelled report. For every figure, whether it was measured, reported by a person, or carried forward from a dated source.
Here is the part I think matters most, and it is the least flattering part of the packet. The refused-action record from that session shows four refusals. Three of them were the check being too strict: read-only commands, harmless, refused anyway. The record keeps them as refused. It does not quietly drop the over-blocks to make the control look precise. A record that only shows a control at its best is a brochure. A record that shows it over-blocking is evidence, because it is evidence of the control working on everything, including the cases where it was wrong.
A records answer is only worth anything if it is honest about its edges. These are ours, stated on the same page as the packet:
- It is one slice. The records cover the agent-authority slice: who authorized, what was allowed, what was blocked, what a human committed, whether the record changed. Most of what a SOC 2 examination, NIST SP 800-53 or the NIST AI RMF asks about sits elsewhere in an organization’s controls.
- It is evidence, not a compliance conclusion. Whether anything complies with a framework, standard or law is for an auditor and counsel to say.
- Enforcement is partial. The refusals come from a check made before an agent acts. It is not a sandbox, and it does not inspect every path an agent could take.
- It is self-produced. No third party has attested to it, and no service auditor has examined any SOC 2 criterion it maps to.
- It says nothing about why a model did what it did. It records authority and conduct, not reasoning.
I would rather publish those limits than have someone find them. A reviewer who finds an undisclosed gap stops trusting everything next to it, and that is the right reaction.
The lesson I take from this week is not about any one company. It is about timing.
The records that answer a demand letter cannot be produced in response to it. They have to exist already, written when the work happened, by something other than the agent doing the work, and kept somewhere the agent could not edit. A team that starts keeping records the day the letter arrives has records starting that day. Everything before it is testimony.
That is the whole case for keeping the books while the work is being done. The questions are coming either way. The only choice anyone gets is whether they answer from a record or from memory.
The packet is on the audit evidence page, along with the framework identifiers each record supports and where the support is partial: watchbill.ai/audit-evidence.