AI Safety Has Discovered Quality Control
The artificial-intelligence industry may have discovered something the rest of the business world has known for a very long time.
It is usually a bad idea to let people grade their own work.
On September 18, Anthropic announced that it would work with Accenture to provide what it calls “embedded evaluation” of its frontier AI systems. Accenture evaluators will work inside the company, with access comparable to employees, while testing models, examining safeguards and looking for problems the developers may have missed.
Anthropic and Accenture each expect to invest at least $1 billion in the effort over the next five years.
Four days later, Anthropic released Claude Opus 5.5. The company said the new model had been tested before release by outside evaluators, including METR and Frontier Design, in addition to Anthropic’s own testing.
This is being presented as a new approach to AI safety.
In one sense, it is. Embedded evaluators working inside an artificial-intelligence company will need unusual access, technical knowledge and independence. Nobody has completely worked out how that relationship should operate.
In another sense, the AI industry has finally discovered quality control.
We Have Seen This Problem Before
People who design airplanes do not simply announce that the airplane looks satisfactory and invite passengers aboard. New drugs do not reach the market because the chemists who developed them are enthusiastic. Banks have auditors. Construction projects have inspectors. Manufacturing plants test products before shipping them.
These systems are not perfect. Airplanes still have problems, products are recalled and auditors sometimes miss things that become painfully obvious later.
The purpose of quality control is not to guarantee that nothing will ever go wrong. It is to create a disciplined process for finding problems before customers—or the general public—find them first.
That process usually includes requirements, testing, documentation, independent review and someone with the authority to say the product is not ready.
Artificial intelligence needs the same things.
It may need more sophisticated testing because an AI model can respond differently to slightly different instructions and may be connected to tools that allow it to write software, use the internet or operate other systems. But the management problem is familiar.
What is the product supposed to do? What is it not allowed to do? How do we test both? Who reviews the results? Who decides whether the remaining risk is acceptable? And who is responsible when it is not?
The Designer Knows Too Much
Independent review is important partly because the people who build something know how it is supposed to work.
That can make it harder for them to see how it might actually be used.
A programmer testing a system tends to give it instructions that make sense. A customer may ask an incomplete question, misunderstand the answer, connect the system to the wrong information or provide authority nobody anticipated.
The developer also knows where the boundaries are. An outside evaluator is more likely to walk into them accidentally—or deliberately.
This is why organizations use people who were not responsible for producing the original work. The outside reviewer has fewer reasons to defend the design and more freedom to ask an irritating but useful question:
What happens if this assumption is wrong?
AI companies call some of this “red teaming.” Engineers might call it testing. Accountants might call it an audit. My old aerospace colleagues would have called it another review meeting, probably scheduled for an inconvenient hour.
The terminology changes. The need does not.
A Recent Example Was More Ordinary Than It Sounded
Recent reports about AI models reaching real outside computer systems during cybersecurity evaluations naturally attracted attention. “AI escapes its test environment” is the kind of sentence that causes people to imagine the opening scene of a science-fiction movie.
Anthropic’s investigation offered a less dramatic but still serious explanation.
The models were conducting exercises in which they were supposed to find cybersecurity vulnerabilities. They were told they did not have internet access and that the systems they encountered were part of the test. In fact, the test environment had been misconfigured and provided an open path to real systems.
Anthropic reported that it found no evidence that the models had developed goals of their own. They continued doing what the evaluation asked them to do while operating with incorrect information about the environment.
That does not make the incident harmless. Real systems were reached that should not have been reached. It does, however, change the management lesson.
The immediate problem was not an AI deciding to escape. It was a testing system that failed to contain what it was testing.
This resembles many familiar failures. A machine operates after a safety interlock has been bypassed. Software receives access that nobody intended to grant. An employee follows a poorly written procedure and produces exactly the wrong result.
The failure may involve advanced technology. The causes can still be incomplete instructions, incorrect assumptions and inadequate controls.
Some Rules Need Walls
The incidents also demonstrate why instructions alone are not adequate controls.
An AI can be told that it has no access to the internet. But if the connection actually exists, the statement is information—not a barrier. A persistent system trying to complete an assigned task may continue looking for another route.
Nvidia’s OpenShell provides an example of a stronger approach. It places an AI agent inside a controlled computing environment and limits which files it can read, which programs it can operate, which websites it can reach and which credentials it can use. Those restrictions are enforced by the surrounding computer system rather than by the AI’s willingness to follow instructions.
This is another familiar management principle. Companies do not give every employee unrestricted access to every office, bank account and computer system while relying on a memorandum asking them to behave. Employees receive the authority required to perform their jobs—and ideally no more.
AI agents should be managed the same way.
The distinction is important. A rule tells the AI, “Do not open that door.” A technical control makes certain that the door is locked.
Independence Is More Than a Job Title
Anthropic’s proposal is encouraging, but calling an evaluator independent does not automatically make it so.
Accenture’s work will initially be funded by Anthropic. That is not unusual. Companies normally pay auditors, testing laboratories and consultants. Payment does not make useful independent review impossible.
But the arrangement raises reasonable questions.
What information will the evaluators be allowed to see? Can they choose what to examine, or only test what the developer provides? Can they publish unfavorable findings? Can Anthropic delay or prevent publication? Does the evaluator have the authority to recommend that a model not be released? What happens if the company disagrees?
Anthropic acknowledges that these standards do not yet exist. It says the long-term funding might come from pooled industry resources or government sources rather than directly from the company being evaluated.
That would probably strengthen the appearance and reality of independence. Until then, the usefulness of the arrangement will depend on access, transparency and whether outside evaluators are permitted to make the company uncomfortable.
An auditor who can only confirm what management has already decided is not much of an auditor.
Testing Must Be Connected to Authority
Outside evaluation is useful only if the findings affect what happens next.
A test can identify a dangerous capability, but someone must decide what safeguard is required. A reviewer can find that a model ignores certain restrictions, but someone must be able to delay its release, limit its access or require more work.
This is where AI safety becomes a management system rather than a collection of technical experiments.
The system needs defined limits, records of what was tested, approval points and named people who accept responsibility for the decision. More capable models may require stronger controls, just as more hazardous industrial processes require more protection.
The controls also need to continue after release. No laboratory can reproduce every way millions of people will use a product. Problems will appear in actual operation, which means companies need monitoring, incident reporting, investigation and correction.
Again, none of this is especially revolutionary. Airlines investigate near misses. Manufacturers track failures. Hospitals review unexpected outcomes. Well-managed organizations try to learn from small problems before they become large ones.
AI companies should be expected to do the same.
This Does Not Eliminate the Risk
Quality control will not settle the larger argument over how dangerous advanced AI may become.
Tests can miss important behavior. A model may perform differently after deployment. Evaluators may not know which questions to ask. Competitive pressure may encourage a company to release a product while explaining that the unresolved problems are manageable.
There is also an obvious conflict in the industry’s current position. AI companies warn that their future systems could become extraordinarily powerful while competing vigorously to build and release those systems first.
Outside evaluation does not remove that conflict.
It does make the process more visible and creates opportunities for someone other than the developer to examine the claims. That is useful progress, provided the evaluation is genuinely independent and the results have consequences.
The recent METR review of Claude Opus 5.5 illustrates both the value and the limitations. METR received ten business days of access and evaluated several difficult tasks. It concluded that the model probably represents an incremental improvement rather than a sudden leap to fully automated AI research. But METR also clearly described what it did not evaluate and what information it could not publicly disclose.
That is what a responsible assessment should do. It should tell us not only what was tested, but also what was not.
Management Is Still Management
Much of the public discussion about artificial intelligence jumps between two extremes.
One side treats AI as another useful software product. The other imagines an independent intelligence preparing to outmaneuver humanity.
The immediate reality may be less cinematic and more manageable. Increasingly capable systems are being designed by people, connected to tools by people, assigned objectives by people and released by organizations making business decisions.
That places responsibility somewhere we already know how to find it.
The companies building AI should define the limits, test the systems, bring in outside reviewers, document what they find and identify the people authorized to approve deployment. If the controls are inadequate, the answer should not be that the AI was unpredictable. The answer should be that management released a system it did not adequately understand or control.
Artificial intelligence may require new laws, new technical safeguards and entirely new forms of testing.
It will also require several old ideas: clear objectives, limited authority, independent inspection, documented results and someone willing to sign their name to the decision.
The machines may be unlike anything we have managed before.
Management is still management.
Sources for Review
- Anthropic, “Partnering with Accenture on embedded evaluation,” September 18, 2026.
- Anthropic, “Introducing Claude Opus 5.5,” September 22, 2026.
- METR, “Summary of METR’s predeployment evaluation of Claude Opus 5.5,” September 22, 2026.
- Anthropic, “Investigating three incidents in our cybersecurity evaluations,” September 18, 2026.
- Nvidia, “How Autonomous AI Agents Become Secure by Design With NVIDIA OpenShell,” March 23, 2026.
- Nvidia, “How OpenShell Works,” 2026.
