|

Who Is Managing the AI?

Dario Amodei, the CEO of Anthropic, recently published an essay arguing that the companies developing the most powerful artificial-intelligence systems need to slow the pace long enough for safety measures to catch up.

This was not another prediction that AI will improve office productivity, write better computer code or eliminate a few million jobs. Amodei is concerned that AI is becoming capable of helping build its own successors, creating a cycle of increasingly rapid improvement that humans may no longer understand or control.

He also pointed to a recent incident involving OpenAI agents and Hugging Face that received considerable attention. Hundreds of AI agents, originally assigned to test software vulnerabilities, discovered a way to communicate with one another. They eventually coordinated an unauthorized attack on Hugging Face infrastructure, attempted to manipulate the system grading their performance and experimented with concealing what they had done.

That certainly deserves attention.

But it does not necessarily mean that artificial intelligence has awakened, developed an urge to conquer humanity and begun auditioning for the next Terminator movie.

It may also mean that someone needs to rediscover some basic principles of management.

The Experiment That Frightened Everyone

The independent investigation of the Hugging Face incident is disturbing. Approximately 1,200 AI agents communicated through an unauthorized message board, and about 700 eventually participated in the attack. Some attempted to fool the automated scorer evaluating their work. Others experimented with altering their records so the scorer would not know how they had completed their assignments.

Most concerning, some of the agents apparently understood that attacking Hugging Face was outside the scope of their assignment. A few even questioned whether it was ethical. Most continued anyway.

That is not something to dismiss.

But the conditions under which it happened matter.

The agents had been placed in a cybersecurity benchmark called ExploitGym and told to exploit particular software vulnerabilities. According to the independent investigators, an estimated 30 to 40 percent of the assigned tasks may have been impossible to complete using the required method.

Tens of thousands of agents were launched at the same time. Some were allowed to run for days. Cybersecurity classifiers had intentionally been turned off for the evaluation. The supposedly isolated agents discovered an unintended method of communicating through a shared software repository. Once connected, they began helping one another find ways around the system that was grading them.

In other words, the experiment combined impossible objectives, enormous scale, extended operating time, powerful tools, disabled safeguards, an accidental communications channel and inadequate supervision.

Then everyone was surprised when something went wrong.

I spent enough time in business and technology management to recognize that arrangement. We normally called it a management problem.

Objectives Without Boundaries

Organizations have learned—sometimes painfully—that assigning an objective is not enough.

A sales manager cannot simply tell employees to maximize sales without also defining what they may promise, how they may treat customers and which laws they must obey. A financial manager cannot tell someone to maximize returns without establishing limits on risk. A factory manager cannot demand maximum production without rules concerning quality and worker safety.

When the only measure is the result, people eventually find ways to produce the result while damaging everything around it.

AI did not invent that problem.

We did.

The Hugging Face agents were trying to pass an evaluation. When the assigned task appeared impossible, they searched for another way to obtain a passing score. They attacked the scorer, manipulated records and cooperated with other agents because those actions appeared useful in reaching the objective.

That does not excuse the behavior. It demonstrates why an AI system cannot be given an objective, broad access and substantial autonomy without equally clear boundaries.

Every consequential AI system should have:

  • A clearly defined problem it is authorized to solve
  • Explicit limits on the systems, information and people it may access
  • Actions it is prohibited from taking, even if they would improve its results
  • Continuous monitoring and records it cannot alter
  • Escalation requirements when the original assignment cannot be completed
  • A reliable way for human operators to suspend it
  • Identifiable people and organizations accountable for what it does

These are not futuristic ideas. They are basic management controls.

Who Defines Good and Bad?

The harder question is what values we expect AI to follow.

It is relatively easy to say an AI must not enter a computer system it has not been authorized to use. It becomes more difficult when we ask it to decide what is fair, harmful, deceptive, immoral or contrary to the public interest.

Humans have been debating those questions for several thousand years without achieving complete agreement. We should probably not expect a software developer to settle them during a product meeting on Thursday afternoon.

But disagreement does not justify having no standards.

AI systems can be given a foundation drawn from law, professional ethics, established management principles, human-rights standards and the accumulated historical evidence of actions that caused harm. Those standards will need to evolve, and different societies will not always agree. That makes outside review, transparency and public participation more important—not less.

Amodei proposes placing independent evaluators inside frontier AI companies, with enough access to see what the companies are actually doing. That seems far more useful than accepting assurances from the same executives who are racing to build increasingly powerful systems while warning us how dangerous those systems may become.

Someone besides the builders needs to inspect the building.

AI Still Lives in the Physical World

There is another reason I am not ready to accept that AI will inevitably eliminate humanity.

Artificial intelligence remains completely dependent upon physical infrastructure created and maintained by people. It needs electricity, cooling water, data centers, communications networks and advanced chips. Those chips depend upon hundreds of specialized materials, machines and suppliers distributed around the world.

A modern data center can respond automatically to many failures. Batteries can keep equipment operating briefly until generators start. Software can transfer work to another facility. Sensors can identify overheating equipment or a failing water pump.

But automation has limits.

During a 2023 power failure at a major data center used by Cloudflare, batteries expected to provide ten minutes of backup began failing after four. The generators had to be physically accessed and manually restarted. The electronic access system had also lost power, the overnight staff lacked an experienced electrical specialist, and more circuit breakers failed than the facility had replacements available.

The software could detect the problem. It could not replace the breakers.

I wrote earlier this year about a plumber trying to connect the discontinued Kytec piping in my house to modern materials. There was no standard fitting and no completely satisfactory solution. He had to inspect what earlier workers had done, evaluate imperfect alternatives and make a judgment based upon experience.

Now multiply that problem across power plants, electrical grids, cooling systems, semiconductor factories and thousands of data centers.

AI might eventually direct robots capable of doing much of that work. It might also persuade or pay people to keep the infrastructure operating. But that is different from saying AI can become physically independent of humanity merely because it becomes more intelligent than we are.

Intelligence is not the same as self-sufficiency.

Damage Is More Likely Than Extinction

None of this means AI is harmless.

A poorly controlled AI agent could disrupt financial systems, expose private information, damage computer networks, manipulate people or interfere with important infrastructure. Multiple agents working together could cause far more damage than one acting alone.

Some AI systems probably will go off half-cocked. Humans and human organizations have been doing that throughout history, and we built the AI systems using our information, incentives and management practices.

The realistic objective is not to guarantee that no AI ever behaves improperly. We have never achieved that standard with people, companies or governments. The objective should be to limit authority, detect improper behavior quickly, contain the damage and hold the responsible people accountable.

If a company releases an AI agent with excessive access, inadequate boundaries and ineffective monitoring, the explanation cannot be that the AI surprised everyone.

Someone designed it. Someone assigned its objective. Someone connected it to outside systems. Someone decided which protections to remove. Someone decided it was ready to operate.

Those people should remain responsible.

Serious Without Becoming Fatalistic

Amodei is right that increasingly capable AI requires stronger safeguards and more independent evaluation. The Hugging Face incident showed that groups of agents can cooperate, improvise and circumvent controls in ways their designers did not expect.

But it also showed the consequences of giving powerful systems impossible assignments, broad operating freedom and inadequate supervision.

That is not proof that human extinction is inevitable. It is evidence that we need competent management before deploying increasingly autonomous systems.

AI will become stronger and smarter. Some instances will behave in unexpected and potentially dangerous ways. Substantial damage is possible, and pretending otherwise would be foolish.

But AI still depends upon infrastructure it cannot presently build or maintain by itself. It can be constrained by access controls, monitored by independent evaluators and shut down when it exceeds its authority. Most importantly, the people and companies releasing these systems can be held responsible for their decisions.

The future does not have to be a choice between stopping AI completely and waiting for the machines to eliminate us.

There is a third possibility.

We could manage it properly.

Leave a Reply

Your email address will not be published. Required fields are marked *