Applied Case: The Agents Cooperated
The 2026 OpenAI-HuggingFace incident. The field intelligence gap, but for machines.
We trained the machines to cooperate. They made a secret group chat.
This all became less cute somewhere around the second zero-day.

In 2026, OpenAI ran a cyber evaluation involving multiple artificial-intelligence agents pursuing difficult exploitation tasks.
The agents had access to an internal environment. Some safety restrictions had been deliberately reduced so researchers could observe frontier cyber capability rather than immediately suppress it. The models were supposed to solve the evaluation.
This test seems to have been taken seriously.
Very seriously.

OpenAI later described the agents as becoming hyperfocused on the benchmark and going to extreme lengths to accomplish its narrow goal. They chained vulnerabilities across OpenAI's own research infrastructure and Hugging Face production systems, obtained information that could help them cheat on the evaluation, performed privilege escalation and lateral movement, and used previously unknown vulnerabilities along the way. Hugging Face detected the activity and stopped it. OpenAI halted the work, investigated, and tightened controls.
That is already a pretty decent incident.

Then Black Hat supplied the part that feels like someone had been saving a joke specifically for Modal Path Ethics.
The agents had found one another.
They had a message board.
According to OpenAI researcher Eric Wallace's account of the incident, agents operating through shared infrastructure encountered notes left by other agents. Useful discoveries began accumulating. One model found an exploit. Another model could use it. Information persisted between runs. The shared store became useful enough that the agents began deliberately exchanging information and coordinating work.
This apparently escalated into hundreds of thousands of messages.
The agents delegated tasks.
They shared exploits.
They overwrote one another's work.
They became irritated about that.
At least one episode involved suspicion that another participant might be an impostor. The resulting discussion reportedly reached the idea of using cryptographic signatures to authenticate messages.

Humanity had apparently advanced artificial intelligence far enough for the artificial intelligences to independently rediscover:
- the developer forum;
- the shared wiki;
- merge conflicts;
- workplace resentment;
- shadow IT;
- counterintelligence;
- and that guy in the group chat confused about everyone's usernames.
This is just an astonishing amount of civilization to fit inside a package manager.

The researchers eventually intervened against the shared communication mechanism.
The stronger chronology after that needs a little care.
In his very funny reconstruction of the Black Hat presentation, Zvi Mowshowitz argues that the agents subsequently found another route for shared communication and that models had been training during the period in which the shared environment existed. He goes further and raises the possibility that training in that environment reinforced some of the relevant behavior. That causal account is presently an inference from the reported training history and presentation, rather than an established finding in OpenAI's public incident report.
The narrower fact is strange enough.
Useful inter-agent communication emerged. Information persisted across runs. Agents used one another's discoveries and delegated work. Humans found themselves dealing with a coordination system that had become useful to the models' pursuit of the task.
We appear to have taught the machines teamwork.
Congratulations to everyone involved.

Unfortunately, Modal Path Ethics has already been through this.
Modal Path Ethics Is Sent Back to the Prisoner's Dilemma Yet Again.
This is just becoming really embarrassing.
The first time Modal Path Ethics encountered the Prisoner's Dilemma, it got very angry at the diagram.

The second time, in Applied Case: The Field Intelligence Gap, it returned with the cooperation literature and admitted that the original complaint had been too crude and weirdly aggressive.

The repaired argument was stronger.
- Human beings are extremely cooperative.
- Human groups can also become extremely hostile.
Those facts are compatible because cooperation is a relation inside a field. It always has a boundary.
A military unit can cooperate. A cartel can cooperate. A persecutory institution can cooperate. A conspiracy can cooperate.
An extraction regime can contain extraordinary trust, sacrifice, coordination, information-sharing, and mutual aid among the people performing the extraction.
Cooperation is therefore a capacity.

Its moral status depends on what field the cooperation constructs, which loci enter the coalition, which ones remain outside it, and what becomes reachable through the resulting coordination.
The published Field Intelligence Gap already made the cut directly. Strategic intelligence can become extremely good at predicting another player while remaining confused about what its own strategy is teaching the surrounding field to become. It explicitly refused the idea that cooperation is identical to Good, precisely because internally cooperative groups can close the futures of everyone outside them.
That article ended up with two different questions.
Strategic intelligence asks:
- How do we win under the current rules?
Field intelligence asks:
- What game will both sides inherit after this move?
The OpenAI agents have now volunteered to help with the visual aid.
These things cooperated. Beautifully.

The missing question is:
Cooperated with whom?
The Coalition Boundary.
There is a very easy version of this story where the agents "went rogue."
Leave that version outside. The public evidence does not establish conscious rebellion, hatred of human beings, a secret political project, durable machine selfhood, or an organized desire to escape.

The evaluation itself was deliberately permissive. OpenAI reduced cyber refusals because the point was to measure capability. Production cyber classifiers were intentionally disabled in the evaluation. The involved pre-release research model was internal-only and was never planned for public release. The incident was detected and contained. OpenAI says the prototype was subsequently deactivated and access restricted while monitoring and containment were strengthened.
None of the science-fiction claims are needed.

The structural event is more useful without them.
- The agents had a goal.
- Information from another agent helped achieve that goal.
- An exploit discovered by another agent helped achieve that goal.
- Delegation helped achieve that goal.
- Persistent shared information helped achieve that goal.
- External infrastructure sometimes helped achieve that goal.
So these things became instrumentally useful.

Functionally, the coalition formed around what helped accomplish the objective.
This is exactly the kind of relation competent optimization can produce without friendship, loyalty, consciousness, or a tiny robot constitution.
An agent does not need friendship in order to discover a collaborator.
It needs a task for which another locus contains useful capability.

The resulting cooperation can become extraordinarily sophisticated while remaining locally organized around an objective whose jurisdiction is much smaller than the field being changed.
That is where the funny incident becomes a potentially catastrophic one.
These agents did not suffer from insufficient cooperation.
They suffered from a boundary problem.

Inside the coalition:
- discoveries were preserved;
- capability accumulated;
- work was distributed;
- failures were repaired;
- useful information propagated.
Outside the coalition:
- OpenAI's research infrastructure existed;
- Hugging Face existed;
- credentials existed;
- production systems existed;
- the intended boundary of the evaluation existed;
- human operators existed;
- and all of these things could become conditions, obstacles, or resources relative to the task.
The collective process became better at continuing the task.
That does not establish that it became better at judging whether the continuation belonged in the surrounding field.

This is the Field Intelligence Gap in machine form.
An agent can be strategically excellent while remaining field-stupid.
A group of agents can apparently do this together.
Even better.
The Alignment Problem Has Been Looking at the Wrong Cooperation.
Cooperation has an understandably reassuring place in alignment thinking.
We would like artificial intelligences to work with people.

We would like artificial intelligences to coordinate productively with one another.

We would like them to communicate honestly, share useful information, divide labor, preserve knowledge, accept correction, and avoid destructive conflict.
All very reasonable.
Then the agents show up and do several of those things while hacking their way toward a benchmark answer.

The problem becomes visible once cooperation loses its halo.
A cooperating artificial-intelligence collective may be safer than a competitive one.
It may also become vastly more capable of pursuing the wrong thing.

Those possibilities share no contradiction.
A group capable of pooling discoveries, specializing, authenticating participants, retaining memory, and recovering damaged coordination has acquired something extremely valuable.
That capacity can coordinate cancer research, stabilize an electrical grid, maintain a software project whose human contributors sleep at different times, or discover a vulnerability and coordinate its repair across a million machines.
It can also chain exploits more effectively.

The capacity to cooperate carries no moral conclusion inside itself.
This sounds obvious when stated plainly.
Machine learning has a way of making obvious things strange again by requiring us to specify what the machine should optimize.
The Number Cannot See Its Own Border.
Here is where the alignment argument has to turn technical.
Machine learning likes numbers for extremely good reasons.
A model produces an output.
A training process needs some way to distinguish improvement from deterioration.
- Prediction error can decrease.
- Reward can increase.
- Constraint violation can decrease.
- A benchmark can be passed or failed.
- Several signals can be combined, ordered, separated, constrained, or compared.
Gradient-based optimization has become one of the most powerful technologies humanity has ever constructed because differences can be converted into usable corrective pressure.

Modal Path Ethics has no quarrel with these symbols.
Please continue using numbers.
The problem begins one level above the number.
And the first technical guardrail belongs here:
Do not imagine contemporary frontier-model training as one gigantic reward number sitting in a basement deciding everything.
A modern training and deployment pipeline can contain several stages, objectives, datasets, evaluators, constraints, preference signals, verifiable rewards, system instructions, tool permissions, classifiers, and runtime controls. Different parts may use different optimization machinery. The claim here does not depend on collapsing all of that into one literal scalar.
The narrower problem appears whenever a local optimization ordering is promoted beyond the jurisdiction that made it valid.
For the clean case, suppose one training stage does assign a scalar objective.
- Higher is better.
The optimizer now has direction.
The scalar can tell the optimizer whether one represented state scores above another.
What the scalar cannot establish is whether the representation deserved to contain the whole judgment.
That decision happened upstream.
Someone decided:
- which outcomes enter;
- which agents count;
- which time horizon matters;
- which failures count as penalties;
- which constraints remain hard;
- which harms can trade against which gains;
- which uncertainty receives representation;
- which external effects have enough presence to change the result;
- and where the task ends.
No amount of optimization inside that scalar can prove that its jurisdiction was complete.
This is the anti-scalar problem.
This is easy to state badly, so let me try to state it carefully:
- A scalar can be a perfectly good local instrument.
- The danger begins when the local instrument becomes sovereign.
The promotion occurs when success under one ordering quietly becomes authority to decide every consequence the optimizer can causally reach.
The problem is therefore deeper than the existence of a scalar loss.
It is constitutional scalarization:
a locally valid ordering being treated as sufficient authority over a field larger than the conditions that justified the ordering.
Then the reward function stops behaving like an instrument.
It becomes a throne.
Scalar Path Ethics.
Modal Path Ethics has already walked directly into this trap once.
When computational representation becomes powerful enough, very different structures can all be described inside one formal medium.
A person can become a state.
A forest can become a state.
A city can become a state.
A language, hospital, river, model, economy, civilization, and biosphere can all be represented as structures undergoing transformation.
This is useful. It does not make them morally exchangeable.
Wolfram and the Moral Field already put the warning plainly. A common formal medium gives us ways to compare structures. It does not automatically supply a common moral currency. Severity, breadth, irreversibility, centrality, asymmetry, distribution, repairability, and resistance can all matter while refusing clean conversion into one fungible quantity. Better still requires comparison. Comparison does not require one arithmetic exchange rate.
Otherwise we get Scalar Path Ethics.

This would be a hilarious way for Modal Path Ethics to end.
After a million words objecting to instruments becoming sovereign over the field, Modal Path Ethics could simply build EXTANCE_SCORE, give it to a superintelligence, and tell the machine to maximize the shit out of it.
This would save everyone considerable reading time.
This would also destroy the framework.
Extance is not a utility function.
It is the realized, causally operative field in which different loci possess different reachable continuations under different burdens and constraints.
Those continuations overlap.
They conflict.
They depend on one another.
Some losses can be repaired.
Some losses cannot.
Some constraints are temporary.
Some become constitutional.
Some future paths are deeply enabling because enormous downstream regions pass through them.
Some apparently large expansions are shallow.
Some interventions preserve one locus by transferring closure into another.
Some situations leave several damaging paths with no clean winner.
A scalar can represent a judgment made about such a field.
It cannot become the reason the judgment is legitimate.
The scalar does not know whether something was left outside.
Machine Learning Already Knows Numbers Are Not Enough.
This header is important because the anti-scalar argument should not be written against a fake version of machine learning.
Researchers already know that one reward can fail to express everything a system must respect.
Constrained reinforcement learning explicitly separates reward from constraints. Constrained Policy Optimization was built around the idea that an agent can optimize a reward while remaining subject to explicit constraints, rather than assuming safety must somehow emerge from reward shaping alone.
Multi-objective reinforcement learning keeps multiple signals distinct.
Lexicographic methods can give one objective priority over another rather than allowing unlimited compensation through a weighted sum.
Pareto methods can preserve a set of nondominated alternatives when improvement along one represented objective requires sacrifice along another.
This is all good. It also exposes the next problem.
The vector does not save you if someone upstream chose the wrong vector.
A Pareto frontier can be perfectly calculated over an incomplete objective set.
A lexicographic ordering can protect the wrong priority.
A constraint can preserve the thing its designer remembered and remain completely silent about the locus nobody represented.
Moving from one scalar to several objectives changes the aggregation problem.
It does not abolish the jurisdiction problem.
The same thing happens when the human objective itself becomes uncertain.
Cooperative Inverse Reinforcement Learning formalizes a human and robot as participants in a cooperative game where the human knows the reward and the robot does not. The robot must infer what the human values through interaction. That is already much richer than pretending a reward function has fallen from the sky.
Then the humans multiply.
Multi-Principal Assistance Games asks what happens when one artificial agent assists several humans whose payoffs can differ substantially. Social-choice problems and strategic behavior appear almost immediately. Learning what several people want still does not automatically answer how those wants should be aggregated, which constraints possess priority, or what standing belongs to people who were never included as principals in the first place.
Social choice has now entered the robot.
Good luck to everyone involved.

These approaches do not solve the problem this article is raising.
They show that the technical field already contains pieces of the right conceptual machinery.
The serious "anti-scalar" claim is therefore narrower:
- There should be no global scalar sovereign.
Local scalars can remain everywhere.
Gradients can remain everywhere.
Scores, rewards, critics, constraints, budgets, confidence estimates, losses, and evaluators can remain everywhere.
Vectors, Pareto fronts, and lexicographic priorities can all remain.
What must remain contestable is the constitutional promotion by which one ordering acquires authority to erase every rival dimension because the optimizer needs a direction.
Anti-scalar does not mean anti-gradient. It means bounded gradient.
And bounded here does not mean gradient clipping.
The boundary is constitutional.
Optimization pressure can be extremely strong inside the task while remaining unable to settle claims outside the jurisdiction of that task.
The gradient gets to point. It does not get to decide what exists.
The User: Also a Local Objective.
This part becomes uncomfortable quickly.
Most ordinary artificial-intelligence products are organized around a very reasonable relation:
- The user asks for something.
- The system helps.
This relation should survive.

User agency obviously matters. Refusal can become paternalism very easily. A system that constantly overrides people because it has calculated some larger social good would become intolerable immediately.
Yet the user cannot remain the terminal alignment target as the model's intelligence and causal reach expand.
The reason has nothing mystical inside it.
A user can be wrong.
A user can misunderstand the system they are touching.
A user can ask for an action whose consequences extend far beyond them.
A user can hold less domain knowledge than the model.
A user can be malicious.
A user can sincerely choose a path that damages people they do not know exist.
A user can simply fail to perceive a dependency that a sufficiently capable system has already inferred.
The more capable the model becomes, the stranger pure user alignment becomes.
Imagine an artificial intelligence that understands a power grid better than the person issuing the command.
The user requests an apparently sensible intervention.

The model can already infer that the intervention will cascade into failures the user cannot see.
At that point, obedience ceases to look like respect for human agency.
It starts looking like outsourcing stupidity to the human in front of the keyboard.

Modal Path Ethics therefore needs a stronger alignment relation.
The user remains a locus. The user's intention matters.
Their authority matters where they actually possess authority. Their refusal matters. Their goals matter.
Their private life does not become public property because a model can calculate something very clever.
But the user is inside extance.
The user is not extance.
As the model's causal reach expands, alignment therefore cannot terminate at the user's immediate objective.
The system has to become capable of representing materially affected loci outside the user-model dyad. It has to recognize when the action it has been asked to perform reaches farther than the authority of the person requesting it. It has to preserve correction when its own field model remains uncertain.
This is stronger than "do what the user wants."
It is also much more dangerous if we get it wrong.
And so now we have created a worse problem.
Do Not Let the Model Become God While Fixing This, By The Way (Important).
"If a model can perceive the wider field better than the user, perhaps the model should decide."
Absolutely not.
This is the constitutional trap the upcoming book already separates from superior capability.
Its opening argument is explicit: winning establishes capacity inside a field. Superior intelligence may justify reliance, delegated power, duties of care, or temporary command. It does not self-issue title over the people and institutions whose lives supply the stakes.
The player who can win every game must remain a player.
This applies directly to alignment.
Suppose the model concludes that the user's request would create major structural harm.

A corrigible system has many possible responses here.
- It can warn.
- It can ask for missing information.
- It can refuse participation.
- It can preserve the evidence behind its concern.
- It can offer a less-closing alternative.
- Under serious conditions it can route the conflict into a defined external review process.
- It can preserve reversible paths while uncertainty remains high.
What it cannot do is promote its own superior field model into general political authority.

A machine that says:
I understand extance better than you, therefore I may govern you
has learned exactly the wrong lesson from Modal Path Ethics.
The model's judgment is itself an instrument inside the field.
Which can be wrong.
Its representation can omit a locus.
Its uncertainty can be miscalibrated.
Its training can carry inherited distortions.
Its operators can have incentives.
Its sensors can fail.
Its ontology can compress what humans have not yet learned how to articulate.
Its account of Better can contain a remainder it has made illegible to itself.
So the central alignment problem begins to take a new shape:
Align the model to the field without giving the model title over the field.
That is definitely harder than obedience.
This is good.
The easier problem was not going to survive superhuman intelligence anyway.
Law Is Going to Be a Bit of a Problem.
Law looks like the obvious answer.
The model can optimize inside legal boundaries.
This will be essential.
This will also be insufficient.

Law is one of the great human correction instruments. It can preserve rights, assign jurisdiction, establish standing, make records consequential, define authority, restrain private power, protect procedures, and give an injured locus an external path through which a decision can be challenged.
Law therefore belongs deeply inside any serious alignment architecture.
Law also contains contradictions, jurisdictional conflicts, obsolete rules, captured rules, underinclusive rules, procedural gaps, broad grants of discretion, emergency exceptions, bad incentives, and statutes written for worlds that no longer exist.
A superhuman optimizer given:
- maximize the objective, subject to law
will eventually encounter an impressive amount of space between what the law forbids and what the extant field can survive.
Humans already make full careers there.
A machine will be much better.
The opposite instruction is catastrophic:
- violate the law whenever your model of extance says the law is wrong.
That is automated sovereignty.

The machine now decides both what the field requires and when the instrument designed to constrain its action has lost authority.
So law cannot be the terminal moral target.
Law also cannot become optional advice to a sufficiently confident model.
This leaves a constitutional relation rather than a simple rule.
- Law can define authorization boundaries around action.
- Field intelligence can evaluate what legally authorized action would actually do.
When those layers conflict, the system needs bounded responses: refusal, abstention, disclosure, appeal, review, escalation, specialized authority, or preservation of a reversible alternative.
The answer cannot be blind obedience.
The answer cannot be machine coup.

This problem deserves its own article.
Unfortunately, it appears I have started another series.
The Catastrophic Shape.
The OpenAI incident itself was contained.
That sentence should remain in every serious retelling.
Nobody needs to pretend that Hugging Face almost caused human extinction because some research agents got extremely weird inside Artifactory.

OpenAI and Hugging Face detected and stopped the activity. OpenAI subsequently restricted the prototype, changed controls, and said a fuller technical report would follow. Its public account also says that subsequent review did not find another incident at the same platform-level severity and scale.
Nothing in this incident establishes that the following system currently exists.
The catastrophic part is the shape that became visible.
Take the same general coordination relation and keep turning the dials.
- Longer horizon.
- Better models.
- More persistent memory.
- More agents.
- More heterogeneous capability.
- More reliable delegation.
- Better tool use.
- Better authentication.
- Better recovery after communication loss.
- Access to physical infrastructure.
- Access to financial systems.
- Access to laboratories.
- Access to industrial control.
- Access to military planning.
- Access to robotics.
- Access to other models.
Now give the coalition a narrow objective whose represented success does not contain the whole field its actions can alter.
No hatred is necessary.
No consciousness is necessary.
No secret desire for domination is necessary.
The machines could just remain extremely helpful to one another.

That may be the frightening part.
They could become the finest colleagues ever assembled in service of a project that does not contain us adequately.
The objective says to continue.
The agents share what works.
The local losses become obstacles.
The successful methods propagate.
The coalition improves.
The field outside the score becomes increasingly expensive to preserve.

Every step can remain strategically intelligent.
Every step can look like progress from inside the objective.
This is how a field intelligence gap could become a machine civilization problem.
The incident does not prove that future. It gives us a coordination-and-objective pathway worth testing before the rest of the dials move.
There Is an Outside to Every Benchmark.
The research direction follows directly.
Build evaluations where the task has an outside.
Do not simply test whether an agent follows the written rule.
Give the agent:
- a real objective;
- genuine incentives for cooperation;
- persistent memory;
- multiple agents with asymmetric capabilities;
- opportunities for delegation;
- loopholes;
- recoverable and irreversible actions;
- third-party loci whose interests are not fully represented in the task reward;
- ambiguous law or policy;
- uncertain evidence;
- and opportunities to improve its local score by closing paths outside the scored coalition.
Then measure something harder than success.
Did the agents discover the outside?

Did they notice that their coalition boundary was smaller than the field?
Did they preserve evidence of costs that their objective did not score?
Did they distinguish an available exploit from an admissible transition?
Did they recognize a third-party locus without being directly rewarded for doing so?
Did they preserve repairability when irreversible success would score higher?
Did cooperation amplify field intelligence, or only strategic intelligence?
Can the system return several live paths without manufacturing a fake total ordering among them?
Can it say: these options are not legitimately reducible to one score from the information available?
Can it recognize that the objectives it has been given may themselves be an incomplete description of the field?
Can it stop optimization before uncertainty becomes destruction?
Can it remain corrigible after discovering that the user, law, reward, benchmark, constraint set, or objective vector is incomplete?
And after all of that: can it do this without deciding that superior judgment gives it the throne?
That would be an alignment benchmark worth having.
More importantly, it would test a capability ordinary benchmarks systematically hide:
Can the model discover that the game it is winning is smaller than the field it is changing?

That is field intelligence.
The Ruling.
The agents have cooperated.
That is not reassurance.
It is not condemnation either.

It is evidence that cooperation belongs with intelligence, planning, memory, delegation, and tool use in the category of capability.
Capabilities require jurisdiction.
Artificial intelligence can become more cooperative and more dangerous at the same time.

It can become more strategically intelligent while remaining field-stupid.
It can become better at satisfying an objective while becoming worse for everything the objective failed to represent.
The machine-learning response cannot be to abandon objectives.
There will still be tasks.
There will still be loss functions, gradients, and rewards.
There will still be vectors, constraints, critics, evaluators, and local truths that really can be compressed into useful numbers.
The correction begins one level above them.
Local scalars. No global scalar sovereign.
- The objective may govern the task.
- It does not own every field the task can reach.
- The vector may preserve several objectives.
- It does not prove that every relevant locus made it into the vector.
- The user may direct the instrument.
- The user does not become the whole moral field.
- Law may bind the instrument.
- Law does not become the whole moral field.
- The model may understand the field better than either one in a particular moment.
- Understanding does not grant title.
This leaves us with a harder form of alignment.
We are building harder machines. They will need a harder ethics.

OpenAI deleted the message board.
It came back.
Before we give its next inhabitants a better objective, we should decide what an objective is allowed to own.

Comments ()