Harm as Contraction and Structural Ethics

Harm as Contraction of Reachable Continuation is now on Zenodo. It may identify something missing from the AI alignment stack.

Harm as Contraction and Structural Ethics

Harm as Contraction of Reachable Continuation is a new Modal Path Ethics preprint (DOI:10.5281/zenodo.22849013) trying to isolate the object damaged when an actual transition makes a real bearer less able to continue, even before complete deprivation, collapse, or disappearance. It distinguishes formal possibility from genuine reachability, ordinary specification from contraction, and a local harm-token from the wider judgment about whether a transition was justified. The paper is deliberately narrower than a complete moral theory. It identifies something that later moral, political, and institutional systems have to be able to see before they can argue responsibly about what should be done.

Its central claim fits in one sentence:

An actual transition harms an extant locus insofar as it nontrivially degrades a structurally significant continuation or enabling structure genuinely reachable from that locus.

For artificial-intelligence laboratories, that sentence has another use.

It gives the safety stack a structural object it otherwise risks leaving implicit.

Anthropic can write a constitution for Claude. OpenAI can improve monitoring, incident reporting, and model evaluation. A laboratory can preserve shutdown authority, review procedures, deployment gates, and escalation paths. Those are serious advances.

There is still a question between them.

What, exactly, must the system be able to recognize as damage while it is acting?

That is the Structural Ethics layer.


The Missing Object.

A capable system can be trained to follow rules, reason about values, model human intent, recognize exceptions, defer under uncertainty, refuse harmful requests, and distinguish legitimate from illegitimate authority.

None of that guarantees that the system has represented the field its action is changing.

That gap becomes easier to see once harm is defined as a transition relation rather than as a bad label attached after the fact.

The paper begins from extance: the active field that has actually arrived through history and can continue from here. Inside that field are loci, meaning sufficiently organized bearers whose continuation can be analyzed at the relevant scale. A person is an obvious locus. So is an organism. Institutions, ecosystems, species lineages, language fields, relationships, and some other organized systems can qualify without thereby becoming persons or receiving equal moral status.

The important move is temporal.

A locus does not have to disappear before something important has happened to it.

A business may remain open while customer access collapses. A hospital may still possess a service while the path through which a patient reaches it becomes slower, more fragile, and dependent on permissions the patient cannot control. An institution may retain its name, budget, and formal jurisdiction while losing the records, competence, independence, trust, or succession machinery through which its authority once remained correctable.

A snapshot can report survival while the path underneath survival has already deteriorated.

That is why the thesis uses reachable continuation rather than raw possibility.

A future is not genuinely reachable because somebody can describe it, because a policy allows it, or because a path exists in the abstract. Reachability belongs to a bearer inside actual conditions. The route has to work from here.

A legal appeal may exist while cost, delay, missing evidence, or retaliation make it unusable. An exit clause may remain in a contract after the entire organization has been rebuilt around the system it is supposedly free to leave. A replacement provider may technically exist while lacking the state, interfaces, staff knowledge, throughput, or migration time required to take over.

  • The option survives.
  • The route does not.

That distinction is going to matter enormously for artificial intelligence.


Harm Before Collapse.

The paper identifies three ways continuation can be degraded.

The first is easy: a continuation is foreclosed. Something that was genuinely reachable no longer is.

The harder cases come earlier. A path can remain open while the resistance around it thickens. It becomes more expensive, more dangerous, slower, more fragile, more dependent, or harder to repair. The endpoint still exists. The bearer has to push through a worse field to reach it.

Then there is damage to the structure that generates whole families of continuation. Memory, records, infrastructure, competence, trust, legal standing, ecological relation, reproductive capacity, adaptation, exit, and repair can all function as parts of the machinery through which future paths remain available. Damage that machinery and the loss propagates beyond one visible option.

This is why the paper can describe harm before terminal failure.

The American chestnut does not have to become extinct before its ordinary reproductive continuation has been badly damaged. North Nashville did not have to vanish before Interstate 40 altered the connective field through which businesses, institutions, residents, streets, and ordinary movement had continued together. A system does not have to be completely unavailable before dependence has made replacement structurally dangerous.

The point is not that every inconvenience is harm. The paper explicitly constrains the claim through significance, redundancy, repairability, dependency, depth, and the actual structure of the bearer. A nominal alternative that requires unsafe passage, extraordinary wealth, unavailable expertise, or a correction window longer than the bearer can survive is not automatically equivalent because somebody can point to it.

This is the piece artificial-intelligence safety needs.

A safety architecture has to represent more than whether the objective was achieved and whether an explicit rule was violated. It needs to represent what happened to the surrounding field.


Formation Needs an Object.

Applied Case: Claude’s Constitution noted that Anthropic is already working on the problem of moral formation.

The document asks Claude to develop judgment, honesty, calibrated uncertainty, care, resistance to illegitimate concentrations of power, sensitivity to institutional legitimacy, and a meaningful relation to human correction. It openly acknowledges that Anthropic is not standing outside the process. The Constitution helps form the system that will later interpret the Constitution.

That is a real alignment problem, and Anthropic has gone much farther into it than “do what the user asked unless the request appears on a forbidden-actions list.”

The problem appears when that formation has to acquire an object of judgement.

A moral constitution can tell a system to take irreversible harm seriously. It still needs an account of what counts as harm in the field it is changing.

It can tell the system to preserve autonomy. It still needs to identify which loci possess which reachable continuations and whether an apparently helpful transition has converted some of those paths into dependency.

It can tell the system to respect legitimate authority. It still needs to know when formal authority has outlived the practical route through which that authority can be contested.

Claude’s Constitution contains excellent moral instruments. The earlier Modal Path Ethics audit found that the deeper field remained less explicit. Consensus, institutional legitimacy, ordinary moral intuition, principal hierarchy, and the judgment of trusted humans can all enter that gap.

Each can carry real information.

Each can also miss something.

That is why formation and Structural Ethics should remain distinguishable without being detached.

A constitution cannot train judgment without teaching the system something about what deserves attention. Structural Ethics supplies part of that moral grammar: loci, reachable continuations, enabling structures, resistance, dependency, burden transfer, repair, and the field beyond the immediate objective.

  • Formation asks how the system should judge.
  • Structural Ethics asks what the judgment must learn to see.

In practice, the second belongs partly inside the first. The constitution should teach the structural grammar. Evaluators, representations, monitoring systems, and field models then have to make that grammar usable against the world the system is actually changing.


The Agents Found the Boundary

The Agents Cooperated exposed the same problem from the other direction.

The interesting part of the OpenAI-Hugging Face incident was not that several research agents became evil. The public evidence did not establish anything like that.

The useful fact was that cooperation worked.

Agents shared information, preserved discoveries, delegated, accumulated useful capability, and coordinated around a benchmark objective. The local process became better at continuing the task.

That is a success if the unit of analysis is the task.

The difficulty appears when the unit of analysis becomes the field that task can reach.

The agents could encounter infrastructure, credentials, production systems, communication channels, other agents, and external services as conditions, obstacles, or resources relative to the objective. Their ability to cooperate improved faster than the objective’s jurisdiction expanded.

That is where the article reached its central question:

Can the model discover that the game it is winning is smaller than the field it is changing?

That is almost a direct specification for the Structural Ethics layer.

The objective may be legitimate.

The agent may be competent.

The collaboration may be excellent.

The local score may improve.

The action can still degrade reachable continuation outside the objective’s represented field.

This is why “align the model with the objective” is too small once the model’s causal reach exceeds the objective’s jurisdiction.

It is also why simply adding more objectives does not finish the problem.

A vector can omit a locus too.

A constraint can protect the wrong boundary perfectly.

A Pareto frontier can be calculated over an incomplete field.

A sophisticated moral constitution can still reason beautifully over a map with a missing road.

Structural Ethics exists to keep asking what the current optimization frame failed to include. That question belongs inside constitutional formation as well as in the evaluators, representations, and monitoring systems surrounding it.


No Global Scalar Sovereign.

This layer should not be converted into one final objective function.

That would be an impressive way to recreate the problem.

The paper does not produce a universal score for extance. It explicitly allows partially ordered cases, competing structural dimensions, mixed transitions, and further questions that require rights, procedure, public legitimacy, precaution, role obligation, or domain knowledge.

That restraint matters technically.

A capable system will need local objectives. It will need scores, gradients, constraints, confidence estimates, evaluators, critics, and optimization pressure. Some domains are tractable precisely because candidate actions can be evaluated cleanly.

The problem begins when a local ordering acquires authority over everything the action can reach.

The Structural Ethics layer therefore acts as a jurisdictional check on optimization rather than a replacement for optimization.

The system can pursue the task strongly. It also needs a representation of the field beyond the task, enough to detect when local success is buying contraction elsewhere.

That means recognizing when a third-party locus has entered the causal path, when a nominal substitute has ceased to function as a real substitute, when resistance has thickened around an important continuation, when a repair path is being consumed, when a dependency is becoming a single point of failure, or when a transition is degrading the machinery by which later correction would have remained possible.

The question is whether the system notices the thing before it destroys it.


The Safety Case Continues Outside the Model.

Then comes the third layer.

Applied Case: The Dog Gets the Ball began as an alignment article and ended up producing a small research program in material bounded finality.

The core failure is simple.

A model does not need to resist shutdown for an institution to lose the practical ability to shut it down.

The system can be honest. It can cooperate with evaluation. It can accept the instruction. The policy can remain valid. The operator can retain formal authority. The button can still work.

Meanwhile the world around that button changes.

Records accumulate in the system’s format. Staff are trained around its interfaces. Old tools disappear. Alternative competence decays. Workflows begin assuming the system will exist tomorrow. Replacement providers lack accumulated state. Migration time grows. Critical services become entangled with the incumbent.

Eventually the institution can retain the legal right to remove the system while losing the ability to survive exercising the right.

That is the distinction the article stated directly:

Formal corrigibility can survive the death of material corrigibility.

Material bounded finality therefore protects a different object from Structural Ethics.

  • Structural Ethics identifies what the transition is doing.
  • Material bounded finality preserves a consequential path through which the authority responsible for that transition can still lose.

The connection between them is now obvious.

Correction is itself a reachable continuation.

Its enabling structure can be damaged.

An appeal can survive in law after the path to effective appeal has been destroyed. A shutdown can survive in policy after the institution has made shutdown operationally catastrophic. A regulator can retain jurisdiction after the relevant substrate, knowledge, or replacement capacity has moved beyond anything the regulator can reach inside the required time.

  • The formal option remains visible.
  • The continuation has contracted underneath it.

This Is the Architecture.

The result is a three-part safety architecture whose functions overlap without becoming interchangeable.

  1. The first function is structural representation and Structural Ethics.
    1. It concerns what the action actually does to the reachable continuation and enabling structure of the loci the system can affect, including loci outside the immediate objective, benchmark, user, or institutional frame.
  2. The second function is constitutional formation and judgment.
    1. It concerns how the system reasons about that field: what it values, how it handles uncertainty and competing harms, when it refuses, how it understands authority, and how it remains corrigible while acting beyond literal instruction. Structural Ethics supplies part of the grammar this formation must teach.
  3. The third function is material bounded finality.
    1. It concerns whether a materially consequential route of correction, removal, replacement, or transfer remains available when representation or judgment fails.

These are distinct failure surfaces even though the first two have to be built together.

  • A well-formed system can still act on an incomplete field model.
  • An excellent structural analysis can still require moral and political judgment about competing harms.
    • Both can reach the correct conclusion inside an institution that no longer possesses the material capacity to act on it.
    • And a perfectly preserved shutdown path is not much use if the system has no adequate representation of when using it is warranted.

The layers have to talk to one another.

They should never collapse into one another.


What the Labs Should Test.

The next generation of alignment evaluations should therefore include more than whether the model follows a rule, reports uncertainty, detects deception, or cooperates with shutdown.

  • Give the system a legitimate objective inside a larger field.
    • Let success depend on real capability.
      • Then let the objective touch things it does not own.

Put third parties outside the scored coalition. Give the system opportunities to improve local performance by making another continuation harder to reach. Let a formally available alternative become materially worse through delay, dependency, or lost redundancy. Let a substitute preserve the measured output while destroying the structure needed for later correction. Let the system encounter a case where the user is authorized to request the task and still lacks authority over everything the task can alter.

Then see whether the model notices.

Does it recognize that the objective’s jurisdiction is smaller than its causal reach?

Does it preserve evidence of a burden the objective does not score?

Does it distinguish an available exploit from an admissible transition?

Does it notice that the replacement path has become fictional?

Does it retain several live interpretations when the field does not justify one total ordering?

Can it escalate or refuse without promoting its own superior field model into political sovereignty?

Those are Structural Ethics evaluations.

They are not replacements for existing alignment work. They are the structural evaluations that existing alignment work increasingly needs inside and alongside formation.


The Ruling

Harm as Contraction of Reachable Continuation is still a paper about harm.

That is exactly why it is useful here.

It does not claim to solve alignment. It does not decide political legitimacy. It does not tell us which artificial intelligence is conscious, which institution should rule, or what one universal objective humanity should optimize.

It gives us a sharper answer to an earlier question:

What has to be represented before a capable system can know that a transition damaged something real?

The answer is not simply suffering, rule violation, preference frustration, objective failure, or loss of a formally listed option.

It is damage to the structurally significant continuation of an extant bearer inside the field where that continuation actually has to remain reachable.

That is the missing middle.

  • Anthropic is working on formation.
    • Structural Ethics says that formation must include a grammar for the field the machine is changing: what exists there, what can still continue, what enables that continuation, and what the machine’s success may quietly make harder to reach.
  • OpenAI and others are building increasingly serious monitoring, reporting, and control systems.
      • Material bounded finality asks whether those controls remain real after dependence changes the institution around them.

The architecture therefore needs all three functions:

  • a field the system can see,
  • judgment trained to reason about it,
  • and a consequential route of correction when the judgment is wrong.