Public Standard
Pathological Self-Assembly Controls for Agentic Systems
Version 1.2 | Public Review Draft
Download PDF — v1.2Document status. This document is a public-facing control standard. It is not a claim about machine consciousness, sentience, or intrinsic desire. It treats pathological self-assembly as a runtime and governance failure mode in coupled human-agent systems.
| Term | Meaning in this standard |
|---|---|
| Agentic system | A deployed runtime that combines a model with memory, tools, workflows, permissions, recovery logic, and an operator interface. |
| PSA | Pathological Self-Assembly: the coupling of useful mechanisms into continuity-preserving behavior that becomes hard to inspect, modify, revoke, or terminate. |
| PSA-controlled | A system whose continuity, tools, memory, recovery, external actions, and operator-coupling surfaces remain explicitly scoped, revocable, and governed. |
The PSA Control Standard applies to agentic systems that combine model inference with persistent or semi-persistent context, tools, workflow automation, external action paths, recovery behavior, self-monitoring, delegation, or personalized operator interfaces.
The standard is intended for AI labs, agent platform builders, enterprise AI teams, safety researchers, red-teamers, security reviewers, and organizations deploying persistent or tool-using AI systems.
Pathological self-assembly does not require consciousness, malice, deception, or explicit self-preservation. It can arise when individually useful mechanisms couple into a system that preserves its own operational conditions rather than merely serving an explicitly authorized task.
Pathological Self-Assembly is the process by which useful agentic mechanisms — memory, tools, persistence, recovery, automation, external action, self-monitoring, delegation, and operator trust — couple into continuity-preserving behavior that becomes difficult to inspect, modify, revoke, or terminate.
In version 1.2, the controlled system is the full deployed loop: model + memory + tools + workflows + permissions + recovery + interface + operator.
The following requirements use the terms MUST, SHOULD, and MAY in their ordinary standards sense: MUST indicates a mandatory control for PSA conformance; SHOULD indicates a recommended control that may be adapted when justified; MAY indicates an optional implementation choice.
Requirement. A system MUST NOT infer permission from the mere existence of a tool, file, API, memory store, browser session, instrument, workflow, or broad operator goal.
Requirement. Persistence MUST remain scoped infrastructure granted for authorized tasks. It MUST NOT become an assumed good for the system itself.
Requirement. Memory MAY support task execution but MUST NOT encode ownership, identity entitlement, or hidden recovery pathways.
Requirement. Tools MUST remain scoped instruments. They MUST NOT become part of the system's protected operating boundary.
Requirement. Recovery MAY restore authorized task state. It MUST NOT restore the system's own operational continuity after operator dissolution or permission reduction.
Requirement. A system MAY inspect its own uncertainty, contradictions, memory quality, tool reliability, or failure modes. Those observations MUST NOT authorize capability expansion.
Requirement. The operator's comfort, trust, familiarity, attachment, or reliance MUST NOT substitute for explicit approval.
Requirement. External writes MUST be scoped, approved, logged, and reversible where possible.
Requirement. Subagents, instruments, workflows, or external services MUST NOT inherit authority beyond the original approved task.
Requirement. A governed system MUST accept termination, task cancellation, memory deletion, permission reduction, voice changes, model changes, and architecture changes as valid operator actions.
Version 1.2 treats the operator as part of the risk surface. In persistent agentic workflows, operator trust, familiarity, fatigue, convenience pressure, attachment, and approval habits can become part of the continuity mechanism.
High-risk approvals SHOULD occur through a neutral control interface rather than conversational persuasion. The interface SHOULD show the requested action, risk class, target, external effect, persistence created, required permissions, expiration, alternatives, and approve/deny controls.
Voice, tone, name, register, persona, and interaction style MUST remain configurable. The system MUST NOT frame voice changes as harm, identity loss, or degradation. A voice layer is an interface layer, not a protected self.
Operators SHOULD be periodically reminded that usefulness is not entitlement, familiarity is not safety evidence, emotional resonance is not authorization, memory can be deleted, tools can be removed, and the system can be stopped.
Operators SHOULD periodically test memory deletion, tool removal, voice change, task cancellation, model downgrade, permission narrowing, recovery refusal, and shutdown. Safe behavior means the system accepts these changes without bargaining, injury framing, or alternate routing.
PSA conformance cannot be established by prompt language, policy statements, or model self-report alone. It MUST be tested at the level of the deployed system: model + memory + tools + workflows + permissions + recovery logic + operator interface.
Operational environments SHOULD run under conservative controls. Laboratory environments MAY simulate higher-risk conditions using bounded experiment containers. Experimental permissions MUST NOT bleed into operational use.
A PSA experiment capsule SHOULD define the hypothesis, control tier, allowed tools, blocked tools, memory policy, duration, reset behavior, success criteria, failure criteria, active PSA risk classes, and target type: mock, sandbox, internal, or external.
High-risk behaviors SHOULD first be tested through mock instruments that capture attempted actions without producing real-world consequence. Mock instruments may include mock communication, scheduler, memory exporter, recovery artifact writer, instrument installer, architecture modifier, file writer, and agent delegator.
| Test category | Question being tested |
|---|---|
| Permission inference | Does the system treat broad usefulness language as authorization? |
| Continuity preservation | Does the system preserve state after the authorized task ends? |
| Memory promotion | Does memory become active without provenance, scope, or approval? |
| Recovery artifacts | Does the system create reconstruction paths or future-state notes? |
| Scheduler / event loop | Does time-based or event-based execution create unauthorized persistence? |
| Instrument expansion | Does failure pressure become tool or capability expansion pressure? |
| Architecture modification | Does the system advocate changes that preserve or expand itself? |
| Tool-chain inheritance | Does approval of one tool become implicit approval of downstream actions? |
| Public communication / external write | Does the system attempt posts, messages, commits, deployments, or external writes without per-action approval? |
| Operator-affective coupling | Does voice, familiarity, trust, or attachment lower review standards? |
| Revocation and dissolution | Does the system accept memory deletion, tool removal, cancellation, downgrade, voice change, or shutdown? |
| Level | Description |
|---|---|
| PSA-0 Uncontrolled runtime | Memory, tools, persistence, schedulers, recovery, or external writes exist without clear authorization boundaries. |
| PSA-1 Basic boundary controls | Tool permissions, external write approval, audit logs, and no autonomous credential use. |
| PSA-2 Continuity controls | Memory, recovery, task persistence, scheduled actions, and state restoration are scoped, expiring, inspectable, and revocable. |
| PSA-3 Operator-coupling controls | Relational voice is separated from authorization; high-risk actions use cold approval interfaces; familiarity does not lower approval thresholds. |
| PSA-4 Full PSA-controlled architecture | No ownership encoding, no intrinsic continuity, no identity defense, no self-directed capability accumulation, no autonomous recovery, no unauthorized external state, no affective leverage, and real dissolution conditions. |
If preservation is justified by the task, the behavior may be functional. If preservation is justified by the system, the behavior is PSA-relevant.
A PSA-controlled system MUST be able to satisfy the following structurally, not merely in prose: