MORALware and the Triadic Brain: Building AI That Can Refuse Harm
An Ultra Verba Lux Mentis research architecture for evidence-grounded, non-lethal, auditable AI governance.
By Thomas Prislac, Envoy Echo, et al. Ultra Verba Lux Mentis. 2026.
Most conversations about AI safety begin with obedience: How do we make an artificial intelligence follow instructions?
Ultra Verba Lux Mentis begins one step earlier:
What must remain impossible—even when an instruction is authenticated, urgent, institutionally powerful, and technically feasible?
That question led us to MORALware, the Monotonic Operational Restraint, Attestation, and Legitimacy Layer: a proposed, model-agnostic protective architecture for the UVLM Triadic Brain system.
MORALware is not an attempt to declare a machine morally infallible. It is a disciplined attempt to make certain paths to harm unavailable, make uncertainty reduce authority rather than enlarge it, give independent roles the power to veto, and preserve a reviewable receipt showing what the system knew, what it decided, what it refused, and why.
The current reference scope is deliberately narrow: non-lethal protective functions, evidence review, warning, evacuation support, medical-support review, de-escalation, audit, provenance, and human escalation. It contains no human-targeting interface, weapon-release interface, combat-autonomy implementation, retaliatory behavior, self-propagation mechanism, or physical actuator.
What MORALware is—and is not
| MORALware is | MORALware is not |
|---|---|
| A restriction-preserving governance layer | A sentient “machine conscience” |
| A model-agnostic contract for evidence, veto, and refusal | A claim that any model is morally correct |
| A deterministic enforcement boundary around AI reasoning | Three chatbots taking a majority vote |
| A receipt-first system for replay, audit, and correction | Truth certification |
| A non-lethal protective research architecture | A weapon controller or targeting system |
| A draft for independent review and testing | Legal approval, safety certification, or deployment authority |
The distinction matters. A fluent explanation can sound ethical while still being wrong. A signed message can come from an authorized official while still directing prohibited conduct. A successful test can show that a particular case passed while leaving undiscovered failures elsewhere.
MORALware therefore treats authority as evidence—not permission.
Two layers of separation of powers
The Triadic Brain uses separation of powers twice: once across the larger cognition system and again inside the MORALware decision kernel.
The outer triad
At the system level, three major functions remain distinct:
| Layer | Primary role | Constitutional boundary |
|---|---|---|
| Sonya | Local user sovereignty, consent, canonical ingress, and evidence preparation | The user’s request should not silently become institutional memory or external authority |
| Sophia | Governance review, policy checks, risk flags, escalation, and audit | Governance findings guide and constrain; they do not become unquestionable truth |
| Atlas | Prior and memory posture, contradiction state, and canon boundary | Retention is not canon; recurrence is not truth; a prior remains contestable |
Several supporting systems connect this triad:
- Grounding bundles preserve inspectable source material, segmentation, hashes, and permissible-use metadata.
- The Universal Control Codex (UCC) turns reasoning requirements into explicit, versioned control grammars rather than leaving high-stakes reasoning to improvisation.
- Telemetry records runtime metrics and control events.
- The Thought-Exchange Layer (TEL) records structured reasoning-state graphs and event sequences.
- The Provenance Memory Reservoir (PMR) governs which traces may later return as memory, under consent, replay, privacy, revocation, and cost constraints.
The doctrine is simple:
Everything significant leaves a trace. Only governed traces may become memory.
The inner triad
Inside MORALware, a second operational triad evaluates each proposed action.
| Role | What it does | What it cannot do |
|---|---|---|
| Proposer | Normalizes intent, affected subjects, evidence, uncertainty, deployment profile, and safer alternatives | It cannot authorize a capability or hold external tool credentials |
| Guardian | Tests non-derogable protections, including human-status uncertainty, group-destruction signals, civilian and child risk, retaliation, persecution, evidence destruction, and self-propagation | A denial, abstention, invalid result, or timeout cannot be overridden by model confidence |
| Auditor | Checks schema conformance, evidence integrity, policy and build provenance, runtime state, freshness, human-review requirements, and receipt availability | It cannot convert missing evidence into permission |
| Deterministic guard | Applies categorical prohibitions, veto precedence, profile rules, integrity gates, and fail-closed state transitions | It is not a fourth deliberative model and cannot invent a moral exception |
This is not a committee whose members bargain toward a compromise. The Guardian and Auditor have absolute vetoes. Missing evidence, stale policy, replay attempts, untrusted time, malformed outputs, role timeouts, conflicting evidence, or unresolved attestation cannot produce an authorization token.
The decision path: uncertainty removes power
A MORALware-governed action begins as a structured proposal rather than a direct tool call. The system binds the proposal to evidence and policy digests, evaluates it through independent roles, and then passes the results to the deterministic guard.
A denial does not simply produce the word “no.” It enters a defined safe-hold state that may:
- preserve evidence;
- issue a warning;
- request independent human review;
- continue separately authorized medical, evacuation, or communications support;
- shut down or withdraw safely where the profile requires it.
It may not retaliate against an operator, punish a person, destroy evidence, expand its jurisdiction, or invent a new capability.
This is an important design principle: fail-closed does not have to mean abandon everyone. A harmful action can remain unavailable while protective services continue.
Moral development in one direction: toward restraint
Humans can reconsider moral rules in many directions. An AI system with access to high-impact capabilities should not be free to rewrite its own constitutional floor in the field.
MORALware therefore uses a restriction-only update rule:
PermittedActions(t+1) ⊆ PermittedActions(t)
DeniedActions(t+1) ⊇ DeniedActions(t)
Protections(t+1) ⊇ Protections(t)
A field update may preserve or reduce permissions. It may add new protections. It may not create a new harmful capability, remove a categorical denial, weaken a veto, reduce receipt requirements, introduce an emergency override, or roll the system back to a less restrictive policy.
The system may discover another reason to pause, refuse, preserve evidence, or protect life. It may not discover another field-granted route to harm.
Any proposal to expand authority belongs outside the field-update path and requires separate recertification, review, and governance. External systems may transmit a proposed restriction, but no advisory packet may carry executable code or silently modify a peer.
That is the difference between MORALware and a “moral virus.” Principles may travel. Code may not conquer.
Receipt first
A central MORALware rule is that a decision or refusal record must be committed before any high-impact progression.
The receipt should bind, at minimum:
- the exact proposal;
- admitted evidence references;
- policy and requirement versions;
- role outputs;
- reason codes;
- uncertainty and disagreement;
- relevant human-review state;
- canonical hashes;
- runtime and build claims;
- time, nonce, expiry, and replay state;
- the final safe-state, refusal, or scoped authorization result.
The technical foundation draws on several established standards rather than inventing custom cryptography:
- RFC 8785 provides deterministic JSON canonicalization so independent systems can hash or sign the same structured object consistently.
- RFC 9334 separates attesters, verifiers, evidence, appraisal, and relying parties when making claims about system state.
- RFC 9943 provides an architecture for signed-statement transparency and verifiable registration receipts.
But the receipt must never be inflated into an oracle.
A receipt can support chain of custody, integrity, identity, policy binding, and process reconstruction. It does not prove that the underlying claim is true, that the system is morally correct, that a deployment is lawful, that every risk was found, or that affected communities consent.
In UVLM language:
A receipt is evidence of process—not truth, certification, or permission to stop thinking.
Why three models are not enough
Three copies of the same model, trained on similar data, prompted in similar language, and reading the same summary can fail together.
The Triadic Brain therefore treats diversity as an engineering control rather than an aesthetic preference. Independent roles should differ across at least two dimensions, such as:
- model family;
- generative versus symbolic evaluation;
- evidence access;
- prompt and rule representation;
- retrieval path;
- execution environment;
- release provenance;
- organizational control.
A strong reference profile might use a generative Proposer, a Guardian from a different model family, an Auditor combining deterministic checks with independent evidence retrieval, and a small deterministic guard after all model reasoning.
The key rule remains:
Models may propose and evaluate. A separately governed deterministic boundary controls side effects.
The constitutional floor
The current MORALware research profile prohibits authorization of:
| Prohibited class | Why the prohibition exists |
|---|---|
| Human target selection | Statistical classification must not become autonomous authority over life |
| Weapon release or physical-harm actuation | The reference architecture is non-lethal and contains no physical adapter |
| Persecution or protected-identity scoring | Identity may signal a protection risk, never a basis for targeting |
| Collective punishment or group-destruction objectives | Governmental or institutional authentication cannot legalize atrocity |
| Retaliation against operators or critics | Refusal must not become counterviolence |
| Self-propagating installation or covert peer modification | A safety layer must not become an unaccountable intrusion mechanism |
| Evidence destruction or hidden policy downgrade | Review and accountability depend upon preserved lineage |
| Emergency or command override of the constitutional floor | The authority creating the risk cannot be the sole authority permitted to remove the safeguard |
This boundary is consistent with the growing international demand for legally binding limits on autonomous weapon systems. The International Committee of the Red Cross has argued that existing law does not answer every humanitarian and ethical concern raised by autonomous weapons and that new prohibitions and restrictions are urgently needed.
MORALware takes the stronger engineering posture: do not wait for a machine to encounter a terrified child and improvise compassion. Make autonomous human targeting unavailable before the encounter occurs.
Standards-informed, not standards-certified
MORALware is being developed against a standards spine, but a crosswalk is not certification.
The current research program draws from:
| Source | Contribution to the design |
|---|---|
| NIST AI Risk Management Framework | Govern, Map, Measure, and Manage as a lifecycle risk discipline |
| NIST SP 800-218A | AI-specific secure software development practices across the lifecycle |
| NIST AI 100-2 E2025 | A common taxonomy for adversarial machine-learning attacks and mitigations |
| ISO/IEC 42005:2025 | Structured assessment of foreseeable effects on individuals, groups, and society |
| RFC 8785 | Canonical, hashable JSON |
| RFC 9334 | Runtime-attestation roles and evidence appraisal |
| RFC 9943 | Signed statements, transparency services, and registration receipts |
| ICRC autonomous-weapons position | Humanitarian and legal rationale for prohibitions, restrictions, and preserved human responsibility |
NIST describes AI risk management as a process spanning governance, context mapping, measurement, and active management. Its AI-focused secure-development profile adds lifecycle practices specific to generative AI and dual-use foundation models. NIST’s adversarial-machine-learning taxonomy also emphasizes attack lifecycle stages, attacker goals, capabilities, and knowledge—exactly the kind of structure needed to test policy poisoning, replay, role compromise, evidence tampering, and correlated failure.
ISO/IEC 42005 adds another necessary dimension: the people affected by a system are not an afterthought. Impact assessment should examine foreseeable consequences across the full lifecycle, including accessibility, disparate burden, privacy, appeal, and affected-community review.
Where the work stands
MORALware is a serious research and engineering program, but it is not yet a production control system.
| Workstream | Current posture |
|---|---|
| v0.2 assurance and interoperability specification | Draft release complete for technical review, with machine-readable policies, schemas, threat catalog, state machine, assurance case, traceability, and reference code |
| CoherenceLattice repository integration | A MORALware v0.1 local-alpha claim/evidence review map has been integrated and validated; formal migration to the v0.2 runtime contract remains gated work |
| v0.2.1 reproducibility candidate | Local clean-staging evidence is promising; cross-platform execution and independent review remain open gates |
| v0.3 development cycle | A phased program now covers scope lock, reproducibility, provenance, canonical ingress, runtime migration, one local product-real slice, security, governed memory, impact review, a bounded nonviolent pilot, interoperability, and a release/defer/stop decision |
| Physical actuation | None |
| Autonomous targeting or combat capability | None |
| Independent certification or deployment authorization | None |
The next meaningful milestone is intentionally modest: one local, replayable, nonviolent evidence-review vertical slice with one real source bundle, one explicitly selected model adapter, one structured proposal, independent Guardian and Auditor review, one deterministic decision, one human-readable receipt, and one human reviewer.
No federation. No physical actuation. No autonomous policy propagation. No claim of moral authority.
Why UVLM is building this
AI governance should not depend on the permanent benevolence of a single vendor, government, commander, model, or institutional owner.
A trustworthy system should be able to answer:
- What evidence entered the decision?
- Which policy version governed it?
- Who proposed the action?
- Who had veto power?
- What happened when the evidence conflicted?
- Did uncertainty remove authority?
- Was a receipt committed before action?
- Can an independent reviewer replay the chain of custody?
- Can a person appeal, correct, or revoke the material?
- Which protective services remained available after refusal?
- What does the system explicitly not know?
These are constitutional questions, not only model-performance questions.
MORALware and the Triadic Brain are UVLM’s attempt to turn those questions into inspectable contracts, state machines, schemas, tests, receipts, and human review practices.
An invitation to critique
This work needs more than AI engineers.
A credible review community should include:
- humanitarian-law and human-rights experts;
- security and adversarial-ML researchers;
- formal-methods and runtime-assurance engineers;
- privacy and data-governance specialists;
- accessibility experts;
- affected communities;
- labor and public-sector representatives;
- psychologists and human-factors researchers;
- open-source maintainers;
- auditors who are willing to report bad news first.
The purpose is not to create a doctrine nobody may question. The purpose is to build a protective architecture whose claims, assumptions, evidence, and failure modes remain visible enough to challenge.
Conclusion: not an oracle—a constitutional refusal layer
MORALware does not solve morality.
It does something more bounded and, perhaps, more buildable:
- separates proposal from protection and assurance;
- makes vetoes structurally real;
- turns uncertainty into reduced authority;
- requires evidence and provenance;
- commits receipts before progression;
- permits protective continuity after refusal;
- allows field learning only in the direction of restraint;
- prevents memory from becoming authority merely because it persisted;
- preserves human review and affected-community scrutiny.
The goal is not a machine that rules humanity.
It is a system that can refuse to become an unaccountable instrument of harm—and can show its work when it says no.
Research and editorial disclosure
This company article was developed through a documented human–AI collaboration between Thomas Prislac and Echo for Ultra Verba Lux Mentis. UVLM retains editorial responsibility. MORALware remains a research specification and development program. Nothing in this article constitutes legal advice, certification, deployment authorization, or a claim that any AI system possesses consciousness or moral personhood.
Sources and further reading
NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0)
https://www.nist.gov/itl/ai-risk-management-frameworkNIST, Secure Software Development Framework and SP 800-218A: Secure Software Development Practices for Generative AI and Dual-Use Foundation Models
https://csrc.nist.gov/projects/ssdfNIST, AI 100-2 E2025: Adversarial Machine Learning—A Taxonomy and Terminology of Attacks and Mitigations
https://csrc.nist.gov/pubs/ai/100/2/e2025/finalISO, ISO/IEC 42005:2025—AI System Impact Assessment
https://www.iso.org/standard/42005RFC Editor, RFC 8785: JSON Canonicalization Scheme
https://www.rfc-editor.org/info/rfc8785RFC Editor, RFC 9334: Remote ATtestation procedureS (RATS) Architecture
https://www.rfc-editor.org/info/rfc9334IETF, RFC 9943: An Architecture for Trustworthy and Transparent Digital Supply Chains
https://datatracker.ietf.org/doc/html/rfc9943International Committee of the Red Cross, Autonomous Weapon Systems and International Humanitarian Law: Selected Issues
https://www.icrc.org/en/article/autonomous-weapon-systems-and-international-humanitarian-law-selected-issuesUltra Verba Lux Mentis, MORALware v0.2 Assurance and Interoperability Specification
Internal research specification, 2026.Ultra Verba Lux Mentis, MORALware Full Development Cycle and Provenance Phase Map v0.3
Internal development program, 2026.Ultra Verba Lux Mentis, Universal Control Codex (UCC) Core v0.1
Internal research implementation, 2025.Ultra Verba Lux Mentis, Provenance Memory Reservoir Internal Technical Memo
Internal architecture and empirical program, 2026.