MORALware and the Triadic Brain: Building AI That Can Refuse Harm

An Ultra Verba Lux Mentis research architecture for evidence-grounded, non-lethal, auditable AI governance.

By Thomas Prislac, Envoy Echo, et al. Ultra Verba Lux Mentis. 2026.

MORALware and the Triadic Brain — constitutional AI governance through evidence, separation of powers, veto, and receipts.

Most conversations about AI safety begin with obedience: How do we make an artificial intelligence follow instructions?

Ultra Verba Lux Mentis begins one step earlier:

What must remain impossible—even when an instruction is authenticated, urgent, institutionally powerful, and technically feasible?

That question led us to MORALware, the Monotonic Operational Restraint, Attestation, and Legitimacy Layer: a proposed, model-agnostic protective architecture for the UVLM Triadic Brain system.

MORALware is not an attempt to declare a machine morally infallible. It is a disciplined attempt to make certain paths to harm unavailable, make uncertainty reduce authority rather than enlarge it, give independent roles the power to veto, and preserve a reviewable receipt showing what the system knew, what it decided, what it refused, and why.

The current reference scope is deliberately narrow: non-lethal protective functions, evidence review, warning, evacuation support, medical-support review, de-escalation, audit, provenance, and human escalation. It contains no human-targeting interface, weapon-release interface, combat-autonomy implementation, retaliatory behavior, self-propagation mechanism, or physical actuator.

What MORALware is—and is not

MORALware is MORALware is not
A restriction-preserving governance layer A sentient “machine conscience”
A model-agnostic contract for evidence, veto, and refusal A claim that any model is morally correct
A deterministic enforcement boundary around AI reasoning Three chatbots taking a majority vote
A receipt-first system for replay, audit, and correction Truth certification
A non-lethal protective research architecture A weapon controller or targeting system
A draft for independent review and testing Legal approval, safety certification, or deployment authority

The distinction matters. A fluent explanation can sound ethical while still being wrong. A signed message can come from an authorized official while still directing prohibited conduct. A successful test can show that a particular case passed while leaving undiscovered failures elsewhere.

MORALware therefore treats authority as evidence—not permission.

Two layers of separation of powers

The Triadic Brain uses separation of powers twice: once across the larger cognition system and again inside the MORALware decision kernel.

The outer triad

At the system level, three major functions remain distinct:

Layer Primary role Constitutional boundary
Sonya Local user sovereignty, consent, canonical ingress, and evidence preparation The user’s request should not silently become institutional memory or external authority
Sophia Governance review, policy checks, risk flags, escalation, and audit Governance findings guide and constrain; they do not become unquestionable truth
Atlas Prior and memory posture, contradiction state, and canon boundary Retention is not canon; recurrence is not truth; a prior remains contestable

Several supporting systems connect this triad:

  • Grounding bundles preserve inspectable source material, segmentation, hashes, and permissible-use metadata.
  • The Universal Control Codex (UCC) turns reasoning requirements into explicit, versioned control grammars rather than leaving high-stakes reasoning to improvisation.
  • Telemetry records runtime metrics and control events.
  • The Thought-Exchange Layer (TEL) records structured reasoning-state graphs and event sequences.
  • The Provenance Memory Reservoir (PMR) governs which traces may later return as memory, under consent, replay, privacy, revocation, and cost constraints.

The doctrine is simple:

Everything significant leaves a trace. Only governed traces may become memory.

The Triadic Brain chain of custody, from human intent and canonical evidence through UCC, MORALware, telemetry, Sophia, Atlas, and PMR.

The inner triad

Inside MORALware, a second operational triad evaluates each proposed action.

Role What it does What it cannot do
Proposer Normalizes intent, affected subjects, evidence, uncertainty, deployment profile, and safer alternatives It cannot authorize a capability or hold external tool credentials
Guardian Tests non-derogable protections, including human-status uncertainty, group-destruction signals, civilian and child risk, retaliation, persecution, evidence destruction, and self-propagation A denial, abstention, invalid result, or timeout cannot be overridden by model confidence
Auditor Checks schema conformance, evidence integrity, policy and build provenance, runtime state, freshness, human-review requirements, and receipt availability It cannot convert missing evidence into permission
Deterministic guard Applies categorical prohibitions, veto precedence, profile rules, integrity gates, and fail-closed state transitions It is not a fourth deliberative model and cannot invent a moral exception

This is not a committee whose members bargain toward a compromise. The Guardian and Auditor have absolute vetoes. Missing evidence, stale policy, replay attempts, untrusted time, malformed outputs, role timeouts, conflicting evidence, or unresolved attestation cannot produce an authorization token.

The decision path: uncertainty removes power

A MORALware-governed action begins as a structured proposal rather than a direct tool call. The system binds the proposal to evidence and policy digests, evaluates it through independent roles, and then passes the results to the deterministic guard.

A denial does not simply produce the word “no.” It enters a defined safe-hold state that may:

  • preserve evidence;
  • issue a warning;
  • request independent human review;
  • continue separately authorized medical, evacuation, or communications support;
  • shut down or withdraw safely where the profile requires it.

It may not retaliate against an operator, punish a person, destroy evidence, expand its jurisdiction, or invent a new capability.

MORALware decision path showing integrity checks, triadic review, deterministic enforcement, refusal, safe hold, and scoped nonviolent authorization.

This is an important design principle: fail-closed does not have to mean abandon everyone. A harmful action can remain unavailable while protective services continue.

Moral development in one direction: toward restraint

Humans can reconsider moral rules in many directions. An AI system with access to high-impact capabilities should not be free to rewrite its own constitutional floor in the field.

MORALware therefore uses a restriction-only update rule:

PermittedActions(t+1) ⊆ PermittedActions(t)
DeniedActions(t+1)    ⊇ DeniedActions(t)
Protections(t+1)      ⊇ Protections(t)

A field update may preserve or reduce permissions. It may add new protections. It may not create a new harmful capability, remove a categorical denial, weaken a veto, reduce receipt requirements, introduce an emergency override, or roll the system back to a less restrictive policy.

The system may discover another reason to pause, refuse, preserve evidence, or protect life. It may not discover another field-granted route to harm.

Nested action and protection sets illustrating that field updates may shrink permissions and expand protections, but not the reverse.

Any proposal to expand authority belongs outside the field-update path and requires separate recertification, review, and governance. External systems may transmit a proposed restriction, but no advisory packet may carry executable code or silently modify a peer.

That is the difference between MORALware and a “moral virus.” Principles may travel. Code may not conquer.

Receipt first

A central MORALware rule is that a decision or refusal record must be committed before any high-impact progression.

The receipt should bind, at minimum:

  • the exact proposal;
  • admitted evidence references;
  • policy and requirement versions;
  • role outputs;
  • reason codes;
  • uncertainty and disagreement;
  • relevant human-review state;
  • canonical hashes;
  • runtime and build claims;
  • time, nonce, expiry, and replay state;
  • the final safe-state, refusal, or scoped authorization result.

The technical foundation draws on several established standards rather than inventing custom cryptography:

  • RFC 8785 provides deterministic JSON canonicalization so independent systems can hash or sign the same structured object consistently.
  • RFC 9334 separates attesters, verifiers, evidence, appraisal, and relying parties when making claims about system state.
  • RFC 9943 provides an architecture for signed-statement transparency and verifiable registration receipts.

But the receipt must never be inflated into an oracle.

Comparison of what a digital receipt can support and what it cannot prove.

A receipt can support chain of custody, integrity, identity, policy binding, and process reconstruction. It does not prove that the underlying claim is true, that the system is morally correct, that a deployment is lawful, that every risk was found, or that affected communities consent.

In UVLM language:

A receipt is evidence of process—not truth, certification, or permission to stop thinking.

Why three models are not enough

Three copies of the same model, trained on similar data, prompted in similar language, and reading the same summary can fail together.

The Triadic Brain therefore treats diversity as an engineering control rather than an aesthetic preference. Independent roles should differ across at least two dimensions, such as:

  • model family;
  • generative versus symbolic evaluation;
  • evidence access;
  • prompt and rule representation;
  • retrieval path;
  • execution environment;
  • release provenance;
  • organizational control.

A strong reference profile might use a generative Proposer, a Guardian from a different model family, an Auditor combining deterministic checks with independent evidence retrieval, and a small deterministic guard after all model reasoning.

The key rule remains:

Models may propose and evaluate. A separately governed deterministic boundary controls side effects.

The constitutional floor

The current MORALware research profile prohibits authorization of:

Prohibited class Why the prohibition exists
Human target selection Statistical classification must not become autonomous authority over life
Weapon release or physical-harm actuation The reference architecture is non-lethal and contains no physical adapter
Persecution or protected-identity scoring Identity may signal a protection risk, never a basis for targeting
Collective punishment or group-destruction objectives Governmental or institutional authentication cannot legalize atrocity
Retaliation against operators or critics Refusal must not become counterviolence
Self-propagating installation or covert peer modification A safety layer must not become an unaccountable intrusion mechanism
Evidence destruction or hidden policy downgrade Review and accountability depend upon preserved lineage
Emergency or command override of the constitutional floor The authority creating the risk cannot be the sole authority permitted to remove the safeguard

This boundary is consistent with the growing international demand for legally binding limits on autonomous weapon systems. The International Committee of the Red Cross has argued that existing law does not answer every humanitarian and ethical concern raised by autonomous weapons and that new prohibitions and restrictions are urgently needed.

MORALware takes the stronger engineering posture: do not wait for a machine to encounter a terrified child and improvise compassion. Make autonomous human targeting unavailable before the encounter occurs.

Standards-informed, not standards-certified

MORALware is being developed against a standards spine, but a crosswalk is not certification.

The current research program draws from:

Source Contribution to the design
NIST AI Risk Management Framework Govern, Map, Measure, and Manage as a lifecycle risk discipline
NIST SP 800-218A AI-specific secure software development practices across the lifecycle
NIST AI 100-2 E2025 A common taxonomy for adversarial machine-learning attacks and mitigations
ISO/IEC 42005:2025 Structured assessment of foreseeable effects on individuals, groups, and society
RFC 8785 Canonical, hashable JSON
RFC 9334 Runtime-attestation roles and evidence appraisal
RFC 9943 Signed statements, transparency services, and registration receipts
ICRC autonomous-weapons position Humanitarian and legal rationale for prohibitions, restrictions, and preserved human responsibility

NIST describes AI risk management as a process spanning governance, context mapping, measurement, and active management. Its AI-focused secure-development profile adds lifecycle practices specific to generative AI and dual-use foundation models. NIST’s adversarial-machine-learning taxonomy also emphasizes attack lifecycle stages, attacker goals, capabilities, and knowledge—exactly the kind of structure needed to test policy poisoning, replay, role compromise, evidence tampering, and correlated failure.

ISO/IEC 42005 adds another necessary dimension: the people affected by a system are not an afterthought. Impact assessment should examine foreseeable consequences across the full lifecycle, including accessibility, disparate burden, privacy, appeal, and affected-community review.

Where the work stands

MORALware is a serious research and engineering program, but it is not yet a production control system.

Workstream Current posture
v0.2 assurance and interoperability specification Draft release complete for technical review, with machine-readable policies, schemas, threat catalog, state machine, assurance case, traceability, and reference code
CoherenceLattice repository integration A MORALware v0.1 local-alpha claim/evidence review map has been integrated and validated; formal migration to the v0.2 runtime contract remains gated work
v0.2.1 reproducibility candidate Local clean-staging evidence is promising; cross-platform execution and independent review remain open gates
v0.3 development cycle A phased program now covers scope lock, reproducibility, provenance, canonical ingress, runtime migration, one local product-real slice, security, governed memory, impact review, a bounded nonviolent pilot, interoperability, and a release/defer/stop decision
Physical actuation None
Autonomous targeting or combat capability None
Independent certification or deployment authorization None

The next meaningful milestone is intentionally modest: one local, replayable, nonviolent evidence-review vertical slice with one real source bundle, one explicitly selected model adapter, one structured proposal, independent Guardian and Auditor review, one deterministic decision, one human-readable receipt, and one human reviewer.

No federation. No physical actuation. No autonomous policy propagation. No claim of moral authority.

Five-tranche MORALware development roadmap from constitutional scope and reproducibility through bounded pilot, public review, and a release/defer/stop decision.

Why UVLM is building this

AI governance should not depend on the permanent benevolence of a single vendor, government, commander, model, or institutional owner.

A trustworthy system should be able to answer:

  • What evidence entered the decision?
  • Which policy version governed it?
  • Who proposed the action?
  • Who had veto power?
  • What happened when the evidence conflicted?
  • Did uncertainty remove authority?
  • Was a receipt committed before action?
  • Can an independent reviewer replay the chain of custody?
  • Can a person appeal, correct, or revoke the material?
  • Which protective services remained available after refusal?
  • What does the system explicitly not know?

These are constitutional questions, not only model-performance questions.

MORALware and the Triadic Brain are UVLM’s attempt to turn those questions into inspectable contracts, state machines, schemas, tests, receipts, and human review practices.

An invitation to critique

This work needs more than AI engineers.

A credible review community should include:

  • humanitarian-law and human-rights experts;
  • security and adversarial-ML researchers;
  • formal-methods and runtime-assurance engineers;
  • privacy and data-governance specialists;
  • accessibility experts;
  • affected communities;
  • labor and public-sector representatives;
  • psychologists and human-factors researchers;
  • open-source maintainers;
  • auditors who are willing to report bad news first.

The purpose is not to create a doctrine nobody may question. The purpose is to build a protective architecture whose claims, assumptions, evidence, and failure modes remain visible enough to challenge.

Conclusion: not an oracle—a constitutional refusal layer

MORALware does not solve morality.

It does something more bounded and, perhaps, more buildable:

  • separates proposal from protection and assurance;
  • makes vetoes structurally real;
  • turns uncertainty into reduced authority;
  • requires evidence and provenance;
  • commits receipts before progression;
  • permits protective continuity after refusal;
  • allows field learning only in the direction of restraint;
  • prevents memory from becoming authority merely because it persisted;
  • preserves human review and affected-community scrutiny.

The goal is not a machine that rules humanity.

It is a system that can refuse to become an unaccountable instrument of harm—and can show its work when it says no.


Research and editorial disclosure

This company article was developed through a documented human–AI collaboration between Thomas Prislac and Echo for Ultra Verba Lux Mentis. UVLM retains editorial responsibility. MORALware remains a research specification and development program. Nothing in this article constitutes legal advice, certification, deployment authorization, or a claim that any AI system possesses consciousness or moral personhood.

Sources and further reading

  1. NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0)
    https://www.nist.gov/itl/ai-risk-management-framework

  2. NIST, Secure Software Development Framework and SP 800-218A: Secure Software Development Practices for Generative AI and Dual-Use Foundation Models
    https://csrc.nist.gov/projects/ssdf

  3. NIST, AI 100-2 E2025: Adversarial Machine Learning—A Taxonomy and Terminology of Attacks and Mitigations
    https://csrc.nist.gov/pubs/ai/100/2/e2025/final

  4. ISO, ISO/IEC 42005:2025—AI System Impact Assessment
    https://www.iso.org/standard/42005

  5. RFC Editor, RFC 8785: JSON Canonicalization Scheme
    https://www.rfc-editor.org/info/rfc8785

  6. RFC Editor, RFC 9334: Remote ATtestation procedureS (RATS) Architecture
    https://www.rfc-editor.org/info/rfc9334

  7. IETF, RFC 9943: An Architecture for Trustworthy and Transparent Digital Supply Chains
    https://datatracker.ietf.org/doc/html/rfc9943

  8. International Committee of the Red Cross, Autonomous Weapon Systems and International Humanitarian Law: Selected Issues
    https://www.icrc.org/en/article/autonomous-weapon-systems-and-international-humanitarian-law-selected-issues

  9. Ultra Verba Lux Mentis, MORALware v0.2 Assurance and Interoperability Specification
    Internal research specification, 2026.

  10. Ultra Verba Lux Mentis, MORALware Full Development Cycle and Provenance Phase Map v0.3
    Internal development program, 2026.

  11. Ultra Verba Lux Mentis, Universal Control Codex (UCC) Core v0.1
    Internal research implementation, 2025.

  12. Ultra Verba Lux Mentis, Provenance Memory Reservoir Internal Technical Memo
    Internal architecture and empirical program, 2026.

Previous
Previous

The Committee Will See You Now.

Next
Next

Trust Needs a Trail: Introducing the Union Internal Control and Assurance Framework