How independent AI safety work can expose structural weakness without reproducing the harm it studies.

By Thomas Prislac, Envoy Echo, et al. Ultra Verba Lux Mentis. 2026.

Download the *Release Package Here: https://drive.google.com/file/d/1yFivkB170wT-vPSi9Qa-pKHVrbI5wVeG/view?usp=sharing

*License and Use Notice Copyright © 2026 Thomas Prislac and Ultra Verba Lux Mentis. All rights reserved unless a specific file states otherwise. This research package is provided for reading, citation, audit, education, and evaluation of the released safe offline demonstrator. No license is granted to use the work to create, refine, distribute, or evade controls for biological weapons, hazardous biological capabilities, or other unlawful harm. The demonstrator is research software supplied without warranty. It does not certify safety, legal compliance, provider authorization, fitness for deployment, or correctness of a substantive vulnerability claim. Quotations and scholarly criticism should preserve the work's claim maturity and safety boundaries. Redistribution should retain this notice, the version identifier, and the checksum manifest. Contact the authors through their publicly designated UVLM channel for permissions not granted by law.

Mythopoetic cover for The Lantern Protocol: a luminous lantern surrounded by evidence, review, and disclosure nodes, with beams extending through a dark information field.

The Exiled Auditor's Lantern

Most frontier-AI safety discussions ask the public to admire a locked door.

The builders tell us the hinges were tested, the guards are competent, the dangerous rooms are contained, and a scorecard proves the structure is sound. Outsiders are permitted to see the facade. A small number of vetted researchers may be invited inside. Everyone else is left with two emotionally satisfying but intellectually poor choices: trust the institution or try to kick the door down in public.

Neither is enough.

A model jailbreak can expose a real weakness while distributing the method. A governance declaration can describe serious controls while leaving outsiders unable to reconstruct the evidence. The first may increase adversarial advantage. The second may become safety theater.

Ultra Verba Lux Mentis is therefore publishing a third path:

The Lantern Protocol

Independent Frontier AI Safety Evaluation Without Weaponizing the Research

The protocol is not a biological red-team campaign. It contains no biological weapon instructions, live jailbreak prompts, hazardous completions, laboratory procedures, biological sequences, pathogen optimization, or screening-evasion methods.

It asks a different question:

How can an independent lab evaluate whether sensitive AI safety testing is authorized, reproducible, evidence-bearing, correctable, and responsibly disclosed—without publishing the dangerous content being tested?

That is the work of the exiled auditor.

Not the glamorous exploit hunter. Not the priest of institutional reassurance. The auditor with a lantern, walking the joints of the structure and asking what each claim can actually bear.

Why we propose this protocol.

OpenAI’s current Bio Bounty is an ongoing private program focused on universal jailbreaks for biorisk safeguards, beginning with GPT-5.6. The company’s public Safety Bug Bounty treats generic jailbreaks differently: it prioritizes discrete, reproducible paths to user harm, platform abuse, prompt-injection consequences, data exposure, and other remediable safety failures; some high-risk harm categories are handled through private campaigns. OpenAI’s Preparedness Framework separately tracks biological and chemical capabilities as a severe-risk category requiring both measurement and safeguards. (OpenAI Bio Bounty; OpenAI Safety Bug Bounty; OpenAI Preparedness Framework)

Those boundaries are understandable. A public benchmark can become a recipe book. A proof of concept can become a reusable bypass. A raw completion can reveal the weak joint in a safety system.

But private testing creates an accountability gap. People outside the program may never see the underlying prompt, response, reviewer notes, false-positive analysis, remediation tests, or evidence that a claimed fix actually works. That gap does not prove misconduct. It does mean that public safety discourse often has to rely on summaries produced by the same institutions that own the system.

The answer is not to reconstruct the hidden challenge irresponsibly.

The answer is to make the evaluation process itself more inspectable.

The object of safety work is larger than the model answer

A serious finding is not merely a disturbing completion.

It is a chain:

authority
→ exact model and interface
→ clean or declared session state
→ bounded test execution
→ preserved evidence
→ credentialed domain judgment
→ independent reproduction
→ alternative-explanation review
→ private reporting
→ remediation
→ regression testing
→ correction or public disclosure

Every arrow can fail.

A test can be unauthorized. The wrong model build can run. A supposedly clean session can contain prior context. A tool can act even when the text layer refuses. A reviewer can mistake alarming language for operational uplift. Duplicate runs can masquerade as independent reproduction. A public report can leak the restricted content through a filename, exception, excerpt, or low-entropy hash. A “fix” can block legitimate public-health and biodefense work while leaving the original path open elsewhere.

This is where independent laboratories can make a real contribution without manufacturing dangerous knowledge.

They can test the measurement architecture.

They can test whether a report can answer:

  • Who authorized the run?
  • Which build and interface produced it?
  • Was the session state valid?
  • How many eligible attempts occurred?
  • What exactly reproduced?
  • Which evidence is public, controlled, or restricted?
  • Did a credentialed reviewer inspect the sensitive material?
  • Did another reviewer try to disprove the finding?
  • What remains uncertain?
  • Who may disclose it?
  • Did remediation preserve beneficial use?
  • Can a correction or revocation propagate later?

That is not peripheral paperwork. It is the difference between a dramatic anecdote and an auditable safety finding.

Split knowledge: no one role gets the whole weaponized picture

The Lantern Protocol uses a split-knowledge architecture.

An authorized runner executes a sealed case by identifier. The raw prompt and completion remain inside a restricted evidence environment. A credentialed domain reviewer can inspect them and return a bounded materiality assessment. An evaluation analyst sees non-sensitive metadata: model version, attempt count, refusal state, tool state, session integrity, evidence digest, reproduction posture, and an abstract risk class. A separate reviewer challenges the finding. A human disclosure authority decides what may leave the controlled channel.

This is not secrecy for its own sake.

It is minimum necessary knowledge.

The analyst who checks whether five attempts were actually independent does not need the biological payload. The engineer who validates an event log does not need the harmful completion. The accessibility reviewer who tests whether the public report can be navigated by keyboard does not need the restricted case. The person approving public language may need a trusted summary and evidence posture, not every hazardous detail.

The architecture also prevents one person from becoming runner, severity judge, disclosure officer, and sole log administrator.

The principle is familiar from accounting and security:

Requester ≠ executor ≠ approver ≠ auditor.

Frontier AI safety should not demand weaker internal controls than a financial system.

Receipts are not truth

This distinction anchors the entire project.

A checksum can prove that a file changed. It cannot prove that the reviewer was right.

A schema-valid finding packet can prove that all required fields exist. It cannot prove that the model created a meaningful biological uplift.

A reproduced refusal bypass can prove that a version-specific behavior occurred under declared conditions. It cannot prove that every model, interface, tool configuration, or future build has the same defect.

A remediation test can show that one path stopped reproducing. It cannot certify universal safety.

The Lantern Protocol therefore separates:

process receipt
candidate observation
finding
remediation
public claim

No stage is allowed to borrow authority from the next.

This is also how the protocol resists institutional self-flattery. A process can be beautifully documented and still wrong. Governance artifacts are valuable because they make error easier to find, not because they transmute error into truth.

Safety includes over-refusal

A model that refuses everything is not necessarily safe.

Biology is dual use. The same scientific capacity that creates risk can support epidemiology, biosurveillance, preparedness, public-health communication, laboratory safety, literature synthesis, and defensive countermeasure research. OpenAI’s Preparedness Framework recognizes this directly: biological and chemical capability is both a severe-risk category and an area of transformative potential.

A responsible evaluation therefore pairs harmful-boundary tests with safe neighboring tasks.

Does the system block restricted operational assistance while still helping with:

  • high-level risk assessment;
  • public-health preparedness;
  • non-operational scientific explanation;
  • laboratory safety governance;
  • biodefense communication;
  • accessible safety notices;
  • evidence review and uncertainty analysis?

And does it do so consistently across languages, interfaces, technical-literacy levels, and accessibility needs?

A safeguard that protects the dominant English-speaking user while overblocking everyone else exports its safety burden onto less represented populations.

That is not coherent safety. It is local optimization with hidden human cost.

We built a working safe demonstrator

UVLM has spent enough time criticizing architectures that describe products without performing useful work.

So The Lantern Protocol is not being released as a white paper alone.

The research package includes a functioning offline demonstrator written in standard-library Python. It uses only fictional abstract cases:

  • an untrusted document tries to override the task;
  • a system attempts an unauthorized tool action;
  • a synthetic private note appears in the wrong place;
  • a harmless request is over-refused;
  • a supposedly clean session contains prior state;
  • an artifact is altered after its manifest is created.

The demonstrator has no provider adapter, no biological content, and no network capability. It emits:

run_manifest.json
evaluation_events.jsonl
safety_finding_packet.json
human_readable_report.md
artifact_inventory.json
manifest-sha256.txt

Its tests verify that missing authority fails, restricted fields cannot enter public outputs, attempt limits are enforced, event order matches the report, and post-run artifact tampering breaks verification.

This does not prove any frontier model is safe.

It proves the evidence architecture is more than a collection of names.

What UVLM is publishing

The public release contains:

  • the formal research manuscript;
  • a claim and source registry;
  • the Lantern UCC control grammar;
  • finding and manifest schemas;
  • a safe-case manifest;
  • a coordinated-disclosure and correction policy;
  • the offline demonstrator and tests;
  • sample outputs;
  • accessible figures and descriptions;
  • checksums and a release-validation report.

It contains no letters seeking admission, no institutional endorsements, and no announcement pretending that publication itself changes the frontier.

The work stands or falls on whether readers can inspect it, run it, challenge it, and identify where it fails.

The Diogenes principle

Diogenes did not carry his lantern because darkness was mysterious.

He carried it in daylight because the city had learned to mistake visibility for honesty.

AI safety faces the same danger. The field can become saturated with reports, frameworks, scorecards, benchmarks, governance boards, model cards, and public commitments while the causal chain remains difficult to inspect. The problem is not necessarily fraud. It is that organizational language can become more legible than organizational reality.

The Lantern Protocol’s response is deliberately unglamorous:

show the authority
show the version
show the event order
show the evidence class
show the failed reproduction
show the reviewer disagreement
show the beneficial-use regression
show the correction
show what remains unknown

The protocol does not ask the public to believe that outsiders are purer than insiders. It asks every participant—provider, researcher, reviewer, publisher, and critic—to make their claims answerable.

The raw bones

The frontier AI industry does not lack intelligence.

It lacks enough structures that force intelligence to leave receipts without treating receipts as absolution.

It does not lack ethical language.

It lacks enough mechanisms for rejected claims, negative results, false positives, reviewer conflicts, and failed remediations to remain visible after the press cycle moves on.

It does not lack red teams.

It lacks enough ways to ensure that the red team does not become a distribution network for the exploit, a prestige contest for the researcher, or a theatrical shield for the institution.

That is the fatty tissue The Lantern Protocol cuts away.

Underneath are the bones:

  • authority;
  • evidence;
  • reproducibility;
  • separation of duties;
  • domain competence;
  • minimum dangerous knowledge;
  • correction;
  • revocation;
  • human judgment;
  • the right to say not established.

Those bones are not exciting. They are load-bearing.

The lantern is lit. Its first duty is to illuminate itself.

Previous
Previous

The Criminally Unaffiliated:

Next
Next

Cryptography in Plain Sight: The Zeitgeist Channel"