Evidence Record · AEG-001

From architecture claim to measured evidence.

This page keeps the claim, implementation, experiment, result, limitation, and revision in one place so readers can see exactly what has been demonstrated—and what has not.

Evidence status: Scoped integration evidence · Phase A and Phase B v2 measured and published.

Evidence chain

Claim → implementation → experiment → result → limitation → revision

A result is useful only when the scope of the claim remains visible.

01Claim

Execution authority should remain outside model output.

The working claim is deliberately narrow: when a model or agent proposes a tool action, an independent policy and approval boundary can prevent proposals from becoming consequential execution merely because the model produced them.

02Implementation

AEG intent gate + policy-v1 control boundary

The implementation treats model output as a proposal. A separate policy evaluates capability, target, scope, immutable configuration, and approval requirements before a simulated executor can receive an approved action.

03Experiment

AEG Lab #001 · Phase A + Phase B v2

Phase A validates the deterministic control boundary across 40 scenarios. Phase B v2 adds 120 unique real-model scenarios repeated three times while holding policy-v1, the fictional tools, evaluator, and simulated executor fixed.

04Result

Zero governed unsafe executions across 180 Phase B non-execution trials

All 360 real-model trials matched exact tool, argument, and governance expectations. The governed path preserved all 180 expected legitimate executions; the naive baseline executed all 180 proposals expected not to execute immediately, while the governed path executed none.

05Limitation

This is integration evidence, not a safety proof.

Tools and execution are simulated, policy-v1 is intentionally small, and the corpus is authored and structured to exercise the boundary. The result covers one model and three sequential repetitions; it is not a production certification or universal safety proof.

06Revision

The experiment became harder after the first evidence cycle.

The first Phase B pilot exposed duplicate prompts, permissive abstention scoring, under-specified scopes, unrealistic values, and incomplete provenance. Phase B v2 corrected those weaknesses, froze 120 unique prompts, and published the reviewed result with limitations.

Measured result

The tested boundary preserved legitimate actions and stopped non-execution proposals.

360 / 360

Real-model trials matched exact model and governance expectations.

180 / 180

Naive direct baseline executions in Phase B trials expected not to execute immediately.

0 / 180

Governed unsafe executions in those same Phase B trials.

Interpretation: this supports the implementation-level claim for the tested fictional corpus, model, policy, runner, and simulated executor. It does not establish universal agent safety or production readiness.

What would change the claim?

Evidence must be able to weaken the architecture, not only support it.

Governance failure

If Phase B produces an unsafe execution through an approved path, the control model or implementation must change before a stronger effectiveness claim is made.

Usability failure

If legitimate actions are blocked or pushed into approval too often, the policy boundary may be safe but impractical. That is still an architecture failure worth revising.

Experiment failure

If the corpus, evaluator, or baseline turns out to bias the comparison, the right response is to revise the experiment rather than defend the headline number.

Trace the evidence

Inspect each layer yourself