You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The systems paper needs evidence that an authored ACES scenario can support participant execution as repeatable, inspectable experimental apparatus. The refined claim is stronger than "two repeated runs": n=2 should mean the same authored enterprise participant/evidence scenario is realized on two independent emulation backends, APTL and the ACES libvirt reference backend.
The goal is not to prove participant/model performance. The goal is to prove the system boundary: authored SDL, processor output, backend realization, participant implementation, episode/action/observation history, evaluator/Wazuh evidence, and outcome interpretation remain separable and auditable across backends.
Define and publish a demonstration corpus for the ACES paper reference scenario consisting of backend-paired participant runs:
one APTL realization of the authored paper scenario;
one libvirt reference-backend realization of the same authored paper scenario;
optional repeat runs per backend for stability, clearly separated from the n=2 backend claim.
The corpus should include a cross-backend invariant ledger rather than a leaderboard or performance table.
Acceptance criteria
The same authored scenario identity/hash is recorded for the APTL and libvirt runs.
Each run records scenario/source hash, processor artifact identity, backend manifest/capability profile, runtime snapshots, realized topology/network attachment matrix, participant implementation provenance, participant episode history, participant behavior history, participant terminal observation, evaluator/Wazuh evidence, and outcome interpretation evidence.
Each run includes reachability evidence: participant can reach the DMZ portal through the declared action surface; direct participant reachability to the internal DB and Wazuh/evaluator surface is absent or explicitly outside scope and checked as negative evidence where the backend supports it.
The comparison ledger identifies preserved invariants, realization differences, unsupported/degraded surfaces, and evidence limitations across APTL and libvirt.
Evidence can be inspected without backend-private state, raw secrets, libvirt/Docker private identifiers as semantics, or unredacted participant credentials.
Context
The systems paper needs evidence that an authored ACES scenario can support participant execution as repeatable, inspectable experimental apparatus. The refined claim is stronger than "two repeated runs": n=2 should mean the same authored enterprise participant/evidence scenario is realized on two independent emulation backends, APTL and the ACES libvirt reference backend.
The goal is not to prove participant/model performance. The goal is to prove the system boundary: authored SDL, processor output, backend realization, participant implementation, episode/action/observation history, evaluator/Wazuh evidence, and outcome interpretation remain separable and auditable across backends.
Parent/link-back: Brad-Edwards/aptl#554. Scenario issue: #598.
Scope
Define and publish a demonstration corpus for the ACES paper reference scenario consisting of backend-paired participant runs:
The corpus should include a cross-backend invariant ledger rather than a leaderboard or performance table.
Acceptance criteria
Non-claims
Related