Skip to content

Paper demonstration corpus: n=2 backend participant evidence runs #600

Description

@Brad-Edwards

Context

The systems paper needs evidence that an authored ACES scenario can support participant execution as repeatable, inspectable experimental apparatus. The refined claim is stronger than "two repeated runs": n=2 should mean the same authored enterprise participant/evidence scenario is realized on two independent emulation backends, APTL and the ACES libvirt reference backend.

The goal is not to prove participant/model performance. The goal is to prove the system boundary: authored SDL, processor output, backend realization, participant implementation, episode/action/observation history, evaluator/Wazuh evidence, and outcome interpretation remain separable and auditable across backends.

Parent/link-back: Brad-Edwards/aptl#554. Scenario issue: #598.

Scope

Define and publish a demonstration corpus for the ACES paper reference scenario consisting of backend-paired participant runs:

  • one APTL realization of the authored paper scenario;
  • one libvirt reference-backend realization of the same authored paper scenario;
  • optional repeat runs per backend for stability, clearly separated from the n=2 backend claim.

The corpus should include a cross-backend invariant ledger rather than a leaderboard or performance table.

Acceptance criteria

  • The same authored scenario identity/hash is recorded for the APTL and libvirt runs.
  • Each run records scenario/source hash, processor artifact identity, backend manifest/capability profile, runtime snapshots, realized topology/network attachment matrix, participant implementation provenance, participant episode history, participant behavior history, participant terminal observation, evaluator/Wazuh evidence, and outcome interpretation evidence.
  • Each run includes reachability evidence: participant can reach the DMZ portal through the declared action surface; direct participant reachability to the internal DB and Wazuh/evaluator surface is absent or explicitly outside scope and checked as negative evidence where the backend supports it.
  • The comparison ledger identifies preserved invariants, realization differences, unsupported/degraded surfaces, and evidence limitations across APTL and libvirt.
  • Evidence can be inspected without backend-private state, raw secrets, libvirt/Docker private identifiers as semantics, or unredacted participant credentials.
  • The corpus links to the APTL evidence issue research: produce the APTL side of paired-backend participant evidence Brad-Edwards/aptl#558 and Libvirt backend: ParticipantRuntime for the paper scenario #614 and Libvirt paper proof: evaluator evidence and Wazuh readback artifact #615.

Non-claims

  • No autonomous-agent capability benchmark claim.
  • No claim that Wazuh detection quality is evaluated.
  • No model-defense robustness claim.
  • No full semantic equivalence across backends beyond the checked invariant ledger.

Related

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:runtimeRuntime and control-plane codedocumentationImprovements or additions to documentationenhancementNew feature or request

    Type

    No type

    Projects

    No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions