Goal
Track how PowerContext expands its benchmark coverage without creating a separate execution stack for every benchmark.
Existing E2E infrastructure, including Harbor, ACP, and containerized execution, should be the starting point where it fits. A benchmark may reuse upstream Harbor content, be represented as a Harbor task, or use another approach when its native structure calls for one. The integration should preserve the benchmark's original intent, inputs, and evaluation semantics.
Principles
We should prefer benchmarks that are recognized or used by comparable memory, context, or agent projects.
Each proposal should explain why the benchmark is relevant to PowerContext, whether the comparison can be fair, and whether the result is clear and measurable. The exact execution and evaluation shape may differ between benchmarks.
Community contributions should focus on a small, runnable validation rather than a complete evaluation. Contributors are not expected to cover the full runtime or inference cost.
After a benchmark is accepted, maintainers may run the complete evaluation or a useful subset. Published results should state what was run and should not present a partial run as a complete benchmark result.
Tracking
This issue tracks benchmark adoption and shared infrastructure. Benchmark-specific design and execution details belong with each benchmark.
Goal
Track how PowerContext expands its benchmark coverage without creating a separate execution stack for every benchmark.
Existing E2E infrastructure, including Harbor, ACP, and containerized execution, should be the starting point where it fits. A benchmark may reuse upstream Harbor content, be represented as a Harbor task, or use another approach when its native structure calls for one. The integration should preserve the benchmark's original intent, inputs, and evaluation semantics.
Principles
We should prefer benchmarks that are recognized or used by comparable memory, context, or agent projects.
Each proposal should explain why the benchmark is relevant to PowerContext, whether the comparison can be fair, and whether the result is clear and measurable. The exact execution and evaluation shape may differ between benchmarks.
Community contributions should focus on a small, runnable validation rather than a complete evaluation. Contributors are not expected to cover the full runtime or inference cost.
After a benchmark is accepted, maintainers may run the complete evaluation or a useful subset. Published results should state what was run and should not present a partial run as a complete benchmark result.
Tracking
This issue tracks benchmark adoption and shared infrastructure. Benchmark-specific design and execution details belong with each benchmark.