Local-first evidence-chain forensics for RAG and AI agents.
ContextTrace is a Python SDK and CLI for tracing a failed answer from the user query through retrieved context, answer claims, citations, verdicts, root cause, repair guidance, and CI regression tests.
query -> retrieved context -> answer claims -> citations -> verdicts -> root cause -> regression test
Use it when a RAG or agent score is not enough: ContextTrace points at the unsupported or contradicted claim, the evidence span and citation involved, why the failure likely happened, and how to keep it from coming back. It is not a hosted dashboard. Traces, reports, judge cache, and SQLite state stay local by default.
Latest stable release: ContextTrace 1.1.0, tested on Python 3.10 through 3.13.
- Versioned public JSON schemas with golden compatibility tests for v1.0 traces.
- Recursive privacy redaction plus streaming-safe, concurrency-isolated framework integrations.
- Bounded single-trace and batch verification with explicit payload, queue, and worker limits.
- A frozen
semantic_v1_calibratedverifier boundary with separate generic, policy, temporal, and legacy calibration rule packs. - Expanded cross-platform packaging, dependency-audit, and 80% coverage gates.
See the 1.1.0 release notes and public artifact schemas for details.
pip install contexttrace
contexttrace initInstall from source for unreleased changes:
git clone https://github.com/samarth1412/Context-Trace.git
cd Context-Trace
pip install -e packages/contexttracecontexttrace verify-demo unsupported_claim --report
contexttrace demo --dataset refund_policy
contexttrace report --last --openDefault local storage:
.contexttrace/contexttrace.db
Create a portable trace with a query, answer, retrieved contexts, and optional citations:
{
"query": "How long does refund processing take?",
"answer": "Refunds are processed within 5 business days.",
"contexts": [
{
"id": "policy",
"text": "Customers may request refunds within 30 days of purchase."
}
]
}Run local evidence checks:
contexttrace inspect trace.json
contexttrace verify trace.json --report
contexttrace diagnose trace.json --report
contexttrace qa trace.json --corpus docs/ --report
contexttrace repair trace.json --corpus docs/ --out repair_plan.mdContextTrace classifies each claim as supported, partially_supported, unsupported, unverifiable, or contradicted, then exposes separate statuses for support, truth, source freshness, citation quality, and likely fix.
Important: supported means grounded by the selected evidence span. It does not mean independently true, current, or authoritative.
diagnose also accepts agent step traces and localizes tool/final-answer
failures:
{
"goal": "Book a meeting with Alex",
"steps": [
{
"type": "tool_call",
"tool": "calendar.search",
"args": {"date": "Friday"},
"result": "No availability"
},
{
"type": "final_answer",
"content": "I booked it for Friday."
}
]
}contexttrace diagnose examples/diagnose_agent_trace.json --report --fail-on high_riskThe diagnosis flags tool_result_contradicted_by_final_answer and suggests
gating final-answer generation on tool-result status.
Turn that diagnosis into a CI regression test:
contexttrace diagnose examples/diagnose_agent_trace.json \
--generate-test \
--test-out tests/contexttrace/test_calendar_agent_diagnosis.py
pytest tests/contexttrace/test_calendar_agent_diagnosis.pyrepair turns diagnosis into an evidence-backed implementation plan. With a
local corpus, it distinguishes retrieval miss, reranking failure, chunking
issue, corpus gap, answer overreach, and stale or conflicting evidence:
contexttrace repair trace.json \
--corpus docs/ \
--out repair_plan.md \
--json-out repair_plan.jsonThe plan records the failed claim, retrieved and corpus evidence, prioritized root-cause-specific changes, and commands to verify the fix. Add only the recaptured passing trace to the generated must-pass regression command.
| Mode | Use When |
|---|---|
lexical |
Fast default checks with no optional dependencies. |
semantic |
Local paraphrase and role-aware contradiction checks. |
local_ml |
Offline hash-embedding similarity, optionally backed by a local SentenceTransformers model. |
nli |
Local claim+span entailment or contradiction with a local Transformers or ONNX NLI model. |
judge |
Higher-accuracy local LLM judging through Ollama, LM Studio, vLLM, or a local OpenAI-compatible server. The judge sees selected evidence spans, not the full answer prose. |
Run the stronger local non-LLM verifier:
contexttrace verify trace.json --mode local_ml --report
contexttrace verify-benchmark --mode local_ml --case-set allOptional neural local-ML support never downloads models automatically:
pip install "contexttrace[local-ml]"
set CONTEXTTRACE_LOCAL_ML_MODEL_PATH=C:/models/bge-small-en-v1.5Run local NLI when you want mechanical claim-versus-span entailment:
pip install "contexttrace[nli]"
set CONTEXTTRACE_NLI_MODEL_PATH=C:/models/deberta-v3-nli
contexttrace verify trace.json --mode nli --report
contexttrace nli-calibrate --case-set all --reportRun a local judge with Ollama:
set CONTEXTTRACE_JUDGE_PROVIDER=ollama
set CONTEXTTRACE_JUDGE_MODEL=llama3.1
contexttrace verify trace.json --mode judge --report
contexttrace judge-calibrate --case-set all --reportRemote judges are blocked while local_only: true is active. To use a remote judge, explicitly disable local-only mode and configure the provider/API key.
# Find whether support existed elsewhere in the corpus.
contexttrace audit trace.json --corpus docs/ --report
# Compare a baseline and current answer after a prompt, model, or retriever change.
contexttrace compare baseline.json current.json --report
# Turn saved failures into replayable endpoint tests.
contexttrace suite create traces/failure.json --out contexttrace-suite.json
contexttrace suite run contexttrace-suite.json --endpoint http://localhost:8000/query --reportCommon root causes include retrieval_miss, reranking_failure, chunking_issue, corpus_gap, answer_overreach, stale_source, citation_mismatch, and should_have_abstained.
support_status, truth_status, and source_status stay separate so a claim can be grounded by a source while the source itself remains stale, wrong, or unassessed.
Source metadata can include source_authority, source_timestamp, source_version, canonical, or canonical_source. ContextTrace uses those local fields to flag grounded_but_stale, grounded_but_conflicted, grounded_by_low_authority_source, or supported_by_canonical_source.
ContextTrace-Bench is the repo-level benchmark for claim-level failure attribution, root-cause diagnosis, citation-error detection, and evidence-span localization.
python benchmarks/contexttrace_bench/run_contexttrace.py \
--mode semantic \
--case-set all \
--enforce-sota-gatesIt writes reproducible JSON, Markdown, leaderboard, static HTML, candidate-input,
and error-analysis artifacts to benchmarks/contexttrace_bench/out/. The default
run targets 500 cases by adding deterministic generated variants to the curated
real-doc cases. Reports include deterministic 95% case-bootstrap confidence
intervals and per-label breakdowns for the headline verifier metrics.
Run the separate public-doc holdout without generated variants:
python benchmarks/contexttrace_bench/run_contexttrace.py \
--mode semantic \
--case-set public_holdout \
--no-generated-cases \
--output-dir benchmarks/contexttrace_bench/out/public_holdoutCurrent 150-case holdout status is tracked in the baseline runbook. The holdout is kept outside the default 500-case run so it can expose external validation gaps without invalidating the main leaderboard rows.
candidate_inputs.jsonl gives external evaluators the exact trace payloads to
score. Candidate prediction JSON files can then be scored with --candidate to
compare ContextTrace against external evaluators or internal baselines on the
same labeled cases. Use benchmarks/contexttrace_bench/adapt_candidate.py to
normalize generic evaluator output into the candidate schema.
For external dataset scaffolding, benchmarks/contexttrace_bench/ragtruth_adapter.py
builds a ContextTrace-style case pack from RAGTruth response.jsonl and
source_info.jsonl. Score it with run_contexttrace.py --case-pack, and use
benchmarks/contexttrace_bench/ragtruth_review.py to generate and apply the
human evidence-span review queue. Treat unreviewed output as review input until
answer-side hallucination spans are manually mapped to source evidence spans.
The public benchmark card is
benchmarks/contexttrace_bench/BENCHMARK_CARD.md.
Methodology and baseline runbooks live in
benchmarks/contexttrace_bench/METHODOLOGY.md
and benchmarks/contexttrace_bench/BASELINES.md.
The public holdout track is documented in
benchmarks/contexttrace_bench/DIAG150.md.
The pinned official CRAG calibration and review track is documented in
benchmarks/contexttrace_bench/CRAG.md.
Human audit criteria for calling that split frozen are in
benchmarks/contexttrace_bench/AUDIT.md.
The evidence and public-claim policy is in
docs/sota-readiness.md.
Run python benchmarks/contexttrace_bench/sota_gate.py for the fail-closed,
machine-readable broad-claim gate; the current result is in
benchmarks/contexttrace_bench/SOTA_STATUS.md.
Treat the default 500-case run as verifier-readiness evidence; publish full
competitor rows and independent external dataset results before making broad
state-of-the-art claims.
Capture one live endpoint response:
contexttrace capture endpoint \
--endpoint http://localhost:8000/query \
--query "What is the refund policy?" \
--answer-path $.answer \
--contexts-path $.contexts \
--citations-path $.citations \
--out traces/refund_trace.json \
--verify \
--reportOr capture artifacts from Python:
from contexttrace import capture_rag_trace, write_rag_trace
trace = capture_rag_trace(
query=question,
answer=answer,
contexts=retrieved_docs,
metadata={"system": "support-rag"},
)
write_rag_trace(trace, "trace.json")from contexttrace import ContextTrace
ct = ContextTrace(project="support-rag")
with ct.trace(query="What is the refund policy?") as trace:
chunks = retriever.search("What is the refund policy?")
trace.log_retrieval(chunks)
trace.log_context(chunks[:5])
answer = llm.generate("What is the refund policy?", chunks[:5])
trace.log_answer(answer, usage={"total_tokens": 1200})
trace.log_citations([
{"claim": "Refunds are available within 30 days.", "source_chunk_id": "chunk_12"}
])
result = trace.evaluate()
print(result["failure"]["failure_type"])pip install "contexttrace[langchain]"
pip install "contexttrace[llamaindex]"
pip install "contexttrace[fastapi]"
pip install "contexttrace[langgraph]"
pip install "contexttrace[otel]"
pip install "contexttrace[all]"Includes LangChain, LlamaIndex, FastAPI, LangGraph, and OpenTelemetry hooks.
ContextTrace makes no network calls unless you point it at an endpoint or configure a judge provider. Local controls include:
local_only: truelog_chunk_text: falselog_answer_text: falsestorage_pathjudge_cache_enabled: truejudge_cache_path: .contexttrace/judge_cache.json
ContextTrace is a diagnostic tool, not a correctness proof. It verifies grounding against provided evidence; it does not certify real-world truth. Claim extraction is rule-based, contradiction detection is conservative, and high-stakes outputs still need human review.
- PyPI: https://pypi.org/project/contexttrace/
- Latest release: https://github.com/samarth1412/Context-Trace/releases/tag/v1.1.0
- Docs: docs
- Artifact schemas: docs/artifact-schemas.md
- Issues: https://github.com/samarth1412/Context-Trace/issues
- Changelog: CHANGELOG.md