Alibaba set out to make it easy to do business anywhere; Draftpaper-loop sets out to make no paper hard to write.
A local research workflow that turns an idea, discipline methods, and real data into auditable scientific figures, a complete manuscript, and a traceable main.pdf.
Draftpaper-loop organizes paper production as an evidence-first research loop. It confirms the research question and feasibility, matches or supplements data and method capabilities, executes analysis, checks whether figures support the claims, and then builds the manuscript, citation audit, discipline review, independent review, and PDF release from one evidence version.
- Create a structured paper project from a research idea, existing data, references, or project code.
- Generate a bilingual research blueprint, claim contract, statistical requirements, and main-figure storyboard for one focused human confirmation.
- Detect single or cross-disciplinary needs, bind
data_connector,method_template, andreview_ruleplugins, and run real project code for data processing, training, statistics, and scientific plotting. - Audit project-local capabilities and prepare traceable rescue tasks from the plugin registry, AcademicForge metadata, or public research-code repositories.
- Trace main figures to a research-plan claim, cohort, data plugin, method plugin, run output, evidence ID, and applicable discipline rules.
- Write Results → Introduction → Data → Methods → Discussion from the same evidence, stage-owned code, formulas, figures, and literature.
- Audit citation support, bibliography format, discipline statistics, Results semantics, and reproducibility before two independent blind reviewers inspect the manuscript.
- Complete authors, affiliations, ORCID, funding, acknowledgments, data/code links, references, and precise paragraph revisions in one packet before releasing a hash-bound
main.pdf.
Current release: v0.33.1. This release adds a capability-negotiated extension host, non-blocking workflow events, scoped artifact reads, and extension status projection while retaining the v0.33.0 evidence and author-completion protections. See Recent Updates for release history; the rest of this README is organized by research task.
| Research stage | Main capability | Key artifacts |
|---|---|---|
| Research design | Idea analysis, journal profile, bilingual blueprint, claim/statistical/figure contracts, and human confirmation | confirmed plan hash, claims, figure storyboard |
| Discipline capability | Single/cross-discipline detection, data/method/review plugin matching, project-local audit, and capability rescue | discipline/capability contracts, plugin bindings |
| Data and methods | Source inventory, feasibility, stage-owned code, verified runs, formula and variable extraction | data/method manifests, run/formula manifests |
| Figures and evidence | Semantic figure/panel contracts, figure-code trace, result validity, and Result Support | main/supporting figures, result manifest, evidence registry |
| Scientific writing | Paper Narrative Engine, section evidence packets, Codex free composition, Scientific Editor | Results, Introduction, Data, Methods, Discussion |
| Literature and citations | Search, Zotero, BibTeX, PDF/summary evidence, citation intent, and audit | library.bib, citation evidence, final audit |
| Review and release | Post-Results discipline review, two blind reviewers, author-completion transaction, compilation, and release hash | reviewer reports, completion packet, main.pdf |
- v0.1-v0.13: paper-project and research-stage foundations. References, journal profiles, research plans, methods/results/discussion writing, artifact tracking, Zotero, observations, scientific plotting, and stage-owned code.
- v0.14-v0.20: discipline plugins and result support. Data connectors, method templates, review rules, plugin sufficiency, AcademicForge/GitHub candidate rescue, cross-discipline execution ledgers, and post-Results discipline review.
- v0.21-v0.28: scientific narrative and evidence semantics. Paper Narrative Engine, section evidence packets, free composition plus Scientific Editor, run/cohort/estimand binding, semantic figure contracts, independent review, and reproducibility bundles.
- v0.28.1-v0.33: transactions, release, and exact recovery. Artifact DAG, unified CommandSpec, scientific non-zero exits, author-completion transactions, stable paragraph locators, cross-journal/cross-platform wheel regression, Result Support v3, and release-hash binding.
Version numbers explain capability origin. Daily use follows the current research question and project state; status, doctor, and run-pipeline recommend the next action.
PowerShell:
git clone https://github.com/xiejhhhhhh/Draftpaper_loop.git
cd Draftpaper_loop
py -3 -m venv .venv
.\.venv\Scripts\python -m pip install -U pip
.\.venv\Scripts\python -m pip install -e ".[plotting]"
.\.venv\Scripts\draftpaper doctor --jsonbash/macOS/Linux:
git clone https://github.com/xiejhhhhhh/Draftpaper_loop.git
cd Draftpaper_loop
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -U pip
python -m pip install -e ".[plotting]"
draftpaper doctor --jsonOpen the repository in Codex and describe the idea, data location, target journal, and existing code:
Use Draftpaper-loop in this repository to create a paper project for my research idea.
Read the idea, data, and existing code first, identify the disciplines and capability gaps, and show me the research blueprint for confirmation.
After confirmation, follow project state through data, methods, figures, writing, citation audit, two independent reviews, and PDF compilation.
Codex handles open-ended research reasoning and prose. The Draftpaper-loop CLI records stage state, evidence binding, write boundaries, and human confirmations.
draftpaper create-project --idea "Your research idea" --field "astronomy machine learning" --target-journal MNRAS
draftpaper status --project .\projects\<project>
draftpaper run-pipeline --project .\projects\<project>run-pipeline stops at research-blueprint, result-support, and final-release checkpoints and recommends the next command. Projects default to projects/<project>/; the compiled manuscript is projects/<project>/latex/main.pdf.
Basic tutorial video: Bilibili
idea, existing data, project code, and literature
-> create an isolated paper project
-> identify disciplines and target journal
-> search/import literature and build citation evidence
-> generate bilingual blueprint, claims, statistics, and main-figure storyboard
-> human confirmation of the research blueprint
-> assess plugin sufficiency and audit project-local capabilities
-> execute data plugins, method plugins, and real project code
-> generate main figures, supporting evidence, tables, and evidence registry
-> assess whether results support the planned claims
-> accept evidence, narrow claims, or supplement data/methods
-> Results -> Introduction -> Data -> Methods -> Discussion
-> post-Results composite-discipline review and semantic repair
-> final author completion and precise revisions
-> final citation audit
-> two independent blind reviewers
-> integrity, journal-format, and PDF compilation checks
-> confirm one release hash
-> latex/main.pdf
load project state
-> select the current stage action
-> execute and produce structured artifacts
-> validate scientific contracts and file hashes
-> record artifacts, runs, evidence, and human decisions
-> mark exact downstream stages stale when inputs change
-> diagnose failure and route to the owning stage
-> repeat until the manuscript is release-ready
Deterministic contracts own project state, cohort/run identity, plugin provenance, figure semantics, formulas, citations, stale propagation, write sets, and release hashes. Codex or another Agent owns open-ended literature interpretation, method design, scientific reasoning, and prose.
The three concentrated human checkpoints are:
- Research blueprint confirmation: inspect the research question, claims, cohorts, data/method requirements, statistical standards, main figure groups, and feasibility boundary together.
- Core result and claim-support confirmation: inspect verified runs, core figures, metrics, uncertainty, and maximum supported claim strength, then select the next route.
- Final manuscript and release confirmation: inspect the completion packet, candidate PDF, final citation audit, two blind-review reports, and release hash together.
- Claim-narrowing route: freeze accepted figures and metrics, reduce claim strength, and rebuild the affected manuscript sections.
- Data/method rescue route: supplement data roles, quality control, methods, or validation, then rerun the affected evidence, figure, and manuscript chain. The current implementation selects one route for the whole result-support checkpoint; per-claim routing remains a later capability.
Result Support v3 prefers current resolved evidence, followed by the selected run manifest and explicitly run-bound result tables. Route commands bind the current checkpoint hash. Cohort, metric, or figure-evidence findings discovered after Results return to the same checkpoint so prose and evidence versions stay aligned.
draftpaper_cli/discipline_modules/<discipline>/ registers three formal research plugin classes:
data_connectors: access, reading, parsing, cleaning, normalization, cohort construction, and data-quality checks.method_templates: statistics, feature engineering, model training, validation, ablation, uncertainty estimation, and scientific plotting.review_rules: discipline- and method-aware checks for statistics, baselines, split/leakage, fit or classification quality, calibration, robustness, figure-claim alignment, and reproducibility.
Plugin manifests declare runtime level, validation level, dependencies, inputs/outputs, fixtures, and provenance. Contracts, mocks, and fixtures verify interfaces; main-figure evidence requires outputs and hashes validated on a real project or live run.
The research plan produces discipline/capability contracts, assigning each claim and figure requirement to primary/secondary disciplines and data/method/review capabilities. Capability gaps follow this order:
- Audit existing project data/method code and outputs into a restricted
project_localbinding. - Search the current registry for reusable plugins.
- Extract candidate capabilities from AcademicForge-style registry metadata.
- Inspect license, structure, reproducibility, and input/output contracts in public research-code repositories.
- Let Codex generate a project-specific implementation and validate it in the current project.
- Send reusable capabilities through generalization, fixtures, overlap review, license review, and human-confirmed promotion.
A figure-stage blocking diagnosis appears after the data and method rescue routes have auditable outcomes and a required input or implementation remains missing.
research-plan claim
-> data requirement -> data plugin/project-local binding
-> method requirement -> method plugin/project-local binding
-> verified run output
-> evidence ID and cohort view
-> figure/panel contract
-> Results claim
-> applicable discipline review rules
Data and method plugins produce the real figure inputs and method outputs. Matching review_rule plugins inspect figures and Results at review-results-with-discipline-rules.
Public research code and AcademicForge use metadata-first candidate pipelines that retain repository, commit, license, dependency, input/output, runtime-level, and provenance records. Candidates enter formal discipline modules through generalize-plugin-candidate, validate-plugin-candidate, package-plugin-contribution, preflight-plugin-contribution, review-plugin-contribution, and human-confirmed promote-plugin-candidate.
workflow_recipe, paper_contract, and shared_capability stay in a support layer. Their verifiable statistical, baseline, ablation, split/leakage, citation-support, and reproducibility conditions may flow into discipline-specific review_rule_candidate records. See the CLI Reference for the complete command chain.
third_party/ stores upstream snapshots, immutable pointers, license notices, and provenance. The wheel-installable paper-fetch fallback lives under draftpaper_cli/_vendor/paper_fetch_skill.
plan-figures reads the confirmed research plan, claims, cohorts, data/method requirements, statistical contract, and existing evidence to produce main-figure groups, supporting/appendix figures, and caption contracts. The first caption sentence summarizes the scientific conclusion of the complete group; later sentences explain each panel, cohort, estimand, uncertainty, and claim boundary.
generate-analysis-code follows the figure contract and bound capabilities. Execution failures produce figure_execution_diagnosis.json/.html and distinguish data gaps, method gaps, dependency issues, runtime errors, insufficient result quality, and required user confirmation, with a recommended repair command.
Paper Narrative Engine reads the Scientific Evidence Registry, result manifest, figure story arc, literature comparison matrix, stage-owned code, and formula trace to build section evidence packets and paragraph goals. Codex composes freely inside those evidence boundaries; post-writing contracts validate numbers, cohorts, runs, models, metric dimensions, citation roles, internal paths, and claim strength.
- Results: interpret main figure groups and key tables, then use supporting/appendix figures for robustness, uncertainty, and boundaries.
- Introduction: organize the research gap from the question, significance, literature evidence, and contribution established by Results.
- Data: reconstruct sources, samples, variables, preprocessing, and missingness from connectors, inventory, cohort registry, and stage code.
- Methods: reconstruct stages from plugins, real implementation, run manifest, formula AST, and figure-code trace; explain formulas, variables, assumptions, and figure relationships.
- Discussion: align local findings with comparison literature and discuss mechanism, innovation, limitations, external validity, and future work.
record-observation stores only stage summaries already shown to the user. Paths, commands, credentials, manifest fields, and local filenames stay in internal context.
Scientific gates, non-zero exit status, one change taxonomy, and the artifact DAG determine which project states may continue. Data, cohort, method, run, metric, main-figure, or claim changes reopen the owning scientific chain; citation-local edits, author metadata, and presentation-only changes receive narrower stale scopes.
Each project records stages, inputs, outputs, confirmations, and recovery reasons through project_passport.yaml, artifact_ledger.jsonl, checkpoint_ledger.jsonl, integrity_ledger.jsonl, and artifact hashes. Use doctor --explain and diagnose-gate-failures to inspect state and the next action.
The literature stage preserves BibTeX, the reference registry, citation evidence, notes, per-paper summaries, and available PDF/full-text evidence. Search results, user-provided papers, and Zotero collections retain origin and selection policy. Writing assigns references to direct support, method provenance, data provenance, comparison context, and background roles.
Zotero example:
$env:ZOTERO_LIBRARY_ID="your_zotero_library_id"
$env:ZOTERO_LIBRARY_TYPE="user"
$env:ZOTERO_API_KEY="your_zotero_api_key"
draftpaper list-zotero-collections
draftpaper search-literature --project <project> --zotero-collection "My Paper References" --zotero-context all --zotero-min-items 20Citation audit runs after the final sections and author completion, and before independent review. It compares manuscript claims with BibTeX metadata, citation evidence, passages, numbers, negation, and causal direction. It reports weak support, misplaced citations, and over-strong wording. Repair prefers tightening or rewriting prose while preserving user-confirmed references and reference coverage.
review-results-with-discipline-rules reads current Results, figure/plugin trace, run outputs, evidence IDs, and composite-discipline rules. It checks metric statements, sample boundaries, baselines, ablations, statistical dimensions, uncertainty, fit/classification standards, and figure interpretation. Semantic findings enter local Results revision; genuine capability gaps return to result support.
After citation audit, the loop builds a frozen anonymous single-manuscript review bundle. Two independent reviewers separately inspect scientific correctness, evidence sufficiency, structure, prose, figures, citations, and reproducibility materials. Both reports bind the same manuscript/evidence/bundle hash; critical or major findings enter revision and re-review, and final confirmation binds the latest reports and compiled PDF.
One manuscript_completion.yaml can add authors, affiliations, ORCID, corresponding author, funding, acknowledgments, keywords, short title, data/code availability, user-confirmed references, and multiple section revisions.
LaTeX line numbers are user-facing hints. Writes also validate stable paragraph_id, expected text, occurrence, and SHA-256. Paragraphs can be relocated after layout drift, while ambiguous, duplicate, or stale targets return the whole packet to preview.
Preview shows the user-declared change class, inferred class, candidate evidence refs, and exact stale scope. Author metadata, acknowledgments, and prose-only refinements stay downstream; new data provenance, executed-method details, metrics, claims, or figure interpretation reopen the matching research or writing stages through current evidence refs.
Completion first builds one diff, candidate LaTeX, and candidate PDF. After acceptance, atomic apply revalidates the packet, project revision, source map, evidence snapshot, and before hash, then writes rollback receipts and exact-text user locks.
draftpaper prepare-manuscript-completion --project <project>
draftpaper preview-manuscript-completion --project <project> --input manuscript_completion.yaml
draftpaper apply-manuscript-completion --project <project> --packet-id <id> --packet-hash <sha256>
draftpaper review-final-manuscript --project <project>
draftpaper confirm-final-manuscript --project <project> --release-hash <sha256>Release order is author completion and precise revision → final citation audit → two independent blind reviews → integrity and compilation checks → release-hash confirmation. See Final Manuscript Completion.
pip install -e .: minimal control plane, project state, references, and basic PDF/image inspection.pip install -e ".[plotting]": NumPy, pandas, Matplotlib, and standard scientific plotting/analysis entry points.pip install -e ".[fulltext]": enhanced PDF and full-text extraction.pip install -e ".[mcp]": local stdio MCP.draftpaper doctor --json: detect the current profile, missing modules, and recovery commands.draftpaper token-report --project <project>: summarize recorded token/cost receipts.
Use .[plotting-full] for complex plotting backends. Agents should normally inspect status and call run-pipeline; common diagnostic and recovery commands are:
draftpaper status --project <project>
draftpaper doctor --project <project> --explain
draftpaper run-pipeline --project <project>
draftpaper detect-artifact-drift --project <project>
draftpaper sync-artifact-stale --project <project>
draftpaper diagnose-gate-failures --project <project>
draftpaper run-integrity-gate --project <project>
draftpaper audit-citations --project <project> --final
draftpaper assess-publication-readiness --project <project>The Python API exposes the same evidence-semantic entry points, including resolve_result_evidence, build_scientific_evidence_registry, validate_figure_semantics, create_evidence_snapshot, and submit_section_draft. MCP is a controlled projection of CommandSpec/CLI handlers for local Agent integration.
CommandSpec, the schema registry, and quality contracts form one command control plane for risk, inputs, outputs, write scope, network behavior, and human checkpoints. Generated references come from the same registry, so the README keeps only common user paths.
| Need | Document |
|---|---|
| Commands, parameters, inputs/outputs, and risk | CLI Reference |
| Minimal, plotting, fulltext, and MCP profiles | Install Profiles |
| Write, network, and confirmation boundaries | Command Risk Matrix |
| Project token and cost receipts | Token and Cost Reporting |
| Final completion and stable paragraph locators | Final Manuscript Completion |
| DPL schema, project state, and artifacts | DPL Schema |
draftpaper_cli/ core Python package, CLI, state, and evidence contracts
draftpaper_cli/discipline_modules/ data connectors, method templates, and review rules
draftpaper_cli/_vendor/ wheel-installable runtime fallbacks
codex_skills/draftpaper-workflow/ Codex workflow skill and Agent contracts
docs/ guides, schemas, audits, and generated references
tests/ unit, adversarial, wheel, cross-platform, and release regressions
third_party/ upstream snapshots, pointers, licenses, and provenance
projects/ local paper projects, ignored by git by default
Within one paper project, data acquisition, cleaning, and cohort code belongs under data/; models, statistics, validation, and plotting code belongs under methods/; results/ stores figures, tables, and metadata; writing/, citation_audit/, review/, and latex/ store prose, audits, reviews, and the final PDF. Large datasets may remain at their source location through private locators, read-only fingerprints, and manifests while workflow artifacts remain owned by the paper project.
Every writing command checks its declared write set before and after execution and uses state revisions, project locks, artifact hashes, and transaction receipts. MCP capability tokens, executable allowlists, path confinement, network policies, and log redaction constrain application behavior. Public multi-tenant isolation, account systems, and hosted billing are separate product-engineering boundaries.
Run baseline verification with:
python tools/validate_capability_truth_matrix.py
python -m pytest tests/test_capability_truth_matrix.py
python -m pytest
python -m buildDraftpaper-loop welcomes reusable discipline-module contributions, especially data connectors, method templates, reviewer rules, fixtures, and project-tested workflow lessons that can be generalized without private paths, credentials, raw data, or project-specific claims.
Current contributors:
- Jinray Xie: overall Draftpaper-loop framework, including literature workflows, data and methods auditing, result-output auditing, rollback mechanisms, and discipline-module contributions for deep learning, astronomy, and geography.
- Chen Wei: astronomy discipline-module supplements and related validation generation.
Draftpaper-loop is source-available for non-commercial research, evaluation, education, and personal paper-writing workflows. Commercial use, paid services, SaaS deployment, enterprise deployment, resale, or integration into commercial products requires separate written authorization from the developer.
Sponsorship or donation supports project maintenance but does not grant commercial use rights. Commercial use still requires separate prior written authorization.
Draftpaper-loop uses the DPL schema family to represent local-first paper-loop state, including project passports, stage manifests, citation evidence, run manifests, result manifests, artifact hashes, claim traces, and loop events.
See LICENSE, NOTICE, COMMERCIAL_LICENSE.md, TRADEMARK.md, COMPLIANCE.md, docs/DPL_SCHEMA.md, and docs/FORENSIC_FINGERPRINTING.md for the current non-commercial source-available terms, attribution notice, commercial authorization scope, project-name/trademark policy, public schema identity, and compliance boundary.
For commercial authorization, contact xiejinhui22@mails.ucas.ac.cn.
Personal homepage: https://xiejhhhhhh.github.io/Jinhui_profile/
Third-party components keep their own licenses.
Building this takes time; a few tokens for maintenance are appreciated!!!
Open the interactive support page / 打开交互式支持页
![]() WeChat Pay 微信支付 |
![]() Alipay 支付宝 |
![]() PayPal International support |
Donation supports maintenance only and does not grant commercial use rights.
Chart powered by star-history/star-history.
- Added the public
dpl.extensionABI 1.0 capability document, entry-point discovery, semantic workflow events, project-confined read grants, non-blocking extension dispatch, write-scope audit receipts, and extension status projection. - Optional commercial or learning sidecars now negotiate ABI capabilities instead of pinning one Core patch. Missing or failing extensions do not change the authoritative scientific command result.
- The public Core contains only neutral extension contracts; private learning content, licenses, customer state, and commercial packs remain outside the Community source tree.
- Completion classification now defaults to strict after a 60-case, three-discipline calibration fixture. Preview preserves the submitted class, inferred class, computed stale scope, evidence validation, and suggested evidence refs; apply revalidates every referenced manifest after preview.
- Result Support v3 selects metrics from current resolved evidence, the selected run manifest, or explicitly run-bound tables; binds downgrade/rescue decisions to one checkpoint hash; reports unbound required data roles; and routes post-Results evidence failures back to the same checkpoint.
- Source and wheel workflow skills, contracts, CLI reference, command risk matrix, schema registry, release manifest, bilingual README, and package identity are released as one v0.33.0 contract.
- v0.32.1 adds the capability truth matrix and keeps the established README H2 structure while making current-version capabilities, final-author completion, release order, and implementation evidence visible.
- v0.32.2 adds canonical completion classification, local evidence-ref suggestions, Result Support signal adapters, checkpoint-hash route handshakes, whole-checkpoint route boundaries, and synchronized Agent command examples in shadow-compatible form.
v0.31.1-v0.32.0 (2026-07-18) -- Architecture, Release Documentation, Cross-platform Identity and Completion Audit
-
v0.32.0reconciles the completeH-01-H-07,M-01-M-12andR01-R11release requirements. It validates one 210-command control plane in source, editable and isolated-wheel installations; runs General/AAS/MNRAS author-completion regression and five-domain adversarial release regression; restores Python 3.10 TOML compatibility; hardens optional bibliography-tool lookup; and publishes explicit completion and sandbox-boundary audits. These regressions validate workflow contracts rather than scientific conclusions or production-hosting isolation. -
v0.31.9verifies release identity across package metadata, release manifest, README, Git tag and the selected wheel. CI covers Windows and Linux on Python 3.10-3.12, adds a macOS control-plane smoke test, rejects stale distributions and validates minimal, plotting, full-text and MCP install profiles independently. -
v0.31.8adds the read-onlytoken-report, separates recorded token/cost receipts from estimates, generates the command risk matrix, and documents local-product, license and hosted-service boundaries. Monetary prices are never inferred when no provider receipt or explicit price table exists. -
v0.31.7generates the exhaustive CLI reference from the authoritativeCommandSpecregistry, so command names, handlers, risk levels and checkpoint metadata no longer need to be duplicated manually. The bilingual README keeps its established detailed-guide structure and links to the generated command reference instead of replacing the original project documentation. -
v0.31.6makes the default wheel a minimal workflow, bibliography and PDF/image-inspection runtime without silently installing NumPy, pandas or Matplotlib. Plotting, enhanced full-text extraction and MCP use explicit extras; scientific plugin NumPy imports are lazy, Doctor reports profile availability and recovery commands, and CI installs minimal/plotting/fulltext/MCP environments independently. A wheel metadata verifier prevents optional stacks from drifting back into core dependencies, while the packaged paper-fetch fallback remains available in minimal installs. -
v0.31.5registers schema families for release fixtures, capability packs, command contracts, release manifests, method run/formula manifests and figure assessments. Release validation checks packaged resource schemas instead of trusting structurally plausible JSON. CI raises first-party coverage from 45% to 65%, and Pyright now includes the command control plane, completion/stale paths, figure façade and split Methods modules. -
v0.31.4routes all 209 CLI commands available at that stage throughCommandSpec;legacy_dispatch_countis zero andcli.pyno longer contains a command-specific fallback chain. Commands awaiting direct typed handlers use an explicit namespace compatibility adapter whose count remains visible rather than being reported as completed migration. Pipeline-stage commands are generated from registry metadata, andassess-figure-contractsnow uses the unified figure façade. -
v0.31.3separates method execution verification, formula extraction and manuscript writing into responsibility modules behind the stabledraftpaper_cli.methodsAPI. A figure-contract façade normalizes gate, semantic, confirmed-blueprint and caption findings into one deterministic{source, code, severity, message}issue list; existing run-manifest strings remain available for compatibility while new consumers usenormalized_issues. -
v0.31.2replaces the 3,700-line plugin-candidate monolith with separate skill-source loading, capability extraction, guarded promotion and contribution-review modules. The establisheddraftpaper_cli.plugin_candidatesimports remain stable, including the audited registry hooks, while focused package tests prevent the monolith from returning and preserve extraction, AcademicForge metadata, provenance, promotion, fixture and contribution behavior. -
v0.31.1connects stable DPL claim and evidence IDs to the real claim-contract and Scientific Evidence Registry fallback paths instead of leaving them in a test helper. Explicit project IDs remain authoritative. Shared JSON/text IO, citation parsing and LaTeX escaping are consolidated for the active claim, evidence, reference and Data-writing paths, while the unused parallel dependency map has been removed from the artifact DAG.
v0.30.1-v0.31.0 (2026-07-18) -- Release Gates, Precise Stale Propagation and Author Completion Transactions
v0.30.1reconciles the public command inventory and release manifest without exposing the experimental manuscript-completion prototype. Core evidence, data quality, result validity, method verification and integrity gates now return nonzero process status for non-passing scientific decisions. MCP artifact reads reject private locators and credential-like files, and journal/registry metadata fetching is routed through an HTTPS, host, DNS and response-size policy.v0.30.2applies one project-local runner contract to manifest and CLI method verification, rejects inline Python and shell runners, confines declared inputs/outputs, records explicit system-binary opt-in, passes only allowlisted scientific environment variables and redacts process logs. Mutating commands now validate declared write roots before running and still verify the actual write set afterward. Literature retrieval distinguishessuccess_with_items,success_empty,provider_error,auth_required,rate_limitedandoffline_fallbackin stage manifests,statusanddoctor.v0.30.3replaces the three competing revision/stale vocabularies with one 13-class change taxonomy. Artifact DAG, external drift sync, section revision, author revision and reviewer repair now emit the same canonical names and derive stale impact from the same contract. Methods and Discussion edits start from their actual section stages, evidence-changing patches reopen the scientific chain, and revision preview/apply is bound to one promoted evidence snapshot.v0.30.4registersdpl.manuscript_completion.v1and addsprepare-manuscript-completionplus the read-onlymanuscript-completion-status. The generated author template covers structured metadata, funding, availability statements, short title, keywords and user-confirmed references; its journal report separates required, recommended, missing, placeholder and not-applicable fields. Completion input cannot directly replace scientific evidence, metrics, claims or figures, and later preview/apply commands remain unpublished until their own transaction gates pass.v0.30.5resolves author edits through stable paragraph ID, expected text/hash and optional occurrence while treating LaTeX line numbers as non-authoritative hints. Line drift is reported and re-anchored only when stable anchors still agree; stale, ambiguous, duplicate-key and duplicate-target packets are rejected as a whole. One packet can resolve multiple sections and user-confirmed references against one source-map/project revision, and project revision files can no longer read content from outside the project boundary.v0.30.6adds completion preview, apply and rollback as one hash-bound transaction flow. Preview produces a unified metadata/section/BibTeX diff, candidate LaTeX and candidate PDF without touching canonical sources. Apply rechecks the packet, project revision, source map, evidence snapshot and before hashes, then atomically writes metadata, references, sections, stale state, ledgers and exact-text user locks. Repeated apply is idempotent, injected failures restore the full write set, and rollback refuses to overwrite any artifact changed after completion.compile_required, unresolved Codex instructions and scientific evidence changes are explicitly non-passing.v0.30.7binds the applied completion manifest, manuscript metadata, every canonical section, BibTeX/reference registry, promoted evidence, result/figure manifests, final citation-audit snapshot, integrity/quality reports, two independent blind reviews andmain.pdfinto one release hash. Non-passing or stale bound artifacts block review and confirmation. The README opening now presents the current evidence-first path and final-author checkpoint directly, with complete English and Chinese completion guides underdocs/.v0.31.0completes multi-journal author-workspace regression with general article, AAS and MNRAS frontmatter semantics and real XeLaTeX candidate compilation. The release tests single and multiple authors and affiliations, ORCID, corresponding-author metadata, funding and repository links, user-confirmed references, multi-paragraph atomic edits, rollback, user locks and final release binding in both source checkout and isolated wheel installations.
v0.28.1-v0.30.0 (2026-07-17) -- Transactional Evidence Architecture and Cross-discipline Release Contracts
v0.28.1makes section revision a real multi-artifact transaction. Accepted edits are installed into canonical section sources, survive repeated LaTeX assembly, propagate real downstream stale state, and roll back canonical, candidate and project-state artifacts together on injected failure. This completes the transaction guarantee first introduced inv0.27.2.v0.28.2migrates all 210 discipline plugins to explicit manifest v2 contracts. Runtime class, validation level, maturity, deployment state, fixture inventory and execution policy are no longer silently inferred; source and wheel registries now agree on 546 fixtures and 175 contract-only, 25 code-generator and 10 fixture-executed runtime records.v0.28.3adds project-level state revisions, locking and rollback acrossproject.json, its YAML mirror and every stage manifest. Write-set violations restore project-local changes where possible, and MCP science/network execution requires a short-lived HMAC capability token bound to the exact project, command and arguments.v0.29.0upgrades the artifact dependency model to real paths, hashes, owners, producer/consumer edges and authoritative/derived roles. Product release numbers are removed from first-party schema IDs and replaced by a versioned schema-family registry.v0.29.1introduces one normalized runtime command-contract view for all 204 CLI commands and makes it authoritative for consistency validation, MCP schemas and doctor diagnostics. The report exposes the migration boundary explicitly: 80 registered handlers and 124 compatibility dispatches remain, so future decomposition is measurable rather than hidden.v0.29.2adds targeted Ruff and pyright gates, coverage, project-scoped dependency auditing with CycloneDX SBOM output, secret scanning, pinned GitHub Action SHAs, reproducible CI tool constraints, source/wheel release identity, SPDX-style license metadata and clean wheel manifest rules.v0.30.0expands the wheel-installable release matrix across scientific-image+ML, geography+ML, astronomy+ML, bioinformatics+medicine and physics/quantum projects. Each fixture now carries six main figure groups through claim, data plugin, method plugin, project run output, evidence ID and review-rule trace, while adversarial cohort/run/unit/model/metric/dimension, blank-figure, plugin and citation cases must still be rejected.
Release checks:
draftpaper validate-command-contracts
python -m draftpaper_cli.release_contract
python -m draftpaper_cli.release_regression --output tmp/release-v0300v0.26.1separates stale artifacts, citation support, manuscript semantics, reproducibility packages, render quality, scientific analysis, figure contracts, and human review into explicit failure domains. Citation or bundle failures no longer fall back toplan-figures; hard-check reports name the failed predicate, artifact, and valid next command. Interactive CLI output is compact, while redirected/API output preserves complete JSON;--compactand--json-fullselect either contract explicitly.v0.26.2adds domain-neutralDataProvenanceContract,CohortRegistry,CohortViewRegistry, and append-only join/filter ledger artifacts. Every analysis view declares its parent cohort, sample unit, count, missingness policy, split, allowed use, and claim boundary. Distinct display and regression subsets may coexist, but silent count or sample-unit reuse is rejected.v0.26.3addsExecutableAnalysisSpec, formula AST, resampling contract, and prespecified run-selection policy. Estimand, cohort view, split, preprocessing scope, implementation entry point, uncertainty semantics, calibration definition, formulas, and variable explanations now share one contract. ECE definition drift, false paired/grouped intervals, and unlocked best-seed headline selection are blocking semantic errors.v0.27.0splits plugin sufficiency into bootstrap and release decisions. Auditable project-local implementations can unblock a new topic before formal reusable plugins exist; final release still requires verified data/method bindings, input/output hashes, execution scope, and a capability passport. AcademicForge/GitHub rescue runs only for a real unresolved capability gap.v0.27.1upgrades the Scientific Evidence Registry key to estimand/cohort-view/analysis-spec/run/model/split/aggregation/dimension and writes section claim maps with sentence hashes and evidence IDs. Figure Contract v2 binds claims, panels, cohort views, estimands, analysis specs, runs, evidence and rendered text; journal-width QA records font, overlap, crop, color and caption checks. Post-Results discipline rules consume compiled claim-level evidence rather than broad role strings.v0.27.2adds an artifact dependency DAG and transactionalapply-section-revision. Citation-only, presentation-only, prose-semantic, cohort, analysis, run, figure and claim-contract changes propagate to different minimal stale sets. A current functional section release is authoritative over optional manifest backfills, while changed accepted hashes still reopen the correct chain.v0.27.3builds anonymous review supplements from the selected run and Python import dependency closure, excludes unrelated historical figures/tables, compiles the closure before release, and archives reviewer reports under the superseded bundle hash when the manuscript, evidence or bundle changes. Review findings carry structured repair classes for prose, rendering, reproducibility, analysis supplements, claim downgrade or new evidence.v0.28.0adds normalized citation intents for direct support, method/tool provenance, comparison context, and dataset/product provenance while retaining all curated references and repairing prose first. Paragraph evidence remains content-addressed and delta-cached. Five cross-domain release fixtures plus adversarial cohort, calibration, run-selection, figure, plugin, and citation regressions guard the generic architecture.- Added
JournalIntentContract: only a confirmed journal may render a submission label; provisional and unset projects remain neutral drafts. Status distinguishesdraft_pdf_ready,review_blocked, andrelease_ready, anddoctor --explainexposes the read-only artifact DAG and recovery rationale.
Key commands:
draftpaper doctor --project <project> --explain
draftpaper apply-section-revision --project <project> --section methods --input revised_methods.tex
draftpaper prepare-independent-manuscript-review --project <project>
draftpaper status --project <project> --compactv0.25.1-v0.26.0 (2026-07-14) -- Confirmed Research Blueprints, Task-aware Statistics, and Project Workspace Isolation
- New papers now default to the configured central
projectsroot instead of following the idea or dataset location.DRAFTPAPER_PROJECTS_ROOT, user config, and--projects-rootcontrol the root; generated slugs are capped at 48 characters and use an eight-character stable project ID. Clean_vNchildren use the same short-name policy. - Added
ProjectWorkspacePolicy,ArtifactOwnershipGuard,path-budget-check,doctor-project-layout, and hash-awareadopt-orphan-artifacts. Managed outputs, logs, and temporary files stay inside the paper project unless an explicit export is requested. - Large datasets remain read-only in place. Public source contracts and fingerprints use logical IDs and relative locators, while machine-local absolute paths move to ignored
external_data_locators.private.json; the default copy policy ismanifest_only. - Added task-aware
StatisticalValidationContractandreview_rule_coverage_report. Classification, regression/fitting, grouped or temporal validation, spatial analysis, representation confounding, anomaly stability, survival, and simulation convergence receive different evidence requirements. Universal F1, R2, p-value, or fit thresholds are forbidden unless a user/journal, cited domain source, or validated discipline plugin supplies them. - The Chinese-first research-plan packet is now a real human checkpoint.
review-research-planpresents the plan, claims, storyboard, panel structure, method/data requirements, statistical contract, plugin preview, and feasibility limits;confirm-research-planfreezes their exact hash.reopen-research-planis mandatory before any scientific contract change. - Key-figure code can consume only the active confirmed blueprint. Every main figure and panel carries the confirmed plan hash, claim, data and method requirement IDs, and statistical validation IDs.
validate-confirmed-figure-alignmentrejects missing, substituted, extra, or semantically changed main figures. - Removed discipline fallback figures from explicit research storyboards. Main-figure count can no longer be satisfied with generic histograms, correlations, or metric summaries; missing scientific story roles return to research-plan revision instead of producing filler.
- Added structured figure-caption contracts. The first sentence summarizes the complete figure group without comma-linked fragments; later sentences explain every panel, cohort, estimand, uncertainty definition, and claim boundary. Short
fig_01.png-style paths keep scientific titles in metadata rather than filenames. - Added the pre-execution support loop. Capability gaps first produce project-local, discipline-plugin, AcademicForge, GitHub, and connector rescue tasks. If rescue remains insufficient, the user chooses data/method supplementation or research-scope downgrade before key-figure code is generated. Post-result claim downgrade remains separate and freezes existing figures and metrics.
- The user-facing core-evidence page is now “Key Results and Claim Support Confirmation” and explains the research question, cohort, method run, statistical evidence, uncertainty, and maximum claim strength being approved. Final release adds a separate hash-bound packet containing
main.pdf, final citation audit, both independent blind-review reports, and the quality report. - The wheel regression now covers scientific-image astronomy+ML, geography+ML, time-domain astronomy+ML, bioinformatics/medicine, and physics/quantum, with task-aware statistical contracts and explicit rule-gap checks. Current source verification is
565 passed; the complete 15-item evidence matrix is indocs/audits/2026-07-14-v0260-acceptance-report.md.
Key workflow additions:
draftpaper create-project --idea "Your research idea" --field "astronomy machine learning"
draftpaper review-research-plan --project <project>
draftpaper confirm-research-plan --project <project> --plan-hash <hash>
draftpaper path-budget-check --project <project>
draftpaper doctor-project-layout --project <project>- Packaged the canonical
draftpaper-workflowskill inside the wheel and addedinstall-skill/skill-doctor. Agent instructions now defer stage order tostatus,verify-next-action, andcontinue, so an old user-level skill cannot silently drive a new CLI through an obsolete workflow. - Upgraded all 186 public commands to CommandSpec v2 metadata: risk class, read/write globs, forbidden paths, resource class, timeout, idempotency, confirmation policy, and typed input/output schemas. Mutating project commands are checked against their actual post-run write set; parent traversal, UNC/device paths, symlink/reparse escapes, protected actions, arbitrary shell/raw writes, SQL, Git push, and credential disclosure are excluded from the MCP surface.
- Added Plugin Execution Contract v2 and deterministic
plugin_catalog_snapshot.json. All 209 plugin manifests receive a valid explicit or compatibility-adapted execution contract, whilecatalog_hashand per-plugin contract hashes flow through sufficiency, binding, execution, figure trace, and discipline review. Contract/fixture-only or mock API/GPU/server plugins still cannot support a scientific result. - Promoted 22 high-frequency local data/method capabilities and 10 review rules through an explicit runnable-profile registry. Each promoted capability executes a deterministic scientific algorithm with minimal, failure, and boundary fixtures; review rules include an applicability boundary and threshold source. Remaining foundation rules stay advisory and external contracts stay mock-only until real validation.
- Made research-plan figure story roles explicit, removed the formal-writing first-eight evidence fallback, and added claim/run/cohort/model/figure/citation-role retrieval. Paragraph evidence is content-addressed, cached, and delta-aware; the held-out repeated-evidence regression reduces actual paragraph input by at least 35% while preserving every evidence ID.
- Added Citation Role Contract v2 (
dataset_provenance, instrument/product definition, processing support, method/tool background, prior-result comparison, mechanism/interpretation, and general background). Roles are assigned before writing; citation audit validates their use afterward and retains the established rewrite-before-delete policy. - Integrity now prefers promoted, run-aware Scientific Evidence Registry counts and reports legacy unbound
sample_composition.csvas a compatibility source. Main-figure narratives explicitly reserve direct scientific signal, comparison, mechanism/ablation, and uncertainty/boundary roles instead of allowing diagnostics to occupy every main slot. - Added
workflow_trace.jsonlandaudit-workflow-runtimewith run/command IDs, attempts, duration, input hashes, process/command/scientific/transaction outcomes, loop detection, duplicate-run detection, and packet metering. Added SQLite-backedsubmit-job,job-status,job-cancel,job-notifications, andrecover-jobs; workers survive the initiating terminal and lost workers becomeorphanedrather than silently retried. - Added a ten-tool local stdio Draftpaper MCP (
python -m draftpaper_cli.mcp.server) with portablemcp-installand deterministicmcp-doctor. It is a thin projection over CommandSpec and CLI handlers, provides bounded artifact selectors, and cannot directly execute human checkpoints or destructive administration. - Expanded wheel-shipped held-out release regressions to astronomy+ML, Euclid-context geography+ML, bioinformatics+medicine, and previously unused physics+quantum capability chains, plus adversarial scope, figure, plugin-runtime, and citation checks. The machine-readable closure ledger is
docs/audits/v025_issue_ledger.json; the 41 AcademicForge collection-level placeholders remain explicitlyrequires_source_inspectionand do not enter sufficiency. - Final source verification: all
555tests pass, including the explicit four-domain wheel-verifier contract regression. Package compilation, isolated v0.25.0 wheel installation, canonical skill hash, 209-plugin catalog, 32 runnable profiles, ten-tool MCP stdio handshake, four held-out domains, and all adversarial release checks pass.
Key setup and diagnostics:
python -m pip install -e .[mcp]
draftpaper install-skill --force
draftpaper skill-doctor
draftpaper mcp-doctor
draftpaper mcp-install --output .mcp.jsonv0.23.1-v0.24.0 (2026-07-13) -- Project Isolation, Native Recovery, and Independent Manuscript Review
- Added clean
_vNproject versioning. The parent stays read-only; only allowlisted reusable assets are imported with hashes and lineage receipts. Passports, stage manifests, sufficiency reports, figure metadata, audits, snapshots, and other derived state are rebuilt in the child instead of being copied as active evidence. - Added a project system of record, command transactions, stage receipts, exact stale propagation, and rebuild plans for derived artifacts. Results lifecycle repair now rebuilds the current result manifest before discipline review, preventing an old review from being applied to changed prose or evidence.
- Added capability packs, routing evaluations, ownership checks, project-local data/method audits, and project implementation contracts. Missing formal plugins no longer create a "method code cannot be generated until the method already exists" deadlock; only an exhausted data/method rescue route can block scientific figure production.
- Added run-aware Figure and Paragraph Evidence Resolvers with task budgets. Doctor evaluates the latest packet for each writing task while retaining lifetime token cost separately, so historical retries do not leave a repaired project permanently in warning state.
- Added a canonical reference registry, work/version duplicate resolution, journal bibliography contracts, and rendered reference proofs. Bibliography formatting is now checked separately from claim-level citation support, and citation repair continues to preserve retained references.
- Added recursive third-party provenance for paper-fetch-skill, AcademicForge, upstream scientific skill repositories, Gbrain design influence, and Superpowers planning influence, including pinned commits, licenses, use modes, and affected paths. Gbrain remains optional design context only.
- Added native Draftpaper-loop
doctor,recover,start,continue,review, andrevisecommands, with all CLI parsers registered under oneCommandSpeccontract. Doctor is deterministic and read-only; recovery never silently accepts figures, downgrades claims, deletes references, or promotes plugins. - Replaced manuscript A/B and original-manuscript quality ratios with a single-manuscript independent review contract. Two reviewers inspect the same frozen anonymous generated manuscript and real figures in independent sessions; release requires no unresolved critical or major findings, and disagreements enter adjudication rather than score averaging.
- Added a post-review manuscript revision workspace with stable paragraph IDs and LaTeX line anchors, structured metadata, custom references, diff/PDF preview, hash-checked acceptance, rollback, and precise stale effects.
- Added locator-safe reproducibility supplements to anonymous review bundles: source/model provenance, safe stage-owned scripts and tests, and dependency-free frozen-output replay can be reviewed without exposing private paths, credentials or identity. Caption corrections can be applied through structured metadata without rewriting scientific figure evidence.
- Added model-comparison semantics to section packets. Writers may describe an incremental or conditional contribution only when preprocessing and model nesting are verified; otherwise the manuscript must report an exact pipeline-performance contrast and keep fold-mean, pooled and resampling estimands distinct.
- Completed general software regressions for astronomy+ML, geography+ML, and bioinformatics+medicine capability chains, then completed an unfamiliar scientific-image manuscript from read-only local data through human core-evidence confirmation, final writing, PDF, integrity, bibliography and citation gates. The final citation audit retained all 17 references across 30 usages with no unsupported or unverifiable use.
- Two fresh independent sessions reviewed the same frozen anonymous final manuscript without an original or prior report. Both returned only minor revisions, scientific-correctness scores were 0.95 and 0.91, unresolved critical/major counts were zero, no adjudication was required, and the release gate passed. Minor recommendations remain visible in the revision queue rather than being silently marked resolved.
- Final verification:
537 passed; package/tests compile cleanly; the isolated v0.24.0 wheel matches 209 plugins, 545 fixtures, six capability packs and six third-party source lineages; all three domain regressions and every adversarial release check pass. GitHub Actions is green for wheel installation and the complete Linux/Windows Python 3.10-3.12 matrix. Eval capture/replay is structural and privacy-preserving and never treats an existing manuscript as a quality baseline.
- Added wheel-shipped, de-identified regression contracts for geography+machine learning, astronomy+machine learning, and bioinformatics+medicine. An ordinary
pip installnow exercises plugin sufficiency, project execution evidence, figure traceability, rendered-pixel and source-table checks, the Scientific Evidence Registry, Results binding, composite review rules, and the real final citation producer. - Added adversarial regressions for wrong runs, cohorts, units, splits, models, metrics or dimensions, forged metadata on blank figures, contract/fixture-only plugins, citation negation, numeric mismatch, and causal reversal.
statusmust also leave every project-file hash unchanged. - Automated keywords, metadata, and file-presence scores can no longer authorize a quality claim. The provisional v0.23.0 comparison protocol is superseded by the v0.24.0 single-manuscript independent-review contract; current projects do not require or receive an original manuscript.
- Isolated wheel verification requires both source and installed registries to discover 209 plugins and 545 fixtures, then runs all three domain and adversarial regressions inside the installed environment. CI also covers Python 3.10-3.12 on Linux and Windows.
- Completed an evidence-first astronomy+machine-learning full run through research contracts, project-local capability rescue, six scientific figure groups, all five free-writing/editor lifecycles, final PDF assembly, and citation re-audit. The final citation snapshot retained all 12 summarized references across 25 usages with zero unsupported or blocking usages.
- The rendered technical comparison found that the generated manuscript preserved the verified cohorts, baseline/Transformer metrics, ablation direction, uncertainty boundary, and six-figure story. Introduction and Discussion were comparable or stronger in structure, while direct observational examples and Data/Methods/Results detail remained thinner than the reference manuscript. The automated functional score was 0.9825, but no formal 95% claim is made without the required independent blind reviews.
- Observable free-writing usage, measured with
tiktokeno200k_base, was 122,238 section-packet input tokens and 5,716 accepted-candidate output tokens (127,954 total): Results 40,075; Introduction 15,691; Data 17,139; Methods 22,920; Discussion 32,129. Deterministic CLI stages used no model tokens; hidden reasoning and platform overhead are not estimated. Seedocs/audits/2026-07-12-astronomy-v0230-full-run.md. - Full local verification after the run:
441 passed. The final quality gate has no unresolved scientific, writing, figure, evidence-snapshot, PDF, or citation error; it remains intentionally unreleased until the required two-reviewer blind comparison is recorded.
- Fixed wheel resources and citation-audit/parity schemas. Formal writing now follows section packet -> free Codex composition -> candidate validation -> Scientific Editor -> explicit acceptance -> release; deterministic fallback remains diagnostic-only.
- Evidence Binding v2 binds quantitative claims to
evidence_id/run/cohort/sample_unit/split/model/metric_dimensionand verifies metric identity and dimension. Wrong scopes or same-value semantic substitutions block writing. - Plugin runtime truth now distinguishes
contract_only,code_generator,fixture_executed,project_validated, andlive_validated; only hashed project or live outputs can satisfy a main-figure capability. - Reviewer Engine v2 consumes a standard EvidenceBundle and executes real thresholds, dimensions, baseline, ablation, uncertainty, and plugin
evaluate_rulechecks. Scientific anomalies route first to local Results semantic repair. - Figure checks inspect rendered pixels, axis/text regions, panels, and source tables. Citation audit adds passage, numeric, negation, causal-direction, and claim-strength checks while preserving references and emitting paragraph-rewrite tasks.
- Added a read-only state kernel, explicit migration, atomic JSON/JSONL writes, cross-platform locks, and separate command registry, artifact repository, schema adapters, plugin runtime, writing coordinator, and release coordinator boundaries.
- Added the discipline-neutral Paper Narrative Engine. It reads true YAML/JSON research artifacts and produces
paper_brief.json,figure_story_arc.json,manuscript_argument_map.json, andsection_claim_allocation.json; main and supporting figures are grouped into finding-level stories with explicit questions, evidence, comparisons, and claim boundaries. - Added section evidence packs, paragraph outlines, and Results synthesis plans. Introduction and Discussion receive gap/comparison matrices; Data and Methods receive lifecycle reconstructions from inventories, stage-owned code, plugin ledgers, formulas, and figure-code traces. Metrics bind through evidence, run, or artifact identifiers rather than title-string guessing.
- Added semantic panel contracts and
prepare-panel-repair. Each panel declares its question, subset, scientific unit, data roles, method output, comparison, statistical check, chart grammar, expected conclusion, and claim boundary; repair remains local to the failed panel chain and cannot substitute weaker evidence. - Added venue writing contracts and functional style profiles that learn section order, information density, caption density, voice, numeric reporting, terminology, and reasoning function without copying wording from a reference manuscript.
- Added a bounded paragraph-local Scientific Editor with at most three auditable iterations. It records paragraph hashes, rejects excessive whole-section churn, preserves evidence and citations, and forbids reference deletion as a shortcut for citation-audit repair.
- Replaced marker-count parity scoring with calibrated scientific dimensions. A release requires validated
codex_free_candidateprose for all core sections, one consistent evidence snapshot, passed Results and figure quality, no blocking evidence conflicts, preserved reference coverage, and a citation audit generated after the final validated draft. Hard scientific correctness remains mandatory; functional quality must reach at least 0.95. - Added cross-disciplinary architecture regressions for geography+machine learning, astronomy+machine learning, and bioinformatics/medicine to verify generality without embedding fixture-specific models, metrics, figure counts, or prose in generic code.
- Added
results_narrative_contract.json, which assigns distinct scientific jobs to main figures: study boundary, pre-model signal, model comparison, component attribution, and error/uncertainty. Each role is bound to run-consistent metrics, its scientific question, and claim boundary, while Codex retains free prose composition. - Added
prepare-section-writingandassess-manuscript-quality. Results are scored for evidence fidelity, narrative coverage, scientific reasoning, claim calibration, and prose diversity; a formal candidate must reach 0.95, while wrong metrics and repetitive template prose enter repair. - Added
assess-figure-publication-quality. A non-empty PNG is no longer sufficient: main figures must satisfy semantic contracts, data/method plugin run provenance, panel completeness, statistical interpretation, pixel dimensions, and legibility. - The generic plotter no longer silently replaces an unknown main-result figure with a data overview. Missing main-figure methods enter plugin rescue; generic diagnostic views remain available only when explicitly planned.
- Added
prepare-results-semantic-repairandassess-paper-quality-parity. The former repairs only affected claims while preserving verified figure narratives; the latter aggregates figures, Results, Introduction, Data, Methods, Discussion, and citation audit, and finalquality-checkcannot pass below 0.95.
- Figure contracts no longer execute discipline
review_rulegates before plotting. Data and method plugins produce figures; only figures with actual plugin traces activate matching discipline rules inreview-results-with-discipline-rules. - Untraceable metrics, internal artifact language, misplaced citations, and missing figure interpretation now enter
repair_required, which prioritizes rewriting or narrowing Results claims instead of treating prose defects as figure-generation failures. - Capability gaps first enter
rescue_requiredand are checked against project-local assets, the registered plugin catalog, AcademicForge, and GitHub research code. The newrecord-plugin-rescue-outcomecommand permitsblocked_unavailableonly after all four routes have auditable search evidence and the required data or method capability still cannot be found.
- Added
audit-project-capabilitiesbetween plugin sufficiency and external rescue. It audits stage-owned local data and method assets, records privacy-safe relative evidence paths plus hashes, and creates constrainedcovered_project_localbindings only for the active project. These bindings never modify global discipline modules or bypass candidate validation and explicit promotion. - Results discipline review now audits manuscript prose as well as figures: it detects untraceable metric claims, internal artifact language, citations used as result evidence, missing figure interpretation, incomplete plugin/run traces, and review-rule evidence conflicts.
- Data and Methods writing contexts now include sanitized descriptions of bound data and method roles, so reusable plugins and audited project-local implementations can improve scientific exposition without exposing paths, commands, manifests, or credentials.
- Added a final
discipline_contract.jsonandresearch_capability_contract.jsonimmediately after research planning. They declare the primary and secondary disciplines, cross-discipline data/method/review ownership, and stable requirements for each planned claim and main figure. - Added
assess-plugin-sufficiency. It performs structured matching against registereddata_connector,method_template, andreview_rulemanifests using roles, method families, outputs, runtime class, validation level, aliases, and discipline compatibility. A mock, plan-only, or external contract never counts as executable support for a main figure. - Added
prepare-plugin-rescue. A capability gap becomes a scoped route through existing plugins, AcademicForge candidate processing, license-aware public research-code discovery, generalization, validation, de-duplication, and explicit human-confirmedpromote-plugin-candidate --write. Project-specific or license-uncertain code remains project-local. - Added
execute-data-pluginsandexecute-method-plugins, plus stage-ownedplugin_execution_ledger.jsonlrecords. Manifest/template hashes, parameters, runtime state, input/output hashes, and fixture-versus-project-result status are preserved; fixture execution is explicitly not treated as scientific evidence. - Added
results/figure_plugin_trace_report.json/.html. Each main figure must trace to a research-plan claim, covered data plugin, covered method plugin, review-rule route, and a verified project run output before Results evidence is accepted. Code generation may proceed only from an explicit pre-run binding chain; missing links route to plugin rescue rather than a substitute figure. - Added
review-results-with-discipline-rulesafterwrite-results. It combines Results prose, complete figure traces, plugin bindings, verified run outputs, and composite-discipline review rules. Only mature, promoted, evidence-bound rules may block scientifically; trace gaps always stop downstream manuscript writing. statusandrun-pipelinenow exposeplugin_sufficiency_requiredandplugin_gap_detectedstates. Cross-discipline regression fixtures cover geography+machine learning, astronomy+machine learning, and bioinformatics+medicine chains.
- Expanded the first-party local foundation catalog to 159 parameterized plugins across statistics, experimental design, classical ML, tabular/array processing, scientific visualization, geography, astronomy, bioinformatics, medicine, chemistry, materials science, physics, and quantum science. Every plugin is placed only under its discipline's
data_connectors,method_templates, orreview_rulesdirectory. - Split previously combined capabilities into independently selectable contracts, including statistical analysis/power/design, Polars, Vaex, Dask local mode, Zarr, Matplotlib, Seaborn, NetworkX, GeoMaster-style local coordinate operations, Astropy FITS/WCS/Time/units/coordinates, Bioinformatics CPU workflows, Pydicom/BIDS/NeuroKit2/survival analysis, Molfeat, FluidSim CPU, and Cirq local simulation.
- Added 12 external foundation contracts for API, remote-server, and GPU routes. They use
mock_validatedfixtures and an explicittask_contract; every one recordslive_execution_performed: false, does not fetch data or contact external services, and surfaces required credentials as user-confirmed inputs. - Real external validation is deliberately project-specific: validate an API, SSH server, or GPU model only when a user-authorized paper project needs it, then upgrade that individual plugin's validation level with provenance instead of claiming a globally live capability.
- Discipline plugin manifests now participate in runtime module registration. Adding a valid directory under
discipline_modules/<discipline>/{data_connectors,method_templates,review_rules}/<plugin>/automatically extends that discipline'sDisciplineModuleSpec; a manifest-only discipline is available without adding a staticmodule.py. - Added explicit plugin execution metadata:
runtime_classdistinguisheslocal_pure_python,local_optional_dependency,remote_api,remote_server,gpu_model,laboratory_hardware, andsupport_only;validation_leveldistinguishesplan_only,mock_validated,fixture_runnable, andlive_validated. A local template never claims that an unavailable package, remote service, cluster, or GPU job has run. - Added 80 first-party, parameterized foundation plugins with
template.py, manifest, and normal/failure/boundary fixtures. They cover local statistics and classic ML, tabular and array processing, scientific visualization, Astropy FITS/WCS/Time/units/coordinates, GeoPandas vector/CRS work, and five new foundations: chemistry, materials science, physics, quantum science, and neuroscience. - Each new discipline starts with 3 local data connectors, 5 method templates, and 5 advisory evidence-bound review rules. Review rules remain contextual candidates until discipline evidence, fixtures, and explicit promotion justify a blocking threshold.
- Candidate promotion now writes one canonical
manifest.jsonnext totemplate.py, fixture files, and provenance. That manifest is immediately consumable by runtime auto-registration; overlapping candidates becomeaugment_existingoverlays that merge aliases, variants, fixture references, and provenance instead of silently creating duplicate plugin contracts.
- Added a metadata-only third-party skill conversion pipeline:
snapshot-skill-source,inspect-skill-source,index-skill-source,classify-skill-source,map-skill-capabilities,extract-skill-capabilities, andcompile-skill-source. It turns AcademicForge-style skill catalogs, personal research skills, or project-distilled skills into candidate reports without copying source code or installing plugins directly. - AcademicForge collection records are expanded against public classification metadata and reconciled with each collection's declared skill count. Detailed records and unresolved placeholders are both indexed, with an explicit silent-loss count; a live metadata-only check currently reconciles all 340 declared skills as 299 detailed registry/classification records plus 41 source-inspection placeholders.
- Clarified that formal discipline modules accept only three promotable subplugin types:
data_connector,method_template, andreview_rule.workflow_recipe,paper_contract, andshared_capabilityremain support-layer candidates rather than being written intodiscipline_modules/<discipline>/. - Added
review_rule_signal_scanand support-layer backflow records so statistical validation, model baseline/ablation, split/leakage, figure-claim consistency, citation support, data/code availability, and reproducibility conditions from workflow, paper-contract, and shared-capability skills can become discipline-specificreview_rule_candidates. - Expanded
ReviewRuleSpec: review rules are evidence-bound scientific quality gates with explicit discipline, method family, data role, evidence binding, criterion type, threshold mode, threshold validation status, support-layer signal refs, fixture refs, aliases/variants, backflow source type, threshold source, failure route, maturity, and human-confirmation metadata. Model-score, goodness-of-fit, and statistical-significance thresholds default to contextual/comparative/human-confirmed rules unless backed by journal guidance, discipline convention, public benchmarks, or user confirmation. - Added runtime review-rule gates.
prepare-method-blueprint, Semantic Figure Contract assessment,assess-result-validity,assess-result-support, and citation audit now write review-rule gate reports so promoted, runnable/mature, evidence-bound rules can block weak or unsupported claims while candidate/foundation rules remain advisory. - Runtime gates now consume
evidence_binding.required_fieldsandevidence_binding.forbidden_conflicts, so promoted rules can block missing evidence or contradictions such as train/test leakage while support-layer and candidate rules remain non-blocking until reviewed. - Review-rule gate reports now include
rescue_tasksandrecommended_next_commands, routing failures to data rescue, method rescue, result downgrade, manuscript repair, citation repair, or human checkpoints instead of leaving weak rules as plain reviewer text. - Added
assess-review-rulesfor direct inspection of the active discipline review rules at stages such asmethod_plan,assess_result_validity,result_support_checkpoint, andcitation_audit. - Added package/preflight/provenance safeguards: contribution packages contain only generalized templates, fixtures, validation reports, and provenance/backflow summaries. Support-layer candidates cannot be submitted as formal discipline plugins directly; contributors should submit the generalized and tested formal candidates that backflow from them.
- Generalized templates now retain source identifiers and matched signal metadata only; bounded source excerpts stay outside contribution packages and are not embedded back into generated Python templates. Review-rule validation checks the positive/negative synthetic fixture contract, while promotion additionally requires runnable-or-higher maturity and explicit human confirmation.
- Added
review-plugin-contribution, a read-only maintainer review helper that summarizes package preflight status, metadata-only source policy, support-layer backflow families, threshold and human-confirmation policy, files to review, and whether the package is ready for human review before any promotion.
-
Added
research_plan/claim_contract.json, generated with the research plan. It records planned claims, active claims, claim strength, linked figures, evidence roles, and claim boundaries so later writing follows an explicit scientific contract instead of silently overclaiming from weak figures. -
Added
apply-result-downgrade. When current verified figures and metrics are scientifically usable but weaker than the original research plan, this command freezes the existing result artifacts inresults/result_evidence_freeze.jsonand a versionedresults/evidence_snapshots/result_freeze_*.json, downgrades only the active claim boundary, and does not rerun data, methods, figures, or metrics. -
Added
prepare-result-rescue. When the user wants to keep the stronger claim, this route prepares connector-aware data supplement tasks, method supplement tasks, and open-source research-code discovery tasks through the discipline plugin layer, then reopens the data/method/figure/evidence/manuscript chain for regeneration and validation. -
Added decision-aware stale propagation. Downgrade stales manuscript and claim-boundary consumers while preserving current result evidence; supplement stales data, method planning, figure contracts, code, method verification, result validity, core evidence, results, and downstream manuscript stages.
-
Updated
statusandrun-pipelineso failed result support becomes a clear human decision point with the two executable routes:apply-result-downgradeorprepare-result-rescue. The pipeline stops there instead of continuing into manuscript writing or blind quality checks. -
Added
assess-result-support, a scientific support checkpoint betweenassess-result-validityandassess-core-evidence. Technical result validity still checks outputs, metrics, and figure execution, while result support checks whether the current evidence can actually sustain the planned research claims. -
Added
results/result_support_checkpoint.json/.md/.html. When figures or metrics only partially support the research plan, the report stops manuscript writing and records two user-facing routes: downgrade the research claims to the current evidence, or supplement data/method evidence and regenerate core figures. -
Updated the staged pipeline so
statusandrun-pipelinestop atresult_supportwhen the support decision fails. Existing failed checkpoints also blockwrite-results, preventing the loop from turning weak or contradictory evidence into manuscript claims.
v0.17.0-v0.17.7 (2026-07-08 to 2026-07-10) -- Evidence-Semantic Manuscript Safeguards and Freer Writing
-
Replaced the permissive fact-ledger writing path with a domain-neutral Scientific Evidence Registry. Evidence records carry role, cohort, sample unit, split, run, model, source, and confidence so contradictory cohort facts block writing instead of becoming prose hallucinations.
-
Added run-aware result-evidence resolution. Metrics are selected from verified method runs and explicitly anchored sibling tables rather than from an arbitrary generic
metrics.csv; Results, Methods, validity checks, and figure metadata share evidence identifiers and provenance. -
Added Semantic Figure Contracts. Main figures now declare scientific question, variable roles, forbidden roles, method outputs, panel structure, plot grammar, metric dimensions, and expected claim. Identifier-versus-identifier plots, mixed-unit axes, missing method outputs, and incomplete panels are rejected before core-evidence confirmation.
-
Added precise change classification and stale propagation. Citation-local and presentation-only changes no longer invalidate evidence generation, whereas data, method, result, figure, and research-design changes reopen the necessary scientific chain.
-
Added human-approved evidence snapshots and explicit reopening. No writer, LaTeX assembly, or final citation audit can silently mix a changed figure, metric, or manuscript section with an older evidence approval.
-
Preserved Codex freedom at sentence level through section evidence packets and post-writing contracts. Freely composed section candidates are accepted only when they respect result-citation, evidence coverage, formula explanation, internal-language, and result-leakage rules.
-
Added
submit-figure-semantic-annotationsfor explicit, auditable legacy-figure mappings; the loop does not infer scientific meaning from an old PNG. Cross-project v0.18.0 release regression remains a separate forthcoming validation step. -
Added
writing/scientific_fact_ledger.json, a shared manuscript-facing ledger for must-preserve scientific facts such as sample sizes, class balance, token coverage, stress-test boundaries, and result metrics. -
Connected the fact ledger to Data and Methods writing briefs, Data/Methods prose, Discussion comparison, and the final quality gate, so cleaning internal paths or raw fields no longer silently removes key scientific facts.
-
Added
results/figure_interpretation_blueprint.jsonso Results writing is driven by main figure groups, scientific questions, primary metrics, claim boundaries, and appendix diagnostics rather than generic artifact summaries. -
Added relevance filtering for Introduction, Data, and Discussion citation insertion, preventing same-discipline but weakly related literature from being written into inappropriate manuscript positions.
-
Extended quality checks with must-preserve fact coverage, and verified the astronomy regression sections against the current
main.pdf: v0.17.0 keeps Results citation-free and subsection-free, increases formula and figure-reference coverage, removes path/raw-field leakage, and preserves the current evidence ledger's key data facts.
v0.16.1-v0.16.9 (2026-07-07 to 2026-07-08) -- Manuscript Writing, Figure Contracts, and Regression Hardening
-
Added the
learn-writing-style-from-draftpath andwriting_style.pyprofile so approved drafts can provide non-verbatim style signals without weakening evidence gates. -
Completed a full astronomy regression on the local time-aware flaring-source project: refreshed plan/data/method/figure stages, verified methods, compiled the APJS/AAS PDF, passed the integrity gate, and passed the final quality gate.
-
Confirmed the regression keeps 6/6 main figure contracts satisfied, records 13 rendered scientific figures, excludes 4 unrendered supporting figures from the result manifest, preserves 12/12 BibTeX references in the assembled manuscript, and keeps Results free of citations and subsections.
-
Reworked Discussion generation so filesystem artifacts, figure paths, table paths, manifest names, and Draftpaper-loop implementation language are sanitized before they can enter manuscript prose.
-
Added discussion comparison preparation and citation-evidence coverage so Discussion can compare results with literature while preserving the post-writing citation audit principle: repair weak claims and placements, do not delete confirmed references.
-
Added regression coverage for discussion artifact sanitization and citation-evidence expansion.
-
Upgraded Data writing to preserve scientific detail such as sample roles, class balance, token coverage, modality availability, and claim boundaries while keeping paths, filenames, script names, and raw field dumps out of manuscript text.
-
Upgraded Methods writing to use method-stage manifests, extracted formulas, formula-variable explanations, and figure-code traces, so Methods prose is organized around sample construction, feature/token construction, model logic, validation, metrics, and ablation evidence.
-
Added AASTeX-safe LaTeX assembly fallbacks for author metadata, table rendering, and missing local bibliography styles so local review PDFs can compile reliably.
-
Rewrote Results generation around
results/result_manifest.yaml, figure metadata, metrics, captions, scientific questions, and claim boundaries instead of generic artifact summaries. -
Results now cites main figures and appendix diagnostics by role, removes literature citations, avoids subsections, and converts internal identifiers such as
row_countorsource_idinto manuscript-facing scientific wording. -
Added idempotent Results behavior so confirmed Results are not rewritten when downstream manuscript stages rerun with unchanged result evidence.
-
Upgraded
inventory-resultsto write a structured v0.16.5 result manifest withmain_figures,appendix_figures,supporting_links,claim_boundaries, internal tables, and figure-code traces. -
Fixed stale-output handling: planned generated figures that were not rendered in the current run are listed under
excluded_unrendered_figuresand no longer enter Results or quality checks merely because an old PNG remains on disk. -
Added regression coverage to ensure unrendered supporting/appendix figures remain in diagnosis/repair context but are excluded from the scientific result inventory.
-
Expanded Data Role aliases for event-level samples, sample groups, current-observation tokens, historical sequence tokens, modality availability, feature matrices, astronomy products, and model-evaluation fields.
-
Strengthened figure contract validation so 5-6 main figure groups are checked separately from supporting or appendix diagnostics, and planned main results cannot be silently replaced by validation artifacts.
-
Connected method feasibility, figure execution diagnosis, result validity, and core evidence checks so missing data or method coverage routes to repair before human confirmation.
-
Reframed the figure contract around 5-6 main figure groups instead of a hard cap on generated PNG files. A main figure group may contain multiple panels or generated artifacts, so extra generated outputs are valid when they serve the planned figure story.
-
Added main/supporting/appendix figure accounting to
figure_plan.jsonandfigure_contracts.json. Supporting diagnostics no longer replace main results, but can be cited as Appendix Figures when they strengthen reliability or validation arguments in Results and Discussion. -
Updated
assess-figure-contractsso it checks the number of main figure groups, allows additional panels and appendix diagnostics, and reports generated/supporting/appendix counts separately. -
Connected figure-contract failures to
repair-figure-dataandrepair-figure-method, so missing data or missing method coverage produces concrete acquisition or method-repair tasks before the loop asks for human confirmation. -
Strengthened Data/Methods manuscript generation for astronomy and machine-learning projects with role-aware evidence numbers, safer product terminology, stage-specific Methods subsection profiles, and instruction-residue filtering.
-
Made final
quality-checkfail when citation reference coverage fails, preserving the no-delete citation audit principle that retained literature summaries must be represented in the manuscript rather than silently dropped. -
Added regression tests for main-figure-group accounting, contract-gate repair tasks, appendix figure citations in Results, citation coverage quality gating, and astronomy Methods prose cleanup.
-
Hardened
verify-methodsafter the local code audit: verification now resolves commands to an argv list, executes withshell=False, rejects shell operators and explicit shell runners, and recordsshell_used=falseinmethods/run_manifest.yaml. -
Added
verify_command_argvand{python}placeholders to generated method-code manifests so project artifacts no longer embed the developer machine's Python executable path. The legacyverify_commandstring remains only for compatibility. -
Moved full method stdout/stderr into project-local
methods/run_logs/files while keeping bounded excerpts and log metadata in the run manifest. -
Clarified the analysis-code output contract: generated analysis code is canonical under
methods/, whilecode/is a compatibility copy for older workflows. -
Removed duplicated core-evidence test setup by sharing one test helper, and added regression coverage for manifest-driven argv execution, shell rejection, log manifests, and canonical/compatibility output separation.
-
Added a Data/Methods writing-brief layer.
build-data-contextandbuild-method-contextnow writedata/data_writing_brief.json/.htmlandmethods/method_writing_brief.json/.htmlbefore manuscript prose is generated. -
Reworked Data and Methods writing so section text is guided by required evidence roles and method stages instead of mechanically dumping context fields into the manuscript.
-
Added research feasibility and research-plan feasibility gates before downstream data/method/figure work. The loop now records whether the proposed study has enough data roles, method intent, and figure-storyboard evidence before code generation.
-
Added data role coverage, method feasibility, method repair, and method degradation reports so missing data or method capability is diagnosed before figure/code execution.
-
prepare-data-acquisitionnow consumes data role coverage and research-plan feasibility gaps, turning missing data roles into connector-aware acquisition tasks instead of waiting for a later reviewer loop. -
Added a figure contract gate.
generate-analysis-codenow refuses blocked main-figure contracts instead of silently replacing planned result figures with validation or diagnostic fallback plots. -
revise-research-plannow writes a human-readable revision packet underresearch_plan/so data/method repair, scope fallback, and regeneration instructions are visible before the plan is rerun. -
assess-result-validitynow reads figure contracts, the figure contract gate, and figure execution diagnosis; blocked or missing planned main figures force a repair route even when tabular metrics pass. -
Connected
statusandrun-pipelineto the new repair-first route: repair data, repair methods, or revise the research plan before narrowing the scientific claim. -
Synchronized the bundled Draftpaper workflow skill and command reference with the new preflight, research-plan feasibility, method feasibility, and figure-contract stages.
-
Fixed data-role canonicalization so short aliases such as
rado not corrupt broader roles such as spectral or remote-sensing features. -
Tightened composite-discipline method blueprints: the full plugin catalog remains available, but the method data contract is now built from templates selected for the current research plan, storyboard, method requirements, and review tasks.
-
Added integrity checks for writing-brief coverage, Methods formula rendering, and formula-variable explanation while keeping prose style open enough for Codex to write natural scientific paragraphs.
-
Local verification:
python -m pytest tests/test_cli_feasibility_commands.py tests/test_research_feasibility.py tests/test_research_plan_feasibility_gate.py tests/test_method_feasibility.py tests/test_figure_contract_gate.py tests/test_orchestrator_research_feasibility_routing.py -
Local verification:
python -m pytest tests/test_data_feasibility.py tests/test_methods.py tests/test_integrity_gate.py tests/test_composite_discipline_modules.py tests/test_method_blueprint.py -
Full local verification:
python -m pytest, 241 tests passed.
v0.15.1-v0.15.12 (2026-07-01 to 2026-07-06) -- Evidence-First Loop, Citation Audit, and CLI Hardening
-
Hardened Data/Methods manuscript generation so local paths, filenames, workflow artifacts, implementation-only script names, and internal manifest language are cleaned before they can enter paper prose.
-
Added astronomy-aware observation-product wording for spectral, response, light-curve, event, image, and exposure products, so technical columns such as PHA, ARF, RMF, and light-curve descriptors are translated into manuscript-facing data descriptions.
-
Expanded method-formula extraction for time-aware classification workflows: Time2Vec-style temporal encoding, sequence position encoding, masked pooling, multimodal classifier logits, cross-entropy, macro-F1, ROC-AUC, confusion matrices, ablation deltas, correlation, and goodness-of-fit formulas now include variable explanations and figure links.
-
Upgraded the integrity gate with manuscript-language linting and Data/Results sample-count consistency checks based on
results/tables/sample_composition.csv, preventing stale or inflated evidence counts from passing silently. -
Repositioned citation audit as a post-writing claim-tightening loop: retained references are preserved, unsupported or weak usages are repaired by narrowing claims, moving citations to better-supported sentences, or adding evidence metadata; citation-bearing claims are no longer deleted by the repair step.
-
Local verification:
python -m pytest -
Current suite: 228 tests
-
Fixed
write-resultsso repeated calls no longer rewriteresults/results.texorresults/results_summary_zh.mdwhen the result manifest and generated text are unchanged. -
Prevented accidental downstream stale propagation after confirmed Results writing: Introduction, Data, Methods, Discussion, LaTeX, and quality stages are no longer marked stale merely because
write-resultsis called again during later manuscript assembly. -
Added a regression test for the evidence-first writing order where Results are confirmed first and later writing stages continue without being invalidated by an unchanged Results rerun.
-
Local verification:
python -m pytest -
Current suite: 224 tests
-
Added
classify-code-ownership,route-stage-code,build-code-provenance,extract-method-formulas, andtrace-figures-to-codeso project-specific or legacy scripts undercode/can be routed into stage-owneddata/scripts/,methods/scripts/,methods/src/, andmethods/plotting/locations. -
Added
data/data_code_manifest.json, extendedmethods/method_code_manifest.json, addedmethods/method_formula_manifest.json,methods/method_formulas.tex, andresults/figure_code_trace.jsonas manuscript-facing provenance inputs. -
Upgraded
build-data-contextandbuild-method-contextso Data and Methods writing must read stage-owned code manifests, formula extraction, and figure-code traces instead of treatingcode/as an undifferentiated script dump. -
Added
prepare-discussion-comparison, which writes a comparison-literature matrix, HTML evidence notes, and innovation/limitations prompts beforewrite-discussion. -
Extended astronomy, geography, and machine-learning discipline modules with stage-owned code-layout and formula/figure-trace constraints.
-
Revalidated the migration path on the local astronomy project: legacy
code/scripts were classified, routed, formula-scanned, and traced to result figures; Methods context remained correctly blocked while the upstream method-plan stage was stale after artifact-drift synchronization. -
Switched GitHub Actions to install
.[dev]and runpython -m pytest, matching the local verification path used during development. -
Fixed direct
csv.DictReader(path.open(...))patterns in built-in discipline templates so CSV inputs are opened through context managers. -
Local verification:
python -m pytest -
Current suite: 219 tests
-
Added
draftpaper_cli/template_registry.pyanddraftpaper validate-template-registryfor validating built-in discipline plugin manifests, template files, fixtures, plugin identifiers, and maturity metadata. -
Added tests so future discipline plugin contributions can be checked before promotion into the formal module library.
-
Added
io_utils,latex_utils, andcitation_utilsto reduce repeated JSON/text loading, LaTeX escaping, BibTeX parsing, and citation-key parsing logic. -
Migrated high-risk manuscript and gate modules, including Methods, Results, Introduction, Discussion, LaTeX assembly, and quality checks, onto shared helpers.
-
Preserved support for common LaTeX citation commands such as
\cite{}and\citep{}. -
Added focused common-utility tests and aligned gate behavior around shared parsing helpers.
-
Kept
methods/as the canonical generated-analysis-code location while retainingcode/as compatibility output. -
verify-methodsnow readsmethods/method_code_manifest.jsonwhen--commandis not supplied, using generated verification metadata, declared outputs, and selected input data. -
Method verification now checks
results/figure_contracts.json; missing or placeholder contracted main figures fail the hard gate. -
generate-analysis-codenow writes manifest-driven verification and plotting-install metadata and recommends a shorter manifest-driven verification command. -
Packaged the paper-fetch runtime under
draftpaper_cli/_vendor/paper_fetch_skillso wheel installs retain the fallback source instead of relying on a source-tree-onlythird_party/path. -
Added a
fulltextoptional extra for heavier article/PDF extraction dependencies while keeping the default install lighter. -
Verified with
python -m pip wheel . --no-depsthat the wheel contains the vendored paper-fetch CLI and third-party license. -
Fixed
verify-methodsCLI semantics so failed method verification returns a non-zero process exit code while still printing the run manifest JSON. Automation can now treat the command as a real hard gate. -
Removed the developer-local historical source path from newly generated
project.jsonfiles and replaced it with a neutrallegacy_mvp_referencenote, improving project portability across machines and public examples. -
Added regression tests for failed
verify-methodsexit codes and portable project scaffolding metadata. -
Refreshed README workflow wording so the documented paper loop stays evidence-first: literature and planning are followed by data/method execution, figure generation, result validity, and core evidence review before manuscript writing.
-
Upgraded
plan-figuresso research-plan storyboard figures become strict main-result contracts written toresults/figure_contracts.jsonand checked throughresults/storyboard_alignment_report.json. -
Upgraded
generate-analysis-codeto writeresults/figure_execution_diagnosis.jsonand.html; missing data and missing method code are diagnosed explicitly instead of silently replacing a failed main figure with a validation, workflow, or supporting plot. -
Added
diagnose-figure-execution,repair-figure-data, andrepair-figure-method. These commands create data/method repair plans using existing connectors, public data/API routes, remote-server workflows, discipline plugins, public research-code repositories, literature implementation repositories, or Codex-generated project-specific method code. -
Upgraded
assess-core-evidence,status,run-pipeline, and the final quality path so unsatisfied main-figure contracts route to data/method repair first; human confirmation is reserved for cases where automated repair still cannot produce the planned core evidence. -
Reordered the main paper pipeline so literature and research planning are followed by data acquisition/integration, method/code execution, figure production, result validity, and a core evidence gate before manuscript-section writing.
-
Added
assess-core-evidence, writingcore_evidence/core_evidence_report.jsonand.htmlto check data supplementation, data integration, method analysis, figure production, figure metadata, and result validity before human figure confirmation. -
Split execution stages from manuscript-writing stages:
dataandmethodsnow prove data/method readiness, whiledata_writingandmethods_writingproducedata.texandmethods.texafter Results evidence exists. -
Updated Results writing to keep continuous prose without per-figure subsections and to write
results/results_summary_zh.mdfor Chinese review of figure-level interpretation. -
Updated orchestration, LaTeX assembly, quality gates, review routing, and the Codex skill wrapper to follow the evidence-first loop.
v0.14.0-v0.14.13 (2026-06-24 to 2026-07-01) -- Discipline Plugins, Citation Repair, IP Protection, and Data Connectors
-
Added the astronomy
remote_fits_zip_streamdata connector for remote-server or instrument-archive workflows where large FITS/ZIP observation products remain external and Draftpaper-loop only persists compact manifests, processed tables, parse-status reports, and provenance records. -
Added a generic public template for event-product manifest construction, ZIP member availability inspection, dense observation-window selection, and streaming data contracts without hard-coded private server addresses, user names, passwords, source identifiers, or project-specific labels.
-
Split training smoke validation into method templates: astronomy now exposes
source_holdout_stream_smoke_test, while machine learning exposesgroup_holdout_training_smoke_test, so event-random metrics remain leakage-risk contrasts and group/source-held-out metrics become the primary validation path when feasible. -
Upgraded
prepare-data-acquisitionto detectfits_zip_streamaccess patterns and route astronomy missing-data tasks toward the proper connector. -
Made DOI and URL fields clickable in per-paper literature summary HTML files.
-
Reworked literature query planning so idea/title text is reduced to short phrases such as method, instrument, data, and task terms before being crossed with data and method queries.
-
Prevented long full-title research ideas from being repeated across every search query while keeping discipline anchors such as high-energy time-domain astronomy and X-ray transient classification.
-
Revalidated the astronomy workflow with
search-literature -> generate-plan; the generated query plan now avoids full-sentence idea duplication while preserving 12 retained references, 6 planned figures, and 1 planned core table. -
Changed
generate-planso the human-facing research plan is written asresearch_plan/research_plan.mdandresearch_plan/research_plan.zh-CN.md;research_plan.htmland separateresearch_questions.*outputs are no longer generated. -
Embedded research questions, figure storyboard, method-plan contract, expected tables, risk checks, and the literature-summary index link directly into the research plan Markdown files.
-
Upgraded the Chinese research-plan renderer so it writes a fluent Chinese planning document from the same blueprint instead of only translating section headings.
-
Added structured literature query planning with context, query ID, combination level, discipline anchor, and query components for introduction/data/method/target-journal searches.
-
Added low-count fallback searches for astronomy-style projects and records query provenance in
references/search_queries.json,references/literature_items.json, and per-paper literature summary HTML files. -
Verified the new flow on the local astronomy project by rerunning
search-literature -> generate-plan; the project now records 12 retained references, 30 citation-evidence rows, 6 planned main figures, and 1 planned core table. -
Added citation-use metadata for claim-level audit records, including citation intent, support status, topic relevance score, claim-alignment score, blocking status, and repair hints.
-
Updated citation repair planning so unsupported or imprecise citation usages preserve retained references and rewrite manuscript claims to match existing evidence; citation audit no longer plans reference deletion or citation-bearing sentence deletion.
-
Added
references/reference_usage_plan.jsonso retained literature summaries are assigned to manuscript sections and must be cited at least once outside Results. -
Added
citation_audit/reference_coverage_report.jsonandcitation_audit/reference_coverage_report.htmlto comparereferences/literature_summaries/against unique cited references; summarized-but-uncited literature is now a blocking citation-audit failure instead of a silent bibliography shrinkage. -
Added regression tests for preserving method/tool/background citations, keeping relevant contextual citations available for rewrite, and reporting reference-coverage gaps separately from unsupported citations.
-
Added public compliance documentation that clarifies the non-commercial source-available boundary, commercial authorization requirement, sponsorship-not-license rule, and prohibited anti-abuse mechanisms such as hidden payloads, telemetry, device fingerprinting, remote license checks, or destructive checks.
-
Added the public DPL schema family documentation and generated project provenance blocks for project metadata, stage manifests, and project passports.
-
Added stable public contract helpers for claim IDs and evidence IDs, plus tests that verify deterministic DPL identifiers and generated project metadata.
-
Added public forensic-fingerprinting guidance at a high level.
-
Updated commercial license and trademark notes so project terms such as Draftpaper-loop, DPL loop engine, project passport, claim trace, and evidence binding cannot be used to imply commercial authorization or official endorsement.
-
Added an independent citation audit and repair loop:
audit-citations,generate-citation-repair-plan,apply-citation-repair,re-audit-citations, andrun-citation-repair-loop. -
Added claim-level local citation support checking against BibTeX and
references/citation_evidence.csv, with RefCheck-style HTML reports undercitation_audit/iterations/and a final pass report atcitation_audit/final_citation_audit_report.html. -
Updated
statusandrun-pipelineso final quality checks are blocked after integrity checks until citation audit passes; failed citation audits now route into the citation repair loop instead of continuing blindly toquality-check. -
Upgraded
quality-checkso direct calls requirecitation_audit/final_citation_audit_report.jsonwithstatus=passed, preventing source-support verification from being skipped. -
Added discipline module maturity metadata with
foundation,runnable, andmatureas the intended progression. -
Added
capture-discipline-learning, which summarizes reusable project lessons from observations, method artifacts, review plans, data contexts, result metadata, and rescue plans intoplugin_candidates/from_loop/...without copying raw data or hidden reasoning. -
Added
classify-plugin-reusability, which separates reusable method/review/data signals from project-specific identifiers, local paths, credentials, fixed regions, and other non-generalizable content before any promotion. -
Upgraded finance, medicine, biology, and engineering from foundation-only specs to runnable foundation modules by adding one standard-library fixture-backed method template per discipline.
-
Added foundation reviewer engines for finance, medicine, biology, and engineering so these disciplines no longer fall back directly to the default reviewer route.
-
Added
inspect-research-repo, which reads a candidate repository checkout as structure/docs/package metadata only and writesrepository_structure.json,file_inventory.csv,package_manifest.json, and HTML inspection reports without copying source code. -
Added
map-repository-workflow, which maps repository file roles into data connector, preprocessing, method, figure, validation, review, environment, and documentation capabilities. -
Added
bootstrap-discipline-foundation, which writes candidate-only discipline foundation suggestions from workflow maps; it does not modify formal modules by default. -
Added foundation discipline modules for finance, medicine, biology, and engineering, each with initial data connector specs, method template specs, and reviewer-rule groups.
-
Added the first minimal public research-code mining chain:
discover-research-repos,score-research-repos, andextract-plugin-candidates. -
The new flow writes metadata-only JSON/HTML reports under
research_code_mining/, ranking repositories by license safety, reproducibility metadata, linked-paper signals, workflow completeness, and reusable capability hints. -
Candidate extraction produces
candidate_manifest.json,candidate_report.html, and an index report while explicitly avoiding repository cloning, third-party source copying, or direct plugin installation. -
This creates a safer front door for future discipline-module expansion: public code can inspire generalized data/method/figure/review templates only after license, privacy, overlap, fixture, and maintainer review.
-
Kept the custom source-available non-commercial license instead of switching to Apache-2.0 or another standard SPDX license, because preserving commercial authorization requirements needs non-standard terms.
-
Cleaned public update wording so plugin seeds are described by reusable capabilities rather than by specific source projects or private validation targets.
-
Added a README policy expectation that future public changelog entries should avoid exposing concrete internal project names, local validation folders, private datasets, or project-specific research directions.
-
Added runtime composite discipline modules for cross-disciplinary papers. The loop now records
primary_discipline,secondary_disciplines,discipline_scores, and ordereddiscipline_modules. -
get_discipline_modulecan now mergedefault, the primary discipline, and secondary disciplines, deduplicating connectors, method templates, and review rules by stable ids. -
Cross-disciplinary projects such as geography + machine learning or astronomy + machine learning can expose both domain data/method plugins and ML modelling/reviewer plugins in
prepare-method-blueprint,prepare-data-acquisition, andplan-figures. -
Plugin candidate manifests now record primary and secondary disciplines while still keeping stable reusable capabilities under their intended home module.
-
Added built-in geography data connectors for Earth Engine precipitation export planning, NetCDF-to-GeoTIFF planning, gridded text-to-raster conversion, and ArcGIS/project-bound zonal statistics manifests.
-
Added geography method templates for monthly remote-sensing index summaries, phenology curve smoothing, NDVI temporal K-means zoning, and cluster statistical diagnostics.
-
Added machine-learning data/model seeds for tabular environmental dataset profiling, sanitized saved-model manifests, RF/XGBoost/GBDT/stacking regression plans, observed-predicted diagnostics, feature importance, PDP/ICE, SHAP planning, and a model statistical-validity reviewer gate.
-
Kept these seeds fixture-backed and dependency-light so reusable plugin code stays generic while project-specific paths, API accounts, data windows, and model binaries remain local.
-
Added astronomy connector and method seeds for photon-event access planning, observation/product manifests, long-term light-curve feature extraction, and event-level sequence input construction.
-
Added machine-learning/deep-learning connector and method seeds for vision catalog alignment, pretrained backbone metadata, self-supervised training plans, checkpoint compatibility diagnostics, embedding health checks, few-label evaluation, and similarity retrieval.
-
Added fixture-backed tests so these discipline plugins can be validated without private data, API credentials, large checkpoints, or GPU training.
-
Kept project-specific paths, credentials, checkpoint binaries, and sample selections out of reusable plugin templates; real projects bind those values locally through Draftpaper-loop project files.
-
Added full
DataConnectorSpecandMethodTemplateSpecschemas for discipline modules. -
Reorganized discipline modules toward a three-layer model:
data_connectors/,method_templates/, andreview_rules/. -
Added geography method templates for
remote_sensing_feature_reconstructionandspatial_block_validation. -
Added machine-learning method templates for
baseline_model,ablation_study, andtrain_validation_test_split_check. -
Added plugin contribution preflight commands:
summarize-plugin-candidates,generalize-plugin-candidate,validate-plugin-candidate,package-plugin-contribution, andwrite-github-contribution-guide. -
Documented fork/PR rules: forks and branches are temporary contribution channels; stable reusable capabilities merge into
mainunder the matching discipline module after privacy, genericity, overlap, fixture, and validation checks.
-
Upgraded
plan-figuresso discipline modules can declareminimum_main_figures,target_main_figures, andrequired_figure_groups; first drafts now aim for at least five generated main figures when data are available. -
Extended discipline modules with data connector catalogs that include packages, import modules, API/download routes, credential requirements, expected data formats, and local feasibility status.
-
Added
ecologyandbioinformaticsmodule skeletons alongsidedefault,geography,astronomy, andmachine_learning. -
Expanded geography/agriculture, astronomy, ecology/environment, machine-learning, and bioinformatics data acquisition routes for research-plan data suggestions and missing-data rescue.
-
Added a stage-owned code layout: data acquisition/preprocessing code belongs under
data/scripts, method/model/statistical/spatial/figure-generation code belongs undermethods/scriptsandmethods/src, andresultskeeps only produced figures, tables, and metadata. -
Added
prepare-method-blueprint, which writesmethods/method_blueprint.json,methods/method_data_contract.json,methods/method_code_plan.json, andmethods/method_formula_plan.json. -
Added the
discipline_modulesframework with default, geography, astronomy, and machine-learning module skeletons for shared data, method, figure, formula, and reviewer constraints. -
Upgraded
generate-analysis-codeso canonical generated code is saved undermethods/whilecode/remains a compatibility launcher/copy for older workflows. -
Added contributor documentation under
docs/discipline_modules/for future discipline-module submissions.
-
Connected reviewer/rescue missing-data advice to
prepare-data-acquisition. -
prepare-data-acquisitionnow readsreview/actionable_analysis_tasks.json,review/review_engineering_plan.json,review/statistical_rescue_plan.json,review/revision_plan.json, andreview/gate_failure_diagnosis.json. -
Added
data/data_acquisition_tasks.jsonanddata/data_acquisition_tasks.html, turning blocked analysis tasks into explicit missing-data requests withneeded_data,optional_data,suggested_connectors, and user-confirmation questions. -
Updated
statusandrun-pipelineso review/rescue execution now recommendsprepare-data-acquisitionafterprepare-analysis-revisionand beforeplan-figures --use-review-tasks. -
Added a shared discipline inference layer used by both data-acquisition planning and reviewer-engineering engines, so Data and review/rescue routes no longer maintain separate discipline guesses.
-
Added
classify-data-access,prepare-data-acquisition, andinventory-data-sourcesfor plan-first data access classification. -
Added generic connector profiles for
local_files,api_access, andremote_server. These detect access patterns without downloading data, writing credentials, or hard-coding astronomy/geography packages into the core Data stage. -
Added acquisition artifacts:
data/data_access_profile.json,data/data_acquisition_plan.json,data/data_acquisition_plan.html,data/data_source_manifest.csv,data/data_access_log.csv,data/data_provenance.json, anddata/data_completeness_report.html. -
Added design and implementation notes under
docs/superpowers/specs/anddocs/superpowers/plans/. -
Validated the generic layer against a local cross-discipline source tree while detecting local-file, API-access, and remote-server data modes without documenting private paths or project-specific identifiers.
v0.11.0-v0.11.1 (2026-06-21 to 2026-06-23) -- Publication Readiness, Statistical Rescue, and Source-Available Protection
-
Added repository-level protection files:
NOTICE,COMMERCIAL_LICENSE.md, andTRADEMARK.md. -
Updated
LICENSEwording from the older DraftPaper CLI identity to Draftpaper-loop and clarified commercial-use examples such as API services, manuscript-production services, and paid course bundles. -
Added copyright and contact headers to first-party Python files while keeping vendored third-party files untouched.
-
Added stable Draftpaper-loop generator provenance to generated LaTeX, HTML reports, generated Python scripts, and JSON reports that include
generated_at. -
Local verification:
python -m unittest discover -s tests -
Current suite: 130 tests
-
Added
assess-publication-readinessto score target-journal submission readiness from saved loop artifacts, including data feasibility, method verification, result validity, figure metadata, integrity, quality, and journal profile state. -
Added
discover-review-workflow-gapsandpropose-review-engineering-planfor discipline-specific reviewer-engineering. The first deterministic engine is geography, covering remote sensing, agricultural geography, spatial scale alignment, remote-sensing QC, spatial autocorrelation, stratified heterogeneity, and weak-fit backtracking; astronomy and machine-learning engines are now reserved with baseline reviewer rules; unmatched projects use a default fallback. -
Added
recommend-statistical-revisionto generate a statistical rescue plan when weak data or unsupported results may be improved through robust statistics, missingness analysis, method rebuilding, explicit success thresholds, domain-aware feature rebuilding, spatial validation, model validation checks, or claim reframing. -
Added review-engineering outputs:
review/review_discipline_profile.json,review/review_workflow_gap_report.json,review/review_workflow_gap_report.html,review/review_engineering_plan.json,review/review_engineering_plan.html, andreview/user_confirmation_requests.json. -
Added
prepare-analysis-revision, which converts reviewer/rescue recommendations intoreview/actionable_analysis_tasks.json, checks data-role completeness inreview/analysis_revision_feasibility.jsonand.html, and writes downstream hints tomethods/analysis_revision_requirements.jsonandresults/revision_figure_plan_delta.json. -
Connected
plan-figures --use-review-tasksandgenerate-analysis-code --use-review-tasksso executable/partial reviewer tasks become revised figures, review-task coverage tables, and review-task metrics; blocked tasks remain explicit data requests. -
Added review outputs:
review/publication_readiness_report.json,review/publication_readiness_report.html,review/codex_archive_review_context.json,review/codex_archive_review_context.html,review/statistical_rescue_plan.json,review/statistical_rescue_plan.html,review/journal_fit_report.html, andreview/claim_evidence_matrix.csv. -
Added a project-archive reviewer context layer so publication readiness reports can produce more natural reviewer-like assessments from saved research plan, literature, data, method, result, figure, journal, integrity, and quality artifacts.
-
Added discipline-aware statistical rescue routes for agricultural remote sensing, spatial/ecological analysis, astronomical time series, and machine-learning validation when those signals appear in the archived project context.
-
Added statistical-metric semantics to result validity and review rescue: p-values are evaluated against alpha thresholds such as 0.05, while R2 is treated as goodness of fit, correlation coefficients are treated as effect sizes, and low figure-level R2/r values trigger data-quality, proxy-variable, outlier, spatial-alignment, and method-rebuild recommendations only when the project methods actually produce those statistical outputs.
-
Upgraded
status,run-pipeline,generate-revision-plan, andre-reviewso quality/integrity failures route through gate diagnosis, reviewer pass, publication readiness, statistical rescue, and one shared stale-stage rerun path. -
Local verification:
python -m unittest discover -s tests -
Current suite: 127 tests
- Added manuscript writing-quality checks for minimum section length, substantive paragraph counts, non-bulleted natural prose, arbitrary bold avoidance, Introduction/Discussion citation presence, Methods formula presence, and Results figure-count requirements.
- Upgraded
write-resultsso result paragraphs cite their supporting figures and tables with LaTeX labels, for exampleFigure~\ref{...}andTable~\ref{...}. - Cleaned manuscript-facing Results and Discussion prose so internal loop terms, gate names, local-file safeguards, manifest references, and Draftpaper-loop implementation wording are not written into the scientific body text.
- Added default LaTeX acknowledgments that disclose Draftpaper-loop assistance and include the project link
https://github.com/xiejhhhhhh/Draftpaper_loop. - Strengthened Methods outputs with formula manifests and
methods/method_formulas.tex, then connected quality checks so Methods drafts cannot silently omit mathematical expressions. - Expanded scientific-figure verification so generated figures must provide PNG/publication metadata, axis labels, text elements, statistics, interpretation summaries, and publication-ready backend evidence.
- Local verification:
python -m unittest discover -s tests - Current suite: 111 tests
- Added plotting dependency extras:
.[plotting]for the recommended research environment and.[plotting-full]for advanced figure/report backends. - Added a project-local scientific plotting runtime copied into generated projects under
code/src/scientific_plotting.py, improving portability across machines. - Upgraded
plan-figuresfrom simple visualization labels to figure specifications withfigure_type, variables, statistical transforms, backend preferences, andno_flowchart_fallback. - Upgraded
generate-analysis-codeso generated pipelines produce empirical SVG figures,results/figure_metadata.json, andresults/figure_quality_report.json. - Added numpy/stdlib-compatible scientific SVG plots for scatter-regression, histogram, class-support, correlation heatmap, and metric-summary outputs.
- Upgraded Results inventory/writing so result claims use figure metadata such as sample size, association direction, R/R2, class support, and metric summaries instead of generic artifact text.
- Upgraded quality checks so generated empirical figures must have scientific metadata, axes/scale evidence, and interpretation summaries; workflow diagrams can no longer silently replace Results figures.
- Local verification:
python -m unittest discover -s tests - Current suite: 103 tests
- Added
record-observationso visible Codex/user analysis summaries can be preserved locally without storing hidden reasoning. - Added
build-data-contextandwrite-dataso the Data section is written from source/content/processing/claim-boundary context rather than file paths. - Added
build-method-contextand upgradedwrite-methodsso Methods is written from method rationale, data role, verification summary, and claim boundary rather than commands or manifest dumps. - Upgraded
run-pipelineso multi-step Data and Methods stages are not marked complete until required context and manuscript outputs exist. - Added quality-gate lint that fails Data/Methods sections containing local filenames, filesystem paths, execution commands, or manifest-style output text.
-
Preserved Zotero-imported references as user-curated evidence outside external-search ranking, recency, abstract/PDF filtering, and the default 30-paper external cap.
-
Added Zotero source/origin/collection/selection-policy metadata to literature review notes, per-paper HTML summaries, and
references/literature_summaries/index.html. -
Added tests confirming Zotero references appear together with searched literature while remaining distinguishable.
-
Added
list-zotero-collectionsfor inspecting Zotero collection names from Codex or the local CLI. -
Added
search-literature --zotero-collectionso a paper project can import user-curated references from one Zotero collection. -
Added
--zotero-context,--zotero-min-items, and--no-zotero-supplementto control citation-evidence context and MVP-style external supplementation. -
Added
references/zotero_collection_manifest.jsonto record the requested collection, matched collection, collection key, usable item count, and supplemental count. -
Documented how to configure
ZOTERO_LIBRARY_ID,ZOTERO_LIBRARY_TYPE, andZOTERO_API_KEYfor Codex-driven loop calls without exposing credentials in outputs.
- Reframed the project from a CLI-first paper drafting tool to a loop-engineered research manuscript system.
- Renamed the README identity from DraftPaper CLI to Draftpaper-loop while keeping
draftpaperas the stable command-line interface. - Added an explicit loop model covering observe, decide, run, verify, persist state, mark stale stages, diagnose failures, and rerun.
- Added
plan-figuresso the loop plans project-specific scientific figures before generating analysis code. - Changed
generate-analysis-codeto followresults/figure_plan.jsoninstead of producing fixed workflow figures. - Added remote/server/API data handling through source manifests and supplied processed/result artifacts with claim-limited conditional passes.
- Added HTML outputs for report-style artifacts such as novelty reports, figure plans, and literature review notes. Later versions moved the main research plan back to Markdown-first outputs for better direct review and editing.
- Moved contact, commercial-use terms, and homepage information to the end of the README after the update log.
- Added
diagnose-gate-failures,review-draft,generate-revision-plan,apply-revision, andre-review. - Added unified revision issues with
source,target_stage,files_to_add_or_edit,required_user_input, andrecommended_commands. - Added
review/commitment_ledger.csvso user revision decisions can be tracked across review cycles. - Connected
statusandrun-pipelineto failed integrity/quality reports so they recommenddiagnose-gate-failuresinstead of repeated blind gate reruns. - Kept
apply-revisionintentionally conservative: it marks affected stages stale and does not rewrite scientific content automatically. - Local verification:
python -m unittest discover -s tests - Current suite: 95 tests
- Added
run-integrity-gate. - Added
integrity/integrity_report.jsonandintegrity/integrity_report.md. - Validates BibTeX citation existence,
citation_evidence.csvtraceability, Results no-citation rules, and result-claim artifact binding. - Connected
statusandrun-pipelineso final quality checks wait for a passed integrity report.
- Added DraftPaper Passport files:
project_passport.yaml,artifact_ledger.jsonl,checkpoint_ledger.jsonl, andintegrity_ledger.jsonl. - Added
status,checkpoint,resume, andrun-pipeline. - Added hash-based drift detection with
detect-artifact-driftandsync-artifact-stale. - Added literature-informed analysis-code generation and early method/result artifact checks for local workflow smoke tests.
- Added
collect-method-planto convert user method notes and literature-informed method summaries intomethods/method_requirements.json. - Added
generate-analysis-codeto create reviewable baseline analysis code from retrieved literature, data inventory, and method requirements. - Added
verify-methodsandmethods/run_manifest.yaml; Methods writing now requires a successful local method-code run. - Added
assess-result-validityso unsupported results can route back to data, methods, or the research plan. - Added
inventory-resultsandwrite-results; Results writing is bound to real figures/tables and rejects citation commands. - Added
write-discussion,assemble-latex,compile-latex-pdf, andquality-check.
- Added the single-paper project directory model with
idea/,references/,research_plan/,introduction/,data/,methods/,results/,discussion/, andlatex/. - Added
create-project,load-project,validate-project,update-stage-status, andmark-stage-stale. - Added free-first literature retrieval through Semantic Scholar, arXiv, Crossref, and optional SerpApi.
- Standardized reference outputs:
references/library.bib,references/literature_items.json,references/citation_evidence.csv, and literature review notes in Markdown plus HTML. - Added target-journal template resolution through
resolve-journal-templateand literature-informedgenerate-plan. - Added a traceable Introduction writer whose citation keys must exist in both BibTeX and citation evidence.


