Skip to content

vadugwi flywheel → refine corpus weights from crowd word-ratings (when data accrues) #32

Description

@deucebucket

The vadugwi Space (deucebucket/vadugwi, repo ~/ai-drive/vadugwi) collects anonymous word ratings into the private dataset deucebucket/vadugwi-data. Use them to refine this engine's corpus — human-reviewed, never auto-applied (weights are GA champions; FP-zero).

When enough data has accrued, run ~/ai-drive/vadugwi/tools/aggregate.py:

  • Known words — where human mean valence diverges from compute_vadug(word).v → weight-correction candidates.
  • Unknown words — mean human rating → proposed force tuple → new-vocabulary candidates for engine/vocabulary.py / engine/forces_curated.py.

Then review and adopt only vetted deltas.

Also: the unknown-word harvest currently falls back to corpora because the pet audit (deucebucket/clanker-audit) is still empty. Verify it auto-upgrades to real slang / suspected_gap tokens once the pet logs live usage (~/ai-drive/vadugwi/tools/harvest_unknowns.py).

Context: built 2026-06-17. Related: #31.

Metadata

Metadata

Assignees

No one assigned

    Labels

    P4-lowMinor, cosmetic, nice-to-havefeatureNew capability requested

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions