-
Notifications
You must be signed in to change notification settings - Fork 0
feat: agent skills for usage, receipts, distillation, and backtesting #11
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from all commits
Commits
Show all changes
11 commits
Select commit
Hold shift + click to select a range
e18c032
feat: agent skills for usage, receipts, distillation, and backtesting
mstrathman c4899ce
chore: spec-polish the skill frontmatter
mstrathman 7dd6a15
feat: receipts reference script and CI conformance checks
mstrathman 48e1498
feat: predict_sha256 makes the receipts workflow pure SQL
mstrathman 7c3c6f6
fix: address all nine CodeRabbit findings on the skills PR
mstrathman 0db3468
test: drop the vacuous or-clause from the replay-attack assertion
mstrathman 18ceff0
fix: address round-two CodeRabbit findings
mstrathman 3aaef19
chore: keep the scale-study harness out of the skills PR
mstrathman 1e49364
fix: verify re-derives the serving model; claims narrowed to behavior
mstrathman cb0f5c6
fix: replay treats stored SQL as untrusted
mstrathman db6046c
fix: complete the untrusted-replay audit in one pass
mstrathman File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,13 @@ | ||
| {"dataset": "Diabetes130US", "n_train": 1500, "n_test": 500, "d": 40, "xgboost": 0.906, "xgb_s": 0.1, "tabicl": 0.916, "tabicl_s": 2.2, "tabicl_device": "mps", "leak_gap": 0.0007, "ic_maxp": 0.9251, "collapsed": 1, "gbt<-tabicl soft": 0.916, "oof_s": 7.6, "oof_gap": 0.0007, "gbt<-tabicl soft-oof": 0.916} | ||
| {"dataset": "Diabetes130US", "n_train": 9999, "n_test": 3334, "d": 40, "xgboost": 0.9070185962807439, "xgb_s": 0.1, "tabicl": 0.9130173965206959, "tabicl_s": 101.9, "tabicl_device": "mps", "leak_gap": -0.0, "ic_maxp": 0.922, "collapsed": 1, "gbt<-tabicl soft": 0.9130173965206959, "oof_s": 116.4, "oof_gap": -0.0, "gbt<-tabicl soft-oof": 0.9130173965206959} | ||
| {"dataset": "airline_satisfaction", "n_train": 1500, "n_test": 500, "d": 21, "xgboost": 0.91, "xgb_s": 0.0, "tabicl": 0.928, "tabicl_s": 1.6, "tabicl_device": "mps", "leak_gap": 0.0693, "ic_maxp": 0.8787, "gbt<-tabicl": 0.908, "distill_s": 0.5, "blob_kb": 103.4, "gbt<-tabicl soft": 0.904, "oof_s": 5.5, "oof_gap": -0.0073, "gbt<-tabicl soft-oof": 0.906} | ||
| {"dataset": "GiveMeSomeCredit", "n_train": 1500, "n_test": 500, "d": 10, "xgboost": 0.922, "xgb_s": 0.0, "tabicl": 0.936, "tabicl_s": 1.2, "tabicl_device": "mps", "leak_gap": 0.0067, "ic_maxp": 0.9407, "gbt<-tabicl": 0.932, "distill_s": 0.3, "blob_kb": 97.1, "gbt<-tabicl soft": 0.926, "oof_s": 4.1, "oof_gap": -0.0113, "gbt<-tabicl soft-oof": 0.93} | ||
| {"dataset": "GiveMeSomeCredit", "n_train": 9999, "n_test": 3334, "d": 10, "xgboost": 0.9289142171565686, "xgb_s": 0.1, "tabicl": 0.9322135572885423, "tabicl_s": 8.4, "tabicl_device": "mps", "leak_gap": 0.021, "ic_maxp": 0.9452, "gbt<-tabicl": 0.9334133173365327, "distill_s": 2.6, "blob_kb": 104.3, "gbt<-tabicl soft": 0.9334133173365327, "oof_s": 29.5, "oof_gap": 0.0009, "gbt<-tabicl soft-oof": 0.9322135572885423} | ||
| {"dataset": "APSFailure", "n_train": 1500, "n_test": 500, "d": 40, "xgboost": 0.994, "xgb_s": 0.1, "tabicl": 0.996, "tabicl_s": 2.6, "tabicl_device": "mps", "leak_gap": -0.004, "ic_maxp": 0.9902, "gbt<-tabicl": 0.99, "distill_s": 1.4, "blob_kb": 76.6, "gbt<-tabicl soft": 0.99, "oof_s": 8.2, "oof_gap": -0.0093, "gbt<-tabicl soft-oof": 0.996} | ||
| {"dataset": "SDSS17", "n_train": 1500, "n_test": 500, "d": 11, "xgboost": 0.97, "xgb_s": 0.1, "tabicl": 0.972, "tabicl_s": 1.4, "tabicl_device": "mps", "leak_gap": 0.0167, "ic_maxp": 0.9892, "gbt<-tabicl": 0.968, "distill_s": 1.1, "blob_kb": 145.8, "gbt<-tabicl soft": 0.968, "oof_s": 4.3, "oof_gap": 0.0007, "gbt<-tabicl soft-oof": 0.964} | ||
| {"dataset": "SDSS17", "n_train": 9999, "n_test": 3334, "d": 11, "xgboost": 0.9679064187162567, "xgb_s": 0.3, "tabicl": 0.9718056388722256, "tabicl_s": 8.7, "tabicl_device": "mps", "leak_gap": 0.0205, "ic_maxp": 0.9906, "gbt<-tabicl": 0.9658068386322736, "distill_s": 9.2, "blob_kb": 159.2, "gbt<-tabicl soft": 0.9661067786442712, "oof_s": 30.3, "oof_gap": 0.0052, "gbt<-tabicl soft-oof": 0.9676064787042592} | ||
| {"dataset": "taiwanese_bankruptcy", "n_train": 1500, "n_test": 500, "d": 40, "xgboost": 0.966, "xgb_s": 0.1, "tabicl": 0.968, "tabicl_s": 2.4, "tabicl_device": "mps", "leak_gap": 0.004, "ic_maxp": 0.9721, "gbt<-tabicl": 0.968, "distill_s": 2.9, "blob_kb": 93.8, "gbt<-tabicl soft": 0.966, "oof_s": 8.2, "oof_gap": 0.0007, "gbt<-tabicl soft-oof": 0.968} | ||
| {"dataset": "online_shoppers_intention", "n_train": 1500, "n_test": 500, "d": 17, "xgboost": 0.882, "xgb_s": 0.1, "tabicl": 0.89, "tabicl_s": 1.5, "tabicl_device": "mps", "leak_gap": 0.0507, "ic_maxp": 0.8863, "gbt<-tabicl": 0.89, "distill_s": 0.4, "blob_kb": 100.9, "gbt<-tabicl soft": 0.888, "oof_s": 4.9, "oof_gap": 0.0153, "gbt<-tabicl soft-oof": 0.886} | ||
| {"dataset": "credit_card_default", "n_train": 1500, "n_test": 500, "d": 23, "xgboost": 0.81, "xgb_s": 0.1, "tabicl": 0.836, "tabicl_s": 1.7, "tabicl_device": "mps", "leak_gap": -0.0247, "ic_maxp": 0.8101, "gbt<-tabicl": 0.798, "distill_s": 0.9, "blob_kb": 93.5, "gbt<-tabicl soft": 0.81, "oof_s": 5.8, "oof_gap": -0.0153, "gbt<-tabicl soft-oof": 0.838} | ||
| {"dataset": "airline_satisfaction", "n_train": 9999, "n_test": 3334, "d": 21, "xgboost": 0.9448110377924415, "xgb_s": 0.3, "tabicl": 0.9511097780443911, "tabicl_s": 139.6, "tabicl_device": "cpu", "leak_gap": 0.0465, "ic_maxp": 0.9842, "gbt<-tabicl": 0.9337132573485303, "distill_s": 3.5, "blob_kb": 105.7, "gbt<-tabicl soft": 0.9340131973605279, "oof_s": 486.7, "oof_gap": 0.0043, "gbt<-tabicl soft-oof": 0.9337132573485303} | ||
| {"dataset": "credit_card_default", "n_train": 9999, "n_test": 3334, "d": 23, "xgboost": 0.8047390521895621, "xgb_s": 0.2, "tabicl": 0.8191361727654469, "tabicl_s": 167.5, "tabicl_device": "cpu", "leak_gap": 0.0205, "ic_maxp": 0.8387, "gbt<-tabicl": 0.8197360527894421, "distill_s": 7.4, "blob_kb": 96.8, "gbt<-tabicl soft": 0.8182363527294542, "oof_s": 571.6, "oof_gap": -0.001, "gbt<-tabicl soft-oof": 0.8194361127774445} |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,25 @@ | ||
| # Scale study: distillation beyond 1500 rows | ||
|
|
||
| TabICL v2 as the teacher on the naturally large TabArena | ||
| datasets, at increasing training sizes. Questions: does student | ||
| retention hold, where does tuned xgboost catch the zero-shot | ||
| teacher, and does the in-context label leak grow with context | ||
| (the gap column; see permissive-teachers.md). | ||
|
|
||
| Device: mps with cpu fallback (per-cell `dev`). | ||
|
|
||
| | dataset | n_train | xgboost | TabICL | gbt<-TabICL | soft | soft-oof | leak gap | oof gap | teacher s | distill s | blob KB | | ||
| |---|---|---|---|---|---|---|---|---|---|---|---| | ||
| | APSFailure | 1500 | 0.994 | 0.996 | 0.990 | 0.990 | 0.996 | -0.0040 | -0.0093 | 2.6 | 1.4 | 76.6 | | ||
| | Diabetes130US | 1500 | 0.906 | 0.916 | - | 0.916 | 0.916 | 0.0007 | 0.0007 | 2.2 | - | - | | ||
| | Diabetes130US | 9999 | 0.907 | 0.913 | - | 0.913 | 0.913 | -0.0000 | -0.0000 | 101.9 | - | - | | ||
| | GiveMeSomeCredit | 1500 | 0.922 | 0.936 | 0.932 | 0.926 | 0.930 | 0.0067 | -0.0113 | 1.2 | 0.3 | 97.1 | | ||
| | GiveMeSomeCredit | 9999 | 0.929 | 0.932 | 0.933 | 0.933 | 0.932 | 0.0210 | 0.0009 | 8.4 | 2.6 | 104.3 | | ||
| | SDSS17 | 1500 | 0.970 | 0.972 | 0.968 | 0.968 | 0.964 | 0.0167 | 0.0007 | 1.4 | 1.1 | 145.8 | | ||
| | SDSS17 | 9999 | 0.968 | 0.972 | 0.966 | 0.966 | 0.968 | 0.0205 | 0.0052 | 8.7 | 9.2 | 159.2 | | ||
| | SDSS17 | 49999 | 0.975 | 0.610 | - | 0.610 | - | -0.0000 | - | 118.9 | - | - | | ||
| | airline_satisfaction | 1500 | 0.910 | 0.928 | 0.908 | 0.904 | 0.906 | 0.0693 | -0.0073 | 1.6 | 0.5 | 103.4 | | ||
| | credit_card_default | 1500 | 0.810 | 0.836 | 0.798 | 0.810 | 0.838 | -0.0247 | -0.0153 | 1.7 | 0.9 | 93.5 | | ||
| | credit_card_default | 9999 | 0.805 | 0.771 | 0.776 | 0.776 | 0.776 | 0.0019 | 0.0054 | 11.7 | 7.4 | 96.2 | | ||
| | online_shoppers_intention | 1500 | 0.882 | 0.890 | 0.890 | 0.888 | 0.886 | 0.0507 | 0.0153 | 1.5 | 0.4 | 100.9 | | ||
| | taiwanese_bankruptcy | 1500 | 0.966 | 0.968 | 0.968 | 0.966 | 0.968 | 0.0040 | 0.0007 | 2.4 | 2.9 | 93.8 | |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,84 @@ | ||
| --- | ||
| name: distill-lifecycle | ||
| description: >- | ||
| Distill a teacher model into a native sqlite-predict student and manage | ||
| it over time: verify quality before serving, re-distill on drift, and | ||
| respect the teacher's license. Use when per-call serving needs to be | ||
| instant and self-contained, when a prediction runs at volume, or when a | ||
| model must travel inside the database file. | ||
| license: MIT OR Apache-2.0 | ||
| metadata: | ||
| version: "0.1.0" | ||
| --- | ||
|
|
||
| # The distillation lifecycle | ||
|
|
||
| Distillation compresses a teacher (your own labels, your existing model's | ||
| predictions, or a licensed foundation model) into a native student the | ||
| serving core executes with no runtime. Measured on the repo's benchmark | ||
| suite: student blobs run tens to hundreds of kilobytes, single-row | ||
| serving is sub-millisecond on CPU, and fits complete in seconds at | ||
| benchmark scale (see `benchmarks/results/` in the sqlite-predict repo). | ||
| The lifecycle judgment is yours. | ||
|
|
||
| ## Distill | ||
|
|
||
| ```sql | ||
| -- tabular: target column holds labels or your model's predictions | ||
| SELECT model_id, holdout_metric FROM distill_predict( | ||
| 'SELECT f1, f2, label FROM training', | ||
| '{"target":"label","student_id":"churn-v1","student_kind":"gbt"}'); | ||
|
|
||
| -- time series: from windows, or with a registered onnx teacher | ||
| SELECT model_id FROM distill_forecast('SELECT series_key, value FROM obs', | ||
| '{"context":96,"horizon":24,"student_id":"traffic-v1"}'); | ||
| ``` | ||
|
|
||
| Students register in `_predict_models` as content-hashed rows; they | ||
| snapshot, fork, and sync with the database. Serve by name: | ||
| `predict(NULL, apply_sql, '{"model":"churn-v1"}')` or | ||
| `forecast(ts, value, 24, '{"model":"traffic-v1"}')`. | ||
|
|
||
| ## Verify before serving | ||
|
|
||
| Never serve a student on faith: | ||
|
|
||
| 1. Read `holdout_metric` from the distill call. It is measured on a | ||
| holdout of your data. Compare it to a floor you trust (the majority | ||
| class, last-value carry-forward, or your current model). | ||
| 2. For forecast students, run `backtest` on the same series and compare | ||
| the student against `theta-classic` and the naive floor (see the | ||
| interpret-backtest skill). | ||
| 3. Know the measured shape of distillation loss: on our benchmark suite | ||
| classification students give up a median half point of accuracy | ||
| against their teacher, but regression tails are worse. Check | ||
| regression students more skeptically. | ||
| 4. Soft-label distillation (`proba`/`classes`) preserves the teacher's | ||
| calibration when the teacher emits probabilities, and usually rescues | ||
| imbalanced datasets where hard labels collapse to one class (the | ||
| distiller refuses loudly on collapse rather than fitting a | ||
| constant). | ||
|
|
||
| ## Watch for drift, re-distill cheaply | ||
|
|
||
| A student is frozen; the world is not. When input distributions move, | ||
| quality decays silently, so schedule verification rather than assuming: | ||
|
|
||
| - Re-run `backtest` (forecast) or score a fresh labeled sample (tabular) | ||
| on a cadence proportional to how fast the data changes. | ||
| - Re-distilling costs seconds. Register the new student under a | ||
| versioned id (`churn-v2`), verify, then switch the serving call. Keep | ||
| the old row until the new one is trusted; retire it after. | ||
|
|
||
| ## Licensing is part of the lifecycle | ||
|
|
||
| A student derives from its teacher, and the teacher's license may | ||
| impose obligations on what you distill and distribute; what applies | ||
| depends on the exact weights, license version, and how the student is | ||
| used. Record the teacher's license alongside each student you register. | ||
| Your own labels and models carry no such limits. Restrictively licensed | ||
| teachers require an explicit `accept_license` opt-in before they will | ||
| run, but that gate is an acknowledgment, not compliance: read the | ||
| license of the exact weights you use before distilling for anything | ||
| beyond evaluation. The project documentation's license notes are | ||
| orientation, not legal advice. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,79 @@ | ||
| --- | ||
| name: interpret-backtest | ||
| description: >- | ||
| Read sqlite-predict backtest() output to choose models, trust or | ||
| distrust prediction intervals, and decide when auto-selection needs | ||
| narrowing. Use before serving forecasts that feed decisions, when | ||
| intervals matter, or when choosing between models on a specific series. | ||
| license: MIT OR Apache-2.0 | ||
| metadata: | ||
| version: "0.1.0" | ||
| --- | ||
|
|
||
| # Interpreting backtest() | ||
|
|
||
| `backtest` answers "how would this model have done on my series" with | ||
| rolling-origin evaluation: it repeatedly hides the tail of the series, | ||
| forecasts it, and scores against what actually happened. | ||
|
|
||
| ```sql | ||
| SELECT * FROM backtest('SELECT ts, value FROM readings', 24, | ||
| '{"model":"theta-classic","folds":5}'); | ||
| ``` | ||
|
|
||
| ## The metrics that matter | ||
|
|
||
| - **MASE** is the headline: error relative to the seasonal-naive floor. | ||
| Below 1.0 beats naive; above 1.0 means the model is losing to | ||
| last-season carry-forward and you should not serve it on this series. | ||
| Compare models by MASE on the same series and folds. | ||
| - **Coverage** is the fraction of actuals that landed inside the | ||
| prediction interval. Compare it to the confidence level you asked for: | ||
| a 0.90 band with 0.60 measured coverage is lying to you. | ||
| - Per-fold rows expose stability. A model that wins on average but | ||
| swings wildly across folds is riskier than a slightly worse, steady | ||
| one. | ||
|
|
||
| ## Interval judgment | ||
|
|
||
| The default Gaussian band can be overconfident on smooth series: the | ||
| sqlite-predict benchmarks measured 0.57 coverage at a nominal 0.90 for | ||
| theta-class models on smooth gluonts series (`benchmarks/results/` in | ||
| the repo). When interval truth matters, request conformal intervals in | ||
| the same call: | ||
|
|
||
| ```sql | ||
| SELECT * FROM backtest('SELECT ts, value FROM readings', 24, | ||
| '{"model":"theta-classic", | ||
| "interval_method":"conformal"}'); | ||
| ``` | ||
|
|
||
| Conformal intervals calibrate to measured out-of-sample residuals and | ||
| substantially improve empirical coverage (they reached the nominal level | ||
| in our benchmarks, but finite folds carry no guarantee). They need | ||
| enough history to calibrate and apply to the statistical models only; a | ||
| short series makes the call refuse rather than fabricate. Always verify | ||
| with `backtest` coverage on your own series rather than trusting either | ||
| method's reputation, ours included. | ||
|
|
||
| ## Auto-selection judgment | ||
|
|
||
| Bare `forecast(ts, value, h)` lets `auto` pick per series by rolling- | ||
| origin MASE. Use `backtest` when you want to see what auto sees: | ||
|
|
||
| - If one model wins consistently across your series, pin it | ||
| (`'{"model":"theta-classic"}'`) and save the selection cost. | ||
| - If the pool is polluted (a distilled student trained for a different | ||
| regime keeps winning on stale patterns), narrow it: | ||
| `'{"candidates":["theta-classic","tsb"]}'`. | ||
| - Intermittent series (many zeros, sporadic demand) are `tsb` territory; | ||
| if auto is not picking it, check whether the series reaches it and | ||
| consider pinning. | ||
|
|
||
| ## Reading degraded outcomes | ||
|
|
||
| `backtest` needs enough history for its folds: expect loud errors or | ||
| degraded statuses on short series rather than fabricated confidence. | ||
| That is the tool working. A series too short to backtest is a series too | ||
| short to trust a model on; fall back to wider intervals and say so in | ||
| whatever the forecast feeds. | ||
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.