Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -57,3 +57,6 @@ bindings/node/node_modules/
bindings/rust/csrc/
bindings/rust/target/
bindings/rust/Cargo.lock

# autogluon working dirs (Mitra benchmark runs dump model copies here)
benchmarks/AutogluonModels/
11 changes: 11 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -139,6 +139,17 @@ permissively licensed (MIT/Apache-2.0), so you ship it inside your product
rather than rent it. And it composes with the rest of the in-database AI
toolbox: sqlite-vec gave SQLite vector search; this gives it prediction.

### Agent skills

For agents, the SQL surface is the whole API, so what ships in
[`skills/`](skills/) is judgment, not wrappers: four
[agentskills.io](https://agentskills.io)-format skills covering core
usage, provenance receipts (record what was asked, which model answered,
and a hash that replays; serving determinism makes this verifiable),
the distillation lifecycle (verify before serving, re-distill on drift,
license discipline), and backtest interpretation. Install them into any
skills-capable agent, or read them as the condensed operator's manual.

## Operations

| Function | Question | Returns |
Expand Down
13 changes: 13 additions & 0 deletions benchmarks/results/scale-arena.jsonl
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
{"dataset": "Diabetes130US", "n_train": 1500, "n_test": 500, "d": 40, "xgboost": 0.906, "xgb_s": 0.1, "tabicl": 0.916, "tabicl_s": 2.2, "tabicl_device": "mps", "leak_gap": 0.0007, "ic_maxp": 0.9251, "collapsed": 1, "gbt<-tabicl soft": 0.916, "oof_s": 7.6, "oof_gap": 0.0007, "gbt<-tabicl soft-oof": 0.916}
{"dataset": "Diabetes130US", "n_train": 9999, "n_test": 3334, "d": 40, "xgboost": 0.9070185962807439, "xgb_s": 0.1, "tabicl": 0.9130173965206959, "tabicl_s": 101.9, "tabicl_device": "mps", "leak_gap": -0.0, "ic_maxp": 0.922, "collapsed": 1, "gbt<-tabicl soft": 0.9130173965206959, "oof_s": 116.4, "oof_gap": -0.0, "gbt<-tabicl soft-oof": 0.9130173965206959}
{"dataset": "airline_satisfaction", "n_train": 1500, "n_test": 500, "d": 21, "xgboost": 0.91, "xgb_s": 0.0, "tabicl": 0.928, "tabicl_s": 1.6, "tabicl_device": "mps", "leak_gap": 0.0693, "ic_maxp": 0.8787, "gbt<-tabicl": 0.908, "distill_s": 0.5, "blob_kb": 103.4, "gbt<-tabicl soft": 0.904, "oof_s": 5.5, "oof_gap": -0.0073, "gbt<-tabicl soft-oof": 0.906}
{"dataset": "GiveMeSomeCredit", "n_train": 1500, "n_test": 500, "d": 10, "xgboost": 0.922, "xgb_s": 0.0, "tabicl": 0.936, "tabicl_s": 1.2, "tabicl_device": "mps", "leak_gap": 0.0067, "ic_maxp": 0.9407, "gbt<-tabicl": 0.932, "distill_s": 0.3, "blob_kb": 97.1, "gbt<-tabicl soft": 0.926, "oof_s": 4.1, "oof_gap": -0.0113, "gbt<-tabicl soft-oof": 0.93}
{"dataset": "GiveMeSomeCredit", "n_train": 9999, "n_test": 3334, "d": 10, "xgboost": 0.9289142171565686, "xgb_s": 0.1, "tabicl": 0.9322135572885423, "tabicl_s": 8.4, "tabicl_device": "mps", "leak_gap": 0.021, "ic_maxp": 0.9452, "gbt<-tabicl": 0.9334133173365327, "distill_s": 2.6, "blob_kb": 104.3, "gbt<-tabicl soft": 0.9334133173365327, "oof_s": 29.5, "oof_gap": 0.0009, "gbt<-tabicl soft-oof": 0.9322135572885423}
{"dataset": "APSFailure", "n_train": 1500, "n_test": 500, "d": 40, "xgboost": 0.994, "xgb_s": 0.1, "tabicl": 0.996, "tabicl_s": 2.6, "tabicl_device": "mps", "leak_gap": -0.004, "ic_maxp": 0.9902, "gbt<-tabicl": 0.99, "distill_s": 1.4, "blob_kb": 76.6, "gbt<-tabicl soft": 0.99, "oof_s": 8.2, "oof_gap": -0.0093, "gbt<-tabicl soft-oof": 0.996}
{"dataset": "SDSS17", "n_train": 1500, "n_test": 500, "d": 11, "xgboost": 0.97, "xgb_s": 0.1, "tabicl": 0.972, "tabicl_s": 1.4, "tabicl_device": "mps", "leak_gap": 0.0167, "ic_maxp": 0.9892, "gbt<-tabicl": 0.968, "distill_s": 1.1, "blob_kb": 145.8, "gbt<-tabicl soft": 0.968, "oof_s": 4.3, "oof_gap": 0.0007, "gbt<-tabicl soft-oof": 0.964}
{"dataset": "SDSS17", "n_train": 9999, "n_test": 3334, "d": 11, "xgboost": 0.9679064187162567, "xgb_s": 0.3, "tabicl": 0.9718056388722256, "tabicl_s": 8.7, "tabicl_device": "mps", "leak_gap": 0.0205, "ic_maxp": 0.9906, "gbt<-tabicl": 0.9658068386322736, "distill_s": 9.2, "blob_kb": 159.2, "gbt<-tabicl soft": 0.9661067786442712, "oof_s": 30.3, "oof_gap": 0.0052, "gbt<-tabicl soft-oof": 0.9676064787042592}
{"dataset": "taiwanese_bankruptcy", "n_train": 1500, "n_test": 500, "d": 40, "xgboost": 0.966, "xgb_s": 0.1, "tabicl": 0.968, "tabicl_s": 2.4, "tabicl_device": "mps", "leak_gap": 0.004, "ic_maxp": 0.9721, "gbt<-tabicl": 0.968, "distill_s": 2.9, "blob_kb": 93.8, "gbt<-tabicl soft": 0.966, "oof_s": 8.2, "oof_gap": 0.0007, "gbt<-tabicl soft-oof": 0.968}
{"dataset": "online_shoppers_intention", "n_train": 1500, "n_test": 500, "d": 17, "xgboost": 0.882, "xgb_s": 0.1, "tabicl": 0.89, "tabicl_s": 1.5, "tabicl_device": "mps", "leak_gap": 0.0507, "ic_maxp": 0.8863, "gbt<-tabicl": 0.89, "distill_s": 0.4, "blob_kb": 100.9, "gbt<-tabicl soft": 0.888, "oof_s": 4.9, "oof_gap": 0.0153, "gbt<-tabicl soft-oof": 0.886}
{"dataset": "credit_card_default", "n_train": 1500, "n_test": 500, "d": 23, "xgboost": 0.81, "xgb_s": 0.1, "tabicl": 0.836, "tabicl_s": 1.7, "tabicl_device": "mps", "leak_gap": -0.0247, "ic_maxp": 0.8101, "gbt<-tabicl": 0.798, "distill_s": 0.9, "blob_kb": 93.5, "gbt<-tabicl soft": 0.81, "oof_s": 5.8, "oof_gap": -0.0153, "gbt<-tabicl soft-oof": 0.838}
{"dataset": "airline_satisfaction", "n_train": 9999, "n_test": 3334, "d": 21, "xgboost": 0.9448110377924415, "xgb_s": 0.3, "tabicl": 0.9511097780443911, "tabicl_s": 139.6, "tabicl_device": "cpu", "leak_gap": 0.0465, "ic_maxp": 0.9842, "gbt<-tabicl": 0.9337132573485303, "distill_s": 3.5, "blob_kb": 105.7, "gbt<-tabicl soft": 0.9340131973605279, "oof_s": 486.7, "oof_gap": 0.0043, "gbt<-tabicl soft-oof": 0.9337132573485303}
{"dataset": "credit_card_default", "n_train": 9999, "n_test": 3334, "d": 23, "xgboost": 0.8047390521895621, "xgb_s": 0.2, "tabicl": 0.8191361727654469, "tabicl_s": 167.5, "tabicl_device": "cpu", "leak_gap": 0.0205, "ic_maxp": 0.8387, "gbt<-tabicl": 0.8197360527894421, "distill_s": 7.4, "blob_kb": 96.8, "gbt<-tabicl soft": 0.8182363527294542, "oof_s": 571.6, "oof_gap": -0.001, "gbt<-tabicl soft-oof": 0.8194361127774445}
25 changes: 25 additions & 0 deletions benchmarks/results/scale.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
# Scale study: distillation beyond 1500 rows

TabICL v2 as the teacher on the naturally large TabArena
datasets, at increasing training sizes. Questions: does student
retention hold, where does tuned xgboost catch the zero-shot
teacher, and does the in-context label leak grow with context
(the gap column; see permissive-teachers.md).

Device: mps with cpu fallback (per-cell `dev`).

| dataset | n_train | xgboost | TabICL | gbt<-TabICL | soft | soft-oof | leak gap | oof gap | teacher s | distill s | blob KB |
|---|---|---|---|---|---|---|---|---|---|---|---|
| APSFailure | 1500 | 0.994 | 0.996 | 0.990 | 0.990 | 0.996 | -0.0040 | -0.0093 | 2.6 | 1.4 | 76.6 |
| Diabetes130US | 1500 | 0.906 | 0.916 | - | 0.916 | 0.916 | 0.0007 | 0.0007 | 2.2 | - | - |
| Diabetes130US | 9999 | 0.907 | 0.913 | - | 0.913 | 0.913 | -0.0000 | -0.0000 | 101.9 | - | - |
| GiveMeSomeCredit | 1500 | 0.922 | 0.936 | 0.932 | 0.926 | 0.930 | 0.0067 | -0.0113 | 1.2 | 0.3 | 97.1 |
| GiveMeSomeCredit | 9999 | 0.929 | 0.932 | 0.933 | 0.933 | 0.932 | 0.0210 | 0.0009 | 8.4 | 2.6 | 104.3 |
| SDSS17 | 1500 | 0.970 | 0.972 | 0.968 | 0.968 | 0.964 | 0.0167 | 0.0007 | 1.4 | 1.1 | 145.8 |
| SDSS17 | 9999 | 0.968 | 0.972 | 0.966 | 0.966 | 0.968 | 0.0205 | 0.0052 | 8.7 | 9.2 | 159.2 |
| SDSS17 | 49999 | 0.975 | 0.610 | - | 0.610 | - | -0.0000 | - | 118.9 | - | - |
| airline_satisfaction | 1500 | 0.910 | 0.928 | 0.908 | 0.904 | 0.906 | 0.0693 | -0.0073 | 1.6 | 0.5 | 103.4 |
| credit_card_default | 1500 | 0.810 | 0.836 | 0.798 | 0.810 | 0.838 | -0.0247 | -0.0153 | 1.7 | 0.9 | 93.5 |
| credit_card_default | 9999 | 0.805 | 0.771 | 0.776 | 0.776 | 0.776 | 0.0019 | 0.0054 | 11.7 | 7.4 | 96.2 |
| online_shoppers_intention | 1500 | 0.882 | 0.890 | 0.890 | 0.888 | 0.886 | 0.0507 | 0.0153 | 1.5 | 0.4 | 100.9 |
| taiwanese_bankruptcy | 1500 | 0.966 | 0.968 | 0.968 | 0.966 | 0.968 | 0.0040 | 0.0007 | 2.4 | 2.9 | 93.8 |
84 changes: 84 additions & 0 deletions skills/distill-lifecycle/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,84 @@
---
name: distill-lifecycle
description: >-
Distill a teacher model into a native sqlite-predict student and manage
it over time: verify quality before serving, re-distill on drift, and
respect the teacher's license. Use when per-call serving needs to be
instant and self-contained, when a prediction runs at volume, or when a
model must travel inside the database file.
license: MIT OR Apache-2.0
metadata:
version: "0.1.0"
---

# The distillation lifecycle

Distillation compresses a teacher (your own labels, your existing model's
predictions, or a licensed foundation model) into a native student the
serving core executes with no runtime. Measured on the repo's benchmark
suite: student blobs run tens to hundreds of kilobytes, single-row
serving is sub-millisecond on CPU, and fits complete in seconds at
benchmark scale (see `benchmarks/results/` in the sqlite-predict repo).
The lifecycle judgment is yours.

## Distill

```sql
-- tabular: target column holds labels or your model's predictions
SELECT model_id, holdout_metric FROM distill_predict(
'SELECT f1, f2, label FROM training',
'{"target":"label","student_id":"churn-v1","student_kind":"gbt"}');

-- time series: from windows, or with a registered onnx teacher
SELECT model_id FROM distill_forecast('SELECT series_key, value FROM obs',
'{"context":96,"horizon":24,"student_id":"traffic-v1"}');
```

Students register in `_predict_models` as content-hashed rows; they
snapshot, fork, and sync with the database. Serve by name:
`predict(NULL, apply_sql, '{"model":"churn-v1"}')` or
`forecast(ts, value, 24, '{"model":"traffic-v1"}')`.

## Verify before serving

Never serve a student on faith:

1. Read `holdout_metric` from the distill call. It is measured on a
holdout of your data. Compare it to a floor you trust (the majority
class, last-value carry-forward, or your current model).
2. For forecast students, run `backtest` on the same series and compare
the student against `theta-classic` and the naive floor (see the
interpret-backtest skill).
3. Know the measured shape of distillation loss: on our benchmark suite
classification students give up a median half point of accuracy
against their teacher, but regression tails are worse. Check
regression students more skeptically.
4. Soft-label distillation (`proba`/`classes`) preserves the teacher's
calibration when the teacher emits probabilities, and usually rescues
imbalanced datasets where hard labels collapse to one class (the
distiller refuses loudly on collapse rather than fitting a
constant).

## Watch for drift, re-distill cheaply

A student is frozen; the world is not. When input distributions move,
quality decays silently, so schedule verification rather than assuming:

- Re-run `backtest` (forecast) or score a fresh labeled sample (tabular)
on a cadence proportional to how fast the data changes.
- Re-distilling costs seconds. Register the new student under a
versioned id (`churn-v2`), verify, then switch the serving call. Keep
the old row until the new one is trusted; retire it after.

## Licensing is part of the lifecycle

A student derives from its teacher, and the teacher's license may
impose obligations on what you distill and distribute; what applies
depends on the exact weights, license version, and how the student is
used. Record the teacher's license alongside each student you register.
Your own labels and models carry no such limits. Restrictively licensed
teachers require an explicit `accept_license` opt-in before they will
run, but that gate is an acknowledgment, not compliance: read the
license of the exact weights you use before distilling for anything
beyond evaluation. The project documentation's license notes are
orientation, not legal advice.
79 changes: 79 additions & 0 deletions skills/interpret-backtest/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,79 @@
---
name: interpret-backtest
description: >-
Read sqlite-predict backtest() output to choose models, trust or
distrust prediction intervals, and decide when auto-selection needs
narrowing. Use before serving forecasts that feed decisions, when
intervals matter, or when choosing between models on a specific series.
license: MIT OR Apache-2.0
metadata:
version: "0.1.0"
---

# Interpreting backtest()

`backtest` answers "how would this model have done on my series" with
rolling-origin evaluation: it repeatedly hides the tail of the series,
forecasts it, and scores against what actually happened.

```sql
SELECT * FROM backtest('SELECT ts, value FROM readings', 24,
'{"model":"theta-classic","folds":5}');
```

## The metrics that matter

- **MASE** is the headline: error relative to the seasonal-naive floor.
Below 1.0 beats naive; above 1.0 means the model is losing to
last-season carry-forward and you should not serve it on this series.
Compare models by MASE on the same series and folds.
- **Coverage** is the fraction of actuals that landed inside the
prediction interval. Compare it to the confidence level you asked for:
a 0.90 band with 0.60 measured coverage is lying to you.
- Per-fold rows expose stability. A model that wins on average but
swings wildly across folds is riskier than a slightly worse, steady
one.

## Interval judgment

The default Gaussian band can be overconfident on smooth series: the
sqlite-predict benchmarks measured 0.57 coverage at a nominal 0.90 for
theta-class models on smooth gluonts series (`benchmarks/results/` in
the repo). When interval truth matters, request conformal intervals in
the same call:

```sql
SELECT * FROM backtest('SELECT ts, value FROM readings', 24,
'{"model":"theta-classic",
"interval_method":"conformal"}');
```
Comment thread
coderabbitai[bot] marked this conversation as resolved.

Conformal intervals calibrate to measured out-of-sample residuals and
substantially improve empirical coverage (they reached the nominal level
in our benchmarks, but finite folds carry no guarantee). They need
enough history to calibrate and apply to the statistical models only; a
short series makes the call refuse rather than fabricate. Always verify
with `backtest` coverage on your own series rather than trusting either
method's reputation, ours included.

## Auto-selection judgment

Bare `forecast(ts, value, h)` lets `auto` pick per series by rolling-
origin MASE. Use `backtest` when you want to see what auto sees:

- If one model wins consistently across your series, pin it
(`'{"model":"theta-classic"}'`) and save the selection cost.
- If the pool is polluted (a distilled student trained for a different
regime keeps winning on stale patterns), narrow it:
`'{"candidates":["theta-classic","tsb"]}'`.
- Intermittent series (many zeros, sporadic demand) are `tsb` territory;
if auto is not picking it, check whether the series reaches it and
consider pinning.

## Reading degraded outcomes

`backtest` needs enough history for its folds: expect loud errors or
degraded statuses on short series rather than fabricated confidence.
That is the tool working. A series too short to backtest is a series too
short to trust a model on; fall back to wider intervals and say so in
whatever the forecast feeds.
Loading