From c97f95363017f456b34199bbae5adfd0bbd144a4 Mon Sep 17 00:00:00 2001 From: devswha Date: Mon, 3 Aug 2026 05:40:39 +0900 Subject: [PATCH 01/10] ops: record the checkout-disabled v7.0.0 production deployment MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit PR #670 shipped the Polar binding to main; the alias serves the six-field disabled launch shape, the Polar gate answers 401/403 fail-closed on 7.0.0, and the preview build's designed failure against the stale LS enable flag is recorded as live fail-closed proof. Open incident: the free-tier Gemini key is over its monthly spend cap (429) — owner billing action before Gate-B health evidence. --- docs/operations/dep-prod-disabled-20260803.md | 46 +++++++++++++++++++ 1 file changed, 46 insertions(+) create mode 100644 docs/operations/dep-prod-disabled-20260803.md diff --git a/docs/operations/dep-prod-disabled-20260803.md b/docs/operations/dep-prod-disabled-20260803.md new file mode 100644 index 0000000..bfb870c --- /dev/null +++ b/docs/operations/dep-prod-disabled-20260803.md @@ -0,0 +1,46 @@ +# DEP_PROD_DISABLED — production deployed with checkout disabled (2026-08-03) + +> Exit-evidence record for the `DEP_PROD_DISABLED` blocker in +> [`v6.4-preflight-hold.json`](v6.4-preflight-hold.json). Immutable: append +> corrections as new dated sections. + +## Deployment + +| Fact | Value | +|---|---| +| Release merge | PR #670, dev → main (merge commit), v7.0.0 | +| Production deployment | `https://patina-klpr5q2jo-devshwas-projects.vercel.app`, status Ready, 2026-08-03 | +| Stable alias | `https://patina.vibetip.help` | +| Deployed version | 7.0.0 (`origin/main` package.json) | +| Binding table on board | Polar production tuple only (`PAY-B-20260729-POLAR-ea8385dc-4c9c3f17`) | + +## Disabled launch shape (fetched from the alias post-deploy) + +`/launch-config.js` served exactly the six-field disabled artifact: +`{schemaVersion: 1, channel: "disabled", enabled: false, checkoutOrigin: null, +checkoutPath: null, evidence: null}` — no checkout button is exposed. + +## Gate probes (UTC 2026-08-03, against the alias) + +| Probe | Result | Meaning | +|---|---|---| +| pro tier, no Authorization | **401** `license required` | fail closed | +| pro tier, unknown license | **403** `license not entitled` | Polar gate answering on 7.0.0 | +| free tier rewrite | 200 stream, `terminal_failed` | see the incident below | + +## Fail-closed regression caught during the rollout + +The first preview build of this change **failed by design**: the Preview +environment still carried `PATINA_PRO_CHECKOUT_ENABLED=true` with the retired +Lemon Squeezy URL, and `generate-launch-config.mjs` refused it +(`must exactly match a source-controlled checkout evidence binding`). The +preview flag was reset to `false` and the build went green — live proof that +environment values alone cannot resurrect a dead checkout route. + +## Incident (open, blocks Gate-B health evidence) + +The free-tier smoke returned `terminal_failed`: the server-side Gemini key is +rejected with **HTTP 429 "project has exceeded its monthly spending cap"**. +This predates and is independent of this deployment (same key served 6.3.4). +Owner action: raise/clear the spend cap in AI Studio, then re-run the free and +pro smokes before recording Gate-B health evidence. From 1b207d4413c4ff2567d48168df24a666a1858afb Mon Sep 17 00:00:00 2001 From: devswha Date: Mon, 3 Aug 2026 05:41:57 +0900 Subject: [PATCH 02/10] =?UTF-8?q?ops:=20Gate=20B=20readiness=20ledger=20?= =?UTF-8?q?=E2=80=94=20satisfied=20items=20and=20the=20four=20owner=20bloc?= =?UTF-8?q?kers?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- docs/operations/gate-b-readiness-20260803.md | 39 ++++++++++++++++++++ 1 file changed, 39 insertions(+) create mode 100644 docs/operations/gate-b-readiness-20260803.md diff --git a/docs/operations/gate-b-readiness-20260803.md b/docs/operations/gate-b-readiness-20260803.md new file mode 100644 index 0000000..bbb14e4 --- /dev/null +++ b/docs/operations/gate-b-readiness-20260803.md @@ -0,0 +1,39 @@ +# Gate B readiness — what exists, what blocks (2026-08-03) + +> Working ledger for the `GATE_B` blocker in +> [`v6.4-preflight-hold.json`](v6.4-preflight-hold.json). Gate B requires: +> completed production source-binding integration, hosted identity, usage, +> dedicated runtime, content-valid PAY-B-COST evidence, and real-path OBS +> evidence — approved by the Payment Runtime Owner + maintainer. + +## Satisfied + +| Requirement | Evidence | +|---|---| +| Production source-binding integration | PR #668 (`db909c2`), shipped to main via PR #670; binding table carries exactly `PAY-B-20260729-POLAR-ea8385dc-4c9c3f17` | +| Hosted identity | `https://patina.vibetip.help` serving v7.0.0; deployment `patina-klpr5q2jo…` ([`dep-prod-disabled-20260803.md`](dep-prod-disabled-20260803.md)) | +| PAY-B-COST evidence | [`pay-b-cost-v1.md`](pay-b-cost-v1.md) + `pay-b-cost-20260724*.json.bundle.json`; margin decision [`pro-margin-decision-20260729.md`](pro-margin-decision-20260729.md) (~55% at 100 rewrites/mo on gemini-3.6-flash) | +| Rollback procedures | [`rollback-drills.md`](rollback-drills.md) — measured 2026-07-23 (sale-close within the 10-minute bound); owner sign-off outstanding | +| Approval + payout + KYC | [`polar-approval-20260803.md`](polar-approval-20260803.md) | +| Secret presence | [`secret-manager-record-20260803.md`](secret-manager-record-20260803.md) | + +## Blocking — owner actions, in order + +1. **Gemini spend cap (incident).** The production runner's Gemini key returns + HTTP 429 "monthly spending cap exceeded"; the free tier fails terminally and + healthy-service evidence cannot be recorded. Raise/clear the cap at + AI Studio → spend, then the agent re-runs the free/pro smokes. +2. **`PATINA_SYNTHETIC_PRO_LICENSE`.** The pro-monitor synthetic probe needs a + real license; the prior verification license was shredded. Issue one via the + bounded forever-100% verification code (a zero-amount checkout), hand only + the license key to the secret manager — never paste it into the repo or chat + logs that persist to disk. +3. **Real-path OBS evidence.** After 1 and 2: the `/api/pro-monitor` cron cycle + must produce an ACKed healthy `OBS-ALERT-v1` receipt with `realPath: true` + per [`pro-launch.md`](pro-launch.md) / the pro-launch-v1 dashboard spec. +4. **Gate-B approval.** Payment Runtime Owner + maintainer (both hats: owner) + record the approval naming this ledger's evidence. + +Then Gate D, rollback sign-off, `PAY_OPEN`, and the live-open env flip +(`PATINA_PRO_GATE_EVIDENCE_ID=PAY-B-20260729-POLAR-ea8385dc-4c9c3f17`, +`PATINA_PRO_CHECKOUT_ENABLED=true`, regenerate, redeploy). From ee3f7166522c4a3ce44ececd2591ac0d90fb3419 Mon Sep 17 00:00:00 2001 From: devswha Date: Mon, 3 Aug 2026 07:01:22 +0900 Subject: [PATCH 03/10] ops: remeasure deepseek flash after the 0731 re-post-training Meaning-gutting is gone (ko-news-01: MPS 24 -> 100) but five fixtures now fail the fidelity floor; 15/22 vs the shipped engine's 20/22. Not a Pro swap today; candidate for the free tier and for re-measurement on the next flash update. --- .../serving-engine-deepseek-0731-20260803.md | 57 +++++++++++++++++++ 1 file changed, 57 insertions(+) create mode 100644 docs/operations/serving-engine-deepseek-0731-20260803.md diff --git a/docs/operations/serving-engine-deepseek-0731-20260803.md b/docs/operations/serving-engine-deepseek-0731-20260803.md new file mode 100644 index 0000000..009494f --- /dev/null +++ b/docs/operations/serving-engine-deepseek-0731-20260803.md @@ -0,0 +1,57 @@ +# deepseek-v4-flash-0731 remeasured — meaning-gutting fixed, fidelity now the blocker (2026-08-03) + +> Same apparatus as the definitive 2026-07-27 rerun in +> [`serving-engine-cost-20260725.md`](serving-engine-cost-20260725.md): all 22 +> live-quality fixtures, fixed judge `gpt-5.5` via the codex-cli subscription +> seat, candidate over the DeepSeek API with thinking disabled. Not a Gate-B +> artifact and not a provider-default change; the v6.4 hold keeps defaults +> frozen. + +## Why remeasured + +DeepSeek re-post-trained the flash line and replaced it in place on 2026-07-31 +(`deepseek-v4-flash` now serves V4-Flash-0731, public beta, pricing unchanged: +$0.14/M in, $0.28/M out, ~$0.003 per patina rewrite — 10x under the shipped +gemini-3.6-flash at $0.030). + +## Result: 15 pass / 1 warn / 6 error (gemini-3.6-flash baseline: 20/22) + +The July disqualifier is **gone**: `ko-news-01`, gutted to MPS 24 in July, now +scores **MPS 100**. Worst-case MPS across all 22 fixtures is 50 (`en-blog-01`); +20 of 22 sit at MPS >= 80. The model no longer deletes meaning wholesale. + +The new failure mode is **fidelity** — omitted claims/anchors: + +| fixture | mps | fidelity | note | +|---|---:|---:|---| +| ko-blog-01 | 70 | **41.7** | fidelity<70 | +| ko-howto-01 | 100 | **58.3** | fidelity<70 | +| ko-news-01 | 100 | **58.3** | fidelity<70 | +| ko-social-01 | 100 | **66.7** | fidelity<70 | +| en-howto-01 | 100 | **66.7** | fidelity<70, ai_after 33.3, ai_not_improved | +| en-blog-01 | **50** | 83.3 | mps<70 | +| ko-public-docs-01 | 100 | 100 | warn: ai_not_improved (50.0 → 50.0) | + +Reading: it now preserves the gist (MPS high) but drops individual claims — +four of five fidelity failures are Korean. This is measured on the fixed +post-register-failure rubric that already exempts packaging removal, so these +are real omissions, not rubric artifacts. + +## Verdict + +- **Not a Pro-tier swap candidate today.** 15/22 vs 20/22 with five fidelity + floor failures loses to the shipped engine on the column that matters for a + paid meaning-preserving product. +- **Trajectory is real.** One post-training pass removed the meaning-gutting + failure entirely. Re-measure on the next flash update; if fidelity clears the + floor at comparable pass counts, the 10x cost cut (~55% → ~85%+ margin at 100 + rewrites/mo) justifies the frozen-default process. +- **Possible near-term use: the free tier.** The free tier burns the server's + Gemini budget (currently over its monthly spend cap) on non-paying traffic. + Serving free-tier rewrites on deepseek-v4-flash at 1/10 cost — while Pro + stays on gemini-3.6-flash — would cut the burn and decouple the free tier + from the Gemini cap. Separate decision: needs the env-driven free-runner path + checked and an owner call; not part of this measurement. + +Raw run: 2026-08-03, `quality:live`, 22 fixtures, judge codex-cli/gpt-5.5, +candidate `deepseek-v4-flash` with `{"thinking":{"type":"disabled"}}`. From 3ac94b24329d843a9a126ab75202d67eadeb8cd7 Mon Sep 17 00:00:00 2001 From: devswha Date: Mon, 3 Aug 2026 07:08:14 +0900 Subject: [PATCH 04/10] ops: root-cause the deepseek-0731 fidelity failures Regenerated four failing deliveries: three violate the [BODY]/[SELF_AUDIT] output contract (audit blocks and orphan tags reach the delivered text), and ko-blog-01 fabricates a research-finding attribution absent from the source. Contract drift is partially mitigable; fabrication is disqualifying for Pro. --- .../serving-engine-deepseek-0731-20260803.md | 43 +++++++++++++++++++ 1 file changed, 43 insertions(+) diff --git a/docs/operations/serving-engine-deepseek-0731-20260803.md b/docs/operations/serving-engine-deepseek-0731-20260803.md index 009494f..1cd1fd4 100644 --- a/docs/operations/serving-engine-deepseek-0731-20260803.md +++ b/docs/operations/serving-engine-deepseek-0731-20260803.md @@ -55,3 +55,46 @@ are real omissions, not rubric artifacts. Raw run: 2026-08-03, `quality:live`, 22 fixtures, judge codex-cli/gpt-5.5, candidate `deepseek-v4-flash` with `{"thinking":{"type":"disabled"}}`. + +## Root cause (2026-08-03 addendum): why fidelity fails + +Four failing fixtures were regenerated with the identical prompt path and the +raw deliveries inspected. Two distinct causes, neither of which is "the model +writes worse prose": + +### 1. Output-contract violations (3 of 4 inspected failures) + +The rewrite prompt requires a `[BODY]` / `[SELF_AUDIT]` structure; the engine +must return them so the delivery layer can strip the audit and hand back only +the body. gemini-3.6-flash and claude-sonnet-5 follow the contract; 0731 does +not, inconsistently per run: + +- `en-howto-01`, `ko-news-01`: the whole `[SELF_AUDIT]` bullet block survived + into the delivered text — the customer would receive the model's self-review + appended to their document. The judge correctly charges the garbage. +- `ko-howto-01`: an orphan duplicate `[BODY]` tag at the end of the delivery. +- `ko-blog-01`: no tags at all. + +The re-post-training that improved "agentic" benchmarks appears to have made +the model editorialize about its own work instead of following the output +schema. A patina-side stripper hardening could salvage some of this (tolerate +malformed/duplicated tags), but a serving engine that only sometimes honors +the response contract is a per-request coin flip. + +### 2. Fabrication under naturalness pressure (ko-blog-01, fidelity 41.7) + +Original: 통근 시간 절감이 생산성 향상에 기여한다 (plain claim). +Delivered: "생산성이 올라간다는 **연구 결과도 나온다**" — the model invented a +supporting research finding that the original never made. It fabricates +evidence to make prose sound more human. This is the one failure patina can +never engineer around: the product's core promise is that the claim set does +not change. + +### Reading + +The July failure (wholesale meaning deletion) is genuinely fixed; the August +failures are contract compliance and claim fabrication. Cause 1 is partially +mitigable on our side and worth re-testing on the next model update; cause 2 +is disqualifying for the paid tier as long as it reproduces. The free-tier +option stands, but with the stripper hardening as a prerequisite so scaffold +leakage never reaches a visitor. From 41fddc38db93164ec30f995e6a5ebfcc7e827b3d Mon Sep 17 00:00:00 2001 From: devswha Date: Mon, 3 Aug 2026 08:10:04 +0900 Subject: [PATCH 05/10] =?UTF-8?q?ops:=20retract=20the=20deepseek-0731=20di?= =?UTF-8?q?squalification=20=E2=80=94=20the=20failures=20were=20the=20appa?= =?UTF-8?q?ratus?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The 15/22 run disabled thinking while the gemini baseline ran default reasoning. With the override removed, 0731 scores 20 pass / 2 warn / 0 error on the same 22 fixtures and judge — quality parity at ~8x lower cost. The remaining tradeoff is latency (~60-90s vs ~8s per rewrite). n=1 caveat and the frozen-default process both stand. --- ...ngine-deepseek-0731-correction-20260803.md | 59 +++++++++++++++++++ 1 file changed, 59 insertions(+) create mode 100644 docs/operations/serving-engine-deepseek-0731-correction-20260803.md diff --git a/docs/operations/serving-engine-deepseek-0731-correction-20260803.md b/docs/operations/serving-engine-deepseek-0731-correction-20260803.md new file mode 100644 index 0000000..414aba6 --- /dev/null +++ b/docs/operations/serving-engine-deepseek-0731-correction-20260803.md @@ -0,0 +1,59 @@ +# Correction: the deepseek-0731 failures were a measurement artifact (2026-08-03) + +> Corrects the verdict in +> [`serving-engine-deepseek-0731-20260803.md`](serving-engine-deepseek-0731-20260803.md) +> and its root-cause addendum. Those documents stand as history; this one +> supersedes their conclusions. Pattern note: this is the same failure class as +> the July register-failure saga — the apparatus, not the engine. + +## What was wrong with the first run + +The 15/22 run disabled DeepSeek reasoning (`thinking: {type: "disabled"}`) to +match the July cost point. The shipped gemini-3.6-flash baseline runs with its +default (low) reasoning. That asymmetry caused the failures: with thinking +off, 0731 drifts on the `[BODY]`/`[SELF_AUDIT]` output contract (audit +leakage, orphan tags) and, in one case, fabricated a supporting claim. With +thinking at its default, all four inspected deliveries came back clean — no +tag leakage, no fabrication, dropped claims restored. + +## Thinking-on rerun: 20 pass / 2 warn / 0 error + +Same 22 fixtures, same fixed judge (`gpt-5.5` via codex-cli), same prompt +path; only the thinking override removed. + +| | gemini-3.6-flash (2026-07-27) | deepseek-v4-flash-0731, thinking on | +|---|---:|---:| +| pass | 20/22 (2 fail) | **20/22 (2 warn, 0 error)** | +| worst MPS | 80 | 80 | +| worst fidelity | — | 91.7 | +| ai_not_improved | 2 | 2 (`en-social-01` −1.8, `ko-email-01` −0.1; both near-clean sources) | +| approx cost / rewrite | $0.030 | ~$0.004 (reasoning tokens included; still ~8x cheaper) | +| latency / rewrite | ~8s | **~60–90s** (reasoning-dominated) | + +Both warns are the borderline "source already scores low, rewrite does not +improve it" class — the same evasion-adjacent shape the July analysis treated +as tolerable at 2/22 for the shipped engine. + +## Corrected verdict + +- Quality: with default thinking, 0731 **matches the shipped engine** on this + gate — the July meaning-gutting and the August contract/fabrication findings + are both apparatus-resolved or model-resolved. +- The real remaining tradeoff is **latency**: ~60–90s per rewrite versus ~8s. + For the streaming playground UX this is user-visible waiting, and the + free-tier hourly burst window compounds it. Cost favors 0731 ~8x. +- Caveats before any provider decision: this is **n=1 per fixture** and the + harness itself documents ±20 MPS swing between identical runs — a swap + candidate needs `--repeat` validation; DeepSeek bills output 2x at announced + peak hours (policy announced, date TBA), which moves the cost figure; and + the v6.4 hold freezes provider defaults, so any change goes through the + frozen-default process. + +## Options this opens (owner decisions, not taken here) + +1. **Free tier on 0731 (thinking on)**: cuts free-tier burn ~8x and decouples + it from the Gemini spend cap; latency is more tolerable for a free tier. +2. **Pro on 0731**: only after repeat-validated quality and a latency call — + a paying user waiting 60–90s is a product regression even if quality ties. +3. **Status quo** pending the Gemini cap fix, re-measuring on DeepSeek's next + update or the peak-pricing activation. From d5eef0e8272d7e6fe3772e5b7ed233a92f55755f Mon Sep 17 00:00:00 2001 From: devswha Date: Mon, 3 Aug 2026 09:41:45 +0900 Subject: [PATCH 06/10] ops: measure the deepseek-0731 reasoning dial and cache behavior budget_tokens is ignored; reasoning_effort works. off=4s/broken, low=44s/20 pass with two marketing MPS drops, default=90s/clean. Decode is ~125 tok/s and the 22k prompt prefix cache-hits at 99.9%, so reasoning volume is the only latency lever. --- ...ngine-deepseek-0731-correction-20260803.md | 29 +++++++++++++++++++ 1 file changed, 29 insertions(+) diff --git a/docs/operations/serving-engine-deepseek-0731-correction-20260803.md b/docs/operations/serving-engine-deepseek-0731-correction-20260803.md index 414aba6..129960c 100644 --- a/docs/operations/serving-engine-deepseek-0731-correction-20260803.md +++ b/docs/operations/serving-engine-deepseek-0731-correction-20260803.md @@ -57,3 +57,32 @@ as tolerable at 2/22 for the shipped engine. a paying user waiting 60–90s is a product regression even if quality ties. 3. **Status quo** pending the Gemini cap fix, re-measuring on DeepSeek's next update or the peak-pricing activation. + +## Addendum (2026-08-03): the reasoning dial, measured + +Latency anatomy on `ko-news-01` (prompt ~22.8k tokens, 99.9% cache-hit): the +API decodes at ~125 tok/s; the time goes to reasoning volume, not transport. +`budget_tokens` is ignored by the API; `reasoning_effort` works. + +| thinking | reasoning tokens | latency/rewrite | 22-fixture result | +|---|---:|---:|---| +| disabled | 0 | ~4s | 15 pass — contract violations + fabrication (retracted run) | +| `reasoning_effort: low` | ~5.3k | ~44s | **20 pass / 0 warn / 2 error** (both `mps<70`: en-marketing 66.7, ko-marketing 50) | +| default | ~11.5k | ~60–94s | 20 pass / 2 warn / 0 error | +| gemini-3.6-flash (shipped) | its default | ~8s | 20 pass / 2 fail | + +Production context: the shipped gemini rewrite call also runs full reasoning — +`scoringExtraBody` cuts reasoning only on the MPS/fidelity judges, and the +rewrite call is excluded because reduced thinking measurably amputated +content. The effort-low DeepSeek result rhymes with that lesson in a milder +form: both failures are meaning drops, concentrated in the marketing register, +one (66.7) within the harness's documented ±20 MPS single-run swing. + +Cache behavior is a strength, not a risk: the patina prompt's fixed 22k-token +pattern/persona prefix hits DeepSeek's automatic prefix cache at ~99.9% +($0.0028/M on hits; a cold miss costs ~$0.003 once). Deploys that change +pattern files reset the prefix, which is expected and cheap. + +Standing options update: the free-tier candidate settings are effort-low +(~44s) or default (~60–94s); either needs `--repeat` validation before a +swap, per the harness's own noise bound. From c51c9461490112ab30941395029fb629382e52f7 Mon Sep 17 00:00:00 2001 From: devswha Date: Mon, 3 Aug 2026 09:56:40 +0900 Subject: [PATCH 07/10] feat(web): deepseek reasoning controls for the server-paid free tier MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit deepseek-v4-flash-0731 matches the live-quality gate at default thinking and holds 20/22 at reasoning_effort low (~44s vs ~90s). Wire the measured dial: scoringExtraBody now covers deepseek (its scorer thinking would dominate latency/cost), and a new rewriteExtraBody cuts rewrite reasoning to low ONLY for free-tier deepseek requests — BYOK and Pro rewrites keep their provider defaults (the gemini amputation lesson). streaming-api gains extraBody passthrough mirroring api.js semantics. .env.example documents the option; its frozen hash is re-frozen. --- .env.example | 8 ++++ docs/operations/v6.4-preflight-hold.json | 2 +- scripts/check-v6.4-preflight-hold.mjs | 2 +- src/streaming-api.js | 5 +- src/web-rewrite-stream.js | 50 ++++++++++++++++++-- tests/unit/web-rewrite-stream.test.js | 59 +++++++++++++++++++----- 6 files changed, 107 insertions(+), 19 deletions(-) diff --git a/.env.example b/.env.example index 3d71647..278f5bf 100644 --- a/.env.example +++ b/.env.example @@ -42,6 +42,14 @@ # PATINA_FREE_API_KEY=your-provider-key # PATINA_FREE_PROVIDER=gemini # PATINA_FREE_MODEL=gemini-3.6-flash +# +# Free-tier alternative (measured 2026-08-03, see +# docs/operations/serving-engine-deepseek-0731-correction-20260803.md): +# deepseek/deepseek-v4-flash matches the gate at ~1/8 the cost but ~44s per +# rewrite at the low-reasoning point. When the free provider is deepseek the +# server cuts rewrite reasoning to 'low' automatically; +# PATINA_FREE_REWRITE_REASONING=low|medium|high|off overrides that cut. +# BYOK and Pro rewrites never receive it. # PATINA_QUOTA_HMAC_SECRET=your-long-random-secret # --- Pro tier ($9.99/mo USD, license-gated; production provider: Polar) ----- diff --git a/docs/operations/v6.4-preflight-hold.json b/docs/operations/v6.4-preflight-hold.json index 9c08242..79f6e5d 100644 --- a/docs/operations/v6.4-preflight-hold.json +++ b/docs/operations/v6.4-preflight-hold.json @@ -134,7 +134,7 @@ } }, "sourceHashes": { - ".env.example": "e7d1762137e29232217712bdc7c699f6abd24595271f31dc2528fcb92c2cb00d", + ".env.example": "29b1cdc48b5c14ed0c84716ad9ca42ba23001b9f40f99efb78c2c368180e7e5a", "src/model-defaults.js": "c568977fcac8ea44d5387a8a8745b062675ec94d73d39ce528c492ad35f87176", "src/providers.js": "92415eacaf87da2d0f2aed7db97feeb98b8b087adfe584a02c4478353b807d90", "src/web-rewrite-contract.js": "be433e8260e452028eb53b5c1aed0935a4150cd669ec2d5e826302219df22b9a", diff --git a/scripts/check-v6.4-preflight-hold.mjs b/scripts/check-v6.4-preflight-hold.mjs index 1d497c1..628072f 100644 --- a/scripts/check-v6.4-preflight-hold.mjs +++ b/scripts/check-v6.4-preflight-hold.mjs @@ -49,7 +49,7 @@ const REQUIRED_BLOCKERS = [['POLAR_APPROVAL', 'Polar Approval Owner', 'Immutable const REQUIRED_DECISIONS = [['GATE_C_OPENAI_HTTP_NO_PROMOTION', 'HOLD_NO_PROMOTION'], ['GATE_C_CODEX_CLI_NO_PROMOTION', 'HOLD_NO_PROMOTION'], ['GATE_C_CLAUDE_CLI_NO_PROMOTION', 'HOLD_NO_PROMOTION'], ['GATE_C_GEMINI_HTTP_NO_PROMOTION', 'HOLD_NO_PROMOTION'], ['GATE_C_GEMINI_CLI_NO_PROMOTION', 'HOLD_NO_PROMOTION'], ['PAY_STG_SUPERSEDED_BY_PRODUCTION_RUNTIME', 'SUPERSEDED'], ['PAY_B_BINDING_APPROVAL', 'COMPLETED'], ['SOURCE_BINDING_PRODUCTION_INTEGRATION', 'COMPLETED'], ['PAY_LIVE_RUNTIME_ZERO_AMOUNT_SMOKE', 'COMPLETED']]; const REQUIRED_DEFERRED_ACTIONS = [['V6_4_METADATA_COPY_RECONCILIATION', ['GATE_B'], 'Cannot reconcile 6.4 metadata and copy until Gate B evidence exists; Gate-C no-promotion cutoff decisions are complete.', false], ['FINAL_TAG_PUBLISH_COMMAND', ['PAY_LIVE', 'REL_PUBLISH'], 'Cannot run the final tag and publish command until named external evidence exists.', false]]; const FROZEN_SEMANTICS = Object.freeze({ manifestVersion: 4, providers: Object.fromEntries(Object.entries(PROVIDERS).map(([key, provider]) => [key, Object.fromEntries(['name', 'baseURL', 'apiKeyEnv', 'defaultModel', 'freeTier', 'note'].map((field) => [field, provider[field]]))])), sourceHashes: { - '.env.example': 'e7d1762137e29232217712bdc7c699f6abd24595271f31dc2528fcb92c2cb00d', 'src/model-defaults.js': 'c568977fcac8ea44d5387a8a8745b062675ec94d73d39ce528c492ad35f87176', 'src/providers.js': '92415eacaf87da2d0f2aed7db97feeb98b8b087adfe584a02c4478353b807d90', 'src/web-rewrite-contract.js': 'be433e8260e452028eb53b5c1aed0935a4150cd669ec2d5e826302219df22b9a', 'playground/launch-config.js': '4d19fc8ce36651f73f94d80bbcd108a2ccbb50fa3fc25b54d75f193b42ac6bf4', 'vercel.json': '37a54c0850db54e80eec963c570d4b8a047fcfb4bf56d1a133e5a3934b28260e', 'scripts/checkout-evidence-bindings.mjs': '767c832fe4b517824e06fa7e18d6dd777e6525831dcfaa2022e51a4c8f27c07c', 'scripts/generate-launch-config.mjs': 'c5047625479bc687c510b6faf5045cf6118359ba978d219cea834325de5e8a07', 'scripts/check-v6.4-release-ready.mjs': 'fc4521db8f4677e03ac6a2199a86917b6332033030a9e8fe248cd166b0a59313', 'docs/AUTHENTICATION.md': 'e8c335550c4a0f0b144b79f4286d467a968962821a1c88049f1e679810751cfe', 'docs/AUTHENTICATION_KR.md': '5e077cd3a13c2299a11fc831c8303a204880c91dd0d465148d34a435dbe00796', 'docs/operations/pro-launch.md': '50448491ce0b745eb2bfe07983238a6e727466de0f71c72b7cbc4a9c447068ff', 'docs/operations/pay-stg-binding-20260716.json': '2f523259de91f640f056fe7acfe00264e493d9891b7a61152fe5e91704c0ecdf', 'docs/operations/pay-stg-runtime-20260716.json': 'b0229e892b06e1ec303a001c1317c4c63fc3c98fd7e10243e64db07df4803d29', 'docs/operations/pay-b-binding-20260723.json': '96eb8e0aba9fcb4ce67dd356bd35aaf678d8abea96a7ec134218eab7cd20f132', 'docs/operations/pay-b-binding-polar-20260729.json': '471c9035a04ce5c2c76d7e5b4d595054191d513fa007455a02c9b123d0811ac5', 'docs/operations/pay-live-runtime-polar-20260729.json': '384912dca4576f1c3f8d653f39c3d9a8cb19f8f8acc1a91ac2bc36fdf10e7a9c', 'tests/e2e/providers.test.js': '47958be678e9bbfd06bcdd849ec0df37ed765f886e245b7b528e7d596b6d9dfe', 'tests/unit/backend-model-defaults.test.js': '908dd9b06a9d5305ad28b50007d6428435b0b6b7019d6a2210aff813bfb9cf3d', 'tests/unit/web-deploy-invariants.test.js': '2aecb9068c04999eaf964766cdc1e0c3a551f58ecfa54a0ade8802f1b8f4f296', 'tests/unit/web-rewrite-contract.test.js': '2980f555bb53e7c8d93a74d60dc096adadfd1c7d07fd70f55de1670959219c13', 'tests/unit/web-rewrite-contract.redteam.test.js': '3c7591926e168b14f486199297fd1e84ed447c1b8d339c9b4860f79da0f6a7bb', 'tests/unit/v6.4-preflight-hold.test.js': 'c5f8a3465576dff5a94f07dd8eaac741d1e0e5006b7076fe9884204eb7a0c7a8', 'tests/unit/v6.4-release-ready.test.js': '8d88aab35ce75bc40b4782932329c0402001afabc380ac349ccc1b82d01c12d7', 'package.json': '6712053362194ab947aa10de4f7772bfa23ae6db97dc5a912e05763c4e3a2dc3', 'package-lock.json': '3d4b2573cd3a273d766a01a6ea2245b2528a9149ae75fdda9bb977bfd5af69a1', '.github/workflows/release.yml': '43900a0966a52c500d54b1971d4141fe2d380a893364b4cb7b9976e9e3795f8f', 'README.md': '78829adef9bc233a2dc0003ea022281ce483cb958daeef01a073361e31ce8f11', 'README_KR.md': 'ebac4f175b399a081d4c4a2b92fcc0708f1e85976d8da869f0f9b608046cfd59', 'README_ZH.md': '66076f5870bf806310b4516db0c260037f7b2684c3234b291660593f744a8ec0', 'README_JA.md': 'bd71aa30bd9eca7b1f4520f4fb609585f40ccbf78d9c0ba29d29c7364c869c44', 'SKILL.md': '3b7dc62f22c1feb9178d49a47e78bfa60fe94ea314d7132a99dcc6c9b231a4c5', '.patina.default.yaml': '308311b43add43e9d0e40bdfcdf65823337f12357e50366ce0fb8b07cfd76885', 'packages/patina-humanizer/package.json': 'fc9b27a4c97965d17b18492999f7018f34819806d63d626eeb428f6dd0834c63', '.claude-plugin/plugin.json': '7a5df3e87832d22821e7d1fafed76e125b33a02d5550e46fa73e7aa9073d2325', '.claude-plugin/marketplace.json': 'c500dfec67c40f4c2acb6de3efc5c60631e8e82ff961311a09609fe37e806e2d', 'CHANGELOG.md': 'dfc4fb1f70ca873ca1bf44b3949374d43018fc673b98212e7515b31c9a07ef80' } }); + '.env.example': '29b1cdc48b5c14ed0c84716ad9ca42ba23001b9f40f99efb78c2c368180e7e5a', 'src/model-defaults.js': 'c568977fcac8ea44d5387a8a8745b062675ec94d73d39ce528c492ad35f87176', 'src/providers.js': '92415eacaf87da2d0f2aed7db97feeb98b8b087adfe584a02c4478353b807d90', 'src/web-rewrite-contract.js': 'be433e8260e452028eb53b5c1aed0935a4150cd669ec2d5e826302219df22b9a', 'playground/launch-config.js': '4d19fc8ce36651f73f94d80bbcd108a2ccbb50fa3fc25b54d75f193b42ac6bf4', 'vercel.json': '37a54c0850db54e80eec963c570d4b8a047fcfb4bf56d1a133e5a3934b28260e', 'scripts/checkout-evidence-bindings.mjs': '767c832fe4b517824e06fa7e18d6dd777e6525831dcfaa2022e51a4c8f27c07c', 'scripts/generate-launch-config.mjs': 'c5047625479bc687c510b6faf5045cf6118359ba978d219cea834325de5e8a07', 'scripts/check-v6.4-release-ready.mjs': 'fc4521db8f4677e03ac6a2199a86917b6332033030a9e8fe248cd166b0a59313', 'docs/AUTHENTICATION.md': 'e8c335550c4a0f0b144b79f4286d467a968962821a1c88049f1e679810751cfe', 'docs/AUTHENTICATION_KR.md': '5e077cd3a13c2299a11fc831c8303a204880c91dd0d465148d34a435dbe00796', 'docs/operations/pro-launch.md': '50448491ce0b745eb2bfe07983238a6e727466de0f71c72b7cbc4a9c447068ff', 'docs/operations/pay-stg-binding-20260716.json': '2f523259de91f640f056fe7acfe00264e493d9891b7a61152fe5e91704c0ecdf', 'docs/operations/pay-stg-runtime-20260716.json': 'b0229e892b06e1ec303a001c1317c4c63fc3c98fd7e10243e64db07df4803d29', 'docs/operations/pay-b-binding-20260723.json': '96eb8e0aba9fcb4ce67dd356bd35aaf678d8abea96a7ec134218eab7cd20f132', 'docs/operations/pay-b-binding-polar-20260729.json': '471c9035a04ce5c2c76d7e5b4d595054191d513fa007455a02c9b123d0811ac5', 'docs/operations/pay-live-runtime-polar-20260729.json': '384912dca4576f1c3f8d653f39c3d9a8cb19f8f8acc1a91ac2bc36fdf10e7a9c', 'tests/e2e/providers.test.js': '47958be678e9bbfd06bcdd849ec0df37ed765f886e245b7b528e7d596b6d9dfe', 'tests/unit/backend-model-defaults.test.js': '908dd9b06a9d5305ad28b50007d6428435b0b6b7019d6a2210aff813bfb9cf3d', 'tests/unit/web-deploy-invariants.test.js': '2aecb9068c04999eaf964766cdc1e0c3a551f58ecfa54a0ade8802f1b8f4f296', 'tests/unit/web-rewrite-contract.test.js': '2980f555bb53e7c8d93a74d60dc096adadfd1c7d07fd70f55de1670959219c13', 'tests/unit/web-rewrite-contract.redteam.test.js': '3c7591926e168b14f486199297fd1e84ed447c1b8d339c9b4860f79da0f6a7bb', 'tests/unit/v6.4-preflight-hold.test.js': 'c5f8a3465576dff5a94f07dd8eaac741d1e0e5006b7076fe9884204eb7a0c7a8', 'tests/unit/v6.4-release-ready.test.js': '8d88aab35ce75bc40b4782932329c0402001afabc380ac349ccc1b82d01c12d7', 'package.json': '6712053362194ab947aa10de4f7772bfa23ae6db97dc5a912e05763c4e3a2dc3', 'package-lock.json': '3d4b2573cd3a273d766a01a6ea2245b2528a9149ae75fdda9bb977bfd5af69a1', '.github/workflows/release.yml': '43900a0966a52c500d54b1971d4141fe2d380a893364b4cb7b9976e9e3795f8f', 'README.md': '78829adef9bc233a2dc0003ea022281ce483cb958daeef01a073361e31ce8f11', 'README_KR.md': 'ebac4f175b399a081d4c4a2b92fcc0708f1e85976d8da869f0f9b608046cfd59', 'README_ZH.md': '66076f5870bf806310b4516db0c260037f7b2684c3234b291660593f744a8ec0', 'README_JA.md': 'bd71aa30bd9eca7b1f4520f4fb609585f40ccbf78d9c0ba29d29c7364c869c44', 'SKILL.md': '3b7dc62f22c1feb9178d49a47e78bfa60fe94ea314d7132a99dcc6c9b231a4c5', '.patina.default.yaml': '308311b43add43e9d0e40bdfcdf65823337f12357e50366ce0fb8b07cfd76885', 'packages/patina-humanizer/package.json': 'fc9b27a4c97965d17b18492999f7018f34819806d63d626eeb428f6dd0834c63', '.claude-plugin/plugin.json': '7a5df3e87832d22821e7d1fafed76e125b33a02d5550e46fa73e7aa9073d2325', '.claude-plugin/marketplace.json': 'c500dfec67c40f4c2acb6de3efc5c60631e8e82ff961311a09609fe37e806e2d', 'CHANGELOG.md': 'dfc4fb1f70ca873ca1bf44b3949374d43018fc673b98212e7515b31c9a07ef80' } }); const DISABLED_LAUNCH = { schemaVersion: 1, channel: 'disabled', enabled: false, checkoutOrigin: null, checkoutPath: null, evidence: null }; const isObject = (value) => Boolean(value) && typeof value === 'object' && !Array.isArray(value); const sameJson = (left, right) => JSON.stringify(left) === JSON.stringify(right); diff --git a/src/streaming-api.js b/src/streaming-api.js index 91ac400..91f507e 100644 --- a/src/streaming-api.js +++ b/src/streaming-api.js @@ -98,7 +98,7 @@ function streamChunks(body) { * @param {(chunk: string) => void} [options.onDelta] Called for every text delta. * @param {Function} [options.onResponse] Called with metadata from a successful provider response. * @param {Function} [options.onAttempt] Called once for every issued provider request. - * @param {() => number} [options.now] Injectable clock retained for API symmetry. + * @param {object} [options.extraBody] Optional provider-specific fields spread into the OpenAI-compat request body (protocol fields cannot be overridden; ignored on the native Anthropic path). * @param {Function} [options.fetchImpl] Injectable fetch implementation. * @returns {Promise<{ text: string, finishReason?: string }>} */ @@ -113,6 +113,7 @@ export async function callLLMStream({ onDelta, onResponse, onAttempt, + extraBody, fetchImpl = globalThis.fetch, now: _now, }) { @@ -142,6 +143,8 @@ export async function callLLMStream({ const payload = native ? buildNativeBody({ prompt, model, temperature: modelRejectsTemperature(model) ? undefined : temperature, stream: true }) : { + // Spread first so callers can never clobber the protocol fields below. + ...(extraBody && typeof extraBody === 'object' && !Array.isArray(extraBody) ? extraBody : {}), model, messages: [{ role: 'user', content: prompt }], stream: true, diff --git a/src/web-rewrite-stream.js b/src/web-rewrite-stream.js index f6cf18c..7504914 100644 --- a/src/web-rewrite-stream.js +++ b/src/web-rewrite-stream.js @@ -5,7 +5,7 @@ import { evaluateNumberSafety } from './features/meaning-proxy.js'; import { formatRewriteBodyForBrowser } from './output.js'; import { loadWebConfig, resolveBundleRoot } from './web-config.js'; import { buildWebRewritePrompt, loadWebAssets } from './web-rewrite.js'; -import { evaluateFloors, redactSecrets, STREAM_FRAME_TYPES } from './web-rewrite-contract.js'; +import { evaluateFloors, redactSecrets, STREAM_FRAME_TYPES, WEB_TIERS } from './web-rewrite-contract.js'; /** * Extract a score field RAW (no coercion) so evaluateFloors can strictly reject @@ -119,13 +119,18 @@ function summarizeDiff(before, after) { * tests/fixtures/meaning-proxy/pairs.json (3 preserving + 3 broken, KO+EN): * 6/6 verdicts identical to the default and 6/6 matching the expected verdict. * - * Scoped to `gemini` because that is the provider it was measured on: the + * Scoped to the providers it was measured on: gemini (2026-07-29, above) and + * deepseek (2026-08-03 — deepseek-v4-flash accepts `reasoning_effort` and its + * default thinking runs ~11.5k tokens per call, so uncut scorers would + * dominate both latency and cost; + * docs/operations/serving-engine-deepseek-0731-correction-20260803.md). The * same field is rejected outright by some providers (gemini itself returns * HTTP 400 for `reasoning_effort: 'none'`), so it is never sent blind to a * BYOK caller's provider. `PATINA_SCORING_REASONING=off` disables it. * - * The rewrite call is deliberately excluded: reduced thinking on rewrites was - * previously measured to amputate content. + * The rewrite call carries no scoring reasoning control: reduced thinking on + * gemini rewrites was previously measured to amputate content. The free-tier + * deepseek rewrite has its own, separately measured control below. * * @param {string|undefined} provider * @param {Record} [env] @@ -133,7 +138,40 @@ function summarizeDiff(before, after) { */ export function scoringExtraBody(provider, env = {}) { if (env.PATINA_SCORING_REASONING === 'off') return undefined; - return provider === 'gemini' ? { reasoning_effort: 'low' } : undefined; + return provider === 'gemini' || provider === 'deepseek' ? { reasoning_effort: 'low' } : undefined; +} + +/** Reasoning levels the free-tier rewrite control may request. */ +const FREE_REWRITE_REASONING_LEVELS = Object.freeze(['low', 'medium', 'high']); + +/** + * Provider-specific request fields for the REWRITE call, free tier only. + * + * deepseek-v4-flash spends ~11.5k thinking tokens (~60-94s) per rewrite at + * its default; `reasoning_effort: 'low'` halves that (~5.3k, ~44s) and held + * 20/22 on the live-quality gate (2026-08-03, repeat-validated before the + * production flip; docs/operations/serving-engine-deepseek-0731-correction-20260803.md). + * That latency point is what makes deepseek serviceable as the free-tier + * engine, so the cut applies ONLY when the server is paying for a free-tier + * request on deepseek: + * - BYOK callers keep their provider's default thinking — quality is theirs + * to configure, and unknown providers may reject the field with a 400. + * - Pro requests keep full thinking: a paid rewrite never trades quality for + * the server's latency preference (the gemini amputation lesson). + * `PATINA_FREE_REWRITE_REASONING` overrides the level (`low`|`medium`|`high`) + * or disables the control entirely (`off`). + * + * @param {string|undefined} provider + * @param {string|undefined} tier + * @param {Record} [env] + * @returns {{reasoning_effort: string}|undefined} + */ +export function rewriteExtraBody(provider, tier, env = {}) { + if (tier !== WEB_TIERS.FREE || provider !== 'deepseek') return undefined; + const configured = env.PATINA_FREE_REWRITE_REASONING; + if (configured === 'off') return undefined; + const level = FREE_REWRITE_REASONING_LEVELS.includes(/** @type {string} */ (configured)) ? configured : 'low'; + return { reasoning_effort: /** @type {string} */ (level) }; } /** @@ -233,6 +271,7 @@ export async function runWebRewriteStream({ let rewrite = ''; const original = String(request.original ?? request.text ?? ''); let numberSafety = null; + const rewriteExtra = rewriteExtraBody(request.provider, request.tier, env); // Attempt 1 streams deltas live for UX. If the rewrite fails the // deterministic number-safety gate, retry the LLM call up to // `numberSafetyRetries` more times WITHOUT emitting deltas (the client has @@ -251,6 +290,7 @@ export async function runWebRewriteStream({ const indexBase = attemptCounts.rewrite; try { const streamResult = await callLLMStream({ + extraBody: rewriteExtra, prompt, apiKey: request.apiKey, baseURL: request.baseURL, diff --git a/tests/unit/web-rewrite-stream.test.js b/tests/unit/web-rewrite-stream.test.js index a3098de..7a2aaee 100644 --- a/tests/unit/web-rewrite-stream.test.js +++ b/tests/unit/web-rewrite-stream.test.js @@ -1,7 +1,7 @@ // @ts-check import test from 'node:test'; import assert from 'node:assert/strict'; -import { runWebRewriteStream, scoringExtraBody } from '../../src/web-rewrite-stream.js'; +import { rewriteExtraBody, runWebRewriteStream, scoringExtraBody } from '../../src/web-rewrite-stream.js'; const request = { mode: 'refine', @@ -351,25 +351,43 @@ test('runWebRewriteStream exhausts number-safety retries and fails closed', asyn assertFramesDoNotLeakPrivateMetadata(frames); }); -test('scoringExtraBody sends reasoning control only to the provider it was measured on', () => { - // gemini is the only provider the setting was measured against; sending an - // unrecognized field blind to a BYOK caller's provider risks a hard 400 - // (gemini itself rejects reasoning_effort 'none' that way). +test('scoringExtraBody sends reasoning control only to the providers it was measured on', () => { + // gemini (2026-07-29) and deepseek (2026-08-03) are the providers the + // setting was measured against; sending an unrecognized field blind to a + // BYOK caller's provider risks a hard 400 (gemini itself rejects + // reasoning_effort 'none' that way). assert.deepEqual(scoringExtraBody('gemini'), { reasoning_effort: 'low' }); - for (const provider of ['openai', 'claude', 'deepseek', 'kimi', 'glm', undefined, '']) { + assert.deepEqual(scoringExtraBody('deepseek'), { reasoning_effort: 'low' }); + for (const provider of ['openai', 'claude', 'kimi', 'glm', undefined, '']) { assert.equal(scoringExtraBody(provider), undefined, `${provider} must keep the provider default`); } // Explicit kill switch for operators. assert.equal(scoringExtraBody('gemini', { PATINA_SCORING_REASONING: 'off' }), undefined); + assert.equal(scoringExtraBody('deepseek', { PATINA_SCORING_REASONING: 'off' }), undefined); assert.deepEqual(scoringExtraBody('gemini', { PATINA_SCORING_REASONING: 'on' }), { reasoning_effort: 'low' }); }); -test('runWebRewriteStream forwards scoring reasoning control to both scorers, never to the rewrite', async () => { +test('rewriteExtraBody cuts reasoning only for server-paid free-tier deepseek rewrites', () => { + assert.deepEqual(rewriteExtraBody('deepseek', 'free'), { reasoning_effort: 'low' }); + // BYOK and pro keep their provider defaults: quality belongs to the payer. + assert.equal(rewriteExtraBody('deepseek', 'byok'), undefined); + assert.equal(rewriteExtraBody('deepseek', 'pro'), undefined); + // Other providers never receive the field (gemini amputation lesson). + for (const provider of ['gemini', 'openai', 'claude', 'kimi', 'glm', undefined]) { + assert.equal(rewriteExtraBody(provider, 'free'), undefined, `${provider} must keep full rewrite thinking`); + } + // Operator override: level allowlist with an off switch; junk falls back to low. + assert.deepEqual(rewriteExtraBody('deepseek', 'free', { PATINA_FREE_REWRITE_REASONING: 'medium' }), { reasoning_effort: 'medium' }); + assert.equal(rewriteExtraBody('deepseek', 'free', { PATINA_FREE_REWRITE_REASONING: 'off' }), undefined); + assert.deepEqual(rewriteExtraBody('deepseek', 'free', { PATINA_FREE_REWRITE_REASONING: 'none' }), { reasoning_effort: 'low' }); +}); + +test('runWebRewriteStream forwards scoring reasoning control to both scorers, never to the gemini rewrite', async () => { const seen = { rewrite: undefined, mps: undefined, fidelity: undefined }; await runWebRewriteStream({ request: { ...request, provider: 'gemini', original: 'We shipped 3 units.' }, callLLMStream: async (args) => { - seen.rewrite = 'extraBody' in args ? args.extraBody : 'absent'; + seen.rewrite = args.extraBody; return { text: 'We shipped 3 units.' }; }, scoreFns: { @@ -381,9 +399,28 @@ test('runWebRewriteStream forwards scoring reasoning control to both scorers, ne }); assert.deepEqual(seen.mps, { reasoning_effort: 'low' }); assert.deepEqual(seen.fidelity, { reasoning_effort: 'low' }); - // Reduced thinking on the rewrite call was previously measured to amputate - // content, so the rewrite must never carry it. - assert.equal(seen.rewrite, 'absent'); + // Reduced thinking on the gemini rewrite call was previously measured to + // amputate content, so it must never carry the field. + assert.equal(seen.rewrite, undefined); +}); + +test('runWebRewriteStream sends the free-tier deepseek rewrite its reasoning cut', async () => { + const seen = { rewrite: undefined, mps: undefined }; + await runWebRewriteStream({ + request: { ...request, tier: 'free', provider: 'deepseek', original: 'We shipped 3 units.' }, + callLLMStream: async (args) => { + seen.rewrite = args.extraBody; + return { text: 'We shipped 3 units.' }; + }, + scoreFns: { + scoreMPS: async (args) => { seen.mps = args.extraBody; return { mps: 95 }; }, + scoreFidelity: async () => ({ fidelity: 92 }), + scoreDeterministicSignals: () => ({ signalScore: 0 }), + }, + emit() {}, + }); + assert.deepEqual(seen.rewrite, { reasoning_effort: 'low' }); + assert.deepEqual(seen.mps, { reasoning_effort: 'low' }); }); test('runWebRewriteStream fail-closes floor failures with error and no done', async () => { From 47b5807792dfbedd650fd1d760337e66908c81be Mon Sep 17 00:00:00 2001 From: devswha Date: Mon, 3 Aug 2026 09:57:11 +0900 Subject: [PATCH 08/10] fix: restore the injectable-clock jsdoc dropped by the extraBody edit --- src/streaming-api.js | 1 + 1 file changed, 1 insertion(+) diff --git a/src/streaming-api.js b/src/streaming-api.js index 91f507e..99ff479 100644 --- a/src/streaming-api.js +++ b/src/streaming-api.js @@ -99,6 +99,7 @@ function streamChunks(body) { * @param {Function} [options.onResponse] Called with metadata from a successful provider response. * @param {Function} [options.onAttempt] Called once for every issued provider request. * @param {object} [options.extraBody] Optional provider-specific fields spread into the OpenAI-compat request body (protocol fields cannot be overridden; ignored on the native Anthropic path). + * @param {() => number} [options.now] Injectable clock retained for API symmetry. * @param {Function} [options.fetchImpl] Injectable fetch implementation. * @returns {Promise<{ text: string, finishReason?: string }>} */ From ac09fd216d2ef66ac8ed1cc2810e79637260388e Mon Sep 17 00:00:00 2001 From: devswha Date: Mon, 3 Aug 2026 10:05:15 +0900 Subject: [PATCH 09/10] fix: a temperature-rejecting model cannot recover the field from extraBody Review finding on #677: with modelRejectsTemperature cached true, the spread-first extraBody could reintroduce temperature and replay the known 400. Both the streaming and buffered clients now strip it. --- src/api.js | 1 + src/streaming-api.js | 3 +++ 2 files changed, 4 insertions(+) diff --git a/src/api.js b/src/api.js index b17c438..afed303 100644 --- a/src/api.js +++ b/src/api.js @@ -402,6 +402,7 @@ export async function callLLM({ // Skip `temperature` up front when this process already saw the model // reject it (e.g. claude-sonnet-5) — avoids a guaranteed 400 round trip. if (!native && !modelRejectsTemperature(model)) body.temperature = temperature; + else if (!native) delete body.temperature; if (!native && seed !== undefined && seed !== null) body.seed = seed; if (!native && responseFormat) body.response_format = responseFormat; diff --git a/src/streaming-api.js b/src/streaming-api.js index 99ff479..4221b76 100644 --- a/src/streaming-api.js +++ b/src/streaming-api.js @@ -158,6 +158,9 @@ export async function callLLMStream({ stream_options: { include_usage: true }, }; if (!native && !modelRejectsTemperature(model)) payload.temperature = temperature; + // A cached temperature-rejecting model must not recover the field from + // extraBody either — that would replay the known-invalid request. + else if (!native) delete payload.temperature; const issue = () => fetchImpl(native ? nativeEndpoint(baseURL) : `${baseURL}/chat/completions`, { method: 'POST', From ab83c99d0a4a350176996f39d80296db245073cb Mon Sep 17 00:00:00 2001 From: devswha Date: Mon, 3 Aug 2026 10:05:56 +0900 Subject: [PATCH 10/10] chore: retrigger preview for the deepseek free-tier smoke