Skip to content

upstream-sync: merge simstudioai/sim main into feat/github-merge-agent (2026-08-04) - #680

Closed
utcarshsrivastava-collab wants to merge 534 commits into
feat/github-merge-agentfrom
upstream-sync/2026-08-04T11-16-45
Closed

upstream-sync: merge simstudioai/sim main into feat/github-merge-agent (2026-08-04)#680
utcarshsrivastava-collab wants to merge 534 commits into
feat/github-merge-agentfrom
upstream-sync/2026-08-04T11-16-45

Conversation

@utcarshsrivastava-collab

@utcarshsrivastava-collab utcarshsrivastava-collab commented Aug 4, 2026

Copy link
Copy Markdown
Collaborator

Upstream sync — 2026-08-04

Merges simstudioai/sim@1b9e0f25 into feat/github-merge-agent.

Sync range: 518 commit(s) since e2fecc86 (merge-base).

Ledger

Verification (advisory — does not block the sync)

bun run check · ❌ bun run lint · ❌ bun run test · ❌ bun run build

Verification is advisory — failures do not block the sync. Review and fix on the draft PR as needed.

bun run check

✅ passed


   • Packages in scope: @sim/audit, @sim/auth, @sim/browser-protocol, @sim/db, @sim/desktop, @sim/desktop-bridge, @sim/emcn, @sim/logger, @sim/pii, @sim/platform-authz, @sim/realtime, @sim/realtime-protocol, @sim/runtime-secrets, @sim/security, @sim/terminal-protocol, @sim/testing, @sim/tsconfig, @sim/utils, @sim/workflow-persistence, @sim/workflow-renderer, @sim/workflow-types, docs, sim, simstudio, simstudio-ts-sdk
   • Running format:check in 25 packages
   • Remote caching disabled

::group::@sim/workflow-renderer:format:check
cache miss, executing 301c97b96789e31b
$ biome format .
Checked 13 files in 42ms. No fixes applied.
::endgroup::
::group::@sim/workflow-types:format:check
cache miss, executing 9801fd5c9dad946e
$ biome format .
Checked 4 files in 52ms. No fixes applied.
::endgroup::
::group::simstudio:format:check
cache miss, executing cd777feb95071583
$ biome format .
Checked 3 files in 53ms. No fixes applied.
::endgroup::
::group::@sim/desktop-bridge:format:check
cache miss, executing 310341c4f3fc1977
$ biome format .
Checked 5 files in 88ms. No fixes applied.
::endgroup::
::group::@sim/security:format:check
cache miss, executing dde4fddfb28e4e8e
$ biome format .
Checked 19 files in 72ms. No fixes applied.
::endgroup::
::group::@sim/realtime-protocol:format:check
cache miss, executing f08ec28fc8f510e6
$ biome format .
Checked 11 files in 101ms. No fixes applied.
::endgroup::
::group::simstudio-ts-sdk:format:check
cache miss, executing d86292ec5adf4fd6
$ biome format .
Checked 6 files in 43ms. No fixes applied.
::endgroup::
::group::@sim/terminal-protocol:format:check
cache miss, executing 5cce5b8f88069bbf
$ biome format .
Checked 3 files in 33ms. No fixes applied.
::endgroup::
::group::@sim/logger:format:check
cache miss, executing 1de180f6c370b3a1
$ biome format .
Checked 6 files in 51ms. No fixes applied.
::endgroup::
::group::@sim/realtime:format:check
cache miss, executing 08a9fd281fc06caa
$ biome format .
Checked 52 files in 415ms. No fixes applied

bun run lint

❌ failed


$ biome check --write --unsafe .
Checked 102 files in 2s. No fixes applied.
::endgroup::
::group::@sim/desktop:lint
cache miss, executing f23b644adac463f3
$ biome check --write --unsafe .
Checked 134 files in 3s. No fixes applied.
::endgroup::
::group::@sim/db:lint
cache miss, executing 5b2cb843cd20bcad
$ biome check --write --unsafe .
Checked 309 files in 9s. No fixes applied.
::endgroup::
�[;31msim:lint�[;0m
cache miss, executing 2d375b30871ee342
$ biome check --write --unsafe .
lib/webhooks/providers/zoho-desk.test.ts:50:9 suppressions/unused ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  ! Suppression comment has no effect. Remove the suppression or make sure you are suppressing the correct rule.
  
    48 │     it('rejects requests without the X-ZDesk-JWT header', async () => {
    49 │       const result = await zohoDeskHandler.verifyAuth?.(
  > 50 │         // biome-ignore lint/suspicious/noExplicitAny: minimal context for the header-only path
       │         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    51 │         makeAuthContext({}, { orgId: '1', webhookId: '2' }) as any
    52 │       )
  

lib/webhooks/providers/zoho-desk.test.ts:61:9 suppressions/unused ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  ! Suppression comment has no effect. Remove the suppression or make sure you are suppressing the correct rule.
  
    59 │         accountId: 'acct-1',
    60 │         userId: 'u1',
  > 61 │         // biome-ignore lint/suspicious/noExplicitAny: partial owner shape is enough for this path
       │         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    62 │       } as any)
    63 │       await zohoDeskHandler.verifyAuth?.(
  

lib/webhooks/providers/zoho-desk.test.ts:67:11 suppressions/unused ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  ! Suppression comment has no effect. Remove the suppression or make sure you are suppressing the correct rule.
  
    65 │           { 'x-zdesk-jwt': 'not-a-real-jwt' },
    66 │           { orgId: '1', externalId: '2', credentialId: 'cred-1' }
  > 67 │           // biome-ignore lint/suspicious/noExplicitAny: minimal context for the fallback path
       │           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    68 │         ) as any
    69 │       )
  

lib/webhooks/providers/zoho-desk.test.ts:78:11 suppressions/unused ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  ! Suppression comment has no effect. Remove the suppression or make sure you are suppressing the correct rule.
  
    76 │           { 'x-zdesk-jwt': 'not-a-real-jwt' },
    77 │           { orgId: '1', externalId: '2', credentialId: 'cred-1', apiDomain: 'https://desk.zoho.eu' }
  > 78 │           // biome-ignore lint/suspicious/noExplicitAny: minimal context for the fast path
       │           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    79 │         ) as any
    80 │       )
  

lib/webhooks/providers/zoho-desk.test.ts:93:11 suppressions/unused ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  ! Suppression comment has no effect. Remove the suppression or make sure you are suppressing the correct rule.
  
    91 │           userId: 'user-1',
    92 │           requestId: 'test',
  > 93 │           // biome-ignore lint/suspicious/noExplicitAny: request is unused on these guard paths
       │           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    94 │           request: {} as any,
    95 │         })
  

lib/webhooks/providers/zoho-desk.test.ts:250:9 suppressions/unused ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  ! Suppression comment has no effect. Remove the suppression or make sure you are suppressing the correct rule.
  
    248 │         accountId: 'acc-1',
    249 │         userId: 'user-1',
  > 250 │         // biome-ignore lint/suspicious/noExplicitAny: partial owner shape for the test
        │         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    251 │       } as any)
    252 │       vi.mocked(refreshAccessTokenIfNeeded).mockResolvedValue('zoho-token')
  

lib/webhooks/providers/zoho-desk.test.ts:273:9 suppressions/unused ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  ! Suppression comment has no effect. Remove the suppression or make sure you are suppressing the correct rule.
  
    271 │         userId: 'user-1',
    272 │         requestId: 'test',
  > 273 │         // biome-ignore lint/suspicious/noExplicitAny: request is unused on this path
        │         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    274 │         request: {} as any,
    275 │       })
  

lib/webhooks/providers/zoho-desk.test.ts:291:9 suppressions/unused ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  ! Suppression comment has no effect. Remove the suppression or make sure you are suppressing the correct rule.
  
    289 │         accountId: 'acc-1',
    290 │         userId: 'user-1',
  > 291 │         // biome-ignore lint/suspicious/noExplicitAny: partial owner shape for the test
        │         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    292 │       } as any)
    293 │       vi.mocked(refreshAccessTokenIfNeeded).mockResolvedValue('zoho-token')
  

lib/webhooks/providers/zoho-desk.test.ts:314:9 suppressions/unused ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  ! Suppression comment has no effect. Remove the suppression or make sure you are suppressing the correct rule.
  
    312 │         userId: 'user-1',
    313 │         requestId: 'test',
  > 314 │         // biome-ignore lint/suspicious/noExplicitAny: request is unused on this path
        │         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    315 │         request: {} as any,
    316 │       })
  

app/workspace/[workspaceId]/home/components/message-content/components/special-tags/choice-blocks.ts:56:7 lint/suspicious/noShadowRestrictedNames ━━━━━━━━━━

  × Do not shadow the global "escape" property.
  
    54 │   let depth = 0
    55 │   let inString = false
  > 56 │   let escape = false
       │       ^^^^^^
    57 │ 
    58 │   for (let i = startIdx; i < text.length; i++) {
  
  i Consider renaming this variable. It's easy to confuse the origin of variables when they're named after a known global.
  

lib/core/config/api-keys.ts:62:14 lint/suspicious/noDuplicateElseIf ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  × This branch can never execute. Its condition is a duplicate or covered by previous conditions in the if-else-if chain.
  
    60 │     if (env.ZAI_API_KEY_2) keys.push(env.ZAI_API_KEY_2)
    61 │     if (env.ZAI_API_KEY_3) keys.push(env.ZAI_API_KEY_3)
  > 62 │   } else if (provider === 'xai') {
       │              ^^^^^^^^^^^^^^^^^^
    63 │     if (env.XAI_API_KEY_1) keys.push(env.XAI_API_KEY_1)
    64 │     if (env.XAI_API_KEY_2) keys.push(env.XAI_API_KEY_2)
  

Checked 12833 files in 41s. Fixed 9 files.
Found 2 errors.
Found 9 warnings.
check ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  × Some errors were emitted while running checks.
  

error: script "lint" exited with code 1

 Tasks:    22 successful, 23 total
Cached:    0 cached, 23 total
  Time:    43.151s 
Failed:    sim#lint


$ turbo run lint
::error::sim#lint: command (/home/runner/work/p2-sim/p2-sim/apps/sim) /home/runner/.bun/bin/bun run lint exited (1)
 ERROR  sim#lint: command (/home/runner/work/p2-sim/p2-sim/apps/sim) /home/runner/.bun/bin/bun run lint exited (1)
 ERROR  run failed: command  exited (1)
error: script "lint" exited with code 1

Command failed: bun run lint
$ turbo run lint
::error::sim#lint: command (/home/runner/work/p2-sim/p2-sim/apps/sim) /home/runner/.bun/bin/bun run lint exited (1)
 ERROR  sim#lint: command (/home/runner/work/p2-sim/p2-sim/apps/sim) /home/runner/.bun/bin/bun run lint exited (1)
 ERROR  run failed: command  exited (1)
error: script "lint" exited with code 1

bun run test

❌ failed


   • Packages in scope: @sim/audit, @sim/auth, @sim/browser-protocol, @sim/db, @sim/desktop, @sim/desktop-bridge, @sim/emcn, @sim/logger, @sim/pii, @sim/platform-authz, @sim/realtime, @sim/realtime-protocol, @sim/runtime-secrets, @sim/security, @sim/terminal-protocol, @sim/testing, @sim/tsconfig, @sim/utils, @sim/workflow-persistence, @sim/workflow-renderer, @sim/workflow-types, docs, sim, simstudio, simstudio-ts-sdk
   • Running test in 25 packages
   • Remote caching disabled

�[;31m@sim/db:test�[;0m
cache miss, executing 44e2af4807877cbd
$ vitest run
/usr/bin/bash: line 1: vitest: command not found
error: script "test" exited with code 127
::group::@sim/workflow-persistence:test
cache miss, executing 4807f8421681162f
::endgroup::
::group::simstudio-ts-sdk:test
cache miss, executing 8263bffcaa397a26
::endgroup::
::group::@sim/runtime-secrets:test
cache miss, executing 817c46839fa6b1f4
::endgroup::
::group::@sim/logger:test
cache miss, executing c0ea2ee57a81889c
::endgroup::
::group::@sim/audit:test
cache miss, executing c554ca230e86b28d
::endgroup::
::group::@sim/emcn:test
cache miss, executing a7d068bf40cafcb7
$ vitest run
::endgroup::
::group::@sim/security:test
cache miss, executing 51ec8e87f33b7b09
$ vitest run
::endgroup::
::group::@sim/testing:test
cache miss, executing 2bd877040331bde6
$ vitest run
::endgroup::
::group::@sim/utils:test
cache miss, executing 16081fd95cdd817a
$ vitest run
::endgroup::
::group::@sim/desktop:test
cache miss, executing 55f1076623d9eec6
::endgroup::
::group::@sim/realtime-protocol:test
cache miss, executing d22b8549a6d7ae4e
$ vitest run
::endgroup::

 Tasks:    0 successful, 12 total
Cached:    0 cached, 12 total
  Time:    897ms 
Failed:    @sim/db#test


$ turbo run test
::error::@sim/db#test: command (/home/runner/work/p2-sim/p2-sim/packages/db) /home/runner/.bun/bin/bun run test exited (127)
 ERROR  @sim/db#test: command (/home/runner/work/p2-sim/p2-sim/packages/db) /home/runner/.bun/bin/bun run test exited (127)
 ERROR  run failed: command  exited (127)
error: script "test" exited with code 127

Command failed: bun run test
$ turbo run test
::error::@sim/db#test: command (/home/runner/work/p2-sim/p2-sim/packages/db) /home/runner/.bun/bin/bun run test exited (127)
 ERROR  @sim/db#test: command (/home/runner/work/p2-sim/p2-sim/packages/db) /home/runner/.bun/bin/bun run test exited (127)
 ERROR  run failed: command  exited (127)
error: script "test" exited with code 127

bun run build

❌ failed


   • Packages in scope: @sim/audit, @sim/auth, @sim/browser-protocol, @sim/db, @sim/desktop, @sim/desktop-bridge, @sim/emcn, @sim/logger, @sim/pii, @sim/platform-authz, @sim/realtime, @sim/realtime-protocol, @sim/runtime-secrets, @sim/security, @sim/terminal-protocol, @sim/testing, @sim/tsconfig, @sim/utils, @sim/workflow-persistence, @sim/workflow-renderer, @sim/workflow-types, docs, sim, simstudio, simstudio-ts-sdk
   • Running build in 25 packages
   • Remote caching disabled

�[;31msim:build�[;0m
cache miss, executing 6cfbabf3ad0b904b
$ bun run build:sandbox-bundles && NODE_OPTIONS='--max-old-space-size=8192' next build
$ bun run ./lib/execution/sandbox/bundles/build.ts
[2026-08-04T13:09:09.639Z] [ERROR] [SandboxBundleBuild] sandbox bundle build failed {
  "message": "Bundle failed",
  "name": "AggregateError"
}
error: script "build:sandbox-bundles" exited with code 1
error: script "build" exited with code 1
::group::@sim/desktop:build
cache miss, executing c5ff5d9ec9a01dfc
$ bun run scripts/ensure-pty-prebuilds.ts && bun run scripts/build.ts
• Fetching node-pty prebuild: darwin-arm64@1.2.0-beta.12
::endgroup::
::group::docs:build
cache miss, executing 250f90d5b4cf5ad1
$ fumadocs-mdx && NODE_OPTIONS='--max-old-space-size=8192' next build
::endgroup::
::group::simstudio-ts-sdk:build
cache miss, executing 9ff1b07e7ffaa0bf
$ tsc
::endgroup::
::group::simstudio:build
cache miss, executing 9d91c9a289475760
$ tsc
::endgroup::

 Tasks:    0 successful, 5 total
Cached:    0 cached, 5 total
  Time:    1.093s 
Failed:    sim#build


$ turbo run build
::error::sim#build: command (/home/runner/work/p2-sim/p2-sim/apps/sim) /home/runner/.bun/bin/bun run build exited (1)
 ERROR  sim#build: command (/home/runner/work/p2-sim/p2-sim/apps/sim) /home/runner/.bun/bin/bun run build exited (1)
 ERROR  run failed: command  exited (1)
error: script "build" exited with code 1

Command failed: bun run build
$ turbo run build
::error::sim#build: command (/home/runner/work/p2-sim/p2-sim/apps/sim) /home/runner/.bun/bin/bun run build exited (1)
 ERROR  sim#build: command (/home/runner/work/p2-sim/p2-sim/apps/sim) /home/runner/.bun/bin/bun run build exited (1)
 ERROR  run failed: command  exited (1)
error: script "build" exited with code 1

Agent usage

child-cluster-1

  • Model: gpt-5.6-luna
  • Iterations: 1
  • Input tokens (direct): 79,521
  • Input tokens (cache read): 1,048,019
  • Input tokens (cache create): 0
  • Input tokens (total): 1,127,540
  • Output tokens: 12,502
  • Cost: $0.051867 (estimated fallback)

Totals

  • Total input tokens: 1,127,540
  • Total output tokens: 12,502
  • Primary models: gpt-5.6-luna
  • Total cost: $0.051867
  • Estimated cost (fallback): $0.051867

Cost by agent

  • child-cluster-1: $0.051867 (estimated fallback)

Draft — mark ready for review when satisfied.

waleedlatif1 and others added 30 commits July 23, 2026 17:17
…se (simstudioai#5906)

* fix(landing): restore single-paragraph hero copy on home and enterprise

* fix(landing): tighten enterprise hero description

* revert(landing): restore original home and enterprise hero headings

* improvement(landing): simplify enterprise hero description

* improvement(landing): restore building-and-managing homepage headline
… pricing (simstudioai#5913)

* improvement(library): add citations and internal links, correct stale pricing

The library had zero third-party citations across 16 posts (~51k words) and
averaged 1.2 internal links per post, with six posts at zero. Auditing every
dollar figure against the vendor's own pricing page while adding the citations
turned up several claims that had gone stale, plus one product that is being
shut down.

Corrections, each verified against the vendor's page on 2026-07-23:
- Relay.app is winding down (signups closed 2026-07-16, free accounts end
  2026-08-15, paid 2026-09-14). It was recommended as the human-in-the-loop
  pick in best-zapier-alternatives; that recommendation now points to a
  platform you can still sign up for
- Make Core is $12/mo billed monthly, not $10.59; annual saves ~15%. One post
  also described $12 as the annual rate
- Pabbly Connect lifetime starts at $349, not $249, and the Standard/Pro/
  Ultimate tier structure quoted no longer exists
- Workato publishes no pricing at all, so the "~$1,000/month" figure is
  replaced with the fact that every deal is quoted through sales
- Dify is at ~149k GitHub stars, not 131k
- n8n cloud Starter is EUR 20/mo billed annually for 2,500 executions

Verified-correct figures were left alone and given a source link: Zapier
$29.99/mo monthly for 750 tasks, Activepieces 10 free flows then $5/flow/mo,
Power Automate $15/user/mo, Lindy $49.99/mo, and the Grand View Research RPA
market figures.

Also: 59 outbound citations and 3-6 internal links per post (zero posts left
without internal links), and MDX external links now carry target=_blank plus
rel=noopener noreferrer, which the landing SEO/GEO rule requires but the MDX
anchor was not applying.

* fix(content): treat only non-Sim hosts as external links in MDX

The first pass classified any http(s) href as external, so the 33 absolute
Sim URLs already in content (sim.ai, sim.ai/slack, www.sim.ai/blog/*, and 14
docs.sim.ai pages) would have opened in a new tab with rel=noopener, which is
wrong for first-party links and contradicted the comment above the check.

Classification now compares the hostname against the apex derived from
SITE_URL, so the apex and any Sim subdomain stay same-tab. SITE_URL is used
rather than getBaseDomain() so a post renders identically in dev, preview,
and production instead of varying with NEXT_PUBLIC_APP_URL. The leading dot
in the suffix check keeps lookalikes such as evil-sim.ai external.
…oard (simstudioai#5907)

* fix(helm): correct chart docs, examples, and dead config across the board

Audit-driven accuracy pass over the chart's entire documentation surface,
verified by rendering every example against the templates:

- migrations run as an init container on the app pod, not a Job — fix the
  README component list, troubleshooting commands, and sim-helm skill refs;
  drop the dead migrations-job NetworkPolicy ingress rule
- referenced-but-never-created resources: document the GKE ManagedCertificate
  creation (values-gcp), comment out the key-file Secret mount that stuck all
  pods in ContainerCreating (values-gcp), enable certManager for the postgres
  TLS issuerRef (values-production), add the cert-manager cluster-issuer
  annotation nginx needs (values-azure)
- values-external-db: networkPolicy.egress is a list, not a map (the map
  rendered an invalid manifest); fill schema-failing placeholder host/username
- realtime >1 replica requires REDIS_URL (Socket.IO Redis adapter) — default
  examples to 1 replica with the scaling note, and warn where autoscaling
  HPAs override replicaCount
- pod anti-affinity selectors matched nothing (simstudio vs sim name label)
- kubernetes.mdx: install commands were missing required CRON_SECRET and
  postgresql password (failed at template time), wrong deployment name in
  port-forward, stale version requirements, unsupported key-remapping claim
- remove unimplemented app.secrets.existingSecret.keys from values + schema;
  fix README PDB default, cronjob list, /metrics caveat, NOTES secret count,
  Azure-only StorageClass in generic examples, dead SOCKET_SERVER_URL and
  GOOGLE_CLOUD_* env, ESO apiVersion mismatch, and skill-reference drift
- bump chart to 1.0.1

* fix(helm): review round 1 — scoped example egress, in-tab secret note, copilot Job wording

- external-db example egress scopes to a placeholder database CIDR instead of
  to: [] (which allowed every destination on 5432, defeating the isolation
  the example teaches)
- kubernetes.mdx cloud tabs state explicitly that they reuse the variables
  generated in the Installation block
- Copilot migrations really do run as a Helm-hook Job — restore Job wording
  there (only the app migrations are an init container)

* fix(helm): template-sweep fixes — telemetry validity, ESO rollout checksums, dead passwordKey knob

- telemetry: memory_limiter gets the required check_interval (collector
  failed startup validation whenever telemetry.enabled=true); the jaeger
  exporter was removed from collector-contrib in v0.86 — export to Jaeger
  via its native OTLP endpoint instead (otlp/jaeger, default port 4317)
- app/realtime rollout checksums now hash the ExternalSecret manifest too,
  mirroring the copilot pattern — with ESO enabled the inline Secret renders
  empty, so remoteRefs changes never rolled the pods
- remove the unimplemented existingSecret.passwordKey knob (values, schema,
  README, dead helpers): nothing consumed it, and a non-default value
  silently produced a DATABASE_URL with an unexpandable placeholder; secrets
  must use the standard POSTGRES_PASSWORD / EXTERNAL_DB_PASSWORD keys
- drop the orphaned sim.migrations.labels helper (its only consumer was the
  dead NetworkPolicy rule removed earlier)
- helm test pod image resolves through sim.image so global.imageRegistry
  mirroring applies; NetworkPolicy realtime-ingress comment reflects actual
  traffic direction; smoke unittest suite loads the newly referenced
  external-secret template

* chore(helm): bump chart to 1.1.0 with upgrade notes

Removing (inert) documented values keys and changing the rollout-checksum
inputs is a values-surface change — per SemVer chart conventions that is
more than a patch. Adds an Upgrading section documenting the one-time pod
roll, the removed no-op keys, and the Jaeger-over-OTLP change.

* fix(helm): review round — external-db NP opt-in with real-CIDR-first flow, prod Jaeger OTLP endpoint

- external-db example ships networkPolicy disabled so a verbatim install
  always reaches the database; the scoped egress rule stays as the documented
  opt-in (set your CIDR first, then enable)
- values-production still pointed telemetry.jaeger at the legacy 14250
  collector port — now Jaeger's OTLP gRPC endpoint to match the otlp/jaeger
  exporter

* feat(helm): autoscaling.realtime.enabled toggle so examples can scale the app without unsafe realtime replicas

Cursor correctly flagged that comment-level warnings didn't stop a verbatim
production/external-db install from running the realtime HPA at minReplicas 2
without REDIS_URL (silent cross-pod event loss). Adds an opt-out toggle
(default true — existing deployments unchanged): the realtime HPA renders only
when autoscaling.realtime.enabled, and the realtime Deployment keeps
spec.replicas under its control when the HPA is excluded. The three
autoscaling examples set it false with the Redis rationale; README and
upgrade notes document the toggle.

* fix(helm): review round — whitelabeled realtime HPA opt-out, external-db isolation on with required CIDR in install flow

- values-whitelabeled now actually sets autoscaling.realtime.enabled: false
  (the earlier batch aborted before reaching this file — Cursor caught it)
- external-db keeps networkPolicy enabled (no isolation regression); the DB
  egress CIDR is marked REQUIRED and wired into both documented install
  commands via --set, so the copy-paste flow sets the real subnet in the same
  breath as the DB host

* fix(helm): use a syntactically valid example CIDR in the external-db install commands

An unreplaced <YOUR_DB_CIDR> literal fails Kubernetes CIDR validation and
aborts the install; every sibling placeholder in the same command (host,
username) installs fine and simply doesn't connect until replaced. The CIDR
now behaves the same way: valid example value (10.20.0.0/24), explicitly
marked as the operator's database subnet in both commands and the values
comment.

* fix(helm): remove text after continuation backslashes in external-db install commands

A trailing comment after the line-continuation backslash (and equally a
comment line spliced mid-command) breaks the shell command when copied —
the remaining --set overrides run as separate commands and required-value
validation fails. Both documented commands now reconstruct to bash-clean
multi-line invocations (verified with bash -n); the CIDR guidance lives in
the networkPolicy section comment.

* fix(docs): cloud-tab installs are alternatives via helm upgrade --install

Following Installation and then a cloud tab ran helm install twice for the
same release and failed on the second. The tabs now state they replace the
generic install and use helm upgrade --install, which is idempotent and also
converts an existing generic install to the cloud values.

* fix(docs): honest conversion caveat for cloud-tab upgrades over an existing install

The cloud values rename the bundled Postgres database to simstudio, which
Postgres only applies at first initialization — an in-place conversion of a
generic install would point DATABASE_URL at a nonexistent database. Document
the two safe paths: keep the original name via --set, or uninstall + delete
PVCs and install fresh.

* fix(helm): external-db header install command declares its secrets and includes CRON_SECRET

The primary documented command failed the chart's required-value validation
(missing app.env.CRON_SECRET with cronjobs default-on) and referenced an
undeclared DB_PASSWORD. It now shows the export lines for every variable it
uses and sets CRON_SECRET; both commands verified with bash -n.

* chore(helm): restore values.schema.json formatting — surgical deletions only

The earlier programmatic edit reformatted the whole file (~390 lines of
whitespace churn hiding the 12 real deleted lines). Re-applied the removal
of the dead keys/passwordKey properties as text-level deletions preserving
the original style.

* fix(helm+docs): explicit secret exports in external-db header; conversion must reuse original secrets

- all five export lines are written out (three were only named in a trailing
  comment, so a verbatim copy passed empty required values)
- the cloud-conversion caveat now leads with reusing the original secret
  values (helm get values) — a regenerated ENCRYPTION_KEY makes previously
  encrypted credentials undecryptable
* feat(agent-stream): add agent-events thinking/tool streaming for chat and canvas

Ship the agent-events-v1 protocol with provider tool loops, dual-gated chat thinking, DeepSeek/Groq/OpenAI reasoning wiring, and ChatGPT-like thinking chrome.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): clear stuck streaming UI and format db snapshot

Biome was failing CI on migrations/meta/0261_snapshot.json. Also settle
assistant streaming/tool flags when SSE ends without a terminal frame,
without clobbering Stop's finalized content.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): satisfy biome format and import order

Auto-format the sim package for CI lint:check, and repair the Anthropic
streaming tool-loop payload after an unsafe delete-to-undefined rewrite.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): keep drained answer on abort and update migration journal test

Treat AbortError from reader.cancel as a cancelled pump result so soft-complete
retains answerText. Point the workspace storage migration journal assertion at
0261_chat_include_thinking.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): keep Stop notice when server emits cancel error

Ignore terminal SSE error frames after the user aborts so
"Client cancelled request" cannot overwrite "Response stopped by user".

Co-authored-by: Cursor <cursoragent@cursor.com>

* improvement(chat): ChatGPT-style thinking shimmer and stick-to-bottom scroll

Add left-to-right shimmer on live thinking label/body, keep scroll working by
shimmering an inner node, and follow the answer only while near the bottom.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): stop pump on client disconnect; soft-complete agents only

Abort the agent stream pump when the projected HTTP body is cancelled so
provider work does not continue after disconnect. Limit AbortError soft-success
to Agent blocks so Function/HTTP cancels still fail in logs.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): persist includeThinking across pause snapshots

Paused chat runs with Include thinking enabled were dropping the flag when
serializing the pause snapshot, so resume always rebuilt streams without
thinking/tool SSE frames.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): keep drained answer text when stream times out

Persist pump answerText onto the streaming execution before throwing on
timeout, and carry that partial content into the failed block output so
logs match what the client already saw.

Co-authored-by: Cursor <cursoragent@cursor.com>

* improvement(chat): auto-collapse tools chrome when tool streaming ends

Match thinking UX: open while tools run, collapse when finished, and keep
the panel open only if the user manually reopens it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): settle canvas stream chrome on failure paths

Clear agentStreamActive and settle running tool chips when blocks error,
timeouts cancel runs, or execution ends without stream:done so the output
panel does not stay on live Thinking/Using tools chrome.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(lint): organize imports in terminal console store

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): mark open tools cancelled on HITL pause

Pause can interrupt a tool loop without tool end events; settling those
chips as success incorrectly showed unfinished tools as complete.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(db): drop branch-local 0261 migration ahead of staging merge

* chore(db): regenerate include_thinking migration as 0266 post staging merge

* fix(providers): resolve type errors in streaming tool loop call sites

* fix(agent-stream): gate agent events opt-in and correct provider loop behavior

- streamToolCalls and provider thinking requests now require run-level
  agentEvents opt-in (canvas on, chat dual-gated, API off) so existing
  runs keep pre-agent-events behavior exactly
- OpenAI reasoning summaries opt-in + strip-and-retry on unverified-org 400
- streaming loops run tool postProcess again (firecrawl/exa async results)
- bedrock live loop falls back to silent path for responseFormat
- deepseek: reasoning_content pass-back unconditional, 'none' sends disabled
- groq: x_groq.usage fallback, reasoning params gated, qwen none disables
- gemini: functionCall parts echoed verbatim, local ids only for events
- truncated turns (max_tokens/length) no longer execute partial tool calls
- MAX_TOOL_ITERATIONS exit flushes last turn text as final answer
- iterations reports actual model calls; shared loop plumbing extracted

* refactor(agent-stream): consolidate protocol, dedupe client/server plumbing, hygiene

- canonical ChatStreamFrame union + type guards consumed by server emitters
  and the chat client; stream_error restored to legacy log-only handling
- strip thinking/tool args from providerTiming on public final envelopes
- shared tool-chip lifecycle module for chat, canvas, and console store
- shared sink-to-execution-events forwarder replaces the copy-pasted
  adapter in the execute route and HITL manager; LIVE_ONLY event set shared
- stream:thinking payload field renamed data->text; canvas thinking batched
- abort reasons carried as AbortError DOMExceptions so raw fetch consumers
  classify correctly; thinking cap renamed to chars and scope-documented
- kimi wired for agent events like the other compat providers
- deleted dead exports/step-N comments; fixtures match real wire shapes;
  loop tests use explicit mocks instead of importOriginal

* test(agent-stream): cover the dual-gated execution path and typed abort reasons

- chat route tests assert agentEvents reaches executeWorkflow only when
  policy and protocol header agree
- execution-limits tests assert AbortError-typed reasons
- executor metadata type carries agentEvents

* fix(deploy-modal): align include-thinking spacing with the modal's 6.5px rhythm

* docs(agent-stream): autogenerate per-model thinking/tool stream support on the Agent block page

- capabilities.thinking.streamed ('full' | 'summary' | 'none') on models.ts,
  explicit for the Anthropic family where visibility varies per generation;
  getThinkingStreamVisibility exposes the derivation for docs and UI alike
- scripts/sync-agent-stream-docs.ts regenerates the support tables between
  markers in workflows/blocks/agent.mdx from the model registry and
  STREAMING_TOOL_CALL_PROVIDERS; --check fails on drift or missing metadata
- wired agent-stream-docs:check into CI next to the other sync gates

* feat(anthropic): request summarized thinking display for omitted-default Claude models

The newest Claude generations (Fable 5, Sonnet 5, Opus 4.8/4.7) default
thinking.display to omitted — empty thinking blocks, no deltas. On
agent-events runs Sim now opts back in with display: 'summarized', driven
by the registry's streamed metadata; legacy runs keep the exact
pre-agent-events request shape. Registry, generated docs, and the family
capability table updated accordingly.

* docs(skills): cover thinking.streamed and agent-stream docs sync in model skills

* chore(deps): upgrade @anthropic-ai/sdk to 0.114.0 and adopt official types

- adaptive thinking, display, and output_config are now SDK-typed; the only
  remaining custom payload field is output_format (beta-header structured
  outputs, which the SDK models as output_config.format instead)
- anthropic stream events narrow on the SDK's discriminated unions instead
  of anonymous casts; compat deltas type content/tool_calls from the OpenAI
  SDK with vendor reasoning fields as an explicit optional extension
- @sim/auth exposes an explicit VerifyAuth contract so its declarations no
  longer reference better-auth's nested zod instance (TS2883 under fresh
  install layouts); realtime consumer aligned
- docs app zod pinned to the repo's exact 4.3.6 so ai SDK types bind the
  same zod instance (docs type-check was latently broken)
- knowledge embedding tests made hermetic against local .env keys and
  hosted rotation fallback

* refactor(providers): replace legacy as-any stream casts with annotated typed casts

* refactor(providers): finish provider audit — remove dead byte-stream helper, annotate remaining legacy casts

Audit of all 26 providers for the agent-events feature confirmed every
streaming execution declares agent-events-v1 and every adapter emits
AgentStreamEvent objects. Cleanup from the audit: the unconsumed legacy
createOpenAICompatibleStream byte helper is deleted, and the remaining
streamResponse-as-any casts (xai, nvidia, kimi, meta, zai, sakana) are
annotated typed casts matching the groq/deepseek fix.

* feat(streaming): stream answer text live during tool loops via turn_end protocol

The live tool loops buffered all answer text per model turn (classification
of intermediate vs final is only known at turn end), so gated surfaces saw
thinking stream, then dead air with the thinking chrome stuck open, then the
whole answer at once.

Loops now emit text deltas live as `turn: 'pending'` plus a `turn_end`
event per turn. The pump buffers pending text and projects it to the byte
path (answerText/logs/memory/legacy clients) only on a final turn_end, so
all settled semantics are unchanged. Gated surfaces render the pending text
as it streams and reconcile with a reset when a turn resolves to tools:

- public chat: live `chunk` frames from the sink + dual-gated `chunk_reset`;
  byte-path frame emission is suppressed to avoid duplicates (kept for
  response-format transformed streams via clientStreamTransformed)
- canvas: forwarder emits live `stream:chunk` + `stream:chunk_reset`; the
  execute route and HITL resume readers stop re-emitting byte chunks; panel
  chat tracks per-block segments and replaces content on flush
- chat client: per-block text segments, chunk_reset handling, and thinking
  chrome now settles on tool start as well as first answer chunk

* fix(streaming): address validated review findings across provider gating and reset reconciliation

Three-reviewer pass over the branch, findings validated against staging:

- agent-handler forwards agentEvents to executeProviderRequest — the flag was
  computed but dropped in the field-by-field copy, so provider-side thinking
  requests (OpenAI summaries, Gemini includeThoughts, Anthropic summarized
  display) never activated on opted-in runs
- openai: restore summary:'auto' alongside explicit reasoning effort — staging
  always paired them; gating summary purely on agentEvents changed legacy
  payloads
- gemini: Gemini 2 + tools + responseFormat falls back to the silent path;
  the live loop never applied the deferred responseSchema for AUTO tools
- openai-compat loop: malformed tool-argument JSON fails the call instead of
  executing with defaulted {} args (staging parsed inside the execution try)
- openai-compat parser: a vendor id arriving after a synthesized start no
  longer renames the call (start/end ids stayed consistent)
- stream-pump: abort closes the byte projection so a drain blocked on
  backpressure cannot deadlock teardown
- chunk_reset removes the block from the client text order (deployed chat +
  panel chat) so a reset block re-registers at arrival position — fixes
  separator/order corruption when parallel blocks stream around a reset
- resume route echoes the negotiated X-Sim-Stream-Protocol response header
  (parity with the chat route); docs: [DONE] wire shape + final-vs-error
  terminal semantics corrected

* chore(deps): exempt pinned @anthropic-ai/sdk 0.114.0 from the release-age gate

CI's bun install --frozen-lockfile blocks 0.114.0 (published 2026-07-23,
younger than the 7-day supply-chain gate). The pin is exact and was vetted
for the agent-events streaming work; following the existing bunfig pattern,
the exclusion ages out on 2026-07-30 and should be dropped then.

* chore(providers): fix double-cast-allowed annotation placement for the strict boundary audit

The audit only recognizes the annotation on the line directly above the cast;
two annotations had drifted behind intervening code lines (groq stream params,
deepseek loop messages) and the OpenAI reasoning-summary widening cast was
never annotated. No behavior change.

* fix(chat): settle straggler tool chips as error when final reports failure

A failed run can still terminate with a `final` frame carrying success: false;
running chips previously settled green regardless of the outcome.

* fix(canvas): wire agent stream chrome into run-from-block

Run-from-block executions emit the same live stream:thinking/stream:tool
events as full runs but registered none of the handlers, so the terminal
never showed thinking or tool chips on that path. The per-run chrome
(batched thinking writes + tool chip lifecycle + settlement on stream done,
block error, and every terminal execution state) is extracted into a shared
createAgentStreamChrome factory consumed by both paths.

---------

Co-authored-by: Bill Leoutsakos <billleoutsakos@Bills-MacBook-Pro.local>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Vikhyath Mondreti <vikhyath@simstudio.ai>
* feat(skills): permissions layer

* chore(db): drop skill_member migration 0261 for regeneration on latest staging

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(db): regenerate skill_member migration as 0262 on latest staging

Same DDL as the dropped 0261 (skill_member table, enums, indexes, skill.workspace_shared)
plus the hand-written write-user backfill, renumbered after staging's 0261.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(db): regenerate skill_member migration as 0263 after staging merge

Staging claimed 0262 (strong_storm); same DDL plus the hand-written
write-user backfill, renumbered on the merged snapshot chain.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(db): drop skill_member migration 0263 for regeneration on latest staging

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(db): regenerate skill_member migration as 0264 after staging merge

Staging claimed 0263 (workflow_fork_sync_excluded); same DDL plus the
hand-written write-user backfill, renumbered on the merged snapshot chain.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* make editing skills full page

* fix disclaimer

* edit access msg

* fix lint

* chore(db): drop skill_member migration 0264 for regeneration on latest staging

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(db): regenerate skill_member migration as 0265 after staging merge

Staging claimed 0264 (fat_ikaris); same DDL plus the hand-written
write-user backfill, renumbered on the merged snapshot chain.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix tests

* simplify system

* fix

* fix lint

* add mship skills docs

* chore(db): drop skill_member migration 0265 for regeneration on latest staging

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(db): regenerate skill_member migration as 0266 after staging merge

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(deps): override zod to 4.3.6 to dedupe nested copies breaking type-check

better-auth 1.6.23 and fumadocs-mdx resolve ^4.3.6 to a nested zod 4.4.3,
which makes @sim/auth's inferred betterAuth types non-portable (TS2883) and
split docs onto a second zod instance. Both ranges accept the repo-wide
pinned 4.3.6, so a single hoisted copy satisfies everything.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix lint

* feat(skills,tools): fullscreen skill create + shared custom tool editor

Moves the rich-markdown and custom-tool editing surfaces out of modals and
onto full-page surfaces, and collapses the duplicated chrome behind shared
components.

Skills
- Add /skills/new, a full-page create surface mirroring the skill detail page
  (CredentialDetailLayout + DetailSection + unsaved-changes guard). "Add to
  Sim" navigates there instead of opening a modal.
- Import moves to a header action (SkillImportButton) backed by a shared
  readSkillFile helper; the GitHub-URL import and its /api/skills/import route
  are removed.
- Skill name validation is now one shared validateSkillName, replacing three
  copies of the kebab-case rule and its messages.
- The skill editor roster renders through the shared MemberRow instead of
  re-deriving its identity block, with a locked role control and a lock-reason
  tooltip explaining inherited workspace-admin access.

Custom tools
- Extract the canvas modal's schema/code editors into a shared
  custom-tool-editor module (fields, wand generation, schema helpers), cutting
  custom-tool-modal.tsx by ~900 lines.
- Settings > Custom tools gains a full-page detail sub-view (SettingsPanel +
  SettingsSection + saveDiscardActions), deep-linkable via ?custom-tool-id.
  Rows are clickable; delete now lives only in the detail view.
- Replace legacy Button/Input/Badge/Label with the chip family, move chip-field
  chrome into CodeEditor behind an error prop, and delete its dead wand button.

Rich markdown field
- maxHeight is now opt-in: omit it on a page and the editor grows with its
  content so the page owns the only scrollbar. Modals pass explicit caps.
- The field variant drops to font-weight 400 to match adjacent chip fields.

* fix(skills): address review round on create navigation, 409 copy, and editor audit

- Skill create navigated using the first element of the upsert response, but
  that endpoint returns the caller's whole skill list (built-ins prepended) —
  match the new skill by its workspace-unique name instead.
- The suggested-skill 409 toast claimed the skill existed but was not shared
  and told the user to ask a skill admin. Every workspace member can already
  see and use every skill, so a 409 only means the name is taken.
- Adding an editor emitted the skill_shared event and SKILL_MEMBER_ADDED audit
  even when onConflictDoNothing skipped the insert on a concurrent add. Gate
  both on the insert actually returning a row.

* chore: format skills-resolver test import

* fix(skills,tools): audit fixes — autocomplete boundary, resize clipping, error routing

Two real regressions introduced while simplifying the extracted editor:

- The schema-param autocomplete's trigger was rewritten to match a trailing
  identifier, but the completion still split on separators. The two disagreed,
  so typing `data.ci` opened the menu and selecting replaced `data.ci` whole —
  eating the member-access prefix. Both now share one SCHEMA_PARAM_WORD regex.
- The uncapped markdown field measured its height only on value change while
  always setting overflow-hidden, so any width change that re-wrapped lines
  clipped the tail with no scrollbar to reach it. Now re-measures via
  ResizeObserver.

Also from the audit:
- Generation writes bypass the code field's change handler, so an open
  autocomplete stayed over a disabled streaming editor; close it on busy.
- Delete failures rendered in the Schema section's error slot on both custom
  tool surfaces; route them to a toast instead.
- Skill create navigated away while still dirty, stranding the unsaved-changes
  guard's history sentinel so Back landed on an empty create form.
- The Description field on skill create never received its error border.
- Drop a double-applied opacity-50 (the editor already dims when disabled), a
  dead try/catch around a non-throwing call that also shadowed the error prop,
  and a stale reference to a /tools page that does not exist.
- Docs still described the removed GitHub-URL import and the old Add Skill
  dialog; rewrite for the create page and file/paste import.

* feat(tools): read-only tool detail, create lands on the new tool, drop dead wand prompt API

- Viewers without edit rights could not open a custom tool at all, while the
  equivalent skill and custom-block surfaces both offer a read-only view. The
  detail page now takes `readOnly`: editors inert, no Save/Discard/Delete, no
  Generate. Creating still requires edit rights.
- Creating a tool bounced back to the list while creating a skill lands on the
  new skill. Tools now do the same. The upsert returns the workspace's whole
  tool list (newest first) rather than just the new row, so the id is matched
  by title instead of by index — the same trap that produced the skill-create
  navigation bug.
- Remove `openPrompt`/`closePrompt` from useWand. `closePrompt`'s last callers
  went away with the custom-tool-modal extraction and `openPrompt` had none
  before it; nothing reads `isPromptVisible` any more either.

* fix(tools): read-only editors, design-system wrench, skills-matching tool identity

- readOnly never reached the editors: the prop gated actions and Generate but
  the schema and code fields were still typable for viewers without edit rights.
  Wire disabled through both fields into CodeEditor.
- The row icon used lucide's Wrench (strokeWidth 2) where @sim/emcn/icons ships
  one drawn for this system (1.55, tuned viewBox), and it inherited body text
  colour instead of --text-icon. Swap it.
- Give the tool detail page the same identity heading as skill detail: tile,
  name, and description at the top left, instead of only a header title.
- Extract ResourceTile so the skills and tools tiles share one definition
  (SkillTile now composes it), and add an opt-in `iconFilled` to
  SettingsResourceRow so the tools list tile matches the skills gallery. Both
  default to today's behaviour for every existing consumer.

* fix(mentions): use the product's own glyph for every @ mention kind

The `@` menu and the inserted chip mapped kinds to arbitrary lucide icons —
`Sparkles` for a skill, a generic `File` for every file — while the rest of the
product has a settled glyph per resource. Mirror CHAT_CONTEXT_KIND_REGISTRY,
which Chat's `@` menu already renders from:

- skill now uses AgentSkillsIcon, the same glyph SkillTile shows everywhere
- workflow / folder / table / knowledge use the @sim/emcn/icons set the sidebar
  and the chat registry use
- file derives its icon from the filename extension, so a .pdf and a .csv are
  distinguishable, matching the file list and Chat's context chips
- integration keeps the block's brand icon from the registry

Also drop the generic placeholder. `kind` is untrusted — the node schema
defaults it to `''` and a hand-written `sim:` link can carry anything — but an
unrecognized kind now yields no icon instead of a meaningless box, which is what
the chat registry does. The menu already guarded a missing icon; the chip now
does too, so this cannot crash on a malformed link.

* chore(db): drop skill_member migration 0266 for regeneration on latest staging

* feat(db): regenerate skill_member migration as 0267 after staging merge

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Waleed Latif <walif6@gmail.com>
…face (simstudioai#5916)

* refactor(skills): one definition of the skill fields across every surface

The Name / Description / Content trio was implemented three times — once in the
canvas modal and once each in the create and detail pages — and had already
drifted: the name hint appeared on create, was missing entirely on detail, and
sat in a different slot in the modal; the content editor was 260px on the pages
and 200px in the modal; only the modal marked the fields required.

Extract SkillFields, the DetailSection form both full-page surfaces render, and
move the shared copy into skill-copy.ts. The modal keeps ChipModalField — that
is required inside a ChipModalBody, so it cannot share the pages' JSX — but now
reads the same placeholder, hint, and max-length constants, so the wording can
no longer diverge.

The detail page's local FieldLockTooltip moves into SkillFields and is now
available to both pages, and the dynamic rich-editor import drops from three
copies to two.

No behavioral change beyond the drift being resolved: the name hint now shows on
the detail page as it already did on create.

* refactor(skills): derive modal saving state, collapse the field message line

From the cleanup passes over this branch:

- SkillModal mirrored `mutation.isPending` into a local `saving` useState, which
  .claude/rules/sim-hooks.md forbids. It also carried a latent bug: the
  `finally { setSaving(false) }` ran after `onSave()` had closed the modal, so it
  set state on an unmounting component, and a throw from `onSave()` would leave
  the modal stuck in a saving state. Derived from the two mutations instead.
- The hint/error line under a field was written three times in SkillFields, in
  three slightly different shapes. One FieldMessage helper now renders all three.
- The name hint is an authoring instruction, so it no longer shows under a field
  that cannot be edited (a built-in skill, or a viewer who is not an editor).
…y chat answers) (simstudioai#5915)

* fix(providers): stop the final regenerated stream from re-calling tools and clobbering the answer

After the silent tool loop settles, OpenAI Responses and Gemini re-issue a
streaming request purely to stream the answer as prose — but with tools still
attached and auto tool choice, a reasoning model can re-decide to call a tool
there. Streamed calls are never executed on this path, so the run ends with a
dead function call, an empty streamed answer, and the stream callback
clobbering the tool loop's settled text with '' (deployed chat rendered
{"content": ""}).

Force tool_choice 'none' / functionCallingConfig NONE on the regeneration and
keep the tool loop's settled answer whenever the stream ends without text.

* feat(openai): settled tool chips on the regenerated answer stream

The silent Responses tool loop has no live stream while tools run, so opted-in
consumers saw no tool chips at all for OpenAI. The loop now records each
executed call and prepends settled tool_call_start/end pairs (name + status
only) to the agent-events stream ahead of the regenerated answer. Runs without
a sink never see these events, so legacy output is unchanged.

* fix(providers): apply the regeneration guard fleet-wide

Audit of every silent tool loop for the same race fixed for OpenAI/Gemini
(final regenerated stream re-calls a tool that is never executed, ending with
an empty answer that clobbers the settled text):

- anthropic (both implementations): tool_choice {type:'none'} on the
  regeneration (tools must stay — history carries tool_use blocks) + keep the
  settled answer when the stream ends without text
- groq: was re-applying the ORIGINAL tool_choice, so forced-tool runs
  re-forced the tool on the regeneration — guaranteed dead call; now 'none'
  + fallback
- deepseek, mistral, cerebras, azure-openai (legacy chat path), openrouter,
  xai: 'auto' -> 'none' + fallback
- bedrock: fallback only — Bedrock's ToolChoice has no 'none' and toolConfig
  is required when history carries toolUse blocks

Already guarded (no change): meta, sakana, nvidia, vllm, litellm, baseten,
together, fireworks, kimi, zai, ollama.
Co-authored-by: Sim Pi Agent <pi@sim.ai>
…wedged saves (simstudioai#5905)

* fix(data-retention): clamp sub-day retention values so saves aren't wedged

A stored value under 12 hours rounded to '0' days on load and was re-sent as 0, which the contract rejects (min 24) — blocking every save on the page, including unrelated fields.

Clamp hours->days into the contract's range on read, and throw instead of emitting 0/NaN on write.

* improvement(docs): rewrite data retention for workspace overrides + PII redaction

- Document PII redaction (Logs / Workflow input / Block outputs stages, entity types, languages, custom regex patterns)
- Replace the stale 'no per-workspace overrides' section with the retention-policies list and override inheritance
- Correct log retention (also covers background job logs) and soft deletion (adds Chat conversations, KB documents)
- Add PII + override screenshots, refresh the main one
…udioai#5909)

* feat(sso): DNS domain verification gating org SSO registration

Add org-scoped domain ownership verification (DNS TXT challenge) as the
security precondition for configuring SSO. Closes the first-come domain-claim
vuln where any org could wire another company's domain to its own IdP.

- New sso_domain table + migration 0266; existing org SSO domains are
  grandfathered as verified so live tenants are unaffected
- Verified-domains settings UI (enterprise-gated) with add/verify/remove
- Register route now requires a verified domain for org-scoped registration;
  personal SSO and already-grandfathered domains are unaffected
- Self-host register script writes the verified sso_domain row directly, so
  script-driven registration stays backwards compatible

* fix(sso): harden domain verification against concurrency + fix CI lint

Addresses review findings on state invariants under concurrent/failed writes:
- Add unique index on (organization_id, domain) so concurrent claims can't
  create duplicate pending rows; POST re-reads and stays idempotent on conflict
- Verify flips the row only if it's still the exact pending challenge checked
  (guards deletion/token-rotation mid-DNS-lookup) and maps the partial unique
  index violation to 409 instead of an unhandled 500
- Wrap the self-host script's provider write + verified-domain upsert in a
  transaction so a failed ownership write can't leave a provider committed
- Format 0266 snapshot/journal with biome (fixes @sim/db lint:check)

* fix(sso): re-check domain verification before provider write (TOCTOU)

The register gate checked the verified sso_domain row only at handler entry,
then ran OIDC discovery before writing the provider. A verified row removed
during that window could still complete registration. Extract the check into a
closure and call it both as an entry fast-fail and authoritatively right before
registerSSOProvider, alongside the existing domain-conflict re-check.

* fix(sso): stop rotating verification token on idempotent re-add

Re-adding a pending domain rotated its verification token, which invalidated a
TXT record the admin may have already published and — under two concurrent
re-adds — could return a token the racing write had already superseded, so the
admin's DNS record would never verify. Return the existing row unchanged
instead; the pending token is always shown in the UI, so it is never lost.

* fix(sso): close register TOCTOU with compensating delete + harden edges

Audit-driven hardening:
- Close the residual register TOCTOU: registerSSOProvider is create-only (throws
  if the providerId exists), so a compensating delete after the write is provably
  safe — it can only remove the just-created row. If verification was revoked
  during the write, roll the provider back and 403.
- Verify is now idempotent under concurrency: a same-org row already flipped to
  verified by a racing request returns 200, not a confusing 409.
- Grandfather backfill + self-host script now match normalizeSSODomain's dominant
  transforms (lower + trim + strip leading wildcard) so a non-canonical legacy
  domain can't miss the runtime gate's lookup. Prod backfill result is unchanged.
- Cleanup: drop dead default export, align card radius to sibling convention.

* fix(sso): redact domain tokens from non-admins + fix script stale-update

Round-5 review findings:
- GET /domains redacted the pending TXT verification token (a management
  secret) to any org member. Now only owner/admins read it; members see the
  list and status without it. Non-Enterprise orgs get an empty list (entitlement
  flag only), never the domains/tokens.
- Self-host script decided update-vs-insert from a read taken OUTSIDE the
  transaction; a provider deleted mid-flight made the UPDATE match zero rows
  silently while the verified-domain upsert still committed (orphaned domain).
  The decision now happens inside the transaction from the UPDATE's row count.

* docs(sso): drop unshipped enforce-SSO / auto-join copy from verified domains

Verified domains currently only gate SSO configuration. Remove the
forward-looking references to enforcing SSO and auto-joining members (deferred
to a later release) from the docs, settings copy, nav description, and schema
comment so we don't promise unshipped features.

* fix(sso): guard rollback to new providers only + Enterprise-gate domain removal

Round-6 review findings:
- The compensating provider rollback now only fires when the provider did not
  exist before this request (providerExistedBefore). registerSSOProvider is
  create-only today so reaching the rollback already implies a fresh create, but
  this makes the safety local and future-proof: if Better Auth ever allowed
  updating an existing provider, a revoked-verification rollback must not delete
  that pre-existing row.
- DELETE /domains now requires an Enterprise plan like add/list/verify, so all
  domain mutations share one entitlement (the UI already hides removal from
  non-Enterprise orgs). Adds a delete-route test.

* fix(sso): roll back the SSO provider by row id, not logical keys

The compensating rollback deleted by (providerId, orgId). providerId is unique,
so if this request's row were deleted and recreated by a concurrent registration
in the narrow window before the rollback, the logical-key delete would remove
that other request's provider. Delete by the primary-key id registerSSOProvider
returns instead, so only the exact row this request created is ever removed.

* chore(sso): final-review polish — trim script read, unify copy, doc migration edge

Cosmetic cleanup from a final 4-track adversarial review (no bugs found in the
new logic):
- Self-host script: narrow the pre-transaction existence read to select({ id })
  instead of SELECT * (it only feeds a log line now).
- Unify invalid-domain copy ("for example acme.com") and the verified-elsewhere
  409 wording ("is already verified by another organization") across routes.
- p-3 shorthand on the domain row card.
- Document the migration's rare two-orgs-share-a-domain grandfather behavior
  (login unaffected; validated no such duplicates in prod).

* fix(sso): apply attribute mapping + make SSO edit work; drop dead guard

Two pre-existing SSO bugs the final review surfaced (prod has one SSO org, RVW,
script-registered with the default mapping, so neither change affects it):

- Attribute mapping was passed at the top level of the register payload, which
  Better Auth ignores — it reads oidcConfig.mapping / samlConfig.mapping. Nest it
  so custom mappings actually apply. (Default mapping is unchanged, so existing
  logins are unaffected.)
- Editing an SSO provider was broken: registerSSOProvider is create-only and
  threw on the existing providerId → generic 500. Route now detects a provider
  the caller already owns and updates it via Better Auth's updateSSOProvider, and
  surfaces Better Auth's own error status/message instead of a blanket 500.

Also drops the now-unnecessary providerExistedBefore guard (the rollback deletes
by the created row's primary-key id and register is create-only) and the earlier
final-review polish (script read, unified copy, migration edge note).

Smoke-test SSO login + edit on staging before merge (auth-path change).

* fix(sso): require null org on personal-mode provider lookups (gate bypass)

The personal branch of both provider-ownership lookups keyed on
(providerId, userId) without requiring organizationId IS NULL. Because org
providers store userId = their creator and providerId is globally unique, an org
admin could send a personal-mode request (no orgId) — which skips the membership
check and the domain-verification gate — yet still match, and then via the new
update path move, their org's provider to an unverified domain. Add
isNull(organizationId) to the personal branch of both clauses so it can only
match a genuinely personal provider, matching the route's own isOwnedByCaller.

Found by an adversarial review of the update path added in 394bda9.

* fix(sso): script updates the observed provider by id, not providerId

Inside the registration transaction the script updated WHERE providerId — the
logical key. If the observed provider was deregistered and a replacement created
with the same providerId before the transaction ran, that update would clobber
the replacement's config and ownership. Update the specific observed row by its
primary-key id instead; if it's gone we insert, which fails cleanly on the
providerId unique constraint rather than overwriting the replacement.

* fix(sso): script upserts provider via delete-then-insert (no unique constraint)

sso_provider.provider_id is a plain (non-unique) index and prod holds legitimate
duplicates, so the previous "update by id, else insert" could create a duplicate
provider when the observed row was deregistered and replaced before the
transaction — the fallback insert would succeed. Delete every row for the
providerId then insert exactly one, inside the transaction, so the providerId
ends up as exactly this config atomically. Linked accounts key on the providerId
string (not the row id), so existing logins are unaffected.

* fix(sso): guard compensating-delete row id so rollback can't silently no-op

* chore(sso): regenerate migration as 0268 after merging staging

Staging landed migrations 0266/0267, colliding with our 0266. Removed our
migration, merged staging, and regenerated cleanly with drizzle-kit as
0268_sso_domain_verification (identical sso_domain table + indexes), then
re-appended the grandfather backfill. api-validation baseline reconciled to 973
(staging 970 + our 3 domain routes). Also make the register-route test's
registerSSOProvider mock return an id so the guarded compensating delete runs.

* refactor(sso): share normalizeSSODomain via @sim/utils so script matches gate

The self-host script canonicalized SSO domains with a minimal inline transform
(lower+trim+wildcard) that diverged from the app's full normalizeSSODomain
(protocol, port, path, trailing dot, email local part) — equivalent spellings
could store a different ownership key than the runtime gate looks up. Move
normalizeSSODomain into @sim/utils/sso-domain (a pure function) so the register
route, the domain-claim route, and the script all use the identical canonicalizer.
The script now skips the verified-domain record when SSO_DOMAIN isn't a valid
registrable domain instead of storing a malformed key.
* feat(providers): add Claude Opus 5 model

* fix(providers): route Opus 5 through adaptive thinking

Opus 5 supports only adaptive thinking (no extended thinking / budget_tokens),
same as Opus 4.8/4.7. Add opus-5 to supportsAdaptiveThinking so thinking
requests use thinking.type: adaptive instead of falling through to budget_tokens
extended thinking, which Opus 4.7+ rejects with a 400.

* docs(agent): regenerate agent-stream docs for Opus 5
…ides (simstudioai#5928)

* content(library): add observability, procurement, and MCP security guides

Three AEO guides, each dated into a recent empty day in the posting
calendar (2026-07-19, 07-21, 07-22) so publication is spread across days
rather than dumped on one.

- ai-agent-observability: what observability is, why APM falls short for
  non-deterministic agents, what to instrument per lifecycle stage, and the
  signals to track
- ai-agents-in-procurement: what procurement agents do, buy-vs-build, where
  they add value, and how to start narrow
- mcp-security: tool poisoning, confused-deputy/OAuth flaws, and supply-chain
  risk, plus how to build and govern MCP servers securely

Every third-party factual claim carries an outbound citation to a verified
primary source (OpenTelemetry semconv, Fiddler, LangChain and PwC surveys,
Icertis/ProcureCon, Ironclad, GEP, the MCPoison/CurXecute CVEs on NVD,
Anthropic's MCP announcement, RFC 8707/9700, OAuth 2.1, the MCP auth spec),
and each post carries 3 internal links to related library posts. One claim
from the source copy — a "November 2025" date on the WhatsApp MCP exfil case
— could not be verified and was dropped; the case itself is cited to Docker's
writeup. Covers come from the autogeneration pipeline via the standard
/library/<slug>/cover.jpg path.

* fix(library): correct two unverifiable claims in the new guides

Accuracy pass on the three new posts turned up two claims that could not be
substantiated:

- HIPAA compliance: the procurement post claimed Sim has "SOC2 and HIPAA
  compliance," but the canonical compliance data (lib/compare/data/sim.ts)
  states SOC2 only, and explicitly that Sim offers self-hosting "beyond SOC2,
  rather than additional certifications." Removed HIPAA; kept SOC2 plus
  self-hosting for data residency. (The same claim exists in ~5 pre-existing
  library posts and should be corrected separately.)
- "more than 18,000 servers were listed on MCP Market": no source
  substantiates this figure. Replaced with "thousands of community-built
  servers," which the ecosystem supports, keeping the Anthropic and MCP Market
  links.

Every other third-party claim was verified against a primary source (both
CVEs on NVD, RFC 8707 title, RFC 9700 as January 2025, and all six survey
statistics).
…uite (simstudioai#5927)

* improvement(ship): pre-flight regenerate artifacts + parallel audit suite

Add a two-phase pre-ship step: (A) regenerate every committed artifact
(agent-stream-docs, skills, contract syncs) in parallel so generated files
never drift into a CI failure, then (B) run lint plus the full audit suite
CI's Lint and Test job enforces, in parallel, aborting ship on any failure.
Regenerate the ship command projections.

* fix(ship): scope Phase A to in-repo generators + propagate failures

Cursor review: (1) mship:generate is an umbrella over the 9 contract
generators and biome-formats the shared generated dir, so running it
alongside its constituents races/corrupts; it also reads an external
copilot-contract source that ENOENTs in a bare worktree. (2) bare wait
swallowed generator exit codes. Narrow Phase A to the always-in-repo
generators (agent-stream-docs, skills) and collect each job's exit
status in both phases; document domain generators as on-demand only.

* fix(ship): make pre-flight status checks exit non-zero on failure

Cursor/Greptile: grep && echo ❌ || echo ✅ always exits 0, so a failed
generator/audit didn't actually gate ship (regressed the sequential
bun-run checks whose non-zero exit agents rely on). Replace both Phase A
and Phase B aggregations with an if-grep that exits 1 on any failure.

* fix(ship): gate lint exit in Phase B pre-flight

Cursor: bare bun run lint before the audit grep meant a non-zero lint
was ignored — if audits then passed, the block exited 0 and ship
continued. Gate lint with || { echo …; exit 1; } so an unfixable lint
error aborts before the audits run.
…end gotcha (simstudioai#5931)

* fix(sso): surface DNS verification failures and the provider auto-append gotcha

Post-merge audit follow-ups for verified domains (simstudioai#5909):

- The host field handed admins the FQDN `_sim-challenge.acme.com`. GoDaddy,
  Namecheap, Hover and most cPanel panels append the zone to whatever is typed,
  yielding `_sim-challenge.acme.com.acme.com` — the record looks right in their
  panel but never verifies, and our 422 tells them to wait 48 hours. Add a hint
  under the field (and a docs callout) telling those admins to enter just the
  label.
- DNS failures were logged at debug, but production log level is ERROR, so a
  blocked-egress or SERVFAIL condition was invisible to us and misreported to the
  admin as "record not found yet". Log infrastructure-class failures at warn with
  the DNS error code; keep the genuinely-absent codes at debug.
- Trim the joined TXT value before comparing: several DNS panels pad the stored
  string, which otherwise fails an exact match forever.
- Lower the resolver to 2s/1 try. c-ares multiplies timeout across servers and
  retries by ~7x, so the previous 5s/2-try config could block a verify request
  for ~35s when resolvers are unreachable.
- Cover `checkDomainTxtRecord` — the function that decides whether the gate opens
  had no tests. Adds exact match, chunk-joined value, match among unrelated
  records, padded value, near-miss, another org's token, absent record,
  infrastructure failure, and empty-response cases.

* fix(sso): log DNS infrastructure failures at error so prod actually surfaces them

Production's default minimum log level is ERROR, so the warn introduced in the
previous commit was still filtered out — the fault stayed invisible exactly as
before. A resolver failure that is not 'record absent' is a genuine
infrastructure error, so ERROR is both the visible and the honest severity.

* fix(sso): state the zone-removal rule instead of a wrong subdomain hint

The hint computed the bare label as the first segment of the challenge host, so
for a subdomain like eng.acme.com it advised entering `_sim-challenge` when the
host relative to the acme.com zone is `_sim-challenge.eng` — following it would
publish the record on the wrong name and verification would never succeed, the
exact failure the hint exists to prevent. Deriving the real zone needs the Public
Suffix List (acme.co.uk defeats naive label-stripping), so state the rule instead:
enter the host with the trailing zone removed. Docs show both the apex and the
subdomain form.
…imstudioai#5935)

.npmrc set ignore-scripts=true globally, which overrode Bun's
trustedDependencies allowlist and blocked every lifecycle script —
including the repo's own root scripts. isolated-vm therefore shipped
unbuilt and each contributor had to npm rebuild it by hand.

The .npmrc line landed in the May 2025 Bun migration, seven months before
isolated-vm was introduced. Two later commits added isolated-vm to
trustedDependencies (root, then apps/sim) and both were silently inert.

- delete .npmrc; lifecycle policy now lives solely in root trustedDependencies
- add isolated-vm to root trustedDependencies
- drop the workspace-level trustedDependencies from apps/sim (Bun only reads
  the root list, so it never had any effect)
- drop the .npmrc entry from CODEOWNERS

Naming any package in trustedDependencies replaces Bun's curated
default-trust list, so only isolated-vm and sharp may run install scripts.
Docker is unaffected — all three Dockerfiles pass --ignore-scripts
explicitly and rebuild isolated-vm against the Node ABI by hand.
…mstudioai#5936)

Reconcile compliance and traction wording across three library guides
against the canonical data in lib/compare/data/sim.ts, and tighten a couple
of related phrasings for consistency.
…O v1 default (simstudioai#5939)

* improvement(helm): hygiene pass — CI gating, strict values schema, ESO v1 default, ci values

Closes the gaps from a best-practices audit of the chart (template-level
conformance was already clean: full label set, 82 unit tests, kubeconform-
valid renders):

- new Helm Chart workflow gates every chart change: helm lint, the 82
  helm-unittest cases, kubeconform validation of default + all-components
  renders (k8s 1.29 strict, CRD catalog), and a render of all 10 example
  values files
- helm/sim/ci/ values files (chart-testing convention) so the chart lints
  and templates cleanly out of the box with dummy secrets
- values.schema.json declares all 30 top-level keys (16 were invisible) and
  sets root additionalProperties: false, so top-level typos fail fast
- externalSecrets.apiVersion defaults to v1: current ESO releases removed
  the v1beta1 compatibility path in 2026, so the old default produced
  rejected manifests on new installs; NOTES/values comments updated
- wait-for-postgres init container gets requests/limits (the only container
  in the chart without them; broke ResourceQuota'd namespaces)
- drop the telemetry Prometheus scrape config for app/realtime — neither
  exposes /metrics, so it was dead config that also rendered a realtime
  target with realtime disabled
- chart 1.2.0 with README upgrade notes

* fix(helm): version-agnostic schema-error assertions in the secret-length suite

Helm v4 phrases schema rejections as 'minLength: got N, want 32' while
v3 says 'String length must be greater than or equal to 32'. The suite
grepped for 'minLength' only, so it passed on local helm v4 and failed on
CI's v3.16.4 — the enforcement itself works on both. Patterns now assert
the key name plus either wording. Caught by the new Helm Chart workflow
on its very first run.

* improvement(helm): kind install test + chart version-bump gate in CI

Benchmarked against the flagship OSS charts (ingress-nginx, argo-cd,
kube-prometheus-stack, grafana, bitnami, cert-manager): sim already exceeds
most of them on validation rigor (strict schema — 4 of 6 ship none; 82 unit
tests vs argo-cd's zero; kubeconform manifest validation none of them run),
but every top community chart repo actually installs the chart on a kind
cluster in chart CI — the one majority practice we lacked. Adds:

- install job: kind cluster, helm install with ci/default + a new
  small-footprint ci/kind-values.yaml overlay (default app requests of 4Gi
  can't schedule on a CI node), --wait, then the chart's helm test hook,
  with pod/event/log diagnostics on failure
- version-bump job (PR-only): fails when helm/sim/** changes without a
  Chart.yaml version increment — argo/kps/grafana all enforce this

* fix(helm): declare naming overrides in the strict schema; SHA-pin CI actions

- nameOverride/fullnameOverride are consumed by sim.name/sim.fullname but
  were never in values.yaml, so root additionalProperties: false rejected
  Helm's standard naming overrides — both now declared (a helper-wide sweep
  confirmed they were the only template-read keys missing), with a smoke
  regression test that installs under the strict schema and asserts the
  override lands in resource names
- CI supply-chain hardening: checkout/setup-helm/kind-action pinned to full
  commit SHAs (tag comments retained) and the kubeconform archive verified
  against its published sha256

* fix(helm): immutable unittest runner and fail-closed version gate in CI

- helm-unittest runs via the project's official docker image pinned by
  immutable sha256 digest instead of a plugin install from a mutable git tag
- the version-bump gate fetches the full base ref (a --depth=1 fetch could
  leave no merge base), computes the merge-base and diff outside the if so
  any git failure fails the job instead of falling into the skip branch

* fix(ci): run the helm-unittest container as the runner UID

The digest-pinned image runs as a non-root user that cannot write into the
runner-owned bind mount (it creates tests/__snapshot__, absent from the
checkout since empty dirs aren't tracked). Standard bind-mount pattern:
--user "$(id -u):$(id -g)" with HOME=/tmp for helm's cache.

* fix(helm): appVersion points at a real GHCR tag (v0.7.44)

The kind install job caught this on its first full run: the default image
tag (Chart.AppVersion 0.6.73) returns 404 on GHCR for all three images —
the registry's tags are v-prefixed — so an unpinned default install could
never pull. Updated appVersion to v0.7.44 (verified 200 for simstudio,
realtime, and migrations manifests), stale values comment refreshed, and
an upgrade note added. The CI kind values deliberately stay tag-free so
the job keeps exercising the true default path.
simstudioai#5922)

* improvement(providers): validation pass, and stream tool loop improvements

* remove deploy options correctly

* fix

* feat(providers): prompt caching capability and usage-based cache pricing

Replace the arbitrary cached-rate heuristic with a single cache-aware pricing
function, and add prompt caching as an opt-in capability for Anthropic.

Pricing: priceModelUsage in cost-policy.ts is now the only place cache
arithmetic happens. Provider adapters normalize their wire shape into
ModelUsage (input always excludes cache buckets); the pricing function never
branches on provider. This removes five divergent behaviors, including the
!!request.context heuristic that gave Router and Evaluator an unearned 10x
input discount, and the overwrite that silently billed Anthropic cache reads
and writes at zero. Also parses OpenAI cache_write_tokens, previously ignored.

Caching: Anthropic gets a capability-gated advanced switch that places
cache_control on the last tool and last system block; system is now always a
TextBlockParam array. OpenAI gets a stable per-block prompt_cache_key with no
UI, since its caching is automatic.

* fix(providers): route OpenAI and Gemini block cost through cache-aware pricing

Cache-aware pricing only reached trace segments. The billable block cost still
called calculateCost on the cache-inclusive prompt total, so OpenAI cache hits
and Gemini implicit-cache hits were charged at the full input rate and GPT-5.6+
cache writes went unbilled.

Both providers now accumulate cache buckets and price through priceModelUsage,
matching the Anthropic token convention where input excludes cache reads and
writes. Cached counts are clamped to the prompt total so an over-reporting
payload cannot bill more input than the request contained.

* fix(streaming): redact tool payloads on selected outputs in public chat

Redaction only ran on the empty-selection branch, but a deployment almost
always selects outputs, so it was dead in the case it exists for. Selecting
toolCalls streamed the raw arguments and results to a public chat client in a
chunk frame, and providerTiming carried thinking content the same way.

Both paths now extract from the sanitized block output rather than the raw log:
the streamed selected output, which is the reachable vector, and the final
envelope. Sanitizing the source rather than per selected path means a newly
selectable field cannot reopen the hole.

* refactor(providers): drop unreachable billing fallbacks

Every provider pricing helper took a policy parameter no caller passed. Worse
than dead: passing one would have double-applied the margin the central layer
already applies. Removed, so providers can only price at list.

Also removed guards that cannot fire. The central fallback normalized cache
buckets no provider can reach it with (all three that report cache usage price
themselves) and did so at a 1x write multiplier no vendor charges.
priceModelUsage re-validated token counts the adapter had already clamped, and
applyModelCostPolicy defaulted a required total field.

Validation now happens once, in the adapter that parses the vendor payload and
is the only layer that knows cache buckets are a subset of the prompt total.
…nner sizing (simstudioai#5945)

* fix(ci): key node_modules sticky disk on the lockfile hash

* improvement(ci): per-image Blacksmith runner sizing + cold-build memory preflight

* fix(deps): exclude @next/swc binaries from the release-age gate

* fix(deps): pin @next/swc binaries so frozen installs get a compiler

* docs(ci): explain ARM runner sizing rationale
…inputs/outputs (simstudioai#5942)

* improvement(whatsapp): validate + improve integration skill for file inputs/outputs

* fix lint

* add whatsapp subblock migration
simstudioai#5671 rewrote the shimmer-sweep keyframes to run 100% -> -100%, which is
already left-to-right, but left the `reverse` from simstudioai#5650 on the animation
shorthand. The two flips cancel into a right-to-left sweep on every
ShimmerText consumer: subagent labels, tool-call rows, and the new
agent-stream thinking chrome.

Drop `reverse` so direction lives in exactly one place — the keyframes.
…imstudioai#5917)

* fix(realtime): evict revoked collaborators from live workflow rooms via periodic read-access re-validation

* fix

* fix(realtime): close join/eviction race and make sweep cleanups independent

Re-authorize immediately before socket.join so an in-flight join cannot
reverse a sweep eviction, and run the sweep's best-effort cleanups
independently so a room-state failure cannot skip the presence broadcast.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): single-flight role resolution and retry failed eviction cleanup

Coalesce concurrent role resolutions per (user, workflow) so a slow stale
read can never overwrite a recorded revocation, and defer failed eviction
room-state cleanups into a per-sweep retry queue so collaborators are not
left with a stale presence entry.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): scope room removal to the target workflow and detect swallowed cleanup failures

Honor the workflowIdHint as the target room in both room managers so
removing a stale room cannot clobber the mapping of a room the socket has
since moved to, and confirm eviction cleanup via the returned workflowId
plus an unswallowed mapping read so Redis failures actually defer into the
retry queue.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): treat unconfirmed removals as failures and refresh role cache on fresh verify

Treat any null removal result as a failed cleanup (the sweep always passes
the target room, so null only means failure — including with expired
mapping keys), move the same-room rejoin guard to a synchronous check
immediately before the removal, and record verifyWorkflowAccess's fresh
decision into the role cache so a re-granted user is not blocked by a
stale cached revocation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): prefer a mid-flight recorded role decision over the in-flight query result

If a fresh authoritative read (join-time verify) records a decision while
a single-flighted resolution's query is in flight, keep the recorded
decision instead of overwriting it with the potentially stale result.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): isolate the security scan from the Redis cleanup lane

Run the revocation scan (local sockets + DB only) and the best-effort
room-state cleanup as independently-guarded lanes so a hanging Redis
command can stall only presence cleanup, never revocation enforcement.
Evictions now enqueue cleanup instead of awaiting it, and the scan no
longer reads presence for a fallback role.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): bound authorization waits in the revocation scan

Race each socket's authorization check against a per-socket timeout and
cap the whole pass with a budget below the sweep interval, so a hanging
DB query skips that socket for the pass (never evicting on uncertainty)
instead of wedging the scan lane and starving subsequent ticks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): round-robin the revocation scan so hung checks cannot starve later sockets

Resume each scan pass after the last target the previous pass processed,
so a fixed prefix of hanging authorization checks can never repeatedly
consume the pass budget and leave sockets behind it unexamined.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
… agent thinking/tool-calls/prompt-caching, opus 5, daytona failover
… state to nuqs, design-system cleanup (simstudioai#5950)

* refactor(settings): fold verified domains into SSO and use shared primitives

Verified domains only gates SSO, so managing it on a separate page meant
discovering the requirement after filling out the whole IdP form and then
navigating away mid-setup. Move it into the SSO page as a section above the
provider config and drop the standalone page, its nav entry, and both route
branches.

Align the surfaces with the shared settings primitives rather than bespoke
chrome, matching whitelabeling/custom-blocks/access-control:
- SSO's local FormField (muted labels) is replaced by the shared SettingRow, so
  its fields read like every other settings page. SettingRow gains optional
  `optional` and `error` props to absorb what FormField did — additive, so
  existing consumers are untouched.
- The domains section is built from SettingsSection, SettingRow,
  SettingsResourceRow, and SettingsEmptyState instead of hand-rolled cards.

Also drop the redundant Upload/Change buttons in whitelabeling: the logo and
wordmark thumbnails were already clickable, so the button was a second control
for the same action. Remove still appears once an image is set.

* fix(settings): move group-detail view state to nuqs and clean up design-system drift

Access control's group detail kept its tab, three search boxes, and three status
filters in useState, so a `?group-id=` link always landed on General and a filter
was lost on reload — the parent already puts the group id in the URL. The three
tabs never render together, so search and status share one param each rather than
carrying three mutually-exclusive keys, and switching tabs resets both. Closing
the detail clears all three alongside group-id in one batched write, so nothing
lingers on the list URL.

Design-system fixes from a cleanup pass over the surfaces this branch touched:
- Restore accessible names lost when the whitelabeling Upload buttons were
  removed. The thumbnail is now the only click target, and it contained just an
  icon, so it announced as an unlabeled button; the icon-only Remove had the same
  problem. Both now carry aria-labels reflecting their state.
- Use Chip, not the legacy Button, for the domain actions — Button is ~26px
  against the 30px ChipInput beside it, so "Add domain" sat visibly short.
- SettingRow now uses the emcn Info component; a bare svg as Tooltip.Trigger was
  neither focusable nor nameable. It also stops re-specifying Label's own default
  styling.
- Use the new SettingRow error prop for the group name instead of a hand-rolled
  error paragraph, which is what the prop was added for.
- Hoist the block-category lookup out of a sort comparator, size-* over h/w, name
  the staleTime constants the rules require, and import RowActionsMenu from its
  barrel.

* fix(settings): alias the old /settings/domains path to SSO

Folding verified domains into the SSO page dropped /settings/domains, so
bookmarks and shared links 404'd instead of landing where domains now live. Both
alias maps already exist for exactly this (organization/'members',
subscription/'billing'); add domains -> sso to each.

* chore(settings): adopt ChipCopyInput, named staleTime constants, and a11y labels

* fix(settings): reset group detail params on open and drop issuer mono styling
waleedlatif1 and others added 11 commits August 3, 2026 18:30
…hain (simstudioai#6235)

Bumps next, @next/env and the @next/swc-* optional deps to 16.3.0 across the
root overrides, apps/sim, apps/docs and packages/emcn.

Two config notes worth keeping:

- `experimental.turbopackFileSystemCacheForBuild`'s default flipped false ->
  true for stable in 16.3.0, so our explicit `false` is now load-bearing rather
  than defensive. Without it this bump would have silently re-enabled a build
  cache measured 3.2x slower on this codebase (simstudioai#6078). Comment updated to say so.
- `experimental.useTypeScriptCli: true` is now pinned. TypeScript 7 ships no
  JavaScript compiler API until 7.1, so Next's default checker cannot run and
  needs the project-local `tsc` CLI instead. 16.2.12 was silently skipping
  build-time type checking entirely because it detected @typescript/native-preview
  and short-circuited the stage ("Finished TypeScript in 138ms"); pinning the flag
  keeps that from drifting back.

TypeScript toolchain cleanup that the upgrade makes possible:

- Drop @typescript/native-preview from apps/sim and apps/docs. It was the
  pre-release channel for TS 7 and is superseded by typescript@7 (nightlies now
  ship as typescript@next), and its presence is what suppressed build type checks.
- Align packages/browser-protocol and packages/terminal-protocol from
  typescript ^5.7.3 to ^7.0.2 so every workspace is on one compiler version.

apps/sim keeps @typescript/typescript6 as a production dependency: the function
sandbox at app/api/function/execute dynamically imports the TS 6 compiler API to
transpile user code, and TS 7 has no API to replace it yet.

Build and dev were benchmarked 3x per version on a byte-identical tree; the
upgrade is performance-neutral (build median 100s -> 99s, dev:full warm 10s -> 9s).
…ath (simstudioai#6232)

* fix(providers): route 6-10MB attachments to the provider large-file path

The inline base64 cap was 10 MB of raw bytes, but the execution payload store
refuses a single value above 8 MiB and base64 inflates by 4/3. Every raw file
over 6 MiB therefore produced a base64 string the store rejected — and the
rejection came from the base64 *cache* write, which threw and failed the run
with "Execution memory limit exceeded" even though the bytes had already been
read successfully. Because shouldUseLargeFilePath only fires above the inline
cap, 6-10 MB attachments had no path at all on any provider: they never reached
the OpenAI or Gemini Files API upload they were supposed to take.

Derive the cap from the payload-store ceiling instead of hardcoding it, and
degrade a refused cache write to "not cached" rather than failing a request
whose bytes are in hand. Every other size guard in the chain compares raw bytes
against maxBytes; only the Redis write sees the encoded size, which is why this
went unnoticed — and why it failed only where Redis is configured.

Also correct the provider ceilings against the vendors' current documentation:

- openai: 50 MiB -> 50,000,000. The gate is `size > maxBytes`, so 50 MiB
  admitted 52,428,800 bytes; the docs say each file must be *under* 50 MB.
- bedrock: had no entry and inherited the inline cap, which is above what
  Converse accepts (3.75 MB per image, 4.5 MB per document).
- groq: 20 MiB -> 20,000,000, and modelled as the request cap the docs
  actually describe rather than a per-file MiB ceiling.
- fireworks: had no entry; its 10 MB budget is on the base64 total, so the
  raw-byte equivalent is 7.5 MB.

Add perRequestMaxBytes for the combined ceilings, enforced before any upload
spend, and cover the OpenAI upload path end to end — it had no test at all.

* chore(deps): upgrade @google/genai to 2.13.0 and @anthropic-ai/sdk to 0.115.0

@google/genai 2.x reworks the Interactions API, which the Gemini deep-research
provider is built on. Migrate it:

- `Interaction.outputs` (a flat content array) is now `steps`, a discriminated
  timeline; the report text lives in the `model_output` steps' text content,
  alongside thought and tool steps we skip.
- `Usage.total_reasoning_tokens` is now `total_thought_tokens`. The old code
  already fell back to that name through a cast, so this just makes the field
  the SDK actually returns the typed one.
- SSE events renamed: `content.delta` -> `step.delta`, `interaction.start` ->
  `interaction.created`, `interaction.complete` -> `interaction.completed`.
  The new event types are discriminated, so the payload casts are gone.

Both `interactions.create` calls also stop annotating their params with
`Interactions.CreateAgentInteractionParams{,Non}Streaming`. In 2.13.0 those
namespace aliases resolve to `CreateAgentInteraction`, whose `stream` is a
plain `boolean` rather than a literal — annotating with them erases the
discriminant and the call resolves to the union-returning overload, so the
result is typed as `Interaction | Stream` at every use. An inline
`stream: true as const` keeps the correct overload.

Neither upgrade required a `minimum-release-age` waiver: 2.15.0 and 0.115.0
were checked and 2.13.0 is the newest genai release clearing the 7-day window.

* chore(deps): upgrade openai to 7.0.0

v5 is the only major with real breaking changes for us; v6 widened a Responses
output type and v7 only raised the Node floor to 22, which apps/sim already
requires. Three things needed fixing:

`ChatCompletionMessageToolCall` became a union of function and custom tool
calls, and the custom variant has no `function` field — 43 unguarded `.function`
accesses across the OpenAI-compatible providers. Narrow once at each
`message.tool_calls` read through a shared `isFunctionToolCall` guard rather
than casting at every use.

That guard deliberately tests for the `function` payload instead of
`type === 'function'`. Many OpenAI-compatible vendors omit `type` on tool calls
entirely — our own fixtures do — so discriminating on it type-checks perfectly
and then silently drops every tool call those providers return.

`ChatCompletionCreateParams.verbosity` narrowed from `string` to a literal
union, and the Responses API's output and input item unions now diverge on
members Sim never emits (computer-use call outputs, whose `status` admits
`failed`, and the `AdditionalTools` escape hatch). Echoing output back as input
is what a tool loop is supposed to do, so that conversion is asserted once in
convertResponseOutputToInputItems and the streaming loop now routes through it
instead of pushing raw output items.

The hand-rolled multipart upload in file-attachments.server.ts can now be
replaced with the SDK's typed `expires_after` — left for a follow-up so this
commit stays a pure upgrade.

* fix(providers): correct defects found auditing the attachment and SDK changes

The mechanical rewrite that added `isFunctionToolCall` to every `tool_calls`
read also rewrote three truthiness guards, where the filtered array was
computed, discarded, and the unfiltered value used in the body. Filter once and
use that value. The helper also landed between `trackForcedToolUsage`'s TSDoc
block and its declaration, leaving that block documenting the wrong function.

Raise the Bedrock ceiling from 3.75 MB to 4.5 MB. Converse caps an image at
3.75 MB and a document at 4.5 MB, and a single `maxBytes` cannot express both.
Taking the lower bound looked conservative but regressed 3.75-4.5 MB documents,
which Converse accepts and which work today. At the document bound every size
that works now still works, and only genuinely-too-large files are rejected
early; oversized images in that band keep surfacing as a Bedrock API error,
exactly as they do without the entry.

Both limits re-verified verbatim against the primary docs: Converse's Message
reference ("Each image's size ... no more than 3.75 MB", "Each document's size
must be no more than 4.5 MB") and Fireworks' vision guide ("Total base64-encoded
images must be less than 10MB").

* fix(providers): keep every attachment that works today working

Auditing the routing change for backwards compatibility turned up two bands it
silently broke.

Lowering the single inline cap made the upload path mandatory above ~6 MB, but
every large-file path reads its bytes back out of cloud object storage. A
deployment without it — local dev, any disk-backed self-host — inlines those
files as base64 today and would have started failing outright with "requires
cloud file storage". Split the one number in two: the inline ceiling stays at
10 MiB, and a separate threshold marks where an upload becomes *preferable*
because the base64 copy no longer fits the payload store. Where no upload path
is reachable, base64 hydration now runs to the inline ceiling as before, and a
missing cloud-storage backend leaves the file for the inline path instead of
throwing.

The two strategies also cross over at different sizes now. `files-api` carries
every type the provider already accepts, so it takes over at the lower
threshold. `remote-url` only fetches images and PDFs, so switching early would
have started rejecting 6-10 MB text documents that inline fine today; it takes
over only once inlining is genuinely impossible.

Revert the Groq ceilings. Its published "20MB" governs a request carrying an
image URL, and on this path the body holds only the URL, so it cannot bind on
the files maxBytes guards. Groq documents no limit on the image it fetches, so
tightening the per-file cap to 20,000,000 and summing raw bytes against the
request cap would both reject uploads that work today on no documented basis.

* fix(providers): drop the provider ceiling changes and close the audit findings

A six-agent line-by-line audit against the vendors' live docs found the
`models.ts` ceiling work was not the strict improvement it was written as, so
all of it is reverted:

- bedrock's 4.5 MB cap broke video. Converse takes image, document AND video
  blocks, and video is allowed 25 MB base64 — a single `maxBytes` cannot
  express three content classes, and every 4.5-10 MiB `.mp4` that works today
  would have started failing.
- openai's combined 50 MB cap is the FILE-input limit. Image inputs are
  governed separately at 512 MB / 1500 images, so summing every attachment
  rejected eight 8 MB PNGs that OpenAI documents as legal.
- fireworks' per-file ceiling was unreachable behind the request budget, while
  the upload picker went on advertising it — a size the UI accepts and
  execution always rejects.
- The whole `perRequestMaxBytes` feature goes with them: it summed raw bytes
  against caps that are variously on encoded bytes, on one content class, or on
  a body that carries only URLs, and it double-counted a file referenced from
  several messages even though the uploader dedupes by key.

Only openai's per-file `maxBytes` stays corrected, to decimal 50,000,000 — the
one number a vendor states unambiguously and writes no MiB against.

Also fixed, all found by the same audit:

The hydration cap stopped short of where `remote-url` actually switches over,
so 6-10 MiB attachments on anthropic/openrouter/xai/groq/together/baseten/vllm
had neither base64 nor a handle and failed outright — the very band this branch
exists to fix. Both decisions now come from one function so they cannot drift.

Eight more sites where the mechanical rewrite computed a filtered array and
then read the unfiltered one (deepseek, sakana, nvidia, kimi), leaving those
providers without the narrowing they appear to have.

`isFunctionToolCall` threw on a null or primitive `tool_calls` entry, because
`in` requires an object — reachable exactly on the self-hosted gateways this
filter was added for. It is now total, and all 32 test mocks match it rather
than being quietly more permissive.

`checkForForcedToolUsage` in utils/litellm/mistral evaluated the response before
the `tool_choice` test, turning a tolerated malformed body into a TypeError on a
path that never used to touch it.

Gemini: `satisfies` restores the excess-property checking the dropped
annotations removed, the poll loop recognises the terminal statuses v2 added
instead of spinning for an hour and reporting a timeout, and the streaming doc
block no longer names five events that were renamed six lines below it.

* fix(providers): report attachment limits in the unit vendors publish

The size ceilings are decimal MB — that is how OpenAI, AWS and Fireworks all
write them — but the error messages divided by 1024², so OpenAI's 50 MB cap
was reported to the user as "48MB". Someone shrinking a 49 MB file to get
under it was chasing a limit that does not exist.

One formatter, used by all three messages, so the file size and the ceiling in
the same sentence are always in the same unit.

* fix(providers): derive the limit unit from the ceiling it belongs to

The previous commit fixed OpenAI's "48MB" by dividing every ceiling by 10⁶ —
which broke the other seven. Only OpenAI's constant is decimal; anthropic,
google, together and openrouter are 50 MiB, baseten and vllm 25 MiB, groq and
xai 20 MiB. Rendering those as decimal MB overstated each by ~5%, so a 21 MB
file on groq was rejected with "(21MB) exceeds the 21MB limit" — a sentence
that contradicts itself and sends the user to shrink a file to a size that is
still over. Same class of bug as the one being fixed, sign flipped.

Both figures now render through one unit taken from the ceiling, so the number
a user is told is the number the vendor publishes and the two sizes in a
sentence are always comparable. Tested against every ceiling in the registry
rather than only the values that happened to round cleanly.

Two more from the same audit:

A file with a missing or zero declared size was stranded on a files-api
provider: hydration bailed on the real byte length while `shouldUseLargeFilePath`
saw `0 > threshold` as false, so it got neither base64 nor a handle and failed
as "may no longer be accessible" — a size failure wearing an access failure's
message. Uploads read the real bytes and enforce the ceiling themselves, so an
unknown size now routes to one.

The oversized-attachment error blamed the provider for a deployment problem:
a files-api provider on a host without cloud storage reported that the provider
"has no large-file upload path", which is not true of the provider.

`isFunctionToolCall` only proves `function` is present, never that it is well
formed, so the trace enricher is defensive again about a hollow payload without
giving up the compile-time gate. The 32 test mocks now match production exactly.

* fix(providers): stop an over-limit size rendering as the limit itself

Deriving the unit from the ceiling fixed the 5% error but left the precision
fixed at two decimals, so a file one byte over a 20 MiB cap still printed
"(20.00MB) exceeds the 20MB agent attachment limit" — the same self-contradicting
sentence, now in a ~5 KB band above every ceiling in the registry. The size
rounds up and the ceiling rounds down, so the two can no longer collide.

The test that was supposed to guard this asserted a file 0.03MB over and an
OpenAI file that was under the limit — neither anywhere near the band — so it
passed while the bug was live. It now walks `limit + 1` for every ceiling, and
goes red against the old rounding.

The reason clause added last commit also claimed a deployment had no cloud file
storage whenever the strategy was not inline. A generated document on a
remote-url provider reaches that same error with storage fully configured,
because a signed URL points at the generation source rather than the rendered
artifact — so it was told something false about its own deployment. That case
now names itself.

* fix(providers): order the attachment failure reason by how general the cause is

The generated-document arm was checked first, so it won over both other causes
and told users two things that were not true.

On an inline-strategy provider — bedrock, mistral, ollama, fireworks, litellm,
vertex, kimi — there is no upload path for any file, generated or not, but the
message blamed the document format and implied a plain PDF would go through.

On openai or google with cloud storage unconfigured it was simply false: a
generated document does take the Files API path there, and that exact file
uploads fine once storage exists. The one actionable fix was hidden from the
operator.

A provider with no upload path cannot be helped by changing the file, and a
deployment with no object storage cannot reach any upload path whatever the
file is, so both now outrank the format-specific case — which is left saying
only what is true of it: a signed URL points at the generation source rather
than the rendered file.

The formatter is unchanged. It was brute-forced over every real ceiling and
three million random pairs with no collision or inversion, but the test's six
ceilings all divide to exact integers, so floor, round and ceil are
indistinguishable on them and the limit-side rounding was unpinned. A ceiling
with a fractional remainder now covers it.
…rbosity, and thinking level (simstudioai#6233)

* improvement(agent): allow variable references in reasoning effort and verbosity

Reasoning Effort and Verbosity were select-only dropdowns, so a workflow could
not sweep them from a variable or an upstream block the way it already can with
the model. Both become editable comboboxes, matching the model field directly
above them and the managed-agent selectors.

- switch both subblocks to `combobox`, keeping their fetched per-model option
  lists intact
- keep them visible when `model` itself holds a reference, since the concrete
  model id is only known at execution time and cannot be matched against the
  static capability list
- normalize the resolved level in the provider chokepoint so a reference that
  resolves to `"High"` or to nothing behaves sanely instead of hitting a
  provider 400

* improvement(agent): allow variable references in thinking level

Extends the same treatment to Thinking Level so all three model-tuning fields
behave consistently, and logs a level a model does not declare.

- switch `thinkingLevel` to `combobox` with the reference-aware condition
- normalize it alongside the other two; an empty resolve now takes the
  deliberate "send nothing" path rather than the incoherent half-state it hit
  before, and stays distinct from an explicit `none`
- warn when a level is not one the model declares, still forwarding it: Sim's
  per-model lists drive the pickers and can lag a provider, and a sweep needs
  the provider's own error rather than a silent fallback to the default

* improvement(agent): report a model level the sanitizer discards

A model bound to a variable or block reference only resolves at execution time,
so a run whose reference landed on a model outside Sim's catalogue had its
requested level cleared with no signal and quietly fell back to that model's
default. Dropping stays the safe default — a provider with no such parameter
rejects the whole request — but it is now reported.

- log the field, model, and value whenever an unsupported-field level is cleared
- cover both diagnostics, including that they stay quiet for a declared level
  and for the `auto` / `none` sentinels

* fix(providers): redact resolved level content from sanitizer diagnostics

The model-level fields accept environment and block references, so an
unrecognized level is not necessarily a mistyped level — it is whatever the
reference resolved to, which can be secret content. The diagnostics added for
dropped and undeclared levels echoed it straight into server logs.

- log a level only when the catalogue declares it somewhere, or it is an
  `auto` / `none` sentinel; anything else is reported by length alone
- stop discarding levels for a model the catalogue has never seen. Absent is
  unknown, not known-incapable, and a reference is exactly how a newly released
  model arrives before Sim catalogues it — the provider decides instead. Models
  the catalogue knows, and every dynamic-provider id, keep the protective drop

* fix(providers): redact the level in Anthropic's unsupported-thinking warning

Forwarding an undeclared level is deliberate, but it means the Anthropic adapter
receives it and interpolates it straight into its "not supported, ignoring"
warning. Since the field is reference-bound, that value can be whatever a
mistyped `{{ENV_VAR}}` or block reference resolved to — so the redaction added
for the sanitizer's own diagnostics was leaking one layer downstream.

- promote the level renderer to `providers/utils` as `describeModelLevel`, the
  single gate every site echoing a caller-supplied level goes through
- use it in Anthropic's warning and in both sanitizer diagnostics

* refactor(providers): drop the sanitizer's level diagnostics

The two warnings logged server-side, where the workflow author who set the level
never sees them, and the surprising case they described — a level discarded for
a model newer than the catalogue — is now fixed at the source rather than
narrated. They also carried the redaction that leaked resolved content before it
was caught, so removing them removes that surface entirely.

Levels still normalize, and still drop for a catalogued model that does not take
the field. `describeModelLevel` stays for Anthropic's unsupported-thinking
warning, which is a pre-existing log this feature newly exposes to resolved
reference content.
* fix(docker): upgrade bun to 1.3.14 to unbreak the Next 16.3.0 server

Bun 1.3.13 cannot load Next 16.3.0's compiled server runtime. The app container
runs the Next server under Bun (`oven/bun:1.3.13-slim`, `bun apps/sim/bootstrap.js`),
so every app-page render threw and `/api/health` returned 500:

  ⨯ Error: Failed to load external module
    next/dist/compiled/next-server/app-page-turbo.runtime.prod.js:
    TypeError: Expected CommonJS module to have a function wrapper.
    If you weren't messing around with Bun's internals, this is a bug in Bun

Isolated to Bun, not Next, by loading that exact module in the real images:

  Next 16.2.12 + Bun 1.3.13 -> loads (why staging was fine before)
  Next 16.3.0  + Bun 1.3.13 -> CJS wrapper error
  Next 16.3.0  + Bun 1.3.14 -> loads

Bun 1.3.14 is the current stable and already fixes it, so this bumps every pin
rather than reverting the framework upgrade, which would only defer the same
latent Bun bug to the next attempt.

Why no gate caught it: local dev machines and this bump's own verification run
Bun 1.3.14, while the container and CI pinned 1.3.13 — and CI only *builds* the
image, it never boots one and probes `/api/health`. A container smoke test in CI
would have caught this before merge; that is worth adding separately.

* fix(docker): align the remaining bun pins with 1.3.14

Two pins were missed in the first pass because the search was scoped to
docker/, package.json and .github/workflows/:

- .devcontainer/Dockerfile still built on oven/bun:1.3.13-alpine
- PI_BUN_VERSION in apps/sim/scripts/pi-sandbox-packages.ts was still 1.3.13,
  despite being documented as mirroring the root packageManager field, so Pi
  sandbox images would have kept installing the Bun release that cannot load
  the Next 16.3.0 server runtime.

Fixed surgically rather than with a repo-wide replace: "1.3.13" also appears
inside SVG path data in apps/sim/components/icons.tsx and
apps/docs/components/icons.tsx, which a blind sed would have corrupted.
…code (simstudioai#6242)

Next 16.3.0's Turbopack optimizer models a bare `return <asyncCall>()` tail
call inside an async function as returning the promise object, then propagates
that always-truthy fact through the caller's `await`. Where the result feeds an
`if (x)` whose every branch returns, it concludes the branch is always taken and
deletes everything after it from the emitted bundle.

Two sites shipped to production that way:

- `POST /api/credentials` lost its entire create path — the transaction, the
  org locks, the insert, the audit, the 201. A first-time create fell into the
  existing-credential branch and threw on `existingCredential.id`, so every new
  credential 500'd.
- `upsertAsyncToolCall` collapsed to `async () => await getAsyncToolCall(id)`.
  The insert is simply gone; it returns null for every new async copilot tool
  call. Silent — no error, no failed request.

A differential scan of 71,266 source string literals across `.next/server` and
`.next/static`, comparing images built from the same commit on 16.2.12 and
16.3.0, found exactly these two and nothing else. That scan cannot see dropped
branches with no distinctive string literal, which is why the version goes back
rather than the two sites being patched alone.

Both are also hardened with `return await`, verified to defeat the miscompile in
a minimal reproduction. The TypeScript toolchain cleanup from the original bump
(dropping @typescript/native-preview, `useTypeScriptCli`) is kept.
@utcarshsrivastava-collab

Copy link
Copy Markdown
Collaborator Author

Open questions — upstream sync 2026-08-04

Two blocking decisions. Everything else in this 518-commit range is resolved from
merge-policy.json + codebase evidence and recorded in
.upstream-sync/ledger/2026-08-04/run.md under ## Grill analysis (20 recorded decisions) —
please skim that section, but you only need to answer the two questions below to unblock
the merge.

Reply on this PR with /upstream-sync resume plus your answers.


Q1 — 🔴 Six upstream migrations will be silently skipped on every existing Arena database. How do we fix it?

This is the most serious finding of the run. It is not a merge conflict, it does not fail
CI, and it does not reproduce on a fresh database — it only breaks already-deployed
environments.

Why. drizzle-orm decides what to apply from a single high-water timestamp
(node_modules/drizzle-orm/pg-core/dialect.cjs:59-69):

select id, hash, created_at from drizzle.__drizzle_migrations order by created_at desc limit 1

then applies a file only when max(created_at) < migration.folderMillis.

The fork and upstream both used idx 0258–0261, and their journal when values interleave.
Arena prod has already applied the fork's four, so its high-water mark is
1784346920598 (0261_local_copilot_user_memory). Six upstream migrations sit below
that and will therefore never run:

upstream tag when
0258_gigantic_lady_mastermind 1783620533559
0259_slack_native_routing 1783722352108
0260_unknown_sinister_six 1783810442774
0261_tranquil_donald_blake 1784043925919
0262_strong_storm 1784224314431
0263_workflow_fork_sync_excluded 1784252317428

What never gets created: the webhook_path_claim table, the
workflow_deployment_operation table + 6 indexes (the deployment state machine from simstudioai#5680 /
simstudioai#5841 / simstudioai#6229), 8 webhook columns (Slack native routing simstudioai#5892), 4 workspace columns, a
paused_executions column (resume path simstudioai#6187), a workspace_files column, workflow.fork_sync_excluded
(simstudioai#5727), and 8 indexes. Result: green CI, then hard runtime failures in production on deploy,
webhook registration, execution resume and file paths.

Aggravating factor: fork migrations 0258, 0259, 0260 are not idempotent (plain
ALTER TABLE … ADD COLUMN / CREATE TABLE, no IF NOT EXISTS; only 0261 is guarded). So
simply renumbering them to 0282+ makes drizzle replay them on prod, where they will fail
on duplicate column/table and wedge the migrations service in docker-compose.p2prod.yml.

Options

(A) Re-stamp both sides — recommended.

  1. Renumber fork 0258→0282, 0259→0283, 0260→0284, 0261→0285 (file + journal idx), with
    when just above upstream's 0281 (1785776917545).
  2. Rewrite those four to be replay-safe (ADD COLUMN IF NOT EXISTS, CREATE TABLE IF NOT EXISTS,
    guarded index/FK adds) so replay on an already-migrated DB is a no-op.
  3. Re-stamp upstream 0258–0263's when into the free window
    (1784346920598, 1784680777879) — 334M ms of headroom, journal order and idx order both
    preserved — so they land above every existing environment's high-water mark.
  • ✅ Works on fresh and existing databases; no manual per-environment SQL.
  • ❌ Edits upstream-owned migration metadata, so step 3 recurs as a conflict in future syncs.

(B) One-off __drizzle_migrations surgery per environment. Keep all upstream files
untouched; on each deployed DB, manually insert/adjust rows so the six upstream files are
applied or considered applied.

  • ✅ Zero upstream drift.
  • ❌ Manual, per-environment, easy to forget on a new environment; needs its own runbook.

(C) One-shot script migration. Use the existing packages/db/script-migrations/
(runScriptMigrations) hook to force-apply the six upstream files idempotently on first boot
after this merge.

  • ✅ Automated, no upstream file edits.
  • ❌ Adds fork-owned code that duplicates upstream SQL; must be retired later.

→ Which option? (Default if you have no preference: A.)
→ Also confirm: can the harness verify migrations against a restored clone of Arena
production
? Verifying only against a fresh DB cannot detect this class of defect, so a clean
CI run is not evidence that this is fixed.


Q2 — 🔴 Upstream deleted the deployed-chat voice/TTS stack. Does Arena keep voice mode?

Upstream removed voice mode wholesale (simstudioai#6215, simstudioai#6218, simstudioai#6220, simstudioai#6224), deleting:

  • app/(interfaces)/chat/components/voice-interface/voice-interface.tsx
  • app/(interfaces)/chat/components/voice-interface/components/particles.tsx
  • app/(interfaces)/chat/components/input/voice-input.tsx
  • app/(interfaces)/chat/hooks/use-audio-streaming.ts
  • app/api/proxy/tts/stream/route.ts
  • hooks/queries/voice-settings.ts
  • lib/api/contracts/media/tts-stream.ts

The fork's fork-only app/(interfaces)/chat/[identifier]/ArenaDeployedChat.tsx (~1600
lines) depends on all of it: it imports VoiceInterface, holds isVoiceFirstMode, streams
through DEFAULT_VOICE_SETTINGS.voiceId = 'EXAVITQu4vr4xnSDxMaL' (ElevenLabs "Bella"), and
reports Prompt Type: Voice to Mixpanel. This looks like a live, customer-facing Arena feature,
which is why I am not defaulting it.

--ours is not a coherent answer on its own: the fork only touched voice-interface.tsx
cosmetically (3 lines of dark-mode tokens), and app/api/proxy/tts/stream/route.ts is
unmodified by the fork, so it will disappear with no conflict marker at all.

Options

(A) Keep voice mode. Skip the seven deletions for these paths, keep the fork's voice stack,
and do take simstudioai#6212 (fix(security): meter and throttle the deployed-chat TTS relay, v0.7.54)
— upstream hardened the relay immediately before deleting it, so that hardening is worth having.
Cost: this subsystem is now fork-owned forever and every future upstream chat refactor must be
re-applied to it by hand.

(B) Drop voice mode, following upstream. Accept all seven deletions and strip voice-first
mode out of ArenaDeployedChat.tsx (state, handlers, ElevenLabs settings, the Mixpanel
Prompt Type: Voice branch). Cost: a visible Arena feature disappears for whoever uses it.

→ (A) or (B)? If (A), please also confirm whether voice should stay on the deployed/public
chat
surface only, or also anywhere else it has been wired.


Not asking about these — recorded as decisions in run.md

Flagged for visibility; tell me if any is wrong and I will change it, but none of them blocks
the merge: models.ts union · HubSpot union · exa_research retired (upstream verified HTTP 410
RESEARCH_RETIRED) · minimal tool registry removed · apps/desktop/ taken · ci.yml fork-first
with upstream desktop + Blacksmith jobs skipped · test-build.yml upstream + the fork's
check:secrets step re-applied · scheduler stays on Trigger.dev (upstream's cron container lands
unused) · Arena branding assets fork-first · legacy folder tables migrated in fork-only files ·
fork billing markup preserved over upstream's attribution refactor · Turbopack cache flags follow
upstream's measurements.

@utcarshsrivastava-collab

utcarshsrivastava-collab commented Aug 4, 2026

Copy link
Copy Markdown
Collaborator Author

/upstream-sync resume

  1. Remove all the fork migrations as that have already been provided manually to the db. We never run migrations container - and those migrations are only kept as part of a documentation. They can be removed.
  2. Drop voice mode.

@utcarshsrivastava-collab

Copy link
Copy Markdown
Collaborator Author

/upstream-sync resume

@utcarshsrivastava-collab

Copy link
Copy Markdown
Collaborator Author

Closing to retry upstream-sync with the release-sliced merge-plan harness on feat/github-merge-agent.

@utcarshsrivastava-collab
utcarshsrivastava-collab deleted the upstream-sync/2026-08-04T11-16-45 branch August 5, 2026 10:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants