Skip to content

upstream-sync: merge simstudioai/sim main into feat/github-merge-agent (2026-08-04) - #679

Closed
utcarshsrivastava-collab wants to merge 534 commits into
feat/github-merge-agentfrom
upstream-sync/2026-08-04T09-49-31
Closed

upstream-sync: merge simstudioai/sim main into feat/github-merge-agent (2026-08-04)#679
utcarshsrivastava-collab wants to merge 534 commits into
feat/github-merge-agentfrom
upstream-sync/2026-08-04T09-49-31

Conversation

@utcarshsrivastava-collab

@utcarshsrivastava-collab utcarshsrivastava-collab commented Aug 4, 2026

Copy link
Copy Markdown
Collaborator

Upstream sync — 2026-08-04

Merges simstudioai/sim@1b9e0f25 into feat/github-merge-agent.

Sync range: 518 commit(s) since e2fecc86 (merge-base).

Ledger

Verification (advisory — does not block the sync)

bun run check · ❌ bun run lint · ❌ bun run test · ❌ bun run build

Verification is advisory — failures do not block the sync. Review and fix on the draft PR as needed.

bun run check

❌ failed


   • Packages in scope: @sim/audit, @sim/auth, @sim/browser-protocol, @sim/db, @sim/desktop, @sim/desktop-bridge, @sim/emcn, @sim/logger, @sim/pii, @sim/platform-authz, @sim/realtime, @sim/realtime-protocol, @sim/runtime-secrets, @sim/security, @sim/terminal-protocol, @sim/testing, @sim/tsconfig, @sim/utils, @sim/workflow-persistence, @sim/workflow-renderer, @sim/workflow-types, docs, sim, simstudio, simstudio-ts-sdk
   • Running format:check in 25 packages
   • Remote caching disabled

::group::simstudio-ts-sdk:format:check
cache miss, executing 69fe2c650ec4955a
$ biome format .
Checked 6 files in 62ms. No fixes applied.
::endgroup::
::group::@sim/workflow-types:format:check
cache miss, executing d2a15219c4795d40
$ biome format .
Checked 4 files in 18ms. No fixes applied.
::endgroup::
::group::@sim/runtime-secrets:format:check
cache miss, executing fa96e3d91d14beaf
$ biome format .
Checked 5 files in 15ms. No fixes applied.
::endgroup::
::group::@sim/audit:format:check
cache miss, executing 16002eca79f888d0
$ biome format .
Checked 7 files in 40ms. No fixes applied.
::endgroup::
::group::@sim/workflow-persistence:format:check
cache miss, executing 276d5109e80f2ab0
$ biome format .
Checked 10 files in 95ms. No fixes applied.
::endgroup::
::group::@sim/workflow-renderer:format:check
cache miss, executing 3bd74cc6d31f2044
$ biome format .
Checked 13 files in 96ms. No fixes applied.
::endgroup::
::group::@sim/desktop-bridge:format:check
cache miss, executing b7ae9c199dd6c422
$ biome format .
Checked 5 files in 105ms. No fixes applied.
::endgroup::
::group::simstudio:format:check
cache miss, executing 1a315830226ea3e8
$ biome format .
Checked 3 files in 32ms. No fixes applied.
::endgroup::
::group::@sim/security:format:check
cache miss, executing 8bcbb4860d2fef77
$ biome format .
Checked 19 files in 123ms. No fixes applied.
::endgroup::
::group::@sim/auth:format:check
cache miss, executing ea836bbe6b5f56d2
$ biome format .
Checked 3 files in 24ms. No fixes applied.
::endgroup::
::group::@sim/platform-authz:format:check
cache miss, executing 00b26842745ffbe8
$ biome format .
Checked 7 files in 84ms. No fixes applied.
::endgroup::
::group::@sim/logger:format:check
cache miss, executing 4654a31aabcd552d
$ biome format .
Checked 6 files in 80ms. No fixes applied.
::endgroup::
::group::@sim/browser-protocol:format:check
cache miss, executing 4843a62c2a478709
$ biome format .
Checked 3 files in 37ms. No fixes applied.
::endgroup::
::group::@sim/terminal-protocol:format:check
cache miss, executing 9ec78829b5024124
$ biome format .
Checked 3 files in 34ms. No fixes applied.
::endgroup::
::group::@sim/realtime:format:check
cache miss, executing a447929d01b45d93
$ biome format .
Checked 52 files in 631ms. No fixes applied.
::endgroup::
::group::@sim/realtime-protocol:format:check
cache miss, executing d0bc88bd3aefedf4
$ biome format .
Checked 11 files in 107ms. No fixes applied.
::endgroup::
::group::@sim/testing:format:check
cache miss, executing 1975b92479a6bbd0
$ biome format .
Checked 73 files in 520ms. No fixes applied.
::endgroup::
::group::@sim/utils:format:check
cache miss, executing 43cb638c5b7de16f
$ biome format .
Checked 24 files in 175ms. No fixes applied.
::endgroup::
::group::@sim/emcn:format:check
cache miss, executing 02ef5e39e5d86528
$ biome format .
Checked 198 files in 1083ms. No fixes applied.
::endgroup::
::group::@sim/desktop:format:check
cache miss, executing 944f8dceb52cc76c
$ biome format .
Checked 134 files in 1497ms. No fixes applied.
::endgroup::
::group::docs:format:check
cache miss, executing 1a5a31d8c5d710bb
$ biome format .
Checked 102 files in 1857ms. No fixes applied.
::endgroup::
::group::@sim/db:format:check
cache miss, executing ae205447e1f91774
$ biome format .
Checked 309 files in 6s. No fixes applied.
::endgroup::
�[;31msim:format:check�[;0m
cache miss, executing 68df7b9187d5e872
$ biome format .
providers/models.ts format ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  × Formatter would have printed the following content:
  
    3491 3491 │    * returned by provider APIs (e.g. `claude-sonnet-4-5-20250514`).
    3492 3492 │    */
    3493      │ - export·function·findCatalogModel(
    3494      │ - ··modelId:·string
    3495      │ - ):·{
         3493 │ + export·function·findCatalogModel(modelId:·string):·{
    3496 3494 │     providerId: string
    3497 3495 │     model: (typeof PROVIDER_DEFINITIONS)[keyof typeof PROVIDER_DEFINITIONS]['models'][number]
  

Checked 12804 files in 17s. No fixes applied.
Found 1 error.
format ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  × Some errors were emitted while running checks.
  

error: script "format:check" exited with code 1

 Tasks:    22 successful, 23 total
Cached:    0 cached, 23 total
  Time:    18.182s 
Failed:    sim#format:check


$ turbo run format:check
::error::sim#format:check: command (/home/runner/work/p2-sim/p2-sim/apps/sim) /home/runner/.bun/bin/bun run format:check exited (1)
 ERROR  sim#format:check: command (/home/runner/work/p2-sim/p2-sim/apps/sim) /home/runner/.bun/bin/bun run format:check exited (1)
 ERROR  run failed: command  exited (1)
error: script "check" exited with code 1

Command failed: bun run check
$ turbo run format:check
::error::sim#format:check: command (/home/runner/work/p2-sim/p2-sim/apps/sim) /home/runner/.bun/bin/bun run format:check exited (1)
 ERROR  sim#format:check: command (/home/runner/work/p2-sim/p2-sim/apps/sim) /home/runner/.bun/bin/bun run format:check exited (1)
 ERROR  run failed: command  exited (1)
error: script "check" exited with code 1

bun run lint

❌ failed

━━━━

  ! Suppression comment has no effect. Remove the suppression or make sure you are suppressing the correct rule.
  
    65 │           { 'x-zdesk-jwt': 'not-a-real-jwt' },
    66 │           { orgId: '1', externalId: '2', credentialId: 'cred-1' }
  > 67 │           // biome-ignore lint/suspicious/noExplicitAny: minimal context for the fallback path
       │           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    68 │         ) as any
    69 │       )
  

lib/webhooks/providers/zoho-desk.test.ts:78:11 suppressions/unused ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  ! Suppression comment has no effect. Remove the suppression or make sure you are suppressing the correct rule.
  
    76 │           { 'x-zdesk-jwt': 'not-a-real-jwt' },
    77 │           { orgId: '1', externalId: '2', credentialId: 'cred-1', apiDomain: 'https://desk.zoho.eu' }
  > 78 │           // biome-ignore lint/suspicious/noExplicitAny: minimal context for the fast path
       │           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    79 │         ) as any
    80 │       )
  

lib/webhooks/providers/zoho-desk.test.ts:93:11 suppressions/unused ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  ! Suppression comment has no effect. Remove the suppression or make sure you are suppressing the correct rule.
  
    91 │           userId: 'user-1',
    92 │           requestId: 'test',
  > 93 │           // biome-ignore lint/suspicious/noExplicitAny: request is unused on these guard paths
       │           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    94 │           request: {} as any,
    95 │         })
  

lib/webhooks/providers/zoho-desk.test.ts:250:9 suppressions/unused ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  ! Suppression comment has no effect. Remove the suppression or make sure you are suppressing the correct rule.
  
    248 │         accountId: 'acc-1',
    249 │         userId: 'user-1',
  > 250 │         // biome-ignore lint/suspicious/noExplicitAny: partial owner shape for the test
        │         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    251 │       } as any)
    252 │       vi.mocked(refreshAccessTokenIfNeeded).mockResolvedValue('zoho-token')
  

lib/webhooks/providers/zoho-desk.test.ts:273:9 suppressions/unused ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  ! Suppression comment has no effect. Remove the suppression or make sure you are suppressing the correct rule.
  
    271 │         userId: 'user-1',
    272 │         requestId: 'test',
  > 273 │         // biome-ignore lint/suspicious/noExplicitAny: request is unused on this path
        │         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    274 │         request: {} as any,
    275 │       })
  

lib/webhooks/providers/zoho-desk.test.ts:291:9 suppressions/unused ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  ! Suppression comment has no effect. Remove the suppression or make sure you are suppressing the correct rule.
  
    289 │         accountId: 'acc-1',
    290 │         userId: 'user-1',
  > 291 │         // biome-ignore lint/suspicious/noExplicitAny: partial owner shape for the test
        │         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    292 │       } as any)
    293 │       vi.mocked(refreshAccessTokenIfNeeded).mockResolvedValue('zoho-token')
  

lib/webhooks/providers/zoho-desk.test.ts:314:9 suppressions/unused ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  ! Suppression comment has no effect. Remove the suppression or make sure you are suppressing the correct rule.
  
    312 │         userId: 'user-1',
    313 │         requestId: 'test',
  > 314 │         // biome-ignore lint/suspicious/noExplicitAny: request is unused on this path
        │         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    315 │         request: {} as any,
    316 │       })
  

lib/core/config/api-keys.ts:62:14 lint/suspicious/noDuplicateElseIf ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  × This branch can never execute. Its condition is a duplicate or covered by previous conditions in the if-else-if chain.
  
    60 │     if (env.ZAI_API_KEY_2) keys.push(env.ZAI_API_KEY_2)
    61 │     if (env.ZAI_API_KEY_3) keys.push(env.ZAI_API_KEY_3)
  > 62 │   } else if (provider === 'xai') {
       │              ^^^^^^^^^^^^^^^^^^
    63 │     if (env.XAI_API_KEY_1) keys.push(env.XAI_API_KEY_1)
    64 │     if (env.XAI_API_KEY_2) keys.push(env.XAI_API_KEY_2)
  

local-copilot/lib/agent/orchestrator.billing.test.ts:70:25 lint/correctness/useYield ━━━━━━━━━━━━━━━

  × This generator function doesn't contain yield.
  
    69 │ vi.mock('@/local-copilot/lib/agent/specialists/parallel-subagents', () => ({
  > 70 │   runParallelSubagents: async function* () {
       │                         ^^^^^^^^^^^^^^^^^^^^
  > 71 │     return { findings: '', results: [], events: [] }
  > 72 │   },
       │   ^
    73 │ }))
    74 │ 
  

local-copilot/lib/agent/orchestrator.billing.test.ts:76:22 lint/correctness/useYield ━━━━━━━━━━━━━━━

  × This generator function doesn't contain yield.
  
    75 │ vi.mock('@/local-copilot/lib/agent/specialists/specialist-pass', () => ({
  > 76 │   runSpecialistPass: async function* () {
       │                      ^^^^^^^^^^^^^^^^^^^^
  > 77 │     return { domain: 'research', findings: '', toolRoundCount: 0, events: [] }
  > 78 │   },
       │   ^
    79 │ }))
    80 │ 
  

tools/figma/figma_to_html_ai.ts:200:5 lint/correctness/noUnreachable ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  × This code will never be reached ...
  
    198 │     }
    199 │ 
  > 200 │     return {
        │     ^^^^^^^^
  > 201 │       success: true,
  > 202 │       output: {
  > 203 │         metadata: data.metadata,
  > 204 │       },
  > 205 │     }
        │     ^
    206 │   },
    207 │ 
  
  i ... because either this statement will throw an exception, ...
  
    115 │     if (!params) {
  > 116 │       throw new Error('Missing required parameters')
        │       ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    117 │     }
    118 │ 
  
  i ... this statement will return from the function, ...
  
    146 │       )
    147 │ 
  > 148 │       return {
        │       ^^^^^^^^
  > 149 │         success: true,
         ...
  > 162 │         },
  > 163 │       }
        │       ^
    164 │     } catch (error) {
    165 │       const errorMessage = error instanceof Error ? error.message : String(error)
  
  i ... or this statement will return from the function beforehand
  
    180 │       cleanedHtml = cleanedHtml.trim() // trim ends
    181 │ 
  > 182 │       return {
        │       ^^^^^^^^
  > 183 │         success: false,
         ...
  > 196 │         error: data.error || 'Figma to HTML conversion failed',
  > 197 │       }
        │       ^
    198 │     }
    199 │ 
  

Checked 12804 files in 43s. Fixed 66 files.
Found 4 errors.
Found 9 warnings.
check ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

  × Some errors were emitted while running checks.
  

error: script "lint" exited with code 1

 Tasks:    22 successful, 23 total
Cached:    0 cached, 23 total
  Time:    44.937s 
Failed:    sim#lint


$ turbo run lint
::error::sim#lint: command (/home/runner/work/p2-sim/p2-sim/apps/sim) /home/runner/.bun/bin/bun run lint exited (1)
 ERROR  sim#lint: command (/home/runner/work/p2-sim/p2-sim/apps/sim) /home/runner/.bun/bin/bun run lint exited (1)
 ERROR  run failed: command  exited (1)
error: script "lint" exited with code 1

Command failed: bun run lint
$ turbo run lint
::error::sim#lint: command (/home/runner/work/p2-sim/p2-sim/apps/sim) /home/runner/.bun/bin/bun run lint exited (1)
 ERROR  sim#lint: command (/home/runner/work/p2-sim/p2-sim/apps/sim) /home/runner/.bun/bin/bun run lint exited (1)
 ERROR  run failed: command  exited (1)
error: script "lint" exited with code 1

bun run test

❌ failed


   • Packages in scope: @sim/audit, @sim/auth, @sim/browser-protocol, @sim/db, @sim/desktop, @sim/desktop-bridge, @sim/emcn, @sim/logger, @sim/pii, @sim/platform-authz, @sim/realtime, @sim/realtime-protocol, @sim/runtime-secrets, @sim/security, @sim/terminal-protocol, @sim/testing, @sim/tsconfig, @sim/utils, @sim/workflow-persistence, @sim/workflow-renderer, @sim/workflow-types, docs, sim, simstudio, simstudio-ts-sdk
   • Running test in 25 packages
   • Remote caching disabled

�[;31m@sim/db:test�[;0m
cache miss, executing 63181e6cac6942a9
$ vitest run
/usr/bin/bash: line 1: vitest: command not found
error: script "test" exited with code 127
::group::@sim/workflow-persistence:test
cache miss, executing 99c24dc26410f451
::endgroup::
::group::@sim/audit:test
cache miss, executing 6c37cfb71f4ab302
::endgroup::
::group::@sim/realtime-protocol:test
cache miss, executing 73b48eb82598a96d
$ vitest run
::endgroup::
::group::@sim/security:test
cache miss, executing 423cdeae94291bca
$ vitest run
::endgroup::
::group::@sim/runtime-secrets:test
cache miss, executing 72e78c533572b708
$ vitest run
::endgroup::
::group::simstudio-ts-sdk:test
cache miss, executing bd07a4cb424f9c8c
$ vitest run
::endgroup::
::group::@sim/testing:test
cache miss, executing 2ecaca1817e52efc
$ vitest run
::endgroup::
::group::@sim/emcn:test
cache miss, executing 5b53501084edb728
$ vitest run
::endgroup::
::group::@sim/desktop:test
cache miss, executing ac91fabd88f3e650
$ vitest run
::endgroup::
::group::@sim/utils:test
cache miss, executing c5c01caeb87a3150
$ vitest run
::endgroup::
::group::@sim/logger:test
cache miss, executing dfaa973b139c131f
$ vitest run
::endgroup::

 Tasks:    0 successful, 12 total
Cached:    0 cached, 12 total
  Time:    1.231s 
Failed:    @sim/db#test


$ turbo run test
::error::@sim/db#test: command (/home/runner/work/p2-sim/p2-sim/packages/db) /home/runner/.bun/bin/bun run test exited (127)
 ERROR  @sim/db#test: command (/home/runner/work/p2-sim/p2-sim/packages/db) /home/runner/.bun/bin/bun run test exited (127)
 ERROR  run failed: command  exited (127)
error: script "test" exited with code 127

Command failed: bun run test
$ turbo run test
::error::@sim/db#test: command (/home/runner/work/p2-sim/p2-sim/packages/db) /home/runner/.bun/bin/bun run test exited (127)
 ERROR  @sim/db#test: command (/home/runner/work/p2-sim/p2-sim/packages/db) /home/runner/.bun/bin/bun run test exited (127)
 ERROR  run failed: command  exited (127)
error: script "test" exited with code 127

bun run build

❌ failed


   • Packages in scope: @sim/audit, @sim/auth, @sim/browser-protocol, @sim/db, @sim/desktop, @sim/desktop-bridge, @sim/emcn, @sim/logger, @sim/pii, @sim/platform-authz, @sim/realtime, @sim/realtime-protocol, @sim/runtime-secrets, @sim/security, @sim/terminal-protocol, @sim/testing, @sim/tsconfig, @sim/utils, @sim/workflow-persistence, @sim/workflow-renderer, @sim/workflow-types, docs, sim, simstudio, simstudio-ts-sdk
   • Running build in 25 packages
   • Remote caching disabled

�[;31msim:build�[;0m
cache miss, executing f91600b78494de8e
$ bun run build:sandbox-bundles && NODE_OPTIONS='--max-old-space-size=8192' next build
$ bun run ./lib/execution/sandbox/bundles/build.ts
[2026-08-04T10:36:47.684Z] [ERROR] [SandboxBundleBuild] sandbox bundle build failed {
  "message": "Bundle failed",
  "name": "AggregateError"
}
error: script "build:sandbox-bundles" exited with code 1
::group::docs:build
cache miss, executing 101dc744b88a53b6
$ fumadocs-mdx && NODE_OPTIONS='--max-old-space-size=8192' next build
::endgroup::
::group::@sim/desktop:build
cache miss, executing 4971da91428f9475
$ bun run scripts/ensure-pty-prebuilds.ts && bun run scripts/build.ts
• Fetching node-pty prebuild: darwin-arm64@1.2.0-beta.12
• Fetching node-pty prebuild: darwin-x64@1.2.0-beta.12
::endgroup::
::group::simstudio:build
cache miss, executing e7361d873ddfd3dc
$ tsc
::endgroup::
::group::simstudio-ts-sdk:build
cache miss, executing a0bd6b9845f4a7d4
$ tsc
::endgroup::

 Tasks:    0 successful, 5 total
Cached:    0 cached, 5 total
  Time:    1.083s 
Failed:    sim#build


$ turbo run build
::error::sim#build: command (/home/runner/work/p2-sim/p2-sim/apps/sim) /home/runner/.bun/bin/bun run build exited (1)
 ERROR  sim#build: command (/home/runner/work/p2-sim/p2-sim/apps/sim) /home/runner/.bun/bin/bun run build exited (1)
 ERROR  run failed: command  exited (1)
error: script "build" exited with code 1

Command failed: bun run build
$ turbo run build
::error::sim#build: command (/home/runner/work/p2-sim/p2-sim/apps/sim) /home/runner/.bun/bin/bun run build exited (1)
 ERROR  sim#build: command (/home/runner/work/p2-sim/p2-sim/apps/sim) /home/runner/.bun/bin/bun run build exited (1)
 ERROR  run failed: command  exited (1)
error: script "build" exited with code 1

Agent usage

child-cluster-1

  • Model: gpt-5.6-luna
  • Iterations: 1
  • Input tokens (direct): 102,935
  • Input tokens (cache read): 1,823,079
  • Input tokens (cache create): 0
  • Input tokens (total): 1,926,014
  • Output tokens: 18,499
  • Cost: $0.079247 (estimated fallback)

Totals

  • Total input tokens: 1,926,014
  • Total output tokens: 18,499
  • Primary models: gpt-5.6-luna
  • Total cost: $0.079247
  • Estimated cost (fallback): $0.079247

Cost by agent

  • child-cluster-1: $0.079247 (estimated fallback)

Draft — mark ready for review when satisfied.

waleedlatif1 and others added 30 commits July 23, 2026 17:17
…se (simstudioai#5906)

* fix(landing): restore single-paragraph hero copy on home and enterprise

* fix(landing): tighten enterprise hero description

* revert(landing): restore original home and enterprise hero headings

* improvement(landing): simplify enterprise hero description

* improvement(landing): restore building-and-managing homepage headline
… pricing (simstudioai#5913)

* improvement(library): add citations and internal links, correct stale pricing

The library had zero third-party citations across 16 posts (~51k words) and
averaged 1.2 internal links per post, with six posts at zero. Auditing every
dollar figure against the vendor's own pricing page while adding the citations
turned up several claims that had gone stale, plus one product that is being
shut down.

Corrections, each verified against the vendor's page on 2026-07-23:
- Relay.app is winding down (signups closed 2026-07-16, free accounts end
  2026-08-15, paid 2026-09-14). It was recommended as the human-in-the-loop
  pick in best-zapier-alternatives; that recommendation now points to a
  platform you can still sign up for
- Make Core is $12/mo billed monthly, not $10.59; annual saves ~15%. One post
  also described $12 as the annual rate
- Pabbly Connect lifetime starts at $349, not $249, and the Standard/Pro/
  Ultimate tier structure quoted no longer exists
- Workato publishes no pricing at all, so the "~$1,000/month" figure is
  replaced with the fact that every deal is quoted through sales
- Dify is at ~149k GitHub stars, not 131k
- n8n cloud Starter is EUR 20/mo billed annually for 2,500 executions

Verified-correct figures were left alone and given a source link: Zapier
$29.99/mo monthly for 750 tasks, Activepieces 10 free flows then $5/flow/mo,
Power Automate $15/user/mo, Lindy $49.99/mo, and the Grand View Research RPA
market figures.

Also: 59 outbound citations and 3-6 internal links per post (zero posts left
without internal links), and MDX external links now carry target=_blank plus
rel=noopener noreferrer, which the landing SEO/GEO rule requires but the MDX
anchor was not applying.

* fix(content): treat only non-Sim hosts as external links in MDX

The first pass classified any http(s) href as external, so the 33 absolute
Sim URLs already in content (sim.ai, sim.ai/slack, www.sim.ai/blog/*, and 14
docs.sim.ai pages) would have opened in a new tab with rel=noopener, which is
wrong for first-party links and contradicted the comment above the check.

Classification now compares the hostname against the apex derived from
SITE_URL, so the apex and any Sim subdomain stay same-tab. SITE_URL is used
rather than getBaseDomain() so a post renders identically in dev, preview,
and production instead of varying with NEXT_PUBLIC_APP_URL. The leading dot
in the suffix check keeps lookalikes such as evil-sim.ai external.
…oard (simstudioai#5907)

* fix(helm): correct chart docs, examples, and dead config across the board

Audit-driven accuracy pass over the chart's entire documentation surface,
verified by rendering every example against the templates:

- migrations run as an init container on the app pod, not a Job — fix the
  README component list, troubleshooting commands, and sim-helm skill refs;
  drop the dead migrations-job NetworkPolicy ingress rule
- referenced-but-never-created resources: document the GKE ManagedCertificate
  creation (values-gcp), comment out the key-file Secret mount that stuck all
  pods in ContainerCreating (values-gcp), enable certManager for the postgres
  TLS issuerRef (values-production), add the cert-manager cluster-issuer
  annotation nginx needs (values-azure)
- values-external-db: networkPolicy.egress is a list, not a map (the map
  rendered an invalid manifest); fill schema-failing placeholder host/username
- realtime >1 replica requires REDIS_URL (Socket.IO Redis adapter) — default
  examples to 1 replica with the scaling note, and warn where autoscaling
  HPAs override replicaCount
- pod anti-affinity selectors matched nothing (simstudio vs sim name label)
- kubernetes.mdx: install commands were missing required CRON_SECRET and
  postgresql password (failed at template time), wrong deployment name in
  port-forward, stale version requirements, unsupported key-remapping claim
- remove unimplemented app.secrets.existingSecret.keys from values + schema;
  fix README PDB default, cronjob list, /metrics caveat, NOTES secret count,
  Azure-only StorageClass in generic examples, dead SOCKET_SERVER_URL and
  GOOGLE_CLOUD_* env, ESO apiVersion mismatch, and skill-reference drift
- bump chart to 1.0.1

* fix(helm): review round 1 — scoped example egress, in-tab secret note, copilot Job wording

- external-db example egress scopes to a placeholder database CIDR instead of
  to: [] (which allowed every destination on 5432, defeating the isolation
  the example teaches)
- kubernetes.mdx cloud tabs state explicitly that they reuse the variables
  generated in the Installation block
- Copilot migrations really do run as a Helm-hook Job — restore Job wording
  there (only the app migrations are an init container)

* fix(helm): template-sweep fixes — telemetry validity, ESO rollout checksums, dead passwordKey knob

- telemetry: memory_limiter gets the required check_interval (collector
  failed startup validation whenever telemetry.enabled=true); the jaeger
  exporter was removed from collector-contrib in v0.86 — export to Jaeger
  via its native OTLP endpoint instead (otlp/jaeger, default port 4317)
- app/realtime rollout checksums now hash the ExternalSecret manifest too,
  mirroring the copilot pattern — with ESO enabled the inline Secret renders
  empty, so remoteRefs changes never rolled the pods
- remove the unimplemented existingSecret.passwordKey knob (values, schema,
  README, dead helpers): nothing consumed it, and a non-default value
  silently produced a DATABASE_URL with an unexpandable placeholder; secrets
  must use the standard POSTGRES_PASSWORD / EXTERNAL_DB_PASSWORD keys
- drop the orphaned sim.migrations.labels helper (its only consumer was the
  dead NetworkPolicy rule removed earlier)
- helm test pod image resolves through sim.image so global.imageRegistry
  mirroring applies; NetworkPolicy realtime-ingress comment reflects actual
  traffic direction; smoke unittest suite loads the newly referenced
  external-secret template

* chore(helm): bump chart to 1.1.0 with upgrade notes

Removing (inert) documented values keys and changing the rollout-checksum
inputs is a values-surface change — per SemVer chart conventions that is
more than a patch. Adds an Upgrading section documenting the one-time pod
roll, the removed no-op keys, and the Jaeger-over-OTLP change.

* fix(helm): review round — external-db NP opt-in with real-CIDR-first flow, prod Jaeger OTLP endpoint

- external-db example ships networkPolicy disabled so a verbatim install
  always reaches the database; the scoped egress rule stays as the documented
  opt-in (set your CIDR first, then enable)
- values-production still pointed telemetry.jaeger at the legacy 14250
  collector port — now Jaeger's OTLP gRPC endpoint to match the otlp/jaeger
  exporter

* feat(helm): autoscaling.realtime.enabled toggle so examples can scale the app without unsafe realtime replicas

Cursor correctly flagged that comment-level warnings didn't stop a verbatim
production/external-db install from running the realtime HPA at minReplicas 2
without REDIS_URL (silent cross-pod event loss). Adds an opt-out toggle
(default true — existing deployments unchanged): the realtime HPA renders only
when autoscaling.realtime.enabled, and the realtime Deployment keeps
spec.replicas under its control when the HPA is excluded. The three
autoscaling examples set it false with the Redis rationale; README and
upgrade notes document the toggle.

* fix(helm): review round — whitelabeled realtime HPA opt-out, external-db isolation on with required CIDR in install flow

- values-whitelabeled now actually sets autoscaling.realtime.enabled: false
  (the earlier batch aborted before reaching this file — Cursor caught it)
- external-db keeps networkPolicy enabled (no isolation regression); the DB
  egress CIDR is marked REQUIRED and wired into both documented install
  commands via --set, so the copy-paste flow sets the real subnet in the same
  breath as the DB host

* fix(helm): use a syntactically valid example CIDR in the external-db install commands

An unreplaced <YOUR_DB_CIDR> literal fails Kubernetes CIDR validation and
aborts the install; every sibling placeholder in the same command (host,
username) installs fine and simply doesn't connect until replaced. The CIDR
now behaves the same way: valid example value (10.20.0.0/24), explicitly
marked as the operator's database subnet in both commands and the values
comment.

* fix(helm): remove text after continuation backslashes in external-db install commands

A trailing comment after the line-continuation backslash (and equally a
comment line spliced mid-command) breaks the shell command when copied —
the remaining --set overrides run as separate commands and required-value
validation fails. Both documented commands now reconstruct to bash-clean
multi-line invocations (verified with bash -n); the CIDR guidance lives in
the networkPolicy section comment.

* fix(docs): cloud-tab installs are alternatives via helm upgrade --install

Following Installation and then a cloud tab ran helm install twice for the
same release and failed on the second. The tabs now state they replace the
generic install and use helm upgrade --install, which is idempotent and also
converts an existing generic install to the cloud values.

* fix(docs): honest conversion caveat for cloud-tab upgrades over an existing install

The cloud values rename the bundled Postgres database to simstudio, which
Postgres only applies at first initialization — an in-place conversion of a
generic install would point DATABASE_URL at a nonexistent database. Document
the two safe paths: keep the original name via --set, or uninstall + delete
PVCs and install fresh.

* fix(helm): external-db header install command declares its secrets and includes CRON_SECRET

The primary documented command failed the chart's required-value validation
(missing app.env.CRON_SECRET with cronjobs default-on) and referenced an
undeclared DB_PASSWORD. It now shows the export lines for every variable it
uses and sets CRON_SECRET; both commands verified with bash -n.

* chore(helm): restore values.schema.json formatting — surgical deletions only

The earlier programmatic edit reformatted the whole file (~390 lines of
whitespace churn hiding the 12 real deleted lines). Re-applied the removal
of the dead keys/passwordKey properties as text-level deletions preserving
the original style.

* fix(helm+docs): explicit secret exports in external-db header; conversion must reuse original secrets

- all five export lines are written out (three were only named in a trailing
  comment, so a verbatim copy passed empty required values)
- the cloud-conversion caveat now leads with reusing the original secret
  values (helm get values) — a regenerated ENCRYPTION_KEY makes previously
  encrypted credentials undecryptable
* feat(agent-stream): add agent-events thinking/tool streaming for chat and canvas

Ship the agent-events-v1 protocol with provider tool loops, dual-gated chat thinking, DeepSeek/Groq/OpenAI reasoning wiring, and ChatGPT-like thinking chrome.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): clear stuck streaming UI and format db snapshot

Biome was failing CI on migrations/meta/0261_snapshot.json. Also settle
assistant streaming/tool flags when SSE ends without a terminal frame,
without clobbering Stop's finalized content.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): satisfy biome format and import order

Auto-format the sim package for CI lint:check, and repair the Anthropic
streaming tool-loop payload after an unsafe delete-to-undefined rewrite.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): keep drained answer on abort and update migration journal test

Treat AbortError from reader.cancel as a cancelled pump result so soft-complete
retains answerText. Point the workspace storage migration journal assertion at
0261_chat_include_thinking.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): keep Stop notice when server emits cancel error

Ignore terminal SSE error frames after the user aborts so
"Client cancelled request" cannot overwrite "Response stopped by user".

Co-authored-by: Cursor <cursoragent@cursor.com>

* improvement(chat): ChatGPT-style thinking shimmer and stick-to-bottom scroll

Add left-to-right shimmer on live thinking label/body, keep scroll working by
shimmering an inner node, and follow the answer only while near the bottom.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): stop pump on client disconnect; soft-complete agents only

Abort the agent stream pump when the projected HTTP body is cancelled so
provider work does not continue after disconnect. Limit AbortError soft-success
to Agent blocks so Function/HTTP cancels still fail in logs.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): persist includeThinking across pause snapshots

Paused chat runs with Include thinking enabled were dropping the flag when
serializing the pause snapshot, so resume always rebuilt streams without
thinking/tool SSE frames.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): keep drained answer text when stream times out

Persist pump answerText onto the streaming execution before throwing on
timeout, and carry that partial content into the failed block output so
logs match what the client already saw.

Co-authored-by: Cursor <cursoragent@cursor.com>

* improvement(chat): auto-collapse tools chrome when tool streaming ends

Match thinking UX: open while tools run, collapse when finished, and keep
the panel open only if the user manually reopens it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): settle canvas stream chrome on failure paths

Clear agentStreamActive and settle running tool chips when blocks error,
timeouts cancel runs, or execution ends without stream:done so the output
panel does not stay on live Thinking/Using tools chrome.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(lint): organize imports in terminal console store

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): mark open tools cancelled on HITL pause

Pause can interrupt a tool loop without tool end events; settling those
chips as success incorrectly showed unfinished tools as complete.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(db): drop branch-local 0261 migration ahead of staging merge

* chore(db): regenerate include_thinking migration as 0266 post staging merge

* fix(providers): resolve type errors in streaming tool loop call sites

* fix(agent-stream): gate agent events opt-in and correct provider loop behavior

- streamToolCalls and provider thinking requests now require run-level
  agentEvents opt-in (canvas on, chat dual-gated, API off) so existing
  runs keep pre-agent-events behavior exactly
- OpenAI reasoning summaries opt-in + strip-and-retry on unverified-org 400
- streaming loops run tool postProcess again (firecrawl/exa async results)
- bedrock live loop falls back to silent path for responseFormat
- deepseek: reasoning_content pass-back unconditional, 'none' sends disabled
- groq: x_groq.usage fallback, reasoning params gated, qwen none disables
- gemini: functionCall parts echoed verbatim, local ids only for events
- truncated turns (max_tokens/length) no longer execute partial tool calls
- MAX_TOOL_ITERATIONS exit flushes last turn text as final answer
- iterations reports actual model calls; shared loop plumbing extracted

* refactor(agent-stream): consolidate protocol, dedupe client/server plumbing, hygiene

- canonical ChatStreamFrame union + type guards consumed by server emitters
  and the chat client; stream_error restored to legacy log-only handling
- strip thinking/tool args from providerTiming on public final envelopes
- shared tool-chip lifecycle module for chat, canvas, and console store
- shared sink-to-execution-events forwarder replaces the copy-pasted
  adapter in the execute route and HITL manager; LIVE_ONLY event set shared
- stream:thinking payload field renamed data->text; canvas thinking batched
- abort reasons carried as AbortError DOMExceptions so raw fetch consumers
  classify correctly; thinking cap renamed to chars and scope-documented
- kimi wired for agent events like the other compat providers
- deleted dead exports/step-N comments; fixtures match real wire shapes;
  loop tests use explicit mocks instead of importOriginal

* test(agent-stream): cover the dual-gated execution path and typed abort reasons

- chat route tests assert agentEvents reaches executeWorkflow only when
  policy and protocol header agree
- execution-limits tests assert AbortError-typed reasons
- executor metadata type carries agentEvents

* fix(deploy-modal): align include-thinking spacing with the modal's 6.5px rhythm

* docs(agent-stream): autogenerate per-model thinking/tool stream support on the Agent block page

- capabilities.thinking.streamed ('full' | 'summary' | 'none') on models.ts,
  explicit for the Anthropic family where visibility varies per generation;
  getThinkingStreamVisibility exposes the derivation for docs and UI alike
- scripts/sync-agent-stream-docs.ts regenerates the support tables between
  markers in workflows/blocks/agent.mdx from the model registry and
  STREAMING_TOOL_CALL_PROVIDERS; --check fails on drift or missing metadata
- wired agent-stream-docs:check into CI next to the other sync gates

* feat(anthropic): request summarized thinking display for omitted-default Claude models

The newest Claude generations (Fable 5, Sonnet 5, Opus 4.8/4.7) default
thinking.display to omitted — empty thinking blocks, no deltas. On
agent-events runs Sim now opts back in with display: 'summarized', driven
by the registry's streamed metadata; legacy runs keep the exact
pre-agent-events request shape. Registry, generated docs, and the family
capability table updated accordingly.

* docs(skills): cover thinking.streamed and agent-stream docs sync in model skills

* chore(deps): upgrade @anthropic-ai/sdk to 0.114.0 and adopt official types

- adaptive thinking, display, and output_config are now SDK-typed; the only
  remaining custom payload field is output_format (beta-header structured
  outputs, which the SDK models as output_config.format instead)
- anthropic stream events narrow on the SDK's discriminated unions instead
  of anonymous casts; compat deltas type content/tool_calls from the OpenAI
  SDK with vendor reasoning fields as an explicit optional extension
- @sim/auth exposes an explicit VerifyAuth contract so its declarations no
  longer reference better-auth's nested zod instance (TS2883 under fresh
  install layouts); realtime consumer aligned
- docs app zod pinned to the repo's exact 4.3.6 so ai SDK types bind the
  same zod instance (docs type-check was latently broken)
- knowledge embedding tests made hermetic against local .env keys and
  hosted rotation fallback

* refactor(providers): replace legacy as-any stream casts with annotated typed casts

* refactor(providers): finish provider audit — remove dead byte-stream helper, annotate remaining legacy casts

Audit of all 26 providers for the agent-events feature confirmed every
streaming execution declares agent-events-v1 and every adapter emits
AgentStreamEvent objects. Cleanup from the audit: the unconsumed legacy
createOpenAICompatibleStream byte helper is deleted, and the remaining
streamResponse-as-any casts (xai, nvidia, kimi, meta, zai, sakana) are
annotated typed casts matching the groq/deepseek fix.

* feat(streaming): stream answer text live during tool loops via turn_end protocol

The live tool loops buffered all answer text per model turn (classification
of intermediate vs final is only known at turn end), so gated surfaces saw
thinking stream, then dead air with the thinking chrome stuck open, then the
whole answer at once.

Loops now emit text deltas live as `turn: 'pending'` plus a `turn_end`
event per turn. The pump buffers pending text and projects it to the byte
path (answerText/logs/memory/legacy clients) only on a final turn_end, so
all settled semantics are unchanged. Gated surfaces render the pending text
as it streams and reconcile with a reset when a turn resolves to tools:

- public chat: live `chunk` frames from the sink + dual-gated `chunk_reset`;
  byte-path frame emission is suppressed to avoid duplicates (kept for
  response-format transformed streams via clientStreamTransformed)
- canvas: forwarder emits live `stream:chunk` + `stream:chunk_reset`; the
  execute route and HITL resume readers stop re-emitting byte chunks; panel
  chat tracks per-block segments and replaces content on flush
- chat client: per-block text segments, chunk_reset handling, and thinking
  chrome now settles on tool start as well as first answer chunk

* fix(streaming): address validated review findings across provider gating and reset reconciliation

Three-reviewer pass over the branch, findings validated against staging:

- agent-handler forwards agentEvents to executeProviderRequest — the flag was
  computed but dropped in the field-by-field copy, so provider-side thinking
  requests (OpenAI summaries, Gemini includeThoughts, Anthropic summarized
  display) never activated on opted-in runs
- openai: restore summary:'auto' alongside explicit reasoning effort — staging
  always paired them; gating summary purely on agentEvents changed legacy
  payloads
- gemini: Gemini 2 + tools + responseFormat falls back to the silent path;
  the live loop never applied the deferred responseSchema for AUTO tools
- openai-compat loop: malformed tool-argument JSON fails the call instead of
  executing with defaulted {} args (staging parsed inside the execution try)
- openai-compat parser: a vendor id arriving after a synthesized start no
  longer renames the call (start/end ids stayed consistent)
- stream-pump: abort closes the byte projection so a drain blocked on
  backpressure cannot deadlock teardown
- chunk_reset removes the block from the client text order (deployed chat +
  panel chat) so a reset block re-registers at arrival position — fixes
  separator/order corruption when parallel blocks stream around a reset
- resume route echoes the negotiated X-Sim-Stream-Protocol response header
  (parity with the chat route); docs: [DONE] wire shape + final-vs-error
  terminal semantics corrected

* chore(deps): exempt pinned @anthropic-ai/sdk 0.114.0 from the release-age gate

CI's bun install --frozen-lockfile blocks 0.114.0 (published 2026-07-23,
younger than the 7-day supply-chain gate). The pin is exact and was vetted
for the agent-events streaming work; following the existing bunfig pattern,
the exclusion ages out on 2026-07-30 and should be dropped then.

* chore(providers): fix double-cast-allowed annotation placement for the strict boundary audit

The audit only recognizes the annotation on the line directly above the cast;
two annotations had drifted behind intervening code lines (groq stream params,
deepseek loop messages) and the OpenAI reasoning-summary widening cast was
never annotated. No behavior change.

* fix(chat): settle straggler tool chips as error when final reports failure

A failed run can still terminate with a `final` frame carrying success: false;
running chips previously settled green regardless of the outcome.

* fix(canvas): wire agent stream chrome into run-from-block

Run-from-block executions emit the same live stream:thinking/stream:tool
events as full runs but registered none of the handlers, so the terminal
never showed thinking or tool chips on that path. The per-run chrome
(batched thinking writes + tool chip lifecycle + settlement on stream done,
block error, and every terminal execution state) is extracted into a shared
createAgentStreamChrome factory consumed by both paths.

---------

Co-authored-by: Bill Leoutsakos <billleoutsakos@Bills-MacBook-Pro.local>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Vikhyath Mondreti <vikhyath@simstudio.ai>
* feat(skills): permissions layer

* chore(db): drop skill_member migration 0261 for regeneration on latest staging

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(db): regenerate skill_member migration as 0262 on latest staging

Same DDL as the dropped 0261 (skill_member table, enums, indexes, skill.workspace_shared)
plus the hand-written write-user backfill, renumbered after staging's 0261.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(db): regenerate skill_member migration as 0263 after staging merge

Staging claimed 0262 (strong_storm); same DDL plus the hand-written
write-user backfill, renumbered on the merged snapshot chain.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(db): drop skill_member migration 0263 for regeneration on latest staging

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(db): regenerate skill_member migration as 0264 after staging merge

Staging claimed 0263 (workflow_fork_sync_excluded); same DDL plus the
hand-written write-user backfill, renumbered on the merged snapshot chain.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* make editing skills full page

* fix disclaimer

* edit access msg

* fix lint

* chore(db): drop skill_member migration 0264 for regeneration on latest staging

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(db): regenerate skill_member migration as 0265 after staging merge

Staging claimed 0264 (fat_ikaris); same DDL plus the hand-written
write-user backfill, renumbered on the merged snapshot chain.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix tests

* simplify system

* fix

* fix lint

* add mship skills docs

* chore(db): drop skill_member migration 0265 for regeneration on latest staging

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(db): regenerate skill_member migration as 0266 after staging merge

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(deps): override zod to 4.3.6 to dedupe nested copies breaking type-check

better-auth 1.6.23 and fumadocs-mdx resolve ^4.3.6 to a nested zod 4.4.3,
which makes @sim/auth's inferred betterAuth types non-portable (TS2883) and
split docs onto a second zod instance. Both ranges accept the repo-wide
pinned 4.3.6, so a single hoisted copy satisfies everything.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix lint

* feat(skills,tools): fullscreen skill create + shared custom tool editor

Moves the rich-markdown and custom-tool editing surfaces out of modals and
onto full-page surfaces, and collapses the duplicated chrome behind shared
components.

Skills
- Add /skills/new, a full-page create surface mirroring the skill detail page
  (CredentialDetailLayout + DetailSection + unsaved-changes guard). "Add to
  Sim" navigates there instead of opening a modal.
- Import moves to a header action (SkillImportButton) backed by a shared
  readSkillFile helper; the GitHub-URL import and its /api/skills/import route
  are removed.
- Skill name validation is now one shared validateSkillName, replacing three
  copies of the kebab-case rule and its messages.
- The skill editor roster renders through the shared MemberRow instead of
  re-deriving its identity block, with a locked role control and a lock-reason
  tooltip explaining inherited workspace-admin access.

Custom tools
- Extract the canvas modal's schema/code editors into a shared
  custom-tool-editor module (fields, wand generation, schema helpers), cutting
  custom-tool-modal.tsx by ~900 lines.
- Settings > Custom tools gains a full-page detail sub-view (SettingsPanel +
  SettingsSection + saveDiscardActions), deep-linkable via ?custom-tool-id.
  Rows are clickable; delete now lives only in the detail view.
- Replace legacy Button/Input/Badge/Label with the chip family, move chip-field
  chrome into CodeEditor behind an error prop, and delete its dead wand button.

Rich markdown field
- maxHeight is now opt-in: omit it on a page and the editor grows with its
  content so the page owns the only scrollbar. Modals pass explicit caps.
- The field variant drops to font-weight 400 to match adjacent chip fields.

* fix(skills): address review round on create navigation, 409 copy, and editor audit

- Skill create navigated using the first element of the upsert response, but
  that endpoint returns the caller's whole skill list (built-ins prepended) —
  match the new skill by its workspace-unique name instead.
- The suggested-skill 409 toast claimed the skill existed but was not shared
  and told the user to ask a skill admin. Every workspace member can already
  see and use every skill, so a 409 only means the name is taken.
- Adding an editor emitted the skill_shared event and SKILL_MEMBER_ADDED audit
  even when onConflictDoNothing skipped the insert on a concurrent add. Gate
  both on the insert actually returning a row.

* chore: format skills-resolver test import

* fix(skills,tools): audit fixes — autocomplete boundary, resize clipping, error routing

Two real regressions introduced while simplifying the extracted editor:

- The schema-param autocomplete's trigger was rewritten to match a trailing
  identifier, but the completion still split on separators. The two disagreed,
  so typing `data.ci` opened the menu and selecting replaced `data.ci` whole —
  eating the member-access prefix. Both now share one SCHEMA_PARAM_WORD regex.
- The uncapped markdown field measured its height only on value change while
  always setting overflow-hidden, so any width change that re-wrapped lines
  clipped the tail with no scrollbar to reach it. Now re-measures via
  ResizeObserver.

Also from the audit:
- Generation writes bypass the code field's change handler, so an open
  autocomplete stayed over a disabled streaming editor; close it on busy.
- Delete failures rendered in the Schema section's error slot on both custom
  tool surfaces; route them to a toast instead.
- Skill create navigated away while still dirty, stranding the unsaved-changes
  guard's history sentinel so Back landed on an empty create form.
- The Description field on skill create never received its error border.
- Drop a double-applied opacity-50 (the editor already dims when disabled), a
  dead try/catch around a non-throwing call that also shadowed the error prop,
  and a stale reference to a /tools page that does not exist.
- Docs still described the removed GitHub-URL import and the old Add Skill
  dialog; rewrite for the create page and file/paste import.

* feat(tools): read-only tool detail, create lands on the new tool, drop dead wand prompt API

- Viewers without edit rights could not open a custom tool at all, while the
  equivalent skill and custom-block surfaces both offer a read-only view. The
  detail page now takes `readOnly`: editors inert, no Save/Discard/Delete, no
  Generate. Creating still requires edit rights.
- Creating a tool bounced back to the list while creating a skill lands on the
  new skill. Tools now do the same. The upsert returns the workspace's whole
  tool list (newest first) rather than just the new row, so the id is matched
  by title instead of by index — the same trap that produced the skill-create
  navigation bug.
- Remove `openPrompt`/`closePrompt` from useWand. `closePrompt`'s last callers
  went away with the custom-tool-modal extraction and `openPrompt` had none
  before it; nothing reads `isPromptVisible` any more either.

* fix(tools): read-only editors, design-system wrench, skills-matching tool identity

- readOnly never reached the editors: the prop gated actions and Generate but
  the schema and code fields were still typable for viewers without edit rights.
  Wire disabled through both fields into CodeEditor.
- The row icon used lucide's Wrench (strokeWidth 2) where @sim/emcn/icons ships
  one drawn for this system (1.55, tuned viewBox), and it inherited body text
  colour instead of --text-icon. Swap it.
- Give the tool detail page the same identity heading as skill detail: tile,
  name, and description at the top left, instead of only a header title.
- Extract ResourceTile so the skills and tools tiles share one definition
  (SkillTile now composes it), and add an opt-in `iconFilled` to
  SettingsResourceRow so the tools list tile matches the skills gallery. Both
  default to today's behaviour for every existing consumer.

* fix(mentions): use the product's own glyph for every @ mention kind

The `@` menu and the inserted chip mapped kinds to arbitrary lucide icons —
`Sparkles` for a skill, a generic `File` for every file — while the rest of the
product has a settled glyph per resource. Mirror CHAT_CONTEXT_KIND_REGISTRY,
which Chat's `@` menu already renders from:

- skill now uses AgentSkillsIcon, the same glyph SkillTile shows everywhere
- workflow / folder / table / knowledge use the @sim/emcn/icons set the sidebar
  and the chat registry use
- file derives its icon from the filename extension, so a .pdf and a .csv are
  distinguishable, matching the file list and Chat's context chips
- integration keeps the block's brand icon from the registry

Also drop the generic placeholder. `kind` is untrusted — the node schema
defaults it to `''` and a hand-written `sim:` link can carry anything — but an
unrecognized kind now yields no icon instead of a meaningless box, which is what
the chat registry does. The menu already guarded a missing icon; the chip now
does too, so this cannot crash on a malformed link.

* chore(db): drop skill_member migration 0266 for regeneration on latest staging

* feat(db): regenerate skill_member migration as 0267 after staging merge

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Waleed Latif <walif6@gmail.com>
…face (simstudioai#5916)

* refactor(skills): one definition of the skill fields across every surface

The Name / Description / Content trio was implemented three times — once in the
canvas modal and once each in the create and detail pages — and had already
drifted: the name hint appeared on create, was missing entirely on detail, and
sat in a different slot in the modal; the content editor was 260px on the pages
and 200px in the modal; only the modal marked the fields required.

Extract SkillFields, the DetailSection form both full-page surfaces render, and
move the shared copy into skill-copy.ts. The modal keeps ChipModalField — that
is required inside a ChipModalBody, so it cannot share the pages' JSX — but now
reads the same placeholder, hint, and max-length constants, so the wording can
no longer diverge.

The detail page's local FieldLockTooltip moves into SkillFields and is now
available to both pages, and the dynamic rich-editor import drops from three
copies to two.

No behavioral change beyond the drift being resolved: the name hint now shows on
the detail page as it already did on create.

* refactor(skills): derive modal saving state, collapse the field message line

From the cleanup passes over this branch:

- SkillModal mirrored `mutation.isPending` into a local `saving` useState, which
  .claude/rules/sim-hooks.md forbids. It also carried a latent bug: the
  `finally { setSaving(false) }` ran after `onSave()` had closed the modal, so it
  set state on an unmounting component, and a throw from `onSave()` would leave
  the modal stuck in a saving state. Derived from the two mutations instead.
- The hint/error line under a field was written three times in SkillFields, in
  three slightly different shapes. One FieldMessage helper now renders all three.
- The name hint is an authoring instruction, so it no longer shows under a field
  that cannot be edited (a built-in skill, or a viewer who is not an editor).
…y chat answers) (simstudioai#5915)

* fix(providers): stop the final regenerated stream from re-calling tools and clobbering the answer

After the silent tool loop settles, OpenAI Responses and Gemini re-issue a
streaming request purely to stream the answer as prose — but with tools still
attached and auto tool choice, a reasoning model can re-decide to call a tool
there. Streamed calls are never executed on this path, so the run ends with a
dead function call, an empty streamed answer, and the stream callback
clobbering the tool loop's settled text with '' (deployed chat rendered
{"content": ""}).

Force tool_choice 'none' / functionCallingConfig NONE on the regeneration and
keep the tool loop's settled answer whenever the stream ends without text.

* feat(openai): settled tool chips on the regenerated answer stream

The silent Responses tool loop has no live stream while tools run, so opted-in
consumers saw no tool chips at all for OpenAI. The loop now records each
executed call and prepends settled tool_call_start/end pairs (name + status
only) to the agent-events stream ahead of the regenerated answer. Runs without
a sink never see these events, so legacy output is unchanged.

* fix(providers): apply the regeneration guard fleet-wide

Audit of every silent tool loop for the same race fixed for OpenAI/Gemini
(final regenerated stream re-calls a tool that is never executed, ending with
an empty answer that clobbers the settled text):

- anthropic (both implementations): tool_choice {type:'none'} on the
  regeneration (tools must stay — history carries tool_use blocks) + keep the
  settled answer when the stream ends without text
- groq: was re-applying the ORIGINAL tool_choice, so forced-tool runs
  re-forced the tool on the regeneration — guaranteed dead call; now 'none'
  + fallback
- deepseek, mistral, cerebras, azure-openai (legacy chat path), openrouter,
  xai: 'auto' -> 'none' + fallback
- bedrock: fallback only — Bedrock's ToolChoice has no 'none' and toolConfig
  is required when history carries toolUse blocks

Already guarded (no change): meta, sakana, nvidia, vllm, litellm, baseten,
together, fireworks, kimi, zai, ollama.
Co-authored-by: Sim Pi Agent <pi@sim.ai>
…wedged saves (simstudioai#5905)

* fix(data-retention): clamp sub-day retention values so saves aren't wedged

A stored value under 12 hours rounded to '0' days on load and was re-sent as 0, which the contract rejects (min 24) — blocking every save on the page, including unrelated fields.

Clamp hours->days into the contract's range on read, and throw instead of emitting 0/NaN on write.

* improvement(docs): rewrite data retention for workspace overrides + PII redaction

- Document PII redaction (Logs / Workflow input / Block outputs stages, entity types, languages, custom regex patterns)
- Replace the stale 'no per-workspace overrides' section with the retention-policies list and override inheritance
- Correct log retention (also covers background job logs) and soft deletion (adds Chat conversations, KB documents)
- Add PII + override screenshots, refresh the main one
…udioai#5909)

* feat(sso): DNS domain verification gating org SSO registration

Add org-scoped domain ownership verification (DNS TXT challenge) as the
security precondition for configuring SSO. Closes the first-come domain-claim
vuln where any org could wire another company's domain to its own IdP.

- New sso_domain table + migration 0266; existing org SSO domains are
  grandfathered as verified so live tenants are unaffected
- Verified-domains settings UI (enterprise-gated) with add/verify/remove
- Register route now requires a verified domain for org-scoped registration;
  personal SSO and already-grandfathered domains are unaffected
- Self-host register script writes the verified sso_domain row directly, so
  script-driven registration stays backwards compatible

* fix(sso): harden domain verification against concurrency + fix CI lint

Addresses review findings on state invariants under concurrent/failed writes:
- Add unique index on (organization_id, domain) so concurrent claims can't
  create duplicate pending rows; POST re-reads and stays idempotent on conflict
- Verify flips the row only if it's still the exact pending challenge checked
  (guards deletion/token-rotation mid-DNS-lookup) and maps the partial unique
  index violation to 409 instead of an unhandled 500
- Wrap the self-host script's provider write + verified-domain upsert in a
  transaction so a failed ownership write can't leave a provider committed
- Format 0266 snapshot/journal with biome (fixes @sim/db lint:check)

* fix(sso): re-check domain verification before provider write (TOCTOU)

The register gate checked the verified sso_domain row only at handler entry,
then ran OIDC discovery before writing the provider. A verified row removed
during that window could still complete registration. Extract the check into a
closure and call it both as an entry fast-fail and authoritatively right before
registerSSOProvider, alongside the existing domain-conflict re-check.

* fix(sso): stop rotating verification token on idempotent re-add

Re-adding a pending domain rotated its verification token, which invalidated a
TXT record the admin may have already published and — under two concurrent
re-adds — could return a token the racing write had already superseded, so the
admin's DNS record would never verify. Return the existing row unchanged
instead; the pending token is always shown in the UI, so it is never lost.

* fix(sso): close register TOCTOU with compensating delete + harden edges

Audit-driven hardening:
- Close the residual register TOCTOU: registerSSOProvider is create-only (throws
  if the providerId exists), so a compensating delete after the write is provably
  safe — it can only remove the just-created row. If verification was revoked
  during the write, roll the provider back and 403.
- Verify is now idempotent under concurrency: a same-org row already flipped to
  verified by a racing request returns 200, not a confusing 409.
- Grandfather backfill + self-host script now match normalizeSSODomain's dominant
  transforms (lower + trim + strip leading wildcard) so a non-canonical legacy
  domain can't miss the runtime gate's lookup. Prod backfill result is unchanged.
- Cleanup: drop dead default export, align card radius to sibling convention.

* fix(sso): redact domain tokens from non-admins + fix script stale-update

Round-5 review findings:
- GET /domains redacted the pending TXT verification token (a management
  secret) to any org member. Now only owner/admins read it; members see the
  list and status without it. Non-Enterprise orgs get an empty list (entitlement
  flag only), never the domains/tokens.
- Self-host script decided update-vs-insert from a read taken OUTSIDE the
  transaction; a provider deleted mid-flight made the UPDATE match zero rows
  silently while the verified-domain upsert still committed (orphaned domain).
  The decision now happens inside the transaction from the UPDATE's row count.

* docs(sso): drop unshipped enforce-SSO / auto-join copy from verified domains

Verified domains currently only gate SSO configuration. Remove the
forward-looking references to enforcing SSO and auto-joining members (deferred
to a later release) from the docs, settings copy, nav description, and schema
comment so we don't promise unshipped features.

* fix(sso): guard rollback to new providers only + Enterprise-gate domain removal

Round-6 review findings:
- The compensating provider rollback now only fires when the provider did not
  exist before this request (providerExistedBefore). registerSSOProvider is
  create-only today so reaching the rollback already implies a fresh create, but
  this makes the safety local and future-proof: if Better Auth ever allowed
  updating an existing provider, a revoked-verification rollback must not delete
  that pre-existing row.
- DELETE /domains now requires an Enterprise plan like add/list/verify, so all
  domain mutations share one entitlement (the UI already hides removal from
  non-Enterprise orgs). Adds a delete-route test.

* fix(sso): roll back the SSO provider by row id, not logical keys

The compensating rollback deleted by (providerId, orgId). providerId is unique,
so if this request's row were deleted and recreated by a concurrent registration
in the narrow window before the rollback, the logical-key delete would remove
that other request's provider. Delete by the primary-key id registerSSOProvider
returns instead, so only the exact row this request created is ever removed.

* chore(sso): final-review polish — trim script read, unify copy, doc migration edge

Cosmetic cleanup from a final 4-track adversarial review (no bugs found in the
new logic):
- Self-host script: narrow the pre-transaction existence read to select({ id })
  instead of SELECT * (it only feeds a log line now).
- Unify invalid-domain copy ("for example acme.com") and the verified-elsewhere
  409 wording ("is already verified by another organization") across routes.
- p-3 shorthand on the domain row card.
- Document the migration's rare two-orgs-share-a-domain grandfather behavior
  (login unaffected; validated no such duplicates in prod).

* fix(sso): apply attribute mapping + make SSO edit work; drop dead guard

Two pre-existing SSO bugs the final review surfaced (prod has one SSO org, RVW,
script-registered with the default mapping, so neither change affects it):

- Attribute mapping was passed at the top level of the register payload, which
  Better Auth ignores — it reads oidcConfig.mapping / samlConfig.mapping. Nest it
  so custom mappings actually apply. (Default mapping is unchanged, so existing
  logins are unaffected.)
- Editing an SSO provider was broken: registerSSOProvider is create-only and
  threw on the existing providerId → generic 500. Route now detects a provider
  the caller already owns and updates it via Better Auth's updateSSOProvider, and
  surfaces Better Auth's own error status/message instead of a blanket 500.

Also drops the now-unnecessary providerExistedBefore guard (the rollback deletes
by the created row's primary-key id and register is create-only) and the earlier
final-review polish (script read, unified copy, migration edge note).

Smoke-test SSO login + edit on staging before merge (auth-path change).

* fix(sso): require null org on personal-mode provider lookups (gate bypass)

The personal branch of both provider-ownership lookups keyed on
(providerId, userId) without requiring organizationId IS NULL. Because org
providers store userId = their creator and providerId is globally unique, an org
admin could send a personal-mode request (no orgId) — which skips the membership
check and the domain-verification gate — yet still match, and then via the new
update path move, their org's provider to an unverified domain. Add
isNull(organizationId) to the personal branch of both clauses so it can only
match a genuinely personal provider, matching the route's own isOwnedByCaller.

Found by an adversarial review of the update path added in 394bda9.

* fix(sso): script updates the observed provider by id, not providerId

Inside the registration transaction the script updated WHERE providerId — the
logical key. If the observed provider was deregistered and a replacement created
with the same providerId before the transaction ran, that update would clobber
the replacement's config and ownership. Update the specific observed row by its
primary-key id instead; if it's gone we insert, which fails cleanly on the
providerId unique constraint rather than overwriting the replacement.

* fix(sso): script upserts provider via delete-then-insert (no unique constraint)

sso_provider.provider_id is a plain (non-unique) index and prod holds legitimate
duplicates, so the previous "update by id, else insert" could create a duplicate
provider when the observed row was deregistered and replaced before the
transaction — the fallback insert would succeed. Delete every row for the
providerId then insert exactly one, inside the transaction, so the providerId
ends up as exactly this config atomically. Linked accounts key on the providerId
string (not the row id), so existing logins are unaffected.

* fix(sso): guard compensating-delete row id so rollback can't silently no-op

* chore(sso): regenerate migration as 0268 after merging staging

Staging landed migrations 0266/0267, colliding with our 0266. Removed our
migration, merged staging, and regenerated cleanly with drizzle-kit as
0268_sso_domain_verification (identical sso_domain table + indexes), then
re-appended the grandfather backfill. api-validation baseline reconciled to 973
(staging 970 + our 3 domain routes). Also make the register-route test's
registerSSOProvider mock return an id so the guarded compensating delete runs.

* refactor(sso): share normalizeSSODomain via @sim/utils so script matches gate

The self-host script canonicalized SSO domains with a minimal inline transform
(lower+trim+wildcard) that diverged from the app's full normalizeSSODomain
(protocol, port, path, trailing dot, email local part) — equivalent spellings
could store a different ownership key than the runtime gate looks up. Move
normalizeSSODomain into @sim/utils/sso-domain (a pure function) so the register
route, the domain-claim route, and the script all use the identical canonicalizer.
The script now skips the verified-domain record when SSO_DOMAIN isn't a valid
registrable domain instead of storing a malformed key.
* feat(providers): add Claude Opus 5 model

* fix(providers): route Opus 5 through adaptive thinking

Opus 5 supports only adaptive thinking (no extended thinking / budget_tokens),
same as Opus 4.8/4.7. Add opus-5 to supportsAdaptiveThinking so thinking
requests use thinking.type: adaptive instead of falling through to budget_tokens
extended thinking, which Opus 4.7+ rejects with a 400.

* docs(agent): regenerate agent-stream docs for Opus 5
…ides (simstudioai#5928)

* content(library): add observability, procurement, and MCP security guides

Three AEO guides, each dated into a recent empty day in the posting
calendar (2026-07-19, 07-21, 07-22) so publication is spread across days
rather than dumped on one.

- ai-agent-observability: what observability is, why APM falls short for
  non-deterministic agents, what to instrument per lifecycle stage, and the
  signals to track
- ai-agents-in-procurement: what procurement agents do, buy-vs-build, where
  they add value, and how to start narrow
- mcp-security: tool poisoning, confused-deputy/OAuth flaws, and supply-chain
  risk, plus how to build and govern MCP servers securely

Every third-party factual claim carries an outbound citation to a verified
primary source (OpenTelemetry semconv, Fiddler, LangChain and PwC surveys,
Icertis/ProcureCon, Ironclad, GEP, the MCPoison/CurXecute CVEs on NVD,
Anthropic's MCP announcement, RFC 8707/9700, OAuth 2.1, the MCP auth spec),
and each post carries 3 internal links to related library posts. One claim
from the source copy — a "November 2025" date on the WhatsApp MCP exfil case
— could not be verified and was dropped; the case itself is cited to Docker's
writeup. Covers come from the autogeneration pipeline via the standard
/library/<slug>/cover.jpg path.

* fix(library): correct two unverifiable claims in the new guides

Accuracy pass on the three new posts turned up two claims that could not be
substantiated:

- HIPAA compliance: the procurement post claimed Sim has "SOC2 and HIPAA
  compliance," but the canonical compliance data (lib/compare/data/sim.ts)
  states SOC2 only, and explicitly that Sim offers self-hosting "beyond SOC2,
  rather than additional certifications." Removed HIPAA; kept SOC2 plus
  self-hosting for data residency. (The same claim exists in ~5 pre-existing
  library posts and should be corrected separately.)
- "more than 18,000 servers were listed on MCP Market": no source
  substantiates this figure. Replaced with "thousands of community-built
  servers," which the ecosystem supports, keeping the Anthropic and MCP Market
  links.

Every other third-party claim was verified against a primary source (both
CVEs on NVD, RFC 8707 title, RFC 9700 as January 2025, and all six survey
statistics).
…uite (simstudioai#5927)

* improvement(ship): pre-flight regenerate artifacts + parallel audit suite

Add a two-phase pre-ship step: (A) regenerate every committed artifact
(agent-stream-docs, skills, contract syncs) in parallel so generated files
never drift into a CI failure, then (B) run lint plus the full audit suite
CI's Lint and Test job enforces, in parallel, aborting ship on any failure.
Regenerate the ship command projections.

* fix(ship): scope Phase A to in-repo generators + propagate failures

Cursor review: (1) mship:generate is an umbrella over the 9 contract
generators and biome-formats the shared generated dir, so running it
alongside its constituents races/corrupts; it also reads an external
copilot-contract source that ENOENTs in a bare worktree. (2) bare wait
swallowed generator exit codes. Narrow Phase A to the always-in-repo
generators (agent-stream-docs, skills) and collect each job's exit
status in both phases; document domain generators as on-demand only.

* fix(ship): make pre-flight status checks exit non-zero on failure

Cursor/Greptile: grep && echo ❌ || echo ✅ always exits 0, so a failed
generator/audit didn't actually gate ship (regressed the sequential
bun-run checks whose non-zero exit agents rely on). Replace both Phase A
and Phase B aggregations with an if-grep that exits 1 on any failure.

* fix(ship): gate lint exit in Phase B pre-flight

Cursor: bare bun run lint before the audit grep meant a non-zero lint
was ignored — if audits then passed, the block exited 0 and ship
continued. Gate lint with || { echo …; exit 1; } so an unfixable lint
error aborts before the audits run.
…end gotcha (simstudioai#5931)

* fix(sso): surface DNS verification failures and the provider auto-append gotcha

Post-merge audit follow-ups for verified domains (simstudioai#5909):

- The host field handed admins the FQDN `_sim-challenge.acme.com`. GoDaddy,
  Namecheap, Hover and most cPanel panels append the zone to whatever is typed,
  yielding `_sim-challenge.acme.com.acme.com` — the record looks right in their
  panel but never verifies, and our 422 tells them to wait 48 hours. Add a hint
  under the field (and a docs callout) telling those admins to enter just the
  label.
- DNS failures were logged at debug, but production log level is ERROR, so a
  blocked-egress or SERVFAIL condition was invisible to us and misreported to the
  admin as "record not found yet". Log infrastructure-class failures at warn with
  the DNS error code; keep the genuinely-absent codes at debug.
- Trim the joined TXT value before comparing: several DNS panels pad the stored
  string, which otherwise fails an exact match forever.
- Lower the resolver to 2s/1 try. c-ares multiplies timeout across servers and
  retries by ~7x, so the previous 5s/2-try config could block a verify request
  for ~35s when resolvers are unreachable.
- Cover `checkDomainTxtRecord` — the function that decides whether the gate opens
  had no tests. Adds exact match, chunk-joined value, match among unrelated
  records, padded value, near-miss, another org's token, absent record,
  infrastructure failure, and empty-response cases.

* fix(sso): log DNS infrastructure failures at error so prod actually surfaces them

Production's default minimum log level is ERROR, so the warn introduced in the
previous commit was still filtered out — the fault stayed invisible exactly as
before. A resolver failure that is not 'record absent' is a genuine
infrastructure error, so ERROR is both the visible and the honest severity.

* fix(sso): state the zone-removal rule instead of a wrong subdomain hint

The hint computed the bare label as the first segment of the challenge host, so
for a subdomain like eng.acme.com it advised entering `_sim-challenge` when the
host relative to the acme.com zone is `_sim-challenge.eng` — following it would
publish the record on the wrong name and verification would never succeed, the
exact failure the hint exists to prevent. Deriving the real zone needs the Public
Suffix List (acme.co.uk defeats naive label-stripping), so state the rule instead:
enter the host with the trailing zone removed. Docs show both the apex and the
subdomain form.
…imstudioai#5935)

.npmrc set ignore-scripts=true globally, which overrode Bun's
trustedDependencies allowlist and blocked every lifecycle script —
including the repo's own root scripts. isolated-vm therefore shipped
unbuilt and each contributor had to npm rebuild it by hand.

The .npmrc line landed in the May 2025 Bun migration, seven months before
isolated-vm was introduced. Two later commits added isolated-vm to
trustedDependencies (root, then apps/sim) and both were silently inert.

- delete .npmrc; lifecycle policy now lives solely in root trustedDependencies
- add isolated-vm to root trustedDependencies
- drop the workspace-level trustedDependencies from apps/sim (Bun only reads
  the root list, so it never had any effect)
- drop the .npmrc entry from CODEOWNERS

Naming any package in trustedDependencies replaces Bun's curated
default-trust list, so only isolated-vm and sharp may run install scripts.
Docker is unaffected — all three Dockerfiles pass --ignore-scripts
explicitly and rebuild isolated-vm against the Node ABI by hand.
…mstudioai#5936)

Reconcile compliance and traction wording across three library guides
against the canonical data in lib/compare/data/sim.ts, and tighten a couple
of related phrasings for consistency.
…O v1 default (simstudioai#5939)

* improvement(helm): hygiene pass — CI gating, strict values schema, ESO v1 default, ci values

Closes the gaps from a best-practices audit of the chart (template-level
conformance was already clean: full label set, 82 unit tests, kubeconform-
valid renders):

- new Helm Chart workflow gates every chart change: helm lint, the 82
  helm-unittest cases, kubeconform validation of default + all-components
  renders (k8s 1.29 strict, CRD catalog), and a render of all 10 example
  values files
- helm/sim/ci/ values files (chart-testing convention) so the chart lints
  and templates cleanly out of the box with dummy secrets
- values.schema.json declares all 30 top-level keys (16 were invisible) and
  sets root additionalProperties: false, so top-level typos fail fast
- externalSecrets.apiVersion defaults to v1: current ESO releases removed
  the v1beta1 compatibility path in 2026, so the old default produced
  rejected manifests on new installs; NOTES/values comments updated
- wait-for-postgres init container gets requests/limits (the only container
  in the chart without them; broke ResourceQuota'd namespaces)
- drop the telemetry Prometheus scrape config for app/realtime — neither
  exposes /metrics, so it was dead config that also rendered a realtime
  target with realtime disabled
- chart 1.2.0 with README upgrade notes

* fix(helm): version-agnostic schema-error assertions in the secret-length suite

Helm v4 phrases schema rejections as 'minLength: got N, want 32' while
v3 says 'String length must be greater than or equal to 32'. The suite
grepped for 'minLength' only, so it passed on local helm v4 and failed on
CI's v3.16.4 — the enforcement itself works on both. Patterns now assert
the key name plus either wording. Caught by the new Helm Chart workflow
on its very first run.

* improvement(helm): kind install test + chart version-bump gate in CI

Benchmarked against the flagship OSS charts (ingress-nginx, argo-cd,
kube-prometheus-stack, grafana, bitnami, cert-manager): sim already exceeds
most of them on validation rigor (strict schema — 4 of 6 ship none; 82 unit
tests vs argo-cd's zero; kubeconform manifest validation none of them run),
but every top community chart repo actually installs the chart on a kind
cluster in chart CI — the one majority practice we lacked. Adds:

- install job: kind cluster, helm install with ci/default + a new
  small-footprint ci/kind-values.yaml overlay (default app requests of 4Gi
  can't schedule on a CI node), --wait, then the chart's helm test hook,
  with pod/event/log diagnostics on failure
- version-bump job (PR-only): fails when helm/sim/** changes without a
  Chart.yaml version increment — argo/kps/grafana all enforce this

* fix(helm): declare naming overrides in the strict schema; SHA-pin CI actions

- nameOverride/fullnameOverride are consumed by sim.name/sim.fullname but
  were never in values.yaml, so root additionalProperties: false rejected
  Helm's standard naming overrides — both now declared (a helper-wide sweep
  confirmed they were the only template-read keys missing), with a smoke
  regression test that installs under the strict schema and asserts the
  override lands in resource names
- CI supply-chain hardening: checkout/setup-helm/kind-action pinned to full
  commit SHAs (tag comments retained) and the kubeconform archive verified
  against its published sha256

* fix(helm): immutable unittest runner and fail-closed version gate in CI

- helm-unittest runs via the project's official docker image pinned by
  immutable sha256 digest instead of a plugin install from a mutable git tag
- the version-bump gate fetches the full base ref (a --depth=1 fetch could
  leave no merge base), computes the merge-base and diff outside the if so
  any git failure fails the job instead of falling into the skip branch

* fix(ci): run the helm-unittest container as the runner UID

The digest-pinned image runs as a non-root user that cannot write into the
runner-owned bind mount (it creates tests/__snapshot__, absent from the
checkout since empty dirs aren't tracked). Standard bind-mount pattern:
--user "$(id -u):$(id -g)" with HOME=/tmp for helm's cache.

* fix(helm): appVersion points at a real GHCR tag (v0.7.44)

The kind install job caught this on its first full run: the default image
tag (Chart.AppVersion 0.6.73) returns 404 on GHCR for all three images —
the registry's tags are v-prefixed — so an unpinned default install could
never pull. Updated appVersion to v0.7.44 (verified 200 for simstudio,
realtime, and migrations manifests), stale values comment refreshed, and
an upgrade note added. The CI kind values deliberately stay tag-free so
the job keeps exercising the true default path.
simstudioai#5922)

* improvement(providers): validation pass, and stream tool loop improvements

* remove deploy options correctly

* fix

* feat(providers): prompt caching capability and usage-based cache pricing

Replace the arbitrary cached-rate heuristic with a single cache-aware pricing
function, and add prompt caching as an opt-in capability for Anthropic.

Pricing: priceModelUsage in cost-policy.ts is now the only place cache
arithmetic happens. Provider adapters normalize their wire shape into
ModelUsage (input always excludes cache buckets); the pricing function never
branches on provider. This removes five divergent behaviors, including the
!!request.context heuristic that gave Router and Evaluator an unearned 10x
input discount, and the overwrite that silently billed Anthropic cache reads
and writes at zero. Also parses OpenAI cache_write_tokens, previously ignored.

Caching: Anthropic gets a capability-gated advanced switch that places
cache_control on the last tool and last system block; system is now always a
TextBlockParam array. OpenAI gets a stable per-block prompt_cache_key with no
UI, since its caching is automatic.

* fix(providers): route OpenAI and Gemini block cost through cache-aware pricing

Cache-aware pricing only reached trace segments. The billable block cost still
called calculateCost on the cache-inclusive prompt total, so OpenAI cache hits
and Gemini implicit-cache hits were charged at the full input rate and GPT-5.6+
cache writes went unbilled.

Both providers now accumulate cache buckets and price through priceModelUsage,
matching the Anthropic token convention where input excludes cache reads and
writes. Cached counts are clamped to the prompt total so an over-reporting
payload cannot bill more input than the request contained.

* fix(streaming): redact tool payloads on selected outputs in public chat

Redaction only ran on the empty-selection branch, but a deployment almost
always selects outputs, so it was dead in the case it exists for. Selecting
toolCalls streamed the raw arguments and results to a public chat client in a
chunk frame, and providerTiming carried thinking content the same way.

Both paths now extract from the sanitized block output rather than the raw log:
the streamed selected output, which is the reachable vector, and the final
envelope. Sanitizing the source rather than per selected path means a newly
selectable field cannot reopen the hole.

* refactor(providers): drop unreachable billing fallbacks

Every provider pricing helper took a policy parameter no caller passed. Worse
than dead: passing one would have double-applied the margin the central layer
already applies. Removed, so providers can only price at list.

Also removed guards that cannot fire. The central fallback normalized cache
buckets no provider can reach it with (all three that report cache usage price
themselves) and did so at a 1x write multiplier no vendor charges.
priceModelUsage re-validated token counts the adapter had already clamped, and
applyModelCostPolicy defaulted a required total field.

Validation now happens once, in the adapter that parses the vendor payload and
is the only layer that knows cache buckets are a subset of the prompt total.
…nner sizing (simstudioai#5945)

* fix(ci): key node_modules sticky disk on the lockfile hash

* improvement(ci): per-image Blacksmith runner sizing + cold-build memory preflight

* fix(deps): exclude @next/swc binaries from the release-age gate

* fix(deps): pin @next/swc binaries so frozen installs get a compiler

* docs(ci): explain ARM runner sizing rationale
…inputs/outputs (simstudioai#5942)

* improvement(whatsapp): validate + improve integration skill for file inputs/outputs

* fix lint

* add whatsapp subblock migration
simstudioai#5671 rewrote the shimmer-sweep keyframes to run 100% -> -100%, which is
already left-to-right, but left the `reverse` from simstudioai#5650 on the animation
shorthand. The two flips cancel into a right-to-left sweep on every
ShimmerText consumer: subagent labels, tool-call rows, and the new
agent-stream thinking chrome.

Drop `reverse` so direction lives in exactly one place — the keyframes.
…imstudioai#5917)

* fix(realtime): evict revoked collaborators from live workflow rooms via periodic read-access re-validation

* fix

* fix(realtime): close join/eviction race and make sweep cleanups independent

Re-authorize immediately before socket.join so an in-flight join cannot
reverse a sweep eviction, and run the sweep's best-effort cleanups
independently so a room-state failure cannot skip the presence broadcast.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): single-flight role resolution and retry failed eviction cleanup

Coalesce concurrent role resolutions per (user, workflow) so a slow stale
read can never overwrite a recorded revocation, and defer failed eviction
room-state cleanups into a per-sweep retry queue so collaborators are not
left with a stale presence entry.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): scope room removal to the target workflow and detect swallowed cleanup failures

Honor the workflowIdHint as the target room in both room managers so
removing a stale room cannot clobber the mapping of a room the socket has
since moved to, and confirm eviction cleanup via the returned workflowId
plus an unswallowed mapping read so Redis failures actually defer into the
retry queue.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): treat unconfirmed removals as failures and refresh role cache on fresh verify

Treat any null removal result as a failed cleanup (the sweep always passes
the target room, so null only means failure — including with expired
mapping keys), move the same-room rejoin guard to a synchronous check
immediately before the removal, and record verifyWorkflowAccess's fresh
decision into the role cache so a re-granted user is not blocked by a
stale cached revocation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): prefer a mid-flight recorded role decision over the in-flight query result

If a fresh authoritative read (join-time verify) records a decision while
a single-flighted resolution's query is in flight, keep the recorded
decision instead of overwriting it with the potentially stale result.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): isolate the security scan from the Redis cleanup lane

Run the revocation scan (local sockets + DB only) and the best-effort
room-state cleanup as independently-guarded lanes so a hanging Redis
command can stall only presence cleanup, never revocation enforcement.
Evictions now enqueue cleanup instead of awaiting it, and the scan no
longer reads presence for a fallback role.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): bound authorization waits in the revocation scan

Race each socket's authorization check against a per-socket timeout and
cap the whole pass with a budget below the sweep interval, so a hanging
DB query skips that socket for the pass (never evicting on uncertainty)
instead of wedging the scan lane and starving subsequent ticks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): round-robin the revocation scan so hung checks cannot starve later sockets

Resume each scan pass after the last target the previous pass processed,
so a fixed prefix of hanging authorization checks can never repeatedly
consume the pass budget and leave sockets behind it unexamined.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
… agent thinking/tool-calls/prompt-caching, opus 5, daytona failover
… state to nuqs, design-system cleanup (simstudioai#5950)

* refactor(settings): fold verified domains into SSO and use shared primitives

Verified domains only gates SSO, so managing it on a separate page meant
discovering the requirement after filling out the whole IdP form and then
navigating away mid-setup. Move it into the SSO page as a section above the
provider config and drop the standalone page, its nav entry, and both route
branches.

Align the surfaces with the shared settings primitives rather than bespoke
chrome, matching whitelabeling/custom-blocks/access-control:
- SSO's local FormField (muted labels) is replaced by the shared SettingRow, so
  its fields read like every other settings page. SettingRow gains optional
  `optional` and `error` props to absorb what FormField did — additive, so
  existing consumers are untouched.
- The domains section is built from SettingsSection, SettingRow,
  SettingsResourceRow, and SettingsEmptyState instead of hand-rolled cards.

Also drop the redundant Upload/Change buttons in whitelabeling: the logo and
wordmark thumbnails were already clickable, so the button was a second control
for the same action. Remove still appears once an image is set.

* fix(settings): move group-detail view state to nuqs and clean up design-system drift

Access control's group detail kept its tab, three search boxes, and three status
filters in useState, so a `?group-id=` link always landed on General and a filter
was lost on reload — the parent already puts the group id in the URL. The three
tabs never render together, so search and status share one param each rather than
carrying three mutually-exclusive keys, and switching tabs resets both. Closing
the detail clears all three alongside group-id in one batched write, so nothing
lingers on the list URL.

Design-system fixes from a cleanup pass over the surfaces this branch touched:
- Restore accessible names lost when the whitelabeling Upload buttons were
  removed. The thumbnail is now the only click target, and it contained just an
  icon, so it announced as an unlabeled button; the icon-only Remove had the same
  problem. Both now carry aria-labels reflecting their state.
- Use Chip, not the legacy Button, for the domain actions — Button is ~26px
  against the 30px ChipInput beside it, so "Add domain" sat visibly short.
- SettingRow now uses the emcn Info component; a bare svg as Tooltip.Trigger was
  neither focusable nor nameable. It also stops re-specifying Label's own default
  styling.
- Use the new SettingRow error prop for the group name instead of a hand-rolled
  error paragraph, which is what the prop was added for.
- Hoist the block-category lookup out of a sort comparator, size-* over h/w, name
  the staleTime constants the rules require, and import RowActionsMenu from its
  barrel.

* fix(settings): alias the old /settings/domains path to SSO

Folding verified domains into the SSO page dropped /settings/domains, so
bookmarks and shared links 404'd instead of landing where domains now live. Both
alias maps already exist for exactly this (organization/'members',
subscription/'billing'); add domains -> sso to each.

* chore(settings): adopt ChipCopyInput, named staleTime constants, and a11y labels

* fix(settings): reset group detail params on open and drop issuer mono styling
waleedlatif1 and others added 12 commits August 3, 2026 17:23
…imstudioai#6231)

* fix(execution): offload buffered event values under budget pressure

An execution buffers EVENT_LIMIT events inside a per-execution byte budget, so
a full ring only fits if events average under budget/EVENT_LIMIT. Values were
only offloaded to object storage at the shared 8 MiB cap, far above that, so a
run emitting large block outputs exhausted its budget within a few dozen events
and stayed pinned at its ceiling for the rest of its life.

Applying that ceiling to every run would be worse than the problem: the SSE
stream carries the compacted event and the terminal renders a ref only as a
preview, so ordinary block outputs would stop being readable live, and every
value would cost an object-storage write on the hot path. Engage the tight
ceiling only once a run has actually buffered past half its budget. A short run
keeps full-fidelity output and pays nothing; a runaway one stops accumulating.

Both bounds derive from the existing budget rather than being asserted, and
preserved UserFile base64 is exempt — it is an explicit request for inline
delivery, already bounded by its own cap and the strip-and-recompact fallback.

Also stop a failed resume-path buffer write from failing the run: it was
awaited bare, so the rejection propagated into the executor callback and failed
work that had already completed. The buffer only backs reconnect replay, so
degrade to live-only delivery the way the execute route does.

* fix(execution): measure pressure at write time and keep terminal status last

Pressure was read from bytes counted once a flush succeeded, but a burst is
compacted long before the scheduled flush runs — so the very batch that
exhausts the budget went through at the loose ceiling and was dropped instead
of offloaded. Count bytes as each event is compacted.

Separately, the terminal-alone retry stamped terminal status while entries
queued ahead of it were still unwritten. Terminal status is the reader's
end-of-run signal: a reconnecting client drains what is in Redis and closes, so
those entries were stranded behind a stream it had already finished with. Drain
the backlog first, then publish the terminal event.

* fix(execution): do not lose the backlog or publish terminal status early

Draining the backlog ahead of the terminal event left the terminal status armed,
so whichever chunk emptied the queue stamped the run complete before its
terminal event was written — the inverse of the ordering the drain was added to
guarantee. Disarm the status for the drain and restore it afterwards.

The drain's result was also discarded: a transient Redis failure requeues its
batch, and the unconditional reassignment that followed dropped those events
even though the budget never rejected them. Keep whatever could not be
persisted, and publish the terminal event alone only once nothing earlier is
still queued — failing otherwise lets the caller degrade, which records the
status without claiming the missing events arrived.

Leave eventId unset on a failed resume-path write. Assigning 0 was persisted by
clients as a reconnect cursor and rewound them to the start of the run.

* fix(execution): keep an event in the buffer when a pressure offload fails

Durable compaction runs before an event is queued, so a storage or metadata
failure dropped it from replay entirely — a reconnecting client would never see
it, even though the live path carried on. Offloading under pressure is only an
optimization that keeps a heavy run from exhausting its budget, so when the
value cannot be persisted, fall back to buffering it inline: exactly what the
run would have done before pressure engaged.
…hain (simstudioai#6235)

Bumps next, @next/env and the @next/swc-* optional deps to 16.3.0 across the
root overrides, apps/sim, apps/docs and packages/emcn.

Two config notes worth keeping:

- `experimental.turbopackFileSystemCacheForBuild`'s default flipped false ->
  true for stable in 16.3.0, so our explicit `false` is now load-bearing rather
  than defensive. Without it this bump would have silently re-enabled a build
  cache measured 3.2x slower on this codebase (simstudioai#6078). Comment updated to say so.
- `experimental.useTypeScriptCli: true` is now pinned. TypeScript 7 ships no
  JavaScript compiler API until 7.1, so Next's default checker cannot run and
  needs the project-local `tsc` CLI instead. 16.2.12 was silently skipping
  build-time type checking entirely because it detected @typescript/native-preview
  and short-circuited the stage ("Finished TypeScript in 138ms"); pinning the flag
  keeps that from drifting back.

TypeScript toolchain cleanup that the upgrade makes possible:

- Drop @typescript/native-preview from apps/sim and apps/docs. It was the
  pre-release channel for TS 7 and is superseded by typescript@7 (nightlies now
  ship as typescript@next), and its presence is what suppressed build type checks.
- Align packages/browser-protocol and packages/terminal-protocol from
  typescript ^5.7.3 to ^7.0.2 so every workspace is on one compiler version.

apps/sim keeps @typescript/typescript6 as a production dependency: the function
sandbox at app/api/function/execute dynamically imports the TS 6 compiler API to
transpile user code, and TS 7 has no API to replace it yet.

Build and dev were benchmarked 3x per version on a byte-identical tree; the
upgrade is performance-neutral (build median 100s -> 99s, dev:full warm 10s -> 9s).
…ath (simstudioai#6232)

* fix(providers): route 6-10MB attachments to the provider large-file path

The inline base64 cap was 10 MB of raw bytes, but the execution payload store
refuses a single value above 8 MiB and base64 inflates by 4/3. Every raw file
over 6 MiB therefore produced a base64 string the store rejected — and the
rejection came from the base64 *cache* write, which threw and failed the run
with "Execution memory limit exceeded" even though the bytes had already been
read successfully. Because shouldUseLargeFilePath only fires above the inline
cap, 6-10 MB attachments had no path at all on any provider: they never reached
the OpenAI or Gemini Files API upload they were supposed to take.

Derive the cap from the payload-store ceiling instead of hardcoding it, and
degrade a refused cache write to "not cached" rather than failing a request
whose bytes are in hand. Every other size guard in the chain compares raw bytes
against maxBytes; only the Redis write sees the encoded size, which is why this
went unnoticed — and why it failed only where Redis is configured.

Also correct the provider ceilings against the vendors' current documentation:

- openai: 50 MiB -> 50,000,000. The gate is `size > maxBytes`, so 50 MiB
  admitted 52,428,800 bytes; the docs say each file must be *under* 50 MB.
- bedrock: had no entry and inherited the inline cap, which is above what
  Converse accepts (3.75 MB per image, 4.5 MB per document).
- groq: 20 MiB -> 20,000,000, and modelled as the request cap the docs
  actually describe rather than a per-file MiB ceiling.
- fireworks: had no entry; its 10 MB budget is on the base64 total, so the
  raw-byte equivalent is 7.5 MB.

Add perRequestMaxBytes for the combined ceilings, enforced before any upload
spend, and cover the OpenAI upload path end to end — it had no test at all.

* chore(deps): upgrade @google/genai to 2.13.0 and @anthropic-ai/sdk to 0.115.0

@google/genai 2.x reworks the Interactions API, which the Gemini deep-research
provider is built on. Migrate it:

- `Interaction.outputs` (a flat content array) is now `steps`, a discriminated
  timeline; the report text lives in the `model_output` steps' text content,
  alongside thought and tool steps we skip.
- `Usage.total_reasoning_tokens` is now `total_thought_tokens`. The old code
  already fell back to that name through a cast, so this just makes the field
  the SDK actually returns the typed one.
- SSE events renamed: `content.delta` -> `step.delta`, `interaction.start` ->
  `interaction.created`, `interaction.complete` -> `interaction.completed`.
  The new event types are discriminated, so the payload casts are gone.

Both `interactions.create` calls also stop annotating their params with
`Interactions.CreateAgentInteractionParams{,Non}Streaming`. In 2.13.0 those
namespace aliases resolve to `CreateAgentInteraction`, whose `stream` is a
plain `boolean` rather than a literal — annotating with them erases the
discriminant and the call resolves to the union-returning overload, so the
result is typed as `Interaction | Stream` at every use. An inline
`stream: true as const` keeps the correct overload.

Neither upgrade required a `minimum-release-age` waiver: 2.15.0 and 0.115.0
were checked and 2.13.0 is the newest genai release clearing the 7-day window.

* chore(deps): upgrade openai to 7.0.0

v5 is the only major with real breaking changes for us; v6 widened a Responses
output type and v7 only raised the Node floor to 22, which apps/sim already
requires. Three things needed fixing:

`ChatCompletionMessageToolCall` became a union of function and custom tool
calls, and the custom variant has no `function` field — 43 unguarded `.function`
accesses across the OpenAI-compatible providers. Narrow once at each
`message.tool_calls` read through a shared `isFunctionToolCall` guard rather
than casting at every use.

That guard deliberately tests for the `function` payload instead of
`type === 'function'`. Many OpenAI-compatible vendors omit `type` on tool calls
entirely — our own fixtures do — so discriminating on it type-checks perfectly
and then silently drops every tool call those providers return.

`ChatCompletionCreateParams.verbosity` narrowed from `string` to a literal
union, and the Responses API's output and input item unions now diverge on
members Sim never emits (computer-use call outputs, whose `status` admits
`failed`, and the `AdditionalTools` escape hatch). Echoing output back as input
is what a tool loop is supposed to do, so that conversion is asserted once in
convertResponseOutputToInputItems and the streaming loop now routes through it
instead of pushing raw output items.

The hand-rolled multipart upload in file-attachments.server.ts can now be
replaced with the SDK's typed `expires_after` — left for a follow-up so this
commit stays a pure upgrade.

* fix(providers): correct defects found auditing the attachment and SDK changes

The mechanical rewrite that added `isFunctionToolCall` to every `tool_calls`
read also rewrote three truthiness guards, where the filtered array was
computed, discarded, and the unfiltered value used in the body. Filter once and
use that value. The helper also landed between `trackForcedToolUsage`'s TSDoc
block and its declaration, leaving that block documenting the wrong function.

Raise the Bedrock ceiling from 3.75 MB to 4.5 MB. Converse caps an image at
3.75 MB and a document at 4.5 MB, and a single `maxBytes` cannot express both.
Taking the lower bound looked conservative but regressed 3.75-4.5 MB documents,
which Converse accepts and which work today. At the document bound every size
that works now still works, and only genuinely-too-large files are rejected
early; oversized images in that band keep surfacing as a Bedrock API error,
exactly as they do without the entry.

Both limits re-verified verbatim against the primary docs: Converse's Message
reference ("Each image's size ... no more than 3.75 MB", "Each document's size
must be no more than 4.5 MB") and Fireworks' vision guide ("Total base64-encoded
images must be less than 10MB").

* fix(providers): keep every attachment that works today working

Auditing the routing change for backwards compatibility turned up two bands it
silently broke.

Lowering the single inline cap made the upload path mandatory above ~6 MB, but
every large-file path reads its bytes back out of cloud object storage. A
deployment without it — local dev, any disk-backed self-host — inlines those
files as base64 today and would have started failing outright with "requires
cloud file storage". Split the one number in two: the inline ceiling stays at
10 MiB, and a separate threshold marks where an upload becomes *preferable*
because the base64 copy no longer fits the payload store. Where no upload path
is reachable, base64 hydration now runs to the inline ceiling as before, and a
missing cloud-storage backend leaves the file for the inline path instead of
throwing.

The two strategies also cross over at different sizes now. `files-api` carries
every type the provider already accepts, so it takes over at the lower
threshold. `remote-url` only fetches images and PDFs, so switching early would
have started rejecting 6-10 MB text documents that inline fine today; it takes
over only once inlining is genuinely impossible.

Revert the Groq ceilings. Its published "20MB" governs a request carrying an
image URL, and on this path the body holds only the URL, so it cannot bind on
the files maxBytes guards. Groq documents no limit on the image it fetches, so
tightening the per-file cap to 20,000,000 and summing raw bytes against the
request cap would both reject uploads that work today on no documented basis.

* fix(providers): drop the provider ceiling changes and close the audit findings

A six-agent line-by-line audit against the vendors' live docs found the
`models.ts` ceiling work was not the strict improvement it was written as, so
all of it is reverted:

- bedrock's 4.5 MB cap broke video. Converse takes image, document AND video
  blocks, and video is allowed 25 MB base64 — a single `maxBytes` cannot
  express three content classes, and every 4.5-10 MiB `.mp4` that works today
  would have started failing.
- openai's combined 50 MB cap is the FILE-input limit. Image inputs are
  governed separately at 512 MB / 1500 images, so summing every attachment
  rejected eight 8 MB PNGs that OpenAI documents as legal.
- fireworks' per-file ceiling was unreachable behind the request budget, while
  the upload picker went on advertising it — a size the UI accepts and
  execution always rejects.
- The whole `perRequestMaxBytes` feature goes with them: it summed raw bytes
  against caps that are variously on encoded bytes, on one content class, or on
  a body that carries only URLs, and it double-counted a file referenced from
  several messages even though the uploader dedupes by key.

Only openai's per-file `maxBytes` stays corrected, to decimal 50,000,000 — the
one number a vendor states unambiguously and writes no MiB against.

Also fixed, all found by the same audit:

The hydration cap stopped short of where `remote-url` actually switches over,
so 6-10 MiB attachments on anthropic/openrouter/xai/groq/together/baseten/vllm
had neither base64 nor a handle and failed outright — the very band this branch
exists to fix. Both decisions now come from one function so they cannot drift.

Eight more sites where the mechanical rewrite computed a filtered array and
then read the unfiltered one (deepseek, sakana, nvidia, kimi), leaving those
providers without the narrowing they appear to have.

`isFunctionToolCall` threw on a null or primitive `tool_calls` entry, because
`in` requires an object — reachable exactly on the self-hosted gateways this
filter was added for. It is now total, and all 32 test mocks match it rather
than being quietly more permissive.

`checkForForcedToolUsage` in utils/litellm/mistral evaluated the response before
the `tool_choice` test, turning a tolerated malformed body into a TypeError on a
path that never used to touch it.

Gemini: `satisfies` restores the excess-property checking the dropped
annotations removed, the poll loop recognises the terminal statuses v2 added
instead of spinning for an hour and reporting a timeout, and the streaming doc
block no longer names five events that were renamed six lines below it.

* fix(providers): report attachment limits in the unit vendors publish

The size ceilings are decimal MB — that is how OpenAI, AWS and Fireworks all
write them — but the error messages divided by 1024², so OpenAI's 50 MB cap
was reported to the user as "48MB". Someone shrinking a 49 MB file to get
under it was chasing a limit that does not exist.

One formatter, used by all three messages, so the file size and the ceiling in
the same sentence are always in the same unit.

* fix(providers): derive the limit unit from the ceiling it belongs to

The previous commit fixed OpenAI's "48MB" by dividing every ceiling by 10⁶ —
which broke the other seven. Only OpenAI's constant is decimal; anthropic,
google, together and openrouter are 50 MiB, baseten and vllm 25 MiB, groq and
xai 20 MiB. Rendering those as decimal MB overstated each by ~5%, so a 21 MB
file on groq was rejected with "(21MB) exceeds the 21MB limit" — a sentence
that contradicts itself and sends the user to shrink a file to a size that is
still over. Same class of bug as the one being fixed, sign flipped.

Both figures now render through one unit taken from the ceiling, so the number
a user is told is the number the vendor publishes and the two sizes in a
sentence are always comparable. Tested against every ceiling in the registry
rather than only the values that happened to round cleanly.

Two more from the same audit:

A file with a missing or zero declared size was stranded on a files-api
provider: hydration bailed on the real byte length while `shouldUseLargeFilePath`
saw `0 > threshold` as false, so it got neither base64 nor a handle and failed
as "may no longer be accessible" — a size failure wearing an access failure's
message. Uploads read the real bytes and enforce the ceiling themselves, so an
unknown size now routes to one.

The oversized-attachment error blamed the provider for a deployment problem:
a files-api provider on a host without cloud storage reported that the provider
"has no large-file upload path", which is not true of the provider.

`isFunctionToolCall` only proves `function` is present, never that it is well
formed, so the trace enricher is defensive again about a hollow payload without
giving up the compile-time gate. The 32 test mocks now match production exactly.

* fix(providers): stop an over-limit size rendering as the limit itself

Deriving the unit from the ceiling fixed the 5% error but left the precision
fixed at two decimals, so a file one byte over a 20 MiB cap still printed
"(20.00MB) exceeds the 20MB agent attachment limit" — the same self-contradicting
sentence, now in a ~5 KB band above every ceiling in the registry. The size
rounds up and the ceiling rounds down, so the two can no longer collide.

The test that was supposed to guard this asserted a file 0.03MB over and an
OpenAI file that was under the limit — neither anywhere near the band — so it
passed while the bug was live. It now walks `limit + 1` for every ceiling, and
goes red against the old rounding.

The reason clause added last commit also claimed a deployment had no cloud file
storage whenever the strategy was not inline. A generated document on a
remote-url provider reaches that same error with storage fully configured,
because a signed URL points at the generation source rather than the rendered
artifact — so it was told something false about its own deployment. That case
now names itself.

* fix(providers): order the attachment failure reason by how general the cause is

The generated-document arm was checked first, so it won over both other causes
and told users two things that were not true.

On an inline-strategy provider — bedrock, mistral, ollama, fireworks, litellm,
vertex, kimi — there is no upload path for any file, generated or not, but the
message blamed the document format and implied a plain PDF would go through.

On openai or google with cloud storage unconfigured it was simply false: a
generated document does take the Files API path there, and that exact file
uploads fine once storage exists. The one actionable fix was hidden from the
operator.

A provider with no upload path cannot be helped by changing the file, and a
deployment with no object storage cannot reach any upload path whatever the
file is, so both now outrank the format-specific case — which is left saying
only what is true of it: a signed URL points at the generation source rather
than the rendered file.

The formatter is unchanged. It was brute-forced over every real ceiling and
three million random pairs with no collision or inversion, but the test's six
ceilings all divide to exact integers, so floor, round and ceil are
indistinguishable on them and the limit-side rounding was unpinned. A ceiling
with a fractional remainder now covers it.
…rbosity, and thinking level (simstudioai#6233)

* improvement(agent): allow variable references in reasoning effort and verbosity

Reasoning Effort and Verbosity were select-only dropdowns, so a workflow could
not sweep them from a variable or an upstream block the way it already can with
the model. Both become editable comboboxes, matching the model field directly
above them and the managed-agent selectors.

- switch both subblocks to `combobox`, keeping their fetched per-model option
  lists intact
- keep them visible when `model` itself holds a reference, since the concrete
  model id is only known at execution time and cannot be matched against the
  static capability list
- normalize the resolved level in the provider chokepoint so a reference that
  resolves to `"High"` or to nothing behaves sanely instead of hitting a
  provider 400

* improvement(agent): allow variable references in thinking level

Extends the same treatment to Thinking Level so all three model-tuning fields
behave consistently, and logs a level a model does not declare.

- switch `thinkingLevel` to `combobox` with the reference-aware condition
- normalize it alongside the other two; an empty resolve now takes the
  deliberate "send nothing" path rather than the incoherent half-state it hit
  before, and stays distinct from an explicit `none`
- warn when a level is not one the model declares, still forwarding it: Sim's
  per-model lists drive the pickers and can lag a provider, and a sweep needs
  the provider's own error rather than a silent fallback to the default

* improvement(agent): report a model level the sanitizer discards

A model bound to a variable or block reference only resolves at execution time,
so a run whose reference landed on a model outside Sim's catalogue had its
requested level cleared with no signal and quietly fell back to that model's
default. Dropping stays the safe default — a provider with no such parameter
rejects the whole request — but it is now reported.

- log the field, model, and value whenever an unsupported-field level is cleared
- cover both diagnostics, including that they stay quiet for a declared level
  and for the `auto` / `none` sentinels

* fix(providers): redact resolved level content from sanitizer diagnostics

The model-level fields accept environment and block references, so an
unrecognized level is not necessarily a mistyped level — it is whatever the
reference resolved to, which can be secret content. The diagnostics added for
dropped and undeclared levels echoed it straight into server logs.

- log a level only when the catalogue declares it somewhere, or it is an
  `auto` / `none` sentinel; anything else is reported by length alone
- stop discarding levels for a model the catalogue has never seen. Absent is
  unknown, not known-incapable, and a reference is exactly how a newly released
  model arrives before Sim catalogues it — the provider decides instead. Models
  the catalogue knows, and every dynamic-provider id, keep the protective drop

* fix(providers): redact the level in Anthropic's unsupported-thinking warning

Forwarding an undeclared level is deliberate, but it means the Anthropic adapter
receives it and interpolates it straight into its "not supported, ignoring"
warning. Since the field is reference-bound, that value can be whatever a
mistyped `{{ENV_VAR}}` or block reference resolved to — so the redaction added
for the sanitizer's own diagnostics was leaking one layer downstream.

- promote the level renderer to `providers/utils` as `describeModelLevel`, the
  single gate every site echoing a caller-supplied level goes through
- use it in Anthropic's warning and in both sanitizer diagnostics

* refactor(providers): drop the sanitizer's level diagnostics

The two warnings logged server-side, where the workflow author who set the level
never sees them, and the surprising case they described — a level discarded for
a model newer than the catalogue — is now fixed at the source rather than
narrated. They also carried the redaction that leaked resolved content before it
was caught, so removing them removes that surface entirely.

Levels still normalize, and still drop for a catalogued model that does not take
the field. `describeModelLevel` stays for Anthropic's unsupported-thinking
warning, which is a pre-existing log this feature newly exposes to resolved
reference content.
* fix(docker): upgrade bun to 1.3.14 to unbreak the Next 16.3.0 server

Bun 1.3.13 cannot load Next 16.3.0's compiled server runtime. The app container
runs the Next server under Bun (`oven/bun:1.3.13-slim`, `bun apps/sim/bootstrap.js`),
so every app-page render threw and `/api/health` returned 500:

  ⨯ Error: Failed to load external module
    next/dist/compiled/next-server/app-page-turbo.runtime.prod.js:
    TypeError: Expected CommonJS module to have a function wrapper.
    If you weren't messing around with Bun's internals, this is a bug in Bun

Isolated to Bun, not Next, by loading that exact module in the real images:

  Next 16.2.12 + Bun 1.3.13 -> loads (why staging was fine before)
  Next 16.3.0  + Bun 1.3.13 -> CJS wrapper error
  Next 16.3.0  + Bun 1.3.14 -> loads

Bun 1.3.14 is the current stable and already fixes it, so this bumps every pin
rather than reverting the framework upgrade, which would only defer the same
latent Bun bug to the next attempt.

Why no gate caught it: local dev machines and this bump's own verification run
Bun 1.3.14, while the container and CI pinned 1.3.13 — and CI only *builds* the
image, it never boots one and probes `/api/health`. A container smoke test in CI
would have caught this before merge; that is worth adding separately.

* fix(docker): align the remaining bun pins with 1.3.14

Two pins were missed in the first pass because the search was scoped to
docker/, package.json and .github/workflows/:

- .devcontainer/Dockerfile still built on oven/bun:1.3.13-alpine
- PI_BUN_VERSION in apps/sim/scripts/pi-sandbox-packages.ts was still 1.3.13,
  despite being documented as mirroring the root packageManager field, so Pi
  sandbox images would have kept installing the Bun release that cannot load
  the Next 16.3.0 server runtime.

Fixed surgically rather than with a repo-wide replace: "1.3.13" also appears
inside SVG path data in apps/sim/components/icons.tsx and
apps/docs/components/icons.tsx, which a blind sed would have corrupted.
…code (simstudioai#6242)

Next 16.3.0's Turbopack optimizer models a bare `return <asyncCall>()` tail
call inside an async function as returning the promise object, then propagates
that always-truthy fact through the caller's `await`. Where the result feeds an
`if (x)` whose every branch returns, it concludes the branch is always taken and
deletes everything after it from the emitted bundle.

Two sites shipped to production that way:

- `POST /api/credentials` lost its entire create path — the transaction, the
  org locks, the insert, the audit, the 201. A first-time create fell into the
  existing-credential branch and threw on `existingCredential.id`, so every new
  credential 500'd.
- `upsertAsyncToolCall` collapsed to `async () => await getAsyncToolCall(id)`.
  The insert is simply gone; it returns null for every new async copilot tool
  call. Silent — no error, no failed request.

A differential scan of 71,266 source string literals across `.next/server` and
`.next/static`, comparing images built from the same commit on 16.2.12 and
16.3.0, found exactly these two and nothing else. That scan cannot see dropped
branches with no distinctive string literal, which is why the version goes back
rather than the two sites being patched alone.

Both are also hardened with `return await`, verified to defeat the miscompile in
a minimal reproduction. The TypeScript toolchain cleanup from the original bump
(dropping @typescript/native-preview, `useTypeScriptCli`) is kept.
@utcarshsrivastava-collab

Copy link
Copy Markdown
Collaborator Author

Upstream sync 2026-08-04 — grill complete, 2 questions blocking the merge

Range e2fecc86..1b9e0f25518 commits, v0.7.29 → v0.7.55.
Measured with git merge-tree: 249 conflicted files (235 content, 7 modify/delete, 7 add/add).
Fork changed 1531 files since baseline, upstream 4903, 428 overlap.

Full analysis: .upstream-sync/ledger/2026-08-04/run.md## Grill analysis.

Good news first: forkFirst held perfectly — upstream touched zero of its 42 prefixes. All the
work is in manualReview paths.


Questions

Both are on the Arena deployed chat (apps/sim/app/(interfaces)/chat/** — 51 fork-modified
files vs 21 upstream). One reply covers both.

Q1 — Keep or drop Arena deployed-chat voice mode?

Upstream simstudioai#6215 + simstudioai#6218 removed voice from the deployed chat (keeping workspace dictation), deleting
voice-interface/ (+ its particles canvas), input/voice-input.tsx, hooks/use-audio-streaming.ts,
hooks/queries/voice-settings.ts, api/proxy/tts/stream/route.ts + its contract, and the anonymous
chatId branch of api/speech/token.

Why it blocks: the fork's fork-only ArenaDeployedChat.tsx imports VoiceInterface and
useAudioStreaming and renders <VoiceInterface> as live code; chat.tsx imports the deleted
voice-settings hook. Either way this fails bun run check until decided, and the child agents
can't resolve input.tsx / chat.tsx / hooks/index.ts / use-chat-streaming.ts / constants.ts
without the answer.

→ A or B? If B, please confirm you accept an unauthenticated platform-key-spending relay on the
public chat endpoint.

Q2 — Confirm the resolution stance for (interfaces)/chat/** + api/chat/**

Assumed default — "fork-first on presentation, upstream-first on auth / security / route contracts":

→ Confirm, or state a different split. A bare "confirmed" is enough.


Already resolved — no answer needed (D1–D15 in the ledger)

# Decision
D1 Migration collision: both sides added 02580261. Renumber fork's four → 02820285 (upstream ends at 0281), rebuild _journal.json (262 + 281 → 285 entries), regenerate snapshots. packages/db/schema.ts = union.
D2 Upstream drops workflow_folder / workspace_file_folders (simstudioai#6051). Only one fork-only consumer breaks: lib/workflows/default-user-workflows/service.ts → port to the generic folder table with resourceType: 'workflow'. No fork-first option; the table ceases to exist.
D3 Exa /research/v1 returns HTTP 410exa_research is already broken in prod. Take upstream's refresh (simstudioai#6074), which auto-routes saved exa_research workflows to the new Agent op and preserves the research[0].text output shape. Re-apply the fork's exaHosting + __costDollars onto upstream's files; fix the fork-only tools/exa-hosting.test.ts.
D4 HubSpot = union. Both sides expanded additively and disjointly; the two files the fork edited (list_contacts, list_associations) upstream never touched. Conflicts only in index.ts / types.ts / blocks/blocks/hubspot.ts. Note the hubspot OAuth provider in auth.ts is byte-identical to upstream's — not a fork customization, just relocated to lib/auth/connectors/providers.ts.
D5 lib/auth/: take upstream's restructure + better-auth 1.6.23 + org session policies, then re-apply every Arena block (ARENA_V3_OAUTH_CALLBACK_ORIGINS, dev embed origins, BETTER_AUTH_COOKIE_DOMAIN cross-subdomain cookies, Arena hub trusted origin, Microsoft endpoint/tenant helpers).
D6 providers/models.ts: union. Keep all azure/* + azure-anthropic/* families and the fork-retained legacy models; take all upstream models + the sunset-tier and prompt-caching fields. Do not tag fork legacy models with sunset tiers (would show amber/red warnings on Arena canvases).
D7 Registries: union. Fork keeps arena, arena-development, chart_generator, cost, development, facebook_ads, figma, google_ads_v1, image_fusion, p2_docs, presentation, semrush, spyfu, unipile. Upstream lands 12 new integrations + the webhook and SimSim Chat renames.
D8 The 7 apps/sim/public/** conflicts are Sim wordmark vs Arena marks → keep fork's.
D9 bunfig.toml: fork side is unchanged from baseline so upstream's minimumReleaseAge = 604800 supply-chain gate lands with no conflict. If install then fails on a young fork dep, add it to minimumReleaseAgeExcludes — do not revert the gate to 0.
D10 ci.yml (fork +237 / upstream +597): upstream base, re-apply the fork's EC2/GHCR deploy jobs.
D11 Root package.json: union — must preserve upstream-sync (this harness), check:credentials, check:secrets, vendor-pricing:*, repair:workflow-room-redis-keys. Drop dev:full:minimal-registry (upstream removed the minimal-registry hatch, simstudioai#6163).
D12 sync-local-draft.ts was moved, not deleted (→ apps/sim/stores/workflows/). Port the fork's flushMergedLocalDraftToServer (fixes deploy silently clearing image-generator provider/model) + its tests + 3 call sites. base-card.tsx was replaced by the shared Resource/resource-tile primitives (simstudioai#6202) — re-attach the fork's clickKnowledgeBaseEvent Mixpanel tracking there (the only one of 16 Arena Mixpanel call sites upstream deleted).
D13 The fork deleted two upstream tests that upstream then modified (home/hooks/use-chat.test.ts, stream/turn-model-serialize.test.ts) → restore upstream's and adapt.
D14 apps/sim/local-copilot/ (77 fork-only files) produces no conflicts but will fail type-check — it imports deep upstream copilot internals that moved. Known break: LOAD_USER_SKILL_TOOL_NAME from the deleted @/lib/mothership/skills (skills moved to lib/skills/ + lib/workflows/skills/). Treat as a post-merge repair cluster.

New CI gates this PR must satisfyregenerateAfterMerge needs to grow beyond mship:generate
to also run tool-metadata:generate, mship-tools:generate, skills:sync,
agent-stream-docs:generate; plus check:tool-registry-boundary, check:migrations,
check:api-validation:strict now gate on fork-added tools/blocks/migrations.

Also filed in extensibility-notes.md: merge-policy.json is missing ~30 fork-owned paths
(most notably apps/sim/local-copilot/, app/api/local-copilot/, public/favicon/**,
public/icon.svg, and 8 fork blocks / 8 fork hooks / 6 fork lib dirs), with recommendations to move
the Azure model catalog and the Arena auth-origin logic into fork-owned sidecars — those two files
alone account for the largest conflicts in the repo.


Reply here with Q1: A or B and Q2: confirmed (or your split), then comment
/upstream-sync resume.

@SaitejaSankoji

Copy link
Copy Markdown
Collaborator

Q1 A, Q2 confirm

@SaitejaSankoji

Copy link
Copy Markdown
Collaborator

/upstream-sync resume

github-actions Bot and others added 10 commits August 4, 2026 10:18
Merge took upstream package.json and dropped check:secrets (and other D11 fork scripts/deps), breaking the secret-scan job. Also fix docs placeholder URL host so the scanner accepts it.

Co-authored-by: Cursor <cursoragent@cursor.com>
Blind upstream checkout dropped fork CI/harness scripts before child agents ran, ignoring grill D11. Deterministically union manifests (upstream base + fork-only keys) and record dropScripts in merge-policy.

Co-authored-by: Cursor <cursoragent@cursor.com>
Align harness with policy intent: only forkFirst/upstreamFirst auto-resolve. Drop the default --ours fallback. Document that agents should extend merge-policy when they learn recurring rules.

Co-authored-by: Cursor <cursoragent@cursor.com>
@utcarshsrivastava-collab
utcarshsrivastava-collab deleted the upstream-sync/2026-08-04T09-49-31 branch August 4, 2026 11:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants