Skip to content

[Feature]: provider/key cooldown recovery notifications #1146

Description

@Cheurteenyt

Area

Provider routing / rate-limit cooldowns / account pools / Dashboard observability

Goal

Notify operators when an OpenCodex-managed provider, API key, or account that entered a temporary cooldown becomes eligible for retry again.

The feature should build on OpenCodex's existing cooldown state rather than requiring users to manually retry requests just to discover whether a temporary limit has cleared.

This is primarily an observability and recovery-state feature.

It must not claim that an upstream quota has definitively reset unless OpenCodex has provider-specific evidence proving that.

Current status

OpenCodex already has substantial rate-limit handling.

For API-key pools, current dev:

  • records a key cooldown after HTTP 429;
  • respects an upstream Retry-After when available;
  • otherwise uses a bounded fallback cooldown;
  • skips keys that remain inside their cooldown;
  • lazily removes cooldown state after its deadline;
  • can rotate to another eligible key.

The current API-key cooldown is therefore conceptually:

429 observed
    ↓
cooldownUntil established
    ↓
key skipped while cooldown is active
    ↓
cooldown deadline passes
    ↓
key becomes eligible for selection again

What is missing is a user-visible signal for that transition.

Important distinction: cooldown expiry is not quota-reset proof

The original proposal suggested periodically calling:

GET /v1/models

and treating HTTP 200 as proof that a provider's rate limit or quota had reset.

That is not a safe generic contract.

A provider may:

  • allow model-list requests while generation remains rate-limited;
  • apply separate limits to discovery and inference;
  • expose a public /models endpoint that ignores the user's generation quota;
  • not expose /models at all;
  • charge or throttle metadata requests differently;
  • use account-, model-, region-, or endpoint-specific quota windows.

Therefore:

GET /v1/models -> 200

must not generically mean:

generation quota reset

and must not automatically clear a generation cooldown.

Recovery states

The implementation should distinguish at least two concepts.

1. Local cooldown expired

OpenCodex's own retry window has ended.

Example:

state: cooldown_expired
provider: opencode-zen

This means:

OpenCodex will allow another request attempt.

It does not mean:

The provider has confirmed that quota is available.

2. Provider recovery confirmed

Use this stronger state only when a provider-specific contract can positively establish that the relevant generation/account limit is usable again.

Example:

state: provider_recovered
provider: example-provider

This requires provider-specific evidence.

Unknown recovery remains unknown.

Do not manufacture certainty from a generic health or model-list request.

V1 scope

The first implementation should require no active upstream polling.

When an existing OpenCodex cooldown transitions from active to expired, emit a bounded recovery event.

For example:

provider_cooldown_expired

or:

{
  "event": "provider_cooldown_expired",
  "provider": "opencode-zen",
  "scope": "api-key",
  "at": "2026-08-09T18:00:00Z"
}

The exact event name can follow existing observability conventions.

This event means that OpenCodex's local suppression period ended and the route/key is eligible to be tried again.

Notification surfaces

At minimum support:

Structured log

Example:

[cooldown] opencode-zen key pool slot became eligible for retry

Use a safe provider/account/key-pool identifier.

Never log raw credentials.

Dashboard

Expose a bounded notification such as:

OpenCode Zen is eligible for retry again.

Do not say:

Quota reset.

unless an authoritative provider-specific source actually established that fact.

If notification history is retained, keep it bounded.

Retry-After behavior

When the original upstream response supplies a trustworthy:

Retry-After

OpenCodex may naturally use that deadline for the cooldown transition.

Example:

429
Retry-After: 45

↓ 45 seconds

cooldown_expired

When no upstream timing exists and OpenCodex uses a fallback cooldown, the notification should remain phrased as:

eligible for retry

rather than claiming the upstream limit reset.

Provider-specific recovery probes

Future provider integrations may define an authoritative recovery probe where the provider exposes one.

Such a probe must be an explicit provider capability.

Conceptually:

provider supports recovery validation
           ↓
provider-specific bounded probe
           ↓
confirmed usable
           ↓
provider_recovered

Do not implement one generic probe URL for every OpenAI-compatible provider.

A recovery probe must define:

  • endpoint;
  • authentication;
  • what resource/quota it actually validates;
  • response semantics;
  • polling bounds;
  • timeout;
  • whether the probe itself consumes quota;
  • safe failure behavior.

Unsupported providers simply use local cooldown-expiry notifications.

Polling safety

If provider-specific polling is added later:

  • polling must be opt-in or capability-gated;
  • use bounded intervals and total lifetime;
  • honor authoritative reset/retry timing before polling;
  • apply jitter where several keys/accounts recover together;
  • stop on shutdown;
  • stop after the configured maximum age;
  • never create an unbounded background timer;
  • never let probe failures damage normal routing.

Do not poll every 15–30 seconds by default across every provider merely because a 429 was observed.

Account and key isolation

Recovery state belongs to the entity that entered cooldown.

Examples:

provider A / key 1
provider A / key 2
account A
account B

A recovery event for one key/account must not clear or report recovery for another.

For API-key pools, use the existing opaque pool-slot identity rather than exposing the key.

For OAuth accounts, use the same safe account-reference policy used by the account-management surfaces.

Interaction with routing

This issue does not introduce a second routing engine.

After a cooldown expires, the existing routing/account-pool logic decides whether that provider/key/account should actually be selected.

A recovery notification means:

eligible again

not:

force the next request onto this route

Session affinity, priority, health, quota evidence, and failover policy remain owned by the relevant routing contracts.

Relationship to #1145

#1145 tracked the specific OpenCode Zen problem where the short-window rate limit was effectively invisible.

That issue is completed.

PR #1330 now:

  • documents the observed Zen burst limit;
  • improves the client-facing 429 message;
  • provides useful retry timing when Zen omits it.

This issue should not reimplement that provider-specific error guidance.

#1146 remains the provider-agnostic cooldown transition notification layer.

Relationship to #695

#695 owns generic OAuth account-pool selection, health, affinity, and failover.

If OAuth account cooldowns later feed this notification surface, their lifecycle should come from that generic account-pool state rather than being recreated here.

Relationship to #1267

#1267 owns broader API-key-pool quota classification and proactive round-robin.

This issue may observe cooldown transitions created by the key-pool machinery, but it does not own key-selection strategy.

Relationship to #657

#657 owns irreversible rejection-triggered reset-credit recovery.

A local cooldown-expiry notification must not trigger reset-credit consumption.

Relationship to #822 / #823

Neither behavior should be triggered automatically merely because a local cooldown timer expired.

Privacy

Notifications and diagnostics must never expose:

  • API keys;
  • access tokens;
  • refresh tokens;
  • authentication headers;
  • raw provider error bodies;
  • unsafe account identifiers.

Use bounded provider names and opaque account/key slot references where needed.

Out of scope

This issue does not require:

Acceptance criteria

Existing foundation

Remaining work

  • A bounded event is emitted when a managed cooldown expires.
  • The event clearly means eligible for retry, not necessarily quota reset.
  • Dashboard users can see the recovery transition.
  • Structured logs expose the transition safely.
  • Recovery state remains scoped to the correct provider/key/account.
  • No credential material is exposed.
  • Notification delivery cannot interfere with routing.
  • No generic /v1/models success is treated as proof of generation-quota recovery.
  • Provider-confirmed recovery is represented separately from local cooldown expiry.
  • Any future active recovery probe is explicitly provider-capability-gated and bounded.
  • Shutdown cleans up any notification/probe timers.
  • Regression tests cover expiry, isolation, duplicate suppression, shutdown, and privacy.
  • Accepted implementation is merged into dev.

Metadata

Metadata

Assignees

No one assigned

    Labels

    account-poolOAuth, credentials, Codex pool, quota, failover, plansenhancementNew feature or requestproxyHTTP proxy, routing, reverse-proxy / management auth

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions