Skip to content

ac12644/langgraph-tenancy

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

10 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

langgraph-tenancy

CI PyPI - Version License: MIT

Tenant isolation for LangGraph persistence — as a drop-in wrapper.

Using LangGraph.js? Same package, same guarantees: ac12644/langgraph-tenancy-js · npm

LangGraph's own threat model says it plainly:

Checkpoint savers index by thread_id. Without application-level auth, any caller with a valid thread_id can access that thread's state. [...] Users embedding LangGraph directly must implement their own access controls.

If you run a multi-tenant product on open-source LangGraph, the only thing between Customer A's agent state and Customer B's is a query filter in your application code. This package replaces that convention with enforcement — plus the operational surface a multi-tenant product needs: per-tenant usage metering, quota enforcement, GDPR-style erasure, migration of pre-tenancy data, and observability hooks for every denial.

Install

pip install langgraph-tenancy

Usage

Wrap your existing checkpointer and store. Nothing else changes.

from langgraph_tenancy import (
    TenantScopedCheckpointer,
    TenantScopedStore,
    InMemoryUsageLedger,
)

ledger = InMemoryUsageLedger()
checkpointer = TenantScopedCheckpointer(PostgresSaver(...), usage_ledger=ledger)
store = TenantScopedStore(InMemoryStore())

graph = builder.compile(checkpointer=checkpointer, store=store)

# tenant_id is now REQUIRED on every invocation
graph.invoke(
    {"messages": ["hello"]},
    config={"configurable": {"thread_id": "t1", "tenant_id": "acme"}},
)

# free per-tenant token metering, extracted from checkpointed messages
ledger.totals("acme")  # TenantUsage(input_tokens=..., output_tokens=..., by_model={...})

What it enforces

Raw LangGraph behavior With langgraph-tenancy
Any caller with a thread_id reads that thread Threads are physically keyed tenant::thread; wrong-thread_id bugs cannot cross tenants
Missing filter → silent unscoped query Missing tenant_idTenantRequiredError, nothing read or written
Missing thread_id → writes keyed "None" Refused with a loud TenancyError
checkpointer.list(None) enumerates every tenant's threads Refused with UnscopedAccessError; list() with a tenant (and no thread) enumerates only that tenant's threads
Store namespaces are convention; any node can read any namespace Every op is rooted at the tenant segment, resolved from the run config automatically
delete_thread("t1") deletes whoever owns t1 Requires an explicit for_tenant("acme").delete_thread("t1") handle
usage_metadata buried in checkpoint blobs, unqueryable Aggregated per tenant (and per model), deduped by message id
Tenant ids are arbitrary strings Restricted to [A-Za-z0-9_-]{1,64} — safe in every backend's key/namespace encoding

Quotas

Give the checkpointer per-tenant limits and it enforces them at the persistence boundary — an over-quota tenant's next run fails at its first checkpoint, before any model call spends money:

from langgraph_tenancy import TenantLimits, QuotaExceededError

checkpointer = TenantScopedCheckpointer(
    inner,
    usage_ledger=ledger,  # quota reads current usage from ledger.totals()
    quota_limits=lambda tenant_id: plans.lookup(tenant_id),
    # e.g. TenantLimits(max_total_tokens=1_000_000, max_messages=10_000)
    # return None for "no limits"
)

Semantics: check-then-write. The turn that crosses a limit still completes (that spend already happened and must not be lost); every run started after that fails with QuotaExceededError. Runs are never killed mid-flight. Supply quota_usage= to read usage from your own billing system instead of the ledger.

Usage metering in production

InMemoryUsageLedger is the reference implementation (bounded memory, deduped by message id). For a real backend, implement record() — and totals() if you want quota enforcement:

class PostgresLedger:
    def record(self, tenant_id: str, record: UsageRecord) -> None:
        db.insert_usage(tenant_id, record)

    def totals(self, tenant_id: str) -> TenantUsage:  # enables quotas
        return db.usage_totals(tenant_id)

A raising ledger does not fail checkpoint writes: the error is reported through on_event (type "ledger_error") and the record is retried on the next checkpoint. Pass ledger_errors="throw" if you'd rather fail the write.

Metering extracts from both checkpoints and pending writes, so it keeps working with DeltaChannel (beta) graphs, whose checkpoints carry only a sentinel instead of the accumulated messages.

Observability

Every denial is a security-relevant event. Wire them to your logger, metrics, or OpenTelemetry:

def on_event(event):  # TenancyEvent
    # event.type: "denied" | "quota_exceeded" | "ledger_error"
    logger.warning("tenancy %s", event)

checkpointer = TenantScopedCheckpointer(inner, on_event=on_event)
store = TenantScopedStore(inner_store, on_event=on_event)

Handler errors are swallowed — observers can never break the data path.

Admin, GDPR, and migration

Everything out-of-band goes through an explicit per-tenant handle:

acme = checkpointer.for_tenant("acme")

acme.list_threads()          # every thread id belonging to acme
acme.delete_thread("t1")     # one thread
acme.purge()                 # GDPR erasure: every acme thread
acme.copy_thread("t1", "t2") # tenant-scoped copy
acme.prune(["t1", "t2"])     # tenant-scoped prune

store.for_tenant("acme").purge()  # every acme store item

# Adopting langgraph-tenancy on an existing deployment? Migrate pre-tenancy
# (unprefixed) threads under their rightful tenant, history intact:
acme.adopt_thread("legacy-thread-id", delete_source=True)

adopt_thread replays every checkpoint — order, parentage, metadata, and pending writes — under the tenant-scoped key, so conversations continue exactly where they left off.

No magic

The entire mechanism is key prefixing plus mandatory-context checks, in a few small files you can audit in ten minutes:

  • thread ids become "{tenant_id}::{thread_id}" before reaching your database; the prefix is stripped from everything returned.
  • store namespaces ("memories",) become ("{tenant_id}", "memories").
  • tenant ids are restricted to [A-Za-z0-9_-]{1,64} — an allowlist, because tenant ids end up inside storage keys and namespace encodings of whatever backend you use, and a blocklist can't anticipate all of them.

It composes with any BaseCheckpointSaver / BaseStore implementation — Postgres, SQLite, Redis, MongoDB, in-memory — because it never touches storage itself.

What it is not

  • Not authentication. You decide which tenant a request belongs to; this package guarantees that decision is enforced everywhere downstream.
  • Not encryption. Combine with EncryptedSerializer for at-rest encryption.
  • Not a replacement for database-level controls in high-assurance setups (RLS, schema-per-tenant) — it's the layer that makes your application unable to leak, whatever the database allows.

Tested

The adversarial test suite — every test attempts a cross-tenant access the raw LangGraph API allows — runs against InMemorySaver and a real PostgresSaver in CI. Coverage includes tenant-wide listing, quota enforcement, GDPR purge, adopt_thread migration, delta-channel metering, and denial events, all proven on actual SQL storage.

Development

uv venv && uv pip install -e ".[test]"
uv run pytest                 # postgres tests skip if no server is reachable

# to run the postgres leg locally:
export LG_TENANCY_PG_URI=postgresql://user@localhost:5432/langgraph_tenancy_test
uv run pytest

License

MIT

Releases

Packages

Contributors

Languages