Cross-session, cross-project memory for Claude Code via Graphiti + Neo4j, exposed as an MCP server.
Different from Claude Code's built-in per-project ~/.claude/projects/.../memory/:
| Built-in auto-memory | pb-graphiti | |
|---|---|---|
| Storage | Flat markdown files | Neo4j graph |
| Scope | Single project (per-cwd) | Cross-project + cross-session |
| Scale | ~50 facts before index gets unreadable | Hundreds+ with native search |
| Supersession | Manual (edit/delete) | Native (X was true until time T, then Y) |
| Cross-fact queries | None | Entity-relation walk via Cypher / MCP search |
| Loaded | Auto every session | Queried on-demand via MCP tools |
Use both. Auto-memory for hard rules that MUST load every session. pb-graphiti for domain facts you want to query when relevant.
.claude-plugin/ plugin + marketplace manifest
.mcp.json MCP client config — URL prompted at enable time (default http://localhost:8765/mcp)
skills/graphiti-usage/ SKILL.md — write/query discipline, group_id scope model
commands/ /pb-graphiti:ingest-{folder,slack,magento-modules,tickets,email} + wipe-channel + automate
hooks/ SessionStart recall + PreCompact consolidation hooks
scripts/ Python helpers used by the ingest commands and hooks (stdlib only)
infra/ docker-compose recipe for the host-side Neo4j + Graphiti stack
Installing the plugin gives Claude the MCP client config + the usage skill. You still need to stand up the server yourself (see Graphiti Server).
Add the marketplace and enable the plugin:
# In your project root (or globally via ~/.claude/settings.json)
claude /plugin marketplace add proxiblue/pb-graphiti
claude /plugin install pb-graphiti@pb-graphitiOr for local development, clone and reference by path in ~/.claude/settings.json:
{
"extraKnownMarketplaces": {
"pb-graphiti": {
"source": { "source": "directory", "path": "/path/to/pb-graphiti" }
}
},
"enabledPlugins": {
"pb-graphiti@pb-graphiti": true
}
}The plugin's MCP config defaults to http://localhost:8765/mcp (overridable at enable time — see URL: prompted at install time below). That endpoint is YOUR responsibility — run the bundled compose stack on your host (or any reachable host):
We build our own docker image from upstream Graphiti source — the image they ship was stale at the time this plugin was made.
cd infra/
# 1. Clone graphiti as ./upstream/ (compose builds the MCP image from it)
git clone https://github.com/getzep/graphiti upstream
# 2. Fill in API keys
cp .env.example .env
$EDITOR .env # at minimum: NEO4J_PASSWORD, ANTHROPIC_API_KEY, VOYAGE_API_KEY
# 3. Up
docker compose up -d
# 4. Verify (note: /mcp, no trailing slash — server 307-redirects /mcp/ → /mcp,
# and many MCP HTTP clients do not follow redirects on POST)
curl -sS -X POST http://localhost:8765/mcp \
-H 'Content-Type: application/json' \
-H 'Accept: application/json, text/event-stream' \
-d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"smoke","version":"0.1"}}}'
# Expect: event: message data: {"jsonrpc":"2.0","id":1,"result":{...}}
# Neo4j browser: http://localhost:7474 (login: neo4j / <your NEO4J_PASSWORD>)ANTHROPIC_API_KEY— entity extraction. Default model isclaude-haiku-4-5(cheap, fast).VOYAGE_API_KEY— embeddings. Voyage's free tier covers typical solo use.OPENAI_API_KEY— leave the dummy value. graphiti-core hardcodes an OpenAI cross-encoder constructor that needs any non-empty string at boot, but the reranker path is rarely called when LLM+embedder are non-OpenAI. Replace with a real OpenAI key only if you start seeing rerank-path failures.
Want different providers (Gemini for LLM, OpenAI embeddings, etc.)? Edit infra/config/config.yaml — the Graphiti config schema lists every supported provider.
When you enable the plugin, Claude Code prompts for graphiti_url. Pick the form that matches where Claude Code runs, not where the docker stack runs:
| Where Claude Code runs | URL to enter |
|---|---|
| Same host as the docker-compose stack (most consumers) | http://localhost:8765/mcp (default) |
| Inside a container that shares the docker host (DDEV, devcontainer) | http://host.docker.internal:8765/mcp |
Change later via /plugin config pb-graphiti or by editing ~/.claude/settings.json → pluginConfigs["pb-graphiti@pb-graphiti"].options.graphiti_url.
Gotchas that will silently break the connection:
- Trailing slash. The server 307-redirects
/mcp/→/mcp; most MCP HTTP clients don't follow 307 on POST. Always use/mcp(no slash). Plugin defaults handle this correctly. - Wrong host for the runtime.
host.docker.internaldoes not resolve on a bare Linux host;localhostinside a container points at the container itself, not the docker host. Pick the URL based on where Claude is actually running.
There is one Neo4j database backing all your projects. Namespacing is per-call, via the group_id field on every add_memory / search_nodes call.
Three tiers:
group_id="initial_ingest"— pinned facts always shown at session start, capped at 20 most recent. Curate manually; no auto-pruning. Use for north-star rules you always want loaded.group_id="fleet"— cross-project facts (methodology, vendor rules, organisation policy).group_id="<project-id>"— project-specific quirks. Resolve from$DDEV_PROJECTorbasename $(git rev-parse --show-toplevel).
From inside a project, ALWAYS query both project + fleet on dynamic recall: group_ids: ["<project-id>", "fleet"]. Surfaces fleet rules everywhere without leaking project A's quirks into project B. initial_ingest is loaded automatically by the SessionStart hook in addition to dynamic recall.
Full discipline (when to write, when to read, scope confirmation rule, cypher to move mis-scoped nodes) lives in skills/graphiti-usage/SKILL.md. The skill auto-surfaces to Claude when the graphiti MCP is reachable.
Two hooks ship with the plugin and turn on the moment you enable it. They make the read/write loop automatic so you stop re-explaining project context every session.
Every time you start a Claude Code session, the plugin queries Graphiti and injects two tiers of context:
- Pinned facts — all episodes in
group_id="initial_ingest", capped at 20 most recent. Use this for fleet-wide rules that should be visible every session regardless of cwd. - Dynamic recall — top 8 entity nodes for
[<project-id>, "fleet"]against a default semantic query.
Project id is resolved deterministically: $DDEV_PROJECT → git toplevel basename → fleet-only fallback. The hook is async, so a slow or unreachable Graphiti never blocks session start — if it fails for any reason, the session proceeds with no injection.
What Claude sees as additionalContext:
## Always-loaded (pinned via group_id='initial_ingest')
1 permanent fact(s):
- **Caveman default response style** — All responses default to caveman style... [src: ~/claude-skills-central/rules/caveman.md]
## Graphiti recall (project=acme-store + fleet)
Top 8 relevant facts from prior sessions:
- **acme-store LIVE branch** [Project] (`acme-store`) — LIVE-equivalent branch is `uat`, not `main`...
- **Vendor X blocked** [Vendor] (`fleet`) — Module quality issues; use Vendor Y instead...
- ...
To pin a fact: call add_memory(group_id="initial_ingest", name="...", episode_body="...", source_description="...") via the graphiti MCP from any session. There's no slash-command wrapper yet — write directly.
A second SessionStart hook runs alongside the Graphiti recall — see "Dual-source recall policy" below.
Claude Code already has a per-project file-based auto-memory at ~/.claude/projects/<sanitized-cwd>/memory/. pb-graphiti does NOT replace it — the two stores carry different content:
- Disk auto-memory — project TODOs, feedback rules, user profile. Stable, slow-changing, what-to-do-next oriented.
- Graphiti — consolidated session episodes from PreCompact + TaskCompleted hooks. Operational/procedural facts, what-happened-when, cross-session continuity.
Neither is a superset. When the user asks to recall past work ("recall last session", "what did we remember about X"), the correct move is to query both, then reconcile.
To make this consistent across sessions, the plugin ships a second SessionStart hook (scripts/session_start_disk_memory_nudge.sh) that:
- Always injects a one-line policy reminder that recall is dual-source.
- If the disk auto-memory dir has
*.mdfiles modified in the last 7 days (excludingMEMORY.md), also injects a count + the latest filename, so the user knows recall is worth invoking.
Auto-memory dir is resolved as $HOME/.claude/projects/<sanitized-cwd>/memory/ (the Claude Code default), or via $CLAUDE_AUTO_MEMORY_DIR if set.
Fires every time a TaskUpdate sets a task to status=completed. The hook agent reviews the recent task-relevant slice of the transcript and writes any worthwhile facts to Graphiti (capped at 3 episodes per fire — task scope is narrow). Most task completions are trivial ("list files", "syntax check"), so most fires output nothing worth saving and exit cheap.
This is the main consolidation trigger on the 1M-context model, where PreCompact effectively never fires (context fills up much slower than compaction would happen on the regular model). Task completions are the natural milestone signal.
Citation: source_description = claude-code-session://<session_id> [task-completed <task-id> <YYYY-MM-DD>] — traceable both to the session AND the specific task that triggered the write.
When the conversation is about to be compacted, the plugin runs an agentic hook that reviews the soon-to-be-lost context and writes any worthwhile facts to Graphiti — without prompting. Scope is auto-resolved by the same rules as SessionStart, defaulting to host when no project context is present.
Facts the hook writes: decisions with rationale, project quirks, incident root causes, vendor verdicts, client preferences, runbook steps. Facts it skips: ephemeral session state, in-progress task lists, anything obtainable from git log / git blame, restatements of CLAUDE.md.
On 1M-context models, PreCompact effectively never fires (sessions don't hit the compaction threshold). The SessionEnd hook below is the practical replacement for that case.
Mirrors PreCompact's behaviour but fires when the session closes regardless of whether compaction happened. This is the primary write path for sessions on 1M-context models where PreCompact stays dormant. Same 8-episode cap, same triage criteria. Citation: claude-code-session://<session_id> [session-end <YYYY-MM-DD>].
If a session does hit compaction AND then terminates, both PreCompact and SessionEnd may fire — the SessionEnd handler is instructed to skip facts already consolidated by an earlier PreCompact in the same session to avoid duplicates.
Fires every time a Task tool subagent returns to its parent session. Captures decisions made by gitnexus-reviewer, tdd-worker, standards-enforcer, devils-advocate, and security-trio agents that would otherwise vanish (HCF plan-tasks change file state, not Claude Code's TaskUpdate, so the TaskCompleted hook doesn't fire on them).
Strict 2-episode cap per fire (subagent fires are frequent — over-writing here clogs the graph). Hard triage: pure-lookup subagents (Explore, general-purpose for read-only enumeration) exit nothing worth saving; clean PASS verdicts from reviewers also exit. Only writes when the subagent surfaced a non-obvious finding, a pushback that drove a change, a vendor verdict, or a runbook step learned the hard way.
Citation: claude-code-session://<session_id> [subagent-stop <subagent_type> <YYYY-MM-DD>] — preserves both the parent session AND which subagent surfaced the fact.
PreCompact / SessionEnd / TaskCompleted / SubagentStop all auto-write without prompting. This bypasses the graphiti-usage skill's Hard Rule 1 (which requires user confirmation before any add_memory call). The rule still applies to add_memory calls Claude makes mid-conversation; it only carves out for non-interactive hook contexts.
Plugin hooks load automatically — you do not edit ~/.claude/settings.json to wire them. Claude Code reads hooks/hooks.json from each active plugin's install dir (~/.claude/plugins/cache/<plugin>/<version>/hooks/hooks.json) at session start and merges those entries into its hook registry alongside whatever is in settings.json.
So a freshly-installed pb-graphiti starts firing SessionStart / PreCompact / TaskCompleted / SessionEnd / SubagentStop with zero post-install wiring. If they're not firing, the issue is one of these (NOT "you forgot to wire them"):
- Session was started BEFORE the plugin install or update. Hook lists are cached at session start — no hot-reload. After
claude plugin install pb-graphiti@pb-graphiti(or update),/quitand restart Claude Code. Then check. - Plugin is not enabled in the current project. Run
/plugin list— pb-graphiti should be listed as active. If not, install/enable it for that project context. DDEV-resident sessions enable per-container (ddev exec claude plugin install pb-graphiti@pb-graphiti). - Claude Code version pre-dates a specific hook type. SessionStart / PreCompact / TaskCompleted have been around for many releases; SessionEnd and SubagentStop are newer. Update Claude Code itself if a specific hook silently does nothing — the older binary may not know that hook name.
- Diagnosis looked in
~/.claude/settings.jsononly. Plugin hooks live in the plugin's own dir (path above), not in settings.json. Agrepof settings.json will never show plugin hooks. Use/plugin show pb-graphiti(or inspect the cache dir directly) to see what the plugin offers.
To verify which hooks are actually live in a session:
# from a Claude Code session shell
cat ~/.claude/plugins/cache/pb-graphiti/*/hooks/hooks.json | jq '.hooks | keys'
# expected: ["PreCompact","SessionEnd","SessionStart","SubagentStop","TaskCompleted"]If you're on a Claude Code build where plugin-hook auto-load is genuinely broken, you can copy the entries out of the plugin's hooks/hooks.json into ~/.claude/settings.json directly. They'll then load from settings.json instead — at the cost of needing to re-sync after every plugin update. This should be a last resort; the auto-load is the supported path.
To turn one off without uninstalling the plugin, edit ~/.claude/settings.json:
{
"disableAllHooks": false,
"hooks": {
"PreCompact": []
}
}Or disable hooks from every plugin with "disableAllHooks": true.
Five slash commands import external content as Graphiti episodes. All go through the same add_memory tool an agent would use ad-hoc; the scripts just batch the calls.
Walks a directory and ingests *.md, *.markdown, *.txt, *.rst (override with --include). Markdown is chunked on ## headings; plain text on paragraph clusters; default target ~1500 words per chunk. Reference time on each episode is the source file's mtime — Graphiti is bi-temporal, so historical docs get correct valid-at metadata. Source description is the file's file:// URI so recalled facts can be cited back to the exact document.
Use for:
- Meeting transcripts (export as markdown first)
- PRDs / ADRs / specs
- Internal wikis, runbooks, postmortems
- Module documentation — pair with GitNexus: GitNexus indexes the code structure, Graphiti indexes the why (purpose, design notes, integration intent from each module's README.md). Example:
Future sessions then recall "we have module X in project Y that does Z" without re-reading the codebase.
/pb-graphiti:ingest-folder app/code --include 'README*.md,readme.md,*.md' --group-id <project-id>
Code-heavy doc trees: pass --suppress-code-entities when ingesting a project's docs/ tree where the content references many NPM / Composer packages, file paths, class names, or infrastructure components. Without it, the graph fills with low-value Component nodes (express, body-parser, nginx, php-fpm, etc.) that duplicate what GitNexus already indexes. The flag is OFF by default because ADR-style docs may legitimately treat code refs as concepts ("we chose class X over class Y because..."). When ON: NPM/Composer packages, file paths, function/method names, and generic infra refs are dropped — but business vendors with intent (Cliniko, Stripe, Telnyx, etc.) and vendor verdicts stay.
IMAP-based email ingest. One episode per email thread (grouped by RFC 2822 References/In-Reply-To, falling back to normalized subject). HTML stripped, quoted replies trimmed, attachments noted in headers only.
Connection details are stored as plugin userConfig — set them once at plugin-enable time (or any time after via /plugin config pb-graphiti or by editing ~/.claude/settings.json → pluginConfigs["pb-graphiti@pb-graphiti"].options):
userConfig key |
What it is | Default |
|---|---|---|
imap_host |
IMAP server hostname (imap.gmail.com, imap.fastmail.com, …) |
(empty — required) |
imap_port |
IMAP port | 993 |
imap_user |
IMAP account login (usually full email address) | (empty — required) |
imap_folder |
Folder to read (INBOX, [Gmail]/All Mail, Archive, …) |
INBOX |
imap_password_env |
Name of the env var holding the password | IMAP_PASSWORD |
The password itself is never stored in settings or on the CLI — only the env var NAME is configured. The actual value lives in your shell environment, sourced from a password manager / pass / shell rc / however you handle secrets. Use an app password (Gmail, Microsoft, Fastmail all offer them) rather than your real account password. The script reads os.environ[<configured-env-name>] at runtime; if the env var is empty, the script refuses to run.
Multiple mailboxes per project aren't supported via userConfig — if you need to ingest from more than one IMAP account, override the values on the CLI (--imap-host, --imap-user, etc.) and the script picks those up.
Two layers, stack them:
- Address allowlist —
--addresses '@client.com,external-consultant@vendor.com'. Matches against EVERY participant (From,To,Cc,Bcc,Reply-To) — including Cc'd parties, so threads where the client is Cc'd alongside an internal recipient still match. Each entry is either a full address or a domain match (@client.comcovers anyone at that domain). Run client-side after IMAP fetch so coverage is reliable across providers. - Keyword gate —
--include-keywords 'projectname,ticket-prefix'. Applied to the rendered thread body (after HTML strip + quote trim). Useful when the address list alone lets in too much chatter — e.g., a client domain that also handles unrelated business.
Both filters AND together (a thread must pass both gates if both are set). Combine an address allowlist with a tight keyword list when correspondence is mixed.
Pass --require-relevance to refuse to run unless an address allowlist or keyword list is set. Prevents the foot-gun of accidentally ingesting an entire mailbox with no scope. Strongly recommended for first-time runs against a new mailbox.
Email ingest ships with a default noise filter that drops messages from auto-systems BEFORE any LLM cost: mailer-daemons, UptimeRobot, GitHub notifications, no-reply senders, undeliverable bounces, OOO autoresponders, calendar invites. Override via:
--exclude-from 'pattern1,pattern2'— additive to defaults; case-insensitive substring on From header--exclude-subjects 'pattern1,pattern2'— additive to defaults; case-insensitive substring on Subject--no-default-noise-filter— disable defaults (rare; "give me literally everything")--exclude-list-file <path>— load patterns from a line-oriented file (additive to defaults + CLI flags)
Smart noise scan during dry-run: --dry-run runs a deterministic pattern scan over the surviving messages and writes editable suggestions to <state-dir>/suggested-excludes-<group>.txt. It detects:
- Sender username patterns:
noreply,no-reply,donotreply,bounce,mailer-daemon,notifications, etc. - Sender domain patterns:
@uptimerobot.com,notifications.*,bounces.*, etc. - Known automated services (flagged for review, NOT auto-suggested): Sentry, Sendgrid, GitHub, Asana, CircleCI, etc. — these often carry useful business signal (Stripe billing, Atlassian tickets) so user judgement applies
- Header-based signals:
Auto-Submitted,Precedence: bulk/list,List-Id,Return-Path: <>(true bounce) - Subject suspect keywords:
digest,weekly summary,build #,delivery status,password reset, etc. - Repeating subject prefix clusters (3+ messages with same normalized subject)
Workflow: dry-run → review file → comment out false positives → re-run dry-run with --exclude-list-file to confirm → live ingest with the same file. Cost saving on big mailboxes: filtering 30-50% noise before LLM can save $10-25 in Haiku per ingest.
--since YYYY-MM-DD is a server-side IMAP filter; --min-words to drop auto-replies (default 10); --min-thread-messages to drop one-offs. Code-entity suppression on by default (same rationale as tickets).
--backlog-mode --max-threads N walks the mailbox newest→oldest using a per-group cursor file at <state-dir>/backlog-cursor-<group>.json. Each invocation fetches one window of --backlog-window-days (default 30) BEFORE the cursor, sorts surviving threads newest-first, writes up to N, then advances the cursor to (oldest-written-msg-date + 1 day). If the window has no survivors the cursor jumps a full window back so empty months don't stall progress.
Designed to pair with the cron wrapper: forward pass first, backlog only fires when the forward pass wrote 0. That gives steady backfill on quiet days, zero overhead on busy days. Default cap = 20 threads/run (PB_BACKLOG_PER_DAY), default window = 30 days (PB_BACKLOG_WINDOW_DAYS). Set PB_BACKLOG_DISABLE=1 to opt out.
Ignore --since when this flag is set — it's overridden by the cursor. State + dedupe file is still the regular email.json; backlog and forward share the same "already written" set, so they can't double-write a thread.
A single IMAP session can only fetch so many messages before the provider drops the socket (Zoho closes around 2,000-3,000 messages; Gmail similar; Microsoft varies). For back-fills covering more than a few months, use:
--batch-days N(default 30) — splits[--since, today]into N-day windows; each window uses a fresh IMAP session. Lower to 7 or 3 for dense mailboxes.--parallel-workers N(default 1) — runs N batch windows concurrently; each opens its own IMAP connection. Bump to 2-4 for big back-fills. Respect provider concurrent-connection caps (Zoho ≈5, Gmail ≈15, Fastmail ≈2/IP).
The dedupe state file means a partial run (e.g., 5 batches succeeded then a connection dropped) can be resumed by simply re-running — already-ingested threads are skipped. If a batch warns "failed" with a socket error, halve --batch-days and re-run.
source_description is mid:<Message-ID> of the thread root — RFC 2392 URI that opens in mail clients respecting the scheme.
Best for: client mailboxes with project-specific subject prefixes, vendor correspondence threads, contract negotiation history.
GitHub issues and pull requests with their comment threads. Driven by the gh CLI (auth via $GH_TOKEN or gh auth login). One episode per ticket containing title, state, labels, author, body, and chronological comment thread. Bot comments (dependabot, github-actions, codecov) filtered out by default.
The since argument accepts a year (2024) or full ISO date (2024-06-15). Repo defaults to the current git origin; override with --repo owner/name. Default labels excluded: dependencies,duplicate,invalid,wontfix — override with --exclude-labels.
Cap: 30k chars per episode (long discussions truncated). source_description is the ticket's html_url, so every recalled fact cites back to the exact thread.
Why this is the highest-value ingest source. Tickets are where decisions happen: vendor verdicts, design rationale, SEO strategies, bug postmortems. After ingesting, search_nodes(group_ids=[<project>, "fleet"], query="why X") typically surfaces the original discussion. Pairs naturally with ingest-magento-modules: modules show what was built, tickets show why.
Magento-aware module-doc ingest — the recommended path for capturing project module documentation. Walks app/code/<Vendor>/<Module>/ (and optionally vendor/*/*/ with --include-vendor) and assembles ONE episode per module containing:
- Canonical
Vendor_Modulename (frometc/module.xml) <sequence>dependencies (module-level — these read as docs)- composer.json: description, version, require, license
- README.md / readme.md content (if present)
- CHANGELOG.md head (last 5 entries, if present)
Deliberately excluded — GitNexus owns these: di.xml preferences, plugin targets, events.xml observers. Putting class wiring into Graphiti creates hundreds of Component nodes per project — graph noise that duplicates GitNexus's structural index. Keep the layers clean: GitNexus = code structure (symbols, callers, signatures); pb-graphiti = the why (purpose, design rationale, vendor verdicts from READMEs and changelogs).
Each episode's source_description is the module's file:// URI, so recalled facts cite back to the exact module directory.
Modules with no README, no composer description, and no CHANGELOG are skipped — a bare module.xml has nothing for Claude to recall. The dry-run reports the skip count.
Takes the .zip Slack produces from Workspace settings → Import/Export Data → Export. Writes one episode per channel-per-day, format=message. Threads inline under their parent. Channel-join/leave/topic noise filtered. Reference time = midnight UTC on the message day.
Slack chats are where the why lives — decisions, vendor verdicts, incident chats. Worth ingesting selectively rather than the whole workspace. Filter flags stack:
| Flag | What it drops | Default |
|---|---|---|
--channels '#x,#y' |
Anything not from listed channels | (all channels in export) |
--since YYYY-MM-DD |
Days before this date | (no time filter) |
--include-keywords '<words>' |
Days that don't mention any of these (substring, case-insensitive) | (no filter) |
--exclude-keywords '<words>' |
Days that mention any of these | (no filter) |
--include-users '<ids or names>' |
Anything not from these users | (all users) |
--exclude-users '<ids or names>' |
Messages from these users (e.g. bots) | (none) |
--min-words N |
Per-message: under N words ('lgtm', 'thanks', emoji-only) | 3 |
--min-day-messages N |
Days with <N surviving messages after per-message filters | 3 |
For client channels that mix project work and chat, --include-keywords is the highest-leverage filter — list the project name, key vendors, ticket prefixes, etc., and entire off-topic days vanish from the plan. Dry-run reports a per-filter dropped-day count so you can tune.
Citations: when --workspace-slug <slug> is set (e.g. acmeco for https://acmeco.slack.com), every rendered message gets its Slack permalink inline in the episode body, and the episode's source_description becomes the channel-archive URL for that date. Recalled facts can then be cited back to the exact thread. Without --workspace-slug, source_description falls back to the structured key slack:<channel>:<YYYY-MM-DD> — informational only.
Your OWN Claude Code session transcripts, distilled. Claude Code writes every session to ~/.claude/projects/<mangled-cwd>/<session-uuid>.jsonl — an append-only log that is ~90% noise: tool_use arguments, tool_result bodies, thinking blocks, hook-injected context, slash-command wrappers. A single real session can be 50+ MB.
This channel does not dump raw transcripts (that would flood the graph with grep-echoes and restatements of CLAUDE.md — exactly what the usage discipline forbids). It condenses each session down to its signal layer first:
- Kept: real user prompts, assistant prose replies, and one compact
[tools: Bash(desc), Edit(file), mcp:search_nodes×2]marker per turn. - Dropped: thinking, tool arguments/results,
isMetamessages,<system-reminder>/<pb-chatroom-inbox>hook blocks, slash-command wrappers,<local-command-stdout>, and (by default) subagent sidechain turns.
A 56 MB / 3,300-line session collapses to ~10 chunks of pure conversation. Graphiti's extraction then pulls the durable facts — decisions with rationale, incident root-causes, vendor verdicts, runbook steps — out of the prose, steered by a custom instruction that explicitly tells it to ignore exploratory dead-ends and step-narration ("signal is a conclusion reached, not a step taken").
| Flag | What it does | Default |
|---|---|---|
--project=<glob> |
Which mangled project-dirs to consider (= form required — the value starts with -) |
all projects |
--group-id <id> |
Graphiti scope. One glob ↔ one scope, or you cross-contaminate | (required) |
--since YYYY-MM-DD |
Skip sessions last modified before this | (none) |
--settle-minutes N |
Skip sessions modified within N min — never half-ingest a live session; $CLAUDE_SESSION_ID always skipped |
30 |
--min-user-turns N |
Skip trivial sessions with fewer than N real prompts | 2 |
--include-sidechains |
Also condense subagent turns (more coverage, more noise) | OFF |
--keep-code-entities |
Do NOT suppress Component extraction; transcripts are path/class-dense so suppression is ON by default (GitNexus owns those) | suppress |
--max-sessions N |
Cap sessions per run | 0 (no cap) |
Scope discipline matters most here. Transcripts are keyed by the cwd they ran in, so a project's sessions belong in that project's scope — map its dir (Claude Code replaces / with -, e.g. ~/workspace/ittools/lcdscreen → --project=-home-lucas-workspace-ittools-lcdscreen*) to --group-id lcdscreen. Use --group-id fleet only for host-level sessions under -home-lucas. Never fold every project into one scope.
Citations: source_description is claude-session://<project-dir>/<session-id>, so a recalled fact points back at the exact session. Reference time = the session's first real user-prompt timestamp (when the work happened, not when ingested).
This is a host-side channel — transcripts live on the host, not inside containers. Run it (and its cron wrapper) from the host.
Every episode written by either ingest command carries a real link in source_description (file:// URI for docs; Slack archive URL for messages). The shipped graphiti-usage skill instructs Claude to surface that source whenever it acts on a recalled fact — same discipline as artefact citation in the investigation protocol. See skills/graphiti-usage/SKILL.md for the recall-side rules.
- Confirm scope first. Both commands prompt for fleet vs project before any write.
- Dry-run first. Both default to
--dry-runin the slash command flow; you see the plan (episode count, sample names) before any episode is created. - Dedupe via state file.
.pb-graphiti-ingest.jsonin the cwd records what's been ingested. Re-runs from the same cwd skip already-ingested items. Ctrl-C is safe — state is flushed after every successful write. Pass--reingestto force a full re-write. - Cost reality. Every episode is one Anthropic Haiku call (entity extraction) + one Voyage embed call. A year of one Slack channel chunked per-day ≈ 365 episodes ≈ ~$0.50 in Haiku. A folder of 50 medium markdown docs chunked per heading ≈ 200-400 episodes ≈ ~$0.50-1.00.
The scripts are stdlib-only Python (no pip install required). Run them directly if you prefer:
python "$(claude plugin path pb-graphiti)/scripts/ingest_folder.py" --help
python "$(claude plugin path pb-graphiti)/scripts/ingest_slack.py" --helpIndicative for a solo developer writing ~5 episodes/day across ~10 projects:
| Service | Tier | Approx monthly cost |
|---|---|---|
| Neo4j Community Edition | Self-hosted (this compose) | $0 |
| Anthropic Haiku (entity extraction) | Pay-as-you-go | <$2 |
| Voyage embeddings | Free tier (200M tokens/month) | $0 |
Heavy fleet writers (50+ episodes/day, 50+ projects) will see Haiku creep toward ~$10/month. Embedder usage stays well inside Voyage free tier.
Manual /pb-graphiti:ingest-* runs work fine but you have to remember to run them. For tickets and email — where new content lands continuously — a cron-driven schedule keeps recall current without thinking about it.
Run /pb-graphiti:automate for an interactive setup. It detects whether you're on a host or inside a DDEV container and picks the right install path. Or set it up by hand using the wrappers below.
scripts/cron/ ships per-source wrappers that:
- Source credentials from
$HOME/.pb-graphiti/env(template atscripts/cron/env.example) - Resolve the group_id automatically:
$PB_GRAPHITI_GROUP_ID→$DDEV_PROJECT→basename $(git rev-parse --show-toplevel)→ error if none - Compute
--sincefromLOOKBACK_DAYS(cross-platform date math — works on Linux GNU date and macOS BSD date) - Write dedupe state to a stable location (
$HOME/.pb-graphiti/state/<source>.json) so cron-from-arbitrary-cwd doesn't lose state - Log to
$HOME/.pb-graphiti/logs/<source>.logwith timestamps (also mirrored to the fleet log dir whenPB_GRAPHITI_FLEET_LOGSis set — see below)
The email wrapper runs in two phases per invocation:
- Forward pass — fetches the last
LOOKBACK_DAYSdays (default 7), dedupe handles overlap. - Backlog pass — fires only when the forward pass wrote 0 threads. Walks the mailbox newest→oldest using
--backlog-modeand the per-group cursor file, capped atPB_BACKLOG_PER_DAY(default 20) threads per run. SetPB_BACKLOG_DISABLE=1to disable.
Recommended schedule:
# Every 6 hours — tickets land continuously, want them in recall by next session
0 */6 * * * /path/to/pb-graphiti/scripts/cron/ingest-tickets.sh
# Daily at 02:00 — email volume is steady, daily is enough; off-hours so it doesn't compete with day work
0 2 * * * /path/to/pb-graphiti/scripts/cron/ingest-email.sh
# Nightly, per project — condense yesterday's session transcripts (host crontab only).
# One line per project: pair a project-dir glob with the graphiti scope it maps to.
30 2 * * * PB_GRAPHITI_GROUP_ID=lcdscreen PB_SESSIONS_PROJECT='-home-lucas-workspace-ittools-lcdscreen*' /path/to/pb-graphiti/scripts/cron/ingest-sessions.sh
45 2 * * * PB_GRAPHITI_GROUP_ID=fleet PB_SESSIONS_PROJECT='-home-lucas' /path/to/pb-graphiti/scripts/cron/ingest-sessions.shDDEV's web container crontab doesn't survive ddev restart unless persisted via .ddev/web-build/. Drop a file at .ddev/web-build/pb-graphiti.cron containing the schedule. The /pb-graphiti:automate command walks you through this. Once installed, the cron stays put across ddev rebuild.
Many DDEV projects already have a Dockerfile.ddev-cron (or similar) that copies *.cron files from .ddev/web-build/ into /etc/cron.d/. The /pb-graphiti:automate command checks for that pattern and reuses it — only creating a new Dockerfile if no existing one already handles the copy.
Critical for DDEV: do NOT put the env file under $HOME. $HOME (/home/<user>) inside the web container is an overlay filesystem that is wiped on ddev restart. Use the host-mounted project directory instead:
PB_GRAPHITI_HOME=/var/www/html/.pb-graphitiPB_GRAPHITI_ENV=/var/www/html/.pb-graphiti/env
Set both as env vars at the top of .ddev/web-build/pb-graphiti.cron so the wrappers find the env file regardless of which user runs them. The cron file template from /pb-graphiti:automate does this automatically. Add /.pb-graphiti/env (or /.pb-graphiti/ for the whole state directory) to your project's .gitignore so secrets don't slip into version control.
In-container vs host:
- In-container gets
$DDEV_PROJECTauto-set, MCP reaches viahost.docker.internal, scripts are mounted at/var/www/html/.claude/plugins-seed/marketplaces/pb-graphiti/. Self-contained per project. Stops when DDEV stops — usually fine. - Host runs regardless of any DDEV state. Right choice for a central mailbox you want polled even when no project is up.
| Source | Recommended cadence | Reason |
|---|---|---|
ingest-tickets |
every 6 hours | New issues / PRs land continuously, want them by next session |
ingest-email |
daily 02:00 | Steady volume; off-hours so IMAP load doesn't compete with day work |
ingest-magento-modules |
manual, after composer update or new module ships |
Modules change rarely |
ingest-folder |
manual | Doc folders update on git push; rerun when something major lands |
ingest-slack |
manual (no live API ingest yet) | Current implementation needs a workspace export .zip; live API integration is future work |
ingest-sessions |
nightly 02:30, host only, one line per project | Transcripts accumulate daily; 30-min settle window skips live sessions; host-side (transcripts aren't in containers) |
$HOME/.pb-graphiti/
├── env # credentials (chmod 600, never commit)
├── state/
│ ├── tickets.json # dedupe hashes from the tickets ingest
│ └── email.json # dedupe hashes from the email ingest
└── logs/
├── ingest-tickets.log
└── ingest-email.log
The state files are the dedupe layer — they record which episodes have been successfully written. A cron run that finds "nothing new" is a no-op except for the IMAP/GitHub fetch overhead. To force a clean re-ingest, delete the relevant state file or pass --reingest on a manual run.
backlog-cursor-<group>.json (only present when --backlog-mode has run) holds the date the next backlog batch fetches BEFORE. Delete it to restart the back-catalogue walk from today.
The graphiti-fleet docker-compose.yml ships a small nginx sidecar (port 7475, host-only) serving:
| Path | What |
|---|---|
/ |
Dashboard SPA — per-group cards (total / +1h / +24h / latest write), recent-activity feed (last 20 episodes), auto-refreshes every 60s. Counters tick green when they grow between refreshes — you literally see ingest progress. |
/api/cypher |
POST-only read proxy to Neo4j. nginx envsubsts NEO4J_BASIC_AUTH into an Authorization header at container start, so the browser never sees credentials. Local-only (127.0.0.1 bound). |
/logs/ |
Directory browser over ${PB_GRAPHITI_FLEET_LOGS:-~/.pb-graphiti-fleet-logs} (read-only). Drill into a project's run logs from the dashboard. |
Visit http://localhost:7475 for the dashboard. Falls back to the static log browser at /logs/ if you want the raw cron-run output.
To surface a project's runs in the /logs/ browser, export PB_GRAPHITI_FLEET_LOGS=/path/on/host in $PB_GRAPHITI_ENV. The wrappers tee each run into $PB_GRAPHITI_FLEET_LOGS/<group_id>/<source>.log in addition to the per-project log. For DDEV-resident crons, mount the host fleet-log dir into the web container first (a one-line .ddev/docker-compose.fleet-logs.yaml bind), then point the env var at the in-container path.
For the dashboard's /api/cypher to work, graphiti-fleet's .env must include NEO4J_BASIC_AUTH=$(printf 'neo4j:%s' "$NEO4J_PASSWORD" | base64). The dashboard returns blank cards if this env var is missing — check docker logs graphiti-log-viewer for 401 responses on /api/cypher.
For operators who want to understand what's on disk, what survives, and how to back it up.
Measured on a real install with ~2 of 10 fleet projects fully ingested:
| Metric | Value |
|---|---|
| Total nodes (episodes + extracted entities) | ~2,700 |
| Total edges | ~18,000 |
| Avg episode body size | ~6.6 KB (max ~30 KB — capped by the ingest scripts) |
| Total episode text | ~4 MB |
| Embedding dimensions | 512 (Voyage voyage-4-lite) → ~4 KB per entity |
| On-disk Neo4j volume | ~550 MB |
Why the volume is much bigger than the raw text: vector embeddings on every Entity node, B-tree + vector indexes, edge metadata, Neo4j's pre-allocated page cache and transaction log.
Linear growth projection: at the density above, expect ~250 MB per fully-ingested project. A full fleet of 10 projects → ~2.5-3 GB. Adding extensive Slack/email/ticket history per project → 5-7 GB total. Still trivial on modern storage.
| Thing | Location | Lifecycle |
|---|---|---|
| Neo4j data files (the graph) | Named Docker volume graphiti-fleet_neo4j_data → host path /var/lib/docker/volumes/graphiti-fleet_neo4j_data/_data |
Persists across container/image rebuilds. Tied to the volume only. |
| Neo4j logs | Named volume graphiti-fleet_neo4j_logs |
Persists; rotate manually if growing. |
| MCP server (Graphiti) | Container only — no state | Rebuilt from ./upstream/ on docker compose build. |
| API keys (Anthropic, Voyage) | infra/.env on host (gitignored) |
Read at container start; loaded into env. Not stored in the data volume. |
| Per-call config (entity types, group_id defaults) | infra/config/config.yaml on host |
Bind-mounted into the MCP container. Changes need docker compose restart graphiti-mcp. |
| Operation | Data survives? |
|---|---|
docker compose down then up -d |
✅ |
docker compose restart |
✅ |
docker compose build (rebuild MCP image) |
✅ |
docker compose pull (newer Neo4j image) |
✅ |
| Container removed + image deleted manually | ✅ — volume is independent |
docker compose down -v |
❌ — -v deletes named volumes |
docker volume rm graphiti-fleet_neo4j_data |
❌ |
docker system prune --volumes |
❌ — wipes all unused volumes |
Host disk failure or /var/lib/docker corruption |
❌ |
The named volume is the protection layer. Anything that targets volumes specifically (the -v flag, volume rm, prune --volumes) loses everything.
Neo4j ships neo4j-admin database dump which produces a single-file snapshot. ~5 seconds at current size. Recipe:
# 1. Dump inside the container (writes to /data/backups within the volume)
docker exec graphiti-neo4j neo4j-admin database dump neo4j --to-path=/data/backups
# 2. Copy the dump out of the container to a host backup directory
docker cp graphiti-neo4j:/data/backups/neo4j.dump "$HOME/backups/graphiti-$(date +%F).dump"
# 3. Optional: prune dumps older than 30 days
find "$HOME/backups/graphiti-*.dump" -mtime +30 -delete 2>/dev/nullWire this into cron to run nightly:
# m h dom mon dow command
0 2 * * * docker exec graphiti-neo4j neo4j-admin database dump neo4j --to-path=/data/backups && docker cp graphiti-neo4j:/data/backups/neo4j.dump "$HOME/backups/graphiti-$(date +\%F).dump" && find $HOME/backups/graphiti-*.dump -mtime +30 -deleteFor off-host durability, pipe the dump through rclone to S3/B2/Drive after step 2.
# 1. Copy the dump file back into the container
docker cp "$HOME/backups/graphiti-2026-06-24.dump" graphiti-neo4j:/data/backups/
# 2. Load the dump (must be done while Neo4j is stopped or against a different DB)
docker compose stop neo4j
docker exec graphiti-neo4j neo4j-admin database load neo4j --from-path=/data/backups --overwrite-destination=true
docker compose start neo4jThe default heap + page cache in infra/docker-compose.yml is 1 GB heap + 512 MB page cache. Once the data volume exceeds the page cache, query latency creeps up (more disk I/O per query). When the data approaches 1 GB:
environment:
- NEO4J_server_memory_heap_initial__size=1G
- NEO4J_server_memory_heap_max__size=2G
- NEO4J_server_memory_pagecache_size=2G # raise to match data sizeRestart Neo4j after editing (docker compose up -d).
Episodes don't carry an explicit "ingestion source" field, but every channel writes a distinct source_description shape, so you can classify them with a string-prefix check:
| Channel | source_description pattern |
|---|---|
/pb-graphiti:ingest-folder |
file:///<absolute-path> (file URI; no trailing slash on the file itself) |
/pb-graphiti:ingest-magento-modules |
file:///<absolute-path-to-module-dir>/ (trailing slash — module directory) |
/pb-graphiti:ingest-slack (with --workspace-slug) |
https://<workspace>.slack.com/archives/<channel-id> (<YYYY-MM-DD>) |
/pb-graphiti:ingest-slack (no slug) |
slack:<channel>:<YYYY-MM-DD> |
/pb-graphiti:ingest-tickets |
https://github.com/<owner>/<repo>/issues/<n> (or /pull/<n>) |
/pb-graphiti:ingest-email |
mid:<Message-ID> |
/pb-graphiti:ingest-sessions |
claude-session://<project-dir>/<session-id> |
| PreCompact consolidation hook | claude-code-session://<session_id> [precompact <YYYY-MM-DD>] |
Manual add_memory |
whatever the caller passed (often opaque) |
Open the Neo4j browser at http://localhost:7474 (or use the HTTP API). Cypher:
MATCH (ep:Episodic)
WITH ep.group_id AS gid,
CASE
WHEN ep.source_description STARTS WITH 'https://github.com/' THEN 'github-tickets'
WHEN ep.source_description STARTS WITH 'mid:' THEN 'email'
WHEN ep.source_description STARTS WITH 'claude-code-session://' THEN 'precompact-hook'
WHEN ep.source_description STARTS WITH 'https://' AND ep.source_description CONTAINS '.slack.com/' THEN 'slack-permalinked'
WHEN ep.source_description STARTS WITH 'slack:' THEN 'slack-opaque'
WHEN ep.source_description STARTS WITH 'file://' AND ep.source_description ENDS WITH '/' THEN 'magento-module'
WHEN ep.source_description STARTS WITH 'file://' THEN 'folder-doc'
ELSE 'other/unknown'
END AS channel
RETURN gid, channel, count(*) AS episodes
ORDER BY gid, episodes DESC;Use case: you want to drop all GitHub-ticket episodes from lcd-mageos (e.g., to re-ingest with tighter filters) without touching modules, Slack, email, or hook-written facts.
Easy path: /pb-graphiti:wipe-channel <channel> — wraps everything below safely (dry-run → confirm → MCP delete → optional orphan prune → state-file reminder). Use the manual cypher only if you want to script the wipe or do something the slash command doesn't cover.
Step 1 — preview what would be deleted. Always run this first.
MATCH (ep:Episodic)
WHERE ep.group_id = 'lcd-mageos'
AND ep.source_description STARTS WITH 'https://github.com/'
RETURN count(ep) AS would_delete;If the count looks right, proceed. If it's surprisingly large or small, recheck your WHERE clauses before running the delete.
Step 2 — delete the episodes.
MATCH (ep:Episodic)
WHERE ep.group_id = 'lcd-mageos'
AND ep.source_description STARTS WITH 'https://github.com/'
DETACH DELETE ep;DETACH DELETE removes the episode node AND its MENTIONS edges to any extracted entities. Entities themselves are NOT deleted in this step — Graphiti's entities are aggregated across many source episodes, so an entity like a vendor name may still be referenced by module or Slack episodes after the tickets are gone. That's intentional.
Step 3 — clean up orphaned entities. After Step 2, some entities may have no remaining MENTIONS edges (i.e., they only ever appeared in ticket episodes). Remove them:
MATCH (e:Entity)
WHERE e.group_id = 'lcd-mageos'
AND NOT EXISTS { MATCH (:Episodic)-[:MENTIONS]->(e) }
DETACH DELETE e;This step is safe: it only deletes entities that no surviving episode references, so entities still mentioned by modules/Slack/email survive automatically.
Step 4 — verify.
MATCH (n)
WHERE n.group_id = 'lcd-mageos'
RETURN labels(n) AS labels, count(*) AS c
ORDER BY c DESC;You should see the Episodic count drop by the ticket count and entity counts reduced by any orphans. Modules / Slack / email episodes and their entities should be untouched.
Substitute the STARTS WITH clause in Step 1 and Step 2:
-- Wipe all Slack ingest (both permalinked and opaque variants)
AND (ep.source_description CONTAINS '.slack.com/' OR ep.source_description STARTS WITH 'slack:')
-- Wipe all email ingest
AND ep.source_description STARTS WITH 'mid:'
-- Wipe all Magento-module ingest (file:// URIs ending in /)
AND ep.source_description STARTS WITH 'file://' AND ep.source_description ENDS WITH '/'
-- Wipe all folder-doc ingest (file:// URIs NOT ending in /)
AND ep.source_description STARTS WITH 'file://' AND NOT ep.source_description ENDS WITH '/'
-- Wipe everything written by the PreCompact consolidation hook
AND ep.source_description STARTS WITH 'claude-code-session://'The .pb-graphiti-ingest.json state file in the cwd where you ran the original ingest contains hashes of every successfully-written episode. After wiping that channel from the graph, the next re-ingest will skip everything (dedupe-hits). Either:
- Delete the state file, OR
- Pass
--reingeston the next ingest run.
Otherwise you'll wipe the graph, re-run the ingest, and see "nothing to do (all up to date)" — confusing.
Every Claude-driven write to the graph carries a distinguishable source_description prefix so you can audit exactly what's been added automatically vs. by bulk ingest:
| Channel | Prefix |
|---|---|
| TaskCompleted hook (per-task consolidation) | claude-code-session://...[task-completed ...] |
| PreCompact hook (on compaction) | claude-code-session://...[precompact ...] |
Ad-hoc Claude add_memory mid-conversation |
claude-code-conversation://... |
To list everything Claude has self-added across all groups:
MATCH (ep:Episodic)
WHERE ep.source_description STARTS WITH 'claude-code-session://'
OR ep.source_description STARTS WITH 'claude-code-conversation://'
RETURN ep.group_id, ep.name, ep.source_description, ep.created_at
ORDER BY ep.created_at DESC
LIMIT 50;Scoped to one project:
MATCH (ep:Episodic) WHERE ep.group_id = 'lcd-mageos'
AND (ep.source_description STARTS WITH 'claude-code-session://'
OR ep.source_description STARTS WITH 'claude-code-conversation://')
RETURN ep.name, ep.source_description, ep.created_at
ORDER BY ep.created_at DESC;Just the consolidation hooks (no ad-hoc):
MATCH (ep:Episodic)
WHERE ep.source_description STARTS WITH 'claude-code-session://'
RETURN ep.group_id, ep.name, ep.source_description, ep.created_at
ORDER BY ep.created_at DESC LIMIT 50;MATCH (ep:Episodic)
WHERE ep.source_description STARTS WITH 'claude-code-session://'
OR ep.source_description STARTS WITH 'claude-code-conversation://'
WITH ep.group_id AS gid,
CASE
WHEN ep.source_description CONTAINS '[task-completed' THEN 'task-completed-hook'
WHEN ep.source_description CONTAINS '[precompact' THEN 'precompact-hook'
WHEN ep.source_description STARTS WITH 'claude-code-conversation://' THEN 'ad-hoc add_memory'
ELSE 'other-claude-write'
END AS hook
RETURN gid, hook, count(*) AS episodes
ORDER BY gid, episodes DESC;These return the actual node objects so the Neo4j Browser draws them as a graph (with MENTIONS edges to the extracted entities). Paste into the Browser, hit run, drag nodes to explore.
Latest 20 auto-ingested episodes with the entities they produced:
MATCH (ep:Episodic)
WHERE ep.source_description STARTS WITH 'claude-code-session://'
OR ep.source_description STARTS WITH 'claude-code-conversation://'
WITH ep ORDER BY ep.created_at DESC LIMIT 20
OPTIONAL MATCH (ep)-[r:MENTIONS]->(e:Entity)
RETURN ep, r, e;Scoped to one project (substitute your group_id):
MATCH (ep:Episodic)
WHERE ep.group_id = 'pvcpipesupplies'
AND (ep.source_description STARTS WITH 'claude-code-session://'
OR ep.source_description STARTS WITH 'claude-code-conversation://')
OPTIONAL MATCH (ep)-[r:MENTIONS]->(e:Entity)
RETURN ep, r, e;Only TaskCompleted writes (the new v0.12.0 hook):
MATCH (ep:Episodic)
WHERE ep.source_description CONTAINS '[task-completed'
OPTIONAL MATCH (ep)-[r:MENTIONS]->(e:Entity)
RETURN ep, r, e
LIMIT 50;Drag the Episodic nodes (they're a distinct label from Entity) to one side and the Entity nodes to the other — the resulting bipartite layout shows what each task completion contributed to the project brain.
# Via the slash command (interactive, recommended):
/pb-graphiti:wipe-channel claude-self-writes
# Or via cypher directly:
MATCH (ep:Episodic) WHERE ep.group_id = '<your-group>'
AND (ep.source_description STARTS WITH 'claude-code-session://'
OR ep.source_description STARTS WITH 'claude-code-conversation://')
DETACH DELETE ep;Open the Neo4j Browser at http://localhost:7474 (login: neo4j / your NEO4J_PASSWORD). All queries below scope by group_id — change 'lcd-mageos' to your project id (or 'fleet', 'initial_ingest').
Run these in the browser command bar — they persist as Browser settings:
:config initialNodeDisplay: 25
:config maxNeighbours: 20
:config maxRows: 500
Without these the Browser silently truncates at 1000 nodes on result graphs and 1000 rows on tabular returns. If your project has 10,000+ nodes, you'll see incomplete pictures and not realize it.
MATCH (n) WHERE n.group_id = 'lcd-mageos'
RETURN labels(n) AS type, count(*) AS count
ORDER BY count DESC;The full project graph is too big to render — start with a small slice and expand by double-clicking nodes in the Browser:
// Top 25 most-connected entities (graph "hubs") for this project
MATCH (e:Entity {group_id: 'lcd-mageos'})
OPTIONAL MATCH (e)-[r]-()
WITH e, count(r) AS connections
ORDER BY connections DESC
LIMIT 25
RETURN e;// Most-recent 30 episodes — pairs with their extracted entities
MATCH (ep:Episodic {group_id: 'lcd-mageos'})
WITH ep ORDER BY ep.created_at DESC LIMIT 30
OPTIONAL MATCH (ep)-[r:MENTIONS]->(e:Entity)
RETURN ep, r, e;// All Vendors
MATCH (e:Entity:Vendor {group_id: 'lcd-mageos'})
RETURN e.name AS name, e.summary AS summary
ORDER BY e.name;
// All Decisions with their rationale
MATCH (e:Entity:Decision {group_id: 'lcd-mageos'})
RETURN e.name, e.summary
ORDER BY e.created_at DESC;
// All Incidents
MATCH (e:Entity:Incident {group_id: 'lcd-mageos'})
RETURN e.name, e.summary, e.created_at
ORDER BY e.created_at DESC;Available entity labels (from infra/config/config.yaml entity_types): Vendor, Decision, Incident, Project, Client, Procedure, Component, Preference, Topic. Plus the generic Entity for anything that didn't get a specialized type.
MATCH (e:Entity {group_id: 'lcd-mageos'})
WHERE toLower(e.name) CONTAINS 'tax'
RETURN labels(e) AS labels, e.name, e.summary
LIMIT 20;MATCH (e:Entity {group_id: 'lcd-mageos'})
WHERE e.name = 'Honeycomb'
OPTIONAL MATCH (ep:Episodic)-[:MENTIONS]->(e)
RETURN e.name AS entity, e.summary AS summary,
collect({name: ep.name, source: ep.source_description, when: ep.created_at}) AS episodes;MATCH (ep:Episodic)
WHERE ep.created_at >= datetime() - duration({hours: 24})
RETURN ep.group_id, ep.name, ep.source_description, ep.created_at
ORDER BY ep.created_at DESC;// All GitHub ticket episodes from lcd-mageos
MATCH (ep:Episodic {group_id: 'lcd-mageos'})
WHERE ep.source_description STARTS WITH 'https://github.com/'
RETURN ep.name, ep.source_description, ep.created_at
ORDER BY ep.created_at DESC
LIMIT 50;// What is "ProxiBlue" connected to in lcd-mageos?
MATCH (start:Entity {group_id: 'lcd-mageos', name: 'ProxiBlue'})-[r1]-(neighbor)-[r2]-(second)
WHERE second.group_id = 'lcd-mageos'
RETURN start, r1, neighbor, r2, second
LIMIT 50;In the Browser this renders as a small subgraph — drag, expand, explore.
// Episode size distribution per project — useful when chasing token cost
MATCH (ep:Episodic)
WITH ep.group_id AS gid, size(ep.content) AS chars
RETURN gid,
count(*) AS episodes,
sum(chars) AS total_chars,
avg(chars) AS avg_chars,
max(chars) AS max_chars
ORDER BY total_chars DESC;// Top 10 biggest individual episodes (candidates for truncation review)
MATCH (ep:Episodic)
RETURN ep.group_id, ep.name, size(ep.content) AS chars, ep.source_description
ORDER BY chars DESC
LIMIT 10;Neo4j Browser caps rendered graphs at 1000 nodes by default. If you run MATCH (n {group_id: 'pvcpipesupplies'}) RETURN n and see exactly 1000 nodes back, that's the cap — your actual count may be higher. Either tighten the query (filter by label, time, name pattern) or raise :config maxRows and accept slower render.
For programmatic use (the HTTP API at /db/neo4j/tx/commit), there's no implicit cap — LIMIT is whatever you write.
pb-gitnexus— structural code graph (gitnexus) for Magento / Mage-OS. Pairs well: gitnexus = code structure, pb-graphiti = domain knowledge.- Built-in Claude Code auto-memory — keep using it for hard rules.
This plugin (pb-graphiti) is Apache License 2.0. The wrapper code, ingest scripts, slash commands, hooks, and configuration are all original work by Lucas van Staden / Proxiblue (https://www.proxiblue.com.au), redistributable under Apache 2.0.
Attribution is required. Redistributors and forks MUST preserve the LICENSE, NOTICE, and copyright notices, and MUST state significant modifications in modified files (per Apache 2.0 §4). Removing attribution or rebranding the work as your own — without keeping the original NOTICE — is a license violation.
The plugin orchestrates and depends on third-party software that is NOT redistributed in this repository — consumers fetch and run those components themselves:
| Component | License | How this plugin uses it |
|---|---|---|
| Graphiti (graphiti-core + MCP server) | Apache 2.0 © Zep AI, Inc. | The plugin's bootstrap instructions ask consumers to git clone https://github.com/getzep/graphiti upstream themselves; the docker-compose builds an image locally from that clone. No Graphiti source is bundled here. |
| Neo4j Community Edition (5.26.0) | GPLv3 (Community); commercial for Enterprise | Pulled as the upstream neo4j:5.26.0 Docker image at run time; not redistributed by this repo. The plugin only uses Community-tier features (no online backup, no clustering, no read replicas). |
| Anthropic Claude API | Commercial — Anthropic Usage Policy | Used as the LLM backend for entity extraction via the consumer's own API key. Not redistributed. |
| Voyage AI embeddings | Commercial — generous free tier | Used as the embedder via the consumer's own API key. Not redistributed. |
Apache 2.0 compliance for our dependencies: because Graphiti is fetched by the consumer (not bundled here), we are NOT redistributing it and the Apache 2.0 redistribution obligations (include LICENSE, NOTICE, modification disclosures) don't apply to this repository for Graphiti. If you fork this plugin and start bundling Graphiti's source directly (e.g., as a git submodule with shipped binaries), you become a redistributor of Graphiti and MUST include Graphiti's LICENSE and NOTICE files in your distribution and disclose any modifications to Graphiti's source.
GPLv3 on Neo4j Community: the plugin invokes Neo4j over its network protocol (Bolt) and HTTP API. Network-protocol use is not "linking" under the GPL — the plugin does not statically or dynamically link to Neo4j code. Consumers run Neo4j as an independent process via Docker. No GPLv3 obligations propagate to this Apache-2.0-licensed code.
Apache 2.0 for the plugin itself: see LICENSE and NOTICE. Apache 2.0 permits use, modification, sublicensing, and redistribution (including commercially), with attribution preserved via the LICENSE and NOTICE files. The license is compatible with depending on Apache 2.0 (Graphiti) and GPLv3 (Neo4j) software via the documented orchestration pattern.
- Use this plugin in any environment, commercial or non-commercial.
- Modify it for your own needs (internal or shipped).
- Embed it in proprietary stacks.
- Distribute it, redistribute modified versions, or fork the repository.
- Keep the
LICENSEandNOTICEfiles in any distribution (including forks). - Preserve the copyright notice in source files that already carry one.
- State significant modifications prominently in any modified files you redistribute.
- Do not use "Proxiblue" or "Lucas van Staden" to endorse derivative works without permission (separate from the trademark clause in Apache 2.0 §6).
If you spot a license compliance gap, open an issue at https://github.com/ProxiBlue/pb-graphiti/issues.
Apache License 2.0 — see LICENSE and NOTICE.
Copyright (c) 2026 Lucas van Staden / Proxiblue.