XE Local AI Engine is the node-side runtime for running local AI workloads while preserving the existing C0re platform contract. The Node Web Server hosts the React management UI, owns the platform
WorkerHub connection, and supervises node-owned llama-server host child processes for local inference.
The current source version is 1.0.0-rc.1, composed in eng/ReleaseVersion.props. Release documentation and
validation evidence live in this repository and must stay current with runtime behavior.
Just want to install and use the app? Start with the User Guide — download, install (Windows & Linux), first run, troubleshooting, and privacy, all in plain language. App downloads are on the Releases page.
Official binaries are portable-only: Windows ships a Velopack Portable.zip with no Setup.exe, and Linux ships a
Velopack AppImage rather than a ZIP. Both formats are self-updating. Release assets are currently unsigned because no
signing certificate exists; verify CHECKSUMS.sha256 and review RELEASE-MANIFEST.json / RELEASE.spdx.json before
running them. Signing is planned.
- Node Web Server (
XE-Local-AI-Engine.Client) — serves the React UI, local APIs under/api/local/v1, local SignalR hubs, SQLite-backed chat state, and the existing platformWorkerHubconnection. - React management UI (
XE-Local-AI-Engine.Client.React) — node-local browser UI for chat, settings, runtime status, logs, and models. - Providers and agents — local provider abstractions, the supervised llama.cpp host runtime (primary/default), Ollama provider (opt-in secondary), and shared agent execution loop.
- Scheduler — Quartz.NET-backed job scheduler with job definitions, run history, cancellation, and live run updates over a local SignalR hub (
Services/Scheduler,src/features/scheduler). - Model-fit / Model Advisor — box-aware GGUF recommendation: profiles the local hardware (RAM / VRAM / GPU vendor), discovers candidate GGUF repos on Hugging Face, estimates each model's memory footprint with a
pure, in-process (I/O-free) formula, and ranks the ones that fit. Exposed as cache-only reads plus a scheduler-driven refresh — the advisor is estimator-only and never spawns a process, so there is no container
or benchmark image anywhere in this path (
Services/ModelFit,src/features/model-fit). See docs/wiki/07-model-fit.md. - Image generation — local text-to-image via stable-diffusion.cpp: the node supervises a resident
sd-serverchild process (one daemon per model, readiness-gated, idle-evicted on its own loopback port range), serializes generation to one job at a time with queue/cancel, and persists produced images encrypted-at-rest. Ships enabled by default (Services/Images,Providers.StableDiffusionCpp,src/features/images). See docs/wiki/14-image-generation.md. - Knowledge Base / RAG — fully offline document knowledge base: upload documents (
.txt/.md/.pdf/.docxand other plaintext types), which are chunked, embedded with a local embedding model, and indexed into local SQLite with selective encryption — source document blobs and display names are encrypted at rest, while the extracted chunk text and its FTS search index are stored unencrypted locally. Retrieval is a hybrid search — lexical FTS5 + semantic vector arms fused with Reciprocal Rank Fusion, with an optional local cross-encoder reranker — surfaced to agents as a tool (Services/Knowledge,Endpoints/Knowledge/V1,src/features/knowledge). See docs/wiki/15-knowledge-base.md. - Agent mode — per-agent definitions plus a governed playbook: manual and analysis-proposed actions, an offline eval gate over golden conversations, relevance-gated action retrieval, and cohort
monitoring (
Services/{Agents,Eval,Insights,Monitoring},XE-Local-AI-Engine.AI.Agent,src/features/agents). - Custom tools — operator-authored HTTP fetches and direct host-program launches that can be assigned to agents. The node-wide feature switch is off by default; the built-in authoring UI initializes new definitions as disabled, while the API persists an acknowledged caller's requested enablement. Every call remains approval-wrapped: a fixed tool may reuse an explicit, version-bound session approval, while a parameterized tool re-prompts for every model-selected argument set. Authoring validates model-visible schemas, executable paths, template placeholders, and HTTP host allow-lists; secret headers and environment values are encrypted at rest and masked on reads (
Services/CustomTools,Endpoints/CustomTools/V1,src/features/customTools). - Development Mode — a default-on, node-local coding workflow with engine-owned detached Git worktrees, deterministic validation, independent review, hash-bound evidence, and explicit final host apply.
The operator registers a trusted local Git repository once, then selects it by an opaque ID and alias; the host path stays internal to the node. The agent works in a managed worktree outside the selected source
repository, and only a reviewed apply whose base and evidence hashes still match may change that source. Generated source, MSBuild targets, source generators, and tests execute as the host user with the host's
filesystem and network access. The Process sandbox and Agent Home controls constrain application-mediated paths and bytes; they are not an operating-system security boundary. Set
Development:Enabled=falseas an emergency switch when that execution posture is not intended. MXC remains future provider work. ADR 0004 (accepted 2026-07-29) moves this feature's execution onto a Docker container provider, behind the same provider seam and as an interim step ahead of MXC. That provider has shipped and is opt-in: setDevelopment:Sandbox:Provider=dockerto select it. Leave it unset — as the shipped configuration does — and Development Mode keeps running on the process provider exactly as described above. On a node that does select it, a running Docker daemon is a hard requirement — there is deliberately no unisolated fallback, so a machine without one gets no Development Mode rather than a quietly weaker one. Docker stays scoped to this feature: chat, embeddings, model acquisition and image generation never require it. See Development Mode container implementation status for the maintained record of what is implemented. - MCP tool extensibility — registered MCP servers whose live tool snapshots are offered to agents through the local tool registry (
Services/Mcp,src/features/mcp). - Tests and fixtures — backend/client persistence tests, integration-style tests, E2E harness, and FakeOllama in-process test server.
- Only the Node Web Server talks to the C0re platform over
WorkerHub. - Worker credentials, cloud-provider credentials, and external endpoint tokens stay local and must not be returned to the browser or written to logs/transcripts.
- Local admin endpoints must be loopback/local-only, authenticated, strict about
Host/Origin, and secret-redacted. - Any future installer or packaging effort must not create background autostart behavior unless a new approved plan changes that contract.
Using the app (non-developers): the User Guide covers download, install, first run, troubleshooting, privacy, and a plain-language glossary.
The contributor deep-dive lives in the Developer Wiki — code-grounded
pages covering architecture, every project, the local llama.cpp runtime and providers, agent mode,
chat, scheduler, model-fit, data/persistence, the API surface, the React client, hosting/deployment,
security/privacy, and testing. Start at docs/wiki/Home.md.
For a baseline-scoped external review, use the
Technical/Security Architecture Dossier.
It describes the implementation at commit 7e64ed589e14eecc0e522e807d2e531a1095d19a as reviewed on
2026-07-28. It is not a certification, compliance mapping, penetration-test report, or operating-
effectiveness assurance package; each chapter labels evidence availability and known gaps.
Supporting notes:
- Architecture Decision Records — repository design decisions and their code-level scope.
- AI runtime developer notes — narrow AI-seam maintenance rules (see the wiki for the full runtime architecture).
- Backend commentary map
Component-specific notes:
- .NET SDK from
global.json - Node.js compatible with
XE-Local-AI-Engine.Client.React/package.json - pnpm via Corepack or a local install
- Python 3 for repository validation and lifecycle scripts
- The Aspire CLI for AppHost development and readiness checks
- On Linux/WSL,
setsid(normally provided byutil-linux) for transactionalscripts/dev-start.shcleanup - A GPU with current drivers is optional; the app self-provisions its llama.cpp runtime and GGUF models at first run. (Neither Docker nor Ollama is required to build or run the engine — llama.cpp is the local runtime and inference needs only a driver; see docs/wiki/03-local-runtime-and-providers.md.)
- Docker: not required by default. The container provider for Development Mode only (ADR 0004) has shipped, but it is opt-in and off unless you configure it: it activates only when you set
Development:Sandbox:Provider=docker, and the shippedappsettings.jsonleaves that key unset. On a node that does set it, Development Mode needs a running daemon, plus its data root inside the WSL2 filesystem on Windows; the real-daemon integration tests need one too (without a daemon they report as blocked or skipped-with-reason, never as a pass). Nothing else in the app gains a Docker dependency. See Development Mode container implementation status.
From the repository root:
scripts/with-build-lock.sh -- dotnet restore XE-Local-AI-Engine.slnx
scripts/with-build-lock.sh -- dotnet build XE-Local-AI-Engine.slnx --configuration Release --no-restore
scripts/with-build-lock.sh -- scripts/assembly-guard.sh guard --test-bins -- \
dotnet test XE-Local-AI-Engine.slnx --configuration Release --no-build --max-parallel-test-modules 1The lock prevents cooperating builds from rewriting test assemblies mid-run; the assembly guard
detects an unwrapped concurrent build. Exit 69 means the lock was not acquired and nothing ran.
Exit 75 means the result was CONTAMINATED and void—rerun it rather than treating it as red or
green.
For the React client:
cd XE-Local-AI-Engine.Client.React
pnpm install --frozen-lockfile
pnpm run lint
pnpm test
pnpm run buildAfter any backend contract change, run pnpm openapi:check from XE-Local-AI-Engine.Client.React/ — it
regenerates the hey-api client and fails on drift. See AGENTS.md for the full
validation command set, including analyzer requirements, build-lock/assembly-guard usage, and the
backend/frontend test suites.
E2E validation is ask-gated because it may require browser/runtime setup:
scripts/run-e2e-local.shUse Aspire for local development and integration checks.
scripts/dev-start.sh # always --isolated and scoped to this worktree's AppHost
scripts/dev-status.sh # filtered status; secrets and dashboard tokens are omitted
scripts/dev-stop.sh # stops only this worktree's registered AppHost
# Bounded integration smoke; refuses to reuse an existing instance and always cleans up its own.
scripts/aspire-readiness-smoke.shThese wrappers make parallel worktrees safe. Do not use aspire stop --all; it crosses checkout
boundaries. See scripts/README-dev-stop.md for the Aspire 13.4
fallback and cleanup contract.
dev-start.sh also owns the node operator secret. On first use it mints a per-checkout, owner-only
XE-Local-AI-Engine.AppHost/.data/node.key (never tracked) and passes it to Aspire's required
node-sqlite-key parameter; later runs reuse it, so encrypted dev data stays readable. If it mints a
key next to dev data written under a different secret, it says so and names what to delete — that
data cannot be decrypted and the node will otherwise crash on the first read.
The desktop package is deliberately asymmetric. Linux remains a self-contained single-file AppImage. Windows is a framework-dependent Velopack Portable ZIP: it contains a small C# launcher apphost plus the managed application DLL, but no .NET runtime. Windows users install the x64 ASP.NET Core Runtime 10.0.10 or a newer .NET 10 servicing patch. If the base .NET runtime is absent, Microsoft's apphost reports the missing framework; if ASP.NET Core is absent or too old, the launcher prints the exact requirement and opens the official .NET 10 download page.
Both packages run as double-click desktop apps: a console window opens with live logs, the default browser opens on the
running site, and closing the console window shuts the whole app down — including the spawned llama-server child,
so there is no orphan process. Closing the browser does not stop the app.
Desktop mode is opt-in via the launcher (env XE_LAUNCH_MODE=desktop or the --desktop flag); headless, Aspire,
and CI runs are unaffected. In desktop mode the host binds HTTP on a free loopback port (127.0.0.1) and skips the
HTTPS-redirect/HSTS pipeline (traffic never leaves the loopback adapter).
Build the React app first; the publish target rejects a missing dist/index.html:
# Web assets (once before either RID)
(
cd XE-Local-AI-Engine.Client.React
pnpm install --frozen-lockfile
pnpm run build
)
# Linux: self-contained single-file payload
dotnet publish XE-Local-AI-Engine.Client -c Release -r linux-x64 -p:PublishProfile=linux-x64
# Windows: framework-dependent app plus the C# launcher, overlaid into one payload directory
WIN_OUT="$PWD/.tmp/publish/win-x64"
dotnet publish XE-Local-AI-Engine.Client -c Release -r win-x64 -p:PublishProfile=win-x64 --output "$WIN_OUT"
dotnet publish XE-Local-AI-Engine.WindowsLauncher -c Release -r win-x64 -p:PublishProfile=win-x64 --output "$WIN_OUT"The Linux profile sets SelfContained=true, PublishSingleFile=true, and
IncludeNativeLibrariesForSelfExtract=true. The Windows application and launcher both set SelfContained=false; only
the launcher sets UseAppHost=true, producing the MIT-licensed Microsoft apphost while keeping coreclr, hostfxr,
and the runtime out of the artifact. Trimming stays off for the reflection-heavy application.
Start the matching published entry point:
- Linux:
publish/linux/run-xe-local-ai-engine.sh—execs the binary in the foreground so the terminal owns it (terminal close →SIGHUP→ graceful shutdown). - Windows: run
XE-Local-AI-Engine.WindowsLauncher.exefrom the combined payload. It validates the adjacent files and ASP.NET Core runtime, sets desktop mode, forwards Velopack arguments, launches the managed DLL, and propagates its exit code.publish/windows/run-xe-local-ai-engine.cmdis retained only for the deprecated self-contained manual flow.
See publish/README.md for the expected layout. Run one instance at a time against the same
user-data directory — a second instance races on the SQLite database.
The tag-triggered .github/workflows/release.yml is the only official release
path. It rejects a tag/version/source mismatch, runs the shared validation workflow, lets the Windows and Linux matrix
jobs build and retain assets only, then splits publication into two protected write transactions. The serialized
prepare-release-draft job creates the draft, merges both Velopack channels, verifies the remote bytes, attaches
detached SPDX/release-manifest/checksum evidence, and re-verifies the complete remote draft after the
open-source-release environment authorizes repository write access. A separately approved publish-release job
re-verifies that same draft, promotes it without rebuilding or replacing any asset, and confirms both public feeds
anonymously.
Windows packing uses Velopack 1.2.0 with --noInst, producing a managed Portable.zip plus feed/full/delta assets and
no Setup.exe. Linux produces a managed AppImage. Public update checks require no GitHub device login or access token:
the main flavor follows stable releases, the tester flavor includes release candidates, and Velopack selects the
independent Windows/Linux OS channel from package metadata.
publish/package-tester-win.ps1 and publish/package-rc.sh are deprecated, reference-only scripts. They describe
superseded manual distribution flows and are not publication alternatives. scripts/lint-release-scripts.sh still
analyzes them so retained reference code does not decay silently.
A
win-x64zip frompackage-rc.shis cross-built on Linux. Smoke-test it on real Windows before handing it to anyone — native-library self-extraction, console-close child cleanup, and browser auto-open cannot be verified off-Windows. The same applies to the two desktop invariants below.
The no-orphan design (terminal/console close reaps
llama-server) and the Windows Job Object path require real-desktop verification with a model loaded; they cannot be exercised in WSL2 or on a headless runner. This baseline documentation review does not include or assert availability of the matching smoke-test transcript.
See docs/velopack-release-install-guide.md for the full release and
update-channel story.
Do not mark release or documentation work complete until matching validation evidence is available. The checklist below defines required release evidence; its presence here does not assert that the evidence was produced, retained, or made available for the documentation baseline.
Required evidence includes:
- the release workflow's (or, for a manual rehearsal, the deprecated packager's) frontend, backend, vulnerability, and package-gate transcript,
- a clean default
scripts/lint-release-scripts.shresult, including its mandatory Pester suite, - a non-vacuous Playwright E2E run (
scripts/run-e2e-local.sh) with no exit-75 contamination, - a passing live GPU smoke run (
scripts/run-gpu-smoke-local.sh) on a GPU box — the only gate that proves the GPU did the work, since a CPU fallback answers correctly, just slowly; treat exit 5 as an infrastructure abort where nothing was judged, not a product failure, - generated schema/sample-manifest validation, including a clean
openapi:check, - pinned runtime binary and package checksums,
- the matching
v<version>source tag on the exact packaged commit, - a real-Windows smoke-test transcript for the exact generated
Portable.zipand a Linux AppImage smoke test, - the generated release assets and their checksums, pushed source-tag verification, and
- confirmation that the release was published to this repository's GitHub Releases.
Run scripts/lint-release-scripts.sh. The Pester suite is part of that default run, not an add-on —
--pester only requests it explicitly, and a missing Pester module is a hard failure, never a silent skip
(a skipped test suite must never read as a pass). See Testing & Validation
and the release guide for the full sequence.
Standalone OS installers/packages (MSI/DEB/RPM) remain deferred. The official distribution is the Velopack-managed Windows Portable ZIP and Linux AppImage.
XE Local AI Engine is licensed under Apache-2.0. See LICENSE and NOTICE.