Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

420 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Pulsar

pulsar

An inference engine for giant Mixture-of-Experts models on hardware that has no business running them. The routed experts live on NVMe and stream per token; everything that makes decisions stays resident in VRAM. No llama.cpp anywhere in the stack.

Successor to NeutronStar, rebuilt as its own engine in Rust + CUDA instead of a C fork. A pulsar is a neutron star that spins fast and emits beams.

What it does today

Eleven model architectures running on consumer GPUs: Hy3 295B (hy-v3, GQA), GLM-5.2 743B (glm-dsa, MLA + DSA sparse attention), Kimi K2.7 1T (deepseek2, MLA + YaRN), MiniMax M3 (partial rotary, swiglu_oai), Gemma 4 26B-A4B (interleaved sliding-window attention, dual GELU FFN), TML Inkling 1T (no rope, learned relative-position bias, shortconv streams, sink router; supported the day after release), Qwen3-235B-A22B (qwen3moe, softmax router; correct output on its first-ever run), DeepSeek-V4-Flash 284B (deepseek4: 4-stream hyper-connection residual with Sinkhorn gates, sink attention over a sliding window plus streaming compressed KV, fp8/fp4 cache quantization-aware sims, token-id hash routing on the early layers; also correct output on its first-ever run), and Qwen3.6-35B-A3B (qwen35moe hybrid: Gated DeltaNet linear attention with O(1) recurrent state on 3 of every 4 layers, sigmoid-gated full attention on the rest - 262k context with KV on only 10 of 40 layers, needle recall verified at 45k tokens with 19.6 tok/s decode at that depth; prefer K-quants for it: the Q4_K_XL decodes at 51.8 tok/s where the smaller Q3_K_XL manages 36, because iq3's codebook lookups are decode compute the simple K-quant shifts don't pay), and gpt-oss 20B (gpt-oss: per-head attention sinks in the softmax denominator, alternating sliding/full attention, biases on both the attention projections and every expert, MXFP4 routed experts), and Laguna-S-2.1 118B (laguna: hybrid attention with a full-window layer every fourth and sliding-window 512 elsewhere, a per-head output gate, and per-layer-type RoPE — YaRN on the full-window layers, plain on the sliding ones; the imatrix IQ2_XXS build decodes faster than the Q4_K_M because more experts stay resident at 36GB), and Ornith-397B (a qwen35moe hybrid at 397B: Q2_K experts on a 512-way sigmoid-gated router with a shared expert, Gated DeltaNet linear attention as above; runs end-to-end on a single 32GB GPU after a load-site fix — see docs/ornith-q2k-cuda-crash-fix.md). Reference box: RTX 5060 Ti 16GB + RTX 4060 Ti 16GB, Ryzen 9900X, 30GB RAM, one Gen5 NVMe.

Model Total Active / token gguf Decode, warm vs ds4, same box
Gemma 4 26B-A4B 26B 4B 16GB (Q4_K_XL) 41 tok/s
gpt-oss 20B 20B 3.6B (top-4 of 32) 12GB (Q8_0 attn + MXFP4 experts) 21.2 tok/s
Qwen3.6-35B-A3B 35B 3B (top-8 of 256 + shared) 22GB (Q4_K_XL) 51.8 tok/s
ThinkingCap-Qwen3.6-27B (dense) 27B 27B 16GB (Q4_K_M) 18.7 tok/s (27.8 w/ nextn MTP)
Laguna-S-2.1 118B 8B (top-10 of 256 + shared) 36GB (IQ2_XXS, imatrix) 17.3 tok/s (22.4 w/ CPU lane)
DeepSeek-V4-Flash 284B ~8B (top-6 of 256 + shared) 87GB (ds4 recipe) 8.2 tok/s (11.3 w/ CPU lane)
Hy3 295B 295B 21B (top-8 of 192) 79GB (IQ2_XXS) 6.0 tok/s (6.9 w/ CPU lane) 0.64–0.70
Qwen3-235B-A22B 235B 22B (top-8 of 128) 83GB (Q2_K_XL) 5.3 tok/s (6.4 w/ CPU lane)
MiniMax M3 428B 23B 134GB (Q2_K_XL) 5.0 tok/s (5.9 w/ CPU lane)
GLM-5.2 744B 40B 211GB (ds4 recipe) 2.7 tok/s (2.8 w/ CPU lane) 0.40
TML Inkling 975B 41B (6 + 2 shared) 296GB (Q2_K_XL) 1.6 tok/s (1.75 w/ CPU lane)
Kimi K2.7 Code† ~1T 32B 339GB (Q2_K_XL) 1.3 tok/s
Ornith-397B†† 397B ~38B (top-10 of 512 + shared) 149GB (Q2_K exp + Q8_0) not benched

All figures are sustained warm decode at n=64, temp 0, second run onward. The resident tier is placed from the popularity census, which builds over the first full run, so measure with a warm census. Shorter generations read higher because the per-token SSD miss rate is still climbing to steady state: Hy3 does 8.2 tok/s at n=32 versus 6.0 at n=64. Gemma is small enough that its whole quantized weight set lives resident in the tier, so warm Gemma is compute-bound, not streaming-bound.

† Measured before the n=64 standardization and not yet re-run (model deleted to free disk); the sustained rate is likely a little lower than shown, for the same reason Hy3 reads 8.2 at n=32 against 6.0 at n=64.

†† Ornith-397B was brought up on a different box (3090 + V100 32GB GPU, 128Gb), not the reference 2×16GB rig, and only a short cold-census 15-token run exists (-p "The capital of France is" -n 15 → "Paris.", correct, 0.82 tok/s cold). That is a bring-up smoke test, not a benchmark: n=15 off a cold census with the disk as the floor. The rate column stays empty until scripts/bench.sh has run on it, per the rule that every number in this table comes from that script and nowhere else. The 149GB gguf size is the experts-Q2_K / rest-Q8_0 mix from pulsar-quant --map "_exps.=q2_k" --default q8_0.

ThinkingCap-27B is pulsar's first fully-dense arch and runs a different mode entirely: no streaming, no tiers - the model fits across both cards, so each layer's whole stack (attention or Gated DeltaNet, KV, FFN triple) is resident on ONE owner card in native K-quant, evaluated with warp-cooperative Q4_K/Q6_K matmuls on q8_K activations, and the residual stream crosses cards twice per 16-token chunk. Speculative decode rides the model's own nextn/MTP layer (PULSAR_MTP=1, depth via PULSAR_MTP_DEPTH, 85% acceptance at depth 3 on greedy); verify rounds snapshot the recurrent GDN state, since unlike KV rows a delta-rule state can't be overwritten after a rejected draft. The port went 9.7 -> 18.3 tok/s base (27.5 with MTP) in one arc: per-layer card ownership, K-quant-native attention, a warp token tile so verify/prefill rows share one weight read (ncu: L1-bound at 94%, DRAM 28% - the next lever is the int8-MMA path the MoE verify unions already use).

DeepSeek-V4-Flash runs its state machines (sliding-window ring, streaming KV compressor, Sinkhorn hyper-connection gates) fully on device and prefills in batched 16-token chunks (~26 tok/s prefill with the CPU lane at shallow positions; past ~2K context the indexer top-k engages and the batched path used to fall back to single-token steps, decaying prefill to decode speed ~8 tok/s. Deep chunks now stay batched with per-token visibility masks - a 3.4K prompt drops 249s to 137s, larger prompts save proportionally more. The float delta vs single-stepping that briefly kept this opt-in was root-caused to matmul_q8_0's batch-size kernel dispatch (dp4a vs int8-MMA accumulate in different orders, tolerance-tested at 1e-3) - the same documented drift class as the tiers and grouped MoE, present in every chunked prefill. PULSAR_NO_DEEP_BATCH=1 restores exact single-stepping. one router readback and one expert union per chunk-layer, with the per-token ring/compressor/attention interleave preserved bit-exactly). Long-context retrieval verified by needle recall at 2.4k ctx through compressed rows.

Long context on the GQA models decodes through a split-K attention kernel: past 4k visible rows the position scan fans out across blocks (2k rows per split, unnormalized online-softmax partials, one combine pass) instead of serializing inside a single block per head. Measured on Qwen3.6 at 45k tokens: needle recall correct, decode 3.7 to 19.6 tok/s (5.3x); at 10.8k tokens recall is also verified. Short contexts take the original kernel unchanged, so bit-exact gates and the 51.8 tok/s bench are untouched. Known remaining lever: each q-head group re-reads its shared kv-head rows, so a kv-head-centric layout has several-fold headroom left at depth.

A CPU expert lane (opt-in, PULSAR_CPU=1) computes host-cache-hit experts on the CPU instead of uploading them: AVX2 iq2_xxs and q2_K x q8_K kernels sustain 42 GB/s across the 9900X's cores, above the 28.7 GB/s the same bytes would cost crossing PCIe, and the dots overlap the GPU resolve. Host-cached experts stop competing for upload bandwidth and VRAM cache slots, so both effects compound: DeepSeek-V4-Flash measures 8.2 to 11.3 tok/s, Hy3 6.0 to 6.9, GLM-5.2 2.7 to 2.8 (re-measured 2026-07-19 after an iq2_xxs correctness fix: the CPU dot had been indexing the encoder-unit grid instead of the dequant lattice, scaling every lane partial by ~1/9 per dot - GLM surfaced it as repetition loops, and a one-row GPU-vs-CPU arbiter now pins the dot to the kernel bit-for-bit-scale; earlier lane numbers were measured with the broken dot and are superseded by these). GLM's pair moved again on 2026-07-25 when a prefill prefetch bug was fixed, and the lane's margin there narrowed from +41% to +6%: much of what the lane had been buying on that model was insulation from the flood's cache thrashing, not upload bandwidth. Covers iq2_xxs, iq2_xs, iq3_xxs, q2_K, q3_K and q4_K expert tensors, which spans the ds4 recipes and the UD-Q2_K_XL mixes: Qwen3-235B 5.3 to 6.4 (+21%), TML Inkling 1.63 to 1.75 (+7%), MiniMax M3 5.0 to 5.9 (+18%; its IQ mix engages on 54 of 57 MoE layers, the three iq4_xs-down layers stay on the GPU). Baselines for dsv4/Hy3/M3 re-measured 2026-07-19 with the triple-aware warm load (PR #2); 235B and Inkling predate it (ggufs rotated off disk).

Decode rate slides with output length on the streaming models: a longer generation routes to a wider set of experts, so the disk-miss fraction creeps up until the working set saturates. Hy3's length scan (measured before the triple warm load; the shape holds, the levels read ~10% low now): 5.7 tok/s at n=64, 4.5 at n=128, 4.2 at n=256, converging toward a ~4 tok/s floor set by how much of the expert working set fits in host RAM (more RAM lifts the whole curve). n=64 is the reported standard; long outputs run nearer the floor. Gemma is exempt: its weights are fully resident, so no disk is in the loop.

Prefill runs the quantized weights through int8 tensor cores: Hy3 28 tok/s (1.8× over dp4a, ds4 0.44), GLM-5.2 15 tok/s (2.7×). Warm start: hot experts bulk-load in ~3s. (ds4 = NeutronStar, the llama.cpp-fork predecessor, on the same box.)

Decode figures are warm-run (second run onward). The first run is cold while the expert-popularity census fills; only after it is written do the host cache and resident tiers load hot, so a cold run reads far more from disk and clocks lower, so don't benchmark the first run. On the reference box Gemma 4 goes 28.7 tok/s cold → 41 tok/s warm (hot experts resident on the second GPU). See the warm-start note under Quick start.

Prefill runs the quantized weights through int8 tensor cores on sm_80+ (mma.m16n8k32 dense GEMM + mmq-style grouped MoE that unpacks each expert superblock to shared memory once per prefill chunk and rescales per quant block in registers), 1.8–2.7× over the dp4a kernels, which remain the path on older GPUs. Decode is single-token and memory-bound, so it is deliberately untouched: ids stay bit-identical to the dp4a path.

GLM runs contexts past its naive 2048-row ceiling via a port of the DSA lightning indexer (top-k row selection per token), validated against the reference engine with a long-context retrieval probe. The indexer's batch scorer runs on tensor cores (f16 keys, m16n8k16 with the relu-weight epilogue fused between heads): 1.9x long-prompt prefill at 4k context, byte-identical ids vs the scalar path, and the index K cache halves to f16 (the reference indexer ships FP8 in production).

On a single RTX 4060 Ti (where NeutronStar set its numbers): Hy3 2.6, GLM 0.56.

Zero-config multi-GPU. At startup pulsar measures each card's H2D bandwidth (labels lie: an x8-labeled slot can train x1, a driver bug can park a Gen5 card at Gen1, only a measurement sees that) and assigns roles by what each card is actually good at:

  • Expert streaming needs link bandwidth → the fastest measured card.
  • Attention residency (MLA models: the whole ~14GB attn stack + KV parked on a second card) only needs capacity, weights cross the bus once at load, then only activations hop (2× 24KB per layer). A bandwidth-crippled card serves attention at full speed.
  • Expert tiers: leftover cards are filled with the hottest expert triples from the warm census, and the MoE kernels run on the card that holds the weights, partial outputs gather back over PCIe. On the reference box the tier serves ~90% of expert computations and nearly doubles Hy3 decode.

Correctness is certified against ds4, not assumed: teacher-forced along ds4's greedy path (15/16 per-position argmax agreement on Hy3, 10/12 on GLM, every miss at a <0.09-logit tie), byte-identical greedy ids across single-GPU vs attn-offload configurations, and bit-exact decode determinism on a fixed code path (--decode-consistency, below).

Requirements

  • Linux (io_uring and CUDA are load-bearing; the workspace compiles on macOS but the engine is stubbed out there)
  • One or more NVIDIA GPUs, GTX 10-series (Pascal, sm_61) or newer, the default build ships native code for 10/16/20, 30, and 40-series plus PTX that JITs on everything else (50-series Blackwell, Volta, Hopper). PULSAR_CUDA_ARCH overrides codegen targets
  • CUDA toolkit with nvcc on PATH, plus a host compiler nvcc accepts (gcc-12 works; newer gcc may need CXX=g++-12 at build time)
  • Rust via rustup
  • The model gguf on a fast NVMe, streaming reads it at up to ~7GB/s, so the disk is the decode speed
  • ~16GB system RAM for the host-side expert cache (more helps; the cache budget is the single biggest knob after the disk)

Get a model

Pulsar reads standard llama.cpp ggufs: ten routed-expert quant formats (q2_K, q3_K, q4_0, q4_K, q5_K, q5_1, q6_K, iq2_xxs, iq2_xs, iq3_xxs, including fused gate_up tensors and non-256-multiple expert widths), K-quant dense tensors (requantized to q8_0 at load), tied embeddings, split -00001-of-000NN shard sets (point -m at the first shard), and both converter dialects (ds4-lineage and upstream). Known-good starters:

# Hy3 295B - 85GB, the friendlier starting point (fromBF16 = current build)
curl -L -C - -o Hy3-ds4-IQ2XXS-AttnQ8-fromBF16.gguf \
  "https://huggingface.co/giannisan/Hy3-ds4-gguf/resolve/main/Hy3-ds4-IQ2XXS-AttnQ8-fromBF16.gguf"

# GLM-5.2 743B - 197GB, needs a second 16GB GPU for the attention stack
curl -L -C - -o GLM-5.2-UD-IQ2_XXS_RoutedIQ2XXS_blk78Q2K.gguf \
  "https://huggingface.co/antirez/GLM-5.2-GGUF/resolve/main/GLM-5.2-UD-IQ2_XXS_RoutedIQ2XXS_blk78Q2K.gguf"

# Kimi K2.7-Code 1T - 339GB in 8 shards (unsloth UD-Q2_K_XL); download
# the folder and point pulsar at shard -00001-of-00008

Put the file on your fastest NVMe - decode speed is read speed.

Quick start

git clone https://github.com/giannisanni/pulsar
cd pulsar

# build (CXX only needed if your default gcc is too new for nvcc)
CXX=g++-12 cargo build --release -p engine

# run: greedy generation (multi-GPU roles auto-detected)
./target/release/pulsar-cli \
    -m /path/to/Hy3-ds4-IQ2XXS-AttnQ8.gguf \
    -p "The capital of France is" -n 64

# or: interactive chat (multi-turn, KV cache retained across turns)
./target/release/pulsar-cli -m /path/to/model.gguf --chat
# opt-in Jinja (embed → cache → HF → catalog; block network with PULSAR_OFFLINE=1)
./target/release/pulsar-cli -m /path/to/model.gguf --chat --jinja-chat

# or: OpenAI-compatible server with a built-in web UI at /
cargo build --release -p serve
./target/release/pulsar-serve -m /path/to/model.gguf --host 127.0.0.1 --port 11435 --ctx 8192
# open http://127.0.0.1:11435/  in a browser for the chat UI
# optional: --jinja-chat  |  --webui-mcp-proxy [--mcp-config FILE]
# see “pulsar-serve flags” below and docs/mcp-server.md / docs/chat_template.md
# or hit the API directly:
curl http://127.0.0.1:11435/v1/chat/completions -d '{
  "messages": [{"role": "user", "content": "Hello!"}],
  "stream": true
}'

# fetch a model's Jinja chat template (HF or llama.cpp catalog; works on
# quantized GGUFs by resolving the base model from metadata / filename):
cargo build --release -p tokenizer --bin get-chat-template
./target/release/get-chat-template microsoft/Phi-3.5-mini-instruct
./target/release/get-chat-template ./Qwen2.5-7B-Instruct-Q4_K_M.gguf --meta
./target/release/get-chat-template CohereForAI/c4ai-command-r-plus tool_use \
    --save command-r-tool_use.jinja

First run is cold. On exit the engine writes a <model>.gguf.warm sidecar (a popularity census of expert slabs); every later run bulk-loads the hot set in a few seconds, and expert tiers (spare GPUs) fill from the same census, so the second run is the fast one.

When no census exists yet, the engine seeds the warm set and tier placement from a built-in per-family hotlist (crates/engine/hotlists/, generated from real routing censuses with the hotlist-gen tool, keyed by layer/expert index so it survives requantized ggufs). Measured on Qwen3.6-35B: first-run decode 19.2 to 25.1 tok/s, with the resident tier active from token one. The real census replaces the seed on exit; PULSAR_NO_HOTLIST=1 restores the plain cold start. Idea borrowed from the static streaming hotlists in antirez's ds4 (MIT).

pulsar-cli flags

One-shot generation (-p / --tokens) or interactive chat (--chat). Linux + CUDA only.

flag meaning
-m FILE model GGUF (required)
-p TEXT prompt text (BOS prepended unless --no-bos)
-f / --prompt-file PATH read prompt from file (long prompts)
--tokens 1,2,3 feed exact token ids instead of text
-n N tokens to generate (default 16; chat uses ≥1024 if -n ≤ 16)
--ctx N context size (default 2048)
--bos / --no-bos force / suppress BOS on one-shot prompts
--chat interactive multi-turn chat (KV retained; ChatMarkers by default)
--jinja-chat with --chat: opt-in Jinja (embed → cache → HF → catalog; same as serve)
--system TEXT system prompt for --chat
--temp F sampling temperature (chat: gguf general.sampling.temp or 0.9; one-shot greedy unless set)
--top-p F nucleus sampling (chat default from gguf or 1.0)
--min-p F min-p sampling (default 0)
--seed N RNG seed (default 42)
--dump-logits FILE write next-token logits as JSON and exit
--teacher-force per-position top-5 JSONL along the given prompt ids
--decode-consistency N decode N steps, fresh-prefill same sequence, compare logits

Also used via env on CLI: PULSAR_DFLASH (draft GGUF for speculative decode), PULSAR_MTP / PULSAR_NGRAM, PULSAR_JINJA_CHAT, PULSAR_OFFLINE, PULSAR_DEBUG_CHAT / PULSAR_DEBUG_IDS, PULSAR_PROFILE — see Tuning knobs.

pulsar-serve flags

OpenAI-compatible HTTP server + web UI. Linux + CUDA only.

flag meaning
-m FILE model GGUF (required)
--host ADDR bind address (default 127.0.0.1)
--port N listen port (default 11435)
--ctx N context size (default 8192)
--jinja-chat opt-in Jinja chat encoding (embed → cache → HF → catalog)
--webui-mcp-proxy enable MCP tool-use (web UI sidebar + agentic loop)
--mcp-config FILE MCP servers JSON path (default ./mcp.json next to cwd)
--prefix-file PATH optional prefix / system context file

Env for serve: PULSAR_JINJA_CHAT, PULSAR_OFFLINE, PULSAR_ALLOWED_HOSTS, PULSAR_CTX_STATE, PULSAR_NO_PREFIX_CACHE, plus engine knobs below. MCP details: docs/mcp-server.md.

get-chat-template flags

No GPU. Resolve / dump a Jinja template (HF id, .gguf, or free-form name).

flag meaning
MODEL_ID or MODEL.gguf positional: what to resolve (required)
VARIANT optional second positional (e.g. tool_use)
--save PATH write template to PATH instead of stdout
--meta source / model_id / size on stderr; template on stdout
--offline embedded GGUF + local cache only (no HF / catalog)
-h / --help usage

Chat templates

Same policy for pulsar-serve and pulsar-cli --chat.

Two encoding paths:

  1. ChatMarkers (default) — hardcoded special-token layouts for Hy3, Kimi, ChatML/Qwen, Gemma, MiniMax, Inkling, DeepSeek, GLM, Laguna, Harmony (gpt-oss), Kimi K3. Carefully tuned for thinking modes and stop sets. Used unless you opt in. No network.
  2. Jinja — HuggingFace-style templates via minijinja. Opt-in only via --jinja-chat or PULSAR_JINJA_CHAT=1, even when the GGUF embeds tokenizer.chat_template. There is no separate --fetch-template flag: opting into Jinja is enough to allow network rollover.

Template resolution (only when --jinja-chat; first hit wins):

  1. Embedded tokenizer.chat_template in the GGUF header
  2. Local disk cache ($PULSAR_TEMPLATE_CACHE / platform cache)
  3. HuggingFace tokenizer_config.json (quant base-model walk)
  4. llama.cpp models/templates catalog on GitHub

Without --jinja-chat, neither binary resolves beyond an offline peek (CLI load log only). With Jinja on, steps 3–4 run unless PULSAR_OFFLINE=1 (then embed + cache only).

# dump / save a template (standalone tool; may use network unless --offline)
./target/release/get-chat-template Qwen/Qwen2.5-7B-Instruct --meta
./target/release/get-chat-template /path/to/model-Q4_K_M.gguf --save out.jinja

# default ChatMarkers (no network)
./target/release/pulsar-serve -m model.gguf
./target/release/pulsar-cli -m model.gguf --chat

# Jinja: embed → cache → HF → llama.cpp catalog
./target/release/pulsar-serve -m model.gguf --jinja-chat
./target/release/pulsar-cli -m model.gguf --chat --jinja-chat

# Jinja offline only (embed + local cache)
PULSAR_OFFLINE=1 ./target/release/pulsar-serve -m model.gguf --jinja-chat
PULSAR_OFFLINE=1 ./target/release/pulsar-cli -m model.gguf --chat --jinja-chat

Gated HF repos need HF_TOKEN. Apply failures log and fall back to ChatMarkers for that request/turn. Full reference: docs/chat_template.md.

Tuning knobs (env vars)

Everything auto-configures; these override. Shared by pulsar-cli and pulsar-serve unless noted.

Device / placement

var default what
PULSAR_GPU measured CUDA index of the expert-streaming (primary) GPU
PULSAR_ATTN_GPU auto (MLA/K3) attention GPU by CUDA index. MLA/K3 auto-offload (off / -1 disables); GQA is opt-in by index (capacity shuffle)
PULSAR_ATTN_VRAM_GB solved (MLA ~5–6) attn VRAM budget when packing weights
PULSAR_ATTN_HOST unset 1 = keep attention weights in pinned host memory
PULSAR_SPLIT auto multi-GPU layer split: N = N leading layers on primary; off = single card
PULSAR_UNIFIED auto unified-memory policy (platform-dependent)
PULSAR_NO_PINNED unset set = use pageable host allocs instead of pinned
PULSAR_QUIET unset set = less device / topology chatter at load
PULSAR_CUDA_ARCH auto build-time NVCC arch list (kernels crate)

KV cache / memory

var default what
PULSAR_KV f32 (serve UI may show auto) GQA K/V storage: f32, fp8, fp16, int8, q8_0, q4_0, turbo8, turbo4 (aliases rotq* / turboq*). Lossy formats are opt-in so default decode stays bit-exact. MLA / Dsv4: latent KV only honors f32 | fp8 | fp16 (turbo* is GQA-only). auto may pick fp8 when f32 KV would not fit. Prefer turbo4 for long GQA context
PULSAR_CACHE_GB measured host RAM budget for the expert LFU cache
PULSAR_DEV_CACHE_GB solved VRAM hot-expert pool (free VRAM − staging − reserve)
PULSAR_BATCH solved prefill chunk size (largest expert-union fit)
PULSAR_TIERS on off = no resident expert tiers (single-device bit-exact path)
PULSAR_NO_HOTLIST unset set = skip built-in family hotlist seed on cold start
PULSAR_NO_PREFETCH unset set = disable cross-layer expert prefetcher
PULSAR_NO_ASYNC_H2D unset set = blocking expert H2D (debug / fallback)
PULSAR_H2D_PREFETCH unset set = extra H2D prefetch aggressiveness

CPU expert lane

var default what
PULSAR_CPU unset 1 or N = host-cache-hit MoE on CPU (N worker threads)
PULSAR_CPU_STEAL on 0 = do not steal VRAM-resident experts onto the CPU lane
PULSAR_CPU_CAP unset max experts per step on the CPU lane
PULSAR_CPU_B solved CPU-lane batch / packing bound
PULSAR_CPU_VERIFY unset set = dump CPU-vs-GPU lane checks (debug)

Speculation / drafts

var default what
PULSAR_MTP unset 1 = MTP / nextn speculative decode when the GGUF has a nextn block (greedy)
PULSAR_MTP_DEPTH 3 draft chain depth for MTP
PULSAR_NGRAM unset draft-free n-gram speculation depth (greedy; disables some serve prefix-cache paths)
PULSAR_DFLASH unset path to DFlash draft GGUF (CLI speculative path)
PULSAR_GRAPHS on 0 = disable CUDA graphs where supported (e.g. Qwen3.5)

Chat templates (serve + CLI chat)

var default what
PULSAR_JINJA_CHAT unset 1 = same as --jinja-chat (opt-in Jinja; may use network)
PULSAR_OFFLINE unset with Jinja: embed + cache only (no HF / llama.cpp catalog)
PULSAR_TEMPLATE_CACHE platform cache directory for downloaded .jinja templates
HF_TOKEN / HUGGING_FACE_HUB_TOKEN unset Bearer for gated HuggingFace tokenizer_config.json
PULSAR_DEBUG_CHAT unset log rendered Jinja prompt text
PULSAR_DEBUG_IDS unset log prompt token id sequences

Serve-only

var default what
PULSAR_ALLOWED_HOSTS localhost-ish comma-separated Host allowlist for the HTTP server
PULSAR_CTX_STATE unset path for persisted context / session state
PULSAR_NO_PREFIX_CACHE unset set = disable multi-request prefix KV reuse

Profiling / debug (operator)

var default what
PULSAR_PROFILE unset set = print per-stage wall-time profile
PULSAR_MTP_DEBUG / PULSAR_MTP_TIMING unset MTP accept / timing logs
PULSAR_DFLASH_DEBUG unset DFlash draft debug
PULSAR_NO_DEEP_BATCH unset Dsv4: restore exact single-step prefill
PULSAR_NO_GROUPED unset disable grouped MoE matmul path
PULSAR_LANE_DBG / PULSAR_ROUTER_HIST / PULSAR_L2_TRACE unset low-level routing / layer traces

¹ Defaults shift with topology: attn offload frees pinned RAM and primary VRAM for larger host / dev expert caches.

Tests

cargo test                                        # host-side (any OS)
CXX=g++-12 cargo test -p kernels --release -- --test-threads=1   # GPU kernel selftests vs CPU references
scripts/check.sh /path/to/model.gguf              # full commit gate (build + selftests + bit-exact decode)

How the streaming works

Per MoE layer, per token (or per prefill chunk as a union across the whole batch), an expert slab resolves through:

  1. Resident tier (spare GPUs): the hottest expert triples live permanently on leftover cards; their MoE compute happens there and only activations cross PCIe. Placement, not cache: no eviction.
  2. VRAM hot-set cache (primary GPU): a fixed pool with touch-count admission: a slab earns a slot only by being hotter than the coldest resident, so the pool holds a stable hot set instead of thrashing.
  3. Host LFU cache: RAM-budgeted, persisted to the .warm sidecar.
  4. io_uring + O_DIRECT: misses are fetched at queue depth 32, and each completion is uploaded to the GPU while the remaining reads are still in flight.

A background thread additionally prefetches the next layer's experts, predicted by running the next layer's router on the current layer's input.

The MoE kernels never consult global state: every launch receives explicit per-(token, slot) device pointers for gate/up/down, and a NULL slot means "not mine", which is what makes per-card partial execution native. Where the bytes came from is the host's problem, resolved before launch.

Fidelity notes

  • All matmuls use ds4's exact math: activations quantized to q8_0/q8_K, integer dp4a dots. Logit-level parity with ds4 is within quantization noise.
  • Batched prefill and single-token decode use different reduction orders, so greedy near-ties (top1−top2 < ~0.5 logits) can flip between them, the same class of drift ds4 has between its CUDA and Metal backends. --decode-consistency N measures it; with PULSAR_BATCH=1 the two paths are identical and the comparison is bit-exact (verified: max |Δlogit| = 0.0).
  • Expert tiers split the per-slot sum across cards, which reorders float adds, same drift class. PULSAR_TIERS=off restores the single-device exact path. Attention offload does NOT drift: ids are byte-identical with and without it.

Status / roadmap

Done: gguf reader · io_uring disk path (parity with C at 4.8GB/s) · hy-v3 + glm-dsa (MLA compact-KV) forward graphs with GPU-vs-CPU kernel selftests · from-gguf BPE tokenizer (gold-vector parity with ds4) · four-tier streaming · warm-cache persistence · batch prefill · cross-layer prefetch · measured-bandwidth GPU role assignment · MLA attention residency on a second GPU · resident expert tiers on spare GPUs · temp/top-p/min-p sampling · interactive chat · OpenAI-compatible server (pulsar-serve: /v1/models, /v1/chat/completions with SSE streaming plus a built-in chat web UI at /, no build step, embedded via include_str!; local single-user, one request at a time). The web UI ships an expert Atlas tab: it plots every routed expert from /experts by where its weights live (VRAM core / RAM-cache ring / disk rim), dot size tracking routing heat — instant on model load, no command or build step. Drop a <model>.atlas.json sidecar in place (from scripts/atlas_build.py: routes ~40 topic probes, diffs per-expert heat, PCA to 2D) and the view upgrades to a topic-affinity galaxy; the UI polls for the sidecar and swaps automatically when it appears.

Done since: DSA lightning indexer (GLM contexts past 2048, batch scorer on tensor cores) · Kimi K2.7/deepseek2 with llama.cpp-exact YaRN · split-gguf loading · MTP + draft-free n-gram speculation (built, measured honestly: net-slower until the host cache outruns the disk; PULSAR_MTP=1 / PULSAR_NGRAM=n to experiment) · style-aware chat templates (Hy3/Kimi/ChatML/Gemma/MiniMax/Inkling/DeepSeek/GLM/Laguna/Harmony/K3) plus opt-in Jinja chat templates on serve and CLI (--jinja-chat; embed → cache → HF → llama.cpp catalog; no separate fetch flag; get-chat-template CLI) · int8 tensor-core prefill (dense GEMM + grouped MoE) · MiniMax M3, Qwen3, Gemma 4, TML Inkling forward graphs · opt-in fp8 e4m3 KV cache (PULSAR_KV=fp8) · TurboQuant rotated block-KV (PULSAR_KV=turbo4|turbo8, orthogonal Π on K/Q spreads outliers so block-quant stops zeroing the other 31 lanes; decode-invariant, V untouched) · pulsar-quant recipe quantizer (BF16 gguf → ds4-style expert mixes, iq2_xxs with imatrix, per-tensor --map rules; removes llama.cpp from the model-prep pipeline; shard streaming: --fetch-cmd/--delete-shards quantize sources bigger than the disk one shard at a time) · DeepSeek-V4-Flash (deepseek4): hyper-connection residual streams, streaming compressed KV + sink attention, indexer QAT top-k, token-id hash routing. Ornith-397B (qwen35moe at 397B): the largest qwen35-moe hybrid yet — Q2_K experts on a 512-way router, GDN linear attention, shared expert. Runs end-to-end after fixing three f32-reader load sites that the generic upload() path left as raw/reqantized bytes (conv1d, gate_inp, gate_inp_shexp); see docs/ornith-q2k-cuda-crash-fix.md. Warm decode bench pending.

Also done, honestly measured: DFlash block-diffusion speculative decoding for Qwen3.6 (the lucebox recipe: a 515MB matched draft proposes 16 tokens conditioned on 5 captured target hidden states, one batched target forward verifies the block, recurrent state snapshots roll back rejections). The machinery works - structured text accepts whole 16-blocks - but it ships opt-in experimental (PULSAR_DFLASH=draft.gguf) because on the reference box it is experimental. Four profiling rounds took it from 6.1 to 39.7 tok/s (resident expert tiers for the hybrid families, grouped tensor-core MoE for verify chunks, recurrence-only fast rollback that replaces the replay forward, a token-tiled K-quant lm head), and on the iq3-heavy Q3_K_XL target it now BEATS sequential decode on reasoning workloads (39.7 vs 36.3, byte-identical output to plain greedy). On the faster Q4_K_XL target sequential decode itself jumps to 51.8 tok/s and DFlash falls behind again: the round's remaining fixed costs (a ~95ms verify floor of per-layer launches and router readbacks, a draft whose cost grows with the feature window) need acceptance ~7+ to amortize, and measured acceptance is 4.3 on math, less on prose. MCP (Model Context Protocol) tool-use in pulsar-serve, opt-in via --webui-mcp-proxy: the server connects to configured MCP servers (rmcp 3.0.1; stdio + streamable-http), exposes their tools to the model as namespaced server__tool specs, and runs a non-stream agentic loop (MAX_TURNS=8) that executes the model's <tool_call> blocks and feeds results back. Each server card's title auto-detects the advertised name from the MCP initialize handshake and shows a rolling connection log with the last handshake latency; the on/off pill matches the CPU Lane / MTP toggles. Full CRUD lives in the web UI sidebar (add/edit/remove servers, enable/disable per tool), persisted to mcp.json. Without the flag: zero behavioral change — every /mcp/* route 404s and the sidebar group stays hidden. End-to-end verified on Qwen3.6-35B against a remote SearXNG MCP (tool_call → dispatch → grounded answer in ~22s). See docs/mcp-server.md.

Not yet:

  • DFlash, remaining: draft context-KV cache ring (lucebox DraftKvCacheRefs - caps the draft cost at long windows), CUDA-graph or fused launches for the verify's per-layer fixed costs, tree verification (DDTree) for higher acceptance per round
  • deepseek4 perf pass: batched prefill (prompts currently process sequentially), resident tiers + cross-layer prefetch for the dsv4 resolve, fewer host syncs on the hyper-connection gates
  • tensor-core unpackers for the remaining expert formats (iq2_xs, iq3_xxs, q4_K, q5_1, q2_K, q3_K, the harness takes one ~40-line unpacker per format)

License

MIT. The CUDA kernels derive from the ds4 lineage (MIT) and carry their attribution: Copyright (c) 2026 The ds4.c authors · Copyright (c) 2023–2026 The ggml authors.

About

SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-GPU: measures PCIe bandwidth, places attention and hot experts where they fit.

Topics

Resources

Stars

204 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages