Reproducible refusal-subspace editing and TP=2 deployment of DeepSeek V4 Flash 0731 on two DGX Sparks
-
Updated
Aug 14, 2026 - Python
Reproducible refusal-subspace editing and TP=2 deployment of DeepSeek V4 Flash 0731 on two DGX Sparks
DeepSeek-V4-Flash + DSpark speculative decoding on a pair of NVIDIA DGX Sparks (vLLM TP=2 over RoCE) — tuned recipe, overlays that halve multi-turn TTFT, contamination-guarded benchmarks, ops runbook
Production-oriented Qwen3.6-35B-A3B-NVFP4-Fast vLLM deployment for NVIDIA DGX Spark / GB10
Measured Muse Glimmer 30B recipe for NVIDIA DGX Spark, with llama.cpp, DFlash parity checks, vision, tools, and GB10 benchmarks.
Reproducible high-speed Qwen3.6-35B-A3B inference on NVIDIA GB10
Benchmark for GB10 - Nvidia DGX Spark
Auditable DeepSeek V4 Flash inference evidence on two NVIDIA GB10 systems
Production-ready local AI agent stack for NVIDIA DGX Spark / ASUS GB10. Gemma 4 31B NVFP4 + bge-m3 embeddings + OpenClaw gateway, one docker compose up.
Auditable DeepSeek V4 Flash inference evidence on two NVIDIA GB10 systems
Empirical NVIDIA GB10 and SM121 microarchitecture characterization
Add a description, image, and links to the nvidia-gb10 topic page so that developers can more easily learn about it.
To associate your repository with the nvidia-gb10 topic, visit your repo's landing page and select "manage topics."