Model-agnostic, hardware-agnostic pure-C++ inference engine — one binary, NPU + GPU + CPU. 94% HF architecture coverage. Reverse-engineered AMD's closed-source Strix Halo NPU stack in 4 days. GGUF/ONNX/1BP. Zero Python. MIT.
vulkan mit-license quantization mamba inference-engine model-agnostic cplusplus-23 ai-inference local-llm one-binary gguf open-source-ai amd-strix-halo npu-inference fused-engine xdna-2 zero-python amd-native 1bp-format ternary-inference
-
Updated
Aug 17, 2026 - C++