Pinned Loading
-
orbitquant
orbitquant PublicExperimental further optimization proposals and implementations for TurboQuant-style KV cache compression
-
awq
awq PublicModel agnostic AWQ quantization CLI - group-wise INT4 quantization with reconstruction verify and runtime export
Python
-
-
transformers
transformers PublicForked from huggingface/transformers
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Python
-
sped
sped Public(WIP) Universal speculative decoding - pair any small draft model with any large target model for 2-5× faster LLM inference, zero quality loss
Python
If the problem persists, check the GitHub status page or contact support.



