Skip to content
View egesabanci's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report egesabanci

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Pinned Loading

  1. orbitquant orbitquant Public

    Experimental further optimization proposals and implementations for TurboQuant-style KV cache compression

  2. awq awq Public

    Model agnostic AWQ quantization CLI - group-wise INT4 quantization with reconstruction verify and runtime export

    Python

  3. reap-mlx reap-mlx Public

    MLX-compatible REAP for pruning MoE models on Apple Silicon

    Python 8

  4. reap-cuda reap-cuda Public

    CUDA-compatible REAP for pruning MoE models on NVIDIA GPUs

    Python

  5. transformers transformers Public

    Forked from huggingface/transformers

    🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

    Python

  6. sped sped Public

    (WIP) Universal speculative decoding - pair any small draft model with any large target model for 2-5× faster LLM inference, zero quality loss

    Python