Skip to content
View giannisanni's full-sized avatar

Block or report giannisanni

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Pinned Loading

  1. pulsar pulsar Public

    SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-GPU: measures PCIe bandwidth, places …

    Rust 205 25

  2. neutronstar neutronstar Public

    Forked from antirez/ds4

    Giant MoE models on a single consumer GPU by streaming experts from SSD. CUDA fork of antirez/ds4: runs GLM-5.2 (743B), Tencent Hy3 (295B), and DeepSeek 4 Flash, with io_uring expert streaming, LFU…

    C 40 3

  3. kokoro-tts-mcp kokoro-tts-mcp Public

    MCP server for Kokoro text-to-speech, with adjustable voices/speed and an optional OpenAI-compatible (kokoro-fastapi) backend.

    Python 10 4

  4. agora agora Public

    Rust

  5. encoder-free-speech encoder-free-speech Public

    Conversational speech generation on a frozen, encoder-free multimodal LLM: QLoRA adapters + CSM-style speech head on a single 16GB GPU. Paper + full pipeline.

    Python

  6. protostar protostar Public

    134M LLM trained from scratch on two consumer GPUs + accretion: replay-gated memory retaining 77% of learned facts across sequential consolidations (vs 12.5% without)

    Python