Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

15 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

tiny-vLLM

A lightweight vLLM implementation built

Key Features

  • Optimization Suite - Prefix caching, Tensor Parallelism, Torch compilation, CUDA graph, etc.
  • 🚀 W8A16,W8A8

Model Download

To download the model weights manually, use the following command:

huggingface-cli download --resume-download Qwen/Qwen3-0.6B \
  --local-dir ~/huggingface/Qwen3-0.6B/ \
  --local-dir-use-symlinks False

Quick Start

See example.py for usage. The API mirrors vLLM's interface with minor differences in the LLM.generate method:

from nanovllm import LLM, SamplingParams
llm = LLM("/YOUR/MODEL/PATH", enforce_eager=True, tensor_parallel_size=1)
sampling_params = SamplingParams(temperature=0.6, max_tokens=256)
prompts = ["Hello, Nano-vLLM."]
outputs = llm.generate(prompts, sampling_params)
outputs[0]["text"]

Benchmark

See bench.py for benchmark.

  • Hardware: NVIDIA A10 (20GB)
  • Model: Qwen3-0.6B
  • Total Requests: 256 sequences
  • Input Length: Randomly sampled between 100–1024 tokens
  • Output Length: Randomly sampled between 100–1024 tokens
Inference Engine Output Tokens Time (s) Throughput (tokens/s)
tiny-vLLM 133966tok, 38.43s 38.43s, 3486.41tok/s

Quantization Results

The repository includes INT8 weight-only and W8A8 quantization paths for Qwen3-0.6B. The following results were measured on an NVIDIA A10 with bench_quant.py and bench_quant_quality.py.

Memory

Quantization clearly reduces model load-time memory:

Mode Memory After Load (allocated / reserved) Runtime Peak Allocated
bf16 1.13 GiB / 1.48 GiB 19.60 GiB
int8 0.74 GiB / 1.05 GiB 19.47 GiB
w8a8 0.74 GiB / 1.05 GiB 19.46 GiB

Load-time memory drops significantly after quantization. Runtime peak memory changes less because KV cache and other runtime buffers dominate the total footprint.

Throughput and TTFT

Mode Total Throughput Prefill Throughput Decode Throughput TTFT
bf16 3069.97 tok/s 38256.06 tok/s 3530.05 tok/s 2.69 s
int8 2634.29 tok/s 12067.17 tok/s 4082.19 tok/s 12.38 s
w8a8 2699.93 tok/s 17385.02 tok/s 3614.86 tok/s 6.95 s

Relative to bf16:

  • INT8 total throughput: 0.858x
  • W8A8 total throughput: 0.879x
  • INT8 decode throughput: 1.156x
  • W8A8 decode throughput: 1.024x

Quantized kernels already help decode, but prefill remains the main bottleneck in the current implementation. As a result, TTFT and end-to-end throughput do not yet beat bf16 on this workload.

Decode-Heavy Scaling

To separate decode behavior from mixed-workload noise, the repository now includes a decode_heavy profile in bench_quant.py. The workload uses short inputs and long outputs, then sweeps max_num_seqs while keeping max_num_batched_tokens fixed.

Observed results on the A10:

Mode Best Total Throughput Best Decode Throughput Best max_num_seqs
bf16 7935.39 tok/s 8538.61 tok/s 256 / 512
int8 8184.68 tok/s 10550.66 tok/s 128 / 512
w8a8 7705.63 tok/s 9086.10 tok/s 256 / 512

Decode throughput improves clearly when the effective decode batch grows from 128 to 256 or 512. INT8 remains the strongest decode path and reaches the highest decode throughput in this repository.

However, throughput does not continue rising past 512. In the current implementation, 512 is an important turning point because decode can no longer fully benefit from the same fast path once batch size grows beyond the captured CUDA graph range. As a result, 768 does not improve end-to-end throughput even though decode is still the dominant phase.

Prefill-Heavy Scaling

The repository also includes a prefill_heavy profile with long inputs and short outputs. This profile sweeps max_num_batched_tokens while keeping max_num_seqs fixed.

Observed results on the A10:

Mode Best Total Throughput Prefill Throughput Range TTFT Range
bf16 764.23 tok/s 46.8k - 47.2k tok/s 3.64 s -> 3.19 s
int8 276.30 tok/s 11.26k - 11.32k tok/s 15.09 s -> 13.77 s
w8a8 421.93 tok/s 19.11k - 19.31k tok/s 8.86 s -> 6.97 s

Increasing max_num_batched_tokens reduces TTFT because prefill is split into fewer waves before the first decode step can start. At the same time, prefill throughput changes only slightly. This indicates that prefill kernels are already near their steady-state efficiency in these tests: larger token capacity mostly reduces scheduling overhead and wave count, rather than unlocking a much faster per-step kernel path.

In other words:

  • Larger prefill token capacity helps latency-to-first-token.
  • It does not materially change the prefill kernel efficiency in the current implementation.
  • Prefill remains the weakest part of the INT8/W8A8 end-to-end story.

Quality Check

Quality was evaluated with artifact validation, logits comparison, greedy generation, and perplexity on WikiText-2.

Artifact check:

  • 112 / 112 target layers passed
  • Mean relative L2 reconstruction error: 0.00949
  • Minimum cosine similarity: 0.999903

Logits check:

  • INT8: cosine mean 0.999398, top-1 agreement 1.0
  • W8A8: cosine mean 0.996619, top-1 agreement 1.0

Perplexity:

Mode Loss PPL PPL Delta vs bf16
bf16 3.6475 38.3786 -
int8 3.6559 38.7034 +0.85%
w8a8 3.6797 39.6332 +3.27%

Greedy generation remained stable in all checks, with INT8 staying closer to bf16 than W8A8 over long decoding runs.

Takeaway

INT8 is currently the strongest trade-off in this repository:

  • Lower load-time memory than bf16
  • Better decode throughput than bf16
  • Very small quality regression

W8A8 is functional and reasonably accurate, but in the current implementation it is less attractive than INT8 because it improves quality less and does not provide a larger end-to-end speedup.

About

No description, website, or topics provided.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages