From faf4a3911fe24f4718ef88b320e639b415533009 Mon Sep 17 00:00:00 2001 From: Viet Nguyen Date: Thu, 13 Aug 2026 12:46:03 +0000 Subject: [PATCH] docs: delete the hand-written metric tables, and correct which whisper build is default MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Every model README carried its numbers twice: once in a table somebody typed, once in the generated `facts` block below it. The generated one is rewritten from the registry; the typed one is not, and it drifted exactly as the script's own docstring predicted — after the registry moved gipformer's Vietnamese WER to 8.5 %, the prose above still said 2.41 % in both languages, in the same file. The typed tables are removed. Facts about the artefact come from the manifest, and the prose keeps what no manifest can state: why the model is in the product. `whisper-tiny`'s README also called int8 "the default variant". It is not, and the measurement says it should not be: on 100 FLEURS clips the two builds are the same speed and int8 is 13.7 WER points worse, so `variant::choose` prefers fp32 and drops to int8 only when memory forces it. The README now says that, with the numbers and the reason the encoder-vs-decoder split produces them. --- models/campplus-sv/README.md | 14 ------ models/gipformer-65m/README.md | 33 ++------------ models/milmmt-46-1b/README.md | 16 ------- models/milmmt-46-4b/README.md | 16 ------- models/sense-voice-small/README.md | 22 ---------- models/silero-vad-v5/README.md | 26 ----------- models/small100/README.md | 16 ------- models/whisper-tiny/README.md | 70 ++++++++++++++++++------------ 8 files changed, 46 insertions(+), 167 deletions(-) diff --git a/models/campplus-sv/README.md b/models/campplus-sv/README.md index 021dda2..45da2ca 100644 --- a/models/campplus-sv/README.md +++ b/models/campplus-sv/README.md @@ -15,13 +15,6 @@ CAM++ speaker embedding — biến một câu nói đã kết thúc thành một (28 281 138 byte; sha256 xem trong `summo-registry/models/campplus-sv.json`). -**Số đo** (nguồn: `summo-registry/models/campplus-sv.json` → `profile`): - -| Chỉ số | Giá trị | -|---|---| -| RAM dùng (RSS) — idle / đỉnh | 60 MB / 180 MB | -| RAM tối thiểu | 512 MB | -| RTF (`cpu_x86_avx512vnni_8t`) | 0.006 | Manifest không có số đo độ trễ (`latency_ms`) hay chất lượng (`quality`) cho model này, nên không có trong bảng trên. @@ -38,13 +31,6 @@ latency path. `task: speaker-embed`, `mode: batch`, via `sherpa-onnx/speaker-emb (28,281,138 bytes; sha256 in `summo-registry/models/campplus-sv.json`). -**Measured numbers** (source: `summo-registry/models/campplus-sv.json` → `profile`): - -| Metric | Value | -|---|---| -| RAM used (RSS) — idle / peak | 60 MB / 180 MB | -| Minimum RAM | 512 MB | -| RTF (`cpu_x86_avx512vnni_8t`) | 0.006 | The manifest has no latency (`latency_ms`) or quality (`quality`) numbers for this model, so none are listed above. diff --git a/models/gipformer-65m/README.md b/models/gipformer-65m/README.md index c3dda29..1638b42 100644 --- a/models/gipformer-65m/README.md +++ b/models/gipformer-65m/README.md @@ -20,19 +20,6 @@ file, kho ): sha256 và dung lượng từng file xem trong `summo-registry/models/gipformer-65m.json`. -**Số đo** (nguồn: `summo-registry/models/gipformer-65m.json` → `profile`): - -| Chỉ số | Giá trị | -|---|---| -| RAM dùng (RSS) — idle / đỉnh | 150 MB / 800 MB | -| RAM tối thiểu | 1024 MB | -| RTF (`cpu_x86_avx512vnni_8t`) | 0.024 | -| RTF (`cpu_x86_avx2_4t`) | 0.06 | -| Độ trễ — partial đầu tiên | 200 ms | -| Độ trễ — hoàn tất (p50) | 300 ms | -| Độ trễ — hoàn tất (p95) | 700 ms | -| Chất lượng — WER (Fleurs VI) | 2.41% | -| Chất lượng — CER (Fleurs VI) | 1.67% | ## English @@ -52,19 +39,6 @@ files, repo ): sha256 and size for each file are in `summo-registry/models/gipformer-65m.json`. -**Measured numbers** (source: `summo-registry/models/gipformer-65m.json` → `profile`): - -| Metric | Value | -|---|---| -| RAM used (RSS) — idle / peak | 150 MB / 800 MB | -| Minimum RAM | 1024 MB | -| RTF (`cpu_x86_avx512vnni_8t`) | 0.024 | -| RTF (`cpu_x86_avx2_4t`) | 0.06 | -| Latency — first partial | 200 ms | -| Latency — finalize (p50) | 300 ms | -| Latency — finalize (p95) | 700 ms | -| Quality — WER (Fleurs VI) | 2.41% | -| Quality — CER (Fleurs VI) | 1.67% | | | | @@ -78,10 +52,11 @@ sha256 and size for each file are in `summo-registry/models/gipformer-65m.json`. | Dung lượng | 73 MB (4 file) | | RAM (idle / đỉnh) | 150 MB / 800 MB | | RAM tối thiểu | 1024 MB | -| RTF · `cpu_x86_avx512vnni_8t` | 0.024 | | RTF · `cpu_x86_avx2_4t` | 0.06 | -| Chất lượng · `wer_fleurs_vi` | 0.0241 | -| Chất lượng · `cer_fleurs_vi` | 0.0167 | +| RTF · `cpu_x86_avx512vnni_4t` | 0.023 | +| RTF · `cpu_x86_avx512vnni_8t` | 0.019 | +| Chất lượng · `wer_fleurs_vi` | 0.085 | +| Chất lượng · `cer_fleurs_vi` | 0.067 | | Độ trễ · first_partial | 200 ms | | Độ trễ · finalize_p50 | 300 ms | | Độ trễ · finalize_p95 | 700 ms | diff --git a/models/milmmt-46-1b/README.md b/models/milmmt-46-1b/README.md index 67e000d..8fba78b 100644 --- a/models/milmmt-46-1b/README.md +++ b/models/milmmt-46-1b/README.md @@ -17,14 +17,6 @@ mradermacher. File gốc (duy nhất, do đây không phải model được mirr (806 057 408 byte; sha256 xem trong `summo-registry/models/milmmt-46-1b.json`). -**Số đo** (nguồn: `summo-registry/models/milmmt-46-1b.json` → `profile`): - -| Chỉ số | Giá trị | -|---|---| -| RAM dùng (RSS) — idle / đỉnh | 900 MB / 1200 MB | -| RAM tối thiểu | 2048 MB | -| Độ trễ — hoàn tất (p50) | 700 ms | -| Độ trễ — hoàn tất (p95) | 1600 ms | Manifest không có số RTF hay chất lượng (`quality`) cho model này (cả hai đều là object rỗng `{}`), và `mode: batch` nên không có "partial đầu tiên" để nói — vì vậy các mục này không xuất hiện @@ -45,14 +37,6 @@ mradermacher. Original file (the only one, since this model is never mirrored): (806,057,408 bytes; sha256 in `summo-registry/models/milmmt-46-1b.json`). -**Measured numbers** (source: `summo-registry/models/milmmt-46-1b.json` → `profile`): - -| Metric | Value | -|---|---| -| RAM used (RSS) — idle / peak | 900 MB / 1200 MB | -| Minimum RAM | 2048 MB | -| Latency — finalize (p50) | 700 ms | -| Latency — finalize (p95) | 1600 ms | The manifest has no RTF or quality (`quality`) numbers for this model (both are empty `{}`), and being `mode: batch` there is no "first partial" to speak of — so those rows are omitted above. diff --git a/models/milmmt-46-4b/README.md b/models/milmmt-46-4b/README.md index 291c16a..51cbbfe 100644 --- a/models/milmmt-46-4b/README.md +++ b/models/milmmt-46-4b/README.md @@ -16,14 +16,6 @@ mradermacher. File gốc: (2 489 893 152 byte; sha256 xem trong `summo-registry/models/milmmt-46-4b.json`). -**Số đo** (nguồn: `summo-registry/models/milmmt-46-4b.json` → `profile`): - -| Chỉ số | Giá trị | -|---|---| -| RAM dùng (RSS) — idle / đỉnh | 2700 MB / 3200 MB | -| RAM tối thiểu | 6144 MB | -| Độ trễ — hoàn tất (p50) | 2100 ms | -| Độ trễ — hoàn tất (p95) | 3600 ms | Manifest không có số RTF hay chất lượng (`quality`) cho model này, và `mode: batch` nên không có "partial đầu tiên" — các mục này không xuất hiện trong bảng trên. @@ -42,14 +34,6 @@ mradermacher. Original file: (2,489,893,152 bytes; sha256 in `summo-registry/models/milmmt-46-4b.json`). -**Measured numbers** (source: `summo-registry/models/milmmt-46-4b.json` → `profile`): - -| Metric | Value | -|---|---| -| RAM used (RSS) — idle / peak | 2700 MB / 3200 MB | -| Minimum RAM | 6144 MB | -| Latency — finalize (p50) | 2100 ms | -| Latency — finalize (p95) | 3600 ms | The manifest has no RTF or quality (`quality`) numbers for this model, and being `mode: batch` there is no "first partial" — those rows are omitted above. diff --git a/models/sense-voice-small/README.md b/models/sense-voice-small/README.md index dc97c74..8cc54bf 100644 --- a/models/sense-voice-small/README.md +++ b/models/sense-voice-small/README.md @@ -23,17 +23,6 @@ vào nguồn gốc. sha256 và dung lượng từng file xem trong `summo-registry/models/sense-voice-small.json`. -**Số đo** (nguồn: `summo-registry/models/sense-voice-small.json` → `profile`): - -| Chỉ số | Giá trị | -|---|---| -| RAM dùng (RSS) — idle / đỉnh | 300 MB / 800 MB | -| RAM tối thiểu | 1536 MB | -| RTF (`cpu_x86_avx512vnni_8t`) | 0.044 | -| RTF (`cpu_x86_avx512vnni_4t`) | 0.062 | -| Độ trễ — partial đầu tiên | 250 ms | -| Độ trễ — hoàn tất (p50) | 300 ms | -| Độ trễ — hoàn tất (p95) | 800 ms | Manifest không có số đo chất lượng (`quality`) — mô tả trong manifest ghi rõ: độ chính xác chưa được `summo-bench` đo; tuyên bố của bên phát hành là model vượt Whisper-small trong khi chạy nhanh @@ -59,17 +48,6 @@ upstream source. sha256 and size for each file are in `summo-registry/models/sense-voice-small.json`. -**Measured numbers** (source: `summo-registry/models/sense-voice-small.json` → `profile`): - -| Metric | Value | -|---|---| -| RAM used (RSS) — idle / peak | 300 MB / 800 MB | -| Minimum RAM | 1536 MB | -| RTF (`cpu_x86_avx512vnni_8t`) | 0.044 | -| RTF (`cpu_x86_avx512vnni_4t`) | 0.062 | -| Latency — first partial | 250 ms | -| Latency — finalize (p50) | 300 ms | -| Latency — finalize (p95) | 800 ms | The manifest has no quality (`quality`) numbers — its description notes accuracy has not yet been measured by `summo-bench`; the publisher's own claim is that it beats Whisper-small while running diff --git a/models/silero-vad-v5/README.md b/models/silero-vad-v5/README.md index e59e9d5..030268f 100644 --- a/models/silero-vad-v5/README.md +++ b/models/silero-vad-v5/README.md @@ -13,19 +13,6 @@ Bộ phát hiện hoạt động giọng nói (VAD) mặc định của Summo. ` (2 327 524 byte; sha256 xem trong `summo-registry/models/silero-vad-v5.json`). -**Số đo** (nguồn: `summo-registry/models/silero-vad-v5.json` → `profile`): - -| Chỉ số | Giá trị | -|---|---| -| RAM dùng (RSS) — idle / đỉnh | 12 MB / 30 MB | -| RAM tối thiểu | 128 MB | -| RTF (`cpu_x86_avx512vnni_8t`) | 0.0063 | -| Độ trễ — partial đầu tiên | 17 ms | -| Độ trễ — hoàn tất (p50) | 91 ms | -| Độ trễ — hoàn tất (p95) | 982 ms | -| Chất lượng — F1 (bộ mười test) | 0.940 | -| Chất lượng — Precision (bộ mười test) | 0.925 | -| Chất lượng — Recall (bộ mười test) | 0.956 | ## English @@ -37,19 +24,6 @@ Summo's default voice activity detector. `task: vad`, `mode: live`, via `onnx/si (2,327,524 bytes; sha256 in `summo-registry/models/silero-vad-v5.json`). -**Measured numbers** (source: `summo-registry/models/silero-vad-v5.json` → `profile`): - -| Metric | Value | -|---|---| -| RAM used (RSS) — idle / peak | 12 MB / 30 MB | -| Minimum RAM | 128 MB | -| RTF (`cpu_x86_avx512vnni_8t`) | 0.0063 | -| Latency — first partial | 17 ms | -| Latency — finalize (p50) | 91 ms | -| Latency — finalize (p95) | 982 ms | -| Quality — F1 (ten-test set) | 0.940 | -| Quality — Precision (ten-test set) | 0.925 | -| Quality — Recall (ten-test set) | 0.956 | | | | diff --git a/models/small100/README.md b/models/small100/README.md index 041f9ef..639b0aa 100644 --- a/models/small100/README.md +++ b/models/small100/README.md @@ -33,14 +33,6 @@ File hiện tại (kho , thư m sha256 và dung lượng từng file xem trong `summo-registry/models/small100.json`. -**Số đo** (nguồn: `summo-registry/models/small100.json` → `profile`): - -| Chỉ số | Giá trị | -|---|---| -| RAM dùng (RSS) — idle / đỉnh | 700 MB / 1100 MB | -| RAM tối thiểu | 1536 MB | -| Độ trễ — hoàn tất (p50) | 244 ms | -| Độ trễ — hoàn tất (p95) | 600 ms | Manifest không có số RTF hay chất lượng (`quality`) cho model này, và `mode: batch` nên không có "partial đầu tiên" — các mục này không xuất hiện trong bảng trên. Số đo chi tiết hơn về hai quyết @@ -79,14 +71,6 @@ Current files (repo , directory sha256 and size for each file are in `summo-registry/models/small100.json`. -**Measured numbers** (source: `summo-registry/models/small100.json` → `profile`): - -| Metric | Value | -|---|---| -| RAM used (RSS) — idle / peak | 700 MB / 1100 MB | -| Minimum RAM | 1536 MB | -| Latency — finalize (p50) | 244 ms | -| Latency — finalize (p95) | 600 ms | The manifest has no RTF or quality (`quality`) numbers for this model, and being `mode: batch` there is no "first partial" — those rows are omitted above. Finer-grained measurements behind two diff --git a/models/whisper-tiny/README.md b/models/whisper-tiny/README.md index 6b86966..e691a90 100644 --- a/models/whisper-tiny/README.md +++ b/models/whisper-tiny/README.md @@ -14,23 +14,29 @@ dưới trước khi chọn model này cho một ngôn ngữ mà một model chu **Nguồn gốc:** OpenAI Whisper, export ONNX bởi csukuangfj. File gốc (kho ): -- `tiny-encoder.int8.onnx`, `tiny-decoder.int8.onnx`, `tiny-tokens.txt` — biến thể mặc định -- `tiny-encoder.onnx`, `tiny-decoder.onnx` — biến thể fp32 +- `tiny-encoder.onnx`, `tiny-decoder.onnx` — biến thể **fp32**, được chọn mặc định khi đủ RAM +- `tiny-encoder.int8.onnx`, `tiny-decoder.int8.onnx` — biến thể **int8**, chỉ dùng khi máy thiếu RAM +- `tiny-tokens.txt` — không gắn biến thể nào, nên đi kèm cả hai sha256 và dung lượng từng file xem trong `summo-registry/models/whisper-tiny.json`. -**Số đo** (nguồn: `summo-registry/models/whisper-tiny.json` → `profile`): +### int8 ở đây **không** nhanh hơn, và kém chính xác hơn nhiều -| Chỉ số | Giá trị | -|---|---| -| RAM dùng (RSS) — idle / đỉnh | 200 MB / 600 MB | -| RAM tối thiểu | 1024 MB | -| RTF (`cpu_x86_avx512vnni_8t`) | 0.107 | -| RTF (`cpu_x86_avx2_4t`) | 0.3 | -| Độ trễ — hoàn tất (p50) | 700 ms | -| Độ trễ — hoàn tất (p95) | 1500 ms | -| Chất lượng — WER (bộ test Whisper, EN) | 4.5% | -| Chất lượng — WER (Fleurs VI) | 65.5% | +100 clip FLEURS vi (21,3 phút), cùng runtime, cùng máy (Xeon Gold 6226R, AVX-512 VNNI): + +| Biến thể | WER | RTF 4 luồng | RTF 8 luồng | Dung lượng | +|---|---:|---:|---:|---:| +| fp32 | **67,6 %** | **0,137** | **0,116** | 146 MB | +| int8 | 81,3 % | 0,138 | 0,120 | 99 MB | + +Tốc độ như nhau, kém **13,7 điểm WER**. Cái int8 mua được ở đây chỉ là 47 MB đĩa và RAM tương ứng. + +Không phải do CPU yếu: riêng encoder, int8 **nhanh hơn 1,24×** trong ONNX Runtime. Phần thắng đó bị +trả lại hết ở decoder tự hồi quy — mỗi clip decoder chạy vài trăm lượt trên các ma trận nhỏ, nơi chi +phí quantize/dequantize quanh mỗi phép nhân lớn hơn phần nhân được rẻ đi. Chi tiết và cách chạy lại: +`docs/benchmarks.md` trong `summo-app`. + +Vì vậy `variant::choose` xếp fp32 trên int8, và chỉ hạ xuống int8 khi RAM trống không đủ. `mode: batch` nên không có "partial đầu tiên" để nói. @@ -46,23 +52,29 @@ mirrored to this repo's Releases. **Upstream:** OpenAI Whisper, ONNX export by csukuangfj. Original files (repo ): -- `tiny-encoder.int8.onnx`, `tiny-decoder.int8.onnx`, `tiny-tokens.txt` — default variant -- `tiny-encoder.onnx`, `tiny-decoder.onnx` — fp32 variant +- `tiny-encoder.onnx`, `tiny-decoder.onnx` — the **fp32** variant, chosen by default where RAM allows +- `tiny-encoder.int8.onnx`, `tiny-decoder.int8.onnx` — the **int8** variant, for machines short of RAM +- `tiny-tokens.txt` — tagged with no variant, so it ships with both sha256 and size for each file are in `summo-registry/models/whisper-tiny.json`. -**Measured numbers** (source: `summo-registry/models/whisper-tiny.json` → `profile`): +### int8 is **not** faster here, and it is much less accurate -| Metric | Value | -|---|---| -| RAM used (RSS) — idle / peak | 200 MB / 600 MB | -| Minimum RAM | 1024 MB | -| RTF (`cpu_x86_avx512vnni_8t`) | 0.107 | -| RTF (`cpu_x86_avx2_4t`) | 0.3 | -| Latency — finalize (p50) | 700 ms | -| Latency — finalize (p95) | 1500 ms | -| Quality — WER (Whisper's own test set, EN) | 4.5% | -| Quality — WER (Fleurs VI) | 65.5% | +100 FLEURS vi clips (21.3 min), same runtime, same machine (Xeon Gold 6226R, AVX-512 VNNI): + +| Variant | WER | RTF 4 threads | RTF 8 threads | Size | +|---|---:|---:|---:|---:| +| fp32 | **67.6 %** | **0.137** | **0.116** | 146 MB | +| int8 | 81.3 % | 0.138 | 0.120 | 99 MB | + +Identical speed, **13.7 points worse**. What int8 buys here is 47 MB of disk and the matching RAM. + +Not a weak CPU: on the encoder alone int8 *is* **1.24× faster** in ONNX Runtime. That win is handed +straight back by the autoregressive decoder, which runs a few hundred times per clip on matrices +small enough that the quantise/dequantise around each multiply costs more than the cheaper multiply +saves. Method and how to re-run: `docs/benchmarks.md` in `summo-app`. + +So `variant::choose` ranks fp32 above int8, and drops to int8 only when free memory forces it. Being `mode: batch`, there is no "first partial" to speak of. @@ -78,10 +90,12 @@ Being `mode: batch`, there is no "first partial" to speak of. | Dung lượng | 256 MB (5 file) | | RAM (idle / đỉnh) | 200 MB / 600 MB | | RAM tối thiểu | 1024 MB | -| RTF · `cpu_x86_avx512vnni_8t` | 0.107 | | RTF · `cpu_x86_avx2_4t` | 0.3 | +| RTF · `cpu_x86_avx512vnni_4t` | 0.138 | +| RTF · `cpu_x86_avx512vnni_8t` | 0.12 | | Chất lượng · `wer_whisper_testset_en` | 0.045 | -| Chất lượng · `wer_fleurs_vi` | 0.655 | +| Chất lượng · `wer_fleurs_vi` | 0.676 | +| Chất lượng · `cer_fleurs_vi` | 0.451 | | Độ trễ · first_partial | 0 ms | | Độ trễ · finalize_p50 | 700 ms | | Độ trễ · finalize_p95 | 1500 ms |