Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 0 additions & 14 deletions models/campplus-sv/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,13 +15,6 @@ CAM++ speaker embedding — biến một câu nói đã kết thúc thành một
<https://huggingface.co/csukuangfj/speaker-embedding-models/resolve/main/3dspeaker_speech_campplus_sv_zh-cn_16k-common.onnx>
(28 281 138 byte; sha256 xem trong `summo-registry/models/campplus-sv.json`).

**Số đo** (nguồn: `summo-registry/models/campplus-sv.json` → `profile`):

| Chỉ số | Giá trị |
|---|---|
| RAM dùng (RSS) — idle / đỉnh | 60 MB / 180 MB |
| RAM tối thiểu | 512 MB |
| RTF (`cpu_x86_avx512vnni_8t`) | 0.006 |

Manifest không có số đo độ trễ (`latency_ms`) hay chất lượng (`quality`) cho model này, nên không có
trong bảng trên.
Expand All @@ -38,13 +31,6 @@ latency path. `task: speaker-embed`, `mode: batch`, via `sherpa-onnx/speaker-emb
<https://huggingface.co/csukuangfj/speaker-embedding-models/resolve/main/3dspeaker_speech_campplus_sv_zh-cn_16k-common.onnx>
(28,281,138 bytes; sha256 in `summo-registry/models/campplus-sv.json`).

**Measured numbers** (source: `summo-registry/models/campplus-sv.json` → `profile`):

| Metric | Value |
|---|---|
| RAM used (RSS) — idle / peak | 60 MB / 180 MB |
| Minimum RAM | 512 MB |
| RTF (`cpu_x86_avx512vnni_8t`) | 0.006 |

The manifest has no latency (`latency_ms`) or quality (`quality`) numbers for this model, so none
are listed above.
Expand Down
33 changes: 4 additions & 29 deletions models/gipformer-65m/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,19 +20,6 @@ file, kho <https://huggingface.co/g-group-ai-lab/gipformer-65M-rnnt>):

sha256 và dung lượng từng file xem trong `summo-registry/models/gipformer-65m.json`.

**Số đo** (nguồn: `summo-registry/models/gipformer-65m.json` → `profile`):

| Chỉ số | Giá trị |
|---|---|
| RAM dùng (RSS) — idle / đỉnh | 150 MB / 800 MB |
| RAM tối thiểu | 1024 MB |
| RTF (`cpu_x86_avx512vnni_8t`) | 0.024 |
| RTF (`cpu_x86_avx2_4t`) | 0.06 |
| Độ trễ — partial đầu tiên | 200 ms |
| Độ trễ — hoàn tất (p50) | 300 ms |
| Độ trễ — hoàn tất (p95) | 700 ms |
| Chất lượng — WER (Fleurs VI) | 2.41% |
| Chất lượng — CER (Fleurs VI) | 1.67% |

## English

Expand All @@ -52,19 +39,6 @@ files, repo <https://huggingface.co/g-group-ai-lab/gipformer-65M-rnnt>):

sha256 and size for each file are in `summo-registry/models/gipformer-65m.json`.

**Measured numbers** (source: `summo-registry/models/gipformer-65m.json` → `profile`):

| Metric | Value |
|---|---|
| RAM used (RSS) — idle / peak | 150 MB / 800 MB |
| Minimum RAM | 1024 MB |
| RTF (`cpu_x86_avx512vnni_8t`) | 0.024 |
| RTF (`cpu_x86_avx2_4t`) | 0.06 |
| Latency — first partial | 200 ms |
| Latency — finalize (p50) | 300 ms |
| Latency — finalize (p95) | 700 ms |
| Quality — WER (Fleurs VI) | 2.41% |
| Quality — CER (Fleurs VI) | 1.67% |

<!-- facts:start -->
| | |
Expand All @@ -78,10 +52,11 @@ sha256 and size for each file are in `summo-registry/models/gipformer-65m.json`.
| Dung lượng | 73 MB (4 file) |
| RAM (idle / đỉnh) | 150 MB / 800 MB |
| RAM tối thiểu | 1024 MB |
| RTF · `cpu_x86_avx512vnni_8t` | 0.024 |
| RTF · `cpu_x86_avx2_4t` | 0.06 |
| Chất lượng · `wer_fleurs_vi` | 0.0241 |
| Chất lượng · `cer_fleurs_vi` | 0.0167 |
| RTF · `cpu_x86_avx512vnni_4t` | 0.023 |
| RTF · `cpu_x86_avx512vnni_8t` | 0.019 |
| Chất lượng · `wer_fleurs_vi` | 0.085 |
| Chất lượng · `cer_fleurs_vi` | 0.067 |
| Độ trễ · first_partial | 200 ms |
| Độ trễ · finalize_p50 | 300 ms |
| Độ trễ · finalize_p95 | 700 ms |
Expand Down
16 changes: 0 additions & 16 deletions models/milmmt-46-1b/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,14 +17,6 @@ mradermacher. File gốc (duy nhất, do đây không phải model được mirr
<https://huggingface.co/mradermacher/MiLMMT-46-1B-v1.0-GGUF/resolve/main/MiLMMT-46-1B-v1.0.Q4_K_M.gguf>
(806 057 408 byte; sha256 xem trong `summo-registry/models/milmmt-46-1b.json`).

**Số đo** (nguồn: `summo-registry/models/milmmt-46-1b.json` → `profile`):

| Chỉ số | Giá trị |
|---|---|
| RAM dùng (RSS) — idle / đỉnh | 900 MB / 1200 MB |
| RAM tối thiểu | 2048 MB |
| Độ trễ — hoàn tất (p50) | 700 ms |
| Độ trễ — hoàn tất (p95) | 1600 ms |

Manifest không có số RTF hay chất lượng (`quality`) cho model này (cả hai đều là object rỗng
`{}`), và `mode: batch` nên không có "partial đầu tiên" để nói — vì vậy các mục này không xuất hiện
Expand All @@ -45,14 +37,6 @@ mradermacher. Original file (the only one, since this model is never mirrored):
<https://huggingface.co/mradermacher/MiLMMT-46-1B-v1.0-GGUF/resolve/main/MiLMMT-46-1B-v1.0.Q4_K_M.gguf>
(806,057,408 bytes; sha256 in `summo-registry/models/milmmt-46-1b.json`).

**Measured numbers** (source: `summo-registry/models/milmmt-46-1b.json` → `profile`):

| Metric | Value |
|---|---|
| RAM used (RSS) — idle / peak | 900 MB / 1200 MB |
| Minimum RAM | 2048 MB |
| Latency — finalize (p50) | 700 ms |
| Latency — finalize (p95) | 1600 ms |

The manifest has no RTF or quality (`quality`) numbers for this model (both are empty `{}`), and
being `mode: batch` there is no "first partial" to speak of — so those rows are omitted above.
Expand Down
16 changes: 0 additions & 16 deletions models/milmmt-46-4b/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,14 +16,6 @@ mradermacher. File gốc:
<https://huggingface.co/mradermacher/MiLMMT-46-4B-v1.0-GGUF/resolve/main/MiLMMT-46-4B-v1.0.Q4_K_M.gguf>
(2 489 893 152 byte; sha256 xem trong `summo-registry/models/milmmt-46-4b.json`).

**Số đo** (nguồn: `summo-registry/models/milmmt-46-4b.json` → `profile`):

| Chỉ số | Giá trị |
|---|---|
| RAM dùng (RSS) — idle / đỉnh | 2700 MB / 3200 MB |
| RAM tối thiểu | 6144 MB |
| Độ trễ — hoàn tất (p50) | 2100 ms |
| Độ trễ — hoàn tất (p95) | 3600 ms |

Manifest không có số RTF hay chất lượng (`quality`) cho model này, và `mode: batch` nên không có
"partial đầu tiên" — các mục này không xuất hiện trong bảng trên.
Expand All @@ -42,14 +34,6 @@ mradermacher. Original file:
<https://huggingface.co/mradermacher/MiLMMT-46-4B-v1.0-GGUF/resolve/main/MiLMMT-46-4B-v1.0.Q4_K_M.gguf>
(2,489,893,152 bytes; sha256 in `summo-registry/models/milmmt-46-4b.json`).

**Measured numbers** (source: `summo-registry/models/milmmt-46-4b.json` → `profile`):

| Metric | Value |
|---|---|
| RAM used (RSS) — idle / peak | 2700 MB / 3200 MB |
| Minimum RAM | 6144 MB |
| Latency — finalize (p50) | 2100 ms |
| Latency — finalize (p95) | 3600 ms |

The manifest has no RTF or quality (`quality`) numbers for this model, and being `mode: batch`
there is no "first partial" — those rows are omitted above.
Expand Down
22 changes: 0 additions & 22 deletions models/sense-voice-small/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,17 +23,6 @@ vào nguồn gốc.

sha256 và dung lượng từng file xem trong `summo-registry/models/sense-voice-small.json`.

**Số đo** (nguồn: `summo-registry/models/sense-voice-small.json` → `profile`):

| Chỉ số | Giá trị |
|---|---|
| RAM dùng (RSS) — idle / đỉnh | 300 MB / 800 MB |
| RAM tối thiểu | 1536 MB |
| RTF (`cpu_x86_avx512vnni_8t`) | 0.044 |
| RTF (`cpu_x86_avx512vnni_4t`) | 0.062 |
| Độ trễ — partial đầu tiên | 250 ms |
| Độ trễ — hoàn tất (p50) | 300 ms |
| Độ trễ — hoàn tất (p95) | 800 ms |

Manifest không có số đo chất lượng (`quality`) — mô tả trong manifest ghi rõ: độ chính xác chưa
được `summo-bench` đo; tuyên bố của bên phát hành là model vượt Whisper-small trong khi chạy nhanh
Expand All @@ -59,17 +48,6 @@ upstream source.

sha256 and size for each file are in `summo-registry/models/sense-voice-small.json`.

**Measured numbers** (source: `summo-registry/models/sense-voice-small.json` → `profile`):

| Metric | Value |
|---|---|
| RAM used (RSS) — idle / peak | 300 MB / 800 MB |
| Minimum RAM | 1536 MB |
| RTF (`cpu_x86_avx512vnni_8t`) | 0.044 |
| RTF (`cpu_x86_avx512vnni_4t`) | 0.062 |
| Latency — first partial | 250 ms |
| Latency — finalize (p50) | 300 ms |
| Latency — finalize (p95) | 800 ms |

The manifest has no quality (`quality`) numbers — its description notes accuracy has not yet been
measured by `summo-bench`; the publisher's own claim is that it beats Whisper-small while running
Expand Down
26 changes: 0 additions & 26 deletions models/silero-vad-v5/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,19 +13,6 @@ Bộ phát hiện hoạt động giọng nói (VAD) mặc định của Summo. `
<https://github.com/snakers4/silero-vad/raw/master/src/silero_vad/data/silero_vad.onnx>
(2 327 524 byte; sha256 xem trong `summo-registry/models/silero-vad-v5.json`).

**Số đo** (nguồn: `summo-registry/models/silero-vad-v5.json` → `profile`):

| Chỉ số | Giá trị |
|---|---|
| RAM dùng (RSS) — idle / đỉnh | 12 MB / 30 MB |
| RAM tối thiểu | 128 MB |
| RTF (`cpu_x86_avx512vnni_8t`) | 0.0063 |
| Độ trễ — partial đầu tiên | 17 ms |
| Độ trễ — hoàn tất (p50) | 91 ms |
| Độ trễ — hoàn tất (p95) | 982 ms |
| Chất lượng — F1 (bộ mười test) | 0.940 |
| Chất lượng — Precision (bộ mười test) | 0.925 |
| Chất lượng — Recall (bộ mười test) | 0.956 |

## English

Expand All @@ -37,19 +24,6 @@ Summo's default voice activity detector. `task: vad`, `mode: live`, via `onnx/si
<https://github.com/snakers4/silero-vad/raw/master/src/silero_vad/data/silero_vad.onnx>
(2,327,524 bytes; sha256 in `summo-registry/models/silero-vad-v5.json`).

**Measured numbers** (source: `summo-registry/models/silero-vad-v5.json` → `profile`):

| Metric | Value |
|---|---|
| RAM used (RSS) — idle / peak | 12 MB / 30 MB |
| Minimum RAM | 128 MB |
| RTF (`cpu_x86_avx512vnni_8t`) | 0.0063 |
| Latency — first partial | 17 ms |
| Latency — finalize (p50) | 91 ms |
| Latency — finalize (p95) | 982 ms |
| Quality — F1 (ten-test set) | 0.940 |
| Quality — Precision (ten-test set) | 0.925 |
| Quality — Recall (ten-test set) | 0.956 |

<!-- facts:start -->
| | |
Expand Down
16 changes: 0 additions & 16 deletions models/small100/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,14 +33,6 @@ File hiện tại (kho <https://huggingface.co/lyphanthuc/small100-onnx>, thư m

sha256 và dung lượng từng file xem trong `summo-registry/models/small100.json`.

**Số đo** (nguồn: `summo-registry/models/small100.json` → `profile`):

| Chỉ số | Giá trị |
|---|---|
| RAM dùng (RSS) — idle / đỉnh | 700 MB / 1100 MB |
| RAM tối thiểu | 1536 MB |
| Độ trễ — hoàn tất (p50) | 244 ms |
| Độ trễ — hoàn tất (p95) | 600 ms |

Manifest không có số RTF hay chất lượng (`quality`) cho model này, và `mode: batch` nên không có
"partial đầu tiên" — các mục này không xuất hiện trong bảng trên. Số đo chi tiết hơn về hai quyết
Expand Down Expand Up @@ -79,14 +71,6 @@ Current files (repo <https://huggingface.co/lyphanthuc/small100-onnx>, directory

sha256 and size for each file are in `summo-registry/models/small100.json`.

**Measured numbers** (source: `summo-registry/models/small100.json` → `profile`):

| Metric | Value |
|---|---|
| RAM used (RSS) — idle / peak | 700 MB / 1100 MB |
| Minimum RAM | 1536 MB |
| Latency — finalize (p50) | 244 ms |
| Latency — finalize (p95) | 600 ms |

The manifest has no RTF or quality (`quality`) numbers for this model, and being `mode: batch`
there is no "first partial" — those rows are omitted above. Finer-grained measurements behind two
Expand Down
70 changes: 42 additions & 28 deletions models/whisper-tiny/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,23 +14,29 @@ dưới trước khi chọn model này cho một ngôn ngữ mà một model chu
**Nguồn gốc:** OpenAI Whisper, export ONNX bởi csukuangfj. File gốc (kho
<https://huggingface.co/csukuangfj/sherpa-onnx-whisper-tiny>):

- `tiny-encoder.int8.onnx`, `tiny-decoder.int8.onnx`, `tiny-tokens.txt` — biến thể mặc định
- `tiny-encoder.onnx`, `tiny-decoder.onnx` — biến thể fp32
- `tiny-encoder.onnx`, `tiny-decoder.onnx` — biến thể **fp32**, được chọn mặc định khi đủ RAM
- `tiny-encoder.int8.onnx`, `tiny-decoder.int8.onnx` — biến thể **int8**, chỉ dùng khi máy thiếu RAM
- `tiny-tokens.txt` — không gắn biến thể nào, nên đi kèm cả hai

sha256 và dung lượng từng file xem trong `summo-registry/models/whisper-tiny.json`.

**Số đo** (nguồn: `summo-registry/models/whisper-tiny.json` → `profile`):
### int8 ở đây **không** nhanh hơn, và kém chính xác hơn nhiều

| Chỉ số | Giá trị |
|---|---|
| RAM dùng (RSS) — idle / đỉnh | 200 MB / 600 MB |
| RAM tối thiểu | 1024 MB |
| RTF (`cpu_x86_avx512vnni_8t`) | 0.107 |
| RTF (`cpu_x86_avx2_4t`) | 0.3 |
| Độ trễ — hoàn tất (p50) | 700 ms |
| Độ trễ — hoàn tất (p95) | 1500 ms |
| Chất lượng — WER (bộ test Whisper, EN) | 4.5% |
| Chất lượng — WER (Fleurs VI) | 65.5% |
100 clip FLEURS vi (21,3 phút), cùng runtime, cùng máy (Xeon Gold 6226R, AVX-512 VNNI):

| Biến thể | WER | RTF 4 luồng | RTF 8 luồng | Dung lượng |
|---|---:|---:|---:|---:|
| fp32 | **67,6 %** | **0,137** | **0,116** | 146 MB |
| int8 | 81,3 % | 0,138 | 0,120 | 99 MB |

Tốc độ như nhau, kém **13,7 điểm WER**. Cái int8 mua được ở đây chỉ là 47 MB đĩa và RAM tương ứng.

Không phải do CPU yếu: riêng encoder, int8 **nhanh hơn 1,24×** trong ONNX Runtime. Phần thắng đó bị
trả lại hết ở decoder tự hồi quy — mỗi clip decoder chạy vài trăm lượt trên các ma trận nhỏ, nơi chi
phí quantize/dequantize quanh mỗi phép nhân lớn hơn phần nhân được rẻ đi. Chi tiết và cách chạy lại:
`docs/benchmarks.md` trong `summo-app`.

Vì vậy `variant::choose` xếp fp32 trên int8, và chỉ hạ xuống int8 khi RAM trống không đủ.

`mode: batch` nên không có "partial đầu tiên" để nói.

Expand All @@ -46,23 +52,29 @@ mirrored to this repo's Releases.
**Upstream:** OpenAI Whisper, ONNX export by csukuangfj. Original files (repo
<https://huggingface.co/csukuangfj/sherpa-onnx-whisper-tiny>):

- `tiny-encoder.int8.onnx`, `tiny-decoder.int8.onnx`, `tiny-tokens.txt` — default variant
- `tiny-encoder.onnx`, `tiny-decoder.onnx` — fp32 variant
- `tiny-encoder.onnx`, `tiny-decoder.onnx` — the **fp32** variant, chosen by default where RAM allows
- `tiny-encoder.int8.onnx`, `tiny-decoder.int8.onnx` — the **int8** variant, for machines short of RAM
- `tiny-tokens.txt` — tagged with no variant, so it ships with both

sha256 and size for each file are in `summo-registry/models/whisper-tiny.json`.

**Measured numbers** (source: `summo-registry/models/whisper-tiny.json` → `profile`):
### int8 is **not** faster here, and it is much less accurate

| Metric | Value |
|---|---|
| RAM used (RSS) — idle / peak | 200 MB / 600 MB |
| Minimum RAM | 1024 MB |
| RTF (`cpu_x86_avx512vnni_8t`) | 0.107 |
| RTF (`cpu_x86_avx2_4t`) | 0.3 |
| Latency — finalize (p50) | 700 ms |
| Latency — finalize (p95) | 1500 ms |
| Quality — WER (Whisper's own test set, EN) | 4.5% |
| Quality — WER (Fleurs VI) | 65.5% |
100 FLEURS vi clips (21.3 min), same runtime, same machine (Xeon Gold 6226R, AVX-512 VNNI):

| Variant | WER | RTF 4 threads | RTF 8 threads | Size |
|---|---:|---:|---:|---:|
| fp32 | **67.6 %** | **0.137** | **0.116** | 146 MB |
| int8 | 81.3 % | 0.138 | 0.120 | 99 MB |

Identical speed, **13.7 points worse**. What int8 buys here is 47 MB of disk and the matching RAM.

Not a weak CPU: on the encoder alone int8 *is* **1.24× faster** in ONNX Runtime. That win is handed
straight back by the autoregressive decoder, which runs a few hundred times per clip on matrices
small enough that the quantise/dequantise around each multiply costs more than the cheaper multiply
saves. Method and how to re-run: `docs/benchmarks.md` in `summo-app`.

So `variant::choose` ranks fp32 above int8, and drops to int8 only when free memory forces it.

Being `mode: batch`, there is no "first partial" to speak of.

Expand All @@ -78,10 +90,12 @@ Being `mode: batch`, there is no "first partial" to speak of.
| Dung lượng | 256 MB (5 file) |
| RAM (idle / đỉnh) | 200 MB / 600 MB |
| RAM tối thiểu | 1024 MB |
| RTF · `cpu_x86_avx512vnni_8t` | 0.107 |
| RTF · `cpu_x86_avx2_4t` | 0.3 |
| RTF · `cpu_x86_avx512vnni_4t` | 0.138 |
| RTF · `cpu_x86_avx512vnni_8t` | 0.12 |
| Chất lượng · `wer_whisper_testset_en` | 0.045 |
| Chất lượng · `wer_fleurs_vi` | 0.655 |
| Chất lượng · `wer_fleurs_vi` | 0.676 |
| Chất lượng · `cer_fleurs_vi` | 0.451 |
| Độ trễ · first_partial | 0 ms |
| Độ trễ · finalize_p50 | 700 ms |
| Độ trễ · finalize_p95 | 1500 ms |
Expand Down
Loading