Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
64 commits
Select commit Hold shift + click to select a range
6203ddd
Add BFCL function-calling evaluation support
dougbtv May 19, 2026
c539b03
fix(bfcl): add /v1 prefix to OpenAI base URL and fix score file glob
dougbtv May 19, 2026
85bd1ad
add Buildkite log group headers for BFCL categories
dougbtv May 19, 2026
989bc28
Add DeepSeek-V4-Flash GSM8K workload on B200
khluu May 19, 2026
ac8ccfd
Remove vllm_bench from DSv4-Flash workload
khluu May 20, 2026
5cf79f6
Add gpt-oss-120b H200 workload for nightly perf & eval
dougbtv May 26, 2026
bcb9667
Add BFCL tool-calling eval to gpt-oss-120b recipe
dougbtv May 26, 2026
868e0c7
Switch gpt-oss-120b to TP=4
dougbtv May 27, 2026
fcfe364
Add BFCL tool-calling eval to all tool-calling-capable models
dougbtv May 27, 2026
baec062
Use ECR pull-through cache for B200 K8s pod images
khluu May 28, 2026
3ac87c4
Delete workloads/deepseek_v3_2_h200.yaml
khluu May 29, 2026
b40984a
add mi355 and kimi k2.5 coverage
micah-wil Jun 1, 2026
7f32755
enable aiter
micah-wil Jun 1, 2026
dd51528
use amd GPU emoji
micah-wil Jun 3, 2026
492e4a5
create k8 plugin for mi355
micah-wil Jun 4, 2026
fddb9c1
Change agent queue to default in pipeline.yaml
dhonnappa-amd Jun 4, 2026
abe2c31
change queue in profile for testing
dhonnappa-amd Jun 4, 2026
dbd8dc7
add image pull secret
Jun 5, 2026
e85197e
restore queue to match upstream
Jun 5, 2026
fdebf9c
update secret name
dhonnappa-amd Jun 5, 2026
842dbcb
fix secret name for amd
dhonnappa-amd Jun 5, 2026
02ba653
use unique queue name
micah-wil Jun 5, 2026
5c9c835
cleanup
micah-wil Jun 5, 2026
5bcfbde
cleanup
micah-wil Jun 5, 2026
af5d554
add mi300
micah-wil Jun 9, 2026
53029ff
add DS V4 config for mi355
micah-wil Jun 9, 2026
b335d91
add quantization flag for ds v4
micah-wil Jun 9, 2026
45a168d
add gpt-oss config for mi355
micah-wil Jun 9, 2026
4800994
enable aiter for gpt oss
micah-wil Jun 9, 2026
a4c14d3
Updated to allow for testing various backend attentions with models. …
stacyroberts Jun 11, 2026
16a2328
added delay between server stop and server start for gpu spin down be…
stacyroberts Jun 11, 2026
d78ee11
clean up
stacyroberts Jun 11, 2026
cda22c5
yamls for a few AFO-LLM workloads, server.sh trying to get compilatio…
stacyroberts Jun 12, 2026
1e4092e
env/compilation-config updated
stacyroberts Jun 15, 2026
006d829
fixed datetime issue to remove warning
stacyroberts Jun 15, 2026
b6386e7
Updates for drain gpu utilization and multiple backends including usi…
stacyroberts Jun 16, 2026
5dc69d0
updates for results storage
stacyroberts Jun 18, 2026
226dcd7
updates to summary output. gen_report for html generation. yamls for …
stacyroberts Jun 19, 2026
deebe2b
selects correct default backend for marker
stacyroberts Jun 19, 2026
9183497
more attention sweep yamls
stacyroberts Jun 19, 2026
ac92383
added recovery from failure. gen_report.py addes header file for easi…
stacyroberts Jun 22, 2026
7d97b59
update config
stacyroberts Jun 22, 2026
d8ba38c
clean up messages on attn backends
stacyroberts Jun 23, 2026
f7e9e3a
[BFCL] Add multiturn test group
tarukumar Jun 9, 2026
cdacc2e
Revert "[BFCL] Add multiturn test group for tool accuracy" (#15) (#16)
khluu Jun 11, 2026
e4864c8
[BFCL] Add test group support (#17)
tarukumar Jun 11, 2026
8181798
Switch vllm_bench from speed_bench to random + ignore_eos (#20)
khluu Jun 11, 2026
5e229a6
[BFCL] Remove parallel tool calling test from gpt-oss model (#21)
tarukumar Jun 12, 2026
1b79cbe
Merge branch 'main' into sroberts-backend-test
stacyroberts Jun 23, 2026
baf3790
Fixed errant call in parse_workload and adding moe sweep
stacyroberts Jun 25, 2026
d60e107
Moved attention sweep logic out of main run.sh file
stacyroberts Jun 25, 2026
141fadf
adding in moe-backend sweep. Still issue with report generation
stacyroberts Jun 26, 2026
302eb8a
New moe sweep file
stacyroberts Jun 26, 2026
ae45227
another updated moe sweep
stacyroberts Jun 26, 2026
b354ac6
AFO-LLM port, two yaml files, one generate report for this test file.
stacyroberts Jun 29, 2026
e0636cd
adding in deepseek afo-llm port yaml files
stacyroberts Jun 30, 2026
352e398
fixed server args issue causing bad token id issue.
stacyroberts Jun 30, 2026
8482a36
vllm-ci yaml files
stacyroberts Jun 30, 2026
3092810
more report scripts
stacyroberts Jul 6, 2026
07998fc
Fixed memory utilization issue for ds, added in the ds vllm_ci yaml f…
stacyroberts Jul 7, 2026
48fbd98
Minor fixes and addition of mxfp4 ds
stacyroberts Jul 7, 2026
1b10a47
Update to kill child processes (tee was causing issues with trailing …
stacyroberts Jul 7, 2026
e194cd4
Merge remote-tracking branch 'vllm-project/main' into sroberts-backen…
stacyroberts Jul 7, 2026
0155613
Corrections to script plus one configuration addition
stacyroberts Jul 14, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitattributes
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
*.sh text eol=lf
*.py text eol=lf
*.yaml text eol=lf
167 changes: 167 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,167 @@
# Byte-compiled / optimized / DLL files
__pycache__/
*.py[cod]
*$py.class

# C extensions
*.so

# Distribution / packaging
.Python
build/
develop-eggs/
dist/
downloads/
eggs/
.eggs/
lib64/
parts/
sdist/
var/
wheels/
share/python-wheels/
*.egg-info/
.installed.cfg
*.egg
MANIFEST

# PyInstaller
# Usually these files are written by a python script from a template
# before PyInstaller builds the exe, so as to inject date/other infos into it.
*.manifest
*.spec

# Installer logs
pip-log.txt
pip-delete-this-directory.txt

# Unit test / coverage reports
htmlcov/
.tox/
.nox/
.coverage
.coverage.*
.cache
nosetests.xml
coverage.xml
*.cover
*.py,cover
.hypothesis/
.pytest_cache/
cover/

# Translations
*.mo
*.pot

# Django stuff:
*.log
local_settings.py
db.sqlite3
db.sqlite3-journal

# Flask stuff:
instance/
.webassets-cache

# Scrapy stuff:
.scrapy

# Sphinx documentation
docs/_build/

# PyBuilder
.pybuilder/
target/

# Jupyter Notebook
.ipynb_checkpoints

# IPython
profile_default/
ipython_config.py

# pyenv
# For a library or package, you might want to ignore these files since the code is
# intended to run in multiple environments; otherwise, check them in:
# .python-version

# pipenv
# According to pypa/pipenv#598, it is recommended to include Pipfile.lock in version control.
# However, in case of collaboration, if having platform-specific dependencies or dependencies
# having no cross-platform support, pipenv may install dependencies that don't work, or not
# install all needed dependencies.
#Pipfile.lock

# poetry
# Similar to Pipfile.lock, it is generally recommended to include poetry.lock in version control.
# This is especially recommended for binary packages to ensure reproducibility, and is more
# commonly ignored for libraries.
# https://python-poetry.org/docs/basic-usage/#commit-your-poetrylock-file-to-version-control
#poetry.lock

# pdm
# Similar to Pipfile.lock, it is generally recommended to include pdm.lock in version control.
#pdm.lock
# pdm stores project-wide configurations in .pdm.toml, but it is recommended to not include it
# in version control.
# https://pdm.fming.dev/latest/usage/project/#working-with-version-control
.pdm.toml
.pdm-python
.pdm-build/

# PEP 582; used by e.g. github.com/David-OConnor/pyflow and github.com/pdm-project/pdm
__pypackages__/

# Celery stuff
celerybeat-schedule
celerybeat.pid

# SageMath parsed files
*.sage.py

# Environments
.env
.venv
env/
venv/
ENV/
env.bak/
venv.bak/

# Spyder project settings
.spyderproject
.spyproject

# Rope project settings
.ropeproject

# mkdocs documentation
/site

# mypy
.mypy_cache/
.dmypy.json
dmypy.json

# Pyre type checker
.pyre/

# pytype static type analyzer
.pytype/

# Cython debug symbols
cython_debug/

# results directory
results/
# test output files
*.txt
*.html

# PyCharm
# JetBrains specific template is maintained in a separate JetBrains.gitignore that can
# be found at https://github.com/github/gitignore/blob/main/Global/JetBrains.gitignore
# and can be added to the global gitignore or merged into this file. For a more nuclear
# option (not recommended) you can uncomment the following to ignore the entire idea folder.
#.idea/
28 changes: 27 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@ Each recipe is one `(model, hardware, set of tasks)` combination. The Buildkite
workloads/ one YAML per (model, hardware) recipe
lib/ orchestrator (run.sh), helpers, GPU profiles
.buildkite/ pipeline bootstrap and step generator
gen_report.py generate HTML benchmark reports from results/
CLAUDE.md agent conventions and detailed Buildkite workflow
```

Expand All @@ -26,7 +27,7 @@ CLAUDE.md agent conventions and detailed Buildkite workflow

A recipe has top-level metadata plus up to three eval blocks:

- **`vllm:`** — *how the server runs.* Defines what model to serve and how (`model`, `serve_args`, optional image/env overrides). Required.
- **`vllm:`** — *how the server runs.* Defines what model to serve and how (`model`, `serve_args`, optional image/env overrides, optional `attention_backends` or `moe_backends` list). Required.
- **`lm_eval:`** — *what accuracy to measure.* Lists lm-evaluation-harness tasks to run against the live server (e.g. `gsm8k`, `aime25`). Each task's score is saved under `results/<name>/<task-name>/`. Optional.
- **`vllm_bench:`** — *what perf to measure.* Lists `vllm bench serve` configs (input/output lengths, concurrency, dataset). Raw JSON is saved and ingested into the perf dashboard. Optional.
- **`bfcl:`** — *function-calling eval.* Runs [BFCL](https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard) test categories against the live server. Some models need `--enable-auto-tool-choice` and `--tool-call-parser` in `serve_args`. Results are transformed to lm_eval format and ingested as `bfcl_<category>` tasks. Optional.
Expand All @@ -47,6 +48,13 @@ vllm: # how the server is brought up
serve_args: >- # appended to `vllm serve <model>`; word-split
-dp 8 --enable-expert-parallel
--trust-remote-code
attention_backends: # optional; list of VLLM_ATTENTION_BACKEND values
- FLASH_ATTN # when set, the full eval suite runs once per
- FLASHINFER # backend; results land in attn-<BACKEND>/ subdirs
# moe_backends: # optional; list of --moe-backend values
# - default # "default" means no --moe-backend flag (vLLM picks)
# - AITER # results land in moe-<BACKEND>/ subdirs
# attention_backends and moe_backends are mutually exclusive in a single workload

lm_eval: # accuracy tasks (optional)
model_args: # workload-level defaults, merged into every task
Expand Down Expand Up @@ -85,6 +93,8 @@ vllm_bench: # perf runs (optional) — fed to the perf dashboard

A few things worth knowing:

- **`vllm.attention_backends`** is an optional list of vLLM attention backend names (`FLASH_ATTN`, `FLASHINFER`, `XFORMERS`, `TRITON_ATTN`, `TRITON_MLA`, `ROCM_FLASH`, `PAGED_ATTENTION`,`ROCM_AITER_FA`,`ROCM_AITER_UNIFIED_ATTN`, `ROCM_ATTN`,`ROCM_AITER_MLA`,`ROCM_AITER_MLA_SPARSE`, `ROCM_AITER_TRITON_MLA`). When set, the orchestrator starts the server once per backend — adding `--attention-backend $ATTN_BACKEND` — and runs the complete eval suite (bench, lm_eval, bfcl) for each. Results are stored under `results/<name>/attn-<BACKEND>/` so every backend gets its own isolated output directory. Without this field, the server starts once with whatever attention backend vLLM selects by default and results go to `results/<name>/` as usual. See `workloads/attn-sweep-gpt-oss-120b-mi355x.yaml` for an example.
- **`vllm.moe_backends`** is an optional list of vLLM `--moe-backend` values (`default`, `AITER`, `TRITON`, `FUSED_MOE`, `deep_gemm_mega_moe`). When set, the orchestrator starts the server once per backend — adding `--moe-backend $MOE_BACKEND` for non-default values — and runs the complete eval suite for each. Results are stored under `results/<name>/moe-<BACKEND>/`. The special value `"default"` means no `--moe-backend` flag (vLLM picks the backend automatically). **Do not include `--moe-backend` in `serve_args` when using this field** — the sweep script owns that flag. `attention_backends` and `moe_backends` are mutually exclusive in a single workload. See `workloads/moe_sweep_deepseek_r1_0528_mi355x.yaml` for an example.
- **`gpu`** must match a key in `lib/gpu_profiles.yaml`. The profile sets the Buildkite queue, default image, HF cache path, and baseline env vars.
- **`nightly`** controls only the nightly schedule. Recipes with `nightly: false` (or omitted) are still triggerable explicitly via the `WORKLOADS` env var.
- **`lm_eval.tasks` is a list** because each entry runs as a separate `lm_eval` invocation — `--num_fewshot` is a single global flag, so different shot counts need separate runs. Each task's results land in `results/<name>/<task-name>/`.
Expand Down Expand Up @@ -138,6 +148,22 @@ A real run needs a GPU host with Docker, vLLM, and lm-eval available:

Locally, you can smoke-test recipe changes without a GPU — see `CLAUDE.md` for the parser stub and shell-syntax checks.

## Benchmark reports

After a run completes, generate interactive HTML reports from the `results/` directory:

```bash
python3 gen_report.py
```

This writes one `benchmark-<model>.html` per model directory found under `results/`, plus a `benchmark-index.html` wrapper. Open `benchmark-index.html` in a browser to tab between all models in one page — each model's report loads on demand when its tab is clicked.

If reports from previous rounds are already present in the directory, `benchmark-index.html` will include them alongside any newly generated ones, so the index always covers every available model regardless of which models were in the current run.

Each per-model report shows attention backend results side by side, with tabs for each input sequence length, color-coded best/worst values per metric, and percentage deltas relative to the default backend.

**MoE backend sweep reports** are generated automatically alongside the attention sweep reports. For each model directory that contains `moe-*` result subdirectories, `gen_report.py` writes a `moe-benchmark-<model>.html` file. A `moe-benchmark-index.html` landing page is also produced, covering all models with MoE sweep data. Open it in a browser to tab between models — same layout and features as the attention backend report.

## Agents

`CLAUDE.md` has conventions for AI agents working in this repo: smoke-testing changes, launching Buildkite builds for a chosen branch/commit, and the AI-assistance disclosure rule for PRs and commits.
Loading