Skills to guide Claude Code, Codex, and other coding agents on using the Weights & Biases AI developer platform to train models and build agents.
- Log metrics and rich media during model training and fine-tuning
- Track model training experiments
- Analyze runs and experiment results to understand how the model is learning
- Tune hyperparameters
- Trace agentic AI applications
- Analyze traces and classify them into failure modes
- Evaluate models with labeled datasets
- Run online evaluations for production monitoring
npx skills add wandb/skillsThen set your W&B API key:
export WANDB_API_KEY=<your-key>
npx skillsis a utility for installing skills into major coding agent CLIs. Use--globalto install for all projects, or--agent <name>to target a specific agent. See the npx skills docs for more details.
The skill helpers are validated against these versions. Older SDKs expose some of the methods used here as unimplemented stubs, so pin at or above the floor rather than relying on whatever is already installed.
| Package | Floor | Validated against |
|---|---|---|
| Python | >=3.13 |
3.13 |
wandb |
>=0.28.1 |
0.28.1 |
wandb-workspaces |
>=0.4.4 |
0.4.4 (pulled in by the [workspaces] extra) |
weave |
>=0.52.41 |
0.52.41 |
The workspaces extra is required for the Workspaces helpers:
uv run --with 'wandb[workspaces]>=0.28.1' --with 'weave>=0.52.41' python your_script.py0.28.1 is a floor rather than a preference: at least one helper branches on
behavior that changed after 0.28.0, so >=0.28.0 is not sufficient.
| Skill | Description | Status |
|---|---|---|
wandb-primary |
Broad W&B project analysis and operations across runs, Artifacts, Registry, Weave, Reports, Workspaces, and Launch. | experimental |
wandb-eval-tables |
Non-destructive conversion of W&B Table artifacts into bounded, verified EvalTable previews. | experimental |
wandb-autoresearch |
Bounded training research through W&B Launch, including readiness checks, serial trials, comparison, and resumable state. | experimental |
We maintain Skill Bench in this repository to evaluate public skill changes across coding agents and task categories. Skill Bench uses W&B Agent Factory as the eval runtime for task definitions, agent profiles, sandbox execution, and structured bench rows.
Pull requests run package validation by default. A maintainer can trigger live Skill Bench runs for larger changes.
Plan a local benchmark without model calls:
python3 -m skillbench.cli plan \
--wbaf-root ../WandBAgentFactory \
--candidate-ref HEAD \
--skill wandb-primary| Category | Tasks | Claude Code (sonnet4.6) |
Codex (gpt-5.3-codex) |
|---|---|---|---|
| Weave analysis | 26 | 97%* | 63%* |
| Weave tooling | 11 | 95%* | 83%* |
| Model training | 8 | 90%* | 85%* |
| LLM finetuning & RL analysis | 14 | 72%* | 86%* |
| Failure & outlier detection | 8 | 86%* | 63%* |
*Pass rates are +/- 3%. Many tasks span multiple categories.
See CONTRIBUTING.md.