Skip to content

Add tiered Grafana dashboards for LLM inference observability - #6

Draft
tgitelman wants to merge 1 commit into
agent-ops-crew:mainfrom
tgitelman:main
Draft

Add tiered Grafana dashboards for LLM inference observability#6
tgitelman wants to merge 1 commit into
agent-ops-crew:mainfrom
tgitelman:main

Conversation

@tgitelman

Copy link
Copy Markdown
  • tier1-failure-saturation-indicators.json: SRE-focused dashboard for immediate failure detection (error rates, latency p50/p90/p99, GPU utilization, scheduler health)
  • tier2-diagnostic-drilldown.json: Diagnostic drill-down dashboard for root cause analysis across well-lit paths (basic serving, routing, prefix caching, P/D disaggregation)

Both dashboards include:

  • Portable datasource variable (DS_PROMETHEUS)
  • Namespace filtering for multi-tenant environments
  • Proper units and legend formats

- tier1-failure-saturation-indicators.json: SRE-focused dashboard for
  immediate failure detection (error rates, latency p50/p90/p99, GPU
  utilization, scheduler health)
- tier2-diagnostic-drilldown.json: Diagnostic drill-down dashboard for
  root cause analysis across well-lit paths (basic serving, routing,
  prefix caching, P/D disaggregation)

Both dashboards include:
- Portable datasource variable (DS_PROMETHEUS)
- Namespace filtering for multi-tenant environments
- Proper units and legend formats
@tgitelman
tgitelman marked this pull request as draft January 18, 2026 08:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant