Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
66 commits
Select commit Hold shift + click to select a range
bded2db
[docs-agent] Update architecture docs: Maestro → kube-applier + hyper…
redhat-chai-bot Jul 24, 2026
72fc018
Apply suggestions from code review
typeid Jul 24, 2026
6fd2ada
Include all open chai-bot PRs in the dashboard regardless of labels
typeid Jul 24, 2026
8e1f517
Fixing missed settings - egress rule for ALB, limit the number of ret…
Makdaam Jul 23, 2026
7a95a1d
Merge pull request #699 from Makdaam/ROSAENG-61610
openshift-merge-bot[bot] Jul 24, 2026
785a97d
Spelling fix.
Makdaam Jul 24, 2026
fe036d8
Merge pull request #702 from Makdaam/ROSAENG-61610
openshift-merge-bot[bot] Jul 24, 2026
dded10d
Adding SRE documentation on SRE-UI ALB.
Makdaam Jul 24, 2026
c9e7f97
Enabling "can_edit" for viewers. This doesn't allow the Viewers to sa…
Makdaam Jul 24, 2026
5ae8e64
Merge pull request #703 from Makdaam/ROSAENG-61610
openshift-merge-bot[bot] Jul 24, 2026
d46ae4e
Merge pull request #704 from Makdaam/ROSAENG-62511
openshift-merge-bot[bot] Jul 24, 2026
2b1e754
Add refresh flag to ArgoCD automatic retries.
Makdaam Jul 24, 2026
896abf9
Merge pull request #705 from Makdaam/ROSAENG-62507
openshift-merge-bot[bot] Jul 24, 2026
0cf2f60
Removing Loki from the exposed SRE UI apps.
Makdaam Jul 24, 2026
64a9e89
Merge pull request #706 from Makdaam/ROSAENG-61610
openshift-merge-bot[bot] Jul 24, 2026
c2a8f00
Merge pull request #700 from redhat-chai-bot/docs/update-maestro-kube…
openshift-merge-bot[bot] Jul 24, 2026
948528c
Merge pull request #701 from typeid/fix_pr_dashboardd
openshift-merge-bot[bot] Jul 24, 2026
702939c
Add ElastiCache Valkey infrastructure for rate limiting
ravitri Jul 27, 2026
d3d742d
Add Platform API rate limiting config, alerting, and documentation
ravitri Jul 27, 2026
ea508f8
Removing bastion access for Valkey
ravitri Jul 27, 2026
b660eb4
Updating platform-api image
ravitri Jul 27, 2026
065970a
Address SRE review: rename alert, fix thresholds, add redisTimeout
ravitri Jul 27, 2026
023fa81
Harden Valkey module: explicit KMS policy, drop open egress, fix docs
ravitri Jul 27, 2026
8d4866b
Fixing check-docs
ravitri Jul 27, 2026
14c8bc1
Surface SSM error details when GitHub token fetch fails
typeid Jul 27, 2026
e4b7220
fix: wire hyperfleet_db_deletion_protection to Terraform in CI
theautoroboto Jul 27, 2026
d863f44
Downgrade RateLimitFailOpenActive severity to warning
ravitri Jul 28, 2026
1e088e0
Merge pull request #714 from theautoroboto/fix/hyperfleet-db-deletion…
openshift-merge-bot[bot] Jul 28, 2026
c7a61da
fix: remove CLM/Maestro/Sentinel references from Grafana dashboards
typeid Jul 28, 2026
cfcd6fa
Merge pull request #680 from ravitri/ratelimit
openshift-merge-bot[bot] Jul 28, 2026
8699f64
feat: add DB state collection and rename collect-logs to dump-env
typeid Jul 28, 2026
b487a0a
Merge pull request #715 from typeid/ROSAENG-62356_fix_tooling
openshift-merge-bot[bot] Jul 29, 2026
5079257
Merge pull request #711 from typeid/fix_error_output_eph
openshift-merge-bot[bot] Jul 29, 2026
8a49e72
Use the quay image
psav Jul 27, 2026
1f7e806
ROSAENG-62642: Also use the konflux image for platform-api
rrp-bot Jul 29, 2026
18e419b
fix: pin AWS Terraform provider to ~> 6.56.0
typeid Jul 29, 2026
a63be5b
fix: add force_destroy to SRE ALB access logs S3 bucket
theautoroboto Jul 29, 2026
d5bd785
ROSAENG-62981: fix teardown swallowing pipeline failures and collecti…
typeid Jul 29, 2026
6402157
Merge pull request #720 from theautoroboto/fix/s3-alb-logs-force-destroy
openshift-merge-bot[bot] Jul 29, 2026
f630a1d
Merge pull request #719 from typeid/fix_terraform_provider
typeid Jul 29, 2026
51b01f2
Bump platform-api to latest
typeid Jul 30, 2026
d4980ef
Merge pull request #718 from rrp-bot/psav/use-quay-image
openshift-merge-bot[bot] Jul 30, 2026
fa534d8
Merge pull request #721 from typeid/fix_provision_error_swallow
typeid Jul 30, 2026
530b02f
ROSAENG-63020: Enable force_destroy on OIDC S3 bucket for ephemeral e…
typeid Jul 30, 2026
211f9b7
Merge pull request #722 from typeid/fix_teardown
openshift-merge-bot[bot] Jul 30, 2026
5fd5a0d
Fix Terraform state drift on default_4xx gateway response
redhat-chai-bot Jul 30, 2026
c285a75
[docs-agent] Update .spec/ references: collect-logs → dump-env
redhat-chai-bot Jul 31, 2026
8c725f7
Fix ephemeral provider pushing to wrong fork repo name
typeid Jul 31, 2026
95add67
NO-JIRA: Bump platform-api and hyperfleet-operator to f7ec7d5
typeid Jul 31, 2026
e6f6068
Merge pull request #728 from redhat-chai-bot/docs/update-spec-refs-20…
openshift-merge-bot[bot] Jul 31, 2026
bb48b87
Merge pull request #726 from redhat-chai-bot/fix/default-4xx-response…
openshift-merge-bot[bot] Jul 31, 2026
55c0244
Merge pull request #732 from typeid/fix/bump-platform-api-image-tag
openshift-merge-bot[bot] Aug 3, 2026
5ed0721
Merge pull request #731 from typeid/fix-ephemeral-fork-repo-name
openshift-merge-bot[bot] Aug 4, 2026
9c8574c
docs: replace stale Maestro/CLM references with current architecture
typeid Jul 29, 2026
6789725
Merge pull request #717 from typeid/cbusse/docs-cleanup-maestro-clm
openshift-merge-bot[bot] Aug 5, 2026
96b3a68
ROSAENG-60886: fix: improve Helm download reliability with retry logi…
theautoroboto Aug 6, 2026
f169286
[docs-agent] Fix stale Maestro reference in testing strategy doc
redhat-chai-bot Aug 7, 2026
585a627
Merge pull request #735 from redhat-chai-bot/docs/fix-stale-maestro-r…
openshift-merge-bot[bot] Aug 7, 2026
7dac4a6
ROSAENG-64794: Add AVP PolicyTemplate permissions to frontend API role
cdoan1 Aug 8, 2026
4f0abbd
Add regional control plane architecture design doc
typeid Jul 31, 2026
36f10a3
Merge pull request #737 from cdoan1/rosaeng-64794-avp-store
openshift-merge-bot[bot] Aug 10, 2026
c84ca0d
Merge pull request #730 from typeid/control-plane-architecture
typeid Aug 11, 2026
16480fa
Update hash4 description to reflect uniqueness enforcement
typeid Aug 11, 2026
864eb2e
Merge pull request #739 from typeid/avoid_dns_collision
typeid Aug 11, 2026
fbd99d5
NO-JIRA: fix: ephemeral config override merge + Go-template variable …
theautoroboto Aug 11, 2026
6c19f34
feat: add refresh-app script to platform image for ArgoCD hard-refresh
rrp-bot Aug 12, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 5 additions & 5 deletions .chai-bot/rosa_hyperfleet_ci_daily_health_report.md
Original file line number Diff line number Diff line change
Expand Up @@ -132,7 +132,7 @@ For each job whose **latest run failed**, produce a **separate threaded reply**
- Fetch scope based on Prow analysis: RC-only, MC + RC, or both if unclear
- If S3 logs are inaccessible, report the specific error — classification ceiling becomes Unclear

3. **Git commit correlation** (Step 5c in ci-troubleshooter) — **MANDATORY.** Identify the last passing run, find all commits between last-good and current-bad, and examine commits touching the failing component. Also check `rosa-hyperfleet-api` for API/CLM failures.
3. **Git commit correlation** (Step 5c in ci-troubleshooter) — **MANDATORY.** Identify the last passing run, find all commits between last-good and current-bad, and examine commits touching the failing component. Also check `rosa-hyperfleet-api` for API/hyperfleet-operator failures.

**S3 log handling:** Always extract tar.gz locally for full analysis. Clean up downloaded files immediately after analysis is complete — never leave S3 logs on disk between runs. See Step 5b in `.claude/agents/ci-troubleshooter.md` for the full procedure.

Expand Down Expand Up @@ -247,10 +247,10 @@ integration: ✅ ✅ ❌ ✅ ✅ ✅ ✅ ✅ ❌

🔧 *Genuine* — E2E test `TestClusterCreation` timed out waiting for hosted cluster to become ready.
Evidence: Prow ✅ | S3 Logs ✅ | Git History ✅ | Trend ✅
Root cause: MC maestro-agent pod in CrashLoopBackOff due to MQTT connection failure — incorrect broker endpoint in ArgoCD values.
S3 Log Evidence: `maestro-agent/pods/agent-xyz/agent/logs/current.log` — 47x `CONNACK refused: not authorized`; pod status: CrashLoopBackOff
Suspect Commits: `a1b2c3d` — `feat(argocd): update maestro broker endpoint` — touches `argocd/config/management-cluster/maestro/`
Consecutive failures (2 days): same root cause as Jun 29 — maestro CONNACK failure with identical error signature.
Root cause: MC kube-applier pod in CrashLoopBackOff due to DynamoDB connectivity failure — incorrect table name in ArgoCD values.
S3 Log Evidence: `kube-applier/pods/kube-applier-xyz/kube-applier/logs/current.log` — 47x `ResourceNotFoundException: Requested resource not found`; pod status: CrashLoopBackOff
Suspect Commits: `a1b2c3d` — `feat(argocd): update kube-applier DynamoDB config` — touches `argocd/config/management-cluster/kube-applier/`
Consecutive failures (2 days): same root cause as Jun 29 — kube-applier DynamoDB failure with identical error signature.

Most recent failure: <url|Build #1234> (Jun 30)
Failing since: Jun 29 (2 consecutive days)
Expand Down
4 changes: 2 additions & 2 deletions .claude/agents/architect.md
Original file line number Diff line number Diff line change
Expand Up @@ -80,8 +80,8 @@ Issues:
3. Database access violates separation of concerns

Recommendations:
1. Subscribe to cluster-events topic for event-driven processing
2. Implement status reporting via REST API (see clm-gcp-environment-validation example)
1. Use controller-runtime watches for event-driven processing
2. Implement status reporting via CR status updates (see hyperfleet-operator controller pattern)
3. Use GET /api/v1/clusters/{id} to fetch data instead of direct DB access

Consider documenting this controller pattern in a new design decision.
Expand Down
86 changes: 50 additions & 36 deletions .claude/agents/ci-troubleshooter.md
Original file line number Diff line number Diff line change
Expand Up @@ -121,16 +121,16 @@ Single `e2e-tests` step — fetch and analyze `<artifacts-url>/e2e-tests/build-l

Single step matching job name — fetch `<artifacts-url>/<job-name>/build-log.txt`.

## Step 5b: Pull Cluster Logs from S3 (MANDATORY for cluster-backed jobs)
## Step 5b: Pull Environment Dumps from S3 (MANDATORY for cluster-backed jobs)

When e2e tests fail, the CI job collects pod logs from the RC and MC clusters and uploads them to S3. These logs are **not** included in the public Prow artifacts (they may contain secrets), but the S3 URIs are printed in the e2e build log.
When e2e tests fail, the CI job dumps environment state from the RC and MC clusters and uploads it to S3. RC dumps include both Kubernetes logs and a PostgreSQL database snapshot; MC dumps contain Kubernetes logs only. These dumps are **not** included in the public Prow artifacts (they may contain secrets), but the S3 URIs are printed in the e2e build log.

**Applies to:** `on-demand-e2e`, `nightly-ephemeral`, `nightly-integration` — jobs that provision clusters and produce S3 log archives. **Does not apply to** validation jobs (`terraform-validate`, `helm-lint`, `check-rendered-files`, `check-docs`) which have no cluster logs — those jobs are classified using Prow build logs and git history only.

**S3 log analysis is mandatory for all cluster-backed job failure classifications.** You MUST download, extract, and analyze S3 logs before classifying any failure from these jobs. A classification of Genuine or Flake is not valid without S3 log evidence. Use the Prow build logs from Step 5 to determine which clusters to fetch logs for:

- **RC-only failure** (e.g., provision failure, API error, ArgoCD sync issue on RC, maestro-server error): fetch **only RC logs** from S3.
- **MC failure or RC↔MC interaction** (e.g., maestro-agent errors, HyperShift issues, hosted cluster failures, connectivity between RC and MC): fetch **both RC and MC logs** from S3 — MC failures often have an RC-side root cause.
- **RC-only failure** (e.g., provision failure, API error, ArgoCD sync issue on RC, hyperfleet-operator error): fetch **only RC logs** from S3.
- **MC failure or RC↔MC interaction** (e.g., kube-applier errors, HyperShift issues, hosted cluster failures, connectivity between RC and MC): fetch **both RC and MC logs** from S3 — MC failures often have an RC-side root cause.
- **Unclear scope**: fetch **both RC and MC logs**.

If S3 logs are inaccessible for any reason (credentials, expired logs, network issues), you MUST still attempt the access and report the specific error. When S3 logs cannot be obtained, the classification ceiling is **⚠️ Unclear** — you cannot claim Genuine or Flake without S3 evidence.
Expand Down Expand Up @@ -161,7 +161,7 @@ There will be one URI per cluster (RC + each MC). The bucket names follow the pa
- RC: `bastion-log-collection-<regional-account-id>-<region>-an`
- MC: `bastion-log-collection-<management-account-id>-<region>-an`

### Fetching the logs
### Fetching the dumps

**Always extract tar.gz archives locally for full analysis.** Download to a temp directory, extract, perform broad grep-based analysis across all namespaces, and clean up after:

Expand All @@ -173,9 +173,9 @@ trap 'rm -rf "$LOGDIR"' EXIT
# Use separate subdirectories for RC and MC to avoid archive name collisions
mkdir -p "$LOGDIR/rc" "$LOGDIR/mc"

aws s3 cp s3://bastion-log-collection-<account>-<region>-an/collect-logs-<id>.tar.gz \
aws s3 cp s3://bastion-log-collection-<account>-<region>-an/dump-env-<id>.tar.gz \
"$LOGDIR/rc/" --profile <PROFILE> && \
tar xzf "$LOGDIR/rc"/collect-logs-*.tar.gz -C "$LOGDIR/rc"
tar xzf "$LOGDIR/rc"/dump-env-*.tar.gz -C "$LOGDIR/rc"

# Perform broad analysis: grep across ALL namespaces, not just suspected ones
grep -rli "error\|fail\|crash\|panic\|fatal\|timeout\|refused\|denied" "$LOGDIR/rc"/inspect-logs/namespaces/ 2>/dev/null
Expand All @@ -189,12 +189,12 @@ Fetch logs based on the failure scope determined from Prow artifacts. Use the ap
LOGDIR=$(mktemp -d /tmp/ci-logs-XXXXXX)
trap 'rm -rf "$LOGDIR"' EXIT
mkdir -p "$LOGDIR/rc" "$LOGDIR/mc"
aws s3 cp s3://bastion-log-collection-720644165472-us-east-1-an/collect-logs-<id>.tar.gz \
aws s3 cp s3://bastion-log-collection-720644165472-us-east-1-an/dump-env-<id>.tar.gz \
"$LOGDIR/rc/" --profile chai-rc-ci && \
tar xzf "$LOGDIR/rc"/collect-logs-*.tar.gz -C "$LOGDIR/rc"
aws s3 cp s3://bastion-log-collection-129678139271-us-east-1-an/collect-logs-<id>.tar.gz \
tar xzf "$LOGDIR/rc"/dump-env-*.tar.gz -C "$LOGDIR/rc"
aws s3 cp s3://bastion-log-collection-129678139271-us-east-1-an/dump-env-<id>.tar.gz \
"$LOGDIR/mc/" --profile chai-mc-ci && \
tar xzf "$LOGDIR/mc"/collect-logs-*.tar.gz -C "$LOGDIR/mc"
tar xzf "$LOGDIR/mc"/dump-env-*.tar.gz -C "$LOGDIR/mc"
# Analyze $LOGDIR/rc/inspect-logs/ and $LOGDIR/mc/inspect-logs/
```

Expand Down Expand Up @@ -222,38 +222,51 @@ Classification ceiling is Unclear — cannot claim Genuine or Flake without S3 e

Do **not** stop the investigation — proceed with whatever information is available from the Prow artifacts and git history. However, **without successfully analyzed S3 log evidence, the maximum classification confidence is ⚠️ Unclear.** You cannot classify as Genuine or Flake without having analyzed S3 logs.

### Analyzing the logs
### Analyzing the dumps

Once extracted, the logs are organized as:
Once extracted, the dump is organized as:

```
```text
inspect-logs/
namespaces/<namespace>/
<resource>.yaml # Resource definitions
<resource>.yaml # Resource definitions (pods, services, etc.)
pods/<pod-name>/<container>/logs/
current.log # Current container log
previous.log # Previous container log (if restarted)
cluster-scoped-resources/ # Cluster-scoped CRs (nodes, etc.)
<group>/<kind>/<name>.yaml
<crd-group>/ # CRD instances collected by oc adm inspect
<kind>.yaml # e.g., hostedclusters, nodepools, applications
db-state/ # RC only — hyperfleet-db state dump
resource-summary.txt # Tabular listing of all kubernetes_resources rows
resources/<kind>/<name>.json # Individual resource objects (spec, status, metadata)
```

Key namespaces and what to look for:

| Cluster | Namespace | What to check |
| ------- | ---------------- | ----------------------------------------------------------- |
| RC | `maestro-server` | Server MQTT connectivity, resource bundle creation |
| RC | `platform-api` | API errors, registration failures |
| RC | `argocd` | Sync failures, application health |
| MC | `maestro-agent` | Agent MQTT connectivity (CONNACK errors), work agent status |
| MC | `argocd` | Sync failures on MC applications |
| MC | `hypershift` | HyperShift operator errors |
| Cluster | Namespace | What to check |
| ------- | -------------- | ------------------------------------------------------------ |
| RC | `hyperfleet` | Operator reconciliation, Manifest CR and hyperfleet-db state |
| RC | `platform-api` | API errors, registration failures |
| RC | `argocd` | Sync failures, application health |
| MC | `kube-applier` | DynamoDB Streams connectivity, resource apply status |
| MC | `argocd` | Sync failures on MC applications |
| MC | `hypershift` | HyperShift operator errors |

**Other dump components** (not Kubernetes namespaces):

| Cluster | Directory | What to check |
| ------- | ----------- | ------------------------------------------------------------------------------------- |
| RC | `db-state/` | Hyperfleet DB contents — resource summary and individual JSON objects (RC dumps only) |

For maestro connectivity issues specifically, check:
For resource distribution issues specifically, check:

```bash
# Agent connection errors
grep -i "connack\|connect\|error\|fail" /tmp/<prefix>-mc01-logs/inspect-logs/namespaces/maestro-agent/pods/*/agent/agent/logs/current.log
# kube-applier errors on MC
grep -i "error\|fail\|dynamo" /tmp/<prefix>-mc01-logs/inspect-logs/namespaces/kube-applier/pods/*/kube-applier/logs/current.log

# Server-side issues
grep -i "error\|fail\|connect" /tmp/<prefix>-regional-logs/inspect-logs/namespaces/maestro-server/pods/*/service/service/logs/current.log
# Operator errors on RC
grep -i "error\|fail\|reconcil" /tmp/<prefix>-regional-logs/inspect-logs/namespaces/hyperfleet/pods/*/manager/logs/current.log
```

### S3 log retention
Expand All @@ -271,11 +284,12 @@ The key question is: **did anything change between the last passing and current
# Provision failure → terraform/, scripts/buildspec/, ci/ephemeral-provider/
# E2E test failure → ci/e2e-tests.sh, ci/e2e-platform-api-test.sh
# ArgoCD sync failure → argocd/
# Maestro failure → argocd/config/*/maestro*
# Hyperfleet operator failure → argocd/config/regional-cluster/hyperfleet*
# kube-applier failure → argocd/config/management-cluster/kube-applier*
# Platform API failure → (check rosa-hyperfleet-api repo)
```

**Cross-repo:** the git commands above cover `rosa-hyperfleet` only. For API/CLM failures, also check recent `rosa-hyperfleet-api` commits via `gh api`. Only check `rosa-hyperfleet-cli` if e2e tests invoke CLI commands.
**Cross-repo:** the git commands above cover `rosa-hyperfleet` only. For API/hyperfleet-operator failures, also check recent `rosa-hyperfleet-api` commits via `gh api`. Only check `rosa-hyperfleet-cli` if e2e tests invoke CLI commands.

If a commit strongly correlates with the failure, this is strong evidence for a Genuine classification even on first occurrence.

Expand Down Expand Up @@ -356,7 +370,7 @@ When today's failure is part of a **consecutive failure streak** (2+ days in a r

1. **Collect failure artifacts from each consecutive failing run** — use the job history to identify the streak, then fetch Prow artifacts and S3 logs (selectively, per Step 5b) for at least the current and previous failing runs.
2. **Compare error signatures** — are the failures the same root cause, or did the root cause shift?
- **Same root cause across streak**: reinforce the diagnosis with the additional evidence. Note the streak length (e.g., "failing for 3 consecutive days with the same maestro-agent CONNACK error").
- **Same root cause across streak**: reinforce the diagnosis with the additional evidence. Note the streak length (e.g., "failing for 3 consecutive days with the same kube-applier DynamoDB connectivity error").
- **Root cause shifted**: clearly state that the root cause changed. Identify when it changed and what the new root cause is. This affects PR management (see Step 9).
3. **Aggregate the signal** — a 3-day streak of the same error is much stronger signal than a single failure. Reflect this confidence in the classification (almost certainly Genuine, not Flake).

Expand Down Expand Up @@ -388,16 +402,16 @@ Present findings in this format:
**S3 Log Evidence:**
<Key error patterns found in extracted S3 logs. Include specific log file paths and grep matches. Example:>

- `inspect-logs/namespaces/maestro-agent/pods/agent-xyz/agent/logs/current.log`: 47 occurrences of `CONNACK refused: not authorized`
- `inspect-logs/namespaces/kube-applier/pods/kube-applier-xyz/kube-applier/logs/current.log`: 47 occurrences of `DynamoDB stream error`
- `inspect-logs/namespaces/hypershift/pods/operator-abc/manager/logs/current.log`: `OOMKilled` at 03:42 UTC
- Pod health scan: 2 pods in CrashLoopBackOff (`maestro-agent`, `work-agent`)
- Pod health scan: 2 pods in CrashLoopBackOff (`kube-applier`, `hyperfleet-operator`)
<If S3 logs could not be fetched, state the error and note the classification ceiling.>

**Suspect Commits:**
<Commits between the last passing run and the current failing run that touch relevant paths. Example:>

- `a1b2c3d` — `fix(terraform): update NAT gateway config` — touches `terraform/modules/eks-cluster/` (relevant: provision failure)
- `e4f5g6h` — `feat(argocd): add maestro-agent resource limits` — touches `argocd/config/management-cluster/maestro/` (relevant: maestro-agent OOMKilled)
- `e4f5g6h` — `feat(argocd): add kube-applier resource limits` — touches `argocd/config/management-cluster/kube-applier/` (relevant: kube-applier OOMKilled)
<If no suspect commits found: "No commits between last passing run (<commit>) and current run (<commit>) touch the failing component's paths.">

**Cross-Day Analysis** (if consecutive failures):
Expand Down Expand Up @@ -426,9 +440,9 @@ Share the root cause and raise a fix PR immediately:

1. **Identify the target repo**:
- `rosa-hyperfleet` — Terraform modules, ArgoCD configs, CI scripts, buildspecs
- `rosa-hyperfleet-api` — Platform API, CLM service code
- `rosa-hyperfleet-api` — Platform API, hyperfleet-operator service code
- `rosa-hyperfleet-cli` — CLI tooling
2. **Create a fix branch** — branch from `main`: `chai-bot/fix-<job>-<short-description>` (e.g., `chai-bot/fix-ephemeral-maestro-mqtt-config`).
2. **Create a fix branch** — branch from `main`: `chai-bot/fix-<job>-<short-description>` (e.g., `chai-bot/fix-ephemeral-kube-applier-config`).
3. **Implement the fix** — make the minimal change needed to address the root cause. Follow the project's development guidelines (run `make pre-push` before committing).
4. **Raise the PR** — use `gh pr create` with:
- Title: `fix(<component>): <short description of the fix>`
Expand Down
2 changes: 1 addition & 1 deletion .claude/skills/add-pre-merge/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@ You are helping the user onboard a new component repository for cross-component

If not provided via `$ARGUMENTS`, ask the user for:

1. **Component name** — the directory name under `argocd/config/regional-cluster/` or `argocd/config/management-cluster/` in this repo (e.g., `platform-api`, `maestro-server`). Validate it exists.
1. **Component name** — the directory name under `argocd/config/regional-cluster/` or `argocd/config/management-cluster/` in this repo (e.g., `platform-api`, `hyperfleet`). Validate it exists.
2. **Org and repo** — the GitHub org/repo for the component (e.g., `openshift-online/rosa-hyperfleet-api`). This determines the CI config path in openshift/release.
3. **Branch** — the branch to configure (default: `main`).
4. **Dockerfile path** — path to the Dockerfile in the component repo (default: `Dockerfile`).
Expand Down
6 changes: 3 additions & 3 deletions .spec/002-spec-to-pr-agent/context-requirements.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,9 +7,9 @@ An agentic workflow that takes a feature specification (from a JIRA ticket or de
## Codebase Research Findings

- **Existing agents**: adversary, architect, ci-troubleshooter, code-reviewer, documentation-updater, scope-creep-craig, tech-spec-beck
- **Ephemeral env targets**: `ephemeral-{provision,teardown,resync,swap-branch,list,shell,bastion-rc,bastion-mc,port-forward-*,e2e,collect-logs}`
- **Ephemeral env targets**: `ephemeral-{provision,teardown,resync,swap-branch,list,shell,bastion-rc,bastion-mc,port-forward-*,e2e,dump-env}`
- **E2E testing**: Tests live in `rosa-hyperfleet-api` repo, run via `ci/e2e-tests.sh`, use `make ephemeral-e2e ID=<env-id>`
- **Component repos**: platform-api, maestro-agent, maestro-server, hyperfleet-adapter, hyperfleet-api, hyperfleet-sentinel
- **Component repos**: platform-api, hyperfleet-operator, hyperfleet-db, kube-applier
- **CLI proxy**: Credential-isolating sidecar for `gh` CLI with deny list for destructive commands
- **Config rendering**: `uv run scripts/render.py` for region configs

Expand Down Expand Up @@ -69,7 +69,7 @@ An agentic workflow that takes a feature specification (from a JIRA ticket or de

### Ephemeral Environment Script (`scripts/dev/ephemeral-env.sh`)

- Commands: provision, teardown, resync, swap-branch, shell, bastion, port-forward, e2e, collect-logs, list
- Commands: provision, teardown, resync, swap-branch, shell, bastion, port-forward, e2e, dump-env, list
- State tracking: `.ephemeral-envs` file with KEY=VALUE pairs (ID, REPO, BRANCH, STATE, REGION, API_URL, CI_BRANCH, CREATED)
- Credentials: Fetched from Vault via OIDC, never persisted to disk
- Container-based execution with AWS credentials and API URL injection
Expand Down
2 changes: 1 addition & 1 deletion .spec/002-spec-to-pr-agent/implementation-plan.md
Original file line number Diff line number Diff line change
Expand Up @@ -293,7 +293,7 @@ Parse the user's request from: $ARGUMENTS
- resync: `make ephemeral-resync ID=<id>`
- list: `make ephemeral-list`
- e2e: `make ephemeral-e2e ID=<id>`
- collect-logs: `make ephemeral-collect-logs ID=<id>`
- dump-env: `make ephemeral-dump-env ID=<id>`
- shell: `make ephemeral-shell ID=<id>`
- swap-branch: `make ephemeral-swap-branch ID=<id> BRANCH=<branch>`
...
Expand Down
Loading