From 4f0abbdd6f663148a3ec15d0655ea27678dbf07f Mon Sep 17 00:00:00 2001 From: Claudio Busse Date: Fri, 31 Jul 2026 12:12:58 +0200 Subject: [PATCH] Add regional control plane architecture design doc Documents the hyperfleet-operator + hyperfleet-db architecture: controller-runtime reconciliation backed by PostgreSQL instead of etcd. Co-Authored-By: Claude Opus 4.6 --- docs/README.md | 51 ++++---- .../regional-control-plane-architecture.md | 112 ++++++++++++++++++ 2 files changed, 138 insertions(+), 25 deletions(-) create mode 100644 docs/design/regional-control-plane-architecture.md diff --git a/docs/README.md b/docs/README.md index ab0057ef1..a2dc9c733 100644 --- a/docs/README.md +++ b/docs/README.md @@ -20,31 +20,32 @@ The architecture consists of three layers within each region: Detailed architecture and rationale for key technical decisions: -| Document | Topic | -| ------------------------------------------------------------------------------ | ------------------------------------------------------------------ | -| [Alerting Architecture](design/alerting-architecture.md) | Fan-out alert routing (AlertManager, PagerDuty, SNS) | -| [AWS IAM Hosted Cluster Auth](design/aws-iam-hosted-cluster-authentication.md) | AWS IAM authentication for hosted clusters (experimental) | -| [DNS Architecture](design/dns-architecture.md) | Hierarchical DNS with zone shards, `deployment_name`, DNSSEC chain | -| [ECS Fargate Bootstrap](design/fully-private-eks-bootstrap.md) | How fully private EKS clusters are bootstrapped via ECS | -| [FIPS-Only EKS Compute](design/fips-eks-compute.md) | FIPS NodeClass/NodePool strategy for FedRAMP workload nodes | -| [GitOps Cluster Configuration](design/gitops-cluster-configuration.md) | ApplicationSet pattern, progressive deployment, config modes | -| [Infrastructure Logging](design/infrastructure-logging.md) | AWS CloudWatch log groups, KMS encryption, Grafana access | -| [Logging Platform](design/logging-platform.md) | Application-level log collection (Vector + Loki) | -| [HyperFleet Architecture](design/hyperfleet-architecture.md) | Operator + kube-applier architecture, component replacement map | -| [MC Metrics Remote Write](design/mc-metrics-remote-write.md) | MC-to-RC metrics forwarding via RHOBS API Gateway | -| [Monitoring Platform](design/monitoring-platform.md) | Metrics pipeline (Prometheus + Thanos) | -| [Pipeline-Based Lifecycle](design/pipeline-based-lifecycle.md) | CodePipeline hierarchy for cluster provisioning | -| [Rate Limiting](design/rate-limiting-architecture.md) | Per-account rate limiting for Platform API | -| [Regional Account Minting](design/regional-account-minting.md) | AWS account structure and minting pipelines | -| [Regional OIDC Ownership](design/regional-oidc-ownership.md) | Shared OIDC bucket per region, cross-account MC writes | -| [Spec-to-PR Agent](design/spec-to-pr-agent.md) | AI agent workflow for spec-driven implementation | -| [SRE UI Access](design/sre-ui-access.md) | ALB + OIDC access to SRE UIs replacing SSM port-forward | -| [Terraform Resource Adoption](design/terraform-resource-adoption.md) | Idempotent import of auto-created AWS resources into Terraform | -| [Testing Strategy](design/testing-strategy.md) | Ephemeral and long-lived test environments | -| [Thanos Metrics Infrastructure](design/thanos-metrics-infrastructure.md) | Thanos S3 storage, operator, and Pod Identity setup | -| [ZOA Architecture](design/zoa-architecture.md) | Zero Operator Access — system components, flows, infrastructure | -| [ZOA Trusted Actions](design/zoa-trusted-actions.md) | TA template format, API design, CLI, dispatch flow | -| [ZOA Security Model](design/zoa-security-model.md) | SA isolation, RBAC, audit trail, threat model, FIPS | +| Document | Topic | +| ------------------------------------------------------------------------------------ | ------------------------------------------------------------------------ | +| [Alerting Architecture](design/alerting-architecture.md) | Fan-out alert routing (AlertManager, PagerDuty, SNS) | +| [AWS IAM Hosted Cluster Auth](design/aws-iam-hosted-cluster-authentication.md) | AWS IAM authentication for hosted clusters (experimental) | +| [DNS Architecture](design/dns-architecture.md) | Hierarchical DNS with zone shards, `deployment_name`, DNSSEC chain | +| [ECS Fargate Bootstrap](design/fully-private-eks-bootstrap.md) | How fully private EKS clusters are bootstrapped via ECS | +| [FIPS-Only EKS Compute](design/fips-eks-compute.md) | FIPS NodeClass/NodePool strategy for FedRAMP workload nodes | +| [GitOps Cluster Configuration](design/gitops-cluster-configuration.md) | ApplicationSet pattern, progressive deployment, config modes | +| [HyperFleet Architecture](design/hyperfleet-architecture.md) | Operator + kube-applier architecture, component replacement map | +| [Infrastructure Logging](design/infrastructure-logging.md) | AWS CloudWatch log groups, KMS encryption, Grafana access | +| [Logging Platform](design/logging-platform.md) | Application-level log collection (Vector + Loki) | +| [MC Metrics Remote Write](design/mc-metrics-remote-write.md) | MC-to-RC metrics forwarding via RHOBS API Gateway | +| [Monitoring Platform](design/monitoring-platform.md) | Metrics pipeline (Prometheus + Thanos) | +| [Pipeline-Based Lifecycle](design/pipeline-based-lifecycle.md) | CodePipeline hierarchy for cluster provisioning | +| [Rate Limiting](design/rate-limiting-architecture.md) | Per-account rate limiting for Platform API | +| [Regional Account Minting](design/regional-account-minting.md) | AWS account structure and minting pipelines | +| [Regional Control Plane Architecture](design/regional-control-plane-architecture.md) | Operator + PostgreSQL control plane (hyperfleet-operator, hyperfleet-db) | +| [Regional OIDC Ownership](design/regional-oidc-ownership.md) | Shared OIDC bucket per region, cross-account MC writes | +| [Spec-to-PR Agent](design/spec-to-pr-agent.md) | AI agent workflow for spec-driven implementation | +| [SRE UI Access](design/sre-ui-access.md) | ALB + OIDC access to SRE UIs replacing SSM port-forward | +| [Terraform Resource Adoption](design/terraform-resource-adoption.md) | Idempotent import of auto-created AWS resources into Terraform | +| [Testing Strategy](design/testing-strategy.md) | Ephemeral and long-lived test environments | +| [Thanos Metrics Infrastructure](design/thanos-metrics-infrastructure.md) | Thanos S3 storage, operator, and Pod Identity setup | +| [ZOA Architecture](design/zoa-architecture.md) | Zero Operator Access — system components, flows, infrastructure | +| [ZOA Trusted Actions](design/zoa-trusted-actions.md) | TA template format, API design, CLI, dispatch flow | +| [ZOA Security Model](design/zoa-security-model.md) | SA isolation, RBAC, audit trail, threat model, FIPS | ### How-To Guides diff --git a/docs/design/regional-control-plane-architecture.md b/docs/design/regional-control-plane-architecture.md new file mode 100644 index 000000000..3d7f0ac11 --- /dev/null +++ b/docs/design/regional-control-plane-architecture.md @@ -0,0 +1,112 @@ +# Regional Control Plane Architecture + +**Last Updated Date**: 2026-07-31 + +## Summary + +The regional control plane manages ROSA HCP cluster lifecycles using Kubernetes-style reconciliation loops (controller-runtime) backed by PostgreSQL instead of etcd. This gives the team the reconciliation model it has deep expertise in while gaining the PITR, scalability, and queryability of a real database. The implementation consists of `hyperfleet-operator` (the controllers) and `hyperfleet-db` (a PostgreSQL-backed controller-runtime library). + +## Context + +- **Problem Statement**: The platform needs cluster lifecycle management that scales beyond etcd's ~8 GB hard ceiling and supports querying fleet state (e.g. "all clusters in degraded state"). The previous architecture (CLM) depended on another team's framework for lifecycle primitives, which constrained development velocity. +- **Constraints**: + - Team must own the full lifecycle stack to move at its own pace + - Must integrate with the existing Regional Cluster (EKS) and Management Cluster topology + - Must support the platform-api as a stateless REST frontend +- **Assumptions**: + - PostgreSQL (via RDS/Aurora) is available in all target AWS regions + - controller-runtime's interfaces (`Manager`, `Client`, `Cache`) are stable and sufficient for the reconciliation model + +## Design + +### Components + +```mermaid +graph TD + Customer["Customer"] -->|SigV4| APIGW["API Gateway"] + APIGW --> PlatformAPI["platform-api\n(stateless REST)"] + PlatformAPI -->|hyperfleet-db client| PG["PostgreSQL\n(RDS/Aurora)"] + + Operator["hyperfleet-operator\n(controller-runtime)"] -->|reconcile loop| PG + Operator -->|writes desires| DynamoDB["DynamoDB\n(→ kube-applier → MCs)"] + Compactor["compactor"] -->|tombstone GC| PG + + DynamoDB ~~~ Compactor + + style PG fill:#f0f0f0,stroke:#333 +``` + +**hyperfleet-operator** is a controller-runtime operator running on the Regional Cluster. It reconciles custom resources that model the cluster lifecycle. + +The operator communicates with Management Clusters via DynamoDB desire documents. A desire is a declarative spec for a Kubernetes resource that should exist on an MC. The operator writes desires to DynamoDB; **kube-applier**, running on each MC, watches for desires and applies them to the local Kubernetes API server. Status flows back the same way — kube-applier writes observed state to DynamoDB status tables, and the operator reads it to update its own resource status. + +**hyperfleet-db** is a Go library that implements controller-runtime's `Manager`, `Client`, and `Cache` interfaces against PostgreSQL. It stores all Kubernetes resources in a single `kubernetes_resources` table. The operator and platform-api both use it: + +- **Operator**: uses the full `Manager` (client + cache + watch) for reconciliation +- **platform-api**: uses `Client` directly (no cache needed for stateless request/response) + +This means the operator's CRDs are not stored in etcd — PostgreSQL is the sole state store. + +**compactor** is a separate process that periodically deletes soft-deleted tombstones from PostgreSQL. It runs alongside the operator and advances a compaction horizon to prevent watchers from seeing gaps in the event stream. + +### Why PostgreSQL + +- **Scales beyond etcd**: No 8 GB ceiling. +- **Point-in-time recovery**: RDS/Aurora PITR provides disaster recovery without custom backup tooling. +- **Fleet querying**: SQL queries over cluster state (e.g. degraded clusters, clusters by region, placement utilization) without building a separate reporting layer. +- **Multi-AZ**: RDS synchronous standby provides zero acknowledged-write loss on failover. + +For detailed internals (schema, invariants, watch mechanism, race catalog), see [hyperfleet-db DESIGN.md](../../../rosa-hyperfleet-api/hyperfleet-db/docs/DESIGN.md). + +## Alternatives Considered + +1. **CLM (Cluster Lifecycle Manager)**: A REST-based lifecycle service with adapters, sentinels, and CloudEvents. The CLM pattern used a stateless API server with GORM, a polling sentinel for change detection, CloudEvents for notification, and adapters for reconciliation. This was rejected because: + - **Velocity**: The framework was owned by another team, creating a dependency that constrained the platform team's development pace. + - **Component count**: 5+ components in the reconcile loop (API server, sentinel, message broker, adapters, status reporters) vs. 2 (operator + PostgreSQL). + - **Operational overhead**: More services to deploy, monitor, and debug during incidents. + + For a detailed comparison of reliability and performance characteristics, see [Architecture Comparison](../../../rosa-hyperfleet-api/hyperfleet-db/docs/ARCHITECTURE_COMPARISON.md). + +2. **Standard controller-runtime with etcd**: Using controller-runtime with its default etcd backend. This was rejected because: + - **8 GB hard ceiling**: etcd's storage limit would require sharding to scale beyond a few thousand clusters, adding significant operational complexity. + - **No fleet querying**: etcd supports key-prefix listing but not the rich queries needed for fleet management. + - **No PITR**: etcd snapshots are coarse-grained; RDS PITR provides second-granularity recovery. + +## Design Rationale + +- **Justification**: The team has deep expertise in Kubernetes-style reconciliation (watch, reconcile, requeue). By implementing controller-runtime's storage interfaces against PostgreSQL, the operator retains that programming model while gaining a real database's PITR, scalability beyond etcd's 8 GB ceiling, and SQL queryability over fleet state. The team owns the full stack. +- **Evidence**: Measured write latency of p50=6.3ms / p99=29ms and throughput of 6,132 writes/s with realistic 15-20KB payloads (Aurora I/O Optimized, db.r6g.8xlarge). See [Architecture Comparison](../../../rosa-hyperfleet-api/hyperfleet-db/docs/ARCHITECTURE_COMPARISON.md) for full benchmarks. +- **Comparison**: CLM's advantages (standard REST API, operational familiarity, existing ecosystem) are real but secondary to the team velocity and component reduction goals that motivated the change. + +## Consequences + +### Positive + +- Team owns the full cluster lifecycle stack with no external framework dependencies +- Fewer components to deploy, monitor, and debug (2 vs. 5+) +- Point-in-time recovery via RDS without custom backup infrastructure +- SQL-based fleet querying without a separate reporting layer + +### Negative + +- No direct `kubectl` access to cluster state — state lives in PostgreSQL, not the Kubernetes API server +- hyperfleet-db is a custom library that must be maintained alongside upstream controller-runtime changes + +## Cross-Cutting Concerns + +### Reliability + +- **Resiliency**: RDS Multi-AZ synchronous standby provides zero acknowledged-write loss on failover. A continuous production verifier checks correctness invariants on live data. See [hyperfleet-db DESIGN.md](../../../rosa-hyperfleet-api/hyperfleet-db/docs/DESIGN.md) for the invariant catalog. +- **Observability**: The operator exposes standard controller-runtime metrics. The compactor logs tombstone deletion counts and compaction horizon advances. + +### Performance + +- This is a low write-rate system — cluster lifecycle events (create, update, delete) and underlying status updates are infrequent relative to the throughput ceiling +- Write latency: p50=6.3ms, p99=29ms with realistic 15-20KB payloads (Aurora I/O Optimized, db.r6g.8xlarge) +- Throughput ceiling: 6,132 writes/s with realistic payloads — orders of magnitude above expected load +- No-op suppression: content-equal writes consume no sequence, version bump, or watch event + +### Cost + +- RDS/Aurora instance per region (sized by cluster count — db.r6g.large at 5,000 clusters, db.r6g.2xlarge at 50,000) +- Eliminates CLM API server, sentinel, and adapter compute costs