392 lines
18 KiB
Markdown
392 lines
18 KiB
Markdown
# Nereus
|
|
|
|
Nereus is an ocean buoy telemetry API that exists so there is something real to
|
|
deploy. It stores buoys and their readings, exposes Prometheus metrics and
|
|
OpenTelemetry traces, and can be told to start failing on demand. The
|
|
application is deliberately small and boring. The project is everything wrapped
|
|
around it: custom images, a two-node k3s cluster, CI/CD on a self-hosted runner,
|
|
blue-green deploys that abort themselves when the new version is bad, and logs,
|
|
metrics and traces that actually join up.
|
|
|
|
This is my final DevOps project for Tokio School. The fictional company in the
|
|
brief is TechWave Solutions, which fits the water theme, so I ran with it.
|
|
|
|
**If you only run one thing, run the k3d lab.** It puts a real multi-node k3s
|
|
inside Docker and proves the automated rollback end to end on any OS with Docker
|
|
and 8 GB of RAM:
|
|
|
|
```bash
|
|
scripts/k3d/lab.sh up
|
|
scripts/k3d/lab.sh prove
|
|
```
|
|
|
|
## What the API does
|
|
|
|
Go 1.26, `chi` for routing, `pgx/v5` against PostgreSQL 17, no ORM. Migrations
|
|
are plain SQL files applied on startup and written to be idempotent, so a pod
|
|
restart or a second replica does not break anything.
|
|
|
|
```
|
|
GET /healthz liveness, 200 whenever the process is up
|
|
GET /readyz readiness, 200 only once the database is reachable
|
|
GET /metrics Prometheus exposition
|
|
|
|
GET /api/v1/buoys list
|
|
POST /api/v1/buoys create
|
|
GET /api/v1/buoys/{id} fetch one
|
|
DELETE /api/v1/buoys/{id} delete
|
|
|
|
GET /api/v1/readings?buoy_id=&from=&to= list, paginated
|
|
POST /api/v1/readings create
|
|
GET /api/v1/readings/aggregate?window=1h deliberately expensive GROUP BY
|
|
```
|
|
|
|
The split between `/healthz` and `/readyz` is the part that matters for
|
|
Kubernetes. The service starts and serves `/healthz` even with the database
|
|
down, so a PostgreSQL outage makes pods unready instead of killing them in a
|
|
restart loop. `/api/v1/readings/aggregate` runs a real grouped query over a time
|
|
window and is not optimised on purpose, because the dashboards need something
|
|
that produces visible latency.
|
|
|
|
`CHAOS_ERROR_RATE` is a float between 0 and 1. Above zero, that fraction of
|
|
`/api/v1/*` requests return HTTP 500 with a JSON body. Health endpoints are never
|
|
affected. This is the entire mechanism for simulating a bad release, and it is
|
|
controlled by an environment variable rather than an admin endpoint so that
|
|
turning it on is a deployment, not a runtime toggle.
|
|
|
|
`apps/loadgen/` is a small Go binary that hits the API continuously with a
|
|
weighted mix of reads, writes and aggregates. It is not decoration. The rollback
|
|
analysis computes an error rate from Prometheus, and with no traffic there are
|
|
no samples and the analysis has nothing to decide on.
|
|
|
|
## Running it
|
|
|
|
There are three levels, and they exist because nobody grading this is going to
|
|
install KVM and download a Fedora image.
|
|
|
|
### Level 1: Docker Compose, about three minutes
|
|
|
|
```bash
|
|
docker compose up --build
|
|
```
|
|
|
|
API, PostgreSQL and an OTel Collector, with the API on `127.0.0.1:18080` and the
|
|
load generator already producing traffic. Add the observability overlay for
|
|
Grafana on `127.0.0.1:13000` and Prometheus on `127.0.0.1:19090`:
|
|
|
|
```bash
|
|
docker compose -f compose.yaml -f observability/compose.yaml up -d
|
|
```
|
|
|
|
### Level 2: k3d, real Kubernetes, any OS with Docker
|
|
|
|
```bash
|
|
scripts/k3d/lab.sh up # one server, two agents, Argo Rollouts, kube-prometheus-stack
|
|
scripts/k3d/lab.sh prove # promote a healthy version, then abort a bad one
|
|
scripts/k3d/lab.sh destroy
|
|
```
|
|
|
|
`prove` is the interesting one. It drives a blue-green Rollout through a real
|
|
AnalysisRun that queries a real Prometheus over the network, first with a query
|
|
that returns a healthy value and then with one that returns a failing value, and
|
|
it fails loudly if the bad revision ever reaches the active Service.
|
|
|
|
Be clear about what is faked here: the harness in
|
|
`scripts/k3d/analysis-harness/` uses a throwaway nginx Rollout and a hardcoded
|
|
metric value. Argo Rollouts, the AnalysisRun, the Prometheus query and the abort
|
|
are all real. Only the number is synthetic. The real gate lives in
|
|
`deploy/rollouts/analysistemplate.yaml` and ran against the real API on the
|
|
Fedora cluster.
|
|
|
|
### Level 3: two Fedora 44 machines
|
|
|
|
This is my actual cluster. `scripts/provision/` holds an Ansible playbook that
|
|
turns two already-installed Fedora 44 hosts into a k3s server and agent:
|
|
|
|
```bash
|
|
cp scripts/provision/inventory.example.yml scripts/provision/inventory.yml
|
|
# edit addresses, CIDRs and SSH user
|
|
K3S_TOKEN="$(openssl rand -hex 32)" scripts/provision/bootstrap.sh scripts/provision/inventory.yml
|
|
```
|
|
|
|
The playbook is idempotent and reruns clean. The token is passed through the
|
|
process environment and never written to the inventory or the repository.
|
|
|
|
`scripts/provision/lab.sh` builds the same two hosts as local QEMU VMs from the
|
|
Fedora 44 cloud image, with 2 vCPUs, 3 GiB of RAM and a 30 GiB thin disk each.
|
|
Those numbers are the lowest I actually tested, not a recommendation. VM state
|
|
and its SSH key live under `~/.local/state/nereus-lab`, outside the repo.
|
|
|
|
## How a bad deploy gets killed
|
|
|
|
This is the part I would defend first, so it is worth spelling out.
|
|
|
|
The API is an Argo Rollouts `Rollout`, not a Deployment. New pods come up
|
|
alongside the old ones and take no production traffic. `nereus-api-active` points
|
|
at the stable version and `nereus-api-preview` at the candidate. A second load
|
|
generator drives the preview Service, because an analysis that measures nothing
|
|
would treat "no samples" as success.
|
|
|
|
Before promotion, a `prePromotionAnalysis` runs this query against Prometheus:
|
|
|
|
```promql
|
|
(
|
|
sum(rate(nereus_http_requests_total{service="nereus-api-preview",status=~"5.."}[1m]))
|
|
or vector(0)
|
|
)
|
|
/
|
|
sum(rate(nereus_http_requests_total{service="nereus-api-preview"}[1m]))
|
|
```
|
|
|
|
Five samples, 20 seconds apart, after a 30 second delay to let the new pods take
|
|
traffic. The success condition is `len(result) == 0 || result[0] < 0.05`, and
|
|
`failureLimit: 0` means a single bad sample aborts. The `or vector(0)` keeps the
|
|
division defined when there are no errors at all, which is otherwise an empty
|
|
result rather than a zero.
|
|
|
|
When it aborts, the active Service selector never moves. `scaleDownDelaySeconds:
|
|
30` keeps the old ReplicaSet warm so the fallback is instant rather than a cold
|
|
start. From `docs/evidence.md`, the two AnalysisRuns from the mechanism proof:
|
|
|
|
```text
|
|
probe-746bbb94df-2-pre Successful vector(0.0)
|
|
probe-556d5b659b-3-pre Failed vector(1.0)
|
|
```
|
|
|
|
To trigger it for real, deploy a version with `CHAOS_ERROR_RATE` above 0.05 and
|
|
watch it refuse to promote.
|
|
|
|
## Architecture
|
|
|
|
```
|
|
Workstation (CachyOS)
|
|
└── libvirt/KVM
|
|
├── nereus-node1 Fedora 44, k3s server
|
|
└── nereus-node2 Fedora 44, k3s agent
|
|
|
|
Mini PC (always on)
|
|
├── Forgejo at git.fiwlabs.dev, behind Traefik with automatic TLS
|
|
├── Forgejo Runner, docker mode, same LAN as the cluster
|
|
├── Container registry, part of Forgejo
|
|
└── Static dashboard at nereus.fiwlabs.dev, nginx behind Traefik
|
|
|
|
Mirror to GitHub, which is the link I hand in.
|
|
```
|
|
|
|
Inside the cluster:
|
|
|
|
```
|
|
Traefik (ships with k3s)
|
|
└── Ingress → nereus-api-active
|
|
├── stable ReplicaSet (real traffic)
|
|
└── preview ReplicaSet (candidate, no traffic)
|
|
|
|
PostgreSQL StatefulSet on a local-path PVC
|
|
loadgen one Deployment against active, one against preview
|
|
|
|
observability/
|
|
kube-prometheus-stack Prometheus, Grafana, node-exporter, kube-state-metrics
|
|
OTel Collector Deployment, receives OTLP traces, exports to Tempo
|
|
OTel Collector DaemonSet, tails pod logs, exports to Loki
|
|
Tempo traces
|
|
Loki logs
|
|
```
|
|
|
|
The web dashboard runs on the Mini PC, not in the cluster, and that is
|
|
deliberate. The cluster is off most of the time. If the page lived inside it,
|
|
the domain would be dead 95% of the time and there would be nothing to show.
|
|
Traefik on the Mini PC serves the static files at `/` and proxies `/api/v1/*`,
|
|
`/healthz` and `/readyz` through to the cluster, so everything is same-origin
|
|
with no CORS and no hardcoded hostnames. With the cluster off, those paths return
|
|
502 and the page renders the last known values greyed out and keeps polling. An
|
|
unreachable API is a normal state for that page, not an error.
|
|
|
|
## Why I built it this way
|
|
|
|
**distroless, not alpine.** The final image is
|
|
`gcr.io/distroless/static:nonroot` with a statically linked binary copied in and
|
|
nothing else. No shell, no package manager, no busybox. If someone gets code
|
|
execution there is nothing to pivot with. It also happens to be tiny: the API
|
|
image is under 8 MB, with the exact figure recorded in `docs/evidence.md`.
|
|
|
|
**k3s, not full Kubernetes.** Same API, one binary, Traefik and a local-path
|
|
provisioner included, and it runs on two 3 GiB VMs. Nothing in this project uses
|
|
an API that k3s does not have.
|
|
|
|
**On-premise, not EKS or AKS.** The brief asks for a managed cloud cluster. I
|
|
did not do that, and the reason is cost and control rather than difficulty: I
|
|
have hardware sitting here and no interest in paying for a load balancer to prove
|
|
a point about Terraform. The manifests are plain Kubernetes and would apply to a
|
|
managed cluster unchanged. This is the deviation from the brief I expect to be
|
|
asked about, so it is stated here rather than buried.
|
|
|
|
**Forgejo Actions, not GitHub Actions.** The brief allows the tool of your
|
|
choice. The runner lives on the same LAN as the cluster, which means the deploy
|
|
job can reach the Kubernetes API without exposing it to the internet or
|
|
punching a hole through a firewall for a hosted runner. The registry is part of
|
|
the same Forgejo instance, so images never leave the network either. The GitHub
|
|
mirror still runs its own workflow so the repo a reviewer opens shows green
|
|
checks, and that workflow consumes no secrets by design.
|
|
|
|
**The cluster is ephemeral, on purpose.** I bring it up, record, and shut it
|
|
down. That is not a gap in the project, it is the evidence that the
|
|
infrastructure as code works. If the whole thing rebuilds from the repository
|
|
with a command, then the repository really is the source of truth.
|
|
|
|
**Everything runs non-root with a read-only root filesystem**, including the
|
|
pieces where that is inconvenient. The log collector was the awkward one. k3s
|
|
writes `/var/log/pods` as `0750 root:root` and each container log as `0640
|
|
root:root`, so a collector running as UID 10001 cannot even traverse the
|
|
directory. Adding the `DAC_READ_SEARCH` capability looks like the fix and does
|
|
nothing, because Kubernetes sets no ambient capabilities and the capability is
|
|
cleared on exec for a non-root user. The container ended up with
|
|
`CapEff: 0000000000000000`, matched no files, and reported no error at all. The
|
|
fix is `supplementalGroups: [0]`, since group root already has read on those
|
|
files. Non-root, no capabilities, read-only mount.
|
|
|
|
## Secrets and security
|
|
|
|
No secret is committed, and none is read by any tooling that logs. Gitleaks
|
|
scans both the working tree and the committed history on every push and a hit
|
|
fails the build. Trivy scans the release configuration and the built images for
|
|
HIGH and CRITICAL findings.
|
|
|
|
The `nereus-db` Secret holds the PostgreSQL password and the DSN, and it is
|
|
created out of band rather than living in `base/kustomization.yaml`. See
|
|
`deploy/secrets/README.md` for how it is created in each environment.
|
|
|
|
The CI pipeline never uses a cluster-admin kubeconfig. It authenticates as a
|
|
`deployer` ServiceAccount with RBAC scoped to the `nereus` namespace, supplied
|
|
as an encrypted `KUBECONFIG_B64` repository secret, written to the workspace with
|
|
mode `0600` and deleted in an `always()` step. The registry pull secret is
|
|
created through a pipe so the generated manifest is never logged or written to
|
|
disk.
|
|
|
|
Containers drop all capabilities, disallow privilege escalation, run with a
|
|
read-only root filesystem and a `RuntimeDefault` seccomp profile, and mount
|
|
`tmpfs` where a writable path is genuinely needed. SELinux stays enforcing on
|
|
both Fedora nodes.
|
|
|
|
## Observability
|
|
|
|
Metrics, logs and traces come from the same request and can be walked between.
|
|
|
|
The API exports five metrics with fixed names, since the dashboards and the
|
|
rollback query depend on them literally:
|
|
|
|
```
|
|
nereus_http_requests_total{method, path, status, version} counter
|
|
nereus_http_request_duration_seconds{method, path, version} histogram
|
|
nereus_db_query_duration_seconds{operation} histogram
|
|
nereus_readings_ingested_total counter
|
|
nereus_buoys_active gauge
|
|
```
|
|
|
|
`path` is always the chi route template, `/api/v1/buoys/{id}` rather than the
|
|
resolved path, because resolved paths give unbounded cardinality and will
|
|
eventually take Prometheus down.
|
|
|
|
Traces go out over OTLP/gRPC to the collector Deployment and on to Tempo. Every
|
|
handler is a span and every database query is a child span. If
|
|
`OTEL_EXPORTER_OTLP_ENDPOINT` is unset, tracing is a no-op, so a missing
|
|
collector can never stop the service from starting.
|
|
|
|
Logs are structured JSON on stdout via `log/slog`. A DaemonSet collector tails
|
|
`/var/log/pods` on each node and ships to Loki. Every request logs one line
|
|
carrying its `trace_id`, and Grafana's Loki data source has a derived field that
|
|
turns that value into a link into Tempo, with the reverse link configured on the
|
|
Tempo side. A verified example, log line and the trace it resolves to, is in
|
|
`docs/evidence.md`.
|
|
|
|
Alert rules for error rate, latency, readiness failures and scrape failure are in
|
|
`observability/alerts/nereus.yaml`. cAdvisor and node-exporter metrics arrive
|
|
with kube-prometheus-stack rather than being installed separately.
|
|
|
|
## CI/CD
|
|
|
|
`.forgejo/workflows/pipeline.yml` runs on pushes and pull requests to `main`.
|
|
|
|
Verification runs the Go tests with the race detector, `go vet`, a pinned
|
|
`golangci-lint`, gitleaks over history and tree, a kustomize build of every
|
|
overlay, and a Trivy config scan. Everything runs as a plain container image
|
|
rather than a marketplace action, because Forgejo resolves actions from
|
|
`code.forgejo.org` and most third-party ones simply are not mirrored there.
|
|
|
|
A push to `main` then builds the three images, tags them with the immutable
|
|
commit SHA, scans them, pushes them to the Forgejo registry, and deploys the
|
|
production overlay. The deploy job waits on the Rollout, which means a failed
|
|
pre-promotion analysis fails the pipeline.
|
|
|
|
## Where each requirement from the brief lives
|
|
|
|
| Requirement | Where |
|
|
|---|---|
|
|
| Custom, optimised Docker images | `build/*.Dockerfile`, multi-stage into distroless, API under 8 MB |
|
|
| Docker Compose | `compose.yaml`, plus `observability/compose.yaml` and `build/compose.mini.yaml` |
|
|
| Kubernetes orchestration | k3s, two Fedora 44 nodes; k3d for a portable equivalent |
|
|
| Deployments, Services, Ingress, ConfigMaps, Secrets | `deploy/base/`, composed with kustomize overlays |
|
|
| Infrastructure as code | `scripts/provision/`, Ansible roles plus an idempotent bootstrap |
|
|
| CI/CD pipeline | `.forgejo/workflows/pipeline.yml`, self-hosted runner |
|
|
| Container registry | Forgejo registry at `git.fiwlabs.dev`, TLS from Traefik |
|
|
| Secure credential handling | Gitleaks, Trivy, namespace-scoped `deployer` kubeconfig, out-of-band `nereus-db` |
|
|
| Blue-green deployment | Argo Rollouts, `deploy/rollouts/rollout.yaml` |
|
|
| Automated rollback | `deploy/rollouts/analysistemplate.yaml`, Prometheus-driven abort |
|
|
| OpenTelemetry Collector | `deploy/observability/otel-collector.yaml` for traces, `otel-log-collector.yaml` for logs |
|
|
| Prometheus | kube-prometheus-stack, scraped through a ServiceMonitor |
|
|
| Grafana dashboards | `observability/dashboards/`, data sources in `deploy/observability/grafana-datasources.yaml` |
|
|
| Loki | `deploy/observability/loki.yaml` |
|
|
| cAdvisor and node-exporter | included with kube-prometheus-stack |
|
|
| Alerting | `observability/alerts/nereus.yaml` |
|
|
|
|
## What is not done
|
|
|
|
Stated plainly so it does not have to be discovered.
|
|
|
|
`terraform/infra/` creates the two Fedora 44 libvirt machines and emits the
|
|
inventory consumed by Ansible. `terraform/platform/` installs Argo Rollouts,
|
|
Sealed Secrets and kube-prometheus-stack. The separation is deliberate:
|
|
Terraform owns machines and cluster-wide controllers, Ansible owns the hosts,
|
|
and Kustomize owns the application.
|
|
|
|
Sealed Secrets and Discord alert routing are configured. The database
|
|
`SealedSecret` is committed for fresh deployments; the Discord resource still
|
|
awaits its owner-provided webhook. `deploy/secrets/README.md` keeps plaintext
|
|
values in local pipelines and commits only encrypted resources. The controller
|
|
key must be backed up outside this repository before it becomes the recovery
|
|
path.
|
|
|
|
The remaining acceptance gap is environmental: the reproducible fresh-VM path
|
|
has not been rerun on a second workstation with no dependencies installed.
|
|
|
|
## Fedora traps worth knowing
|
|
|
|
Three things broke the cluster in ways that took longer to find than to fix, all
|
|
of them handled by the playbook now.
|
|
|
|
| Symptom | Cause | Fix |
|
|
|---|---|---|
|
|
| Pods on different nodes cannot reach each other, DNS fails strangely | firewalld blocks VXLAN | open `8472/udp` between nodes |
|
|
| kubelet fails with permission errors | SELinux enforcing without the policy | `dnf install k3s-selinux` |
|
|
| k3s complains at startup | zram swap enabled | `systemctl disable --now zram-generator` |
|
|
|
|
## Repo layout
|
|
|
|
```
|
|
apps/api/ Go service
|
|
apps/loadgen/ traffic generator
|
|
apps/web/ static dashboard, runs on the Mini PC
|
|
build/ Dockerfiles and the Mini PC compose file
|
|
deploy/base/ kustomize base
|
|
deploy/overlays/ dev and prod
|
|
deploy/rollouts/ Rollout and AnalysisTemplate
|
|
deploy/observability/ Loki, Tempo, collectors, Grafana data sources
|
|
observability/ dashboards, alert rules, collector configs, local compose
|
|
scripts/k3d/ portable rollback lab
|
|
scripts/provision/ Ansible roles, bootstrap, QEMU lab
|
|
terraform/ libvirt machines, Ansible inventory and cluster-wide Helm releases
|
|
docs/ roadmap, evidence, CI/CD and Mini PC notes
|
|
```
|
|
|
|
`AGENTS.md` holds the rules the AI agents working in this repo had to follow.
|
|
`PLAN.md` is my own working plan, in Spanish, and it is a worklog rather than
|
|
documentation.
|