# Jeff Slavin | Lead Technical Writer - full site content Generated at build time from the Markdown sources of https://jslavin-docs.github.io/ - see llms.txt for the curated index. (c) Jeff Slavin. Portfolio content is all rights reserved. ======================================================================== Page: https://jslavin-docs.github.io/ Description: Technical writing portfolio for Jeff Slavin, focused on GitOps, DevSecOps, APIs, cloud, edge, IoT, and Docs as Code. ======================================================================== # Jeff Slavin TECHNICAL WRITER · DOCUMENTATION LEAD ## GitOps, DevSecOps, API, cloud, edge, and IoT documentation that helps technical users move faster. Clear user and developer documentation, with a focus on Docs as Code, information architecture, OpenAPI, and content structured for self-service and AI retrieval. [Portfolio](https://jslavin-docs.github.io/portfolio/) [Resume](https://jslavin-docs.github.io/resume/) [Email](mailto:jslavin.docs@gmail.com) [LinkedIn](https://www.linkedin.com/in/jeff-slavin) [GitHub](https://github.com/jslavin-docs) 40% retail client hardware cost reduction supported by secure AWS edge platform documentation 60% search-success improvement after Confluence-to-self-service portal redesign 30% migration support-ticket reduction for GCP IoT Core-to-ClearBlade guidance Weeks → days partner integration acceleration through the OpenAPI documentation behind SDK generation ## Specialties ### Developer and API docs REST API references, OpenAPI specifications, SDK docs, webhook guides, integration guides, and developer onboarding. ### Cloud and edge documentation AWS, GCP, Kubernetes, EKS, Terraform, Docker, Argo CD, Prometheus, Grafana, and Linux. ### Docs as Code and automation Git, GitHub Actions, Bitbucket, Jenkins, MkDocs, GitBook, Confluence, Scroll Sites, Mermaid, Postman, Swagger, Python, Bash, YAML, JSON, and Markdown. ### Documentation strategy Information architecture, content strategy, style guides, taxonomy, metadata, AI-ready content structuring, knowledge management, and cross-functional documentation leadership. ### Security and operations Cybersecurity, IAM, zero trust, DevSecOps, GitOps, SRE, CI/CD, IaC, runbooks, incident response, and operational standards. ### Technical domains Cloud computing, edge computing, IoT, SaaS, API integration, GenAI, blockchain, Web3, and cloud networking. ## Featured work EKS · Argo CD · GitOps · DevSecOps ### [GitOps and DevSecOps documentation](https://jslavin-docs.github.io/portfolio/#gitops-devsecops-documentation) EKS and Argo CD operator guidance showing deployment controls, secret-handling rules, verification checks, and rollback decisions. OpenAPI · REST · Webhooks ### [API and developer documentation](https://jslavin-docs.github.io/portfolio/#api-developer-documentation) Interactive API references, query schema docs, webhook payload guides, and SDK-generation workflows. Migration · Quick Start ### [End-user and onboarding documentation](https://jslavin-docs.github.io/portfolio/#end-user-onboarding-documentation) Step-by-step migration and onboarding content that helps technical users self-serve. Edge · Architecture · Tutorial ### [Edge computing documentation](https://jslavin-docs.github.io/portfolio/#edge-computing-documentation) Platform overviews, architecture guidance, and deployment tutorials for edge environments. ## Certifications and education | Credential | Details | |---|---| | PMP | Project Management Professional | | ITIL Expert | IT Infrastructure Library | | MBA | Florida Atlantic University | | BBA, Management Information Systems | Florida Atlantic University | ======================================================================== Page: https://jslavin-docs.github.io/portfolio/ Description: Selected technical writing samples by Jeff Slavin covering GitOps, DevSecOps, APIs, developer onboarding, IoT platform migration, and edge computing. ======================================================================== # Portfolio A focused set of technical writing samples covering GitOps, DevSecOps, APIs, developer onboarding, IoT platform migration, and edge computing. Each sample highlights the technical scope, intended audience, and business context. ## GitOps & DevSecOps Documentation Amazon EKS · Argo CD · GitOps · DevSecOps ### [NovaDeploy GitOps Administration Guide – Portfolio Cut](https://jslavin-docs.github.io/writing-samples/novadeploy-gitops-admin-guide-portfolio-cut/) A 5–7-minute cut of a fictional operator runbook: architecture, deployment steps, stop checkpoints, verification, and rollback. Covers source-of-truth rules for GitOps and Terraform, as well as keeping plaintext secrets out of Git, PRs, logs, and tickets. Includes a worked secret-refresh walkthrough and the CI check that enforces the Reloader restart guardrail. Amazon EKS · IRSA · KMS · ESO · Reloader ### [NovaDeploy GitOps Administration Guide – Full Runbook](https://jslavin-docs.github.io/writing-samples/novadeploy-gitops-admin-guide-full-version/) The complete fictional operator runbook: SecretStore and ESO setup, as well as Terraform patterns for IAM, KMS, and Secrets Manager. Includes a rotation readiness gate, CI guardrails, verification commands, break-glass rollback, and an evidence checklist. Start with the portfolio cut for a shorter read. ## API & Developer Documentation OpenAPI 3.0 · SDK generation ### [OpenAPI 3.0 Interactive API Reference – ClearBlade IoT Enterprise](https://docs.clearblade.com/iotenterprise/apis) Documented the OpenAPI 3.0 specification behind the interactive API reference and SDK generation for ClearBlade IoT Enterprise, reducing partner integration time from weeks to days. REST API · Query schema ### [REST API Query Schema Technical Reference](https://docs.clearblade.com/iotenterprise/rest-api-query-schema) Detailed ClearBlade's REST API query schema, including SQL-to-JSON translation, filter-condition logic, and TypeScript interfaces for developers querying collections through the platform. Webhooks · Payloads · cURL ### [Webhooks Configuration & Payload Guide](https://docs.clearblade.com/iotenterprise/webhooks) Created webhook documentation covering authentication methods, cURL examples, JSON payload structures, path wildcards, and REST API management for programmatic webhook configuration. ## End-User & Onboarding Documentation Quick start · GCP IoT Core migration ### Quick Start Guide – ClearBlade IoT Core Authored a step-by-step quick start with code samples and telemetry testing as part of the migration guides that moved 250+ enterprises off GCP IoT Core onto ClearBlade, cutting support tickets by 30%. [Read sample →](https://docs.clearblade.com/iotcore/quick-start) [Read case study →](https://jslavin-docs.github.io/case-studies/gcp-iot-core-migration/) ## Edge Computing Documentation Architecture · Platform overview ### [Edge Computing Platform Overview & Architecture](https://docs.clearblade.com/iotenterprise/edge) Developed ClearBlade Edge documentation for installation, configuration, and synchronization, helping teams independently deploy and manage edge environments. Hands-on tutorial · Deployment ### [Hands-On Edge Deployment Tutorial](https://docs.clearblade.com/iotenterprise/edge-tutorial) Created a hands-on tutorial for deploying, installing, starting, and upgrading ClearBlade IoT Enterprise edge instances, from initial setup through operation. ## Public Contributions GitHub · Contribution activity ### [ClearBlade Contribution Activity](https://github.com/jslavin-clearblade?tab=overview&from=2023-01-01&to=2023-12-31) Maintained documentation across 23 ClearBlade repositories in 2023, with 135 public contributions covering the native code library reference, SDK READMEs, and developer kit documentation. Git · Docs as Code ### [Native Code Library Commits – ClearBlade](https://github.com/ClearBlade/native-libraries/commits/master/?author=jslavin-clearblade) Authored 28 commits to ClearBlade's public native-libraries documentation repository, each with a reviewable diff of the changes. ## Built for AI Retrieval The site has machine-readable files so AI assistants can give accurate answers about this work: - **[skill.md](https://jslavin-docs.github.io/skill.md)** — an instruction file that keeps AI answers grounded in the published content. - **[llms.txt](https://jslavin-docs.github.io/llms.txt)** — index of the site's pages. - **[llms-full.txt](https://jslavin-docs.github.io/llms-full.txt)** — the site's complete Markdown content in one file. - Every page is also available as plain Markdown. ======================================================================== Page: https://jslavin-docs.github.io/resume/ Description: Resume for Jeff Slavin, a technical writer and documentation lead with experience in API, cloud, edge, IoT, and DevSecOps documentation, information architecture, Docs as Code, and AI-ready content. ======================================================================== # Jeff Slavin **Lead Technical Writer | API, Cloud & Edge | Docs as Code & GitOps** Miami, FL · [jslavin.docs@gmail.com](mailto:jslavin.docs@gmail.com) · [LinkedIn](https://www.linkedin.com/in/jeff-slavin) · [GitHub](https://github.com/jslavin-docs) [View resume (PDF)](https://jslavin-docs.github.io/assets/Jeff-Slavin-Resume.pdf) --- ## Summary Technical writer and documentation lead with 10+ years across API, cloud, edge, IoT, and DevSecOps. Owns documentation end to end for users and developers, spanning information architecture, Docs as Code, and content structured for self-service and AI retrieval. Cuts support tickets by 30% and takes partner integrations from weeks to days. --- ## Experience ### Artisan Studios **Senior Technical Writer** Nov 2024 – Present · Remote - Author design documentation for a secure AWS edge platform (Kubernetes, Argo CD, GitOps) whose rollout cut retail client hardware costs by 40% - Create process flow and network topology diagrams from working sessions with solutions architects and developers, shortening architecture review cycles by 25% - Codify triage, diagnostics, and escalation for 65 Prometheus alerts in the DevSecOps runbook used by on-call engineers - Write the 160-page user guide for managing clusters, nodes, namespaces, workloads, and applications across distributed edge locations - Document the access control model, secrets management, and certificate handling across user and machine-to-machine authentication ### ClearBlade **Lead Technical Writer** Oct 2022 – Nov 2024 · Remote - Owned three product documentation sets end to end, standardizing topic-type architecture and templates on the developer platforms - Implemented a Git-based Docs as Code workflow with CI/CD publishing, maintaining SDK and library documentation across 23 repositories - Documented the OpenAPI specification behind the interactive API reference and SDK generation, reducing partner integration time from weeks to days - Authored the migration guides that moved 250+ enterprises off GCP IoT Core onto ClearBlade, cutting support tickets by 30% - Rebuilt Confluence content into a self-service portal in Scroll Sites, with a new taxonomy that raised search success rates by 60% - Restructured Confluence topics to be self-contained and consistently labeled, then revised them to improve Rovo's RAG-based AI answers across question types - Advised the technical support organization on end-user documentation practices, shaping how the team delivered guidance to customers ### ODEM **Technical Writer** Mar 2018 – Oct 2022 · Remote - Built the knowledge base for a Web3 credentialing platform that issued 4,800+ blockchain diplomas, covering onboarding, wallet setup, and credential verification - Detailed 14 smart contract use cases with swimlane diagrams, platform screenshots, and a glossary for the 28-page Program Staking and Token Architecture white paper - Wrote the Ethereum-to-Algorand migration guides for ODEM token and credential holders ### MobiWork **Technical Writer** Feb 2017 – Nov 2017 · Delray Beach, FL - Delivered a mobile-first knowledge base for a workforce management SaaS platform, covering field technician workflows, scheduling, and client onboarding ### AT&T / Randstad Technologies **Technical Writer** Jun 2014 – Dec 2016 · Remote - Standardized ITIL-compliant service assurance workflows for network operations, developing the interface agreements and job aids that route cross-team escalations --- ## Skills ### Documentation Strategy Content Strategy · Information Architecture · Taxonomy and Metadata · Style Guide Development · AI-Ready Content Structuring ### Deliverables REST API and SDK Documentation · Knowledge Bases · Quick Starts and Tutorials · Migration Guides · Administrator and User Guides · Runbooks ### Cloud and Infrastructure AWS · GCP · Amazon EKS · Kubernetes · Docker · Terraform · Argo CD · Prometheus · Grafana · Linux ### Documentation Tools MkDocs · GitBook · Confluence · Scroll Sites · Zendesk · Mermaid · Lucidchart ### Development and Automation Git · GitHub Actions · Bitbucket · Jenkins · Jira · Postman · OpenAPI · Swagger · Python · Bash · YAML · JSON · Markdown ### Technical Domains Cloud and Edge Computing · IoT · API Integration · Cybersecurity · IAM · Zero Trust Architecture · Cloud Networking · GenAI · RAG ### Methodologies Docs as Code · GitOps · DevOps · DevSecOps · CI/CD · IaC · SRE · Agile/Scrum · ITIL --- ## Certifications PMP · ITIL Expert --- ## Education Florida Atlantic University MBA · BBA, Management Information Systems ======================================================================== Page: https://jslavin-docs.github.io/case-studies/gcp-iot-core-migration/ Description: How the ClearBlade IoT Core quick start was designed to prevent silent migration failures during Google's shutdown of Cloud IoT Core. ======================================================================== # Case Study: Designing a Quick Start So Migrations Don't Fail Silently *Lead Technical Writer, ClearBlade (Oct 2022 – Nov 2024)* When Google shut down Cloud IoT Core, every connected fleet had to move by a hard deadline. I wrote the quick start that became the front door to the replacement, designed so the whole path is proven on a sample device before anyone trusts a migration reported as complete. ## The problem In August 2022, Google announced it would retire Cloud IoT Core. On August 16, 2023, the service shut down: MQTT and HTTP bridges closed and Google's own docs went offline with it. Every connected fleet had to move, whether they wanted to or not. ClearBlade built ClearBlade IoT Core as a direct replacement. More than 250 Google Cloud customers ultimately migrated. That number is the company's, for the whole program. It's what the docs had to support: production fleets, a fixed external deadline, and operators forced to migrate. ## The constraint The product goal was to minimize customer changes. Customers keep their Google Cloud project, their Pub/Sub topics (Google's message pipeline), and their device credentials. The device-side change is pointing to ClearBlade's MQTT endpoint (MQTT is the messaging protocol used by the devices). That dictates the onboarding order: provisioning runs through Google Cloud Marketplace, a GCP service account authorizes the connection, and Pub/Sub permissions let telemetry flow. The guide has to start in GCP because that's how the product works. I also called out that the temporary IAM permissions role used for migration could be removed when done. ## When migrations only look complete In a forced migration, the dangerous failure is silent success. A migration tool can report "complete" even when: * the device registry doesn't exist yet * a credential is expired or mismatched * Pub/Sub wiring was never verified The fleet looks migrated and isn't. This happened: a user of the open-source migration tool [publicly reported](https://github.com/ClearBlade/clearblade-iot-core-migration/issues/2) a run that completed against an empty registry and did nothing, and asked for the docs to address it. ## What I wrote I authored the step-by-step quick start that is the front door to ClearBlade IoT Core. I built in code samples and live telemetry testing. The core steps follow the Google quick start customers already knew: * **Create the registry first** so devices can't migrate into a void * **Generate the device keypair in-flow** so you don't assume it exists * **Make Pub/Sub wiring explicit** so it isn't inherited invisibly from the old service * **End on proof, not config** so the guide is done when you see the sample device's messages arrive in Pub/Sub, not when a command returns success Each step is aimed at a specific way a migration could look finished without being finished: you set that piece up and see it work on a sample device before the fleet moves. ## Result The [quick start](https://docs.clearblade.com/iotcore/quick-start) remains the published entry point. It's one page in a larger set of quick starts, how-tos, reference, and migration tooling that supported 250+ customers off a discontinued service. The migration guides in that set cut support tickets by 30%. ======================================================================== Page: https://jslavin-docs.github.io/writing-samples/novadeploy-gitops-admin-guide-portfolio-cut/ Description: Concise portfolio cut of a fictional NovaDeploy GitOps administration guide for Amazon EKS, Argo CD, ESO, secrets, verification, and rollback workflows. ======================================================================== # NovaDeploy Platform: GitOps Administration Guide - Portfolio Cut *Deploying Services to Amazon EKS with Argo CD* Version 1.0 | Status: Portfolio cut | Written by: Jeff Slavin This is the short version of a runbook: the steps an operator follows to release a change to a live service, refresh the passwords it depends on without exposing them, and undo the release if it fails. [Read the full runbook.](https://jslavin-docs.github.io/writing-samples/novadeploy-gitops-admin-guide-full-version/) !!! note "Portfolio Notice" NovaDeploy is a fictional platform created for portfolio purposes. This sample contains no proprietary employer, client, or production information. !!! info "Scope and Audience" **Scope:** Deploy a fictional production service and refresh its secrets (such as passwords) on Amazon Elastic Kubernetes Service (EKS) using Argo CD, Terraform, AWS Identity and Access Management (IAM), AWS Key Management Service (KMS), External Secrets Operator (ESO), and Reloader. **Audience:** Platform engineers, DevOps/SRE practitioners, engineering managers, and technical writing reviewers. ??? abstract page-contents "Contents" - [1. At-a-Glance Deployment Path](#1-at-a-glance-deployment-path) - [2. Decision Walkthrough: API Gateway Secret Refresh](#2-decision-walkthrough-api-gateway-secret-refresh) - [3. Core Guardrails](#3-core-guardrails) - [4. Architecture Overview](#4-architecture-overview) - [5. Repository and Sync Policy](#5-repository-and-sync-policy) - [6. Verification Pattern](#6-verification-pattern) - [7. Implementation Excerpt: CI Reloader Guardrail](#7-implementation-excerpt-ci-reloader-guardrail) - [8. Rollback Matrix](#8-rollback-matrix) ## 1. At-a-Glance Deployment Path !!! success "Standard Deployment Path" Check tools and controllers -> update Git and Terraform-managed IAM/KMS metadata -> open a pull request (PR) -> pass CI (the automated checks) and platform review -> merge to protected main -> sync manually or automatically -> verify health, secrets, and rollout -> roll back if needed. !!! warning "Stop Checkpoints" Stop if controllers are unhealthy, CI fails, a workload using secrets lacks the Reloader annotation on root metadata, an ExternalSecret is not Ready, Argo CD is not Synced/Healthy, or a check requires printing a secret value. | Step | Operator Action | Evidence | | --- | --- | --- | | 1 | Check local tools and controller health. | Tool versions; Argo CD, ESO, and Reloader Running/Ready | | 2 | Update configuration in Git and Terraform-managed cloud metadata. | PR diff contains no plaintext secrets | | 3 | Pass CI and platform review. | lint, helm template, kubeconform, secret scan, Reloader guardrail | | 4 | Merge to main and sync. | argocd app get shows Synced / Healthy | | 5 | Verify release and secrets without exposing values. | rollout status, ExternalSecret Ready=True, key names present, secret-mounted | | 6 | Close or roll back. | Closed ticket with final health evidence, or revert PR, approval, and health evidence after rollback | --- ## 2. Decision Walkthrough: API Gateway Secret Refresh This example verifies a secret refresh, including workload restarts, while protecting secret values and keeping Git as the authoritative configuration. | Stage | Evidence Snapshot | What It Proves | | --- | --- | --- | | PR opened | `PR #1842` diff includes the chart, production values, ExternalSecret, root Reloader annotation, and `nova/api-gateway/db` reference; secret scan reports no plaintext values. | The change is Git-tracked and safe to inspect. | | CI completed | `lint`, `helm template`, `kubeconform`, `secret scan`, and Reloader guardrail pass. | The workload has the required restart control before merge. | | Argo CD before sync | `api-gateway` is `Synced / Healthy` at commit `7c4e91a`. | The starting state is stable. | | Argo CD after sync | `api-gateway` syncs to `9f28b6c` and returns `Synced / Healthy`. | The cluster matches the merged Git configuration. | | ExternalSecret verified | `Ready=True` and `SecretSynced`. | ESO created or updated the Kubernetes Secret object. | | Secret checked safely | Secret object exists; key-name output shows `DATABASE_PASSWORD`. | Expected keys are present without printing or decoding values. | | Reloader rollout confirmed | Rollout succeeds; pods are newer than the Secret refresh; last-reloaded annotation is present. | The refresh triggered a controlled rolling restart, not a manual pod delete. | | Rollback decision | No rollback: Argo CD is Healthy, ExternalSecret is Ready, mount prints only `secret-mounted`, and smoke tests pass. | Git remains authoritative. Failed checks would trigger a Git revert; Argo CD history rollback is for approved emergencies only. | --- ## 3. Core Guardrails Apply these controls during deployment, verification, and recovery. The full runbook includes commands, Terraform examples, and emergency procedures. | Control | Rule | Why It Matters | | --- | --- | --- | | GitOps source of truth | `main` is protected; every change requires a PR and passing CI. | Argo CD can restore the Git configuration and preserve an audit trail. | | Terraform source of truth | IAM, KMS, Secrets Manager metadata, rotation config, and Lambda permissions stay in Terraform. | Cloud permissions remain reviewable, reproducible, and importable after break-glass work. | | No plaintext secrets | Secret values never enter Git, Terraform state, PRs, CI logs, tickets, or chats. | Reviewers can validate controls without exposing credentials. | | Separate IAM roles for service accounts (IRSA) | The workload role never reads Secrets Manager; the dedicated ESO reader role is limited to `nova//*`. | Application pods do not receive broad secret-read permissions. | | Reloader safety | Workloads using secrets carry `reloader.stakater.com/auto: "true"` on root workload metadata. | Secret refreshes trigger controlled rolling restarts. | | Argo CD compatibility | Application defines `ignoreDifferences` for the Reloader annotation and sets `RespectIgnoreDifferences=true`. | Argo CD does not undo Reloader restart patches during sync. | | Rotation gate | Keep `var.rotation_enabled=false` until KMS, Lambda, ESO, Reloader, and mount checks pass. | Enable rotation only when workloads can safely use refreshed secrets. | --- ## 4. Architecture Overview Git defines the intended cluster configuration; Terraform defines cloud control-plane resources. ```mermaid %%{init: {"theme": "base", "flowchart": {"htmlLabels": true, "nodeSpacing": 115, "rankSpacing": 85, "curve": "basis"}, "themeVariables": {"fontFamily": "Roboto, Arial, sans-serif", "fontSize": "16px", "primaryTextColor": "#111827", "secondaryTextColor": "#111827", "tertiaryTextColor": "#111827", "lineColor": "#374151", "edgeLabelBackground": "#ecfdf5"}}}%% flowchart TD subgraph gitops["GitOps path"] direction TB pr["Developer PR
opens change"] ci["CI guardrails
block unsafe diff"] main["Protected main
receives merge"] argocd["Argo CD sync
applies desired state"] eks["Amazon EKS
runs target state"] pr --> ci --> main --> argocd --> eks end subgraph cloud["Terraform-owned cloud controls"] direction TB tf["Terraform
declares cloud state"] iam["IAM roles
scope access"] kms["KMS policy
controls decrypt"] smMeta["Secrets Manager
metadata and rotation"] sm["AWS secret path
stores values"] tf --> iam --> kms --> smMeta --> sm end subgraph runtime["Runtime secret sync and refresh"] direction TB eso["ESO
syncs approved value"] k8sSecret["Kubernetes Secret
object updated"] reloader["Reloader
detects data change"] apiPatch["Kubernetes API server
metadata patch"] rollout["Workload controller
rolls pods safely"] eso --> k8sSecret --> reloader reloader -->|"Patch .spec.template
metadata"| apiPatch apiPatch -->|"Native rolling update"| rollout end eks -. "Hosts ESO + Reloader
and workloads inside EKS" .-> eso sm -->|"Scoped read only
dedicated ESO reader
IRSA role
path nova/<service>/*"| eso style gitops fill:#eef2ff,stroke:#4338ca,stroke-width:2px,color:#312e81 style cloud fill:#fffbeb,stroke:#d97706,stroke-width:2px,color:#78350f style runtime fill:#ecfdf5,stroke:#0d9488,stroke-width:2px,color:#134e4a classDef gitopsNode fill:#f5f7ff,stroke:#4338ca,color:#111827,stroke-width:2px classDef cloudNode fill:#fff7ed,stroke:#b45309,color:#111827,stroke-width:2px classDef runtimeNode fill:#f0fdfa,stroke:#0d9488,color:#111827,stroke-width:2px class pr,ci,main,argocd,eks gitopsNode class tf,iam,kms,smMeta,sm cloudNode class eso,k8sSecret,reloader,apiPatch,rollout runtimeNode ``` !!! note "Accessible Diagram Summary" The diagram shows three flows: GitOps deployment, Terraform cloud controls, and secret refresh. Reviewed changes pass CI, merge to protected main, and reach Amazon EKS through Argo CD. Terraform defines IAM, KMS, Secrets Manager metadata, rotation configuration, and the approved AWS secret path. Amazon EKS runs ESO, Reloader, application pods, and other controllers. Only ESO reads Secrets Manager, using a dedicated IRSA role limited to `nova//*`. Application pods do not receive broad Secrets Manager read access. ESO syncs the approved value into a Kubernetes Secret. Reloader detects the change and patches workload Pod template metadata, triggering a rolling restart by the workload controller. --- ## 5. Repository and Sync Policy ```text nova-gitops/ apps/ # Argo CD Application manifests clusters/production/ # AppProject, root app, namespaces, policy baseline charts// # Service Helm chart envs/production/values/ # Production value overrides secrets/external/ # ExternalSecret CRs only; no plaintext secrets infra/iam/.tf # IAM, KMS, Secrets Manager metadata, rotation config scripts/check-reloader-annotations.sh .github/workflows/ # lint, render, kubeconform, secret scan, guardrails ``` Argo CD watches protected `main`. Automatic pruning deletes resources removed from Git; self-healing corrects differences from Git. Sync windows must block routine syncs outside approved times unless an incident-approved manual-sync override is enabled. Deleting a resource from Git requires a PR showing the removal, passing CI, platform approval, and a merge through protected `main` before Argo CD can prune it. The cluster baseline pre-creates production namespaces; service Applications do not rely on `CreateNamespace=true`. !!! warning "Auto-Prune Boundary" Enable `prune: true` only within the production AppProject and sync windows. Without these controls, a bad merge, wrong path, or unauthorized destination can trigger automatic deletion. Set both `ignoreDifferences` and `RespectIgnoreDifferences=true` so Argo CD ignores the Reloader-managed field during comparison and sync. ```yaml # Excerpt from Application.spec. # Required surrounding control: this Application belongs to the restricted # production AppProject, which limits source repos, destinations, resource # kinds, and sync windows. project: novadeploy-production ignoreDifferences: - group: apps kind: Deployment jsonPointers: - /spec/template/metadata/annotations/reloader.stakater.com~1last-reloaded-from - group: apps kind: StatefulSet jsonPointers: - /spec/template/metadata/annotations/reloader.stakater.com~1last-reloaded-from - group: apps kind: DaemonSet jsonPointers: - /spec/template/metadata/annotations/reloader.stakater.com~1last-reloaded-from syncPolicy: automated: prune: true selfHeal: true # Reverts manual drift back to reviewed Git state. syncOptions: - ServerSideApply=true - RespectIgnoreDifferences=true ``` --- ## 6. Verification Pattern After every sync, check health, rollout status, ExternalSecret readiness, Secret existence and key names, mount success, and Reloader state. Never decode, print, paste, or include secret values in tickets. ```bash argocd app get --refresh argocd app wait --health kubectl rollout status deployment/ -n kubectl get externalsecret -app-secrets -n kubectl describe externalsecret -app-secrets -n kubectl get secret -app-secrets -n kubectl get secret -app-secrets -n \ -o go-template='{{range $k, $_ := .data}}{{printf "%s\n" $k}}{{end}}' # Expected: ExternalSecret Ready=True and expected key names are present. # Never print or decode values. ``` | Check | Pass Criteria | Forbidden Evidence | | --- | --- | --- | | Argo CD state | Synced / Healthy | Manual kubectl patch not represented in Git | | ExternalSecret | Ready=True and SecretSynced reason | Secret value output | | Kubernetes Secret | Object exists; expected key names are present | Decoded data or base64 payload | | Mount check | Disposable pod prints only secret-mounted | cat/print of mounted file content | | Reloader rollout | Pods recreated after Secret refresh; app remains healthy | Secret payload in logs, tickets, or screenshots | --- ## 7. Implementation Excerpt: CI Reloader Guardrail This CI check fails a PR if a workload using secrets lacks the required Reloader annotation. The full runbook includes ServiceAccount, SecretStore, IAM, KMS, and rotation examples. ```bash set -euo pipefail rendered="$(mktemp)" trap 'rm -f "$rendered"' EXIT release_name="" # Application.metadata.name target_namespace="" # Application.spec.destination.namespace helm template "${release_name}" charts/ \ --namespace "${target_namespace}" \ -f envs/production/values/.yaml \ > "$rendered" python3 - "$rendered" <<'PY' import sys, yaml WORKLOADS = {"Deployment", "StatefulSet", "DaemonSet"} SECRET_KEYS = {"secretKeyRef", "secretRef", "secretName", "secret"} # "secret": projected volume sources def uses_secret(node): if isinstance(node, dict): return bool(SECRET_KEYS & node.keys()) or any(uses_secret(v) for v in node.values()) if isinstance(node, list): return any(uses_secret(v) for v in node) return False missing = [] with open(sys.argv[1], encoding="utf-8") as rendered: for obj in yaml.safe_load_all(rendered): if not isinstance(obj, dict) or obj.get("kind") not in WORKLOADS: continue meta = obj.get("metadata") or {} pod = obj.get("spec", {}).get("template", {}).get("spec", {}) annotations = meta.get("annotations") or {} if uses_secret(pod) and annotations.get("reloader.stakater.com/auto") != "true": missing.append(f'{obj["kind"]}/{meta.get("name", "")}') if missing: print('ERROR: secret-consuming workloads missing reloader.stakater.com/auto="true":', file=sys.stderr) print("\n".join(f" - {item}" for item in missing), file=sys.stderr) sys.exit(1) PY ``` Use the Argo CD Application name as the Helm release name, unless `source.helm.releaseName` overrides it. Match `--namespace` to `spec.destination.namespace`. The script checks the restart requirement, not the full secret lifecycle. --- ## 8. Rollback Matrix !!! warning "Rollback Principle" Use Git revert by default to keep Git authoritative and preserve an audit trail. Reserve Argo CD history rollback for approved emergencies (break-glass), followed by a Git revert within 24 hours. | Scenario | Strategy | Operator Note | | --- | --- | --- | | Bad image tag promoted | Git revert | Revert the image-bump commit, pass CI, merge, then sync or wait for automation. | | Wrong Helm values or Application manifest | Git revert | Revert the change in Git so it remains authoritative. | | Application unreachable and service-level agreement (SLA) at risk | Argo CD history rollback | Use only if Argo CD and the Kubernetes API are reachable and Git revert cannot meet the SLA. Follow the break-glass sequence below. | | GitHub or CI outage blocks revert | Argo CD history rollback | Roll back to the last-good revision while Git or CI is unavailable, and record non-secret evidence. Follow the break-glass sequence below. | | Secret value misconfiguration | Secrets Manager rollback + ESO re-sync | Roll back through the approved secret process. Use Git revert only for SecretStore, ExternalSecret, IAM, KMS, or rotation-config changes. | | Cluster unreachable | Infrastructure troubleshooting | Do not use Argo CD. Troubleshoot EKS control plane, networking, IAM, and node health first. | For Argo CD history rollback: 1. Record the root and target Applications' current sync-policy settings in the incident ticket. 2. Suspend the App-of-Apps root app. 3. Disable auto-sync on the target Application with `argocd app set --sync-policy none`. 4. Roll back to the last known good revision and verify health. 5. Keep both suspended until the matching Git revert merges, then restore their prior sync policies. ======================================================================== Page: https://jslavin-docs.github.io/writing-samples/novadeploy-gitops-admin-guide-full-version/ Description: Full fictional NovaDeploy GitOps administration guide covering Amazon EKS, Argo CD, IAM, KMS, ESO, Reloader, CI guardrails, verification, and rollback workflows. ======================================================================== # NovaDeploy Platform: GitOps Administration Guide *Deploying Services to Amazon EKS with Argo CD* Version 1.0 | Status: Full runbook | Written by: Jeff Slavin This fictional runbook covers service deployment to Amazon Elastic Kubernetes Service (EKS) with Argo CD, including access controls, encryption, secrets, GitOps, and rollback. [Read the portfolio cut.](https://jslavin-docs.github.io/writing-samples/novadeploy-gitops-admin-guide-portfolio-cut/) !!! note "Portfolio Notice" NovaDeploy is a fictional portfolio platform. This sample contains no proprietary employer, client, or production information. !!! info "Document Purpose" This runbook demonstrates documentation leadership through clear operator guidance: one source of truth, clear stop points, auditable checks, and safe rollback paths. !!! info "Scope and Audience" **Scope:** Deploy and recover NovaDeploy services on Amazon EKS with Argo CD. Covers GitOps, AWS Identity and Access Management (IAM), AWS Key Management Service (KMS), Secrets Manager, External Secrets Operator (ESO), Reloader, verification, and rollback. Excludes application-code changes, broader incident response, and service-specific business logic. **Audience:** Platform engineers, DevOps/SRE operators, cloud engineers, and documentation reviewers working with GitOps-managed Kubernetes services. ??? abstract page-contents "Contents" - [1. Quick Start and Stop Conditions](#1-quick-start-and-stop-conditions) - [2. Deployment Guardrails](#2-deployment-guardrails) - [3. Architecture Overview](#3-architecture-overview) - [4. Prerequisites and Tooling](#4-prerequisites-and-tooling) - [5. IAM, KMS, SecretStore, and ESO Setup](#5-iam-kms-secretstore-and-eso-setup) - [6. GitOps Repository Layout](#6-gitops-repository-layout) - [7. Argo CD Application and Sync Policy](#7-argo-cd-application-and-sync-policy) - [8. Deployment Verification](#8-deployment-verification) - [9. Rollback and Recovery](#9-rollback-and-recovery) - [10. Appendices](#10-appendices) ## 1. Quick Start and Stop Conditions Follow this workflow for standard, non-emergency production deployments. Later sections provide the implementation details. | Step | Action | What to Do | Stop Condition | | --- | --- | --- | --- | | 1 | Validate readiness | Run controller health checks, local tool checks, and the guardrail table before editing the deployment pull request (PR). | Stop if Argo CD, ESO, or Reloader is unhealthy. | | 2 | Change declared state | Update Helm values, Argo CD Application resources, ExternalSecret custom resources (CRs), or Terraform-owned IAM/KMS metadata. | Do not commit, paste, or type plaintext secrets into Git, a PR, or a shell. | | 3 | Open PR | Pass CI (the automated checks): lint, helm template, kubeconform, secret scan, and Reloader annotation guardrail. | Stop if any workload consumes a Secret without the root Reloader annotation. | | 4 | Merge to main | Merge after approval. Argo CD watches main and reconciles the application. | No direct pushes and no direct kubectl edits. | | 5 | Sync and verify | Wait for automated sync or run argocd app sync ``; then run health, smoke, and secret-mount checks. | Do not use --force for normal deployment hotfixes. | | 6 | Close or recover | Close the ticket only after Synced/Healthy, smoke-test success, and non-secret evidence is recorded. | Use Git revert by default; use Argo CD history only for approved service-level agreement (SLA) emergencies. | !!! info "Secret Handling Rule" No plaintext secrets in Git, ConfigMaps, literal environment variables, Terraform state, PRs, logs, chats, or tickets. Secret values live in AWS Secrets Manager. ESO syncs values into Kubernetes Secret objects. Reloader propagates changes by controlled rolling restart, not by exposing secret values. ## 2. Deployment Guardrails These production safety rules apply throughout the guide. | Guardrail | Required Evidence | Pass Criteria | | --- | --- | --- | | Git is source of truth | main branch protected; all changes through PR; CI passes before merge | Manual cluster drift is rejected or reverted through Argo CD self-heal. | | Terraform owns cloud controls | IAM roles, policies, KMS keys, Secrets Manager metadata, rotation config, and Lambda permissions are managed in Terraform | Use the AWS CLI for read-only checks of Terraform-managed configuration. Direct CLI changes to that configuration require approved break-glass procedures and subsequent reconciliation with Terraform. | | No plaintext secrets | Secret scan, PR review, and no aws_secretsmanager_secret_version for production values | Secret values never enter Git, Terraform state, PR comments, CI logs, chats, or tickets. | | IRSA separation | Two IAM roles for service accounts (IRSA) per service: workload ServiceAccount has non-secret AWS permissions only; dedicated ESO reader ServiceAccount assumes `nova--eso-read` | Only ESO reads AWS Secrets Manager for service-scoped paths. | | Namespace-scoped SecretStore | ExternalSecret uses secretStoreRef.kind: SecretStore in the workload namespace | Avoid ClusterSecretStore for app secrets unless a platform exception is approved. | | Reloader compatibility | Root workload metadata contains `reloader.stakater.com/auto: "true"`; the Application defines `ignoreDifferences` for the Reloader annotation and sets `RespectIgnoreDifferences=true`. | Reloader can patch pod templates without Argo CD immediately removing its annotation. | | Rotation gate | var.rotation_enabled remains false until KMS policy, Lambda role, ESO readiness, Reloader role-based access control (RBAC), and mount checks pass | Enable rotation only after every dependency is verified in staging and approved for production. | ### 2.1 Rotation Readiness Gate Enable production rotation only after each item passes in staging and the production change is approved. - Run the cluster health check in Section 4.2. - Confirm ESO can reconcile the target ExternalSecret and create/update the Kubernetes Secret. - Confirm the KMS key policy permits the ESO reader role and the rotation Lambda execution role when rotation is enabled. - Confirm Reloader can get/list/watch Secrets and ConfigMaps and patch workloads in the workload namespace. - Confirm every secret-consuming Deployment, StatefulSet, or DaemonSet has reloader.stakater.com/auto: "true" on root workload metadata. - Confirm the Argo CD Application ignores Reloader last-reloaded annotations and sets RespectIgnoreDifferences=true. - Run the secret mount check in Section 8.2. ## 3. Architecture Overview Git defines the desired cluster state; Terraform defines cloud control-plane resources; AWS Secrets Manager stores secret values; ESO syncs them into Kubernetes Secret objects. ```mermaid %%{init: {"theme": "base", "flowchart": {"htmlLabels": true, "nodeSpacing": 115, "rankSpacing": 85, "curve": "basis"}, "themeVariables": {"fontFamily": "Roboto, Arial, sans-serif", "fontSize": "16px", "primaryTextColor": "#111827", "secondaryTextColor": "#111827", "tertiaryTextColor": "#111827", "lineColor": "#374151", "edgeLabelBackground": "#ecfdf5"}}}%% flowchart TD subgraph gitops["GitOps path"] direction TB pr["Developer PR
opens change"] ci["CI guardrails
block unsafe diff"] main["Protected main
receives merge"] argocd["Argo CD sync
applies desired state"] eks["Amazon EKS
runs target state"] pr --> ci --> main --> argocd --> eks end subgraph cloud["Terraform-owned cloud controls"] direction TB tf["Terraform
declares cloud state"] iam["IAM roles
scope access"] kms["KMS policy
controls decrypt"] smMeta["Secrets Manager
metadata and rotation"] sm["AWS secret path
stores values"] tf --> iam --> kms --> smMeta --> sm end subgraph runtime["Runtime secret sync and refresh"] direction TB eso["ESO
syncs approved value"] k8sSecret["Kubernetes Secret
object updated"] reloader["Reloader
detects data change"] apiPatch["Kubernetes API server
metadata patch"] rollout["Workload controller
rolls pods safely"] eso --> k8sSecret --> reloader reloader -->|"Patch .spec.template
metadata"| apiPatch apiPatch -->|"Native rolling update"| rollout end eks -. "Hosts ESO + Reloader
and workloads inside EKS" .-> eso sm -->|"Scoped read only
dedicated ESO reader
IRSA role
path nova/<service>/*"| eso style gitops fill:#eef2ff,stroke:#4338ca,stroke-width:2px,color:#312e81 style cloud fill:#fffbeb,stroke:#d97706,stroke-width:2px,color:#78350f style runtime fill:#ecfdf5,stroke:#0d9488,stroke-width:2px,color:#134e4a classDef gitopsNode fill:#f5f7ff,stroke:#4338ca,color:#111827,stroke-width:2px classDef cloudNode fill:#fff7ed,stroke:#b45309,color:#111827,stroke-width:2px classDef runtimeNode fill:#f0fdfa,stroke:#0d9488,color:#111827,stroke-width:2px class pr,ci,main,argocd,eks gitopsNode class tf,iam,kms,smMeta,sm cloudNode class eso,k8sSecret,reloader,apiPatch,rollout runtimeNode ``` !!! note "Accessible Diagram Summary" The diagram shows three paths: GitOps, Terraform-owned cloud controls, and runtime secret sync. GitOps moves a reviewed PR through CI, protected main, Argo CD, and Amazon EKS. Terraform defines IAM, KMS, Secrets Manager metadata, rotation config, and the approved AWS secret path. Amazon EKS hosts ESO, Reloader, application pods, and other runtime controllers. Only ESO reads AWS Secrets Manager, using a dedicated IRSA role scoped to `nova//*`. Application pods do not receive broad Secrets Manager read access. ESO syncs the approved value into a Kubernetes Secret. Reloader detects the change and patches workload Pod template metadata through the Kubernetes API server, triggering a rolling restart by the workload controller. ## 4. Prerequisites and Tooling The platform team pins exact versions in the infrastructure repository. Check compatibility before opening a deployment PR. | Tool / Resource | Requirement | Purpose | | --- | --- | --- | | AWS CLI | v2; approved role | EKS auth, read-only validation, and break-glass evidence | | kubectl | Compatible with cluster | Health, rollout, RBAC, and Secret-object checks | | Helm | 3.x; platform-pinned | Chart rendering during local validation and CI | | Argo CD CLI | Compatible with server | Application status, sync, wait, history, rollback | | Terraform | Version pinned by infra repo | IAM, KMS, Secrets Manager metadata, rotation config | | External Secrets Operator | Platform-pinned; custom resource definitions (CRDs) installed | Syncs AWS Secrets Manager values to Kubernetes Secret objects | | ESO controller RBAC | `create` on `serviceaccounts/token` for ServiceAccounts referenced by `auth.jwt.serviceAccountRef` | Allows ESO to request short-lived projected tokens through the Kubernetes TokenRequest API | | Reloader | Platform-pinned; reload strategy = annotations | Triggers rolling restarts when watched Secrets/ConfigMaps change | | Python + PyYAML | Python 3.x and PyYAML | Fast CI guardrail for rendered workload annotations | | jq | 1.6 or later | Safe JSON construction during approved secret seeding | | Approved password manager or privileged access management (PAM) CLI | Platform-approved client, authenticated with multi-factor authentication (MFA) | Supplies the initial secret value to the seeding workflow without exposing it to a shell | ### 4.1 Local Tool Validation Run these checks before editing the GitOps repository. If any command fails, fix local access or tooling before opening the PR. ```bash aws --version kubectl version --client=true helm version --short argocd version --client terraform version python3 -c "import yaml; print('PyYAML available')" jq --version ``` ### 4.2 Cluster Health Check Run this before every release cycle. All controllers must be healthy before sync, rollback, or rotation work proceeds. Save the block as a script and run it with `bash`; pasted into a live shell, `set -e` and `exit 1` close the session on the first failed check. The RBAC checks use `kubectl auth can-i --as` to test controller ServiceAccount permissions. The operator or CI identity must be allowed to impersonate those accounts; most production operator roles should not have broad impersonation rights. If you lack permission, have a platform administrator or approved CI identity run this block and attach the non-secret results to the deployment ticket. ```bash #!/usr/bin/env bash set -euo pipefail fail() { printf 'ERROR: %s\n' "$1" >&2 exit 1 } require_can_i_as() { local subject="$1" local message="$2" shift 2 if ! kubectl auth can-i "$@" --as "${subject}" --quiet; then fail "${message}" fi } kubectl wait --for=condition=Ready node --all --timeout=120s \ || fail "One or more cluster nodes are not Ready." kubectl wait --for=condition=Available deployment --all \ -n argocd --timeout=120s \ || fail "One or more Argo CD deployments are unavailable." kubectl rollout status statefulset/argocd-application-controller \ -n argocd --timeout=120s \ || fail "Argo CD application controller is not ready." kubectl wait --for=condition=Available deployment --all \ -n --timeout=120s \ || fail "One or more ESO deployments are unavailable." kubectl wait --for=condition=Available deployment --all \ -n --timeout=120s \ || fail "Reloader is unavailable." RELOADER_STRATEGY=$(kubectl get deploy \ -n \ -o jsonpath='{.spec.template.spec.containers[*].args}{" "}{.spec.template.spec.containers[*].env[?(@.name=="RELOAD_STRATEGY")].value}' \ | tr '",[]' ' ') printf '%s\n' "${RELOADER_STRATEGY}" \ | grep -Eqi '(^|[[:space:]])(--)?reload-strategy[=[:space:]]+annotations([[:space:]]|$)|(^|[[:space:]])annotations([[:space:]]|$)' \ || fail "Reloader is not configured with the required annotations reload strategy." ESO_SA=$(kubectl get deploy \ -n \ -o jsonpath='{.spec.template.spec.serviceAccountName}') RELOADER_SA=$(kubectl get deploy \ -n \ -o jsonpath='{.spec.template.spec.serviceAccountName}') test -n "${ESO_SA}" \ || fail "Could not resolve the ESO controller ServiceAccount." test "${RELOADER_SA}" = "" \ || fail "Reloader is using '${RELOADER_SA}', not the expected ''." ESO_SUBJECT="system:serviceaccount::${ESO_SA}" RELOADER_SUBJECT="system:serviceaccount::${RELOADER_SA}" require_can_i_as \ "${ESO_SUBJECT}" \ "ESO cannot create TokenRequest objects for the referenced ServiceAccount in ." \ create serviceaccounts/-eso-secret-reader \ --subresource=token -n for verb in get list watch; do require_can_i_as \ "${RELOADER_SUBJECT}" \ "Reloader cannot ${verb} Secrets in ." \ "${verb}" secrets -n require_can_i_as \ "${RELOADER_SUBJECT}" \ "Reloader cannot ${verb} ConfigMaps in ." \ "${verb}" configmaps -n done for workload in deployments.apps statefulsets.apps daemonsets.apps; do for verb in get list update patch; do require_can_i_as \ "${RELOADER_SUBJECT}" \ "Reloader cannot ${verb} ${workload} in ." \ "${verb}" "${workload}" -n done done ``` The strategy check normalizes the `jsonpath` array for `args` to match both `--reload-strategy=annotations` and `--reload-strategy annotations`. Confirm the flag and environment-variable names against the platform-pinned Reloader chart version before relying on this check. | Component | Pass Criteria | | --- | --- | | Nodes | All schedulable nodes report Ready and no unexpected NoSchedule taints. | | Argo CD | server, repo-server, application-controller, and dex are Running. | | ESO | Controller and webhook are Running; the controller can `create` `serviceaccounts/token` for the namespace that contains the referenced ServiceAccount; ExternalSecret status becomes Ready after apply. | | Reloader | Live deployment uses the annotations reload strategy and can read watched Secrets/ConfigMaps and update each supported workload type. | | EKS API-data encryption | Clusters below Kubernetes 1.28 must have explicit Secrets envelope encryption configured; EKS clusters running 1.28 or later receive default envelope encryption for all Kubernetes API data. See the [AWS default envelope encryption documentation](https://docs.aws.amazon.com/eks/latest/userguide/envelope-encryption.html). | ## 5. IAM, KMS, SecretStore, and ESO Setup !!! info "Section Summary" Create two narrowly scoped IRSA roles per service: a workload role for non-secret AWS access, and an ESO reader role for service-scoped Secrets Manager reads and KMS decryption through Secrets Manager. Terraform manages the cloud resources; Kubernetes manifests bind the matching ServiceAccounts. ### 5.1 Role Model | Role / Account | Used By | Allowed Access | Explicitly Not Allowed | | --- | --- | --- | --- | | `nova--prod` | Workload ServiceAccount `` | Only the non-secret AWS APIs the application needs, such as S3 or DynamoDB | No Secrets Manager read permissions | | `nova--eso-read` | ServiceAccount `-eso-secret-reader` | `secretsmanager:GetSecretValue`, `DescribeSecret`, and `ListSecretVersionIds` for `nova//*`, plus KMS decrypt through Secrets Manager | No paths outside `nova//*`; no trust for other service accounts | | rotation Lambda role | Approved Secrets Manager rotation Lambda | Rotation-only actions and KMS use through Secrets Manager when rotation is enabled | Not present in KMS policy while var.rotation_enabled=false | ### 5.2 Terraform Pattern Confirm the EKS OpenID Connect (OIDC) issuer before provisioning IRSA. This check is read-only; create IAM resources through Terraform. ```bash aws eks describe-cluster \ --name \ --region \ --query "cluster.identity.oidc.issuer" \ --output text ``` Create or update service IAM resources through `infra/iam/.tf`. This example restricts the ESO reader role to the dedicated ESO secret-reader ServiceAccount. ```hcl locals { oidc_provider = replace(var.oidc_provider_url, "https://", "") eso_sa_sub = "system:serviceaccount:${var.namespace}:${var.service}-eso-secret-reader" } data "aws_iam_policy_document" "eso_assume_role" { statement { actions = ["sts:AssumeRoleWithWebIdentity"] principals { type = "Federated" identifiers = ["arn:aws:iam::${var.account_id}:oidc-provider/${local.oidc_provider}"] } condition { test = "StringEquals" variable = "${local.oidc_provider}:aud" values = ["sts.amazonaws.com"] } condition { test = "StringEquals" variable = "${local.oidc_provider}:sub" values = [local.eso_sa_sub] } } } resource "aws_iam_role" "eso_read" { name = "nova-${var.service}-eso-read" assume_role_policy = data.aws_iam_policy_document.eso_assume_role.json } ``` Attach the secret-read policy only to `nova--eso-read`, never to the workload role. Pass the Amazon Resource Name (ARN) of the KMS key that encrypts the service secret through the typed `secrets_kms_key_arn` input. ```hcl variable "secrets_kms_key_arn" { description = "ARN of the KMS key that encrypts this service's Secrets Manager secrets" type = string } data "aws_iam_policy_document" "eso_read" { statement { actions = [ "secretsmanager:GetSecretValue", "secretsmanager:DescribeSecret", "secretsmanager:ListSecretVersionIds" ] resources = [ "arn:aws:secretsmanager:${var.region}:${var.account_id}:secret:nova/${var.service}/*" ] } statement { actions = ["kms:Decrypt"] resources = [var.secrets_kms_key_arn] condition { test = "StringEquals" variable = "kms:ViaService" values = ["secretsmanager.${var.region}.amazonaws.com"] } } } resource "aws_iam_role_policy" "eso_read" { name = "nova-${var.service}-eso-read" role = aws_iam_role.eso_read.id policy = data.aws_iam_policy_document.eso_read.json } ``` !!! info "KMS Source of Truth" The same KMS key ARN must be used by the Secrets Manager secret, the `nova--eso-read` IAM policy, the KMS key policy, and the rotation Lambda role policy. If the platform uses a shared externally managed key, verify the key policy before merge. ### 5.3 ServiceAccount and SecretStore Manifests Prefer keeping ServiceAccount manifests in charts/ to version-control IAM bindings. Keep the workload and ESO reader ServiceAccounts separate. ```yaml apiVersion: v1 kind: ServiceAccount metadata: name: -eso-secret-reader namespace: annotations: eks.amazonaws.com/role-arn: arn:aws:iam:::role/nova--eso-read ``` Define each service's SecretStore in its workload namespace. Use ClusterSecretStore for application secrets only with a platform-approved cross-namespace exception. ```yaml apiVersion: external-secrets.io/v1 kind: SecretStore metadata: name: aws-secrets-manager- namespace: spec: provider: aws: service: SecretsManager region: auth: jwt: serviceAccountRef: name: -eso-secret-reader --- apiVersion: external-secrets.io/v1 kind: ExternalSecret metadata: name: -app-secrets namespace: spec: refreshInterval: 5m secretStoreRef: name: aws-secrets-manager- kind: SecretStore target: name: -app-secrets creationPolicy: Owner data: - secretKey: DATABASE_PASSWORD remoteRef: key: nova//db property: password ``` remoteRef.key must match the Terraform-managed Secrets Manager name pattern: `nova//`. Do not use kubectl, Git, or Terraform to set production secret values. ### 5.4 Approved Secret-Seeding Workflow Seeding sets the initial secret value and is the only approved human write path for production secrets. Terraform manages Secrets Manager metadata, KMS policy, and rotation config; Section 2 forbids `aws_secretsmanager_secret_version` for production values. An approved administrator creates the first AWSCURRENT version once, using this procedure. !!! warning "Where seeding is allowed to happen" If your organization forbids production secrets on workstations, seed through the PAM session broker, an approved bastion or jump host, or a CI job. The CI job must assume the seeding role through OIDC and read the value from the approved secret broker. The commands are the same; only the host and identity change. Record the path used in the deployment ticket. 1. A platform administrator retrieves the initial value from the approved password manager or PAM workflow. 2. The administrator opens a private session with MFA. **The secret value is never typed, pasted, echoed, or interpolated into a shell.** It flows from the password manager through the input stream to AWS and appears nowhere else. !!! danger "Do not disable session controls" Keep shell history and terminal recording enabled. PAM session capture is required, and step 7 relies on its audit trail. Both forms below keep the value out of process arguments (`argv`) and shell history without disabling these controls. 3. Use one of these approved forms. Both read directly from the password-manager CLI, keeping the value out of the shell prompt, `argv`, and history. **Form A, no file on disk (preferred).** Process substitution passes a file descriptor to the AWS CLI without writing a plaintext copy to the filesystem. ```bash # Requires bash or zsh. Use Form B in a POSIX shell. aws secretsmanager put-secret-value \ --secret-id nova// \ --secret-string file://<( read "" \ | jq -Rn '{password: input}') ``` **Form B, temporary file with restricted permissions.** Set the umask *before* creating the file; applying `chmod` afterward leaves a window in which it is world-readable. ```bash umask 077 # remove group/other permissions from new files if SECURE_DIR="$(mktemp -d "/nova-seed.XXXXXX")"; then SECRET_FILE="${SECURE_DIR}/-secret.json" read "" \ | jq -Rn '{password: input}' > "${SECRET_FILE}" ls -l "${SECRET_FILE}" # confirm -rw------- before continuing else echo "Could not create the secure directory. Stop here." >&2 fi ``` !!! danger "Never place the value in argv" Do not use `--secret-string "$( read ...)"` or hand-write the JSON in a heredoc. Command substitution exposes plaintext in process arguments to `ps` and local processes while the command runs. A heredoc requires pasting the value into the terminal, which step 2 forbids. !!! note "If no password-manager CLI is available" Run Form B's `umask` line and `if` block with `: > "${SECRET_FILE}"` in place of the password-manager pipeline; that creates the empty file, and `ls -l` must show `-rw-------`. Save the value to that path with the password manager's own save-to-file function, then confirm `-rw-------` again. The file must be JSON with a `password` string field, for example `{"password":""}`, not a raw password or a vendor export. Do not route it through the terminal, the clipboard, or an editor buffer. These examples accept single-line passwords only. The command safely escapes quotes and backslashes but reads only the first line. Do not build the JSON by hand. Use an approved multiline-safe workflow for values containing line breaks. 4. Create the first AWSCURRENT version with `put-secret-value` and a file reference. Form A already does this; for Form B, use the file from step 3. ```bash aws secretsmanager put-secret-value \ --secret-id nova// \ --secret-string "file://${SECRET_FILE}" ``` 5. Verify with `describe-secret` only. Do not use `get-secret-value` during deployment verification. ```bash aws secretsmanager describe-secret \ --secret-id nova// \ --query "{Name:Name,VersionIdsToStages:VersionIdsToStages,KmsKeyId:KmsKeyId}" ``` 6. Form B only: remove the temporary file and directory immediately after seeding. ```bash shred -u "${SECRET_FILE}" 2>/dev/null || rm -f "${SECRET_FILE}" rmdir "${SECURE_DIR}" unset SECRET_FILE SECURE_DIR ``` !!! note "shred is not a guarantee" On copy-on-write filesystems and SSDs with wear leveling, `shred` cannot reliably overwrite the original blocks. Prefer Form A. Memory-backed storage such as `/dev/shm` can still write secrets to disk through swap. 7. Record only non-secret evidence in the deployment ticket: secret ARN/name, KMS key ID, AWSCURRENT version ID, seeding path used (workstation, PAM, bastion, or CI), approver, timestamp, and rotation-readiness status. ## 6. GitOps Repository Layout One GitOps repository defines NovaDeploy cluster state: manifests, Helm overrides, ExternalSecret resources, cluster baselines, and infrastructure modules. Application source code lives in separate repositories. ```text nova-gitops/ apps/ # Argo CD Application manifests clusters/production/ # AppProject, root app, namespaces, policy baseline charts// # Service Helm chart envs/production/values/ # Production value overrides secrets/external/ # ExternalSecret CRs only; no plaintext secrets infra/iam/.tf # IAM, KMS, Secrets Manager metadata, rotation config scripts/check-reloader-annotations.sh .github/workflows/ # lint, render, kubeconform, secret scan, guardrails ``` | Path | Owner | Review Focus | | --- | --- | --- | | `apps/` | Platform engineering | Application project, destination, sync policy, `ignoreDifferences` | | `clusters/production/` | Platform engineering | AppProject, sync windows, namespace baseline | | `charts//` | Service team + platform reviewer | Workload metadata annotations, probes, resources, service accounts | | `envs/production/values/` | Service team | Image tag, config values, environment-specific overrides | | `secrets/external/` | Platform engineering | ExternalSecret references only; no secret values | | `infra/iam/.tf` | Platform engineering | IAM trust boundaries, KMS policy, Secrets Manager metadata, rotation gates | | `scripts/` | Platform engineering | Guardrail correctness, fail-closed behavior, and portability | | `.github/workflows/` | Platform engineering | Required checks, pinned actions, and least-privilege workflow permissions | ## 7. Argo CD Application and Sync Policy This Application manifest combines automated sync, server-side apply, and Reloader compatibility. Create production namespaces through clusters/production/ so NetworkPolicy, ResourceQuota, LimitRange, labels, and admission policies exist before workload sync. !!! warning "Auto-Prune Boundary" Enable `prune: true` in production only when the production AppProject and sync windows constrain the Application. Without these controls, a bad merge, path mistake, or unauthorized destination can trigger automatic deletion. Configure both settings: `ignoreDifferences` excludes the Reloader-managed annotation when Argo CD compares live and desired state; `RespectIgnoreDifferences=true` also applies that exclusion during sync. ```yaml apiVersion: argoproj.io/v1alpha1 kind: Application metadata: name: api-gateway namespace: argocd spec: project: novadeploy-production source: repoURL: https://github.com/novadeploy/nova-gitops targetRevision: main path: charts/api-gateway helm: valueFiles: - ../../envs/production/values/api-gateway.yaml destination: server: https://kubernetes.default.svc namespace: api-gateway syncPolicy: automated: prune: true selfHeal: true syncOptions: - ServerSideApply=true - RespectIgnoreDifferences=true ignoreDifferences: - group: apps kind: Deployment jsonPointers: - /spec/template/metadata/annotations/reloader.stakater.com~1last-reloaded-from - group: apps kind: StatefulSet jsonPointers: - /spec/template/metadata/annotations/reloader.stakater.com~1last-reloaded-from - group: apps kind: DaemonSet jsonPointers: - /spec/template/metadata/annotations/reloader.stakater.com~1last-reloaded-from ``` | Setting | Meaning | Operational Note | | --- | --- | --- | | prune: true | Resources removed from Git are removed from the cluster on sync. | Treat deletions as production changes; require review. | | selfHeal: true | Manual drift is reverted to Git state. | Do not hotfix production with direct kubectl edits. | | ServerSideApply=true | Kubernetes tracks field ownership during apply. | Useful for resources not fully managed by Argo CD. Applies with `--force-conflicts`, so a conflicting field is taken over, not blocked. | | RespectIgnoreDifferences=true | Argo CD respects ignoreDifferences during sync. | Prevents sync from removing Reloader last-reloaded annotations. | | No CreateNamespace=true | Production namespaces are not created ad hoc by service apps. | Cluster baseline creates namespaces with required policy first. | ### 7.1 Sync Windows Sync windows are defined in the AppProject. This project has a matching `allow` window, so automated and manual syncs are blocked outside that schedule. An optional `deny` window can block shorter periods within the allowed schedule; it is not needed for nights or weekends. | Window / State | Effect | Operator Action | | --- | --- | --- | | Monday-Friday 09:00-17:00 UTC; matching allow window active | Routine automated or manual sync is permitted after PR approval. | Sync normally, then complete Section 8 verification. | | All other times; no matching allow window active | Routine sync is blocked by default, including Monday-Thursday overnight. | Wait for the next allow window unless an approved incident requires a manual override. | | Approved emergency manual override | `manualSync` is temporarily enabled on the matching allow window. | Enable the override by window ID, perform one manual sync, verify, and disable the override immediately. | ## 8. Deployment Verification Close the deployment ticket only after all checks pass and the evidence contains no secret values. Validate object state, expected key names, mount success, rollout state, and recent pod creation time only. ### 8.1 Health and Secret Checks ```bash argocd app get --refresh argocd app wait --health kubectl rollout status deployment/ -n kubectl get externalsecret -app-secrets -n kubectl describe externalsecret -app-secrets -n kubectl get secret -app-secrets -n kubectl get secret -app-secrets -n \ -o go-template='{{range $k, $_ := .data}}{{printf "%s\n" $k}}{{end}}' # Expected: ExternalSecret Ready=True and expected key names are present. # Never print or decode values. ``` ### 8.2 Secret Mount Check This disposable pod checks that the Secret mounts and is readable without exposing values. It prints only `secret-mounted` on success. The manifest meets the restricted Pod Security profile and namespace resource controls in `clusters/production/` (Section 7). These namespaces reject pods without a `securityContext` or `resources` block before attempting a mount. Run this check in the Secret's workload namespace; mounts cannot be checked across namespaces. ```bash cat <<'EOF' | kubectl apply -n -f - apiVersion: v1 kind: Pod metadata: name: secret-mount-check spec: restartPolicy: Never automountServiceAccountToken: false securityContext: runAsNonRoot: true runAsUser: 1000 runAsGroup: 1000 fsGroup: 1000 seccompProfile: type: RuntimeDefault containers: - name: secret-mount-check image: busybox:1.36 command: ["sh", "-ec", "test -s /mnt/secrets/DATABASE_PASSWORD && test -r /mnt/secrets/DATABASE_PASSWORD && echo secret-mounted"] securityContext: allowPrivilegeEscalation: false capabilities: drop: ["ALL"] resources: requests: cpu: 10m memory: 16Mi limits: cpu: 50m memory: 32Mi volumeMounts: - name: app-secrets mountPath: /mnt/secrets readOnly: true volumes: - name: app-secrets secret: secretName: -app-secrets defaultMode: 0440 EOF kubectl wait --for=jsonpath='{.status.phase}'=Succeeded pod/secret-mount-check \ -n --timeout=60s kubectl logs secret-mount-check -n kubectl delete pod secret-mount-check -n --ignore-not-found ``` | Field | Why it is required | | --- | --- | | `runAsNonRoot: true` plus `runAsUser` | Restricted Pod Security requires a non-root user. With `busybox:1.36`, `runAsNonRoot` alone causes `CreateContainerConfigError` because the image defaults to root. An explicit user ID (UID) is required. | | `allowPrivilegeEscalation: false`, `capabilities.drop: ["ALL"]`, `seccompProfile.type: RuntimeDefault` | Remaining restricted-profile requirements. Omitting any one rejects the pod at admission. | | `resources` requests and limits | The namespace baseline applies ResourceQuota and LimitRange. A pod with no resources block is rejected by a quota covering requests or limits unless a LimitRange supplies defaults. | | `automountServiceAccountToken: false` | A disposable debug pod in a namespace built on IRSA separation must not receive a projected ServiceAccount token. | | `fsGroup: 1000` with `defaultMode: 0440` | Kubernetes sets the mounted Secret's group ownership to `fsGroup`, allowing the non-root UID to read it regardless of cluster defaults. | !!! warning "Mount mode and UID are coupled" The default Secret mode, 0644, allows any UID to read the file. With mode 0400 and no `fsGroup`, a non-root container cannot read it: `test -s` passes because the file exists and is not empty, but `test -r` fails. This permissions problem can look like a failed mount. Set `defaultMode` and `fsGroup` explicitly, as above, to avoid that ambiguity. !!! note "kubectl compatibility" `--for=jsonpath` requires kubectl 1.23 or later. The pod exits immediately after the check, so `--for=condition=Ready` can miss it and time out even after success. On older clients, poll instead: `kubectl get pod secret-mount-check -n -o jsonpath='{.status.phase}'`. **Pass criteria.** The pod reaches Succeeded, the log contains exactly `secret-mounted`, and the pod is deleted. If admission rejects the pod, first align the manifest with the namespace's Pod Security level and resource controls. That rejection does not establish a problem with the Secret. ### 8.3 Reloader Confirmation ```bash kubectl rollout status deployment/ -n kubectl get deploy -n \ -o go-template='{{ index .spec.template.metadata.annotations "reloader.stakater.com/last-reloaded-from" }}{{ "\n" }}' kubectl get pods -n -l app= \ --sort-by=.metadata.creationTimestamp argocd app get --refresh # Expected: pods were recreated after the Secret refresh; app remains Synced / Healthy. ``` ## 9. Rollback and Recovery !!! warning "Rollback Principle" Use Git revert by default to preserve the source of truth and audit trail. Argo CD history rollback requires an approved emergency exception (break-glass). The matching Git revert must then merge to bring Git back in line with the cluster. | Scenario | Strategy | Operator Note | | --- | --- | --- | | Bad image tag promoted | Git revert | Revert the image-bump commit, pass CI, merge, then sync or wait for automation. | | Wrong Helm values or Application manifest | Git revert | Revert the change in Git so it continues to define the desired state. | | Application unreachable and SLA at risk | Argo CD history rollback | Use only if Argo CD and the Kubernetes API are reachable and Git revert cannot meet the SLA. Follow Section 9.2. | | GitHub or CI outage blocks revert | Argo CD history rollback | Roll back to the last-good revision while Git or CI is unavailable, and record non-secret evidence. Follow Section 9.2. | | Secret value misconfiguration | Secrets Manager rollback + ESO re-sync | Roll back through the approved secret process. Use Git revert only for SecretStore, ExternalSecret, IAM, KMS, or rotation-config changes. | | Cluster unreachable | Infrastructure troubleshooting | Do not use Argo CD. Troubleshoot EKS control plane, networking, IAM, and node health first. | ### 9.1 Git Revert 1. Identify the bad commit SHA and the last-good commit in nova-gitops. 2. Create a revert branch from protected main. 3. Check whether the bad commit is single-parent or a merge commit. 4. Open a PR, require emergency approval, merge, then sync or wait for automation. 5. Run the full Section 8 verification path before closing the incident. ```bash git checkout main && git pull git checkout -b revert/ git show -s --format=%P # If one parent SHA is returned: git revert --no-edit # If two or more parent SHAs are returned, keep main as parent 1: git revert -m 1 --no-edit git push origin revert/ ``` If the allow schedule blocks an approved emergency sync, use this manual-sync override. Automated sync remains blocked outside the allow window. ```bash argocd proj windows list argocd proj windows enable-manual-sync argocd app sync argocd app wait --health argocd proj windows disable-manual-sync ``` ### 9.2 Argo CD History Rollback Use only when a Git revert cannot meet the SLA deadline. If an App-of-Apps root app manages child Application CRs, suspend it during the approved incident window; otherwise, it may re-enable the child app and re-sync the broken commit. Run the break-glass sequence in this order: 1. Confirm Argo CD and the Kubernetes API are reachable. 2. Record the root and target Applications' current sync-policy settings in the incident ticket. Do not export and re-apply the full live Application object: it includes server-managed fields and may bypass the Git-managed definition. 3. Suspend the App-of-Apps root app, then disable auto-sync on the target Application. 4. Roll back the target Application to the last-good revision and wait for health. 5. Keep both suspended until the matching Git revert merges. ```bash argocd app get --refresh argocd app get --refresh argocd app list --selector app.kubernetes.io/part-of= argocd app set --sync-policy none argocd app set --sync-policy none argocd app history argocd app rollback argocd app wait --health # Leave target auto-sync disabled until the mandatory Git revert merges. ``` Restore Git and cluster consistency after the incident: 1. Open a Jira ticket tagged [gitops-debt]. 2. Complete the matching Git revert within 24 hours. 3. After the revert PR merges, restore the target and root Applications with `argocd app set` using their exact pre-incident, Git-declared policies. 4. Refresh both Applications and confirm they return to their pre-incident sync policies and the reverted Git revision. The example below assumes both Applications normally use automated sync, prune, and self-heal. Remove any flag that was not enabled in the Git-managed definition. ```bash argocd app set \ --sync-policy automated \ --auto-prune \ --self-heal argocd app set \ --sync-policy automated \ --auto-prune \ --self-heal argocd app get --refresh argocd app get --refresh ``` ## 10. Appendices ### 10.1 CI Reloader Annotation Guardrail This guardrail checks each rendered workload, so one correct Deployment annotation cannot hide a missing root annotation on another workload that uses secrets. Match the Helm release name to `Application.metadata.name` and `--namespace` to `Application.spec.destination.namespace`. If `source.helm.releaseName` is set, use that override. ```bash #!/usr/bin/env bash set -euo pipefail rendered="$(mktemp)" trap 'rm -f "$rendered"' EXIT release_name="" # Application.metadata.name target_namespace="" # Application.spec.destination.namespace helm template "${release_name}" charts/ \ --namespace "${target_namespace}" \ -f envs/production/values/.yaml \ > "$rendered" python3 - "$rendered" <<'PY' import sys import yaml WORKLOADS = {"Deployment", "StatefulSet", "DaemonSet"} SECRET_KEYS = {"secretKeyRef", "secretRef", "secretName", "secret"} # "secret": projected volume sources def uses_secret(node): if isinstance(node, dict): return bool(SECRET_KEYS & node.keys()) or any( uses_secret(value) for value in node.values() ) if isinstance(node, list): return any(uses_secret(value) for value in node) return False missing = [] with open(sys.argv[1], encoding="utf-8") as rendered: for obj in yaml.safe_load_all(rendered): if not isinstance(obj, dict) or obj.get("kind") not in WORKLOADS: continue meta = obj.get("metadata") or {} pod = obj.get("spec", {}).get("template", {}).get("spec", {}) annotations = meta.get("annotations") or {} if ( uses_secret(pod) and annotations.get("reloader.stakater.com/auto") != "true" ): missing.append( f'{obj["kind"]}/{meta.get("name", "")}' ) if missing: print( 'ERROR: secret-consuming workloads missing ' 'reloader.stakater.com/auto="true":', file=sys.stderr, ) print( "\n".join(f" - {item}" for item in missing), file=sys.stderr, ) sys.exit(1) PY ``` With the temporary file and `set -euo pipefail`, a failed `helm template` stops the job before Python runs. The script checks Deployment, StatefulSet, and DaemonSet pod specs; Ingress `tls.secretName` values do not cause false positives. Use kubeconform and admission policy for broader structural checks. ### 10.2 Evidence Checklist | Evidence Item | Acceptable Example | Forbidden Evidence | | --- | --- | --- | | Argo CD state | Screenshot or text showing Synced / Healthy | None | | ExternalSecret state | Ready=True, SecretSynced reason, recent refresh time | Secret value output | | Kubernetes Secret | Object exists and expected key names are present | Decoded data or base64 content | | Mount check | Disposable pod reached Succeeded; log shows only secret-mounted | cat/print of mounted file content | | Reloader rollout | Rollout status and pod creation times after Secret refresh | Secret payload | | Rollback | Revert PR link, approval, commit SHA, app health after sync | Manual kubectl patch not represented in Git | ### 10.3 Common Placeholders | Placeholder | Meaning | Example | | --- | --- | --- | | `` | AWS account ID | `123456789012` | | `` | AWS region for EKS, Secrets Manager, and KMS | `us-east-1` | | `` | Kubernetes namespace for the workload | `api-gateway` | | `` | NovaDeploy service name | `api-gateway` | | `` | EKS cluster name | `nova-prod` | | `` | Argo CD Application name | `api-gateway` | | `` | Argo CD AppProject name | `novadeploy-production` | | `` | Git commit SHA being reverted | `3d1a7f0` | | `` | Argo CD history revision number | `42` | | `` | App-of-Apps root Application name | `novadeploy-production-root` | | `` | Service-scoped Secrets Manager secret suffix | `db` | | `` | Approved password manager or PAM command-line client | `op` | | `` | Item reference within the approved password manager | `op://Platform/nova-api-gateway-db/password` | | `` | Argo CD sync-window ID | `7` | | `` | Workload Kubernetes ServiceAccount | `api-gateway-sa` | | `` | Namespace where ESO runs | `external-secrets` | | `` | ESO controller Deployment name | `external-secrets` | | `` | Namespace where Reloader runs | `reloader` | | `` | Reloader Deployment name | `reloader` | | `` | Expected Reloader ServiceAccount name | `reloader` |