로그인 회원가입
모두인포
전체IT경제사회생활스포츠시사여행연애

FastAPI + PostgreSQL + Redis + Nginx Ingress + Prometheus + Grafana + Loki + Tempo + Argo CD + Argo Rollouts 종합 실전 프로젝트

lmkfox · 2026-09-21 · 조회 0
FastAPI + PostgreSQL + Redis + Nginx Ingress + Prometheus + Grafana + Loki + Tempo + Argo CD + Argo Rollouts 종합 실전 프로젝트
FastAPI + PostgreSQL + Redis + Nginx Ingress + Prometheus + Grafana + Loki + Tempo + Argo CD + Argo Rollouts 종합 실전 프로젝트 1. 들어가며 지금까지 Kubernetes를 공부하면서 다음과 같은 기술을 각각 살펴봤다. - FastAPI - PostgreSQL - Redis - Docker - Kubernetes - Helm - Nginx Ingress - Prometheus - Grafana - Loki - Tempo - GitHub Actions - Argo CD - Argo Rollouts - GitOps - Canary Deployment - 자동 Abort/Rollback 하지만 실제 시스템 엔지니어와 DevOps 엔지니어가 운영하는 환경에서는 이 기술들이 각각 독립적으로 존재하지 않는다. 중요한 것은 각각의 기술을 설치하는 것이 아니라 하나의 운영 흐름으로 연결하는 것이다. 이번 프로젝트에서는 다음과 같은 가상의 웹 서비스를 구축한다. 사용자 | v Internet | v Nginx Ingress | v FastAPI Backend | +----------------+ | | v v PostgreSQL Redis 그리고 운영 시스템에는 다음 구성요소를 추가한다. GitHub | | Git Push v GitHub Actions | | Docker Build v Container Registry | | Image Tag v GitOps Repository | | Helm values 변경 v Argo CD | v Kubernetes Cluster | +------------+------------+ | | v v Argo Rollouts Kubernetes Canary Deployment Resources | v 10% -> 25% -> 50% -> 100% | v Prometheus | +------+------+ | | v v 5xx Error P95 Latency | +------ 실패 ------+ | v Abort / Rollback 운영 관측 FastAPI | +--> Prometheus --> Grafana | +--> Loki --> Grafana | +--> Tempo --> Grafana 최종적으로 목표하는 것은 다음과 같다. 개발자가 Git에 코드를 Push하면 GitHub Actions가 이미지를 만들고, GitOps Repository가 변경되며, Argo CD가 Kubernetes에 배포하고, Argo Rollouts가 Canary 방식으로 트래픽을 증가시키면서 Prometheus가 오류율과 응답시간을 검사한다. 문제가 발생하면 Canary를 자동으로 중단하고 Stable 버전으로 트래픽을 되돌린다. 이것이 이번 종합 프로젝트의 핵심이다. 2. 최종 프로젝트 아키텍처 전체 구조를 먼저 이해해야 한다. Developer | | git push v GitHub Repository | v +-------------------+ | GitHub Actions | +-------------------+ | Docker Build | v Container Registry | image: SHA | v GitOps Repository | Helm values.yaml | v +-----------+ | Argo CD | +-----------+ | GitOps Sync | v +--------------------+ | Kubernetes Cluster | +--------------------+ | Argo Rollouts | +-------------+-------------+ | | Stable Canary | | +-------------+-------------+ | v Nginx Ingress | v Users Application FastAPI | +---------+---------+ | | v v PostgreSQL Redis Observability FastAPI -----> Prometheus -----> Grafana | +---------> Loki -----------> Grafana | +---------> Tempo -----------> Grafana Deployment Verification Argo Rollouts | v Prometheus Analysis | +---- 5xx Error Rate | +---- P95 Latency | +---- Success | +---- Failure | v Abort / Rollback 3. 프로젝트 목표 이번 프로젝트의 최종 목표는 다음과 같다. Application FastAPI 기반 REST API Database PostgreSQL Cache Redis Container Docker Kubernetes Application 배포 및 운영 Helm Kubernetes 설정 패키징 Ingress 외부 HTTP/HTTPS 접근 Monitoring Prometheus Visualization Grafana Logging Loki Tracing Tempo CI GitHub Actions CD Argo CD Progressive Delivery Argo Rollouts Deployment Strategy Canary Deployment Automatic Verification Prometheus Analysis Automatic Failure Handling Abort / Rollback 4. 프로젝트 Repository 구성 실제 운영에서는 Application Repository와 GitOps Repository를 분리하는 것을 권장한다. 예를 들어 다음과 같이 구성한다. fastapi-app/ ├── app/ │ ├── main.py │ ├── database.py │ ├── redis.py │ └── metrics.py │ ├── requirements.txt ├── Dockerfile └── .github/ └── workflows/ └── build.yml GitOps Repository는 별도로 둔다. k8s-gitops/ ├── environments/ │ ├── dev/ │ └── values.yaml │ ├── prod/ │ └── values.yaml │ └── helm/ └── fastapi/ ├── Chart.yaml ├── values.yaml └── templates/ ├── rollout.yaml ├── service.yaml ├── ingress.yaml ├── configmap.yaml ├── secret.yaml ├── postgres.yaml ├── redis.yaml └── analysis-template.yaml 이렇게 분리하면 역할이 명확해진다. Application Repository 개발자가 관리 | +-- FastAPI +-- Dockerfile +-- 테스트 +-- GitHub Actions GitOps Repository 운영자가 관리 | +-- Helm +-- Kubernetes +-- Deployment +-- Ingress +-- Rollout +-- Config 5. FastAPI 애플리케이션 만들기 간단한 API를 만든다. from fastapi import FastAPI from prometheus_fastapi_instrumentator import Instrumentator app = FastAPI() Instrumentator().instrument(app).expose(app) @app.get("/") def root(): return { "message": "Hello Kubernetes" } @app.get("/health") def health(): return { "status": "ok" } 실행하면 다음과 같다. GET / { "message": "Hello Kubernetes" } Health Check: GET /health 결과: { "status": "ok" } Prometheus Metrics: GET /metrics 6. PostgreSQL 연결 실제 서비스에서는 데이터베이스가 필요하다. 예를 들어 SQLAlchemy를 사용할 수 있다. from sqlalchemy import create_engine DATABASE_URL = "postgresql://appuser:password@postgres:5432/appdb" engine = create_engine(DATABASE_URL) Kubernetes에서는 PostgreSQL Service 이름을 사용한다. postgres 따라서 애플리케이션에서는 다음 주소로 접근한다. postgres:5432 Kubernetes에서는 Service Discovery가 제공되므로 IP를 직접 입력하지 않는다. FastAPI | | DNS v postgres | v PostgreSQL Pod 7. Redis 연결 Redis도 마찬가지다. import redis r = redis.Redis( host="redis", port=6379, decode_responses=True ) FastAPI에서 다음 API를 만들 수도 있다. @app.get("/cache/{key}") def get_cache(key: str): value = r.get(key) return { "key": key, "value": value } 8. Dockerfile FastAPI 애플리케이션을 Docker Image로 만든다. FROM python:3.12-slim WORKDIR /app COPY requirements.txt . RUN pip install --no-cache-dir -r requirements.txt COPY app ./app EXPOSE 8000 CMD [ "uvicorn", "app.main:app", "--host", "0.0.0.0", "--port", "8000" ] 빌드: docker build -t fastapi-app:latest . 실행: docker run -p 8000:8000 fastapi-app:latest 확인: curl http://localhost:8000/health 9. Kubernetes Namespace 운영 환경을 분리하기 위해 Namespace를 만든다. apiVersion: v1 kind: Namespace metadata: name: app 적용: kubectl apply -f namespace.yaml 10. PostgreSQL Kubernetes 구성 개발용 프로젝트에서는 PostgreSQL을 Kubernetes에 함께 구성할 수 있다. apiVersion: apps/v1 kind: Deployment metadata: name: postgres namespace: app spec: replicas: 1 selector: matchLabels: app: postgres template: metadata: labels: app: postgres spec: containers: - name: postgres image: postgres:17 env: - name: POSTGRES_DB value: appdb - name: POSTGRES_USER value: appuser - name: POSTGRES_PASSWORD valueFrom: secretKeyRef: name: postgres-secret key: password ports: - containerPort: 5432 Service: apiVersion: v1 kind: Service metadata: name: postgres namespace: app spec: selector: app: postgres ports: - port: 5432 targetPort: 5432 실제 Production에서는 PostgreSQL을 Kubernetes에 직접 구성하기보다 Managed Database나 별도의 PostgreSQL HA 환경을 사용하는 경우가 많다. 11. Redis Kubernetes 구성 apiVersion: apps/v1 kind: Deployment metadata: name: redis namespace: app spec: replicas: 1 selector: matchLabels: app: redis template: metadata: labels: app: redis spec: containers: - name: redis image: redis:7 ports: - containerPort: 6379 Service: apiVersion: v1 kind: Service metadata: name: redis namespace: app spec: selector: app: redis ports: - port: 6379 targetPort: 6379 12. Helm Chart 구성 이제 Kubernetes YAML을 Helm Chart로 관리한다. helm/ └── fastapi/ ├── Chart.yaml ├── values.yaml │ └── templates/ ├── rollout.yaml ├── service.yaml ├── ingress.yaml ├── configmap.yaml ├── secret.yaml ├── postgres.yaml ├── redis.yaml └── analysis-template.yaml Chart.yaml: apiVersion: v2 name: fastapi description: FastAPI Application type: application version: 1.0.0 appVersion: "1.0" 13. Helm values.yaml 가장 중요한 파일 중 하나다. image: repository: ghcr.io/example/fastapi tag: "main-abc123" replicaCount: 4 service: port: 8000 ingress: enabled: true host: api.example.com resources: requests: cpu: 100m memory: 256Mi limits: cpu: 500m memory: 512Mi 운영 환경에서 이미지 Tag를 다음과 같이 사용하는 것을 권장한다. main-abc123 또는: sha-abc123 다음과 같은 방식은 피하는 것이 좋다. latest 왜냐하면 latest는 어떤 버전인지 명확하지 않기 때문이다. 14. Argo Rollouts 일반 Kubernetes Deployment 대신 Argo Rollout을 사용한다. apiVersion: argoproj.io/v1alpha1 kind: Rollout metadata: name: fastapi namespace: app spec: replicas: 4 selector: matchLabels: app: fastapi template: metadata: labels: app: fastapi spec: containers: - name: fastapi image: "{{ .Values.image.repository }}:{{ .Values.image.tag }}" ports: - containerPort: 8000 readinessProbe: httpGet: path: /health port: 8000 initialDelaySeconds: 5 periodSeconds: 5 livenessProbe: httpGet: path: /health port: 8000 initialDelaySeconds: 10 periodSeconds: 10 resources: {{- toYaml .Values.resources | nindent 12 }} strategy: canary: canaryService: fastapi-canary stableService: fastapi-stable steps: - setWeight: 10 - pause: duration: 5m - setWeight: 25 - pause: duration: 5m - setWeight: 50 - pause: duration: 10m - setWeight: 100 핵심은 이것이다. 10% | v 25% | v 50% | v 100% 15. Stable / Canary Service Argo Rollouts에서는 Stable과 Canary를 구분할 수 있다. apiVersion: v1 kind: Service metadata: name: fastapi-stable namespace: app spec: selector: app: fastapi ports: - port: 8000 targetPort: 8000 Canary Service: apiVersion: v1 kind: Service metadata: name: fastapi-canary namespace: app spec: selector: app: fastapi ports: - port: 8000 targetPort: 8000 하지만 여기서 중요한 점이 있다. Service 두 개만 만들어 놓는다고 HTTP 트래픽이 자동으로 10% / 90%로 분할되는 것은 아니다. 실제 트래픽 비율을 제어하려면 Nginx Ingress와 Argo Rollouts의 Traffic Routing 연동이 필요하다. 16. Nginx Ingress 외부 사용자는 Kubernetes 내부 Service에 직접 접근하지 않는다. Internet | v Nginx Ingress | v FastAPI Ingress 예시는 다음과 같다. apiVersion: networking.k8s.io/v1 kind: Ingress metadata: name: fastapi namespace: app spec: ingressClassName: nginx rules: - host: api.example.com http: paths: - path: / pathType: Prefix backend: service: name: fastapi-stable port: number: 8000 Production에서는 HTTPS를 적용한다. Client | HTTPS | v Nginx Ingress | HTTP | v FastAPI 17. Prometheus Metrics FastAPI에 Prometheus Instrumentation을 적용한다. from prometheus_fastapi_instrumentator import Instrumentator Instrumentator().instrument(app).expose(app) 그러면 다음과 같은 Metrics를 얻을 수 있다. http_requests_total http_request_duration_seconds 예를 들어: http_requests_total{ method="GET", status="200" } 또는: http_request_duration_seconds_bucket 이 데이터를 Prometheus가 수집한다. 18. Prometheus 구조 FastAPI | | /metrics v Prometheus | +----------------+ | | v v Grafana Argo Rollouts Grafana는 사람이 보는 용도이고 Prometheus는 시스템이 판단하는 용도로 사용할 수 있다. Grafana | +-- 운영자 확인 Prometheus | +-- 자동 판단 19. Grafana Dashboard Grafana에서는 다음과 같은 Dashboard를 구성한다. +------------------------------------------------+ | FastAPI Production Dashboard | +------------------------------------------------+ | | | Request Rate 125 req/s | | | +------------------------------------------------+ | | | HTTP 5xx Error Rate 0.12 % | | | +------------------------------------------------+ | | | P95 Latency 180 ms | | | +------------------------------------------------+ | | | CPU Usage 35 % | | Memory Usage 48 % | | | +------------------------------------------------+ | | | Pod Status 4 / 4 | | | +------------------------------------------------+ 운영자가 장애를 확인할 때 가장 먼저 보는 화면이 된다. 20. Loki 로그 수집 Metrics만으로는 장애 원인을 알기 어려운 경우가 있다. 예를 들어: HTTP 500 증가 라는 사실은 Prometheus로 알 수 있다. 하지만 왜 500이 발생했는지는 로그를 확인해야 한다. Prometheus | | 500 증가 v Grafana | v Loki | v Application Log 예: ERROR database connection failed ERROR redis timeout ERROR unexpected exception Loki에서는 LogQL을 사용할 수 있다. {namespace="app", app="fastapi"} |= "ERROR" 또는: {namespace="app", app="fastapi"} |= "500" 21. Tempo Distributed Tracing 로그만으로도 부족할 수 있다. 예를 들어 하나의 요청이 다음과 같이 이동한다고 가정한다. Client | v Ingress | v FastAPI | +----> Redis | +----> PostgreSQL 응답시간이 5초라고 하자. 문제는 어디일까? Ingress 20ms FastAPI 100ms Redis 30ms PostgreSQL 4.8sec 이런 문제는 Trace를 보면 빠르게 찾을 수 있다. Tempo를 사용하면 하나의 Request가 어느 구간에서 오래 걸렸는지 확인할 수 있다. Trace | +-- FastAPI 5.0s | +-- Redis 20ms | +-- PostgreSQL 4.8s 따라서 다음 세 가지를 함께 사용하는 것이 중요하다. Metrics = 문제가 발생했는가? Logs = 무슨 오류가 발생했는가? Traces = 어디에서 문제가 발생했는가? 22. Grafana에서 Metrics + Logs + Traces 연결 운영 환경에서는 세 시스템을 따로 보는 것보다 Grafana에서 연결하는 것이 좋다. Grafana | +------------+------------+ | | | v v v Prometheus Loki Tempo | | | Metrics Logs Traces 예를 들어 Dashboard에서 5xx 오류를 발견한다. HTTP 5xx | v 해당 시간대 Log | v Trace ID | v Database Query 이런 방식으로 장애 원인을 좁혀갈 수 있다. 23. Argo CD 설치 이제 GitOps를 적용한다. Argo CD의 역할은 간단하다. Git Repository | v Argo CD | v Kubernetes Git에 정의된 상태와 실제 Kubernetes 상태를 비교한다. Desired State Git | +-- image: sha-123456 Actual State Kubernetes | +-- image: sha-123456 같으면: Synced 다르면: OutOfSync 24. Argo CD Application 예를 들어: apiVersion: argoproj.io/v1alpha1 kind: Application metadata: name: fastapi namespace: argocd spec: project: default source: repoURL: https://github.com/example/k8s-gitops.git targetRevision: main path: helm/fastapi destination: server: https://kubernetes.default.svc namespace: app syncPolicy: automated: prune: true selfHeal: true 이제 Git의 변경사항을 Argo CD가 자동으로 Kubernetes에 반영한다. 25. GitHub Actions 개발자가 Application Repository에 코드를 Push한다. git add . git commit -m "update api" git push origin main GitHub Actions가 실행된다. Git Push | v GitHub Actions | +-- Test | +-- Docker Build | +-- Docker Push | +-- GitOps values update 26. GitHub Actions 예제 name: Build and Deploy on: push: branches: - main jobs: build: runs-on: ubuntu-latest steps: - name: Checkout uses: actions/checkout@v4 - name: Login Registry uses: docker/login-action@v3 with: registry: ghcr.io username: ${{ github.actor }} password: ${{ secrets.GITHUB_TOKEN }} - name: Build Image run: | docker build \ -t ghcr.io/example/fastapi:${{ github.sha }} . - name: Push Image run: | docker push \ ghcr.io/example/fastapi:${{ github.sha }} 여기까지는 CI다. 27. GitOps Repository 자동 변경 다음 단계에서는 GitHub Actions가 GitOps Repository의 Helm values를 변경한다. 기존: image: repository: ghcr.io/example/fastapi tag: "abc123" 변경: image: repository: ghcr.io/example/fastapi tag: "def456" 그리고 GitOps Repository에 Commit한다. Application Repository | v GitHub Actions | v Docker Image | v GitOps Repository | v Helm values.yaml 28. Argo CD 자동 배포 GitOps Repository에 변경사항이 발생하면 Argo CD가 감지한다. GitOps Repository | | change v Argo CD | | Sync v Kubernetes Argo CD는 새로운 이미지가 지정된 Helm Chart를 적용한다. image: tag: "def456" 그러면 Argo Rollouts가 새로운 ReplicaSet을 생성한다. 29. Canary Deployment 시작 새로운 버전이 배포되었다고 가정한다. Stable v1.0 새 버전: Canary v1.1 처음에는: Stable 90% Canary 10% 다음: Stable 75% Canary 25% 다음: Stable 50% Canary 50% 최종: Stable 0% Canary 100% 구조: Ingress | +------+------+ | | v v Stable Canary 90% 10% | | v v v1.0 v1.1 실제 HTTP 트래픽 비율을 제어하려면 Nginx Ingress와 Argo Rollouts의 Traffic Routing 기능을 구성해야 한다. 30. Canary에서 가장 중요한 것 단순히 Canary를 배포하는 것만으로는 부족하다. 다음 질문에 답해야 한다. 새 버전이 정상인가? 따라서 자동 검증을 추가한다. 검증 기준: HTTP 5xx Error Rate P95 Latency 예를 들어: 5xx < 1% P95 < 500ms 이면 성공이다. 31. AnalysisTemplate Argo Rollouts가 Prometheus에 질의할 수 있도록 AnalysisTemplate을 만든다. apiVersion: argoproj.io/v1alpha1 kind: AnalysisTemplate metadata: name: fastapi-analysis namespace: app spec: metrics: - name: error-rate interval: 1m count: 5 successCondition: result[0] < 0.01 failureLimit: 1 provider: prometheus: address: http://prometheus.monitoring:9090 query: | sum( rate( http_requests_total{ app="fastapi", status=~"5.." }[5m] ) ) / sum( rate( http_requests_total{ app="fastapi" }[5m] ) ) - name: p95-latency interval: 1m count: 5 successCondition: result[0] < 0.5 failureLimit: 1 provider: prometheus: address: http://prometheus.monitoring:9090 query: | histogram_quantile( 0.95, sum( rate( http_request_duration_seconds_bucket{ app="fastapi" }[5m] ) ) by (le) ) 32. Canary 단계에 Analysis 연결 실제 운영에서는 단순히 마지막에 한 번 검사하는 것보다 각 단계에서 검증하는 것이 중요하다. 개념적으로: 10% | +--> Analysis | | | +-- PASS | v 25% | +--> Analysis | | | +-- PASS | v 50% | +--> Analysis | | | +-- PASS | v 100% Rollout 구성 예: strategy: canary: steps: - setWeight: 10 - pause: duration: 5m - analysis: templates: - templateName: fastapi-analysis - setWeight: 25 - pause: duration: 5m - analysis: templates: - templateName: fastapi-analysis - setWeight: 50 - pause: duration: 10m - analysis: templates: - templateName: fastapi-analysis - setWeight: 100 33. 정상적인 배포 정상적인 경우 다음과 같이 진행된다. Git Push | v GitHub Actions | v Docker Build | v Registry | v GitOps Update | v Argo CD | v Argo Rollouts | v 10% | | Prometheus | 5xx = 0.2% | P95 = 180ms | v 25% | | Prometheus | 5xx = 0.3% | P95 = 190ms | v 50% | | Prometheus | 5xx = 0.4% | P95 = 210ms | v 100% | v Deployment 성공 34. 장애가 발생한 경우 이번에는 새로운 버전에 버그가 있다고 가정한다. Canary 10%에서: HTTP 5xx = 8% P95 = 2.1s 정상 기준: 5xx < 1% P95 < 500ms 따라서 Analysis가 실패한다. Canary 10% | v Prometheus | +---- 5xx = 8% | +---- P95 = 2.1s | v Analysis Failed | v Argo Rollouts | v Abort 35. Abort 이후 Canary를 계속 진행하지 않는다. Before Stable 90% Canary 10% Abort: Canary | X | v Abort 그리고 Traffic은 Stable 쪽으로 돌아간다. Stable v1.0 | +---- 100% 사용자는 문제가 있는 Canary 버전에 계속 노출되지 않는다. 36. Abort와 Git Rollback은 구분해야 한다 여기서 매우 중요한 개념이 있다. Argo Rollouts의 Abort와 Git Rollback은 같은 의미가 아니다. Runtime Abort 현재 Canary 배포를 중단한다. v1.1 Canary | v Abort | v v1.0 Stable Git Rollback GitOps Repository의 이미지 버전을 이전 버전으로 되돌린다. Git v1.0 | v v1.1 Rollback: v1.1 | v v1.0 즉: Runtime Recovery != Git Repository Rollback 운영 환경에서는 두 개를 명확하게 구분해야 한다. 37. Loki를 이용한 장애 분석 Canary가 실패했다면 다음 단계는 로그 분석이다. Grafana에서 Loki를 조회한다. {namespace="app", app="fastapi"} ERROR만 찾는다. {namespace="app", app="fastapi"} |= "ERROR" 예: 2026-09-21 10:10:01 ERROR PostgreSQL connection timeout 2026-09-21 10:10:03 ERROR database query failed 2026-09-21 10:10:05 ERROR HTTP 500 이제 원인을 추적할 수 있다. 38. Tempo를 이용한 Trace 분석 다음으로 Tempo에서 Trace를 확인한다. Request | v FastAPI | +---- Redis | +---- PostgreSQL Trace: FastAPI 2.3s | +-- Redis 20ms | +-- PostgreSQL 2.2s 따라서 문제의 원인이 PostgreSQL 쿼리 지연이라는 것을 확인할 수 있다. 39. 장애 대응 전체 흐름 이제 전체 운영 프로세스를 하나로 연결해 보자. 장애 발생 | v Prometheus | 5xx 증가 | v Alertmanager | Slack / Email | v Grafana | +---------+---------+ | | v v Loki Tempo | | Logs Trace | | +---------+---------+ | v 원인 분석 | v Kubernetes 상태 | v Argo Rollouts | v Abort / Rollback | v Stable 복구 40. 전체 CI/CD 흐름 이제 개발부터 운영까지 연결한다. Developer | | git push v GitHub | v GitHub Actions | +---- Unit Test | +---- Docker Build | +---- Security Scan | +---- Push Image | v Container Registry | v GitOps Repository | +---- Helm values 변경 | v Argo CD | +---- Sync | v Argo Rollouts | +---- Canary 10% | +---- Prometheus Analysis | +---- Canary 25% | +---- Prometheus Analysis | +---- Canary 50% | +---- Prometheus Analysis | +---- Canary 100% | v Production 41. 전체 Observability 흐름 Application의 상태는 세 가지 관점으로 확인한다. Application | +--------------+--------------+ | | | v v v Metrics Logs Traces | | | v v v Prometheus Loki Tempo | | | +--------------+--------------+ | v Grafana 각각의 역할은 다음과 같다. 시스템 역할 Prometheus Metrics Grafana Visualization Loki Logs Tempo Traces Alertmanager Alert Argo Rollouts Progressive Delivery Argo CD GitOps Deployment 42. Kubernetes Self-Healing 애플리케이션 장애와 배포 장애는 구분해야 한다. 예를 들어 Pod 하나가 죽으면 Kubernetes가 다시 생성할 수 있다. Pod | X | v Kubernetes | v New Pod Replica: Desired = 4 Current = 3 Kubernetes: Current = 4 이것이 Kubernetes의 기본 Self-Healing이다. 43. Readiness Probe 애플리케이션이 실행 중이어도 실제 서비스가 불가능할 수 있다. 예: FastAPI Process = Running PostgreSQL = Down 이 경우 단순히 Process가 살아 있다는 것만으로는 정상이라고 판단하면 안 된다. Readiness Probe를 사용한다. readinessProbe: httpGet: path: /health port: 8000 periodSeconds: 5 Readiness가 실패하면 Kubernetes Service의 트래픽 대상에서 제외할 수 있다. 44. Liveness Probe 프로세스 자체가 비정상 상태에 빠졌다면 Liveness Probe가 사용할 수 있다. livenessProbe: httpGet: path: /health port: 8000 initialDelaySeconds: 10 periodSeconds: 10 실패가 지속되면 Kubernetes가 Container를 재시작한다. Application | X | Liveness Failed | v Container Restart 45. Resource Requests / Limits 운영 환경에서는 CPU와 Memory를 반드시 관리해야 한다. resources: requests: cpu: 100m memory: 256Mi limits: cpu: 500m memory: 512Mi 의미: requests | +-- 스케줄링 기준 limits | +-- 최대 사용량 제한 46. HPA 트래픽이 증가하면 Pod 수를 자동으로 증가시킬 수 있다. Traffic | v CPU 80% | v HPA | v Pod 증가 예: 2 Pods | v 4 Pods | v 8 Pods 하지만 Canary Deployment와 HPA를 함께 사용할 경우 Replica 수 변화가 Canary 검증에 영향을 줄 수 있으므로 설계를 신중하게 해야 한다. 47. Database Migration의 중요한 문제 자동 배포 환경에서 가장 조심해야 하는 부분 중 하나가 Database Migration이다. 예를 들어 v1.1에서 다음 컬럼을 삭제했다고 하자. ALTER TABLE users DROP COLUMN old_column; 그런데 Rollback으로 v1.0으로 돌아가면 v1.0 애플리케이션이 old_column을 필요로 할 수 있다. 그러면: Application Rollback | v v1.0 Application | X | DB Schema는 v1.1 따라서 Database는 일반적인 Application Rollback과 다르게 관리해야 한다. 48. Expand / Contract Migration 운영 환경에서는 Expand / Contract 패턴을 사용할 수 있다. Step 1 새로운 컬럼 추가 old_column new_column Step 2 새 애플리케이션이 new_column 사용 v1.1 | +-- new_column Step 3 충분히 검증 Step 4 이전 버전 의존성 제거 Step 5 기존 컬럼 삭제 old_column 삭제 이렇게 해야 Application Rollback과 DB Schema 문제가 충돌하는 것을 줄일 수 있다. 49. Security Production 환경에서는 다음 보안 요소도 필요하다. GitHub | +-- Secret 관리 | +-- OIDC | +-- Token 최소 권한 Kubernetes | +-- RBAC | +-- NetworkPolicy | +-- Secret | +-- Pod Security Container | +-- Non-root | +-- Image Scan | +-- Minimal Image 특히 Kubernetes Secret을 Git Repository에 평문으로 저장하면 안 된다. 50. CI Pipeline에 Security Scan 추가 GitHub Actions에서 이미지 취약점 검사를 추가할 수 있다. Git Push | v Test | v Build | v Security Scan | +---- FAIL | v Push Image 취약점이 기준 이상이면 Production 배포를 중단할 수 있다. 51. GitOps의 핵심 원칙 GitOps에서는 Git이 원하는 상태의 기준점이 된다. Git | | Desired State v Argo CD | v Kubernetes | | Actual State v Compare 예를 들어 Git에는: replicas: 4 실제 Kubernetes: replicas: 3 Argo CD가 차이를 발견한다. OutOfSync 그리고 Self Heal을 사용하면 원하는 상태로 되돌릴 수 있다. 52. 전체 프로젝트에서 각 기술의 역할 이제 모든 기술의 역할을 정리해 보자. 기술 역할 FastAPI Application PostgreSQL Persistent Database Redis Cache Docker Container Kubernetes Container Orchestration Helm Kubernetes Packaging Nginx Ingress HTTP/HTTPS Routing Prometheus Metrics Grafana Dashboard Loki Log Tempo Distributed Trace Alertmanager Alert GitHub Actions CI Argo CD GitOps CD Argo Rollouts Canary / Progressive Delivery 이것이 이번 프로젝트의 전체 기술 스택이다. 53. 실제 장애 시나리오 실제 운영을 가정해 보자. 개발자가 v1.2를 배포했다. v1.1 -> v1.2 GitHub Actions: BUILD PASS TEST PASS IMAGE PUSH PASS GitOps: image.tag = v1.2 Argo CD: SYNC Argo Rollouts: Canary 10% 그런데 새로운 코드에 문제가 있다. PostgreSQL Query Slow Prometheus: P95 = 1.8s 기준: P95 < 500ms Analysis: FAILED Argo Rollouts: ABORT 결과: v1.1 = Stable v1.2 = Canary v1.2 | X | v Abort Traffic | v v1.1 운영자는 Grafana에서 확인한다. Prometheus | +-- P95 증가 Loki: database query timeout Tempo: PostgreSQL Span = 1.7s 최종적으로 DB Query 문제가 원인임을 확인한다. 54. 이 프로젝트의 핵심 운영 패턴 전체 시스템을 하나의 문장으로 표현하면 다음과 같다. Code ↓ GitHub ↓ GitHub Actions ↓ Docker Image ↓ GitOps ↓ Argo CD ↓ Argo Rollouts ↓ Canary ↓ Nginx Ingress ↓ FastAPI ↓ PostgreSQL / Redis ↓ Prometheus / Loki / Tempo ↓ Grafana ↓ Analysis ↓ Success or Failure ↓ Abort / Recovery 이것이 현대적인 Kubernetes 기반 DevOps/SRE 운영 구조의 하나의 대표적인 형태다. 55. 최종 프로젝트 디렉터리 전체 프로젝트는 다음과 같이 구성할 수 있다. project/ │ ├── application/ │ │ ├── app/ │ │ ├── main.py │ │ ├── database.py │ │ ├── redis.py │ │ └── metrics.py │ │ │ ├── requirements.txt │ ├── Dockerfile │ │ │ └── .github/ │ └── workflows/ │ └── build.yml │ │ └── gitops/ │ ├── environments/ │ │ │ ├── dev/ │ │ └── values.yaml │ │ │ └── prod/ │ └── values.yaml │ └── helm/ │ └── fastapi/ │ ├── Chart.yaml ├── values.yaml │ └── templates/ │ ├── rollout.yaml ├── service.yaml ├── ingress.yaml ├── configmap.yaml ├── secret.yaml ├── postgres.yaml ├── redis.yaml └── analysis-template.yaml 56. 전체 배포 순서 실제 실습은 다음 순서로 진행하는 것이 좋다. 1단계 Kubernetes Cluster 구축 Kubernetes 2단계 Nginx Ingress 설치 Nginx Ingress Controller 3단계 Prometheus + Grafana 설치 Monitoring 4단계 Loki 설치 Logging 5단계 Tempo 설치 Tracing 6단계 Argo CD 설치 GitOps 7단계 Argo Rollouts 설치 Progressive Delivery 8단계 FastAPI Application 구축 FastAPI 9단계 PostgreSQL 구축 Database 10단계 Redis 구축 Cache 11단계 Helm Chart 작성 Helm 12단계 Nginx Ingress 연결 External HTTP 13단계 Prometheus Metrics 연결 Metrics 14단계 Loki Logging 연결 Logs 15단계 Tempo Tracing 연결 Traces 16단계 GitHub Actions 구성 CI 17단계 GitOps Repository 구성 GitOps 18단계 Argo CD 연결 CD 19단계 Argo Rollouts Canary 구성 10% 25% 50% 100% 20단계 Prometheus Analysis 구성 5xx P95 21단계 자동 Abort 테스트 Failure ↓ Abort 22단계 Loki / Tempo 장애 분석 Logs + Traces 23단계 Grafana Dashboard 구성 Observability 24단계 전체 Pipeline 테스트 GitHub ↓ Actions ↓ Registry ↓ GitOps ↓ Argo CD ↓ Rollouts ↓ Canary ↓ Prometheus ↓ Abort / Promote 57. 최종적으로 구현되는 운영 환경 모든 구성이 완료되면 다음과 같은 환경이 만들어진다. Developer | Git Push | v +----------------+ | GitHub Actions | +----------------+ | Test / Build / Scan | v Container Registry | v GitOps Repository | v +---------+ | Argo CD | +---------+ | v +----------------------+ | Kubernetes Cluster | | | | +----------------+ | | | Nginx Ingress | | | +-------+--------+ | | | | | v | | +----------------+ | | | Argo Rollouts | | | +-------+--------+ | | | | | +---+---+ | | | | | | v v | | Stable Canary | | | | | | +---+---+ | | | | | v | | FastAPI | | / \ | | / \ | | v v | | PostgreSQL Redis | +----------------------+ Observability | +---------------+---------------+ | | | v v v Prometheus Loki Tempo | | | +---------------+---------------+ | v Grafana | v Operator / SRE 58. 이 프로젝트에서 배우는 것 이 하나의 프로젝트를 완성하면 단순히 Kubernetes 명령어를 아는 수준을 넘어 다음의 전체 흐름을 이해할 수 있다. Application FastAPI PostgreSQL Redis Container Docker Orchestration Kubernetes Packaging Helm Networking Nginx Ingress CI GitHub Actions GitOps Argo CD Progressive Delivery Argo Rollouts Monitoring Prometheus Grafana Logging Loki Tracing Tempo Reliability Canary Health Check Automatic Analysis Abort Rollback 즉, Application 개발부터 Kubernetes 배포, GitOps, Observability, Canary Deployment, 자동 장애 대응까지 하나의 Production 운영 흐름을 직접 구현하는 프로젝트 가 된다. 59. 마무리 지금까지 Kubernetes를 공부하면서 각각의 기술을 따로 보면 상당히 복잡해 보인다. 하지만 전체 흐름을 하나로 연결하면 역할이 명확해진다. GitHub Actions | | Build v Docker Image | v GitOps | v Argo CD | v Argo Rollouts | v Canary Deployment | v Nginx Ingress | v FastAPI | +------ PostgreSQL | +------ Redis | +------ Prometheus | +------ Loki | +------ Tempo | v Grafana 그리고 장애가 발생하면: Prometheus | v Analysis Failure | v Argo Rollouts | v Abort | v Stable Version 이 구조의 핵심은 단순히 **"자동 배포"**가 아니다. 중요한 것은 다음 네 가지가 연결되어 있다는 점이다. 1. 빠르게 배포한다. ↓ 2. 조금씩 트래픽을 증가시킨다. ↓ 3. Metrics / Logs / Traces로 상태를 관찰한다. ↓ 4. 문제가 발생하면 자동으로 배포를 중단한다. 이것이 CI/CD와 GitOps, Observability, Progressive Delivery를 하나의 운영 체계로 연결하는 핵심이다.

댓글

아직 댓글이 없습니다.

로그인 후 댓글을 남길 수 있습니다.