IMG-LOGO
공지사항 :

FastAPI + PostgreSQL + Redis 애플리케이션의 Argo Rollouts Canary 자동 배포 실습

lmkfox - 2026-09-18 07:04:43 5 Views 0 Comment

FastAPI + PostgreSQL + Redis 애플리케이션의 Argo Rollouts Canary 자동 배포 실습

1. 들어가며

지금까지 Kubernetes를 다음과 같이 단계적으로 발전시켰다.

Kubernetes
   ↓
Helm
   ↓
Argo CD
   ↓
GitOps
   ↓
Prometheus
   ↓
Grafana
   ↓
Loki
   ↓
Tempo
   ↓
Alertmanager
   ↓
자동 복구

이번에는 여기서 한 단계 더 나아가 실제 애플리케이션 배포 과정 전체를 연결한다.

실습 대상은 다음과 같다.

FastAPI
PostgreSQL
Redis

배포 자동화 흐름은 다음과 같이 만든다.

Developer
    ↓
GitHub Push
    ↓
GitHub Actions
    ↓
Docker Image Build
    ↓
Container Registry
    ↓
GitOps Repository
    ↓
Helm values.yaml 변경
    ↓
Argo CD
    ↓
Argo Rollouts
    ↓
Canary 10%
    ↓
Prometheus 검증
    ↓
Canary 25%
    ↓
Prometheus 검증
    ↓
Canary 50%
    ↓
Prometheus 검증
    ↓
100%

그리고 새로운 버전에서 문제가 발생하면:

Prometheus
    ↓
HTTP 5xx 증가
또는
P95 Latency 증가
    ↓
Analysis 실패
    ↓
Argo Rollouts Abort
    ↓
이전 Stable Version 유지
    ↓
Rollback

즉 이번 실습의 최종 목표는 다음과 같다.

코드 변경부터 Canary 배포, Prometheus 기반 자동 검증, 실패 시 자동 중단 및 복구까지 하나의 GitOps Pipeline으로 연결하는 것


2. 최종 아키텍처

전체 구조부터 이해하는 것이 중요하다.

                         Developer
                             │
                             │ git push
                             ▼
                    ┌─────────────────┐
                    │ GitHub          │
                    │ Application Repo│
                    └────────┬────────┘
                             │
                             ▼
                    ┌─────────────────┐
                    │ GitHub Actions  │
                    │                 │
                    │ Test            │
                    │ Docker Build    │
                    │ Push Image      │
                    └────────┬────────┘
                             │
                             ▼
                    ┌─────────────────┐
                    │ Container       │
                    │ Registry        │
                    └────────┬────────┘
                             │
                             │ image tag
                             ▼
                    ┌─────────────────┐
                    │ GitOps Repo     │
                    │                 │
                    │ Helm values     │
                    └────────┬────────┘
                             │
                             ▼
                        ┌─────────┐
                        │ Argo CD │
                        └────┬────┘
                             │
                             ▼
                    ┌────────────────┐
                    │ Argo Rollouts  │
                    └───────┬────────┘
                            │
                ┌───────────┴───────────┐
                ▼                       ▼
             Stable                  Canary
             v1.0                     v1.1
                │                       │
                └───────────┬───────────┘
                            │
                            ▼
                       Kubernetes
                            │
              ┌─────────────┼─────────────┐
              ▼             ▼             ▼
           FastAPI       PostgreSQL      Redis
              │
              ▼
         Prometheus
              │
       ┌──────┴──────┐
       ▼             ▼
   HTTP 5xx       P95 Latency
       │             │
       └──────┬──────┘
              ▼
        Argo Analysis
              │
        ┌─────┴─────┐
        ▼           ▼
      PASS         FAIL
        │           │
        ▼           ▼
    Next Step     Abort
                    │
                    ▼
                 Rollback

3. 실습 환경

예제 환경은 다음과 같이 구성한다.

Kubernetes Cluster

Namespace:

tax-prod

Application:

backend

구성:

FastAPI
PostgreSQL
Redis

Monitoring:

Prometheus
Grafana

GitOps:

Argo CD

Progressive Delivery:

Argo Rollouts

CI:

GitHub Actions

Registry:

Container Registry

예제에서는 Registry 주소를 다음처럼 사용한다.

registry.example.com

실제 환경에서는 Docker Hub, GHCR 또는 사내 Registry 주소로 변경하면 된다.


4. Repository를 두 개로 분리한다

GitOps 구조에서는 Application Repository와 GitOps Repository를 분리하는 방식이 관리하기 좋다.

Application Repository

tax-backend
├── app
│   └── main.py
│
├── requirements.txt
├── Dockerfile
└── .github
    └── workflows
        └── build.yml

여기에는 애플리케이션 소스가 들어간다.


GitOps Repository

tax-gitops
├── charts
│   └── tax-app
│       ├── Chart.yaml
│       ├── values.yaml
│       └── templates
│           ├── rollout.yaml
│           ├── service.yaml
│           ├── postgres.yaml
│           ├── redis.yaml
│           └── analysis.yaml
│
└── argocd
    └── tax-app.yaml

여기에는 Kubernetes의 Desired State가 들어간다.


5. FastAPI 애플리케이션 만들기

간단한 FastAPI 애플리케이션을 만든다.

from fastapi import FastAPI
from prometheus_client import Counter, Histogram, generate_latest
from starlette.responses import Response
import time

app = FastAPI()

REQUEST_COUNT = Counter(
    "http_requests_total",
    "Total HTTP requests",
    ["method", "path", "status"]
)

REQUEST_LATENCY = Histogram(
    "http_request_duration_seconds",
    "HTTP request latency",
    ["method", "path"]
)


@app.get("/health")
def health():
    return {"status": "ok"}


@app.get("/api/expenses")
def expenses():

    start = time.time()

    result = {
        "items": [
            {
                "id": 1,
                "amount": 10000
            }
        ]
    }

    duration = time.time() - start

    REQUEST_LATENCY.labels(
        "GET",
        "/api/expenses"
    ).observe(duration)

    REQUEST_COUNT.labels(
        "GET",
        "/api/expenses",
        "200"
    ).inc()

    return result


@app.get("/metrics")
def metrics():

    return Response(
        generate_latest(),
        media_type="text/plain"
    )

실제 애플리케이션에서는 Middleware를 이용해 모든 HTTP 요청을 자동으로 측정하는 방법이 더 적합하다.

이번 예제에서는 구조를 단순하게 보여주기 위해 직접 Metrics를 생성한다.


6. PostgreSQL 연결

FastAPI에서 PostgreSQL을 사용할 수 있도록 한다.

예를 들어:

import os
import psycopg

DATABASE_URL = os.getenv(
    "DATABASE_URL"
)

def get_db():

    conn = psycopg.connect(
        DATABASE_URL
    )

    return conn

환경변수:

DATABASE_URL

예:

postgresql://taxuser:password@postgres:5432/taxdb

운영 환경에서는 Password를 YAML에 평문으로 저장하지 않는다.


7. Redis 연결

Redis 역시 Service 이름을 통해 접근한다.

import redis
import os

REDIS_HOST = os.getenv(
    "REDIS_HOST",
    "redis"
)

r = redis.Redis(
    host=REDIS_HOST,
    port=6379,
    decode_responses=True
)

Kubernetes에서는:

redis

가 Redis Service DNS가 된다.

즉:

FastAPI
   │
   │ redis:6379
   ▼
Redis Service
   │
   ▼
Redis Pod

8. Dockerfile

FastAPI Container를 만든다.

FROM python:3.12-slim

WORKDIR /app

COPY requirements.txt .

RUN pip install --no-cache-dir \
    -r requirements.txt

COPY app ./app

EXPOSE 8000

CMD [
  "uvicorn",
  "app.main:app",
  "--host",
  "0.0.0.0",
  "--port",
  "8000"
]

9. requirements.txt

fastapi
uvicorn[standard]
psycopg[binary]
redis
prometheus-client

10. Helm Chart 구조

GitOps Repository에서는 Helm Chart를 사용한다.

charts/tax-app
│
├── Chart.yaml
├── values.yaml
│
└── templates
    ├── rollout.yaml
    ├── service.yaml
    ├── postgres.yaml
    ├── redis.yaml
    ├── analysis.yaml
    └── secret.yaml

11. Chart.yaml

apiVersion: v2

name: tax-app

description: Tax application

type: application

version: 1.0.0

appVersion: "1.0.0"

12. values.yaml

기본 설정을 만든다.

replicaCount: 4

image:
  repository: registry.example.com/tax-backend
  tag: "1.0.0"
  pullPolicy: IfNotPresent

service:
  port: 80
  targetPort: 8000

resources:

  requests:
    cpu: 250m
    memory: 256Mi

  limits:
    cpu: 1000m
    memory: 512Mi

13. Rollout YAML

Deployment 대신 Rollout을 사용한다.

apiVersion: argoproj.io/v1alpha1
kind: Rollout

metadata:
  name: backend

spec:

  replicas: {{ .Values.replicaCount }}

  selector:

    matchLabels:
      app: backend

  template:

    metadata:
      labels:
        app: backend

    spec:

      containers:

        - name: backend

          image: "{{ .Values.image.repository }}:{{ .Values.image.tag }}"

          imagePullPolicy: {{ .Values.image.pullPolicy }}

          ports:

            - containerPort: 8000

          env:

            - name: DATABASE_URL

              valueFrom:
                secretKeyRef:
                  name: backend-secret
                  key: database-url

            - name: REDIS_HOST

              value: redis

          resources:
{{ toYaml .Values.resources | indent 12 }}

          readinessProbe:

            httpGet:
              path: /health
              port: 8000

            initialDelaySeconds: 5
            periodSeconds: 5

          livenessProbe:

            httpGet:
              path: /health
              port: 8000

            initialDelaySeconds: 15
            periodSeconds: 10

  strategy:

    canary:

      canaryService: backend-canary

      stableService: backend-stable

      steps:

        - setWeight: 10

        - pause:
            duration: 5m

        - setWeight: 25

        - pause:
            duration: 5m

        - setWeight: 50

        - pause:
            duration: 10m

      analysis:

        templates:

          - templateName: backend-analysis

14. Stable Service

apiVersion: v1
kind: Service

metadata:
  name: backend-stable

spec:

  selector:
    app: backend

  ports:

    - port: 80
      targetPort: 8000

15. Canary Service

apiVersion: v1
kind: Service

metadata:
  name: backend-canary

spec:

  selector:
    app: backend

  ports:

    - port: 80
      targetPort: 8000

중요한 점은 실제 트래픽을 어떻게 Stable/Canary로 분배할지 Ingress/Gateway 및 Argo Rollouts traffic routing 설정까지 함께 설계해야 한다는 것이다.

단순 Service 두 개만 만든다고 일반적인 Kubernetes 환경에서 HTTP 요청이 자동으로 10/90으로 분배되는 것은 아니다.


16. PostgreSQL

개발/실습 환경에서는 PostgreSQL을 Kubernetes에 직접 구성할 수 있다.

apiVersion: apps/v1
kind: Deployment

metadata:
  name: postgres

spec:

  replicas: 1

  selector:
    matchLabels:
      app: postgres

  template:

    metadata:
      labels:
        app: postgres

    spec:

      containers:

        - name: postgres

          image: postgres:17

          env:

            - name: POSTGRES_DB
              value: taxdb

            - name: POSTGRES_USER
              value: taxuser

            - name: POSTGRES_PASSWORD
              valueFrom:
                secretKeyRef:
                  name: postgres-secret
                  key: password

          ports:

            - containerPort: 5432

실제 Production에서는 PostgreSQL을 Kubernetes에 직접 배포하기 전에 관리형 DB나 PostgreSQL Operator 등을 검토하는 것이 좋다.


17. Redis

Redis Deployment:

apiVersion: apps/v1
kind: Deployment

metadata:
  name: redis

spec:

  replicas: 1

  selector:
    matchLabels:
      app: redis

  template:

    metadata:
      labels:
        app: redis

    spec:

      containers:

        - name: redis

          image: redis:7

          ports:

            - containerPort: 6379

Service:

apiVersion: v1
kind: Service

metadata:
  name: redis

spec:

  selector:
    app: redis

  ports:

    - port: 6379
      targetPort: 6379

18. Prometheus AnalysisTemplate

이제 핵심이다.

Argo Rollouts가 Prometheus에 질문하도록 AnalysisTemplate을 만든다.

apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate

metadata:
  name: backend-analysis

spec:

  metrics:

    - name: error-rate

      interval: 1m

      count: 5

      failureLimit: 1

      successCondition: result[0] < 0.01

      provider:

        prometheus:

          address: http://prometheus-operated.monitoring.svc:9090

          query: |
            (
              sum(
                rate(
                  http_requests_total{
                    app="backend",
                    status=~"5.."
                  }[5m]
                )
              )
              /
              sum(
                rate(
                  http_requests_total{
                    app="backend"
                  }[5m]
                )
              )
            )

    - name: p95-latency

      interval: 1m

      count: 5

      failureLimit: 1

      successCondition: result[0] < 0.5

      provider:

        prometheus:

          address: http://prometheus-operated.monitoring.svc:9090

          query: |
            histogram_quantile(
              0.95,
              sum(
                rate(
                  http_request_duration_seconds_bucket{
                    app="backend"
                  }[5m]
                )
              ) by (le)
            )

여기서 두 가지를 검사한다.

HTTP 5xx Error Rate
P95 Latency

19. Error Rate 기준

설정:

successCondition: result[0] < 0.01

의 의미는:

Error Rate < 1%

이다.

예:

0.2%

이면 성공.

0.8%

이면 성공.

3.0%

이면 실패할 수 있다.


20. P95 Latency 기준

다음 조건:

successCondition: result[0] < 0.5

는:

P95 < 0.5 seconds

즉:

P95 < 500ms

를 의미한다.

예:

220ms

정상.

480ms

정상.

1,200ms

실패할 수 있다.


21. Analysis 실행 방식에서 주의할 점

여기서 한 가지 중요한 문제가 있다.

Prometheus Query가 전체 Backend의 Metrics를 조회한다면 Stable과 Canary가 섞일 수 있다.

예:

Stable
v1
90%

Canary
v2
10%

Prometheus:

전체 backend

를 조회하면:

v1 + v2

가 합쳐진 결과가 나온다.

그러면 Canary 자체의 오류율을 정확하게 판단하기 어렵다.

따라서 실제 운영에서는 Canary Pod를 식별할 수 있는 Label 또는 Traffic Routing 기준을 Prometheus Query에 반영해야 한다.


22. Canary Metrics를 분리하는 이유

예를 들어:

Stable
Error = 0.2%

Canary
Error = 15%

인데 전체 요청을 합치면:

전체 Error = 1.68%

정도가 될 수 있다.

전체 기준만 보면 문제가 명확하지 않을 수 있다.

따라서:

Canary Error Rate

를 별도로 측정해야 한다.


23. Canary 식별 Label

예를 들어 Pod에 다음과 같은 Label을 추가할 수 있다.

metadata:

  labels:

    app: backend

    rollout: backend

실제 환경에서는 Rollout이 생성하는 ReplicaSet/Pod의 revision 식별 정보나 Service/Gateway traffic routing 정보를 활용해 현재 Canary ReplicaSet만 대상으로 하는 Query를 설계하는 것이 더 정확하다.

핵심은:

Prometheus
    ↓
Canary만 조회

하는 것이다.


24. GitHub Actions

이제 CI를 만든다.

파일:

.github/workflows/build.yml

예:

name: Build Backend

on:

  push:

    branches:
      - main

jobs:

  build:

    runs-on: ubuntu-latest

    steps:

      - name: Checkout

        uses: actions/checkout@v4

      - name: Login Registry

        uses: docker/login-action@v3

        with:

          registry: ghcr.io

          username: ${{ github.actor }}

          password: ${{ secrets.GITHUB_TOKEN }}

      - name: Build

        run: |

          docker build \
            -t ghcr.io/${{ github.repository }}/backend:${{ github.sha }} \
            .

      - name: Push

        run: |

          docker push \
            ghcr.io/${{ github.repository }}/backend:${{ github.sha }}

여기까지는 CI다.


25. Image Tag 전략

운영 환경에서는 다음과 같은 Tag 전략을 사용할 수 있다.

1.0.0
1.0.1
1.1.0

또는 Git SHA:

a81f7c2

실전에서는 Git SHA나 immutable version을 사용하는 방식이 안전하다.

예:

backend:a81f7c2

다음과 같이 동일 Tag를 계속 덮어쓰는 방식은 피하는 것이 좋다.

backend:latest

왜냐하면 정확히 어떤 이미지가 배포되었는지 추적하기 어렵기 때문이다.


26. GitOps values 변경

CI가 이미지를 Registry에 Push한 후 GitOps Repository의 values-prod.yaml을 변경한다.

기존:

image:
  tag: "1.0.0"

변경:

image:
  tag: "1.1.0"

Commit:

Update backend image to 1.1.0

이제 GitOps Repository가 새로운 Desired State를 갖는다.


27. GitHub Actions에서 GitOps Repository 변경

개념적인 구조는 다음과 같다.

- name: Update GitOps

  run: |

    git clone \
      https://github.com/example/tax-gitops.git

    cd tax-gitops

    sed -i \
      "s/tag: .*/tag: \"${IMAGE_TAG}\"/" \
      charts/tax-app/values-prod.yaml

    git add .

    git commit \
      -m "Update backend image ${IMAGE_TAG}"

    git push

운영 환경에서는 GitHub App 또는 제한된 권한의 Deploy Key 등을 사용해 GitOps Repository에 최소 권한으로 접근하도록 구성하는 것이 좋다.


28. Argo CD Application

Argo CD Application:

apiVersion: argoproj.io/v1alpha1
kind: Application

metadata:
  name: tax-app
  namespace: argocd

spec:

  project: default

  source:

    repoURL:
      https://github.com/example/tax-gitops.git

    targetRevision:
      main

    path:
      charts/tax-app

    helm:

      valueFiles:

        - values-prod.yaml

  destination:

    server:
      https://kubernetes.default.svc

    namespace:
      tax-prod

  syncPolicy:

    automated:

      prune: true

      selfHeal: true

    syncOptions:

      - CreateNamespace=true

29. GitOps 배포 시작

Developer:

git push origin main

GitHub Actions:

Test
 ↓
Docker Build
 ↓
Registry Push
 ↓
GitOps Update

GitOps:

values-prod.yaml

변경.

Argo CD:

OutOfSync
 ↓
Sync

Kubernetes:

Rollout
 ↓
New ReplicaSet

30. Argo Rollouts Canary 시작

기존:

Stable
v1.0.0
████████████████████

새 버전:

Canary
v1.1.0
██

Traffic:

Stable = 90%
Canary = 10%

그리고 Prometheus Analysis가 실행된다.

Canary 10%
     ↓
Prometheus
     ↓
Error Rate
     ↓
P95 Latency

31. 10% 단계

첫 번째 단계:

- setWeight: 10

- pause:
    duration: 5m

즉:

v1 = 90%
v2 = 10%

상태를 일정 시간 관찰한다.


32. 10%에서 정상일 경우

Prometheus:

Error Rate = 0.2%
P95 = 180ms

기준:

Error Rate < 1%
P95 < 500ms

이므로 다음 단계로 이동한다.

10%
 ↓
25%

33. 25% 단계

v1 = 75%
v2 = 25%

다시 검증한다.

Prometheus
   │
   ├── Error Rate
   │
   └── P95 Latency

정상이면:

25%
 ↓
50%

34. 50% 단계

v1 = 50%
v2 = 50%

이 단계에서는 신규 버전이 상당한 트래픽을 처리한다.

예:

Requests

v1 = 5,000
v2 = 5,000

다시 Metrics를 검사한다.


35. 100% Promotion

모든 검증이 성공하면:

v1 = 0%
v2 = 100%

이제 새로운 버전이 Production Version이 된다.

Stable
v1.1.0
████████████████████

36. 실패 시나리오

이번에는 의도적으로 새로운 버전에 오류를 넣어보자.

예를 들어:

@app.get("/api/expenses")
def expenses():

    raise Exception(
        "Test deployment failure"
    )

새로운 Image:

backend:1.2.0

를 빌드한다.


37. Canary 10%

Argo CD가 새로운 Rollout을 시작한다.

v1.1.0
90%

v1.2.0
10%

Canary에서 오류가 발생한다.

HTTP 500

Prometheus:

Error Rate = 12%

기준:

< 1%

실패.


38. Analysis Failure

Argo Rollouts:

Analysis
   ↓
FAIL

상태:

Abort

그리고 Canary Promotion을 중단한다.

10%
 ↓
FAIL
 ↓
ABORT

39. Stable Version 유지

Canary가 실패했으므로 Stable Version은 그대로 유지한다.

Stable
v1.1.0
████████████████████

Canary
v1.2.0
        X

사용자 트래픽은 Stable Version으로 계속 전달된다.

User
 ↓
Stable
 ↓
v1.1.0

이것이 Canary 배포의 핵심이다.


40. P95 Latency 장애 테스트

이번에는 HTTP 오류 대신 성능 문제를 만든다.

import time

@app.get("/api/expenses")
def expenses():

    time.sleep(3)

    return {
        "items": []
    }

Canary에서 응답 시간이 증가한다.

P95

Stable
200ms

Canary
3,000ms

Prometheus Query:

P95 = 3.0

기준:

< 0.5

실패한다.

Analysis
   ↓
FAIL
   ↓
Abort

41. Error Rate + Latency를 동시에 검사

실전에서는 두 조건을 모두 확인하는 것이 좋다.

Canary
   │
   ├── Error Rate
   │
   └── P95 Latency

예:

Error Rate < 1%
AND
P95 < 500ms

둘 중 하나라도 기준을 초과하면 배포를 중단한다.


42. FailureLimit

예:

failureLimit: 1

이 값은 Analysis 실패를 몇 번까지 허용할 것인지와 관련된다.

운영 환경에서는 너무 민감하게 설정하면 일시적인 Metrics 변동 때문에 배포가 중단될 수 있다.

반대로 너무 관대하면 실제 장애를 놓칠 수 있다.

따라서:

Traffic 규모
Metrics 안정성
서비스 중요도
배포 위험도

를 고려해 설정해야 한다.


43. Prometheus 데이터가 없는 경우

새로운 Canary Pod가 아직 충분한 요청을 처리하지 않았다면 Prometheus Query 결과가 없을 수 있다.

이 경우 Analysis가 예상과 다르게 동작할 수 있다.

따라서 다음을 고려해야 한다.

Initial Delay
Interval
Count
Traffic Volume

예:

initialDelay: 2m

개념적으로:

Canary 시작
   ↓
2분 대기
   ↓
Metrics 수집
   ↓
Analysis

44. 작은 트래픽 환경의 문제

Canary가 10%라고 해도 전체 요청이 매우 적다면 통계적으로 의미 있는 결과가 나오지 않을 수 있다.

예:

5분간 전체 요청 = 10건
Canary 요청 = 1건

이때:

Error Rate = 100%

이라는 결과가 나올 수도 있지만 표본 자체가 너무 작다.

따라서 실제 운영에서는:

Minimum Request Count
Observation Window

등을 고려해야 한다.


45. Grafana에서 Rollout 관찰

Grafana에서는 다음 Dashboard를 구성한다.

==================================================
Backend Progressive Delivery
==================================================

Version

Stable : 1.1.0
Canary : 1.2.0

Canary Weight
10%

--------------------------------------------------

HTTP Request Rate

Stable  ███████████████
Canary  ██

--------------------------------------------------

5xx Error Rate

Stable  0.2%
Canary  0.4%

--------------------------------------------------

P95 Latency

Stable  180ms
Canary  220ms

--------------------------------------------------

Pod

Stable  4
Canary  1

--------------------------------------------------

Rollout

Analysis : PASS
Next     : 25%
==================================================

46. Loki로 Canary 로그 확인

Metrics에서 이상이 발견되면 Loki를 검색한다.

예:

{namespace="tax-prod", app="backend"}
|= "ERROR"

특정 Revision을 구분할 수 있다면:

{namespace="tax-prod", app="backend", revision="abc123"}
|= "ERROR"

결과:

DatabaseError
Connection refused
Timeout

등을 확인할 수 있다.


47. Tempo로 Canary Trace 확인

Tempo를 사용하면 특정 요청의 전체 호출 흐름을 확인할 수 있다.

HTTP Request
      │
      ▼
FastAPI
      │
 ┌────┴────┐
 ▼         ▼
Redis   PostgreSQL

예를 들어 Canary에서:

FastAPI = 20ms
Redis   = 10ms
Postgres = 1,800ms

라면 PostgreSQL 호출이 지연의 원인인지 추가로 조사할 수 있다.


48. 자동 Rollback의 전체 흐름

이제 가장 중요한 부분이다.

Git Push
    ↓
GitHub Actions
    ↓
Docker Build
    ↓
Registry
    ↓
GitOps values 변경
    ↓
Argo CD
    ↓
Argo Rollouts
    ↓
Canary 10%
    ↓
Prometheus
    │
    ├── Error Rate OK
    └── P95 OK
    ↓
Canary 25%
    ↓
Prometheus
    ↓
Canary 50%
    ↓
Prometheus
    ↓
100%

장애 발생:

Canary 10%
    ↓
HTTP 5xx 증가
    ↓
Prometheus
    ↓
Analysis Failure
    ↓
Argo Rollouts Abort
    ↓
Canary 중단
    ↓
Stable Version 유지

이것이 자동화된 Progressive Delivery다.


49. GitOps Repository까지 되돌리는 방법

Runtime Rollback과 GitOps Rollback은 구분해야 한다.

예를 들어 Git:

1.2.0

인데 Runtime이:

1.1.0

이면 Git과 Kubernetes 상태가 달라질 수 있다.

따라서 최종적으로 Git도 이전 버전으로 되돌리는 것이 좋다.

Git
1.2.0
 ↓
revert
 ↓
1.1.0
 ↓
Argo CD
 ↓
Kubernetes

즉:

Rollback 후 Git을 Source of Truth와 일치시키는 작업이 필요하다.


50. 자동 Rollback을 구현할 때의 현실적인 구조

Production에서는 다음과 같이 단계적으로 자동화하는 것을 권장할 수 있다.

Level 1

Kubernetes Self-Healing
Level 2

Readiness/Liveness
Level 3

HPA
Level 4

Canary Analysis
Level 5

Automatic Abort
Level 6

Automatic Rollback

그리고 Database 변경이나 데이터 손상 가능성이 있는 작업은 별도의 승인과 보호 장치를 두는 것이 중요하다.


51. Blue/Green으로 변경하려면

같은 애플리케이션을 Blue/Green으로 변경할 수도 있다.

strategy:

  blueGreen:

    activeService:
      backend-active

    previewService:
      backend-preview

    autoPromotionEnabled:
      false

구조:

                 User
                   │
                   ▼
             Active Service
                   │
                   ▼
                 v1.1

새 버전:

Preview Service
      │
      ▼
    v1.2

Prometheus:

Preview
   ↓
Analysis

검증 성공:

Promotion

문제:

Abort

52. Canary와 Blue/Green 선택

실제 서비스에서는 다음과 같이 생각할 수 있다.

Canary

가 적합한 경우:

트래픽을 점진적으로 이동하고 싶다
작은 비율부터 실제 사용자 트래픽으로 검증하고 싶다
Blue/Green

이 적합한 경우:

새로운 환경을 별도로 준비할 수 있다
빠른 Traffic 전환이 중요하다
Rollback을 빠르게 수행해야 한다

서비스 특성에 따라 선택한다.


53. CI/CD와 GitOps의 역할 분리

최종적으로 역할을 명확하게 나누는 것이 중요하다.

GitHub Actions

역할:

Test
Build
Scan
Push Image
GitOps Repository

역할:

Desired State
Argo CD

역할:

Git → Kubernetes
Argo Rollouts

역할:

Progressive Delivery
Prometheus

역할:

Deployment Metrics
Grafana

역할:

Visualization
Loki

역할:

Logs
Tempo

역할:

Traces

54. 운영 배포 프로세스

실제 Production에서는 다음과 같은 프로세스를 구성할 수 있다.

Developer
   ↓
Pull Request
   ↓
Code Review
   ↓
CI Test
   ↓
Security Scan
   ↓
Build Image
   ↓
Push Registry
   ↓
Update GitOps
   ↓
Argo CD
   ↓
Argo Rollouts
   ↓
Canary
   ↓
Prometheus Analysis

정상:

10%
 ↓
25%
 ↓
50%
 ↓
100%

실패:

Analysis Failure
 ↓
Abort
 ↓
Stable Version

55. 장애 발생 후 분석

자동 Rollback이 수행됐다고 해서 장애 원인 분석이 끝난 것은 아니다.

다음과 같이 분석한다.

Prometheus
    ↓
왜 Rollback 되었는가?
    ↓
Grafana
    ↓
어떤 Metric이 증가했는가?
    ↓
Loki
    ↓
어떤 Error가 발생했는가?
    ↓
Tempo
    ↓
어떤 요청에서 문제가 발생했는가?

예:

5xx 증가
   ↓
Database timeout
   ↓
PostgreSQL Query 2초
   ↓
새로운 Index 누락

이렇게 Root Cause를 찾아야 한다.


56. 장애 이후 Git 수정

원인을 찾았다면 코드를 수정한다.

Bug
 ↓
Code Fix
 ↓
Git Commit
 ↓
GitHub Actions
 ↓
New Image
 ↓
GitOps
 ↓
Argo CD
 ↓
Argo Rollouts

그리고 다시 Canary 배포한다.

v1.2.1
 ↓
10%
 ↓
Analysis
 ↓
25%
 ↓
50%
 ↓
100%

57. 최종 자동화 Pipeline

이번 실습에서 구축한 전체 Pipeline은 다음과 같다.

┌─────────────────────────────────────────────┐
│                 Developer                   │
└──────────────────────┬──────────────────────┘
                       │
                       ▼
                 Git Push / PR
                       │
                       ▼
┌─────────────────────────────────────────────┐
│              GitHub Actions                 │
│                                             │
│ Test → Build → Security → Docker Push      │
└──────────────────────┬──────────────────────┘
                       │
                       ▼
                Container Registry
                       │
                       ▼
┌─────────────────────────────────────────────┐
│               GitOps Repository             │
│                                             │
│ Helm values-prod.yaml                       │
│ image.tag = new version                     │
└──────────────────────┬──────────────────────┘
                       │
                       ▼
                  ┌──────────┐
                  │ Argo CD  │
                  └────┬─────┘
                       │
                       ▼
              ┌─────────────────┐
              │ Argo Rollouts   │
              └───────┬─────────┘
                      │
          ┌───────────┼───────────┐
          ▼           ▼           ▼
        10%          25%         50%
          │           │           │
          ▼           ▼           ▼
      Prometheus  Prometheus  Prometheus
          │           │           │
          └───────────┼───────────┘
                      │
                 Analysis
                      │
                ┌─────┴─────┐
                ▼           ▼
              PASS         FAIL
                │           │
                ▼           ▼
             Promote      Abort
                │           │
                ▼           ▼
              100%       Rollback

58. 장애가 없을 때

정상적인 배포는 다음과 같다.

v1.0
 ↓
Canary v1.1 10%
 ↓
Prometheus PASS
 ↓
25%
 ↓
Prometheus PASS
 ↓
50%
 ↓
Prometheus PASS
 ↓
100%
 ↓
Deployment Complete

59. HTTP 5xx 장애가 발생할 때

v1.0
 ↓
Canary v1.1 10%
 ↓
HTTP 5xx ↑
 ↓
Prometheus
 ↓
Analysis FAIL
 ↓
Abort
 ↓
Stable v1.0

60. P95 Latency 장애가 발생할 때

v1.0
 ↓
Canary v1.1
 ↓
P95 2,000ms
 ↓
Prometheus
 ↓
Analysis FAIL
 ↓
Abort
 ↓
Stable v1.0

61. 최종적으로 구축된 시스템

이번 실습을 통해 다음과 같은 운영 플랫폼이 만들어졌다.

                    ┌─────────────┐
                    │ Developer   │
                    └──────┬──────┘
                           │
                           ▼
                    ┌─────────────┐
                    │ GitHub      │
                    └──────┬──────┘
                           │
                           ▼
                 ┌──────────────────┐
                 │ GitHub Actions    │
                 │                  │
                 │ Test             │
                 │ Build            │
                 │ Scan             │
                 │ Push             │
                 └────────┬─────────┘
                          │
                          ▼
                  ┌───────────────┐
                  │   Registry    │
                  └───────┬───────┘
                          │
                          ▼
                  ┌───────────────┐
                  │   GitOps      │
                  │   Helm        │
                  └───────┬───────┘
                          │
                          ▼
                    ┌──────────┐
                    │ Argo CD  │
                    └────┬─────┘
                         │
                         ▼
                  ┌──────────────┐
                  │Argo Rollouts │
                  └──────┬───────┘
                         │
              ┌──────────┴──────────┐
              ▼                     ▼
           Stable                 Canary
             v1                     v2
              │                     │
              └──────────┬──────────┘
                         │
                         ▼
                    Kubernetes
                         │
          ┌──────────────┼──────────────┐
          ▼              ▼              ▼
       FastAPI       PostgreSQL       Redis
          │
          ▼
     Prometheus
          │
     ┌────┴────┐
     ▼         ▼
  5xx Rate   P95 Latency
     │         │
     └────┬────┘
          ▼
      Analysis
          │
     ┌────┴────┐
     ▼         ▼
   PASS       FAIL
     │         │
     ▼         ▼
 Promote     Abort
     │         │
     ▼         ▼
   100%     Rollback

62. 시스템 엔지니어가 반드시 이해해야 하는 핵심

이번 실습에서 가장 중요한 것은 명령어 자체가 아니다.

각 구성요소의 역할을 이해하는 것이다.

GitHub Actions
    =
"새로운 프로그램을 만들고 검증한다"
GitOps
    =
"어떤 버전을 운영할 것인지 선언한다"
Argo CD
    =
"Git의 선언 상태를 Kubernetes에 반영한다"
Argo Rollouts
    =
"새 버전을 안전하게 조금씩 배포한다"
Prometheus
    =
"배포된 버전이 실제로 정상인지 측정한다"
Grafana
    =
"측정 결과를 사람이 이해할 수 있도록 보여준다"
Loki
    =
"문제가 발생했을 때 로그를 찾는다"
Tempo
    =
"문제가 발생한 요청의 전체 호출 경로를 추적한다"

63. 최종 정리

이번 실습의 핵심 Pipeline은 다음 한 줄로 정리할 수 있다.

GitHub Actions
→ Docker Registry
→ GitOps
→ Argo CD
→ Argo Rollouts
→ Canary
→ Prometheus Analysis
→ Promote 또는 Abort/Rollback

정상적인 경우:

10%
 ↓
25%
 ↓
50%
 ↓
100%

장애가 발생하면:

Canary
 ↓
HTTP 5xx ↑
또는
P95 Latency ↑
 ↓
Prometheus
 ↓
Analysis Failure
 ↓
Argo Rollouts Abort
 ↓
Stable Version 유지/복구

그리고 장애가 끝난 후에는:

Grafana
 ↓
Loki
 ↓
Tempo
 ↓
Root Cause Analysis
 ↓
Code Fix
 ↓
CI
 ↓
GitOps
 ↓
Argo CD
 ↓
Argo Rollouts

으로 다시 연결된다.

결국 우리가 만든 구조는 단순한 CI/CD가 아니다.

              CI/CD
                ↓
              GitOps
                ↓
        Progressive Delivery
                ↓
          Observability
                ↓
          Automated Analysis
                ↓
          Automated Rollback
                ↓
             SRE 운영

이라는 하나의 운영 체계가 된다.

특히 중요한 것은 "배포가 성공했다"의 의미가 단순히 Pod가 Running 상태가 되는 것이 아니라는 것이다.

실제 운영에서는:

Pod Ready
   +
HTTP 5xx 정상
   +
P95 Latency 정상
   +
Logs 정상
   +
Trace 정상

까지 확인해야 진정한 배포 성공에 가까워진다.

이 구조를 기반으로 하면 다음 단계에서는 Argo Rollouts + Istio 또는 NGINX Ingress/Gateway API를 이용한 실제 HTTP 트래픽 10% → 25% → 50% → 100% 제어, Canary 전용 Prometheus Metrics, Grafana 배포 Dashboard, Loki/Tempo 연계, 그리고 GitHub Actions에서 이미지 생성부터 Production 자동 Rollback까지 실제 Kubernetes 클러스터에서 동작하는 완성형 실습 환경으로 확장할 수 있다.


댓글