OBSERVABILITY · 2026

서비스 오류와 성능 모니터링

Service error and performance monitoring

서버 오류만 보던 운영 체계를 분산 트레이싱, 구조화 로그, CPU·메모리, 실제 사용자의 Web Vitals까지 확장했습니다.

Expanded operations from server errors to traces, structured logs, infrastructure signals, and real-user Web Vitals.

규모·성과

1 → 4 핵심 운영 알람 확장

담당 업무

분석 · 설계 · 구현, 1인 전담

사용 기술

OpenTelemetry · Grafana · Prometheus · Sentry · Next.js

배경

BFF와 백엔드 사이의 장애 원인을 한 흐름으로 보기 어려웠고, 운영 알람도 메모리에 치우쳐 CPU 병목과 사용자 체감 성능 저하를 미리 알 수 없었습니다.

과정

  • 기존 Sentry 알림은 유지하고 OpenTelemetry 트레이싱을 추가해 전환 위험을 줄였습니다.
  • 트레이스에서 관련 구조화 로그로 바로 이동할 수 있게 조사 흐름을 연결했습니다.
  • 브라우저 Web Vitals를 자체 API와 Prometheus Histogram으로 수집해 서버 지표와 함께 알렸습니다.

결과

  • BFF-백엔드 호출 흐름과 로그를 한 화면에서 조사
  • 메모리 중심 알람을 CPU·LCP·INP 영역으로 확장
  • 인프라 상태와 실제 사용자 경험을 함께 보는 기반 확보
At a glance

1 → 4 core operational alerts

Role

Sole owner across analysis, design, and implementation

Stack

OpenTelemetry · Grafana · Prometheus · Sentry · Next.js

Context

Failures across the BFF and backend were difficult to trace as one request flow. Monitoring focused on memory, leaving CPU pressure and real-user performance regressions unseen.

Process

  • Kept existing Sentry alerts and added OpenTelemetry traces to reduce migration risk.
  • Linked traces directly to related structured logs for a simpler investigation flow.
  • Collected browser Web Vitals through an internal API and Prometheus histograms alongside server signals.

Outcome

  • Investigated BFF-to-backend flows and logs in one view
  • Expanded alerts from memory to CPU, LCP, and INP
  • Connected infrastructure health with real user experience