Skip to content

fix: 배포 헬스체크를 로그 기반 완료 신호 감지 방식으로 변경#196

Merged
unam98 merged 2 commits into
devfrom
fix/log-based-healthcheck
Jul 24, 2026
Merged

fix: 배포 헬스체크를 로그 기반 완료 신호 감지 방식으로 변경#196
unam98 merged 2 commits into
devfrom
fix/log-based-healthcheck

Conversation

@unam98

@unam98 unam98 commented Jul 21, 2026

Copy link
Copy Markdown
Collaborator

작업 배경

  • OTel agent 적용 후 부팅 시간이 3s→20~30s로 늘어나면서, 고정 타임아웃(15s+10x10s) 기반 헬스체크가 정상 배포도 "실패"로 오판하는 구조적 문제 발견

변경 사항

영역 내용
scripts/deploy.sh 고정 폴링 방식 제거, 프로세스 생존 여부(진짜 실패) + 로그 기반 완료 신호(Started ServerApplication) 감지 방식으로 교체

영향 범위

  • 진짜 크래시(프로세스 종료)는 즉시 감지, 정상 부팅 중이면 시간이 얼마나 걸리든 계속 대기 (임의의 초/횟수 추측 제거)
  • 5분 상한선은 여전히 존재하나, 이는 "정상 부팅 시간" 추측이 아니라 진짜 행(hang) 상황에 대한 안전장치로 성격이 다름
  • CI/CD 워크플로우(dev-cd.yml 등)는 변경 없음

Test Plan

  • bash -n scripts/deploy.sh 문법 검증
  • 머지 후 실제 재배포하여 헬스체크가 정확히 판정하는지 확인

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Bug Fixes
    • Improved deployment startup verification by waiting for the application to report that it has started.
    • Added detection and reporting for startup crashes and readiness timeouts.
    • Displays recent application logs when startup fails.
    • Performs a final health check after successful startup.

기존 방식(15초 대기 + 10회x10초 폴링)은 OTel agent 계측으로 부팅 시간이
늘어나면서(3s→20~30s) 정상 배포도 실패로 오판하는 문제가 있었음.

이제 프로세스 생존 여부로 진짜 실패(크래시)를 즉시 감지하고, 살아있는 동안은
Spring Boot가 실제로 찍는 완료 로그("Started ServerApplication")를 기다림.
임의의 타임아웃 추측 대신 실제 이벤트 기반 판단이라 부팅 시간 변동에 안전하고,
5분 상한선은 진짜 행(hang) 상황에 대한 안전장치로만 남김.
@unam98 unam98 self-assigned this Jul 21, 2026
@coderabbitai

coderabbitai Bot commented Jul 21, 2026

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 135a36a1-b0c6-4c7d-a5df-429374a52e4f

📥 Commits

Reviewing files that changed from the base of the PR and between c0da68e and b2a2de2.

📒 Files selected for processing (1)
  • scripts/deploy.sh
🚧 Files skipped from review as they are similar to previous changes (1)
  • scripts/deploy.sh

📝 Walkthrough

Walkthrough

Deployment startup now uses a log-based readiness signal, monitors process liveness for failures, enforces a 240-second timeout, and performs one health check after readiness succeeds.

Changes

Deployment readiness

Layer / File(s) Summary
Log-driven startup and health verification
scripts/deploy.sh
Starts the Java application with nohup, waits up to 240 seconds for the startup log signal, detects crashes and timeouts, then performs one health endpoint request.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed 제목이 배포 헬스체크를 로그 기반 완료 신호 감지로 바꾸는 핵심 변경을 정확히 요약합니다.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/log-based-healthcheck

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (1)
scripts/deploy.sh (1)

33-41: 🩺 Stability & Availability | 🔵 Trivial | ⚡ Quick win

Unbounded log growth from switching to append mode.

LOG_FILE is now appended (>> "$LOG_FILE") instead of truncated, needed to track START_LINE for readiness detection. Over many deploys with no rotation, nohup.out will grow indefinitely (worse with OTel agent output) and can eventually fill disk on the instance.

Also, LOG_FILE re-hardcodes /home/ec2-user/app/nohup.out even though APP_DIR=/home/ec2-user/app is already defined and the script cds into it — consider LOG_FILE="$APP_DIR/nohup.out" for consistency.

♻️ Suggested tweaks
-LOG_FILE=/home/ec2-user/app/nohup.out
+LOG_FILE="$APP_DIR/nohup.out"
 START_LINE=$(wc -l < "$LOG_FILE" 2>/dev/null || echo 0)

Add basic rotation (e.g., via logrotate config for nohup.out, or truncate/archive it on each deploy once START_LINE bookkeeping is no longer needed for the previous run).

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@scripts/deploy.sh` around lines 33 - 41, Update the deployment script’s
LOG_FILE to derive from APP_DIR instead of hard-coding the application path, and
add basic nohup.out rotation or archival before appending new output. Preserve
START_LINE-based readiness detection by rotating only when the previous run’s
log bookkeeping is no longer needed.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@scripts/deploy.sh`:
- Around line 75-77: Update the health-check curl invocation in the deployment
script to include an explicit request timeout, ensuring an unavailable or
unresponsive actuator endpoint cannot block indefinitely. Keep the existing
health response logging behavior and do not change deployment success gating.
- Around line 43-73: Reduce MAX_WAIT in the readiness loop to leave sufficient
time below the AppSpec hook’s 300-second timeout for the earlier shutdown delay,
health check, and cleanup diagnostics; use a margin such as 240–270 seconds
while preserving the existing timeout failure path and 5-minute overall
hang-protection intent.

---

Nitpick comments:
In `@scripts/deploy.sh`:
- Around line 33-41: Update the deployment script’s LOG_FILE to derive from
APP_DIR instead of hard-coding the application path, and add basic nohup.out
rotation or archival before appending new output. Preserve START_LINE-based
readiness detection by rotating only when the previous run’s log bookkeeping is
no longer needed.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 82ea5247-ff4d-497f-86fd-985aeee49e1a

📥 Commits

Reviewing files that changed from the base of the PR and between e845749 and c0da68e.

📒 Files selected for processing (1)
  • scripts/deploy.sh

Comment thread scripts/deploy.sh Outdated
Comment thread scripts/deploy.sh Outdated
- MAX_WAIT을 appspec.yml의 AfterInstall 훅 타임아웃(300초)과 정확히
  일치시키지 않고 240초로 낮춰서, CodeDeploy의 외부 강제종료가
  스크립트 자체의 정상 종료 로직보다 먼저 발동하지 않도록 여유 확보
- 헬스체크 curl에 --max-time 10 추가, 포트가 아직 안 열려있을 때
  무한 대기하지 않도록 방어
@unam98
unam98 merged commit 253bc7d into dev Jul 24, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants