<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="ko_KR"><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://stan-dev.cloud/feed.xml" rel="self" type="application/atom+xml" /><link href="https://stan-dev.cloud/" rel="alternate" type="text/html" hreflang="ko_KR" /><updated>2026-10-08T17:41:57+09:00</updated><id>https://stan-dev.cloud/feed.xml</id><title type="html">Stan 기술블로그</title><subtitle>사내 인프라를 운영하며 만난 장애와 그때의 판단을 기록합니다. self-hosted Sentry, Keycloak, Nginx 리버스 프록시, 레거시 모니터링 콘솔, 사내 코딩 어시스턴트.</subtitle><author><name>이진용</name><email>kouig14@gmail.com</email></author><entry><title type="html">코드는 다 있는데 알림이 안 온다</title><link href="https://stan-dev.cloud/smon-alert-integration/" rel="alternate" type="text/html" title="코드는 다 있는데 알림이 안 온다" /><published>2026-09-17T09:46:00+09:00</published><updated>2026-09-17T09:46:00+09:00</updated><id>https://stan-dev.cloud/smon-alert-integration</id><content type="html" xml:base="https://stan-dev.cloud/smon-alert-integration/"><![CDATA[<blockquote>
  <p>앞선 <a href="/npm-reverse-proxy-realm-access/">리버스 프록시 두 편</a>이 트래픽 앞단 얘기였다면, 이번엔 <strong>인수받은 사내 모니터링 콘솔</strong>에 알림을 붙이다가 겪은 일이다. 결론부터: 알림은 켜졌고, 그 과정에서 <strong>2주 전 내 커밋이 심어둔 지뢰</strong>를 발견해 제거했다.</p>

  <p><em>(사내 인프라 특성상 서버·도메인·사람 이름·수치 일부는 일반화했다. 시스템 이름은 “콘솔”로 부른다.)</em></p>
</blockquote>

<hr />

<h2 id="배경--인수받은-콘솔">배경 — 인수받은 콘솔</h2>

<p>콘솔은 사내 서버들의 CPU/메모리/디스크를 수집해 대시보드로 보여주는 Java 8 서블릿 앱이다. 특징이 몇 가지 있다.</p>

<ul>
  <li><strong>SVN 커밋 = 즉시 운영 배포.</strong> 서버 쪽 훅이 바뀐 <code class="language-plaintext highlighter-rouge">.java</code>를 컴파일해서 돌아가는 Tomcat에 밀어넣는다. 스테이징은 없다.</li>
  <li>나는 <strong>서버 접근 권한이 없다.</strong> 할 수 있는 건 SVN 커밋과 관리자 웹 화면뿐.</li>
  <li>자체 프레임워크, 자체 마이그레이션 도구, 자체 스케줄러. 문서는 있지만 3천 줄짜리 XML 한 파일에 SQL이 다 들어있는 식.</li>
</ul>

<p>이런 시스템을 인수받고 첫 과내가 “알림을 Slack으로 받게 해달라”였다.</p>

<h2 id="코드를-읽었더니--이미-다-있었다">코드를 읽었더니 — 이미 다 있었다</h2>

<p>먼저 “알림 기능을 만들어야 하나?”부터 확인했다. 코드를 읽어보니:</p>

<ul>
  <li><code class="language-plaintext highlighter-rouge">Slack.java</code> — Incoming Webhook으로 POST, 3회 재시도</li>
  <li><code class="language-plaintext highlighter-rouge">AlertNotifier.java</code> — 임계치 초과(FIRING) / 해소(RESOLVED) 단계별 발송, 실패하면 DB 컬럼(<code class="language-plaintext highlighter-rouge">notified_at</code>)을 비워둬서 다음 사이클에 재시도</li>
  <li>60초마다 도는 평가 배치가 규칙을 읽어 위 둘을 호출</li>
</ul>

<p><strong>전부 있었다.</strong> 없는 건 딱 하나, 서버에 webhook 주소가 설정돼 있지 않았던 것. 주소가 없으니 <code class="language-plaintext highlighter-rouge">Slack.post()</code>는 매번 조용히 <code class="language-plaintext highlighter-rouge">false</code>를 돌려주고, “못 보냈다”는 표시만 몇 주째 쌓이고 있었다.</p>

<p>여기서 첫 번째 교훈: <strong>“기능을 만들어달라”는 요청의 절반은 이미 있는 기능을 켜는 일</strong>이다. 코드를 먼저 읽자.</p>

<h2 id="켜기-전에-로컬에서-4번-울려보기">켜기 전에 로컬에서 4번 울려보기</h2>

<p>서버 권한이 없고 스테이징도 없으니, 켜기 전에 <strong>내 PC에 전체 스택</strong>(JDK 8 + Tomcat 9 + MariaDB)을 올렸다. 그리고 실제 알림 파이프라인을 네 가지 시나리오로 돌렸다.</p>

<table>
  <thead>
    <tr>
      <th>시나리오</th>
      <th>만든 방법</th>
      <th>결과</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>🔴 FIRING</td>
      <td>가짜 호스트에 5초마다 CPU 99% 샘플 주입 + 규칙 <code class="language-plaintext highlighter-rouge">cpu &gt; 1</code></td>
      <td>감지 → 0.5초 뒤 Slack 도착</td>
    </tr>
    <tr>
      <td>✅ RESOLVED</td>
      <td>규칙 임계값을 200으로</td>
      <td>다음 사이클에 도착</td>
    </tr>
    <tr>
      <td>🔴 DOWN</td>
      <td>샘플 주입 중단 → 2분 미수신</td>
      <td>도착</td>
    </tr>
    <tr>
      <td>✅ RECOVERED</td>
      <td>주입 재개</td>
      <td>도착</td>
    </tr>
  </tbody>
</table>

<p>여기서 작은 함정 하나. 처음엔 샘플을 한 번만 넣고 60초 배치를 기다렸는데 <strong>아무 일도 안 일어났다.</strong> 규칙의 “유지 시간”이 0초면 평가기는 <strong>최근 15초 창</strong>만 보는데, 60초 배치가 돌 때쯤엔 샘플이 이미 창 밖이었던 것이다. 수집기처럼 계속 넣어야 한다. 이 함정은 나중에 운영에서 그대로 다시 만난다(5장).</p>

<h2 id="켜는-방법을-두고--그냥-커밋하면-안-돼요">켜는 방법을 두고 — “그냥 커밋하면 안 돼요?”</h2>

<p>주소를 어디에 넣을지가 문제였다. 이 앱은 시크릿을 <code class="language-plaintext highlighter-rouge">System property &gt; 환경변수 &gt; 코드 기본값</code> 순서로 읽는다. 선택지는:</p>

<ol>
  <li><strong>코드 기본값에 박아서 커밋</strong> — 된다. 커밋하면 배포되니까. 하지만 시크릿이 SVN 히스토리에 영원히 남고, 프로젝트 문서에 “시크릿은 SVN에 넣지 않는다”가 명시돼 있고, 두 달 전 똑같은 짓을 했다가 롤백한 전례가 있었다(파트4 예정).</li>
  <li><strong>서버 환경변수</strong> — 정석. 서버 권한 있는 분께 한 줄 부탁.</li>
  <li><strong>관리 화면에서 넣게 코드 수정</strong> — 오너가 서버 권한이 없는 구조에선 사실 이게 맞는 설계지만, 문서 규칙과 충돌해 합의 필요.</li>
</ol>

<p>2번으로 요청했더니 배포 담당자가 “설정 파일(<code class="language-plaintext highlighter-rouge">jnut.properties</code>)에 넣고 알려달라”고 답했다. 그 파일은 SVN 안에 있어서 결국 1번과 비슷하게 시크릿이 저장소에 들어가는 셈이었는데, <strong>배포 담당의 지시 + webhook은 노출돼도 “그 채널에 글쓰기”만 되고 재발급으로 즉시 무효화되는 낮은 등급 시크릿</strong>이라 감수하기로 했다. 다만 이 판단은 문서에 남겼다. 나중에 누가 “왜 여기 있지?” 할 테니까.</p>

<p>(참고: 키 이름이 환경변수 방식과 설정파일 방식이 다르다 — <code class="language-plaintext highlighter-rouge">SMON_SLACK_WEBHOOK_URL</code> vs <code class="language-plaintext highlighter-rouge">smon.slack.webhookUrl</code>. 이걸 틀리면 조용히 안 읽힌다. 로컬에서 설정파일 방식으로 한 번 더 울려보고 커밋했다.)</p>

<h2 id="재시작-직후--알림-폭탄">재시작 직후 — 알림 폭탄</h2>

<p>재시작하자마자 Slack이 울리기 시작했다. <code class="language-plaintext highlighter-rouge">DOWN</code>, <code class="language-plaintext highlighter-rouge">RECOVERED</code>, <code class="language-plaintext highlighter-rouge">DOWN</code>, <code class="language-plaintext highlighter-rouge">RECOVERED</code>… 수백 건.</p>

<p>원인은 1장의 그 “못 보냈다는 표시”이다. 코드는 실패한 알림을 <strong>다음 사이클에 최대 100건씩 재시도</strong>하도록 설계돼 있고, 몇 주치 적체가 그대로 방출된 것이다. 설계상 정상이고, 사이클당 50개 이벤트(FIRING+RESOLVED 100메시지)씩 오래된 것부터 나간다.</p>

<p><strong>예상은 했는데 준비를 안 했다.</strong> “쌓인 게 쏟아질 수 있다”까진 알았지만 재시작 전에 건수를 확인하거나, 오래된 것들에 미리 “보낸 것”으로 표시하는 SQL을 같이 부탁했어야 했다. 한 줄이면 됐다:</p>

<div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">UPDATE</span> <span class="n">alert_events</span>
   <span class="k">SET</span> <span class="n">firing_notified_at</span>   <span class="o">=</span> <span class="n">COALESCE</span><span class="p">(</span><span class="n">firing_notified_at</span><span class="p">,</span> <span class="n">NOW</span><span class="p">(</span><span class="mi">3</span><span class="p">)),</span>
       <span class="n">resolved_notified_at</span> <span class="o">=</span> <span class="k">CASE</span> <span class="k">WHEN</span> <span class="k">state</span><span class="o">=</span><span class="s1">'resolved'</span> <span class="k">THEN</span> <span class="n">COALESCE</span><span class="p">(</span><span class="n">resolved_notified_at</span><span class="p">,</span> <span class="n">NOW</span><span class="p">(</span><span class="mi">3</span><span class="p">))</span> <span class="k">ELSE</span> <span class="n">resolved_notified_at</span> <span class="k">END</span>
 <span class="k">WHERE</span> <span class="n">fired_at</span> <span class="o">&lt;</span> <span class="n">NOW</span><span class="p">(</span><span class="mi">3</span><span class="p">)</span> <span class="o">-</span> <span class="n">INTERVAL</span> <span class="mi">1</span> <span class="n">HOUR</span><span class="p">;</span>
</code></pre></div></div>

<p>결국 채널 음소거 걸고 밤새 흘려보냈다. 다음날 아침엔 조용했고요.</p>

<h2 id="운영-검증--0초의-함정-다시">운영 검증 — 0초의 함정, 다시</h2>

<p>다음날 관리자 화면에서 테스트 규칙을 만들었다. <code class="language-plaintext highlighter-rouge">CPU &gt; 1</code>, 유지 시간 0초. 아무 일도 안 일어났다. 이벤트 자체가 안 생김.</p>

<p>2장의 그 함정이다. 운영 수집기는 15초보다 뜸하게 보내니, “0초 = 최근 15초 창”은 운영에선 <strong>거의 항상 비어 있는 창</strong>이다. 유지 시간을 120초로 바꾸니 바로 <code class="language-plaintext highlighter-rouge">🔴 FIRING</code> 이 도착했다.</p>

<p>이건 사용법 문제라기보다 <strong>설계 허점</strong>이다. 화면이 “0초”를 허용하는데 그 규칙은 운영에서 영영 안 울린다. 최소 창을 수집 주기 기반으로 잡거나, 화면에 경고를 띄우는 게 맞다. 개선 항목으로 남겼다.</p>

<p>그리고 하나 더 — “규칙을 삭제하면 RESOLVED가 오겠지” 했는데 <strong>안 온다.</strong> 삭제 SQL이 이벤트를 해소 처리하면서 “해소 알림도 이미 보냈음”으로 같이 표시하더라. 관리자가 직접 끈 건 알리지 않는 설계였다. 나는 함수 이름만 보고 “온다”고 말했다가 SQL을 읽고 정정했다. <strong>이름 말고 SQL을 읽자.</strong></p>

<h2 id="그런데-로컬에서-배치가-전부-멈췄다">그런데 로컬에서 배치가 전부 멈췄다</h2>

<p>시간을 2장로 되돌려서. 로컬 스택을 올렸을 때 사실 첫 반응은 “배치가 하나도 안 돈다”였다. 로그엔:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>batch execution requires a compatible application schema (reason=DRIFT)
</code></pre></div></div>

<p>이 콘솔은 앱을 기동할 때마다 “DB 스키마가 패키지된 카탈로그와 정확히 같은가”를 검사하고, 다르면(DRIFT) <strong>배치를 전부 거부</strong>한다. 검사 항목 중엔 “대시보드 위젯 설정 행 전체의 SHA-256이 이 값이어야 한다”는 핀이 있었고, 그게 어긋나 있었다.</p>

<p>누가 어긋나게 했나? <strong>2주 전의 저였다.</strong></p>

<h2 id="내가-심은-지뢰">내가 심은 지뢰</h2>

<p>2주 전에 “CPU/메모리/디스크가 85%를 넘으면 대시보드 카드를 깜빡이게 해달라”는 과제로 위젯 설정 원본(매니페스트)에 <code class="language-plaintext highlighter-rouge">displayThreshold</code> 필드를 추가해 커밋했다. 테스트도 통과했고… 라고 생각했는데, 사실 <strong>관련 유닛테스트를 안 돌렸다.</strong> 돌렸으면 바로 알았을 것이다 — 그 테스트가 정확히 “매니페스트를 바꿨으면 검증기의 해시 핀도 같이 올려라”를 검사하니까.</p>

<p>결과는 두 가지였다.</p>

<ol>
  <li><strong>운영에선 깜빡임이 켜진 적이 없었다.</strong> 매니페스트는 “원본”이고, DB 복사본은 부트스트랩 잡이 돌 때만 갱신된다. 그 잡은 8월에 한 번 돌고 삭제된 뒤였다. 아무도 눈치 못 챈 건 그동안 85%를 넘은 서버가 없었기 때문.</li>
  <li><strong>새로 설치하거나 리셋하면 배치가 전면 정지하는 지뢰</strong>가 저장소에 심겼다. 운영은 (복사본이 옛것 그대로라) 검사를 통과하고 있었을 뿐.</li>
</ol>

<p>기능이 안 켜진 것보다 2번이 무섭다. 언젠가 누가 새 서버를 세우는 날 터지는 지뢰니까.</p>

<h2 id="제거--선임의-패턴을-그대로">제거 — 선임의 패턴을 그대로</h2>

<p>이 콘솔의 마이그레이션 도구는 Flyway와 비슷하되 더 엄격하다. 마이그레이션마다 SQL과 <strong>검증기 SQL</strong>이 쌍으로 있고, 검증기는 “적용된 상태”를 증명해야 하며, 후속 마이그레이션이 이전 검증기를 “대체(supersede)”할 수 있다. 8월에 선임이 같은 상황(위젯 v3→v4)을 처리한 커밋이 있어서, 그 패턴을 그대로 따랐다.</p>

<ul>
  <li>새 마이그레이션(rank 11): DDL 없음. 부트스트랩 잡을 다시 큐에 넣는 <code class="language-plaintext highlighter-rouge">INSERT</code> 한 문장 — 운영 복사본을 원본으로 갱신하라는 지시서.</li>
  <li>새 검증기: 이전 검증기 복사 + 위젯 해시를 새 값으로. 카탈로그에 “rank 9를 대체한다”고 등록.</li>
  <li>그리고 <strong>동기화해야 할 곳이 여섯 군데</strong>: 카탈로그, 신규 설치용 스냅샷 SQL, 배치 스케줄러의 DB측 게이트(<code class="language-plaintext highlighter-rouge">이력이 정확히 N개일 때만 실행</code>), SQL을 Java 상수로 임베딩한 파일(배포기가 <code class="language-plaintext highlighter-rouge">.java</code>만 배포하니까), 스냅샷 해시 핀, 유닛테스트 4개.</li>
</ul>

<p>파일 16개. 로직 변경은 0줄. <strong>전부 “설정 추가 + 숫자 맞추기”</strong> 인데, 한 군데라도 빠지면 어딘가의 검사가 실패한다.</p>

<h2 id="운영을-내-pc에-복제하기">운영을 내 PC에 복제하기</h2>

<p>커밋하면 바로 배포되는 환경이라, 커밋 전에 <strong>운영과 똑같은 DB 상태</strong>를 로컬에 만들어 순서대로 돌려봤다.</p>

<ol>
  <li>옛 스냅샷으로 DB 생성 → <strong>옛 빌드</strong>로 Tomcat 기동 → 부트스트랩이 v4 콘텐츠 설치</li>
  <li>위젯에서 <code class="language-plaintext highlighter-rouge">displayThreshold</code> 조각을 제거하고 부트스트랩 잡을 삭제 → <strong>운영과 동일한 상태</strong>. 옛 카탈로그로 <code class="language-plaintext highlighter-rouge">status</code> → <code class="language-plaintext highlighter-rouge">compatible=true</code> (지금 운영과 같음)</li>
  <li><strong>새 빌드</strong> 배포 → <code class="language-plaintext highlighter-rouge">status</code> → rank 11 = <code class="language-plaintext highlighter-rouge">apply</code>, drift 없음 ← 운영이 배포 직후 보게 될 화면</li>
  <li><code class="language-plaintext highlighter-rouge">apply</code> → 즉시 <code class="language-plaintext highlighter-rouge">compatible=true</code>, 잡 큐잉 → 앱이 알아서 부트스트랩 → 위젯 3개에 필드 반영 → 검증 11개 전부 통과</li>
</ol>

<p>3번 단계에서 <strong>또 하나의 실수가 잡혔다.</strong> 처음엔 rank 11이 <code class="language-plaintext highlighter-rouge">apply</code>가 아니라 <code class="language-plaintext highlighter-rouge">baseline</code>(“이미 적용된 걸로 기록만”)으로 나왔다. 옛 위젯 상태인데 새 검증기가 통과한다? 파일을 다시 보니 — 해시를 <strong>설명 주석에서만</strong> 바꾸고, 실제로 <strong>비교하는 리터럴 한 줄은 옛 값 그대로</strong>였다. 유닛테스트는 주석 마커만 검사해서 통과했고요. 리터럴을 고치고, 테스트가 리터럴도 검사하도록 보강했다.</p>

<p>이 시뮬레이션이 없었으면 어떻게 됐을까 커밋 → 운영에서 rank 11이 <code class="language-plaintext highlighter-rouge">baseline</code>으로 기록 → 부트스트랩은 큐잉되지 않음 → <strong>깜빡임은 여전히 안 켜지고, 저장소는 “고쳤다”고 믿는 상태.</strong> 지뢰가 한 층 더 깊어졌을 것이다.</p>

<h2 id="커밋-그리고-회색-버튼">커밋, 그리고 회색 버튼</h2>

<p>커밋(r187)하고 관리자 화면의 “스키마 마이그레이션”에서 Apply 하면 끝… 인데 <strong>Apply 버튼이 회색</strong>이었다. 확인 문구(<code class="language-plaintext highlighter-rouge">APPLY 1 ONLINE MIGRATIONS TO &lt;db&gt; &lt;해시12자리&gt;</code>)를 정확히 붙여넣어도 회색.</p>

<p>이 화면은 <strong>최근 5분 이내 로그인</strong>이 아니면 Apply를 거부하고, 한 번 거부되면 문구를 맞춰도 버튼이 계속 비활성이다. 로그아웃 → 로그인 → Plan부터 다시 → 파란 버튼 → Apply → “스키마 마이그레이션과 독립 검증을 완료했다”.</p>

<p>마지막으로 테스트 규칙 하나를 만들어 <code class="language-plaintext highlighter-rouge">🔴 FIRING</code>이 오는지 봤다. 왔다. 이건 두 가지를 동시에 증명한다 — 배치가 살아있다는 것, 그리고 스케줄러의 XML 게이트(<code class="language-plaintext highlighter-rouge">이력 = 11</code>)까지 재시작 없이 배포됐다는 것. 후자는 이 배포 파이프라인이 “XML도 자동 반영하는가”라는 문서의 오래된 미확인 항목이었는데, 이걸로 답이 나왔다.</p>

<h2 id="마치며">마치며</h2>

<p>이틀 동안 한 일을 한 줄로 줄이면 “설정 한 줄 넣고, 파일 16개짜리 마이그레이션 하나 추가”이다. 하지만 그 사이에 있었던 것들:</p>

<ul>
  <li><strong>코드를 먼저 읽어서</strong> 만들 필요가 없는 기능을 만들지 않았고</li>
  <li><strong>로컬 전체 스택</strong>으로 켜기 전에 4번 울려봤고</li>
  <li>그 로컬이 <strong>내 지뢰를 밟아줬고</strong></li>
  <li><strong>운영을 복제한 시뮬레이션</strong>이 두 번째 실수를 잡아줬고</li>
  <li>실수 두 개를 모두 <strong>테스트가 다음번엔 잡도록</strong> 남겼다</li>
</ul>

<h3 id="교훈-요약">교훈 요약</h3>
<ol>
  <li>“기능을 만들어달라”는 요청 → 이미 있는지 코드부터. 없는 건 대개 설정 한 줄.</li>
  <li>스테이징이 없으면 <strong>내가 스테이징을 만든다.</strong> 커밋=배포인 환경에선 더더욱.</li>
  <li>매니페스트/스키마 같은 “원본”을 바꿀 땐 그걸 고정하는 <strong>검증기·테스트·게이트가 어디 있는지</strong> 먼저 찾는다. 테스트는 <strong>돌린다.</strong></li>
  <li>시뮬레이션은 “될 것 같은” 걸 확인하는 게 아니라 <strong>“안 되는 걸 찾는” 도구</strong>다. 두 번 살았다.</li>
  <li>함수 이름 말고 SQL을 읽는다. 주석 말고 리터럴을 본다.</li>
  <li>재시작 전에 “그동안 못 한 일이 쌓여 있진 않나”를 묻는다. 재시도 큐는 성실하다.</li>
</ol>

<p>다음 글은 시간을 거슬러 올라가 두 달 전 이야기 — “급하니까 시크릿을 코드에 박았다가 4일 만에 되돌린” Keycloak realm 사건이다. 이번 글의 3장에서 “전례가 있다”고 한 바로 그 일이고, 이 시스템의 “배포기는 <code class="language-plaintext highlighter-rouge">.java</code>만 옮긴다”는 전제를 얻은 사건이기도 하다.</p>]]></content><author><name>이진용</name><email>kouig14@gmail.com</email></author><category term="Java" /><category term="Tomcat" /><category term="MariaDB" /><category term="Slack" /><category term="마이그레이션" /><summary type="html"><![CDATA[앞선 리버스 프록시 두 편이 트래픽 앞단 얘기였다면, 이번엔 인수받은 사내 모니터링 콘솔에 알림을 붙이다가 겪은 일이다. 결론부터: 알림은 켜졌고, 그 과정에서 2주 전 내 커밋이 심어둔 지뢰를 발견해 제거했다. (사내 인프라 특성상 서버·도메인·사람 이름·수치 일부는 일반화했다. 시스템 이름은 “콘솔”로 부른다.)]]></summary></entry><entry><title type="html">리버스 프록시로 Keycloak 렐름별 접근제어 실전</title><link href="https://stan-dev.cloud/npm-keycloak-realm-access/" rel="alternate" type="text/html" title="리버스 프록시로 Keycloak 렐름별 접근제어 실전" /><published>2026-09-17T09:44:00+09:00</published><updated>2026-09-17T09:44:00+09:00</updated><id>https://stan-dev.cloud/npm-keycloak-realm-access</id><content type="html" xml:base="https://stan-dev.cloud/npm-keycloak-realm-access/"><![CDATA[<blockquote>
  <p><a href="/npm-reverse-proxy-realm-access/">지난 글</a>에서는 Nginx Proxy Manager(NPM)로 리버스 프록시의 기초 — 도메인 라우팅, IP 화이트리스트, 경로 단위 접근제어의 <strong>개념</strong> — 을 <code class="language-plaintext highlighter-rouge">whoami</code> 거울 백엔드로 익혔다. 이번엔 그 원리를 <strong>진짜 Keycloak</strong> 앞단에 적용하면서 만난 함정들, 그리고 서버 접근 권한 없이 <strong>실제 운영 서버의 노출 구조를 역추론</strong>한 이야기이다.</p>

  <p><em>(사내 인프라 특성상 실제 도메인/IP/렐름 이름/설정값은 일반화한 예시로 대체했다.)</em></p>
</blockquote>

<hr />

<h2 id="이번-글의-목표">이번 글의 목표</h2>

<p>지난 글 마지막에서 이렇게 예고했다.</p>

<blockquote>
  <p>“다음 글에서는 이 원리를 실제 Keycloak 앞단에 적용해, 렐름별로 내부/외부 노출을 관리하는 구성을 다뤄본다.”</p>
</blockquote>

<p>목표를 한 문장으로 정리하면:</p>

<blockquote>
  <p><strong>하나의 Keycloak(<code class="language-plaintext highlighter-rouge">auth.example.com</code>)에서, 렐름마다 내부/외부 노출을 다르게 — 리버스 프록시만으로 — 통제할 수 있는가?</strong></p>
</blockquote>

<p>Keycloak은 렐름을 URL 경로로 노출한다(<code class="language-plaintext highlighter-rouge">/realms/&lt;렐름&gt;/...</code>). 렐름이 URL에 보인다는 건 프록시가 경로를 보고 렐름별로 막고 열 수 있다는 뜻이다. 이론은 지난 글에서 <code class="language-plaintext highlighter-rouge">whoami</code>로 증명했다. 그런데 <strong>진짜 Keycloak을 붙이는 순간, <code class="language-plaintext highlighter-rouge">whoami</code>로는 절대 못 봤던 함정들이 쏟아졌다.</strong></p>

<hr />

<h2 id="왜-whoami로는-부족했나">왜 <code class="language-plaintext highlighter-rouge">whoami</code>로는 부족했나</h2>

<p><code class="language-plaintext highlighter-rouge">whoami</code>는 <strong>어떤 경로로 요청하든 200으로 자기 정보를 뱉는</strong> 백엔드이다. 학습용 거울로는 최고지만, 바로 그 성질 때문에 <strong>경로에 민감한 실제 앱의 문제를 하나도 못 보여준다.</strong> <code class="language-plaintext highlighter-rouge">/realms/foo</code>를 치든 <code class="language-plaintext highlighter-rouge">/aaa/bbb</code>를 치든 whoami는 똑같이 200이니까.</p>

<p>그래서 테스트 환경에 <strong>일회용 Keycloak</strong>을 진짜로 띄웠다. (운영과 격리된 테스트 NPM과 같은 docker-compose 네트워크에 올려, 호스트 포트 노출 없이 컨테이너 이름으로 프록시)</p>

<div class="language-yaml highlighter-rouge"><div class="highlight"><pre class="highlight"><code>  <span class="na">keycloak</span><span class="pi">:</span>
    <span class="na">image</span><span class="pi">:</span> <span class="s">quay.io/keycloak/keycloak:26.0</span>
    <span class="na">command</span><span class="pi">:</span> <span class="s">start-dev</span>          <span class="c1"># 개발 모드 (내장 H2 DB)</span>
    <span class="na">environment</span><span class="pi">:</span>
      <span class="na">KC_BOOTSTRAP_ADMIN_USERNAME</span><span class="pi">:</span> <span class="s">admin</span>
      <span class="na">KC_BOOTSTRAP_ADMIN_PASSWORD</span><span class="pi">:</span> <span class="s">admin</span>
</code></pre></div></div>

<p>NPM에서 <code class="language-plaintext highlighter-rouge">auth.test.local</code> → <code class="language-plaintext highlighter-rouge">keycloak:8080</code>으로 프록시. 그리고 첫 삽질이 곧바로 터졌다.</p>

<hr />

<h2 id="함정---custom-location이-keycloak을-통째로-죽였다">함정 ① — “Custom Location”이 Keycloak을 통째로 죽였다</h2>

<p>지난 글에서 경로 단위 제어는 NPM의 <strong>Custom Locations</strong>로 한다고 배웠다. 그대로 <code class="language-plaintext highlighter-rouge">/realms/master</code>에 IP 제한을 걸어봤다.</p>

<p>결과: <strong>로그인 페이지가 404, 그리고 호스트 전체가 죽었다.</strong> 다른 경로까지 전부 404. 프록시 설정 파일(nginx conf)이 아예 생성되지 않았다.</p>

<p>원인을 파보니 이랬다.</p>

<blockquote>
  <p>NPM의 Custom Location은 단순히 “이 경로에 규칙 추가”가 아니라, <strong>그 경로를 백엔드로 다시-프록시하는 별도 블록</strong>을 만든다. 경로가 재작성되면서, 경로에 민감한 실제 앱(Keycloak)은 엉뚱한 경로를 받아 깨지고 nginx 문법 검사도 실패한다.</p>
</blockquote>

<p><code class="language-plaintext highlighter-rouge">whoami</code>는 아무 경로나 200이라 이 재작성 부작용이 <strong>안 보였던</strong> 것이다. 실제 앱을 붙이고 나서야 드러난 맹점이었다.</p>

<p><strong>복구:</strong> 그 Custom Location을 삭제하니 conf가 다시 생성되고 200으로 돌아왔다. 그리고 여기서 중요한 디버깅 습관 하나:</p>

<blockquote>
  <p>NPM UI에서 뭘 눌렀든, <strong>진짜 결과물은 컨테이너 안의 nginx 설정 파일</strong>이다. 원인 확정은 항상 실제 생성된 conf를 직접 열어본다.</p>
</blockquote>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c"># NPM이 실제로 만든 nginx 설정을 직접 확인</span>
docker <span class="nb">exec</span> &lt;npm컨테이너&gt; sh <span class="nt">-c</span> <span class="se">\</span>
  <span class="s1">'grep -rnE "location|allow|deny|proxy_pass" /data/nginx/proxy_host/'</span>
</code></pre></div></div>

<hr />

<h2 id="올바른-방법--advanced-탭에-raw-nginx-필터만-얹기">올바른 방법 — Advanced 탭에 raw nginx, “필터만 얹기”</h2>

<p>정답은 <strong>호스트 레벨 Advanced 탭</strong>(per-location 톱니가 아니라, 호스트 전체에 raw nginx를 넣는 곳)이었다. 핵심 발상의 전환:</p>

<blockquote>
  <p><strong>다시-프록시하지 말고, 프록시는 원래대로 흘려보내되 그 위에 IP 필터만 얹는다.</strong></p>
</blockquote>

<div class="language-nginx highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">location</span> <span class="n">/realms/internal-realm</span> <span class="p">{</span>
    <span class="kn">allow</span> <span class="mi">10</span><span class="s">.0.0.0/8</span><span class="p">;</span>   <span class="c1"># 사내 대역(예시)</span>
    <span class="kn">deny</span> <span class="s">all</span><span class="p">;</span>
    <span class="kn">proxy_pass</span> <span class="s">http://keycloak:8080</span><span class="p">;</span>   <span class="c1"># ← 끝에 슬래시 없음 = 경로 보존</span>
    <span class="kn">proxy_set_header</span> <span class="s">Host</span> <span class="nv">$host</span><span class="p">;</span>
    <span class="kn">proxy_set_header</span> <span class="s">X-Forwarded-For</span> <span class="nv">$proxy_add_x_forwarded_for</span><span class="p">;</span>
    <span class="kn">proxy_set_header</span> <span class="s">X-Forwarded-Proto</span> <span class="nv">$scheme</span><span class="p">;</span>
<span class="p">}</span>
</code></pre></div></div>

<p><code class="language-plaintext highlighter-rouge">proxy_pass</code>에 <strong>경로(끝 슬래시)를 붙이지 않는 것</strong>이 포인트이다. 이러면 nginx가 경로를 자르거나 재작성하지 않고 그대로 백엔드에 넘겨, Keycloak이 안 깨진다. (이 “끝 슬래시” 얘기는 뒤에서 <strong>훨씬 무서운 형태로</strong> 다시 나온다.)</p>

<p>여기서 배운 운영 팁:</p>

<blockquote>
  <p>NPM Advanced의 nginx 문법 오류는 <strong>컨테이너 런타임 로그(<code class="language-plaintext highlighter-rouge">docker logs</code>)에 안 뜬다.</strong> 저장 직후 깨졌는지 확인하려면 <code class="language-plaintext highlighter-rouge">docker exec &lt;npm&gt; nginx -t</code>로 직접 검사해야 한다. 그리고 Advanced는 append가 아니라 <strong>전체 replace</strong> — 같은 <code class="language-plaintext highlighter-rouge">location</code>을 중복 정의하면 nginx가 거부하고 호스트가 죽는다.</p>
</blockquote>

<hr />

<h2 id="렐름별-차등-접근제어-달성">렐름별 차등 접근제어 달성</h2>

<p>Keycloak 관리 콘솔에서 렐름 두 개를 만들었다 — <code class="language-plaintext highlighter-rouge">public-realm</code>(외부 공개용), <code class="language-plaintext highlighter-rouge">internal-realm</code>(내부 전용). 그리고 Advanced 탭에 렐름별 <code class="language-plaintext highlighter-rouge">location</code> 블록을 넣고 정책을 다르게 걸었다.</p>

<ul>
  <li><code class="language-plaintext highlighter-rouge">public-realm</code> → 규칙 없음(기본 location으로 흘려 열림)</li>
  <li><code class="language-plaintext highlighter-rouge">internal-realm</code>, <code class="language-plaintext highlighter-rouge">master</code> → 사내 IP만 <code class="language-plaintext highlighter-rouge">allow</code>, 나머지 <code class="language-plaintext highlighter-rouge">deny</code></li>
</ul>

<p>응답 코드로 검증:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c"># (응답 코드만 관찰)</span>
public-realm   : 200   ← 열림
internal-realm : 403   ← 프록시가 차단
master         : 403   ← 프록시가 차단
</code></pre></div></div>

<p><strong>같은 도메인(<code class="language-plaintext highlighter-rouge">auth.test.local</code>), 렐름 경로별로 다른 노출 정책.</strong> 정책이 바뀌면 해당 렐름 location의 <code class="language-plaintext highlighter-rouge">allow</code> 목록만 갈아끼우면 내부↔외부가 토글된다. 목표였던 “리버스 프록시만으로 렐름별 내/외부 통제”를 실물로 증명한 순간이었다.</p>

<blockquote>
  <p>참고로 응답 코드 판독이 이 실습의 언어였다: <strong>403 = 프록시가 막음 / 200·302·404 = Keycloak까지 도달(안 막힘)</strong>. 그리고 로그인 화면 정적자원(<code class="language-plaintext highlighter-rouge">/resources/...</code>)은 <code class="language-plaintext highlighter-rouge">/realms</code> 밖에 있어서, <strong>렐름 경로만 막고 공용 경로는 열어둬야</strong> 허용된 사용자의 로그인 화면이 안 깨진다.</p>
</blockquote>

<hr />

<h2 id="함정---어제-만든-렐름이-사라졌다">함정 ② — “어제 만든 렐름이 사라졌다”</h2>

<p>다음 날 서버를 다시 켜니 렐름이 <strong>전부 없어져 있었다.</strong> NPM 설정(프록시 규칙)은 멀쩡한데 Keycloak 렐름만 증발.</p>

<p>원인:</p>

<blockquote>
  <p>Keycloak을 <code class="language-plaintext highlighter-rouge">start-dev</code>로 띄우면 내장 H2 데이터베이스가 <strong>컨테이너 내부</strong>(<code class="language-plaintext highlighter-rouge">/opt/keycloak/data</code>)에 저장된다. 컨테이너를 삭제/재생성하면 그 데이터가 날아간다. NPM 설정은 호스트 바인드 마운트(<code class="language-plaintext highlighter-rouge">./data</code>)라 살아남지만, Keycloak엔 볼륨을 안 잡아둔 게 화근이었다.</p>
</blockquote>

<p>수정은 간단했다. 그 데이터 경로에 named volume을 하나 붙이는 것 (메인 compose는 안 건드리고 override 파일에만 추가):</p>

<div class="language-yaml highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># docker-compose.override.yml</span>
<span class="na">services</span><span class="pi">:</span>
  <span class="na">keycloak</span><span class="pi">:</span>
    <span class="na">volumes</span><span class="pi">:</span>
      <span class="pi">-</span> <span class="s">kc-data:/opt/keycloak/data</span>
<span class="na">volumes</span><span class="pi">:</span>
  <span class="na">kc-data</span><span class="pi">:</span>
</code></pre></div></div>

<p>그리고 <strong>정말 영속되는지 증명</strong>했다. 예전에 렐름을 날렸던 바로 그 동작 — 컨테이너 완전 삭제 후 재생성:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>docker compose <span class="nb">rm</span> <span class="nt">-sf</span> keycloak     <span class="c"># 컨테이너 파괴</span>
docker compose up <span class="nt">-d</span> keycloak      <span class="c"># 재생성 (볼륨 다시 마운트)</span>
<span class="c"># → 렐름 재조회: 살아있음. 볼륨 붙이기 전이었다면 여기서 사라졌을 것.</span>
</code></pre></div></div>

<p>교훈:</p>

<blockquote>
  <p>“재부팅하면 사라진다”가 아니라 <strong>“컨테이너를 삭제/재생성하면 사라진다”</strong> 가 정확한 표현. 단순 재시작(<code class="language-plaintext highlighter-rouge">restart</code>)은 컨테이너를 유지하므로 데이터가 남는다. 개발 모드라도 데이터 경로에 볼륨만 잡으면 영속된다.</p>
</blockquote>

<hr />

<h2 id="함정--하이라이트--슬래시-하나가-접근제어를-뚫었다">함정 ③ (하이라이트) — 슬래시 하나가 접근제어를 뚫었다</h2>

<p>영속화를 끝내고 마지막 검증을 돌리는데, 이상한 게 나왔다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>internal-realm  <span class="o">(</span>슬래시 있음<span class="o">)</span>  /realms/internal-realm/   → 403   ✅ 막힘
internal-realm  <span class="o">(</span>슬래시 없음<span class="o">)</span>  /realms/internal-realm    → 301   ❌ ???
</code></pre></div></div>

<p><strong>슬래시를 뺀 요청이 403(차단)이 아니라 301(리다이렉트)로 빠져나갔다.</strong> 301이 나왔다는 건 요청이 <strong>차단되지 않고 백엔드까지 도달했다</strong>는 뜻 — 즉 <strong>접근제어가 새고 있었다.</strong></p>

<p>301이 어디로 가는지 봤더니 <code class="language-plaintext highlighter-rouge">.../realms/internal-realm/</code> — <strong>끝에 슬래시를 붙이라는 리다이렉트</strong>였다. 범인은 nginx의 <code class="language-plaintext highlighter-rouge">auto_redirect</code>였고, 이걸 이해하려면 <strong>nginx가 요청을 처리하는 순서</strong>를 알아야 한다.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>① location 매칭  →  ②rewrite  →  ③ access(allow/deny)  →  ④ content(proxy_pass)
                                       ↑
                              우리 deny 규칙은 여기서 평가된다
</code></pre></div></div>

<p>문제의 블록은 이렇게 <strong>끝에 슬래시가 붙어</strong> 있었다.</p>

<div class="language-nginx highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">location</span> <span class="n">/realms/internal-realm/</span> <span class="p">{</span>   <span class="c1"># ← 슬래시로 끝남</span>
    <span class="kn">allow</span> <span class="mi">10</span><span class="s">.0.0.0/8</span><span class="p">;</span> <span class="kn">deny</span> <span class="s">all</span><span class="p">;</span>
    <span class="kn">proxy_pass</span> <span class="s">http://keycloak:8080</span><span class="p">;</span> <span class="c1"># ← 경로 없음</span>
<span class="p">}</span>
</code></pre></div></div>

<p>이 “<strong>슬래시로 끝나는 location + 경로 없는 proxy_pass</strong>” 조합은 nginx의 <code class="language-plaintext highlighter-rouge">auto_redirect</code>를 켠다. 그러면 슬래시 없는 <code class="language-plaintext highlighter-rouge">/realms/internal-realm</code> 요청이 들어올 때, nginx는 <strong>①매칭 단계에서 곧바로 301(슬래시 붙여 다시 오라)을 쏘고 끝낸다.</strong> ③access 단계(우리 <code class="language-plaintext highlighter-rouge">deny</code>)까지 <strong>가지도 않다.</strong></p>

<ul>
  <li><code class="language-plaintext highlighter-rouge">/realms/internal-realm/</code> (슬래시 O) → 블록에 정확히 매칭 → ③access → <code class="language-plaintext highlighter-rouge">deny</code> → <strong>403</strong> ✅</li>
  <li><code class="language-plaintext highlighter-rouge">/realms/internal-realm</code> (슬래시 X) → ①에서 301 자동 리다이렉트 → <strong>deny 평가 자체가 안 일어남</strong> ❌</li>
</ul>

<p>nginx가 왜 이런 “친절”을 베푸냐면 — <code class="language-plaintext highlighter-rouge">location /app/</code>처럼 하위 경로로 앱을 서빙할 때 사용자가 <code class="language-plaintext highlighter-rouge">/app</code>으로 오면 백엔드의 상대경로 링크가 깨지니까, 슬래시를 붙여주려는 편의 기능이다. 그런데 이 좋은 의도가 <strong>하필 접근제어를 건너뛰는 부작용</strong>을 냈다.</p>

<p><strong>수정은 딱 한 글자 — location에서 끝 슬래시 제거:</strong></p>

<div class="language-nginx highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">location</span> <span class="n">/realms/internal-realm</span> <span class="p">{</span>    <span class="c1"># ← 슬래시 없음 = auto_redirect 조건 탈락</span>
    <span class="kn">allow</span> <span class="mi">10</span><span class="s">.0.0.0/8</span><span class="p">;</span> <span class="kn">deny</span> <span class="s">all</span><span class="p">;</span>
    <span class="kn">...</span>
<span class="err">}</span>
</code></pre></div></div>

<p>이제 슬래시가 없으니 301이 안 뜨고, 슬래시 없는 요청도 그냥 prefix 매칭되어 ③access까지 도달 → 403. 슬래시 유무와 무관하게 견고해졌다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>internal-realm  <span class="o">(</span>슬래시 없음<span class="o">)</span>  → 403   ✅ 고쳐짐
internal-realm  <span class="o">(</span>슬래시 있음<span class="o">)</span>  → 403   ✅
public-realm                  → 200   ✅ 그대로 열림
</code></pre></div></div>

<blockquote>
  <p><strong>한 줄 요약:</strong> 슬래시로 끝나는 proxy location은 “슬래시 붙이라”는 301을 <strong>access(deny) 단계 전에</strong> 쏴서, 슬래시 없는 URL이 접근제어를 통째로 우회한다. 렐름 ACL을 걸 땐 <strong>location에 끝 슬래시를 붙이지 마라.</strong></p>
</blockquote>

<p>이게 이번 여정에서 제일 값진 삽질이었다. <code class="language-plaintext highlighter-rouge">whoami</code>(아무 경로나 200)로는 죽었다 깨어나도 못 봤을 버그니까.</p>

<hr />

<h2 id="실전-확인--서버-권한-없이-진짜-서버는-어떻게-돼있나-정찰하기">실전 확인 — 서버 권한 없이 “진짜 서버”는 어떻게 돼있나 정찰하기</h2>

<p>여기까지는 <strong>격리된 테스트 환경</strong>에서 “된다”를 증명한 것이었다. 그럼 <strong>실제 운영 Keycloak(<code class="language-plaintext highlighter-rouge">auth.example.com</code>)은 지금 어떻게 노출돼 있을까</strong></p>

<p>문제는 내게 그 서버나 운영 프록시의 <strong>설정을 열어볼 권한이 없다</strong>는 것이었다. 가진 건 브라우저 접근과 관리 콘솔 계정뿐. 그래서 <strong>블랙박스 정찰</strong> — 밖에서 HTTP 응답 코드만 관찰해 노출 구조를 역추론했다.</p>

<p>접근제어가 IP 기반이라면, <strong>“어디서 요청하느냐”가 곧 정책 판정기</strong>이다. 그래서 같은 요청을 <strong>사내망</strong>과 <strong>외부망(휴대폰 LTE)</strong> 두 곳에서 쏴서 비교했다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">for </span>r <span class="k">in </span>master internal-realm partner-realm<span class="p">;</span> <span class="k">do
  </span><span class="nb">echo</span> <span class="nt">-n</span> <span class="s2">"</span><span class="nv">$r</span><span class="s2">: "</span><span class="p">;</span> curl <span class="nt">-s</span> <span class="nt">-o</span> /dev/null <span class="nt">-w</span> <span class="s2">"%{http_code}</span><span class="se">\n</span><span class="s2">"</span> <span class="se">\</span>
    https://auth.example.com/realms/<span class="nv">$r</span>
<span class="k">done</span>
</code></pre></div></div>

<table>
  <thead>
    <tr>
      <th>경로</th>
      <th>내부(사내망)</th>
      <th>외부(LTE)</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">/realms/master</code></td>
      <td>200</td>
      <td><strong>403</strong></td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">/realms/internal-realm</code></td>
      <td>200</td>
      <td><strong>403</strong></td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">/realms/partner-realm</code></td>
      <td>200</td>
      <td><strong>403</strong></td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">/</code> (루트)</td>
      <td>403</td>
      <td>403</td>
    </tr>
  </tbody>
</table>

<p><strong>외부에서 전부 403</strong> = 실제 <code class="language-plaintext highlighter-rouge">auth.example.com</code>은 <strong>렐름 구분 없이 통째로 내부 전용</strong>이었다. 어떤 렐름도 외부엔 열려있지 않았다. 즉 우리가 테스트로 만든 “렐름별 차등”은 <strong>현행 운영엔 아직 없는, “가능성”으로서의 미래 기능</strong>이었던 것이다.</p>

<h2 id="첫-결론--경로-단위-acl로-추정">첫 결론 — “경로 단위 ACL로 추정”</h2>

<h3 id="내부에서도-루트가-403인데">내부에서도 루트가 403인데?</h3>

<p>표를 보면 이상한 게 있다. 내부에서 렐름은 200인데 <strong>루트 <code class="language-plaintext highlighter-rouge">/</code>는 내부에서도 403</strong>이다. 내부는 열린 거 아니었나?</p>

<p>여기서 <strong>대조군</strong>을 하나 놓는 게 결정적이었다. 우리 테스트 Keycloak의 루트를 <strong>프록시 우회로 직접</strong> 찔러봤다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>curl <span class="nt">-s</span> <span class="nt">-o</span> /dev/null <span class="nt">-w</span> <span class="s2">"%{http_code}</span><span class="se">\n</span><span class="s2">"</span> http://127.0.0.1:8480/   <span class="c"># → 302</span>
</code></pre></div></div>

<p><strong>Keycloak의 루트는 원래 302(리다이렉트)를 준다.</strong> 그런데 실제 서버 루트는 403. 즉 그 403은 Keycloak이 준 게 아니라 <strong>프록시가 루트를 따로 막은 값</strong>이라는 뜻이다. (루트 <code class="language-plaintext highlighter-rouge">/</code>는 어차피 Keycloak이 서빙하는 실경로가 아니라, host/path 구분용으론 애초에 애매한 프로브였다는 점도 배웠다.)</p>

<p>여기서 조심스러운 <strong>역추론</strong>:</p>

<blockquote>
  <p>같은 내부 IP인데 <code class="language-plaintext highlighter-rouge">/realms/*</code>는 200이고 <code class="language-plaintext highlighter-rouge">/</code>는 403 → 운영 프록시는 <strong>경로를 구분해서 처리</strong>하고 있다(경로 단위 ACL로 추정). 다만 “호스트 전체를 막고 특정 경로만 여는 구조”인지, “운영 Keycloak 버전이 루트에서 자체적으로 403을 주는 구조”인지까지는 <strong>설정 원문 없이 100% 확정 불가</strong>.</p>
</blockquote>

<p>중요한 건, <strong>어느 쪽이든 결론은 안 바뀐다</strong>는 점이다: 외부에선 전부 닫힘, 내부 전용. 정찰의 목적은 달성했다.</p>

<p>이 결론으로 사수에게 보고서 초안을 냈다. 그리고 6일 뒤에 뒤집혔다.</p>

<hr />

<h2 id="6일-뒤-재검증--결론이-뒤집혔다">6일 뒤 재검증 — 결론이 뒤집혔다</h2>

<p>보고서 최종본을 쓰기 전에 같은 명령을 다시 돌렸다. 표가 달라졌다.</p>

<table>
  <thead>
    <tr>
      <th>경로</th>
      <th>1차 정찰 (내부)</th>
      <th>재검증 (내부)</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">/</code> (루트)</td>
      <td>403</td>
      <td><strong>302</strong> (→ <code class="language-plaintext highlighter-rouge">/admin/</code>)</td>
    </tr>
  </tbody>
</table>

<p>내부 루트가 403이 아니라 302였다 — 대조군에서 확인한 “Keycloak 본연의 응답” 그대로. <strong>재현이 안 된 것이다.</strong></p>

<p>그러면 1차의 403은 무엇이었나. 두 가지 가능성이 있다.</p>

<ul>
  <li>(A) 1차 측정 자체가 오류였다 (다른 네트워크, 세션 상태 등)</li>
  <li>(B) 그 사이 실제 운영 설정이 바뀌었다</li>
</ul>

<p>설정 원문과 변경 이력을 볼 권한이 없으니 (A)/(B) 확정은 불가능하다. 그러나 <strong>지금 재현되는 값은 302</strong>다. 그리고 302라면 루트도 “프록시가 따로 막은 것”이 아니라 그냥 Keycloak이 응답한 것이므로, <strong>경로 간 차등이 없다.</strong></p>

<p>새 데이터로 표를 다시 그리면 내부에선 루트(302)·모든 렐름(200)·관리콘솔(200)·관리 REST(401=인증필요)까지 전부 도달 가능하고, 외부에서만 전 경로 403이다. 즉 이 데이터가 가리키는 구조는 “경로 단위 ACL”이 아니라 <strong>호스트 단위 하나 — 통째로 내부만 허용</strong>이다.</p>

<blockquote>
  <p>블랙박스 정찰 데이터는 <strong>재현될 때까지 결론이 아니다.</strong> 1차의 추정은 논리적으론 그럴듯했지만 근거 데이터 한 줄이 재현되지 않으면서 통째로 뒤집혔다. 설정 원문을 못 볼 때는 <strong>같은 명령을 다른 시점·다른 위치에서 반복 실행해 재현성을 확인하는 것</strong>이 유일한 검증 도구다.</p>
</blockquote>

<p>그리고 정정하는 방식도 배웠다. 틀린 항목을 소리 없이 지우지 않고 “1차엔 X였는데 재검증에서 Y로 나왔고, 권한 없이는 측정 오류인지 설정 변경인지 확정 불가”까지 같이 남겼다. 정정의 근거가 있어야 다음 결론도 신뢰받는다.</p>

<h3 id="덤--과녁이-틀린-요청">덤 — 과녁이 틀린 요청</h3>

<p>표를 다시 그리다 하나 더 드러났다. 지시 예시에 “master 렐름을 외부로 열자”는 항목이 있었는데, 이 요청 자체가 과녁이 어긋나 있었다.</p>

<ul>
  <li><code class="language-plaintext highlighter-rouge">/realms/master</code>는 master 렐름의 <strong>OIDC 엔드포인트</strong>(메타데이터·공개키·로그인)다. 관리 권한이 아니라 인증 표면이고, 오히려 이 경로를 막으면 <strong>관리콘솔 로그인 자체가 깨진다</strong>(콘솔이 이 경로로 인증을 걸기 때문).</li>
  <li>실제 관리자 권한 표면은 <strong><code class="language-plaintext highlighter-rouge">/admin/*</code></strong> — 콘솔 UI(<code class="language-plaintext highlighter-rouge">/admin/master/console/</code>)와 관리 REST(<code class="language-plaintext highlighter-rouge">/admin/realms/*</code>).</li>
</ul>

<p>그래서 되물었다 — “master” 관련 요구가 실제로 <strong>관리자 접근(<code class="language-plaintext highlighter-rouge">/admin</code>)</strong>을 말하는 건지, master 렐름을 쓰는 <strong>일반 로그인</strong>을 말하는 건지. 이 되물음이 없었다면 잘못된 과녁에 정확한 화살을 쏘는 상황이 됐을 것이다.</p>

<blockquote>
  <p>이름이 비슷한 두 경로가 완전히 다른 역할일 수 있다. 정책 요청을 받을 때 “무엇을 통제하려는가”를 표면 단위로 되물어야 한다.</p>
</blockquote>

<hr />

<h2 id="두-개의-시각-콘솔논리-vs-프록시경로">두 개의 시각: 콘솔(논리) vs 프록시(경로)</h2>

<p>이 정찰에서 개념적으로 제일 크게 남은 건 <strong>두 계층을 분리해서 보는 습관</strong>이었다.</p>

<table>
  <thead>
    <tr>
      <th>관리 콘솔(논리 계층)</th>
      <th>리버스 프록시(경로 계층)</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>렐름·클라이언트·유저 목록</td>
      <td>URL 경로별 노출/차단</td>
    </tr>
    <tr>
      <td>“어떤 렐름이 존재하는가”</td>
      <td>“그 렐름 경로가 밖에서 보이는가”</td>
    </tr>
    <tr>
      <td>Keycloak admin이 봄</td>
      <td>curl 응답 코드가 봄</td>
    </tr>
  </tbody>
</table>

<p>관리자 콘솔에 로그인할 수 있어도, <strong>“이 렐름이 외부에 열렸나”는 콘솔에 안 나온다.</strong> 그건 프록시 계층의 일이니까. 반대로 프록시는 ‘렐름’이라는 개념 자체를 모르고 <strong>URL 경로만</strong> 본다. 그래서 실제 노출 지도를 그리려면 <strong>콘솔(존재 목록) + 내부/외부 블랙박스(노출 여부)</strong> 를 겹쳐 봐야 했다.</p>

<hr />

<h2 id="마치며">마치며</h2>

<p><code class="language-plaintext highlighter-rouge">whoami</code> 거울로 개념을 익히고, 진짜 Keycloak을 붙이자마자 세 번 넘어졌다.</p>

<ol>
  <li><strong>Custom Location이 경로 민감 백엔드를 죽인다</strong> → Advanced 탭에 raw nginx로 “필터만 얹기”. <code class="language-plaintext highlighter-rouge">proxy_pass</code>는 경로 없이.</li>
  <li><strong>개발 모드 데이터는 볼륨 없으면 컨테이너와 함께 사라진다</strong> → 데이터 경로에 볼륨. “재시작”과 “컨테이너 재생성”은 다르다.</li>
  <li><strong>슬래시로 끝나는 proxy location은 접근제어를 우회시킨다</strong> → 끝 슬래시 제거. nginx <code class="language-plaintext highlighter-rouge">auto_redirect</code> 301이 <code class="language-plaintext highlighter-rouge">deny</code>보다 먼저 터진다.</li>
</ol>

<p>셋 다 <code class="language-plaintext highlighter-rouge">whoami</code>로는 죽었다 깨어나도 못 봤을 버그다. 학습용 거울은 아무 경로나 200이라, 경로에 민감한 실제 앱의 문제를 하나도 보여주지 않는다.</p>

<p>그리고 서버 권한 하나 없이도 밖에서 응답 코드만 비교해 운영 노출 구조를 역추론할 수 있다는 걸 배웠다 — 다만 <strong>그 결론은 재현될 때까지 결론이 아니었다.</strong> 6일 뒤 재검증에서 뒤집힌 게 이 글에서 가장 값진 부분이다.</p>

<hr />

<h3 id="참고">참고</h3>

<ul>
  <li>대상 도구: Nginx Proxy Manager, Keycloak 26 (개발 모드, 내장 H2)</li>
  <li>배포 환경: 격리된 테스트 docker-compose (호스트 포트 노출 없이 컨테이너 이름으로 프록시)</li>
  <li>미확인 항목(서버 원문 열람 권한 확보 시 확정): 운영 프록시의 Access List 원문, 접근제어가 호스트 단위인지 경로 단위인지 최종 확정, 1차↔재검증 내부 루트 응답 불일치(403→302)의 원인</li>
</ul>]]></content><author><name>이진용</name><email>kouig14@gmail.com</email></author><category term="Keycloak" /><category term="Nginx" /><category term="NPM" /><category term="접근제어" /><category term="트러블슈팅" /><summary type="html"><![CDATA[지난 글에서는 Nginx Proxy Manager(NPM)로 리버스 프록시의 기초 — 도메인 라우팅, IP 화이트리스트, 경로 단위 접근제어의 개념 — 을 whoami 거울 백엔드로 익혔다. 이번엔 그 원리를 진짜 Keycloak 앞단에 적용하면서 만난 함정들, 그리고 서버 접근 권한 없이 실제 운영 서버의 노출 구조를 역추론한 이야기이다. (사내 인프라 특성상 실제 도메인/IP/렐름 이름/설정값은 일반화한 예시로 대체했다.)]]></summary></entry><entry><title type="html">NPM 리버스 프록시 기초</title><link href="https://stan-dev.cloud/npm-reverse-proxy-realm-access/" rel="alternate" type="text/html" title="NPM 리버스 프록시 기초" /><published>2026-08-18T14:19:00+09:00</published><updated>2026-08-18T14:19:00+09:00</updated><id>https://stan-dev.cloud/npm-reverse-proxy-realm-access</id><content type="html" xml:base="https://stan-dev.cloud/npm-reverse-proxy-realm-access/"><![CDATA[<blockquote>
  <p><strong>TL;DR</strong>
사수 과제로 Nginx Proxy Manager(NPM)를 다루며 리버스 프록시의 개념부터 접근제어 두 축(IP·경로)까지 실습했다.
whoami를 뒤에 세운 “거울 백엔드”로 두면 프록시가 뒤로 뭘 넘기는지 눈으로 보이고, <code class="language-plaintext highlighter-rouge">X-Forwarded-For</code>가 왜 필요한지 자연스럽게 이해된다.
프록시는 요청에 드러난 것만(도메인·경로·헤더) 본다. 토큰 안 정보(예: 어느 렐름에서 로그인)는 앱의 몫. 사수가 말한 “렐름 단위 접근제어의 범위”가 정확히 이 경계선이었다.</p>
</blockquote>

<hr />

<h2 id="배경">배경</h2>

<p>사수가 이런 과제를 줬다.</p>

<blockquote>
  <p>“특정 도메인으로 들어올 때 접근 방식을 NPM에서 다뤄봐. 도메인 단위가 아니라 <strong>렐름(realm) 단위로 접근제어</strong>할 수 있는 범위까지 공부해봐.”</p>
</blockquote>

<p>용어부터 막혔다. NPM? 렐름 단위 접근제어? 하나씩 풀어봐야 했다.</p>

<p>리버스 프록시를 이렇게 이해했다 — 건물 안내 데스크. 바깥에서 오는 요청은 전부 안내 데스크로 온다. 데스크는 “어느 도메인 찾나?”를 보고 뒤에 있는 진짜 서버로 안내한다. 손님(브라우저)은 뒤에 서버가 몇 대인지, 뭐가 있는지 알 필요가 없다. <strong>정문은 하나, 안내만 잘하면 됨.</strong></p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>사용자 ──"auth.example.com 달라"──▶ [안내 데스크 / 리버스 프록시] ──▶ 실제 서버
</code></pre></div></div>

<p><strong>Nginx Proxy Manager(NPM)</strong> 는 이 안내 데스크를 웹 GUI로 쉽게 만들게 해주는 도구다. nginx 설정 파일을 손으로 안 짜도, 클릭 몇 번으로 “이 도메인 오면 저 서버로” 규칙을 만들 수 있다.</p>

<blockquote>
  <p>이 글에서 남기고 싶은 건 두 가지다.</p>
  <ol>
    <li>whoami 같은 거울 백엔드는 프록시의 동작을 눈으로 보게 해준다 — 이론이 아니라 헤더로 확인하는 학습법.</li>
    <li>프록시가 “볼 수 있는 것”의 경계선이 곧 “렐름 단위 접근제어의 범위”였다. 이 경계 위에서만 프록시가 렐름을 제어할 수 있다.</li>
  </ol>
</blockquote>

<hr />

<h2 id="실습-환경--운영-서버-위에-테스트-얹기">실습 환경 — 운영 서버 위에 테스트 얹기</h2>

<p>운영 중인 서버(이미 여러 서비스가 도는)에 테스트 NPM을 설치하려니 첫 번째 교훈이 왔다.</p>

<blockquote>
  <p>NPM은 기본적으로 80/443/81 포트를 먹으려 한다. 그런데 운영 서비스가 이미 그 포트를 쓰고 있을 수 있다.</p>
</blockquote>

<p>시작은 항상 “뭘 건드리면 안 되는지” 파악부터다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c"># 실제로 호스트가 열어놓은 포트 확인</span>
<span class="nb">sudo </span>ss <span class="nt">-tlnp</span>

<span class="c"># 도커 컨테이너들의 포트 매핑 확인</span>
docker ps <span class="nt">--format</span> <span class="s2">"table {{.Names}}</span><span class="se">\t</span><span class="s2">{{.Ports}}"</span>
</code></pre></div></div>

<p>여기서 알게 된 것: <code class="language-plaintext highlighter-rouge">docker ps</code>에서 <code class="language-plaintext highlighter-rouge">9000/tcp</code>처럼 IP 없이 포트만 있는 건 컨테이너 내부 전용이라 호스트와 안 겹친다. <code class="language-plaintext highlighter-rouge">0.0.0.0:8080-&gt;8080</code>처럼 <code class="language-plaintext highlighter-rouge">0.0.0.0:</code>이 붙은 것만 실제 호스트 포트를 점유한다.</p>

<p>포트 충돌을 피하려고 테스트 NPM은 대체 포트로 격리했다.</p>

<div class="language-yaml highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="na">services</span><span class="pi">:</span>
  <span class="na">npm-test</span><span class="pi">:</span>
    <span class="na">image</span><span class="pi">:</span> <span class="s1">'</span><span class="s">jc21/nginx-proxy-manager:latest'</span>
    <span class="na">restart</span><span class="pi">:</span> <span class="s">unless-stopped</span>
    <span class="na">ports</span><span class="pi">:</span>
      <span class="pi">-</span> <span class="s1">'</span><span class="s">8090:80'</span>              <span class="c1"># 프록시 HTTP</span>
      <span class="pi">-</span> <span class="s1">'</span><span class="s">8453:443'</span>             <span class="c1"># 프록시 HTTPS</span>
      <span class="pi">-</span> <span class="s1">'</span><span class="s">127.0.0.1:8181:81'</span>    <span class="c1"># 관리자 UI — localhost만</span>
    <span class="na">volumes</span><span class="pi">:</span>
      <span class="pi">-</span> <span class="s">./data:/data</span>
      <span class="pi">-</span> <span class="s">./letsencrypt:/etc/letsencrypt</span>
</code></pre></div></div>

<p>관리자 UI(81)는 <code class="language-plaintext highlighter-rouge">127.0.0.1</code>에만 묶었다. 반쯤 설정된 프록시 관리자가 인터넷에 노출되면 안 되니까. 접속은 SSH 터널로.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c"># 내 노트북에서</span>
ssh <span class="nt">-L</span> 8181:127.0.0.1:8181 사용자@서버
<span class="c"># → 브라우저 http://localhost:8181</span>
</code></pre></div></div>

<hr />

<h2 id="whoami--프록시-뒤-서버가-보는-값의-실체">whoami — 프록시 뒤 서버가 보는 값의 실체</h2>

<p>라우팅을 연습하려면 뒤에 서버가 있어야 하는데, 진짜 앱은 복잡하고 위험하다. 그래서 연습용 백엔드로 <code class="language-plaintext highlighter-rouge">traefik/whoami</code>를 썼다.</p>

<p>whoami는 딱 한 가지만 한다. 자기한테 온 요청을 전부 화면에 그대로 뱉어준다. 리버스 프록시가 뒤로 뭘 어떻게 넘기는지 눈으로 보는 거울이다.</p>

<div class="language-yaml highlighter-rouge"><div class="highlight"><pre class="highlight"><code>  <span class="na">whoami</span><span class="pi">:</span>
    <span class="na">image</span><span class="pi">:</span> <span class="s">traefik/whoami</span>
    <span class="na">restart</span><span class="pi">:</span> <span class="s">unless-stopped</span>
</code></pre></div></div>

<p>같은 docker-compose에 넣으면 같은 네트워크라 컨테이너 이름으로 프록시할 수 있다(호스트 포트도 필요 없음). NPM에서 Proxy Host 하나 만들고:</p>

<ul>
  <li>Domain: <code class="language-plaintext highlighter-rouge">whoami.test.local</code></li>
  <li>Forward Hostname: <code class="language-plaintext highlighter-rouge">whoami</code>, Port: <code class="language-plaintext highlighter-rouge">80</code></li>
</ul>

<p>요청을 쏴본다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>curl <span class="nt">-H</span> <span class="s2">"Host: whoami.test.local"</span> http://127.0.0.1:8090
</code></pre></div></div>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Hostname: 322fcb1f40dc
RemoteAddr: 172.20.0.2:38868      ← whoami가 본 "날 부른 사람" = 프록시(NPM)
Host: whoami.test.local           ← 프록시가 라우팅 판단에 쓴 도메인
X-Forwarded-For: 172.20.0.1       ← 프록시가 "원래 요청자는 이 IP였어"라고 적어준 값
X-Forwarded-Proto: http
</code></pre></div></div>

<p>여기서 뒤에 나올 접근제어의 핵심이 하나 나온다.</p>

<blockquote>
  <p>리버스 프록시 뒤의 서버는 <code class="language-plaintext highlighter-rouge">RemoteAddr</code>로 원래 요청자가 아니라 프록시를 본다. 진짜 요청자 IP는 프록시가 <code class="language-plaintext highlighter-rouge">X-Forwarded-For</code> 헤더에 담아 넘겨준다.</p>
</blockquote>

<p>whoami를 하나 더(<code class="language-plaintext highlighter-rouge">whoami2</code>) 띄우고 도메인을 갈라봤다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>curl <span class="nt">-s</span> <span class="nt">-H</span> <span class="s2">"Host: whoami.test.local"</span>  http://127.0.0.1:8090 | <span class="nb">grep </span>Hostname
<span class="c"># Hostname: 322fcb1f40dc</span>
curl <span class="nt">-s</span> <span class="nt">-H</span> <span class="s2">"Host: whoami2.test.local"</span> http://127.0.0.1:8090 | <span class="nb">grep </span>Hostname
<span class="c"># Hostname: 7d6be2c4c16d   ← 다른 컨테이너</span>
</code></pre></div></div>

<p>같은 정문(8090)으로 들어왔는데 도메인만 보고 서로 다른 뒷단으로 갈렸다. 이게 리버스 프록시의 심장이다.</p>

<hr />

<h2 id="커스텀-에러-페이지--advanced-탭의-첫-사용">커스텀 에러 페이지 — Advanced 탭의 첫 사용</h2>

<p>NPM의 Advanced(고급) 설정 — 이 버전에선 톱니바퀴 아이콘 — 에 nginx 설정을 직접 넣어 커스텀 에러 페이지를 만들 수 있다.</p>

<div class="language-nginx highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">error_page</span> <span class="mi">502</span> <span class="mi">503</span> <span class="mi">504</span> <span class="n">/maintenance_50x.html</span><span class="p">;</span>

<span class="k">location</span> <span class="p">=</span> <span class="n">/maintenance_50x.html</span> <span class="p">{</span>
    <span class="kn">internal</span><span class="p">;</span>
    <span class="kn">default_type</span> <span class="nc">text/html</span><span class="p">;</span>
    <span class="kn">return</span> <span class="mi">200</span> <span class="s">'&lt;h1&gt;서비스</span> <span class="s">점검</span> <span class="s">중이다&lt;/h1&gt;'</span><span class="p">;</span>
<span class="p">}</span>
</code></pre></div></div>

<p>뒷단 컨테이너를 <code class="language-plaintext highlighter-rouge">docker stop</code> 하고 요청하면, 기본 502 대신 이 페이지가 뜬다. Advanced 탭은 뒤에서 나올 경로 단위 접근제어에도 쓰이는 도구라 손에 익혀둘 가치가 있었다.</p>

<hr />

<h2 id="접근제어-두-축--ipaccess-list-vs-경로custom-location">접근제어 두 축 — IP(Access List) vs 경로(Custom Location)</h2>

<p>접근제어에는 축이 두 개다.</p>

<p><strong>축 하나. IP 화이트리스트 (Access List)</strong></p>

<p>NPM의 Access List는 손님 명단이다. 명단에 있는 IP만 통과, 나머지는 403.</p>

<p>명단을 만들고 Proxy Host에 붙인 뒤, 응답 코드만 확인해봤다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>curl <span class="nt">-s</span> <span class="nt">-o</span> /dev/null <span class="nt">-w</span> <span class="s2">"%{http_code}</span><span class="se">\n</span><span class="s2">"</span> <span class="nt">-H</span> <span class="s2">"Host: whoami.test.local"</span> http://127.0.0.1:8090
<span class="c"># 명단에 없을 때 → 403</span>
<span class="c"># 우리 IP를 명단에 추가하면 → 200</span>
</code></pre></div></div>

<p>403 → 200. 이게 “특정 IP만 들여보낸다”의 실체다. 사수 과제의 “특정 외부만 접근 가능하게”가 바로 이거였다.</p>

<p><strong>축 둘. 경로(렐름) 단위 — 이번 여정의 하이라이트</strong></p>

<p>여기서 과제의 진짜 목표가 나온다. 지금까지는 도메인 전체를 열고 닫았다. 이번엔 같은 도메인 안에서 경로별로 다르게 제어한다.</p>

<p>왜 “렐름 단위”라는 말이 나왔을까. Keycloak 같은 인증 서버는 렐름(realm, 인증 경계 단위)을 URL 경로로 노출한다.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>auth.example.com/realms/internal-realm/...
auth.example.com/realms/public-realm/...
                   └── 렐름 이름이 경로에 그대로 박혀있음
</code></pre></div></div>

<p>렐름이 URL에 보인다는 건, 리버스 프록시가 경로를 보고 렐름별로 접근제어를 걸 수 있다는 뜻이다. NPM의 Custom Locations로 특정 경로에만 규칙을 건다.</p>

<div class="language-nginx highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># /realms/public-realm → 아무나 (규칙 없음, 기본 열림)</span>
<span class="c1"># /realms/internal-realm → 사내 IP만</span>
<span class="k">location</span> <span class="n">/realms/internal-realm</span> <span class="p">{</span>
    <span class="kn">allow</span> <span class="mi">10</span><span class="s">.0.0.0/8</span><span class="p">;</span>   <span class="c1"># 사내 대역(예시)</span>
    <span class="kn">deny</span> <span class="s">all</span><span class="p">;</span>
    <span class="c1"># ... proxy_pass ...</span>
<span class="p">}</span>
</code></pre></div></div>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>curl <span class="nt">-s</span> <span class="nt">-o</span> /dev/null <span class="nt">-w</span> <span class="s2">"public : %{http_code}</span><span class="se">\n</span><span class="s2">"</span> ... /realms/public-realm
<span class="c"># public : 200</span>
curl <span class="nt">-s</span> <span class="nt">-o</span> /dev/null <span class="nt">-w</span> <span class="s2">"internal: %{http_code}</span><span class="se">\n</span><span class="s2">"</span> ... /realms/internal-realm
<span class="c"># internal: 403  (사내 IP 아님)</span>
</code></pre></div></div>

<p>같은 도메인인데 경로에 따라 한쪽은 열리고 한쪽은 막혔다. 이게 경로(렐름) 단위 접근제어다. 보안 정책이 바뀌어 “이 렐름은 이제 외부에도 열자”가 되면, 해당 경로의 Access List만 바꾸면 된다.</p>

<blockquote>
  <p>참고: 여기까진 whoami 백엔드로 “개념”만 증명한 것. 실제 Keycloak을 붙였을 땐 이 Custom Location 방식이 오히려 앱을 죽였다. 그 이야기는 다음 편(Keycloak 실전 삽질 3종 편)에서.</p>
</blockquote>

<hr />

<h2 id="진짜로-배운-것">진짜로 배운 것</h2>

<p><strong>첫째.</strong> 거울 백엔드는 프록시의 “행동”을 이론이 아니라 헤더로 보게 해준다.
whoami 같은 백엔드가 없었다면 <code class="language-plaintext highlighter-rouge">X-Forwarded-For</code>가 왜 필요한지, <code class="language-plaintext highlighter-rouge">RemoteAddr</code>가 왜 프록시로 보이는지, 도메인 라우팅이 실제로 어떻게 갈리는지 — 다 문서로만 읽고 넘어갔을 것이다. 거울 백엔드는 학습 도구지만 나중에 실제 앱에서 헤더 문내가 터졌을 때도 첫 진단 도구가 된다.</p>

<p><strong>둘째.</strong> 프록시가 볼 수 있는 것의 경계선이 곧 접근제어의 범위다.</p>

<table>
  <thead>
    <tr>
      <th>프록시가 보는 것</th>
      <th>프록시가 못 보는 것</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>도메인 (Host 헤더)</td>
      <td>“이 사용자가 어느 렐름에서 로그인했나”</td>
    </tr>
    <tr>
      <td>경로 (URL)</td>
      <td>→ 토큰(JWT) 안에 있음</td>
    </tr>
    <tr>
      <td>IP, 헤더</td>
      <td>→ 앱만 안다</td>
    </tr>
  </tbody>
</table>

<p>그래서 렐름 단위 제어는 두 경우로 갈린다.</p>

<ul>
  <li><strong>인증 서버 자신(<code class="language-plaintext highlighter-rouge">auth.example.com</code>)</strong> — 렐름이 URL 경로에 드러남 → 프록시가 제어 가능</li>
  <li><strong>앱(<code class="language-plaintext highlighter-rouge">app.example.com</code>)</strong> — 어느 렐름에 붙는지는 앱의 OIDC 설정과 토큰 안에 있음 → URL만으론 구분 불가 → 프록시가 아니라 앱이 처리</li>
</ul>

<p>사수가 말한 “렐름 단위로 접근제어할 수 있는 범위”가 정확히 이 경계선이었다. 굳이 앱 도메인에서 프록시가 렐름을 보게 하려면 토큰을 까서 읽는 <code class="language-plaintext highlighter-rouge">auth_request</code> / oauth2-proxy 같은 훨씬 무거운 세팅이 필요하고, 보통은 앱이 직접 한다.</p>

<hr />

<h2 id="체크리스트--다음번-리버스-프록시-붙일-때">체크리스트 — 다음번 리버스 프록시 붙일 때</h2>

<ol>
  <li><strong>운영 서버 위에 테스트 얹을 땐 포트 충돌부터</strong> — <code class="language-plaintext highlighter-rouge">sudo ss -tlnp</code>, <code class="language-plaintext highlighter-rouge">docker ps</code>로 확인. <code class="language-plaintext highlighter-rouge">0.0.0.0:</code> 붙은 것만 호스트 점유.</li>
  <li><strong>관리자 UI는 localhost + SSH 터널</strong> — 반쯤 설정된 프록시 관리자가 인터넷에 노출되면 안 된다.</li>
  <li><strong>거울 백엔드(<code class="language-plaintext highlighter-rouge">traefik/whoami</code>)를 붙여라</strong> — 프록시가 뒤로 뭘 넘기는지 헤더로 보인다. 헤더 문제 첫 진단 도구.</li>
  <li><strong>접근제어 두 축을 분리해서 생각한다</strong> — 도메인 전체는 Access List(IP), 경로별은 Custom Location. 어느 축으로 할지 먼저 정한다.</li>
  <li><strong>프록시가 볼 수 있는 것만 프록시가 제어한다</strong> — 토큰 안 정보는 앱의 몫. 렐름이 URL에 있으면 O, 토큰에 있으면 X.</li>
</ol>

<hr />

<h2 id="참고">참고</h2>

<ul>
  <li>대상 도구: Nginx Proxy Manager(NPM, <code class="language-plaintext highlighter-rouge">jc21/nginx-proxy-manager</code>), <code class="language-plaintext highlighter-rouge">traefik/whoami</code></li>
  <li>배포 환경: 운영 서버 위에 테스트 컨테이너 격리 (포트 8090/8453/8181)</li>
  <li>후속 예고: 이 원리를 실제 Keycloak 앞단에 적용한 삽질 3종(Custom Location 죽음·볼륨 소멸·슬래시 우회) — Keycloak 실전 삽질 3종 편에서.</li>
</ul>]]></content><author><name>이진용</name><email>kouig14@gmail.com</email></author><category term="Nginx" /><category term="NPM" /><category term="Docker" /><category term="리버스프록시" /><summary type="html"><![CDATA[TL;DR 사수 과제로 Nginx Proxy Manager(NPM)를 다루며 리버스 프록시의 개념부터 접근제어 두 축(IP·경로)까지 실습했다. whoami를 뒤에 세운 “거울 백엔드”로 두면 프록시가 뒤로 뭘 넘기는지 눈으로 보이고, X-Forwarded-For가 왜 필요한지 자연스럽게 이해된다. 프록시는 요청에 드러난 것만(도메인·경로·헤더) 본다. 토큰 안 정보(예: 어느 렐름에서 로그인)는 앱의 몫. 사수가 말한 “렐름 단위 접근제어의 범위”가 정확히 이 경계선이었다.]]></summary></entry><entry><title type="html">TabbyML 사내 파일럿 온보딩 가이드를 쓰며</title><link href="https://stan-dev.cloud/tabbyml-onboarding-guide/" rel="alternate" type="text/html" title="TabbyML 사내 파일럿 온보딩 가이드를 쓰며" /><published>2026-08-12T23:10:00+09:00</published><updated>2026-08-12T23:10:00+09:00</updated><id>https://stan-dev.cloud/tabbyml-onboarding-guide</id><content type="html" xml:base="https://stan-dev.cloud/tabbyml-onboarding-guide/"><![CDATA[<blockquote>
  <p><strong>TL;DR</strong>
사내 파일럿용 TabbyML(오픈소스 코드 자동완성 서버) 온보딩 가이드를 쓰며 관찰한 두 가지.
하나. 예측 가능한 함정은 트러블슈팅 섹션이 아니라 설치 단계 인라인에 못 박아야 한다. 트러블슈팅까지 가기 전에 이미 함정을 만나고, 그 순간 도구 자체에 대한 신뢰가 무너지기 때문이다.
둘. 성능이 아직 부족한 도구는 “왜 느린지”의 원인 층(CPU 추론 → GPU 전환 예정 등)까지 같이 밝혀야 피드백이 “느려서 못 씀”이 아니라 “GPU 전까진 Chat 위주로 판단” 같은 조건부 판단으로 돌아온다.</p>
</blockquote>

<hr />

<h2 id="배경">배경</h2>

<p>사내에 TabbyML(오픈소스 코드 자동완성 서버)을 올려두고 파일럿 테스터 한두 명한테 먼저 써보게 하려고 온보딩 가이드를 쓰게 됐다. 정식 롤아웃 전 단계라 문서의 목적은 “설치 순서 안내”보다는 “이 시점에 이 도구를 처음 붙이는 사람이 헤매지 않게 하는 것”에 가까웠다.</p>

<p>가이드를 다 쓰고 보니 절반이 “이걸 조심해야 한다”였다. 이 글은 그 관찰에 대한 짧은 기록이다.</p>

<hr />

<h2 id="왜-함정-경고가-설치-순서보다-먼저-오는가">왜 함정 경고가 설치 순서보다 먼저 오는가</h2>

<p>정식 릴리스가 된 상용 도구라면 설치 순서만 잘 적어두면 대부분 문제가 없다. 그런데 파일럿 단계의 오픈소스는 상황이 다르다. 실제로 이번 TabbyML 온보딩에서 명시적으로 걸린 함정이 두 개 있었다.</p>

<p><strong>함정 하나. <code class="language-plaintext highlighter-rouge">config.toml</code>이 아닌 VS Code <code class="language-plaintext highlighter-rouge">settings.json</code>에 서버 주소를 넣으면 안 먹힌다.</strong></p>

<p>VS Code에서 확장을 설치하면 자연스럽게 <code class="language-plaintext highlighter-rouge">settings.json</code>을 열어서 <code class="language-plaintext highlighter-rouge">tabby.endpoint</code> 같은 키를 넣게 된다. 그런데 실제로 유효한 위치는 <code class="language-plaintext highlighter-rouge">~/.tabby-client/agent/config.toml</code>이다. 확장 설정과 에이전트 설정이 파일이 다르고, 확장 UI에서 이 사실을 안내해주지 않는다.</p>

<p><strong>함정 둘. 서버 웹 UI가 알려주는 Endpoint URL이 잘못 나온다.</strong></p>

<p>TabbyML 서버(내부망 <code class="language-plaintext highlighter-rouge">192.168.x.x:8080</code>)에 로그인하면 “Endpoint URL / Token” 페이지가 뜬다. 여기서 표시되는 Endpoint URL이 <code class="language-plaintext highlighter-rouge">localhost:8080</code>으로 나온다. 서버 자체가 자기 자신의 외부 접근 주소를 모르기 때문에 벌어지는 일인데, 사용자 입장에서는 UI가 알려주는 값을 그대로 복사해서 붙일 가능성이 높다. 그러면 당연히 접속이 안 된다.</p>

<p>둘 다 “설치는 다 했는데 왜 안 되지?” 단계에서 만나는 함정이다. 온보딩 문서에 미리 못 박아두지 않으면 몇 번 헤매고 나서 물어보게 된다. 그런데 그 헤매는 순간이 곧 도구 자체에 대한 신뢰가 흔들리는 지점이기도 하다.</p>

<p>그래서 이번 가이드에서는 설치 3단계 안에 두 함정 모두를 인라인으로 붙였다. 원문에서 뽑아보면 이런 식이다.</p>

<div class="language-markdown highlighter-rouge"><div class="highlight"><pre class="highlight"><code>【설치 2단계 · 서버 정보 입력】
<span class="p">   -</span> 경로: <span class="sb">`~/.tabby-client/agent/config.toml`</span>
<span class="p">   -</span> ⚠️ VS Code <span class="sb">`settings.json`</span>의 <span class="sb">`tabby.endpoint`</span>가 아님 — 여기 넣으면 안 먹힘
</code></pre></div></div>

<div class="language-markdown highlighter-rouge"><div class="highlight"><pre class="highlight"><code>【토큰 발급 페이지】
<span class="p">   -</span> ⚠️ 이 페이지에 표시되는 Endpoint URL은 <span class="sb">`localhost:8080`</span>으로 <span class="gs">**잘못**</span> 나옴
     → 무시하고 위 <span class="sb">`192.168.x.x:8080`</span> 사용
</code></pre></div></div>

<p>두 경고 모두 트러블슈팅 섹션으로 빼놓지 않은 이유는 하나다. 함정을 만난 사람이 트러블슈팅 페이지를 열어보는 시점은 이미 늦기 때문이다.</p>

<hr />

<h2 id="성능-아직-부족한-도구는-왜-느린지까지-같이-말한다">성능 아직 부족한 도구는 “왜 느린지”까지 같이 말한다</h2>

<p>지금 TabbyML 자동완성은 후보 하나당 1초 안팎, 세 개면 3초 가까이 걸린다. GitHub Copilot을 쓰던 사람 입장에서는 “얘 뭔가 고장난 거 아닌가” 싶을 속도다.</p>

<p>여기서 그냥 “속도가 좀 느립니다” 정도로 언급하고 넘어가면, 도구가 원래 이 정도인 줄 알게 된다. 그러면 파일럿 피드백이 “느려서 못 쓰겠다”로 수렴한다. 실제로는 추론을 GPU로 옮기면 개선될 여지가 있는 상태인데도 그렇다.</p>

<p>그래서 가이드 3장에 이 항목을 넣었다.</p>

<blockquote>
  <p>지금 자동완성이 <strong>느리게</strong> 느껴질 수 있음 (후보 하나당 ~1초, 3개면 3초 가량) → 고장 아니고, 백엔드 서버가 현재 CPU로 추론 중이라 그럼. GPU 전환 후 개선 예정.</p>

  <p>Chat 기능은 상대적으로 빠름 (2초 이내) — 코드 설명, 질문 등에 먼저 써보길 추천</p>
</blockquote>

<p>두 가지를 같이 넣었다.</p>

<ul>
  <li>지금 느린 이유는 CPU 추론이고, GPU로 옮기면 개선된다 (한계의 층을 밝힘)</li>
  <li>상대적으로 잘 돌아가는 건 Chat이니 거기부터 써보라 (기대치 재조정)</li>
</ul>

<p>이렇게 하면 피드백이 “속도가 안 나온다”에서 “GPU 붙기 전까진 Chat 위주로 쓰겠다” 정도로 방향이 잡힌다. 도구 자체에 대한 최종 판단은 GPU 전환 이후로 자연스럽게 미뤄진다.</p>

<hr />

<h2 id="문서-전체-골격">문서 전체 골격</h2>

<p>가이드 목차는 이렇게 갔다.</p>

<ol>
  <li>설치 &amp; 연결 (3단계) — 함정 두 개를 인라인 경고로 붙임</li>
  <li>토큰 발급 방법 — 잘못된 Endpoint URL 함정을 한 번 더 반복</li>
  <li>미리 알아두면 좋은 것 — 성능 기대치 조정, Chat 우선 추천, 캐시 재생 현상</li>
  <li>피드백 요청 (사용 3~5일 후) — 5개 항목 (속도·품질·Chat·설정 난이도·기타)</li>
</ol>

<p>문서 전체 길이는 A4 한 장 안쪽. 파일럿 단계에서는 이 정도가 균형점이다 — 더 길면 안 읽고, 더 짧으면 함정에 빠진다.</p>

<hr />

<h2 id="참고">참고</h2>

<ul>
  <li>대상 도구: TabbyML (오픈소스 코드 자동완성 서버, 자체 호스팅)</li>
  <li>배포 환경: 사내 서버, 추론은 별도 LLM 서버에 위임(당시 CPU 추론)</li>
  <li>후속: 서버 구축 과정(임베딩 CUDA 크래시 · 자동완성 안 뜨는 원인 4겹)은 다음 두 편에</li>
</ul>]]></content><author><name>이진용</name><email>kouig14@gmail.com</email></author><category term="TabbyML" /><category term="문서화" /><summary type="html"><![CDATA[TL;DR 사내 파일럿용 TabbyML(오픈소스 코드 자동완성 서버) 온보딩 가이드를 쓰며 관찰한 두 가지. 하나. 예측 가능한 함정은 트러블슈팅 섹션이 아니라 설치 단계 인라인에 못 박아야 한다. 트러블슈팅까지 가기 전에 이미 함정을 만나고, 그 순간 도구 자체에 대한 신뢰가 무너지기 때문이다. 둘. 성능이 아직 부족한 도구는 “왜 느린지”의 원인 층(CPU 추론 → GPU 전환 예정 등)까지 같이 밝혀야 피드백이 “느려서 못 씀”이 아니라 “GPU 전까진 Chat 위주로 판단” 같은 조건부 판단으로 돌아온다.]]></summary></entry><entry><title type="html">Sentry Session Replay 용량 실측</title><link href="https://stan-dev.cloud/sentry-replay-storage/" rel="alternate" type="text/html" title="Sentry Session Replay 용량 실측" /><published>2026-07-30T21:00:00+09:00</published><updated>2026-07-30T21:00:00+09:00</updated><id>https://stan-dev.cloud/sentry-replay-storage</id><content type="html" xml:base="https://stan-dev.cloud/sentry-replay-storage/"><![CDATA[<blockquote>
  <p><strong>TL;DR</strong>
자체 호스팅 Sentry의 Session Replay 데이터가 서버 디스크에서 얼마나 잡아먹는지 실측.
관련 DB 테이블(<code class="language-plaintext highlighter-rouge">replays_replayrecordingsegment</code>, <code class="language-plaintext highlighter-rouge">sentry_file</code>)은 스키마만 있고 row가 0건 — 이 배포에서는 리플레이가 Django ORM을 안 거치고 파일시스템에 직접 저장된다. DB 쿼리는 시간 낭비였다.
폴더는 UUID 그대로 샤딩(체크섬 샤딩 아님)이라 UI의 리플레이 ID로 바로 잡을 수 있다. 리플레이 1개당 평균 약 213KB, 총 10.9MB / 66GB 여유 — 지금 페이스로는 걱정할 단계 아님.</p>
</blockquote>

<hr />

<h2 id="배경">배경</h2>

<p>프론트팀에서 “에러마다 리플레이 영상이 남는데 얼마나 용량을 잡아먹는지” 물어봤다. 특정 프로젝트에 리플레이가 대략 50개 쌓여 있는 상황이었다. 답을 주려고 DB부터 뒤졌는데, 스키마는 있고 데이터가 없어서 잠깐 헷갈렸다. 결론은 “이 배포에서는 리플레이가 Django ORM을 안 거치고 파일시스템에 직접 저장된다”였다.</p>

<hr />

<h2 id="시도-1-db-쿼리--스키마는-있는데-데이터가-없었다">시도 1: DB 쿼리 — 스키마는 있는데 데이터가 없었다</h2>

<p>리플레이 관련 테이블부터 찾아봤다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>docker compose <span class="nb">exec </span>postgres psql <span class="nt">-U</span> postgres <span class="nt">-d</span> postgres
</code></pre></div></div>

<p>(참고로 DB 이름이 <code class="language-plaintext highlighter-rouge">sentry</code>가 아니라 <code class="language-plaintext highlighter-rouge">postgres</code>였다. <code class="language-plaintext highlighter-rouge">-l</code>로 데이터베이스 목록을 먼저 확인해야 한다.)</p>

<p>관련 테이블은 이렇게 나왔다.</p>

<ul>
  <li><code class="language-plaintext highlighter-rouge">replays_replayrecordingsegment</code></li>
  <li><code class="language-plaintext highlighter-rouge">sentry_file</code></li>
  <li><code class="language-plaintext highlighter-rouge">sentry_fileblob</code></li>
</ul>

<p>각각 <code class="language-plaintext highlighter-rouge">SELECT count(*)</code> 걸어봤는데 <strong>세 테이블 모두 row가 0건</strong>이었다. 그런데 실제로 Sentry UI에서 리플레이는 정상적으로 재생됐다. 이 시점에 좀 헷갈렸다 — 재생이 되면 데이터는 어딘가 있어야 한다.</p>

<p>결론은 명확했다. 이 Sentry 배포는 리플레이 레코딩을 Django File ORM을 경유하지 않고 스토리지에 직접 쓰고 읽는 구조였다. <code class="language-plaintext highlighter-rouge">sentry_file</code> 테이블은 다른 종류의 파일(첨부, 아바타 등)용이지 리플레이 데이터를 여기서 트래킹하지 않는다.</p>

<blockquote>
  <p>DB 스키마가 있다고 실제로 그 경로로 데이터가 흐르는 건 아니다. “스키마 존재”와 “쓰기 경로”는 다른 층의 정보다.</p>
</blockquote>

<p>이 층에서 뽑을 수 없다고 판단하고 파일시스템으로 넘어갔다.</p>

<hr />

<h2 id="시도-2-파일시스템-직접-확인--uuid-패턴을-찾기">시도 2: 파일시스템 직접 확인 — UUID 패턴을 찾기</h2>

<p>Sentry 데이터 볼륨의 마운트 위치부터 확인했다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>docker volume inspect sentry-data <span class="nt">--format</span> <span class="s1">'{{ .Mountpoint }}'</span>
<span class="c"># /var/lib/docker/volumes/sentry-data/_data</span>
</code></pre></div></div>

<p>이 안에 <code class="language-plaintext highlighter-rouge">/files/</code> 폴더가 있고, 그 밑이 여러 레벨의 샤딩 폴더로 나뉜다. 그런데 여기에 함정이 하나 있었다.</p>

<p><strong>리플레이 폴더는 UUID 패턴(<code class="language-plaintext highlighter-rouge">/data/files/&lt;샤드&gt;/&lt;서브샤드&gt;/&lt;UUID&gt;/&lt;세그먼트번호&gt;</code>), 일반 파일 폴더는 체크섬 샤딩 패턴(2자/4자/나머지).</strong></p>

<p>처음에 <code class="language-plaintext highlighter-rouge">/data/files/7a/e111/...</code> 폴더를 발견했을 때 이게 리플레이인 줄 알았다. 그런데 <code class="language-plaintext highlighter-rouge">file</code> 명령으로 내용물을 확인해보니 <strong>1024x1024 PNG 이미지</strong>였다. 생성 날짜도 리플레이 테스트 훨씬 이전이었다. 조직/프로젝트 아바타로 추정되는 파일이었고, 이건 체크섬 샤딩(2자/4자) 규칙을 따르고 있어서 리플레이와는 완전히 다른 경로였다.</p>

<p>실제 리플레이 폴더는 UI에 표시되는 리플레이 ID를 그대로 폴더 이름으로 쓰고 있었다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nb">sudo </span>find /var/lib/docker/volumes/sentry-data/_data/files <span class="nt">-type</span> d <span class="nt">-iname</span> <span class="s2">"&lt;리플레이ID&gt;*"</span>
<span class="nb">sudo du</span> <span class="nt">-sb</span> &lt;나온 경로&gt;   <span class="c"># 리플레이 하나의 정확한 크기</span>
</code></pre></div></div>

<p>UI의 리플레이 ID(예: <code class="language-plaintext highlighter-rouge">7a7fbf98...</code>)를 그대로 <code class="language-plaintext highlighter-rouge">find</code>에 넣으면 폴더가 바로 잡힌다. 그 안에 세그먼트 번호로 파일이 몇 개 들어있고, 세션 길이(이벤트/브레드크럼 개수)에 비례해서 크기가 바뀐다.</p>

<blockquote>
  <p>폴더 이름 패턴이 스토리지 종류를 알려준다. UUID 형태면 리플레이/이벤트, 짧은 문자열의 계층 샤딩이면 체크섬 기반 일반 파일. <code class="language-plaintext highlighter-rouge">file</code> 명령으로 내용물까지 한 번 확인해두면 오판이 없다.</p>
</blockquote>

<hr />

<h2 id="샘플-측정-결과">샘플 측정 결과</h2>

<p>이번 배포에서는 리플레이가 전부 <code class="language-plaintext highlighter-rouge">/data/files/90/</code> 밑에 몰려 있었다(다른 샤드는 없었음). 그래서 이 폴더 전체를 재면 됐다.</p>

<table>
  <thead>
    <tr>
      <th>리플레이</th>
      <th>세그먼트 수</th>
      <th>크기</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>새로 만든 테스트 리플레이</td>
      <td>4개</td>
      <td>~49.5KB</td>
    </tr>
    <tr>
      <td>기존 리플레이 A</td>
      <td>19개</td>
      <td>~93.1KB</td>
    </tr>
    <tr>
      <td>기존 리플레이 B</td>
      <td>?</td>
      <td>~156.1KB</td>
    </tr>
  </tbody>
</table>

<p>세션 길이(이벤트/브레드크럼 개수)에 비례해서 커진다.</p>

<p>전체 합계.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>files/90 폴더 전체: 10,881,645 bytes
÷ 51개 (기존 50개 + 테스트 1개)
= 리플레이 1개당 평균 약 213KB
</code></pre></div></div>

<ul>
  <li>리플레이 총량: 약 10.9MB</li>
  <li>서버 전체 디스크 여유: 66GB (98GB 중 28GB 사용)</li>
  <li><strong>결론: 지금 페이스로는 용량 걱정할 단계 전혀 아님</strong></li>
</ul>

<hr />

<h2 id="참고">참고</h2>

<ul>
  <li>환경: self-hosted Sentry (Docker Compose, Postgres + 파일 스토리지)</li>
  <li>볼륨: <code class="language-plaintext highlighter-rouge">sentry-data</code>, <code class="language-plaintext highlighter-rouge">sentry-seaweedfs</code></li>
  <li>이 글의 결론은 배포 시점 기준. Sentry 버전이나 스토리지 백엔드 설정이 바뀌면 실제 쓰기 경로가 달라질 수 있음</li>
  <li>선행 글: <a href="/sentry-csrf-reverse-proxy/">CSRF 6단계</a> · <a href="/sentry-slack-integration/">Slack 통합</a> · <a href="/sentry-permission-model/">권한 모델</a> · <a href="/sentry-ingest-exposure/">ingest 외부 노출</a></li>
</ul>]]></content><author><name>이진용</name><email>kouig14@gmail.com</email></author><category term="Sentry" /><category term="Docker" /><category term="PostgreSQL" /><summary type="html"><![CDATA[TL;DR 자체 호스팅 Sentry의 Session Replay 데이터가 서버 디스크에서 얼마나 잡아먹는지 실측. 관련 DB 테이블(replays_replayrecordingsegment, sentry_file)은 스키마만 있고 row가 0건 — 이 배포에서는 리플레이가 Django ORM을 안 거치고 파일시스템에 직접 저장된다. DB 쿼리는 시간 낭비였다. 폴더는 UUID 그대로 샤딩(체크섬 샤딩 아님)이라 UI의 리플레이 ID로 바로 잡을 수 있다. 리플레이 1개당 평균 약 213KB, 총 10.9MB / 66GB 여유 — 지금 페이스로는 걱정할 단계 아님.]]></summary></entry><entry><title type="html">Sentry ingest 엔드포인트만 외부에 여는 방법</title><link href="https://stan-dev.cloud/sentry-ingest-exposure/" rel="alternate" type="text/html" title="Sentry ingest 엔드포인트만 외부에 여는 방법" /><published>2026-07-29T21:00:00+09:00</published><updated>2026-07-29T21:00:00+09:00</updated><id>https://stan-dev.cloud/sentry-ingest-exposure</id><content type="html" xml:base="https://stan-dev.cloud/sentry-ingest-exposure/"><![CDATA[<blockquote>
  <p><strong>TL;DR</strong>
자체 호스팅 Sentry에 이벤트 수집(ingest) 경로만 외부로 열었다. 관리 UI는 내부망 그대로 두고, 새 서브도메인(<code class="language-plaintext highlighter-rouge">ingest.sentry.x.xxx.kr</code>)을 분리해서 화이트리스트 경로만 허용.
지난 글에서 “외부 노출 안 됨”이라 판단한 게 실은 부정확했다 — 실측해보니 타임아웃이 아니라 403이었고, IP ACL이 애플리케이션 층에서 차단하고 있었을 뿐이다.
DSN 게이트로 열리는 write-only 엔드포인트는 인바운드 웹훅과 리스크 성격이 다르다. “내부망 전용” 원칙을 뒤집을지 여부는 리스크 등급별로 나눠서 판단해야 한다.</p>
</blockquote>

<hr />

<h2 id="배경">배경</h2>

<p>지난 글에서 자체 호스팅 Sentry에 Slack 연동을 붙였을 때, 아웃바운드(알림 발송)까지만 하고 인바운드(슬래시 커맨드·인터랙티브 버튼)는 접었다. 그때 결론이 “우리 Sentry는 내부망에만 열려 있어서 슬랙에서 오는 콜백을 못 받는다”였다.</p>

<p>이번에 실기기 앱이 보내는 에러 리포트가 403으로 막힌다는 문의가 들어왔다. 앱 SDK가 호출하는 이벤트 수집 경로(<code class="language-plaintext highlighter-rouge">/api/{project_id}/envelope/</code>, <code class="language-plaintext highlighter-rouge">/store/</code>)만 외부에서 허용해달라는 요청이었다. 작업하다 보니 지난 글에서 내가 했던 “외부 노출 안 됨” 판단이 실제로는 부정확했다는 걸 알게 됐다. 노출은 되어 있었고, IP ACL로 차단되고 있었을 뿐이었다.</p>

<hr />

<h2 id="어떤-상황이었나">어떤 상황이었나</h2>

<p>모바일팀에서 문의가 왔다.</p>

<blockquote>
  <p>Sentry가 내부망 IP만 허용하고 있어서, 실기기 앱이 보내는 에러 리포트가 전부 403으로 막힘. 관리 UI는 지금처럼 내부망 제한 유지하고, 아래 이벤트 수집 경로만 외부 허용 가능한지?</p>
  <div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>POST /api/{project_id}/envelope/
POST /api/{project_id}/store/
</code></pre></div>  </div>
</blockquote>

<p>요청 자체는 명확했다. 관리 대시보드는 지금처럼 내부망에서만 접근하고, 이벤트 수집 경로만 인터넷 어디에서든 도달 가능하게 만들면 된다.</p>

<hr />

<h2 id="개념-정리--dsn-게이트가-뭔지">개념 정리 — DSN 게이트가 뭔지</h2>

<p>Sentry 클라이언트는 이벤트를 보낼 때 <strong>DSN(Data Source Name)</strong> 이라는 문자열을 쓴다. 형태는 대략 이렇다.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>https://&lt;public_key&gt;@&lt;host&gt;/&lt;project_id&gt;
</code></pre></div></div>

<p>여기서 <code class="language-plaintext highlighter-rouge">&lt;public_key&gt;</code>는 프로젝트별로 발급되는 공개 키다. 이 키를 요청 헤더(또는 쿼리)에 실어서 <code class="language-plaintext highlighter-rouge">/api/{project_id}/envelope/</code>나 <code class="language-plaintext highlighter-rouge">/store/</code>에 POST한다. 서버는 키가 유효하고 프로젝트에 매칭되는지만 확인하고 이벤트를 받는다.</p>

<p>특징이 두 가지 있다.</p>

<ul>
  <li><strong>write-only</strong>: 이 경로로는 이벤트를 “보낼” 수만 있다. 데이터를 조회하거나 설정을 바꿀 수 없다.</li>
  <li><strong>원래 인터넷에 노출되는 게 정상</strong>: SaaS Sentry(<code class="language-plaintext highlighter-rouge">sentry.io</code>)나 일반적인 사내 배포에서도 이 경로는 인터넷에 열려 있다. 인증은 DSN 키가 담당한다.</li>
</ul>

<p>지난 글에서 다뤘던 슬랙 인바운드(Event Subscriptions/Interactivity) 웹훅은 성격이 다르다. 그쪽은 서명 검증만 있고 콜백 URL이 노출되면 재생 공격 여지가 상대적으로 크다. 반면 이 ingest 경로는 프로젝트별 키가 게이트를 잡고 있고, 최악의 경우에도 “가짜 이벤트가 프로젝트에 쌓인다” 수준이다.</p>

<p>리스크 성격이 다르다는 걸 짚어두고 시작한다.</p>

<hr />

<h2 id="층별-추적기">층별 추적기</h2>

<h3 id="1차-외부-도달성-실측--지난-판단-정정">[1차] 외부 도달성 실측 — 지난 판단 정정</h3>

<p>먼저 정말로 외부에서 도달이 안 되는지 다시 확인했다. 폰에서 와이파이를 끄고 데이터망으로 <code class="language-plaintext highlighter-rouge">https://sentry.x.xxx.kr</code>에 접속했다.</p>

<p><strong>타임아웃이 아니라 403 응답이 왔다.</strong></p>

<p>이 차이가 결정적이다. 타임아웃이면 네트워크 경로 자체가 막힌 것이고, 403이면 서버까지는 요청이 도달했는데 애플리케이션/프록시 레이어에서 반려한 것이다. 즉 라우터 포트포워딩과 DNS는 이미 열려 있고, NPM(Nginx Proxy Manager)의 Access List가 IP로 걸러내고 있는 상태였다.</p>

<p>지난 글에서 “외부 인터넷에 전혀 노출 안 됨”이라 적어둔 건 부정확한 서술이었다. 정확히는 “노출은 되어 있는데 IP ACL로 차단”이 맞다. 슬랙 인바운드를 접었던 결정 자체는 유효하지만(어차피 DSN 같은 게이트 없이 서명만으로 여는 건 별개 판단이 필요), 그때 이유 설명은 이 글에서 정정해둔다 — <a href="/sentry-slack-integration/">지난 글</a>에도 정정 박스를 달아뒀다.</p>

<blockquote>
  <p>응답 코드는 문제의 층을 알려준다. “안 된다”를 “타임아웃이라 안 된다”와 “403이라 안 된다”로 나눠 보는 순간, 다음 조치가 완전히 달라진다.</p>
</blockquote>

<h3 id="2차-npm에서-access-list-확인">[2차] NPM에서 Access List 확인</h3>

<p>192.168.x.x:81(NPM 관리 UI)에 접속해서 <code class="language-plaintext highlighter-rouge">sentry.x.xxx.kr</code> 프록시 호스트의 Access List 탭을 확인했다. 예상대로 “Local”이라는 사내 IP 화이트리스트가 걸려 있었다.</p>

<p>라우터 포트포워딩은 이미 되어 있는 상태였다. 그러니 이 단계는 손댈 게 없었다.</p>

<h3 id="3차--판단-기존-호스트-수정-vs-새-서브도메인-분리">[3차 — 판단] 기존 호스트 수정 vs 새 서브도메인 분리</h3>

<p>두 가지 선택지가 있었다.</p>

<ul>
  <li>(A) 기존 <code class="language-plaintext highlighter-rouge">sentry.x.xxx.kr</code> 호스트의 Advanced 탭에 커스텀 nginx를 넣어서 특정 경로만 예외 처리</li>
  <li>(B) 새 서브도메인 <code class="language-plaintext highlighter-rouge">ingest.sentry.x.xxx.kr</code>을 파고, 이건 처음부터 Public + 경로 화이트리스트만 허용</li>
</ul>

<p>(A)의 문제가 두 개였다.</p>

<ul>
  <li>NPM Access List는 호스트 단위로 걸리는 UI라 경로별 예외를 깔끔하게 표현하기 어렵다</li>
  <li>기존 관리 UI가 그대로 사는 호스트를 직접 수정하다가 실수하면 관리 UI가 통째로 인터넷에 노출되는 사고가 날 수 있다</li>
</ul>

<p>(B)로 갔다. 새 호스트를 파면 기존 <code class="language-plaintext highlighter-rouge">sentry.x.xxx.kr</code>은 손대지 않아도 된다. 실수의 폭발 반경이 좁아진다.</p>

<blockquote>
  <p>위험한 변경은 기존 리소스를 뜯어고치지 말고, 새 리소스로 격리해서 붙이는 게 안전하다. 문제가 생겨도 원상복구 대신 새 리소스만 걷어내면 된다.</p>
</blockquote>

<h3 id="4차-npm-새-proxy-host-설정--옵션을-최대한-껐다">[4차] NPM 새 Proxy Host 설정 — 옵션을 최대한 껐다</h3>

<p>새 프록시 호스트를 만들면서 옵션을 뺐다.</p>

<ul>
  <li>Cache Assets: 끔</li>
  <li>Websockets Support: 끔</li>
</ul>

<p>기존 관리 UI 호스트는 둘 다 켜져 있었다. 대시보드 정적 자산과 실시간 업데이트 때문에 필요한 설정이라 그건 맞다. 그런데 이 ingest 호스트는 POST API 두 개만 통과시키는 게 목적이다. 옵션을 켜면 NPM이 자동으로 추가하는 location 블록이 늘어나서, 내가 짜려는 “이 두 경로만 허용, 나머지 403” 규칙과 우선순위가 꼬일 여지가 있다.</p>

<p>Access List는 “Publicly Available”로 지정했다. 기존 “Local”에는 절대 연결하지 않았다(연결하는 순간 외부에서 오는 정상 요청이 다시 차단된다).</p>

<p>SSL 탭에서 Let’s Encrypt로 신규 발급, Force SSL을 켰다.</p>

<p>Advanced 탭에 넣은 커스텀 nginx는 이렇다.</p>

<div class="language-nginx highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">location</span> <span class="p">~</span> <span class="sr">^/api/[0-9]+/(envelope|store)/</span> <span class="p">{</span>
    <span class="kn">proxy_pass</span> <span class="s">http://192.168.x.x:9000</span><span class="p">;</span>
    <span class="kn">proxy_set_header</span> <span class="s">Host</span> <span class="s">sentry.x.xxx.kr</span><span class="p">;</span>
    <span class="kn">proxy_set_header</span> <span class="s">X-Forwarded-For</span> <span class="nv">$remote_addr</span><span class="p">;</span>
    <span class="kn">proxy_set_header</span> <span class="s">X-Forwarded-Proto</span> <span class="s">https</span><span class="p">;</span>
    <span class="kn">proxy_set_header</span> <span class="s">X-Request-Id</span> <span class="nv">$request_id</span><span class="p">;</span>
<span class="p">}</span>

<span class="k">location</span> <span class="n">/</span> <span class="p">{</span>
    <span class="kn">return</span> <span class="mi">403</span><span class="p">;</span>
<span class="p">}</span>
</code></pre></div></div>

<p>두 가지 포인트가 있다.</p>

<ul>
  <li><code class="language-plaintext highlighter-rouge">Host</code> 헤더는 반드시 <code class="language-plaintext highlighter-rouge">sentry.x.xxx.kr</code>로 고정한다. 지난 글(Sentry CSRF 트러블슈팅 편)에서 배운 것과 같은 이유다 — Sentry 내부 <code class="language-plaintext highlighter-rouge">system.url-prefix</code> 설정과 일치시켜야 한다.</li>
  <li>Rate limiting(<code class="language-plaintext highlighter-rouge">limit_req</code>)은 일단 뺐다. NPM에서 이건 호스트별 Advanced가 아니라 별도 글로벌 커스텀 설정 파일(<code class="language-plaintext highlighter-rouge">/data/nginx/custom/http_top.conf</code> 같은 것)을 컨테이너에 마운트해야 한다. 복잡도가 붙으니 기본 동작을 먼저 확인하고 2단계 작업으로 미뤘다.</li>
</ul>

<h3 id="5차--함정-외부망-테스트-폰-단독은-신뢰도가-낮았다">[5차 — 함정] 외부망 테스트, 폰 단독은 신뢰도가 낮았다</h3>

<p>서버 자체에서 curl로 확인해봤다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>curl <span class="nt">-I</span> https://ingest.sentry.x.xxx.kr/
<span class="c"># → HTTP/2 403                                (관리 UI 경로 차단 정상)</span>

curl <span class="nt">-X</span> POST https://ingest.sentry.x.xxx.kr/api/999999/store/
<span class="c"># → {"detail":"missing authorization information"}  (403 아님 = 실제 백엔드까지 도달)</span>
</code></pre></div></div>

<p>여기까지는 예상대로였다. 그런데 이건 지난 글에서 배운 대로 “셀프 curl 함정”이다 — 같은 서버 또는 내부망에서 도메인을 부르는 건 진짜 외부 경로 검증이 아니다. 그래서 폰에서도 테스트했다.</p>

<p>폰 브라우저/앱으로 두 경로를 쳐봤는데 <strong>둘 다 403</strong>이 나왔다. <code class="language-plaintext highlighter-rouge">/</code>는 의도된 403이지만 <code class="language-plaintext highlighter-rouge">/api/.../store/</code>까지 403이면 이상하다.</p>

<p>폰 단독 테스트를 그만두고, 맥북을 폰 핫스팟에 테더링한 뒤 <code class="language-plaintext highlighter-rouge">curl -v</code>로 재테스트했다. 그 결과 서버 자체 테스트와 동일했다 — <code class="language-plaintext highlighter-rouge">/</code>는 403, <code class="language-plaintext highlighter-rouge">/store/</code>는 인증 에러 JSON이 정상으로 나왔다.</p>

<p>원인을 명확히 특정하진 못했다. 폰 브라우저/앱의 캐시일 가능성이 크다. 다만 결론은 확실했다.</p>

<blockquote>
  <p>진짜 외부망 검증은 폰 단독 도구보다 노트북+핫스팟+<code class="language-plaintext highlighter-rouge">curl -v</code> 조합이 신뢰도가 높다. 폰 앱은 응답 헤더를 상세히 보여주지 않고, 캐시 층도 통제하기 어렵다.</p>
</blockquote>

<hr />

<h2 id="원인-층위-정리">원인 층위 정리</h2>

<table>
  <thead>
    <tr>
      <th>층</th>
      <th>상태</th>
      <th>조치</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>라우터 포트포워딩</td>
      <td>이미 열려있음</td>
      <td>손대지 않음</td>
    </tr>
    <tr>
      <td>DNS</td>
      <td><code class="language-plaintext highlighter-rouge">sentry.x.xxx.kr</code>은 이미 있음 → <code class="language-plaintext highlighter-rouge">ingest.</code> 서브도메인 A 레코드 추가</td>
      <td>신규</td>
    </tr>
    <tr>
      <td>NPM Access List</td>
      <td>기존 호스트에 “Local” 걸림</td>
      <td>유지, 새 호스트는 “Publicly Available”</td>
    </tr>
    <tr>
      <td>NPM 프록시 호스트</td>
      <td>기존 호스트는 관리 UI용</td>
      <td>새 호스트를 별도로 팜</td>
    </tr>
    <tr>
      <td>nginx 라우팅</td>
      <td>화이트리스트 미적용</td>
      <td><code class="language-plaintext highlighter-rouge">envelope|store</code>만 허용, 나머지 403</td>
    </tr>
    <tr>
      <td>Sentry 내부 <code class="language-plaintext highlighter-rouge">Host</code> 헤더</td>
      <td><code class="language-plaintext highlighter-rouge">system.url-prefix</code>와 일치해야 함</td>
      <td>프록시에서 <code class="language-plaintext highlighter-rouge">Host: sentry.x.xxx.kr</code> 강제</td>
    </tr>
    <tr>
      <td>DSN 클라이언트 설정</td>
      <td>기존은 <code class="language-plaintext highlighter-rouge">sentry.x.xxx.kr</code> 호스트</td>
      <td>호스트 부분만 <code class="language-plaintext highlighter-rouge">ingest.</code>로 교체 (키/project_id 유지)</td>
    </tr>
  </tbody>
</table>

<hr />

<h2 id="최종-설정-스니펫">최종 설정 스니펫</h2>

<p>새 프록시 호스트의 Advanced 탭 커스텀 nginx.</p>

<div class="language-nginx highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">location</span> <span class="p">~</span> <span class="sr">^/api/[0-9]+/(envelope|store)/</span> <span class="p">{</span>
    <span class="kn">proxy_pass</span> <span class="s">http://192.168.x.x:9000</span><span class="p">;</span>
    <span class="kn">proxy_set_header</span> <span class="s">Host</span> <span class="s">sentry.x.xxx.kr</span><span class="p">;</span>
    <span class="kn">proxy_set_header</span> <span class="s">X-Forwarded-For</span> <span class="nv">$remote_addr</span><span class="p">;</span>
    <span class="kn">proxy_set_header</span> <span class="s">X-Forwarded-Proto</span> <span class="s">https</span><span class="p">;</span>
    <span class="kn">proxy_set_header</span> <span class="s">X-Request-Id</span> <span class="nv">$request_id</span><span class="p">;</span>
<span class="p">}</span>

<span class="k">location</span> <span class="n">/</span> <span class="p">{</span>
    <span class="kn">return</span> <span class="mi">403</span><span class="p">;</span>
<span class="p">}</span>
</code></pre></div></div>

<p>외부망 확인용 curl.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c"># 관리 UI 경로(차단되어야 정상)</span>
curl <span class="nt">-I</span> https://ingest.sentry.x.xxx.kr/
<span class="c"># → HTTP/2 403</span>

<span class="c"># 이벤트 수집 경로(403이 아니어야 정상 = 백엔드까지 도달)</span>
curl <span class="nt">-X</span> POST https://ingest.sentry.x.xxx.kr/api/999999/store/
<span class="c"># → {"detail":"missing authorization information"}</span>
</code></pre></div></div>

<p>앱 쪽 DSN 변경은 호스트 부분만 교체하면 된다. 예: <code class="language-plaintext highlighter-rouge">https://&lt;public_key&gt;@sentry.x.xxx.kr/&lt;project_id&gt;</code> → <code class="language-plaintext highlighter-rouge">https://&lt;public_key&gt;@ingest.sentry.x.xxx.kr/&lt;project_id&gt;</code>. <strong>키/<code class="language-plaintext highlighter-rouge">project_id</code>는 그대로 유지한다</strong> — DSN 재발급이 아니라 호스트 스왑이다.</p>

<hr />

<h2 id="다음번-순서">다음번 순서</h2>

<ol>
  <li><strong>외부 도달성 실측</strong> — 다른 망(데이터망/핫스팟)에서 응답 코드까지 확인. 타임아웃인지 403인지에 따라 원인이 달라진다.</li>
  <li><strong>위험한 변경은 새 리소스로 격리</strong> — 기존 호스트/설정을 직접 뜯어고치지 말고, 새 서브도메인/호스트를 파서 붙인다. 실수의 폭발 반경을 좁힌다.</li>
  <li><strong>경로 화이트리스트 + 나머지 403</strong> — 이 두 경로만 허용한다면, <code class="language-plaintext highlighter-rouge">location / { return 403; }</code>을 명시적으로 넣어서 미매칭 요청을 반려한다. <code class="language-plaintext highlighter-rouge">Host</code> 헤더는 원본과 일치시키는 걸 잊지 않는다.</li>
  <li><strong>외부망 최종 검증은 노트북+핫스팟+<code class="language-plaintext highlighter-rouge">curl -v</code></strong> — 서버 자체 curl은 셀프 함정, 폰 단독은 캐시 함정. 노트북을 폰 핫스팟에 물려서 <code class="language-plaintext highlighter-rouge">curl -v</code>로 헤더까지 확인한다.</li>
</ol>

<hr />

<h2 id="참고">참고</h2>

<ul>
  <li>선행 글: <a href="/sentry-csrf-reverse-proxy/">CSRF 6단계 편</a> (<code class="language-plaintext highlighter-rouge">Host</code> 헤더·<code class="language-plaintext highlighter-rouge">system.url-prefix</code>), <a href="/sentry-slack-integration/">Slack 통합 편</a> (이 글이 정정한 인바운드 판단)</li>
  <li>환경: self-hosted Sentry (Docker Compose), Nginx Proxy Manager, 사내망 게이트웨이</li>
</ul>]]></content><author><name>이진용</name><email>kouig14@gmail.com</email></author><category term="Sentry" /><category term="Nginx" /><category term="NPM" /><category term="네트워크" /><summary type="html"><![CDATA[TL;DR 자체 호스팅 Sentry에 이벤트 수집(ingest) 경로만 외부로 열었다. 관리 UI는 내부망 그대로 두고, 새 서브도메인(ingest.sentry.x.xxx.kr)을 분리해서 화이트리스트 경로만 허용. 지난 글에서 “외부 노출 안 됨”이라 판단한 게 실은 부정확했다 — 실측해보니 타임아웃이 아니라 403이었고, IP ACL이 애플리케이션 층에서 차단하고 있었을 뿐이다. DSN 게이트로 열리는 write-only 엔드포인트는 인바운드 웹훅과 리스크 성격이 다르다. “내부망 전용” 원칙을 뒤집을지 여부는 리스크 등급별로 나눠서 판단해야 한다.]]></summary></entry><entry><title type="html">Sentry Admin은 왜 사라졌나</title><link href="https://stan-dev.cloud/sentry-permission-model/" rel="alternate" type="text/html" title="Sentry Admin은 왜 사라졌나" /><published>2026-07-28T21:00:00+09:00</published><updated>2026-07-28T21:00:00+09:00</updated><id>https://stan-dev.cloud/sentry-permission-model</id><content type="html" xml:base="https://stan-dev.cloud/sentry-permission-model/"><![CDATA[<blockquote>
  <p><strong>TL;DR</strong>
Sentry에서 조직 Admin 역할을 <strong>신규로는 더 이상 부여할 수 없다.</strong> 팀 스코프의 Team Admin이 그 자리를 대체했다.
배경엔 “권한을 좁혀서 최소 범위로 준다”는 설계 방향이 있고, 이 방향은 UI 곳곳에 반영되어 있다 — 예를 들어 Integrations(조직 스코프)와 Alerts(프로젝트/팀 스코프)는 같은 “알림 관련”으로 보여도 접근 권한이 완전히 다르다.
핵심 열쇠는 <strong>프로젝트에 독립적인 권한 체계가 없다</strong>는 것 — 프로젝트 접근은 팀 소속으로만 결정된다.</p>
</blockquote>

<hr />

<h2 id="배경">배경</h2>

<p>이전 글에서 Slack 통합을 붙이고 며칠 뒤, 프론트팀 리더(Owner 권한 보유)한테서 문의가 왔다.</p>

<blockquote>
  <p>“팀원 한 명에게 Admin 권한을 주려는데, Settings → Members에서 Admin 역할이 선택 자체가 안 돼요.”</p>
</blockquote>

<p>처음엔 UI 버그로 의심했는데, 파고들어 보니 의도된 변경이었다. 그리고 이 변경이 단순한 UI 정리가 아니라 Sentry가 권한을 어떻게 바라보는지에 대한 설계 방향을 드러내고 있었다.</p>

<hr />

<h2 id="어떤-상황이었나">어떤 상황이었나</h2>

<p>프론트팀 리더는 Owner 권한을 갖고 있었고, 팀원 한 명에게 관리 권한을 넘기고 싶어했다. 자연스럽게 Settings → Members로 가서 대상 계정의 역할을 Admin으로 바꾸려 했는데, <strong>Admin 항목이 아예 선택 불가 상태</strong>였다.</p>

<p>문서를 확인해봤더니 이건 버그가 아니라 정책이었다. <a href="https://docs.sentry.io/organization/membership/">Sentry 공식 문서</a>는 Business·Enterprise 플랜에서 Org Admin 역할이 Team Admin으로 대체돼 신규 부여가 안 된다고 명시한다 — <em>“This role can no longer be assigned.”</em> 우리 self-hosted 인스턴스에서도 같은 동작이었고, 기존에 이미 Admin이던 계정만 그대로 유지된다.</p>

<hr />

<h2 id="개념-정리--sentry의-두-층-권한-모델">개념 정리 — Sentry의 두 층 권한 모델</h2>

<p>Sentry의 권한 구조는 이렇게 정리된다.</p>

<table>
  <thead>
    <tr>
      <th>층</th>
      <th>역할</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>조직 (전체 범위)</td>
      <td>Owner / Manager / Member / Billing / (Admin, 레거시)</td>
    </tr>
    <tr>
      <td>팀 (해당 팀 범위)</td>
      <td>Team Admin</td>
    </tr>
    <tr>
      <td>프로젝트</td>
      <td><strong>없음</strong> (팀 소속을 통해서만 접근 결정)</td>
    </tr>
  </tbody>
</table>

<p>핵심은 <strong>프로젝트에 독립적인 권한 체계가 없다</strong>는 것이다. “이 프로젝트에만 적용되는 admin”이라는 개념 자체가 Sentry에 없다. 프로젝트에 대한 권한은 그 프로젝트가 속한 팀의 소속으로만 결정된다.</p>

<p>Admin(레거시)이 신규 부여가 막힌 이유는 이 자리를 Team Admin이 대체하기 때문이다. 조직 전체에 걸친 관리자가 필요하면 Manager를, 특정 팀에 한정된 관리자가 필요하면 Team Admin을 쓴다.</p>

<hr />

<h2 id="층별-추적기">층별 추적기</h2>

<h3 id="1차-admin-신규-부여-시도가-막혀-있다">[1차] Admin 신규 부여 시도가 막혀 있다</h3>

<p>Settings → Members에서 대상 계정의 역할을 Admin으로 바꾸려니 선택지 자체가 회색으로 잠겨 있음. 문서 확인 결과 의도된 변경.</p>

<blockquote>
  <p>같은 이름의 역할(Admin)이라도 스코프가 바뀌면 완전히 다른 것이 된다. 조직 Admin은 사라졌고, 그 이름은 이제 팀 스코프의 Team Admin으로만 이어진다.</p>
</blockquote>

<h3 id="2차-team-admin으로-대체">[2차] Team Admin으로 대체</h3>

<p>대안은 Team Admin. 팀 페이지(Teams → 해당 팀 → Members)에서 팀 단위로 부여한다. “이 팀의 관리자”이므로 다른 팀에는 영향이 없다.</p>

<p>Manager와의 차이는 스코프다. Manager는 조직 전체 관리자, Team Admin은 그 팀 안에서만 관리자. 지금 이 요청은 “특정 팀의 팀원이 그 팀 안에서 관리 권한을 갖는 것”이라 Team Admin이 맞는 그림이었다.</p>

<h3 id="3차-이-프로젝트만-스코프-요청은-어떻게-처리하는가">[3차] “이 프로젝트만” 스코프 요청은 어떻게 처리하는가</h3>

<p>여기가 개념적으로 흥미로운 지점이다. 프로젝트에 독립 권한이 없으니 “이 프로젝트만 관리 권한”이라는 요청은 문자 그대로는 처리할 수 없다. 대신 팀 구조를 통해 표현한다.</p>

<ul>
  <li><strong>팀이 그 프로젝트 하나만 갖고 있으면</strong>: Team Admin이 사실상 “그 프로젝트 전용 admin”이 된다. 우회로가 자연스럽게 성립.</li>
  <li><strong>팀이 여러 프로젝트를 갖고 있으면</strong>: Team Admin을 주는 순간 그 팀의 모든 프로젝트에 관리 권한이 생긴다. 프로젝트 단위로 격리하려면 해당 프로젝트만 소유하는 별도 팀을 새로 만드는 구조 변경이 필요.</li>
</ul>

<p>팀이 어떤 프로젝트를 갖고 있는지 확인하는 곳: Settings → Teams → 해당 팀 → Projects 탭.</p>

<p>이번 케이스에서는 프론트팀의 Projects 탭에 <code class="language-plaintext highlighter-rouge">&lt;프로젝트&gt;</code> 하나만 소속되어 있음을 확인. Team Admin을 줘도 다른 프로젝트로 권한이 새지 않는다는 걸 확인하고 부여했다.</p>

<blockquote>
  <p>프로젝트에 독립 권한이 없다는 건 “권한을 좁힐 수 없다”가 아니라 “권한을 팀 구조로 표현하라”는 뜻이다. 프로젝트 하나만 소유한 팀을 만들면 사실상 프로젝트 전용 권한이 된다.</p>
</blockquote>

<h3 id="4차--후속-team-admin으로도-integrations-페이지엔-못-들어간다">[4차 — 후속] Team Admin으로도 Integrations 페이지엔 못 들어간다</h3>

<p>며칠 뒤 같은 팀원한테서 문의가 왔다. “Team Admin 받았는데 Settings → Integrations 페이지가 안 보여요. 알림도 팀 단위로 관리되는 건가요?”</p>

<p>확인해보니 두 개가 완전히 다른 스코프였다.</p>

<table>
  <thead>
    <tr>
      <th>페이지</th>
      <th>스코프</th>
      <th>접근 권한</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Settings → Integrations (Slack 워크스페이스 연결 자체 관리)</td>
      <td>조직 전체</td>
      <td>Owner / Manager</td>
    </tr>
    <tr>
      <td>Alerts (Alert Rule에서 어느 채널로 보낼지 설정)</td>
      <td>프로젝트/팀</td>
      <td>Team Admin 이상</td>
    </tr>
  </tbody>
</table>

<p>같은 “Slack 관련” 화면인데 스코프가 다르다. Integrations는 조직이 어떤 외부 서비스를 연결할지 자체를 다루므로 조직 관리자가, Alerts는 이미 연결된 것을 어떻게 활용할지를 다루므로 팀 관리자가 처리한다.</p>

<blockquote>
  <p>“Slack 알림” 하나 다루는 데도 두 스코프가 섞인다 — 워크스페이스 연결 자체(조직) vs 그 워크스페이스를 어떻게 사용할지(프로젝트/팀). 문의 받았을 때 어느 쪽 얘기인지부터 구분하지 않으면 엉뚱한 곳에서 헤맨다.</p>
</blockquote>

<hr />

<h2 id="정리">정리</h2>

<p>요청 유형별로 어느 층에서 처리하는지 한 번에 정리하면 이렇다.</p>

<table>
  <thead>
    <tr>
      <th>요청</th>
      <th>처리 층</th>
      <th>필요 권한</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>조직 전체 관리</td>
      <td>조직</td>
      <td>Manager (Admin 신규 부여 불가)</td>
    </tr>
    <tr>
      <td>특정 팀 안에서만 관리</td>
      <td>팀</td>
      <td>Team Admin</td>
    </tr>
    <tr>
      <td>“이 프로젝트만” 관리</td>
      <td>팀 구조 확인 후 결정</td>
      <td>Team Admin (팀이 그 프로젝트만 소유해야 안전)</td>
    </tr>
    <tr>
      <td>Slack 워크스페이스 연결 자체</td>
      <td>조직 (Integrations)</td>
      <td>Owner / Manager</td>
    </tr>
    <tr>
      <td>Alert Rule 관리 (채널·조건 등)</td>
      <td>프로젝트/팀 (Alerts)</td>
      <td>Team Admin</td>
    </tr>
  </tbody>
</table>

<p>문의는 늘 “권한 좀 주세요”라는 한 문장으로 오지만, 실제로는 이 표의 어느 행인지부터 구분해야 정확한 응답이 나온다.</p>

<hr />

<h2 id="최종-결정-실제-케이스">최종 결정 (실제 케이스)</h2>

<ul>
  <li>프론트팀 Projects 탭 확인 → <code class="language-plaintext highlighter-rouge">&lt;프로젝트&gt;</code> 하나만 소속</li>
  <li>팀원에게 <strong>Team Admin 부여</strong> → 사실상 그 프로젝트 전용 관리자와 동일한 효과</li>
  <li>안내: 앞으로 이 팀에 다른 프로젝트가 추가되면, 그 시점부터 이 Team Admin들이 새 프로젝트에도 자동으로 권한을 갖게 된다. 팀에 프로젝트 추가할 때마다 이 점 재검토 필요.</li>
  <li>Integrations 자체 관리는 여전히 Owner/Manager가 처리하고, Alert Rule은 Team Admin으로 충분하다는 것도 함께 안내.</li>
</ul>

<hr />

<h2 id="권한-요청-받았을-때-순서">권한 요청 받았을 때 순서</h2>

<p>같은 유형의 요청을 받을 때, 다음 순서를 밟으면 대부분의 헛발질이 앞단에서 걸러진다.</p>

<ol>
  <li><strong>“조직 전체가 필요한가, 팀/프로젝트 한정이면 되는가”부터 확인한다.</strong> 조직 전체면 Manager, 팀/프로젝트 한정이면 Team Admin.</li>
  <li><strong>“이 프로젝트만” 요청이면 팀 구조부터 확인한다</strong> (Teams → 해당 팀 → Projects 탭). 팀이 그 프로젝트만 소유하면 Team Admin 안전, 여러 개 소유면 새 팀을 만들거나 요청 자체를 재조정.</li>
  <li><strong>Team Admin 부여 후 그 팀에 프로젝트가 추가될 때마다 재검토한다.</strong> 팀 소속이 확장되면 자동으로 권한도 확장된다는 사실을 인수인계 문서에 명시.</li>
  <li><strong>“권한이 없다” 문의 오면 Integrations인지 Alerts인지 먼저 구분한다.</strong> 둘은 완전히 다른 스코프.</li>
  <li><strong>가능한 한 좁은 범위(Team Admin)로 준다.</strong> 필요할 때 넓히는 게, 넓게 줬다 좁히는 것보다 훨씬 쉽다.</li>
</ol>

<hr />

<h3 id="참고">참고</h3>

<ul>
  <li><strong>환경</strong>: Sentry self-hosted (Docker Compose) · 내부망 전용 도메인</li>
  <li><strong>관련 문서</strong>: <a href="https://docs.sentry.io/organization/membership/">Sentry Membership</a> · <a href="https://docs.sentry.io/organization/membership/#team-level-roles">Team-level Roles</a></li>
  <li><strong>선행 글</strong>: 이 이슈는 <a href="/sentry-slack-integration/">Slack 통합 편</a>에서 파생된 후속. Slack 통합을 붙인 뒤 팀 리더가 팀원에게 관리 권한을 넘기려다 발견한 케이스.</li>
</ul>]]></content><author><name>이진용</name><email>kouig14@gmail.com</email></author><category term="Sentry" /><category term="권한" /><summary type="html"><![CDATA[TL;DR Sentry에서 조직 Admin 역할을 신규로는 더 이상 부여할 수 없다. 팀 스코프의 Team Admin이 그 자리를 대체했다. 배경엔 “권한을 좁혀서 최소 범위로 준다”는 설계 방향이 있고, 이 방향은 UI 곳곳에 반영되어 있다 — 예를 들어 Integrations(조직 스코프)와 Alerts(프로젝트/팀 스코프)는 같은 “알림 관련”으로 보여도 접근 권한이 완전히 다르다. 핵심 열쇠는 프로젝트에 독립적인 권한 체계가 없다는 것 — 프로젝트 접근은 팀 소속으로만 결정된다.]]></summary></entry><entry><title type="html">self-hosted Sentry에 Slack 통합을 붙이며</title><link href="https://stan-dev.cloud/sentry-slack-integration/" rel="alternate" type="text/html" title="self-hosted Sentry에 Slack 통합을 붙이며" /><published>2026-07-27T21:00:00+09:00</published><updated>2026-07-27T21:00:00+09:00</updated><id>https://stan-dev.cloud/sentry-slack-integration</id><content type="html" xml:base="https://stan-dev.cloud/sentry-slack-integration/"><![CDATA[<blockquote>
  <p><strong>TL;DR</strong>
사내 self-hosted Sentry에 Slack 알림을 붙였다. WebHooks 레거시 플러그인은 2025년 중반부터 신규 설치가 막혀 있어 조직 설정의 Integrations를 써야 했다.
Event Subscriptions URL 검증이 계속 튕겼고, 원인은 <strong>리버스 프록시의 IP 화이트리스트가 Slack의 요청을 앱 앞에서 막고 있었기 때문</strong>이다. (당시엔 “외부 경로 자체가 없다”로 판단했다가 이후 실측으로 정정했다 — §4 정정 참고.)
결론은 인바운드 3종(Event Subscriptions·Slash Command·Interactivity)을 의도적으로 포기하고 아웃바운드 알림만 유지 — 원래 요청은 그것만으로 충족된다.</p>
</blockquote>

<hr />

<h2 id="배경">배경</h2>

<p>이전 글에서 세팅한 self-hosted Sentry(<code class="language-plaintext highlighter-rouge">sentry.x.xxx.kr</code>) 위에, 이번엔 프론트팀에서 요청이 들어왔다.</p>

<blockquote>
  <p>“우리 프로젝트 에러가 뜨면 Slack 채널로 알림을 받고 싶어요.”</p>
</blockquote>

<p>익숙한 요청이라 시작은 단순해 보였다. Slack App 하나 만들고, Sentry에 자격증명 등록하고, 프로젝트에 알림 규칙 하나 만들면 끝나는 줄 알았다. 실제로는 URL 검증에서 몇 시간을 태웠고, 마지막엔 <strong>기능 중 일부를 의도적으로 포기하는 것</strong>이 답이었다.</p>

<hr />

<h2 id="어떤-상황이었나">어떤 상황이었나</h2>

<p>프론트팀에서 특정 프로젝트의 Sentry 이슈가 발생하면 Slack 채널에 알림을 받고 싶다는 요청이었다. 익숙한 흐름이라 프로젝트 설정 페이지로 바로 들어갔다.</p>

<p><code class="language-plaintext highlighter-rouge">sentry.x.xxx.kr/settings/&lt;조직&gt;/projects/&lt;프로젝트&gt;/plugins/</code></p>

<p>여기서 WebHooks를 켜면 될 줄 알았는데, <strong>항목 자체가 뜨지 않았다.</strong> 다른 레거시 플러그인들도 대부분 비활성 표시(<code class="language-plaintext highlighter-rouge">This Plugin is deprecated and not available to install on new projects</code>)로 잠겨 있었다.</p>

<p>버그인 줄 알고 이슈 트래커부터 뒤졌더니, 이건 의도된 변경이었다.</p>

<hr />

<h2 id="webhooks는-왜-사라졌나--개념-정리">WebHooks는 왜 사라졌나 — 개념 정리</h2>

<p>Sentry는 2025년 중반에 레거시 플러그인의 신규 설치를 차단했다 (관련 PR #91550). 이후 대부분의 플러그인은 “deprecated” 상태로 잠기고, 새 프로젝트에는 설치가 안 된다.</p>

<p>권장 흐름은 <strong>조직 설정의 Integrations</strong>를 통과하는 구조로 재편됐다. Slack의 경우 두 층으로 나뉜다.</p>

<table>
  <thead>
    <tr>
      <th>층</th>
      <th>담당</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>조직 Integrations</td>
      <td>Slack OAuth로 워크스페이스에 붙임 (조직당 1회)</td>
    </tr>
    <tr>
      <td>프로젝트 Alert Rule</td>
      <td>“Send a Slack notification” 액션으로 채널 지정</td>
    </tr>
  </tbody>
</table>

<p>이 구조가 왜 나은가. <strong>한 조직에서 Slack을 한 번만 연결해두면 여러 프로젝트가 재사용</strong>할 수 있고, 자격증명을 프로젝트 플러그인마다 흩어놓지 않아도 된다. 조직 자원과 프로젝트 자원의 경계를 분리한 방향이다.</p>

<p>이제 이 흐름을 따라가면 되는데 — 함정은 여기서부터였다.</p>

<hr />

<h2 id="층별-추적기">층별 추적기</h2>

<h3 id="1차-slack-app-만들고-자격증명-반영">[1차] Slack App 만들고 자격증명 반영</h3>

<p><code class="language-plaintext highlighter-rouge">api.slack.com/apps</code>에서 App 생성 → Basic Information에서 Client ID / Client Secret / Signing Secret 확보 → 서버 <code class="language-plaintext highlighter-rouge">sentry.conf.py</code>에 등록.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">SENTRY_OPTIONS</span><span class="p">[</span><span class="s">'slack.client-id'</span><span class="p">]</span>       <span class="o">=</span> <span class="s">'&lt;masked&gt;'</span>
<span class="n">SENTRY_OPTIONS</span><span class="p">[</span><span class="s">'slack.client-secret'</span><span class="p">]</span>   <span class="o">=</span> <span class="s">'&lt;masked&gt;'</span>
<span class="n">SENTRY_OPTIONS</span><span class="p">[</span><span class="s">'slack.signing-secret'</span><span class="p">]</span>  <span class="o">=</span> <span class="s">'&lt;masked&gt;'</span>
</code></pre></div></div>

<p><code class="language-plaintext highlighter-rouge">docker compose restart</code> 후 반영 여부를 실제로 확인.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>docker compose <span class="nb">exec </span>web sentry shell <span class="nt">-c</span> <span class="se">\</span>
  <span class="s2">"from sentry import options; print(options.get('slack.signing-secret'))"</span>
</code></pre></div></div>

<p>값이 정상 출력되면 옵션이 실제로 로드된 것. 이전 CSRF 작업에서 배운 감각이 여기서 쓰였다.</p>

<blockquote>
  <p>설정이 반영됐다고 믿기 전에 어디서 어떻게 읽고 있는지를 먼저 확인한다.</p>
</blockquote>

<h3 id="2차-slack-app-세부-url스코프-설정">[2차] Slack App 세부 URL·스코프 설정</h3>

<p>Slack App 설정에 Sentry 쪽 엔드포인트를 붙였다.</p>

<ul>
  <li>Slash Commands: <code class="language-plaintext highlighter-rouge">/sentry</code> → <code class="language-plaintext highlighter-rouge">https://sentry.x.xxx.kr/extensions/slack/commands/</code></li>
  <li>Interactivity Request URL: <code class="language-plaintext highlighter-rouge">https://sentry.x.xxx.kr/extensions/slack/action/</code></li>
  <li>Options Load URL: <code class="language-plaintext highlighter-rouge">https://sentry.x.xxx.kr/extensions/slack/options-load/</code></li>
  <li>Bot Token Scopes 12개 (<code class="language-plaintext highlighter-rouge">channels:read</code>, <code class="language-plaintext highlighter-rouge">chat:write</code>, <code class="language-plaintext highlighter-rouge">chat:write.customize</code>, <code class="language-plaintext highlighter-rouge">chat:write.public</code>, <code class="language-plaintext highlighter-rouge">commands</code>, <code class="language-plaintext highlighter-rouge">groups:read</code>, <code class="language-plaintext highlighter-rouge">im:history</code>, <code class="language-plaintext highlighter-rouge">im:read</code>, <code class="language-plaintext highlighter-rouge">links:read</code>, <code class="language-plaintext highlighter-rouge">links:write</code>, <code class="language-plaintext highlighter-rouge">team:read</code>, <code class="language-plaintext highlighter-rouge">users:read</code>)</li>
</ul>

<p>여기까지는 문서 순서 그대로.</p>

<h3 id="3차--함정-시작-event-subscriptions-url-검증-실패">[3차 — 함정 시작] Event Subscriptions URL 검증 실패</h3>

<p>Event Subscriptions에 <code class="language-plaintext highlighter-rouge">https://sentry.x.xxx.kr/extensions/slack/event/</code>를 넣고 저장하니 이 에러가 떴다.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Your request URL responded with an HTTP error.
Your URL didn't respond with the value of the challenge parameter.
</code></pre></div></div>

<p>Slack이 검증용 <code class="language-plaintext highlighter-rouge">challenge</code> 값을 POST하고 서버가 그 값을 응답 바디로 되돌려줘야 통과하는 방식인데, 응답이 이상하다는 뜻이다.</p>

<p>먼저 서버 내부와 도메인 양쪽에서 직접 찔러봤다.</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>curl localhost:9000/extensions/slack/event/            <span class="c"># 401</span>
curl https://sentry.x.xxx.kr/extensions/slack/event/   <span class="c"># 401</span>
</code></pre></div></div>

<p>둘 다 401 Unauthorized로 동일. 이 시점에선 “nginx 라우팅 문제는 아니고, 요청은 Sentry 앱까지 도달한다”고 결론 냈다.</p>

<p>컨테이너 로그도 봤다.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>INFO   slack.event.url_verification  ...
ERROR  slack.action.auth             ...
</code></pre></div></div>

<p>URL 인식은 되는데 서명 검증에서 막힘. 다만 이건 서명 헤더 없이 수동으로 curl을 찌른 결과라 401은 정상이었다.</p>

<p>여기서 이상해서 로그 tail을 걸어두고 Slack App 화면에서 Save를 다시 눌렀다.</p>

<p><strong>로그에 아무것도 찍히지 않았다.</strong></p>

<blockquote>
  <p>서명이나 인증에서 막힌다고 보여도, 진짜 원인이 그 앞단(요청이 서버에 도달하는지)에 있을 수 있다. 도달성부터 확인한다.</p>
</blockquote>

<h3 id="4차--원인-규명-인바운드가-막힌-세계">[4차 — 원인 규명] 인바운드가 막힌 세계</h3>

<p>셀프 curl은 앱까지 도달하는데, 진짜 Slack이 보내는 요청은 로그에 아예 찍히지 않는다. 답은 하나였다.</p>

<p><code class="language-plaintext highlighter-rouge">sentry.x.xxx.kr</code>은 <strong>내부망 리버스 프록시로만 열려 있고, 외부 인터넷에 노출되지 않은 도메인</strong>이다. DNS는 풀리더라도 해당 IP까지 가는 경로는 사내망 밖에서는 존재하지 않는다. Slack 클라우드 서버는 아무리 우리를 부르려 해도 우리 게이트웨이까지 못 온다.</p>

<p><strong>함정은 확인 과정 자체에 있었다.</strong> “외부 접근이 되는지” 확인한다고 셀프 서버에서 도메인 curl을 찔렀는데, 그건 결국 같은 사내망 안에서 도는 요청이었다. 진짜 외부 경로를 검증한 게 아니었다. 이 구분이 안 돼서 몇 시간을 태웠다.</p>

<blockquote>
  <p>셀프 curl은 애플리케이션이 요청을 어떻게 처리하는지를 확인하는 도구지, 인터넷에서 우리 서버에 도달할 수 있는지를 확인하는 도구가 아니다.</p>
</blockquote>

<blockquote>
  <p><strong>📌 정정 — ingest 외부 노출 편 이후</strong></p>

  <p>위에서 “사내망 밖에서는 경로가 존재하지 않는다”고 쓴 건 부정확했다. 이후 앱 SDK용 ingest 경로를 외부에 열면서 외부망(LTE)에서 실측해보니 <strong>타임아웃이 아니라 403이 돌아왔다.</strong> 요청은 서버까지 도달하고 있었고, 리버스 프록시(NPM)의 IP 화이트리스트(Access List)가 애플리케이션 앞에서 차단하고 있었을 뿐이다.</p>

  <p>그러니 Slack의 URL 검증 요청이 앱 로그에 안 찍힌 진짜 이유는 “못 온 것”이 아니라 <strong>프록시에서 403으로 잘려 앱까지 못 간 것</strong>이다. 인바운드 3종을 포기한 결정 자체는 그대로 유효하다 — 화이트리스트를 Slack IP 대역에 여는 건 별개의 보안 판단이 필요하고, 그 판단은 이 글의 범위 밖이다. 실측 과정은 <a href="/sentry-ingest-exposure/">ingest 엔드포인트 외부 노출 편</a>에.</p>
</blockquote>

<h3 id="5차--방향-판단-무엇을-포기할-것인가">[5차 — 방향 판단] 무엇을 포기할 것인가</h3>

<p>여기서 갈래가 갈렸다.</p>

<ul>
  <li><strong>A. 서버를 외부에 노출한다</strong> — 공인 IP나 별도 프록시로 Slack이 도달할 수 있게 만든다. 대신 세무 데이터를 다루는 서비스의 외부 노출은 편의를 넘는 별도의 보안 결정이 된다.</li>
  <li><strong>B. 인바운드가 필요한 기능만 포기한다</strong> — 원래 목적이 인바운드 없이도 충족되는지 다시 본다.</li>
</ul>

<p>Slack 통합의 기능들을 방향으로 나눠봤다.</p>

<table>
  <thead>
    <tr>
      <th>기능</th>
      <th>방향</th>
      <th>우리 환경에서</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Event Subscriptions</td>
      <td>Slack → Sentry</td>
      <td>❌ 인바운드 필요</td>
    </tr>
    <tr>
      <td>Slash Commands (<code class="language-plaintext highlighter-rouge">/sentry</code>)</td>
      <td>Slack → Sentry</td>
      <td>❌ 인바운드 필요</td>
    </tr>
    <tr>
      <td>Interactivity (버튼·모달)</td>
      <td>Slack → Sentry</td>
      <td>❌ 인바운드 필요</td>
    </tr>
    <tr>
      <td><strong>에러 알림 전송</strong></td>
      <td><strong>Sentry → Slack</strong></td>
      <td><strong>✅ 아웃바운드로 충분</strong></td>
    </tr>
  </tbody>
</table>

<p>원래 요청은 “에러 알림”이었다. Sentry가 Slack API를 호출해서 메시지를 밀어넣는 방향이라, Slack이 우리 서버에 접근할 필요가 자체가 없다. <strong>필요한 방향이 이미 확보되어 있었다.</strong></p>

<p>B로 갔다. 세무 데이터를 다루는 서비스의 외부 노출을 인바운드 편의 몇 개 때문에 감수하는 건 균형이 안 맞는다. Event Subscriptions·Slash Command·Interactivity는 검증 실패 상태 그대로 방치.</p>

<blockquote>
  <p>통합 기능은 “필요한 방향이 이미 확보되어 있는가”부터 좁혀서 본다. 인바운드가 안 되는 환경에서는 인바운드 기능 포기가 종종 정답이다.</p>
</blockquote>

<h3 id="6차-oauth-콜백-url-등록">[6차] OAuth 콜백 URL 등록</h3>

<p>방향을 정리하고 Sentry Settings → Integrations → Slack → Add Workspace를 눌렀더니 다음 에러.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>redirect_uri did not match any configured URIs.
Passed URI: https://sentry.x.xxx.kr/extensions/slack/setup/
</code></pre></div></div>

<p>원인은 Slack App의 OAuth Redirect URLs에 Sentry 콜백 주소가 등록되어 있지 않았다는 것.</p>

<p>한 가지 흥미로운 지점 — <strong>OAuth 리다이렉트는 Slack 서버가 직접 우리 서버를 호출하는 게 아니라 사용자 브라우저를 우리 서버로 튕겨보내는 방식이다.</strong> 즉 브라우저 경유라 인바운드가 안 되는 내부망 환경에서도 정상 작동한다. 방향으로 보면 이것도 “우리가 원인 제공하고 사용자 브라우저가 우리한테 오는 것”에 가깝다.</p>

<p>Slack App → OAuth &amp; Permissions → Redirect URLs에 <code class="language-plaintext highlighter-rouge">https://sentry.x.xxx.kr/extensions/slack/setup/</code>를 추가하고 재시도하니 워크스페이스 연결 성공.</p>

<h3 id="7차-alert-rule-생성과-첫-발송-실패">[7차] Alert Rule 생성과 첫 발송 실패</h3>

<p>프로젝트 → Settings → Alerts → Create Alert Rule에서 “Issues” 템플릿 선택.</p>

<ul>
  <li><strong>WHEN</strong>: <code class="language-plaintext highlighter-rouge">A new issue is created</code></li>
  <li><strong>THEN</strong>: <code class="language-plaintext highlighter-rouge">Send a Slack notification</code> → 채널 <code class="language-plaintext highlighter-rouge">#&lt;프로젝트&gt;-alerts</code></li>
  <li><strong>Action interval</strong>: 24시간 (동일 이슈 도배 방지)</li>
  <li><strong>Owner</strong>: 프론트팀 (팀에서 직접 규칙 수정할 수 있도록)</li>
</ul>

<p>테스트 발송 시 다음 에러.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Slack: The resource "&lt;프로젝트&gt;-alerts" does not exist or has not been granted access
</code></pre></div></div>

<p>채널이 아직 Slack에 만들어져 있지 않았다. 채널을 공개(Public)로 새로 만들자 <code class="language-plaintext highlighter-rouge">chat:write.public</code> 스코프 덕에 봇 초대 없이도 발송이 됐다. 재테스트에서 실제 에러 알림(TypeError)이 채널에 도착했다.</p>

<hr />

<h2 id="왜-이렇게-걸렸나">왜 이렇게 걸렸나</h2>

<p>한 번에 정리하면 층별로 사건이 흩어져 있었다.</p>

<table>
  <thead>
    <tr>
      <th>층</th>
      <th>사건</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>벤더 정책</td>
      <td>1차 진입 — WebHooks 레거시 신규 설치 차단</td>
    </tr>
    <tr>
      <td>자격증명 반영</td>
      <td>1차 — <code class="language-plaintext highlighter-rouge">SENTRY_OPTIONS</code>에 등록 후 <code class="language-plaintext highlighter-rouge">sentry shell</code>로 검증</td>
    </tr>
    <tr>
      <td>Slack App URL·스코프</td>
      <td>2차 — 문서대로 세팅</td>
    </tr>
    <tr>
      <td><strong>접근 제어</strong></td>
      <td><strong>3~4차 — IP 화이트리스트에 막힌 인바운드</strong></td>
    </tr>
    <tr>
      <td><strong>판단</strong></td>
      <td><strong>5차 — 기능 포기가 정답</strong></td>
    </tr>
    <tr>
      <td>OAuth 콜백</td>
      <td>6차 — Redirect URL 누락</td>
    </tr>
    <tr>
      <td>Slack 채널 상태</td>
      <td>7차 — 채널 미생성</td>
    </tr>
  </tbody>
</table>

<p>문제 자체는 “Slack 통합 안 됨”이라는 하나였지만, 원인은 벤더 정책·네트워크·판단·설정 여러 층에 흩어져 있었다.</p>

<blockquote>
  <p>이전 글에서도 같은 감각이 있었다 — 증상 하나로 뭉쳐 있어도 원인은 층별로 다르게 발현된다.</p>
</blockquote>

<hr />

<h2 id="최종-설정-기록용">최종 설정 (기록용)</h2>

<p><code class="language-plaintext highlighter-rouge">sentry/sentry.conf.py</code></p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">SENTRY_OPTIONS</span><span class="p">[</span><span class="s">'slack.client-id'</span><span class="p">]</span>       <span class="o">=</span> <span class="s">'&lt;masked&gt;'</span>
<span class="n">SENTRY_OPTIONS</span><span class="p">[</span><span class="s">'slack.client-secret'</span><span class="p">]</span>   <span class="o">=</span> <span class="s">'&lt;masked&gt;'</span>
<span class="n">SENTRY_OPTIONS</span><span class="p">[</span><span class="s">'slack.signing-secret'</span><span class="p">]</span>  <span class="o">=</span> <span class="s">'&lt;masked&gt;'</span>
</code></pre></div></div>

<p>Slack App 설정</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>OAuth Redirect URLs:
  https://sentry.x.xxx.kr/extensions/slack/setup/

Bot Token Scopes:
  channels:read, chat:write, chat:write.customize, chat:write.public,
  commands, groups:read, im:history, im:read, links:read, links:write,
  team:read, users:read

Event Subscriptions : 비활성 (IP 화이트리스트로 인바운드 차단, 의도적 포기)
Slash Commands      : 비활성
Interactivity       : 비활성
</code></pre></div></div>

<p>Sentry Alert Rule</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>WHEN : A new issue is created
THEN : Send a Slack notification
       → 워크스페이스: &lt;사내 워크스페이스&gt;
       → 채널:       #&lt;프로젝트&gt;-alerts
Action interval : 24h
Owner           : 프론트팀
</code></pre></div></div>

<hr />

<h2 id="다음번-순서">다음번 순서</h2>

<p>같은 유형의 통합을 다시 붙일 때, 시작 전에 다음 순서를 따르면 이번 케이스의 3~4차 우회는 대부분 앞단에서 걸러진다.</p>

<ol>
  <li><strong>통합 페이지의 기능 목록을 방향별로 분류한다</strong> (아웃바운드 / 인바운드 / OAuth 리다이렉트). 우리 환경에서 되는 것과 안 되는 것을 시작 전에 갈라둔다.</li>
  <li><strong>원래 요구사항이 어느 방향만으로 충족되는지 확인한다.</strong> 아웃바운드로 되면 인바운드는 안 뚫는다. 노출은 편의가 아니라 보안 결정이다.</li>
  <li><strong>외부 도달성 검증은 반드시 진짜 외부 네트워크에서</strong> (개인 LTE, 외부 클라우드 인스턴스 등). 셀프 curl은 앱 처리만 확인해준다.</li>
  <li><strong>시크릿 반영은 <code class="language-plaintext highlighter-rouge">sentry shell</code>로 실제 로드 값까지 조회해서 검증한다.</strong> 재기동만으로 반영을 단정하지 않는다.</li>
</ol>

<hr />

<h3 id="참고">참고</h3>

<ul>
  <li><strong>환경</strong>: Sentry self-hosted (Docker Compose) · Nginx Proxy Manager · IP 화이트리스트 적용 도메인</li>
  <li><strong>관련 문서</strong>: <a href="https://docs.sentry.io/organization/integrations/notification-incidents/slack/">Sentry Slack Integration 공식 문서</a> · <a href="https://api.slack.com/apis/events-api">Slack API — Event Subscriptions</a> · <a href="https://github.com/getsentry/sentry/pull/91550">Sentry PR #91550 (레거시 플러그인 신규 설치 차단)</a></li>
  <li><strong>후속</strong>: 이 세션에서 파생된 권한 이슈(조직 Admin이 왜 신규 부여가 안 되는가, Team Admin과 Integrations 스코프)는 <a href="/sentry-permission-model/">다음 편</a>에</li>
</ul>]]></content><author><name>이진용</name><email>kouig14@gmail.com</email></author><category term="Sentry" /><category term="Slack" /><category term="네트워크" /><category term="트러블슈팅" /><summary type="html"><![CDATA[TL;DR 사내 self-hosted Sentry에 Slack 알림을 붙였다. WebHooks 레거시 플러그인은 2025년 중반부터 신규 설치가 막혀 있어 조직 설정의 Integrations를 써야 했다. Event Subscriptions URL 검증이 계속 튕겼고, 원인은 리버스 프록시의 IP 화이트리스트가 Slack의 요청을 앱 앞에서 막고 있었기 때문이다. (당시엔 “외부 경로 자체가 없다”로 판단했다가 이후 실측으로 정정했다 — §4 정정 참고.) 결론은 인바운드 3종(Event Subscriptions·Slash Command·Interactivity)을 의도적으로 포기하고 아웃바운드 알림만 유지 — 원래 요청은 그것만으로 충족된다.]]></summary></entry><entry><title type="html">리버스 프록시 뒤 Sentry에 Keycloak SSO를 붙이며 만난 CSRF 6단계</title><link href="https://stan-dev.cloud/sentry-csrf-reverse-proxy/" rel="alternate" type="text/html" title="리버스 프록시 뒤 Sentry에 Keycloak SSO를 붙이며 만난 CSRF 6단계" /><published>2026-07-26T21:00:00+09:00</published><updated>2026-07-26T21:00:00+09:00</updated><id>https://stan-dev.cloud/sentry-csrf-reverse-proxy</id><content type="html" xml:base="https://stan-dev.cloud/sentry-csrf-reverse-proxy/"><![CDATA[<blockquote>
  <p><strong>TL;DR</strong>
사내 Sentry(self-hosted)에 Keycloak OIDC를 연결하고 리버스 프록시로 HTTPS까지 붙였다.
도메인을 붙이는 순간 <strong>CSRF Validation Failed</strong>가 떨어졌고, 원인을 하나씩 벗겨내는 데 6단계가 걸렸다.
가장 큰 함정은 다섯 번째였다 — <code class="language-plaintext highlighter-rouge">sentry.conf.py</code>에 아무리 URL을 박아둬도 <strong><code class="language-plaintext highlighter-rouge">config.yml</code>의 <code class="language-plaintext highlighter-rouge">system.url-prefix</code>가 이긴다.</strong></p>
</blockquote>

<hr />

<h2 id="배경--왜-이-작업을-하고-있었나">배경 — 왜 이 작업을 하고 있었나</h2>

<p>인프라 엔지니어로 출근한 첫날 임무는 단순했다.</p>

<blockquote>
  <p>“PC 서버 하나에 Sentry 깔아봐.”</p>
</blockquote>

<p>세무 데이터를 다루는 회사라 외부 클라우드(SaaS Sentry)로 데이터를 못 뺀다. 그래서 self-hosted로 가고, 계정은 사내 Keycloak과 SSO로 묶고, 사내망에서 도메인으로 접근하도록 리버스 프록시(Nginx Proxy Manager) 뒤에 앉힌다 — 여기까지가 요구사항이었다.</p>

<p>세팅 자체는 순서대로 밟으면 어렵지 않았다. 문제는 <strong>도메인을 붙인 다음</strong>이었다.</p>

<hr />

<h2 id="어떤-문제였나">어떤 문제였나</h2>

<p>세팅을 다 끝내고 <code class="language-plaintext highlighter-rouge">https://sentry.x.xxx.kr</code>로 접속했더니 로그인 페이지 대신 이 화면이 떴다.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Forbidden (403)
CSRF verification failed. Request aborted.
</code></pre></div></div>

<p>Sentry 컨테이너까지는 정상, <code class="language-plaintext highlighter-rouge">192.168.x.x:9000</code>으로 직접 붙으면 로그인도 됐다. 도메인 + HTTPS를 거친 요청만 CSRF에서 튕겼다.</p>

<blockquote>
  <p>IP 직결은 되는데, 도메인은 안 된다.</p>

  <p>이 한 줄이 출발점이다.</p>
</blockquote>

<hr />

<h2 id="설정이-두-곳에-있다">설정이 두 곳에 있다</h2>

<p>CSRF 검증은 요청의 Origin이 신뢰 목록에 있는지 보는 것이다. Sentry(Django 기반)는 그 신뢰 기준을 두 파일에서 관리한다.</p>

<table>
  <thead>
    <tr>
      <th>파일</th>
      <th>담당</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">sentry/config.yml</code></td>
      <td><code class="language-plaintext highlighter-rouge">system.url-prefix</code> — Sentry가 자기 자신을 부르는 절대 URL</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">sentry/sentry.conf.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">CSRF_TRUSTED_ORIGINS</code> — CSRF 검증이 허용하는 Origin 목록</td>
    </tr>
  </tbody>
</table>

<p>이 둘이 한쪽이라도 도메인과 안 맞으면 CSRF에서 튕긴다. <strong>그리고 겹치는 부분에서 우선순위가 있다는 걸 나중에 알게 된다.</strong></p>

<hr />

<h2 id="6단계-추적기">6단계 추적기</h2>

<h3 id="1차-csrf_trusted_origins에-도메인이-없다">[1차] <code class="language-plaintext highlighter-rouge">CSRF_TRUSTED_ORIGINS</code>에 도메인이 없다</h3>

<p><strong>증상</strong> : 도메인 접속 시 CSRF 에러.</p>

<p><strong>원인</strong> : <code class="language-plaintext highlighter-rouge">sentry.conf.py</code>의 <code class="language-plaintext highlighter-rouge">CSRF_TRUSTED_ORIGINS</code>에 <code class="language-plaintext highlighter-rouge">sentry.x.xxx.kr</code>이 없었다.</p>

<p><strong>조치</strong> :</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">CSRF_TRUSTED_ORIGINS</span> <span class="o">=</span> <span class="p">[</span>
    <span class="s">"http://192.168.x.x:9000"</span><span class="p">,</span>
    <span class="s">"https://sentry.x.xxx.kr"</span><span class="p">,</span>
<span class="p">]</span>
</code></pre></div></div>

<p>여전히 튕겼다.</p>

<h3 id="2차-sentry가-자기를-http라고-오해한다">[2차] Sentry가 자기를 HTTP라고 오해한다</h3>

<p><strong>증상</strong> : Origin은 등록했는데도 CSRF 에러.</p>

<p><strong>원인</strong> : 요청 구조는 <code class="language-plaintext highlighter-rouge">유저 → HTTPS → nginx → HTTP → Sentry</code>다. Sentry 입장에선 자기가 HTTP로 서비스된다고 인식한다. 브라우저가 보낸 <code class="language-plaintext highlighter-rouge">Origin: https://sentry.x.xxx.kr</code>과 서버가 인식하는 <code class="language-plaintext highlighter-rouge">http://sentry.x.xxx.kr</code> 사이에 <strong>스킴이 어긋난다.</strong></p>

<p>리버스 프록시가 <code class="language-plaintext highlighter-rouge">X-Forwarded-Proto: https</code> 헤더로 “원래는 HTTPS야”라고 알려줘야 하는데, 중간 nginx 설정이 이 헤더를 <code class="language-plaintext highlighter-rouge">$scheme</code>(즉 http)로 덮어썼다.</p>

<p><strong>조치</strong> : Sentry 내부 <code class="language-plaintext highlighter-rouge">nginx.conf</code>에서 값을 강제로 박고, 애플리케이션 측이 그 헤더를 신뢰하도록 설정.</p>

<div class="language-nginx highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">proxy_set_header</span> <span class="s">X-Forwarded-Proto</span> <span class="s">https</span><span class="p">;</span>
</code></pre></div></div>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">SECURE_PROXY_SSL_HEADER</span> <span class="o">=</span> <span class="p">(</span><span class="s">'HTTP_X_FORWARDED_PROTO'</span><span class="p">,</span> <span class="s">'https'</span><span class="p">)</span>
<span class="n">SOCIAL_AUTH_REDIRECT_IS_HTTPS</span> <span class="o">=</span> <span class="bp">True</span>
</code></pre></div></div>

<p>여기서 CSRF는 넘어갔다. 이제 로그인은 되는데 — 곧바로 세션이 풀렸다.</p>

<h3 id="3차-secure-쿠키가-잘린다">[3차] <code class="language-plaintext highlighter-rouge">Secure</code> 쿠키가 잘린다</h3>

<p><strong>증상</strong> : 로그인 성공 → 새로고침 → 다시 로그인 화면.</p>

<p><strong>원인</strong> : <code class="language-plaintext highlighter-rouge">SESSION_COOKIE_SECURE = True</code>, <code class="language-plaintext highlighter-rouge">CSRF_COOKIE_SECURE = True</code>가 켜져 있었다. 이 옵션이 켜지면 쿠키는 <strong>HTTPS 연결에서만</strong> 전송된다. 그런데 nginx 뒤 Sentry는 내부적으로 HTTP로 통신해서, 응답에 실린 <code class="language-plaintext highlighter-rouge">Secure</code> 쿠키가 브라우저까지 도달은 하지만 다음 요청에서 다시 붙을 때 애플리케이션 서버 앞에서 잘렸다.</p>

<p><strong>조치</strong> : 두 옵션을 주석 처리.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># SESSION_COOKIE_SECURE = True
# CSRF_COOKIE_SECURE = True
</span></code></pre></div></div>

<blockquote>
  <p>“HTTPS니까 켜두자”의 직관은 리버스 프록시 뒤에선 함정이다. 브라우저는 HTTPS로 오지만 서버 사이는 HTTP다.</p>
</blockquote>

<h3 id="4차-sed-치환이-따옴표를-날려버렸다">[4차] <code class="language-plaintext highlighter-rouge">sed</code> 치환이 따옴표를 날려버렸다</h3>

<p><strong>증상</strong> : <code class="language-plaintext highlighter-rouge">CSRF_TRUSTED_ORIGINS</code>를 고쳤는데 재기동 후에도 반영이 안 됨.</p>

<p><strong>원인</strong> : 배포 스크립트에서 <code class="language-plaintext highlighter-rouge">sed</code>로 값을 넣었는데, 이스케이프를 잘못 짜서 파이썬 문자열의 양쪽 따옴표가 사라졌다. 파이썬 파서가 문자열로 인식을 못 해 아예 로드에 실패했다.</p>

<p><strong>조치</strong> : <code class="language-plaintext highlighter-rouge">sed</code> 치환을 걷어내고 파이썬 스크립트로 정확히 다시 썼다. 로그가 흘려버릴 뻔한 실패였는데, 컨테이너 로그에서 파싱 에러 한 줄로 잡아냈다.</p>

<h3 id="5차--핵심-configyml이-sentryconfpy를-이긴다">[5차 — 핵심] <code class="language-plaintext highlighter-rouge">config.yml</code>이 <code class="language-plaintext highlighter-rouge">sentry.conf.py</code>를 이긴다</h3>

<p><strong>증상</strong> : <code class="language-plaintext highlighter-rouge">sentry.conf.py</code>에서 <code class="language-plaintext highlighter-rouge">SENTRY_OPTIONS["system.url-prefix"]</code>를 <code class="language-plaintext highlighter-rouge">https://sentry.x.xxx.kr</code>로 지정했는데도, 로그인 리다이렉트가 계속 IP + 포트(<code class="language-plaintext highlighter-rouge">http://192.168.x.x:9000/...</code>)로 나갔다. 브라우저는 도메인으로 접속하는데 응답이 IP로 튕겨내는 상황.</p>

<p><strong>원인</strong> : <code class="language-plaintext highlighter-rouge">sentry/config.yml</code>에 <code class="language-plaintext highlighter-rouge">system.url-prefix</code> 값이 하드코딩돼 있었고, 이 값이 <code class="language-plaintext highlighter-rouge">sentry.conf.py</code>보다 <strong>우선</strong>으로 로드됐다.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>config.yml  &gt;  sentry.conf.py
</code></pre></div></div>

<p>이걸 몰라서 파이썬 파일만 계속 고쳤다. <code class="language-plaintext highlighter-rouge">SENTRY_OPTIONS</code>를 아무리 덮어써도 원래 값이 살아 있는 이유가 이거였다.</p>

<p>어떻게 알아냈냐면 — 문서가 아니라 동작이었다. 파이썬 파일을 고치고 재기동해도 리다이렉트 URL이 안 바뀌길래, <code class="language-plaintext highlighter-rouge">config.yml</code> 쪽 값을 바꿔봤더니 그제야 바뀌었다. 두 파일이 같은 키를 갖고 있을 때 어느 쪽이 이기는지는 실행해봐야 보이는 순서였다. “왜 값이 안 먹지”에 붙잡히기 전에 “지금 어디서 이 값을 읽고 있지”를 먼저 물었어야 했다.</p>

<p><strong>조치</strong> : <code class="language-plaintext highlighter-rouge">config.yml</code>도 함께 수정.</p>

<div class="language-yaml highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="na">system.url-prefix</span><span class="pi">:</span> <span class="s1">'</span><span class="s">https://sentry.x.xxx.kr'</span>
</code></pre></div></div>

<blockquote>
  <p>설정을 두 곳에서 관리하는 시스템은 <strong>어디가 이기는지</strong>를 먼저 확인해야 한다. 우선순위를 모르면 아래쪽만 계속 고치게 된다.</p>
</blockquote>

<h3 id="6차-realm을-잘못-잡았다">[6차] Realm을 잘못 잡았다</h3>

<p><strong>증상</strong> : Sentry에서 “Keycloak으로 로그인”을 누르면 로그인 페이지까지는 가는데, 사내 계정이 인식되지 않았다.</p>

<p><strong>원인</strong> : Keycloak 관리 콘솔에서 <code class="language-plaintext highlighter-rouge">master</code> realm의 <code class="language-plaintext highlighter-rouge">sentry</code> 클라이언트를 수정했다. 실제 사내 계정은 <code class="language-plaintext highlighter-rouge">TEAM-&lt;사내&gt;</code> realm에 있었다. 좌측 상단 realm 드롭다운을 못 봤다.</p>

<p><strong>조치</strong> : <code class="language-plaintext highlighter-rouge">TEAM-&lt;사내&gt;</code> realm의 <code class="language-plaintext highlighter-rouge">sentry</code> 클라이언트에서 값을 다시 세팅.</p>

<ul>
  <li>Redirect URI: <code class="language-plaintext highlighter-rouge">https://sentry.x.xxx.kr/auth/sso/*</code></li>
  <li>Web Origins: <code class="language-plaintext highlighter-rouge">+</code></li>
  <li><code class="language-plaintext highlighter-rouge">sentry.conf.py</code>의 <code class="language-plaintext highlighter-rouge">OIDC_DOMAIN</code>도 <code class="language-plaintext highlighter-rouge">https://key.x.xxx.kr/realms/TEAM-&lt;사내&gt;</code>으로 정정.</li>
</ul>

<hr />

<h2 id="왜-이렇게-많은-단계가-필요했나">왜 이렇게 많은 단계가 필요했나</h2>

<p>한 번에 정리해 보면 여섯 개가 서로 다른 층에서 발생했다.</p>

<table>
  <thead>
    <tr>
      <th>층</th>
      <th>사건</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>애플리케이션 (신뢰 목록)</td>
      <td>1차 — <code class="language-plaintext highlighter-rouge">CSRF_TRUSTED_ORIGINS</code> 누락</td>
    </tr>
    <tr>
      <td>프록시 (헤더 전달)</td>
      <td>2차 — <code class="language-plaintext highlighter-rouge">X-Forwarded-Proto</code> 덮어쓰기</td>
    </tr>
    <tr>
      <td>브라우저 ↔ 서버 (쿠키 정책)</td>
      <td>3차 — <code class="language-plaintext highlighter-rouge">Secure</code> 쿠키 충돌</td>
    </tr>
    <tr>
      <td>배포 도구 (파일 편집)</td>
      <td>4차 — <code class="language-plaintext highlighter-rouge">sed</code> 이스케이프 실패</td>
    </tr>
    <tr>
      <td>설정 우선순위</td>
      <td>5차 — <code class="language-plaintext highlighter-rouge">config.yml</code> 대 <code class="language-plaintext highlighter-rouge">sentry.conf.py</code></td>
    </tr>
    <tr>
      <td>IAM (realm 격리)</td>
      <td>6차 — realm 오인</td>
    </tr>
  </tbody>
</table>

<p>같은 “CSRF 에러”라는 증상이 5개 층을 걸쳐 다르게 발현됐다. 하나의 로그 메시지만 보면 여섯 개가 다 똑같이 보인다.</p>

<blockquote>
  <p>증상은 하나여도 원인은 층별로 다르다. 로그 메시지가 아니라 <strong>구조</strong>를 봐야 한다.</p>
</blockquote>

<hr />

<h2 id="최종-설정-기록용">최종 설정 (기록용)</h2>

<p><code class="language-plaintext highlighter-rouge">sentry/sentry.conf.py</code></p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">SENTRY_OPTIONS</span><span class="p">[</span><span class="s">"system.url-prefix"</span><span class="p">]</span> <span class="o">=</span> <span class="s">"https://sentry.x.xxx.kr"</span>

<span class="n">CSRF_TRUSTED_ORIGINS</span> <span class="o">=</span> <span class="p">[</span>
    <span class="s">"http://192.168.x.x:9000"</span><span class="p">,</span>
    <span class="s">"https://sentry.x.xxx.kr"</span><span class="p">,</span>
<span class="p">]</span>

<span class="n">SOCIAL_AUTH_REDIRECT_IS_HTTPS</span> <span class="o">=</span> <span class="bp">True</span>
<span class="n">SECURE_PROXY_SSL_HEADER</span> <span class="o">=</span> <span class="p">(</span><span class="s">'HTTP_X_FORWARDED_PROTO'</span><span class="p">,</span> <span class="s">'https'</span><span class="p">)</span>

<span class="c1"># 리버스 프록시 뒤 내부 HTTP 구간과 충돌하므로 켜지 않는다.
# SESSION_COOKIE_SECURE = True
# CSRF_COOKIE_SECURE = True
</span>
<span class="n">OIDC_CLIENT_ID</span> <span class="o">=</span> <span class="s">"sentry"</span>
<span class="n">OIDC_CLIENT_SECRET</span> <span class="o">=</span> <span class="s">"&lt;masked&gt;"</span>
<span class="n">OIDC_SCOPE</span> <span class="o">=</span> <span class="s">"openid email profile"</span>
<span class="n">OIDC_DOMAIN</span> <span class="o">=</span> <span class="s">"https://key.x.xxx.kr/realms/TEAM-&lt;사내&gt;"</span>
<span class="n">OIDC_ISSUER</span> <span class="o">=</span> <span class="s">"Keycloak"</span>
</code></pre></div></div>

<p><code class="language-plaintext highlighter-rouge">sentry/config.yml</code></p>

<div class="language-yaml highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="na">system.url-prefix</span><span class="pi">:</span> <span class="s1">'</span><span class="s">https://sentry.x.xxx.kr'</span>
</code></pre></div></div>

<p><code class="language-plaintext highlighter-rouge">sentry/nginx.conf</code> (Sentry 내부 프록시)</p>

<div class="language-nginx highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">proxy_set_header</span> <span class="s">Host</span> <span class="s">sentry.x.xxx.kr</span><span class="p">;</span>
<span class="k">proxy_set_header</span> <span class="s">X-Forwarded-For</span> <span class="nv">$remote_addr</span><span class="p">;</span>
<span class="k">proxy_set_header</span> <span class="s">X-Forwarded-Proto</span> <span class="s">https</span><span class="p">;</span>
<span class="k">proxy_set_header</span> <span class="s">X-Request-Id</span> <span class="nv">$request_id</span><span class="p">;</span>
</code></pre></div></div>

<hr />

<h2 id="만약-다시-짠다면">만약 다시 짠다면</h2>

<p>지금 이 세팅을 다시 한다면 두 가지를 순서에 끼워 넣을 것이다.</p>

<ul>
  <li><strong>도메인 붙이기 전에 <code class="language-plaintext highlighter-rouge">config.yml</code>부터 열어본다.</strong> 이번엔 IP 접근으로 세팅을 다 끝내고 나서 마지막에 도메인을 붙였는데, 그 순서가 5차 사건의 원인이었다. 애초에 <code class="language-plaintext highlighter-rouge">system.url-prefix</code>를 최종 도메인으로 박은 상태에서 세팅을 진행했으면 CSRF의 절반은 아예 발생하지 않았을 것이다.</li>
  <li><strong>리버스 프록시 뒤 세팅은 로컬에서 한 번 그림을 그린 뒤 붙인다.</strong> <code class="language-plaintext highlighter-rouge">유저 → HTTPS → nginx → HTTP → Sentry</code> 구조에서 스킴이 바뀌는 지점이 어디인지, 쿠키가 어디까지 실려 가는지, <code class="language-plaintext highlighter-rouge">X-Forwarded-*</code>가 어디서 세팅·덮어써지는지를 종이 위에 한 번 정리하면 2차·3차·5차 사건은 대부분 사전에 걸러진다.</li>
</ul>

<p>다만 그건 “리버스 프록시 뒤 서비스”라는 구조가 익숙해진 다음의 이야기다. 이번엔 그 구조 자체가 처음이라 층별로 하나씩 부딪히면서 배운 셈이다. <strong>명령어는 잊혀지겠지만, “이 문제가 왜 생기고 어디를 봐야 하는가”의 감각은 남을 것이다.</strong></p>

<hr />

<h3 id="참고">참고</h3>

<ul>
  <li><strong>환경</strong> : Sentry self-hosted 26.6.0 · Docker Compose 20+ 컨테이너 · Ubuntu Server · Nginx Proxy Manager</li>
  <li><strong>연동</strong> : Keycloak (OIDC, Authorization Code Flow) · <code class="language-plaintext highlighter-rouge">sentry-auth-oidc</code> 플러그인</li>
  <li><strong>관련 문서</strong> : <a href="https://develop.sentry.dev/self-hosted/">Sentry Self-Hosted</a> · <a href="https://github.com/siemens/sentry-auth-oidc">sentry-auth-oidc</a> · <a href="https://docs.djangoproject.com/en/4.2/ref/settings/#std-setting-CSRF_TRUSTED_ORIGINS">Django <code class="language-plaintext highlighter-rouge">CSRF_TRUSTED_ORIGINS</code></a></li>
</ul>]]></content><author><name>이진용</name><email>kouig14@gmail.com</email></author><category term="Sentry" /><category term="Keycloak" /><category term="Nginx" /><category term="OIDC" /><category term="트러블슈팅" /><summary type="html"><![CDATA[TL;DR 사내 Sentry(self-hosted)에 Keycloak OIDC를 연결하고 리버스 프록시로 HTTPS까지 붙였다. 도메인을 붙이는 순간 CSRF Validation Failed가 떨어졌고, 원인을 하나씩 벗겨내는 데 6단계가 걸렸다. 가장 큰 함정은 다섯 번째였다 — sentry.conf.py에 아무리 URL을 박아둬도 config.yml의 system.url-prefix가 이긴다.]]></summary></entry><entry><title type="html">diff 기반 업데이트로 saveSettings를 다시 짠다면</title><link href="https://stan-dev.cloud/diff-based-update/" rel="alternate" type="text/html" title="diff 기반 업데이트로 saveSettings를 다시 짠다면" /><published>2026-06-23T11:07:00+09:00</published><updated>2026-06-23T11:07:00+09:00</updated><id>https://stan-dev.cloud/diff-based-update</id><content type="html" xml:base="https://stan-dev.cloud/diff-based-update/"><![CDATA[<blockquote>
  <p>단기 처방 <code class="language-plaintext highlighter-rouge">flush()</code>로 막아두고 PR에 약속해뒀던 “다음 스프린트의 리팩토링”을 어떤 모양으로 닫을지 설계로 정리한 글</p>
</blockquote>

<p>전편에서 단기로는 명시적 <code class="language-plaintext highlighter-rouge">flush()</code>, 장기로는 <strong>“전체 삭제 후 재삽입”을 그만두고 변경분만 처리하는 diff 기반 업데이트로 전환할 것</strong>이라고 <a href="https://github.com/SWYP-Backend/BangCheck/pull/179">PR #179</a>와 코드 주석에 적어뒀다.</p>

<p>그 약속은 닫지 못했다. 팀 프로젝트가 5월에 종료돼 이 개선은 실 적용까지 가지 못했고, 그래서 이 글은 구현기가 아니라 <strong>설계 기록</strong>이다 — 어떤 모양이어야 같은 사건이 다시 안 돌아오는지를 남겨둔다.</p>

<hr />

<h2 id="전편-요약--무엇이-한계로-남아있었나">전편 요약 — 무엇이 한계로 남아있었나</h2>

<p>전편의 코드는 단순한 “전체 삭제 후 재삽입”이었다.</p>

<div class="language-java highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nd">@Transactional</span>
<span class="kd">public</span> <span class="kt">void</span> <span class="nf">saveSettings</span><span class="o">(</span><span class="nc">Long</span> <span class="n">userId</span><span class="o">,</span> <span class="nc">List</span><span class="o">&lt;</span><span class="nc">SettingRequest</span><span class="o">&gt;</span> <span class="n">requests</span><span class="o">)</span> <span class="o">{</span>
    <span class="n">settingRepository</span><span class="o">.</span><span class="na">deleteByUserId</span><span class="o">(</span><span class="n">userId</span><span class="o">);</span>
    <span class="n">settingRepository</span><span class="o">.</span><span class="na">flush</span><span class="o">();</span>   <span class="c1">// 단기 처방 — DELETE를 먼저 DB로</span>

    <span class="nc">List</span><span class="o">&lt;</span><span class="nc">UserChecklistSetting</span><span class="o">&gt;</span> <span class="n">settings</span> <span class="o">=</span> <span class="n">requests</span><span class="o">.</span><span class="na">stream</span><span class="o">()</span>
            <span class="o">.</span><span class="na">map</span><span class="o">(</span><span class="n">req</span> <span class="o">-&gt;</span> <span class="nc">UserChecklistSetting</span><span class="o">.</span><span class="na">of</span><span class="o">(</span><span class="n">userId</span><span class="o">,</span> <span class="n">req</span><span class="o">))</span>
            <span class="o">.</span><span class="na">toList</span><span class="o">();</span>

    <span class="n">settingRepository</span><span class="o">.</span><span class="na">saveAll</span><span class="o">(</span><span class="n">settings</span><span class="o">);</span>
<span class="o">}</span>
</code></pre></div></div>

<p><code class="language-plaintext highlighter-rouge">flush()</code>로 운영 장애는 막혔지만, 코드는 여전히 <strong>Hibernate ActionQueue 동작 순서에 의존</strong>하고 있다. 이 코드는 다음 세 가지 함정을 그대로 갖고 있다.</p>

<table>
  <thead>
    <tr>
      <th>함정</th>
      <th>이유</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>① 매 저장마다 N행 DELETE + N행 INSERT</td>
      <td>변경 없는 항목까지 다시 쓴다</td>
    </tr>
    <tr>
      <td>② <code class="language-plaintext highlighter-rouge">flush()</code> 한 줄을 무심코 지우면 다시 409</td>
      <td>“왜 flush가 있는지”가 코드만 보고는 안 보임</td>
    </tr>
    <tr>
      <td>③ 감사 로그·트리거가 변경 없는 행에도 발화</td>
      <td>DB 쪽 부수효과를 헷갈리게 만듦</td>
    </tr>
  </tbody>
</table>

<blockquote>
  <p>단기 처방은 운영 장애를 멈추기 위한 것이지, <em>코드를 안전하게 만든 것은 아니다</em>.</p>
</blockquote>

<hr />

<h2 id="delete-insert가-잘못이라고-단정하기-전에--왜-처음엔-그렇게-짰나">delete-insert가 잘못이라고 단정하기 전에 — 왜 처음엔 그렇게 짰나</h2>

<p>리팩토링 글은 “예전 코드는 틀렸고 새 코드는 옳다”는 톤으로 빠지기 쉽다. 그건 정직하지 않다. 처음 delete-insert를 택한 데에는 이유가 있었다.</p>

<ul>
  <li>저장 요청마다 활성/비활성 항목 수가 달라진다 — 단순 UPDATE로 풀 수 없다</li>
  <li>항목별 diff를 그때그때 짜는 게 초기에는 복잡했다</li>
  <li>“전부 지우고 다시 박는다”는 시그니처가 <em>동기화 깨질 여지를 적게</em> 보였다</li>
</ul>

<p>이 판단 자체는 첫 구현 시점에선 합리적이었다. 다만 그 다음에 두 가지를 같이 결정해뒀어야 했다.</p>

<ol>
  <li>UNIQUE 제약을 가진 테이블에 “delete 후 insert” 패턴은 <strong>ORM 동작 순서가 부메랑이 될 수 있는 모양</strong>이다.</li>
  <li>그 모양을 알았다면 <em>처음부터 diff 기반으로 짤 것인지, 단기 단순화로 갈 것인지</em>를 명시적으로 결정해두는 게 맞다.</li>
</ol>

<p>전편 PR에서 <em>“단기로 flush, 장기로 diff”</em> 라고 분리해서 남겨둔 이유가 이거다. 결정을 미룬 게 아니라, <strong>결정을 미룬다는 사실을 결정해두는</strong> 절차다.</p>

<hr />

<h2 id="diff-기반-업데이트가-하는-일--세-단계로-쪼개기">diff 기반 업데이트가 하는 일 — 세 단계로 쪼개기</h2>

<p>“변경분만 처리한다”는 말은 코드로 옮기기 전엔 단순해 보인다. 실제로 짜보면 <strong>현재 상태 조회 → 차집합 계산(교집합은 손대지 않는다) → 최소한의 INSERT/DELETE</strong> 세 단계가 명확히 분리돼 있어야 테스트가 가능해진다.</p>

<div class="language-java highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nd">@Transactional</span>
<span class="kd">public</span> <span class="kt">void</span> <span class="nf">saveSettings</span><span class="o">(</span><span class="nc">Long</span> <span class="n">userId</span><span class="o">,</span> <span class="nc">List</span><span class="o">&lt;</span><span class="nc">SettingRequest</span><span class="o">&gt;</span> <span class="n">requests</span><span class="o">)</span> <span class="o">{</span>
    <span class="c1">// 1단계 — 현재 상태 조회</span>
    <span class="nc">Set</span><span class="o">&lt;</span><span class="nc">Long</span><span class="o">&gt;</span> <span class="n">currentItemIds</span> <span class="o">=</span>
            <span class="n">settingRepository</span><span class="o">.</span><span class="na">findItemIdsByUserId</span><span class="o">(</span><span class="n">userId</span><span class="o">);</span>

    <span class="c1">// 2단계 — 차집합 계산</span>
    <span class="nc">Set</span><span class="o">&lt;</span><span class="nc">Long</span><span class="o">&gt;</span> <span class="n">requestedItemIds</span> <span class="o">=</span> <span class="n">requests</span><span class="o">.</span><span class="na">stream</span><span class="o">()</span>
            <span class="o">.</span><span class="na">map</span><span class="o">(</span><span class="nl">SettingRequest:</span><span class="o">:</span><span class="n">itemId</span><span class="o">)</span>
            <span class="o">.</span><span class="na">collect</span><span class="o">(</span><span class="nc">Collectors</span><span class="o">.</span><span class="na">toSet</span><span class="o">());</span>

    <span class="nc">Set</span><span class="o">&lt;</span><span class="nc">Long</span><span class="o">&gt;</span> <span class="n">toAdd</span> <span class="o">=</span> <span class="k">new</span> <span class="nc">HashSet</span><span class="o">&lt;&gt;(</span><span class="n">requestedItemIds</span><span class="o">);</span>
    <span class="n">toAdd</span><span class="o">.</span><span class="na">removeAll</span><span class="o">(</span><span class="n">currentItemIds</span><span class="o">);</span>

    <span class="nc">Set</span><span class="o">&lt;</span><span class="nc">Long</span><span class="o">&gt;</span> <span class="n">toRemove</span> <span class="o">=</span> <span class="k">new</span> <span class="nc">HashSet</span><span class="o">&lt;&gt;(</span><span class="n">currentItemIds</span><span class="o">);</span>
    <span class="n">toRemove</span><span class="o">.</span><span class="na">removeAll</span><span class="o">(</span><span class="n">requestedItemIds</span><span class="o">);</span>

    <span class="c1">// 3단계 — 최소한의 INSERT/DELETE만</span>
    <span class="k">if</span> <span class="o">(!</span><span class="n">toRemove</span><span class="o">.</span><span class="na">isEmpty</span><span class="o">())</span> <span class="o">{</span>
        <span class="n">settingRepository</span><span class="o">.</span><span class="na">deleteByUserIdAndItemIdIn</span><span class="o">(</span><span class="n">userId</span><span class="o">,</span> <span class="n">toRemove</span><span class="o">);</span>
    <span class="o">}</span>
    <span class="k">if</span> <span class="o">(!</span><span class="n">toAdd</span><span class="o">.</span><span class="na">isEmpty</span><span class="o">())</span> <span class="o">{</span>
        <span class="nc">List</span><span class="o">&lt;</span><span class="nc">UserChecklistSetting</span><span class="o">&gt;</span> <span class="n">newSettings</span> <span class="o">=</span> <span class="n">toAdd</span><span class="o">.</span><span class="na">stream</span><span class="o">()</span>
                <span class="o">.</span><span class="na">map</span><span class="o">(</span><span class="n">itemId</span> <span class="o">-&gt;</span> <span class="nc">UserChecklistSetting</span><span class="o">.</span><span class="na">of</span><span class="o">(</span><span class="n">userId</span><span class="o">,</span> <span class="n">itemId</span><span class="o">))</span>
                <span class="o">.</span><span class="na">toList</span><span class="o">();</span>
        <span class="n">settingRepository</span><span class="o">.</span><span class="na">saveAll</span><span class="o">(</span><span class="n">newSettings</span><span class="o">);</span>
    <span class="o">}</span>
<span class="o">}</span>
</code></pre></div></div>

<p>여기서 중요한 건 코드 라인 수가 늘어난 게 아니라, <strong>세 단계가 서로 독립적으로 테스트 가능한 형태</strong>가 됐다는 점이다. 차집합 계산은 순수 함수(<code class="language-plaintext highlighter-rouge">Set</code> 연산)라 도메인 의존성이 없고, INSERT/DELETE는 비어있는 컬렉션을 받았을 때 안 나가도록 분리된다.</p>

<hr />

<h2 id="actionqueue-함정이-사라지는-이유">ActionQueue 함정이 사라지는 이유</h2>

<p>새 코드는 <code class="language-plaintext highlighter-rouge">flush()</code>를 쓰지 않는다. 그 이유를 한 줄로 정리하면 이렇다.</p>

<blockquote>
  <p>같은 <code class="language-plaintext highlighter-rouge">(user_id, item_id)</code> 키로 DELETE와 INSERT가 동시에 큐에 쌓이는 모양 자체가 사라진다.</p>
</blockquote>

<p><code class="language-plaintext highlighter-rouge">toRemove</code>와 <code class="language-plaintext highlighter-rouge">toAdd</code>는 차집합이라 교집합이 없다. 같은 키가 두 큐에 동시에 들어갈 수 없다. ActionQueue가 INSERT를 먼저 실행하든 DELETE를 먼저 실행하든 결과가 동일하다.</p>

<p>이게 단기 처방과 장기 개선의 결정적 차이다.</p>

<table>
  <thead>
    <tr>
      <th>구분</th>
      <th>단기 처방 (<code class="language-plaintext highlighter-rouge">flush()</code>)</th>
      <th>장기 개선 (diff)</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>ActionQueue 순서 영향</td>
      <td>우회한다</td>
      <td><strong>영향을 받지 않는 모양으로 바꾼다</strong></td>
    </tr>
    <tr>
      <td>코드에서 의도 가시성</td>
      <td>주석 필요</td>
      <td>코드 자체로 자명</td>
    </tr>
    <tr>
      <td>변경 없는 행 처리</td>
      <td>매번 DELETE → INSERT</td>
      <td>손대지 않는다</td>
    </tr>
    <tr>
      <td>감사 로그·트리거</td>
      <td>매번 발화</td>
      <td>변경분에만 발화</td>
    </tr>
  </tbody>
</table>

<blockquote>
  <p>ORM 내부 동작에 의존하는 코드를 작성하는 것보다, <em>그 동작이 영향을 줄 수 없는 모양</em>으로 코드를 바꾸는 게 더 안전하다.</p>
</blockquote>

<hr />

<h2 id="검증-설계--어떤-케이스를-테스트로-박을-것인가">검증 설계 — 어떤 케이스를 테스트로 박을 것인가</h2>

<p>diff 기반 업데이트는 <em>분기</em>가 많다. 단위 테스트는 다음 다섯 케이스로 설계했다.</p>

<table>
  <thead>
    <tr>
      <th>#</th>
      <th>시나리오</th>
      <th>기대 동작</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>1</td>
      <td>처음 저장 (현재 0개 → 요청 N개)</td>
      <td>INSERT N건, DELETE 0건</td>
    </tr>
    <tr>
      <td>2</td>
      <td>전부 그대로 (현재 = 요청)</td>
      <td><strong>쿼리 0건</strong></td>
    </tr>
    <tr>
      <td>3</td>
      <td>일부만 추가 (현재 ⊂ 요청)</td>
      <td>INSERT 차집합만, DELETE 0건</td>
    </tr>
    <tr>
      <td>4</td>
      <td>일부만 제거 (요청 ⊂ 현재)</td>
      <td>INSERT 0건, DELETE 차집합만</td>
    </tr>
    <tr>
      <td>5</td>
      <td>첫 사건 재현 (같은 요청 두 번)</td>
      <td><strong>두 번째 호출 = 쿼리 0건, 409 안 남</strong></td>
    </tr>
  </tbody>
</table>

<p>특히 5번 — 전편에서 운영 장애를 낸 정확한 시나리오 — 가 새 구조에선 <strong>두 번째 호출에서 아무 쿼리도 나가지 않는</strong> 상태가 된다. ActionQueue가 끼어들 여지 자체가 없다.</p>

<p>5번 케이스는 이런 형태다 — Hibernate <code class="language-plaintext highlighter-rouge">Statistics</code>로 쿼리 수를 직접 센다.</p>

<div class="language-java highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nd">@Test</span>
<span class="nd">@DisplayName</span><span class="o">(</span><span class="s">"같은 요청을 두 번 보내도 두 번째 호출은 쿼리 0건"</span><span class="o">)</span>
<span class="kt">void</span> <span class="nf">saveSettings_idempotent_secondCallProducesNoQueries</span><span class="o">()</span> <span class="o">{</span>
    <span class="c1">// given</span>
    <span class="nc">List</span><span class="o">&lt;</span><span class="nc">SettingRequest</span><span class="o">&gt;</span> <span class="n">sameRequest</span> <span class="o">=</span> <span class="nc">List</span><span class="o">.</span><span class="na">of</span><span class="o">(</span>
            <span class="k">new</span> <span class="nf">SettingRequest</span><span class="o">(</span><span class="mi">1L</span><span class="o">),</span>
            <span class="k">new</span> <span class="nf">SettingRequest</span><span class="o">(</span><span class="mi">2L</span><span class="o">),</span>
            <span class="k">new</span> <span class="nf">SettingRequest</span><span class="o">(</span><span class="mi">3L</span><span class="o">)</span>
    <span class="o">);</span>
    <span class="n">checklistService</span><span class="o">.</span><span class="na">saveSettings</span><span class="o">(</span><span class="n">userId</span><span class="o">,</span> <span class="n">sameRequest</span><span class="o">);</span>
    <span class="n">statistics</span><span class="o">.</span><span class="na">clear</span><span class="o">();</span>

    <span class="c1">// when</span>
    <span class="n">checklistService</span><span class="o">.</span><span class="na">saveSettings</span><span class="o">(</span><span class="n">userId</span><span class="o">,</span> <span class="n">sameRequest</span><span class="o">);</span>

    <span class="c1">// then</span>
    <span class="n">assertThat</span><span class="o">(</span><span class="n">statistics</span><span class="o">.</span><span class="na">getPrepareStatementCount</span><span class="o">())</span>
            <span class="o">.</span><span class="na">as</span><span class="o">(</span><span class="s">"두 번째 호출은 현재 상태 조회 한 번 외에 INSERT/DELETE 0건"</span><span class="o">)</span>
            <span class="o">.</span><span class="na">isEqualTo</span><span class="o">(</span><span class="mi">1L</span><span class="o">);</span> <span class="c1">// SELECT만 1건</span>
<span class="o">}</span>
</code></pre></div></div>

<blockquote>
  <p>같은 입력에 같은 결과를 내는 것을 <em>함수의 멱등성</em>이라고 부른다면, 저장 API에 멱등성을 부여하는 셈이다.</p>
</blockquote>

<hr />

<h2 id="before--after-쿼리-패턴-비교">Before / After 쿼리 패턴 비교</h2>

<p>아래는 쿼리 패턴 기준으로 계산한 값이다(N=10). 실측이 아니라 설계 단계의 비교이고, 실측은 Hibernate <code class="language-plaintext highlighter-rouge">Statistics</code>로 같은 표를 채우면 된다.</p>

<table>
  <thead>
    <tr>
      <th>시나리오</th>
      <th>Before (delete-insert + flush)</th>
      <th>After (diff)</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>첫 저장 (N=10)</td>
      <td>DELETE 1 + INSERT 10 = <strong>11</strong></td>
      <td>SELECT 1 + INSERT 10 = <strong>11</strong></td>
    </tr>
    <tr>
      <td>같은 요청 두 번째</td>
      <td>DELETE 1 + INSERT 10 = <strong>11</strong></td>
      <td>SELECT 1 = <strong>1</strong></td>
    </tr>
    <tr>
      <td>일부만 추가 (10→12)</td>
      <td>DELETE 1 + INSERT 12 = <strong>13</strong></td>
      <td>SELECT 1 + INSERT 2 = <strong>3</strong></td>
    </tr>
    <tr>
      <td>일부만 제거 (10→8)</td>
      <td>DELETE 1 + INSERT 8 = <strong>9</strong></td>
      <td>SELECT 1 + DELETE 1 = <strong>2</strong></td>
    </tr>
  </tbody>
</table>

<p>전반적으로 변경이 적을수록 쿼리가 적게 나간다. 운영 트래픽 패턴이 “조금씩 자주 저장”하는 형태라면 부하 차이가 더 벌어진다.</p>

<hr />

<h2 id="만약-지금-짠다면-어디까지-갈-것인가">만약 지금 짠다면 어디까지 갈 것인가</h2>

<p>이 설계 위에 두 단계 더 갈 수 있다.</p>

<ul>
  <li><strong>API 시그니처를 도메인 차원에서 다시 정의한다.</strong> 지금은 <em>“활성 항목 전체 목록”</em> 을 받는 시그니처라 클라이언트가 매번 전체 리스트를 보낸다. <em>“켜기 / 끄기 단일 항목”</em> 시그니처로 분리하면 변경 범위가 더 명확해지고 트래픽도 가벼워진다. 다만 “여러 항목을 한 번에 토글하는 UX”가 필요하면 현 시그니처가 더 자연스럽다 — 도메인 요구 우선.</li>
  <li><strong><code class="language-plaintext highlighter-rouge">@DataJpaTest</code>로 ActionQueue 케이스를 회귀 테스트로 박는다.</strong> 전편 사건은 통합 테스트가 없어서 운영에 도달한 뒤 발견됐다. ActionQueue처럼 ORM 내부 동작이 얽힌 시나리오는 <em>문서로 남기는 것</em> 보다 <em>테스트로 굳히는 것</em> 이 훨씬 강한 안전망이다.</li>
</ul>

<p>다만 그건 “체크리스트 설정”이 도메인 차원에서 <em>전체 교체 자원</em>인지 <em>부분 변경 자원</em>인지를 먼저 정한 다음의 이야기다. 이 글에선 <em>전체 교체 시그니처를 유지한 채 내부 구현만 diff로 바꾸는</em> 선까지로 범위를 좁혔다.</p>

<hr />

<h2 id="마무리">마무리</h2>

<p>운영 장애 한 건이 결국 <em>멱등성</em>과 <em>변경분 처리</em>라는 두 가지 일반 원칙으로 연결됐다. 구현은 못 했지만, 다음에 같은 모양의 테이블을 만나면 처음부터 이 형태로 짤 것이다. 그게 이 글에서 남기고 싶은 부분이다.</p>

<blockquote>
  <p>백엔드·ORM 시리즈는 여기까지. 이후 글은 인프라 담당자로 옮긴 뒤의 기록이다 — <a href="/sentry-csrf-reverse-proxy/">Sentry CSRF 6단계 편</a>부터.</p>
</blockquote>

<hr />

<h3 id="참고">참고</h3>

<ul>
  <li>단기 처방 PR: <a href="https://github.com/SWYP-Backend/BangCheck/pull/179">#179 fix: saveSettings 409 — deleteByUserId 직후 flush 호출</a> — 장기 개선 방향도 여기 주석으로 남겼다</li>
  <li>사건 회고: <a href="https://github.com/std-yong/bangcheck-portfolio/blob/main/docs/incidents/03-actionqueue-flush-409.md">docs/incidents/03-actionqueue-flush-409.md</a> — 정리 레포</li>
  <li>프로젝트: BangCheck (Spring Boot 3.x · JPA · MySQL 8) · 원본 레포 <a href="https://github.com/SWYP-Backend/BangCheck">SWYP-Backend/BangCheck</a> (팀 5인, 본인은 체크리스트 도메인 담당)</li>
</ul>]]></content><author><name>이진용</name><email>kouig14@gmail.com</email></author><category term="JPA" /><category term="Hibernate" /><category term="설계" /><category term="테스트" /><summary type="html"><![CDATA[단기 처방 flush()로 막아두고 PR에 약속해뒀던 “다음 스프린트의 리팩토링”을 어떤 모양으로 닫을지 설계로 정리한 글]]></summary></entry></feed>