Skip to content

번역 문서 사용 통계와 출처 추가 가이드 (Step 2) - #57

Merged
9bow merged 19 commits into
PyTorchKR:codex/hf-glossary-prfrom
Jwaminju:codex/usage-statistics-pr
Sep 15, 2026
Merged

9bow merged 19 commits into
PyTorchKR:codex/hf-glossary-prfrom
Jwaminju:codex/usage-statistics-pr

Conversation

@Jwaminju

@Jwaminju Jwaminju commented Sep 13, 2026 •

Copy link
Copy Markdown
Collaborator

번역 문서 사용 통계와 출처 추가 가이드

Step 1에 의존하는 후속 PR

이 PR은 Step 1 PR #56에 의존하는 stacked PR입니다. Base는 원본 저장소의 codex/hf-glossary-pr이며, #56의 HEAD와 동일한 커밋 fe26403을 가리킵니다. 현재 diff에는 Step 2 전용 8개 커밋·24개 파일만 포함됩니다.

Step 2만 비교하기에서 통계 기능과 가이드 변경만 볼 수 있습니다. 아래 커밋별 리뷰 순서도 Step 2 전용 커밋입니다.

#56을 먼저 병합하고, 이 PR의 base를 main으로 변경한 뒤 최종 병합해 주세요. 현재 base는 리뷰용 Step 1 브랜치이므로 지금 병합하면 main이 아닌 해당 브랜치에 반영됩니다. #56에 후속 커밋이 생기면 원본의 기준 브랜치도 명시적으로 동기화해야 합니다(포크와 자동 동기화되지 않습니다). Squash 병합처럼 커밋 이력이 달라지는 경우 Step 2 전용 커밋만 최신 main 위로 옮겨 중복 변경을 제거해야 합니다. 자동 병합은 설정하지 않았습니다.

변경 사항

  • 기존 용어 상세 페이지에 ‘번역 문서에서의 쓰임’을 추가합니다. Step 2 자체는 Step 1 대비 홈·카드·사전 데이터·대표 번역·의미 순서를 바꾸지 않습니다.
  • usage/sources.json에서 출처·커뮤니티·정확한 커밋·문서 경로·어댑터·제외 조건을 관리합니다. UI도 이 설정에서 생성한 출처 목록을 사용합니다.
  • 집계·검증 스크립트와 Python 의존성 파일은 scripts/usage-statistics/에 모았습니다. 사전에 없는 미사용 검색 후보 3개는 제거했으며 기존 통계는 바뀌지 않습니다.
  • 문서 blob SHA와 포함 조건을 기준으로 증분 갱신하고 출처별 상태를 Git에 저장합니다. --source로 선택한 출처만 갱신하며 다른 출처의 체크아웃 없이 기존 결과를 보존할 수 있습니다.
  • 집계 규칙·후보가 서로 다른 출처 스냅샷을 섞으려 하면 저장 전에 실패합니다. 미수집은 null/—, 실제 미출현은 0, 한글 없는 후보는 집계 제외로 구분합니다.
  • docs/usage-statistics/README.md에 설계·데이터 계약·집계 규칙을, adding-source.md에 출처 추가 절차와 검증 체크리스트를 작성했습니다. CONTRIBUTING.md에서 연결하며 AGENTS.md는 변경하지 않습니다.
  • 출처 전달 양식, 체크아웃 준비 명령, 새 문서 형식의 실제 구현 지점·데이터 계약을 보강했습니다. 전체 제외 상태도 재집계 검증을 통과할 수 있으므로 포함 문서 수·제외 사유의 게시 전 검증을 필수로 명시했습니다. UI는 커뮤니티별 별도 표가 아니라 한 표에 출처별 컬럼을 추가하는 방식입니다.
  • 상단 참여 커뮤니티는 해당 용어의 출현 근거가 있는 출처만 표시합니다. 근거가 없으면 이름을 숨기며, 출처 컬럼·숫자·집계 범위는 유지합니다.

범위와 해석

현재 데이터는 Transformers·smolagents·HF Blog 고정 커밋에서 나온 것입니다. 실제 PyTorch 문서는 아직 등록하거나 집계하지 않았습니다. 같은 상대 경로의 한국어·영문 Markdown을 연결하는 공통 어댑터를 제공하며, PyTorch의 실제 문서가 RST/MDX 등이라면 형식에 맞는 어댑터·테스트를 먼저 추가해야 합니다.

표기 횟수는 한국어 문자열의 출현 빈도이며 특정 영문 용어의 번역 대응·권장 번역·사용자 선호도를 뜻하지 않습니다. 영문 대응의 존재는 포함 범위를 정하기 위한 조건입니다.

현재 결과: 용어 263개 중 194개 출현 확인, 문서 254개 스캔 중 205개 포함. 출처 공통화 이전의 Step 2와 전체 용어별 횟수·문서 수·근거 문장이 동일함을 비교했습니다. 원문은 이전에 수집해 둔 커밋이므로 최신 원격 문서 전체의 통계라고 주장하지 않습니다.

커밋별 리뷰 순서

  1. 45c3016 — 출처 설정·공통 집계·증분 갱신·검증·테스트
  2. 99f6702 — 출처별 HF 스냅샷과 공개 결과 (대량 생성 파일)
  3. ec5805c — 출처 목록에 따라 확장되는 상세 페이지 UI
  4. ad3961e — 설계 및 출처 기여 가이드
  5. fa246b1 — 미사용 후보 agentic·guardrail·open-vocabulary 제거
  6. aaca72b — 스크립트를 기능 폴더로 이동하고 실행·테스트·문서 경로 갱신
  7. f2821a8 — 새 출처 온보딩·범위 검증·어댑터 구현 안내와 커뮤니티 표시 정책 보강 (문서만 변경)
  8. Show only communities with term evidence — 용어별 근거 기준으로 커뮤니티 이름 필터링, 회귀 테스트·문서 갱신

병합용 임시 도구·내부 검토 기록은 포함하지 않습니다. 통계 재현에 필요한 집계 도구와 상태 파일은 이번 기능의 유지보수 자료이므로 포함합니다.

검증

  • Python 테스트 21개, 출처 링크 테스트 3개, 커뮤니티 표시 테스트 1개(네 경우) 통과
  • 출처 단독 갱신, 새 출처의 미수집 표시, 다른 출처 스냅샷 보존
  • 추가·수정·삭제·이동·제외·재포함, 읽기 실패 시 기존 결과 보존
  • 전체 205개 문서를 다시 센 결과와 캐시 결과의 횟수·근거 일치
  • 동일 입력 재실행: 문서 205개 재사용, filesChanged: 0
  • 사전 데이터 및 AGENTS.md가 Step 1과 동일함을 확인
  • 데이터/통계 정합성·TypeScript·프로덕션 빌드 통과
  • 로컬 Step 1·Step 2 HTTP 연결 및 새 공개 통계 JSON 확인
npm run test:usage
python3 scripts/usage-statistics/update_usage_counts.py --check-full --sources-dir /path/to/document-checkouts
npm run build

자동 브라우저 조작·스크린샷 기반 시각 검증은 수행하지 않았습니다. 새 서버·DB·스케줄러·배포 설정·Node 의존성은 추가하지 않습니다. 집계용 Python 파서만 scripts/usage-statistics/requirements.txt로 고정합니다.

@Jwaminju
Jwaminju changed the base branch from main to codex/hf-glossary-pr September 13, 2026 12:32
Jwaminju and others added 13 commits September 13, 2026 21:38
reST prose and sphinx-gallery text blocks are extracted in rst_source.py;
usage_core keeps the shared matching rules and only takes a block extractor.
Source roots may now be a list so one source can cover sibling document
directories, each paired with one original root.
tutorials-kr@84b7db6e paired with pytorch/tutorials@c4d9d93 over the five
*_source directories: 269 documents scanned, 250 included, 14 without an
English counterpart and 5 code-only .py files excluded. Hugging Face
snapshots and dictionary data are unchanged.
Record what paired-sphinx includes, why adding a format keeps the counting
rule version, and the pinned commits, scope and exclusion counts of the
PyTorch tutorials source.
Documents that live at a repository root are configured with '.', so hub-kr
pairs with pytorch/hub by file name. The exclude list mirrors what
pytorch.kr's _config.yml drops from the hub collection: 58 scanned, 46
included, 2 cards whose English originals upstream removed.
pytorch.kr posts name their original in frontmatter and quote it paragraph
by paragraph, so org_link plus the translation category is the inclusion
evidence. Posts pair by URL slug because the Korean and English date
prefixes differ, and the reason records whether the English Markdown is
still in the repository (paired-translation) or only on the web
(linked-translation).
pytorch.kr@dbc281dc with pytorch/pytorch.github.io@9104164e, the last
commit before the upstream blog left the repository: 48 posts scanned, 45
included (14 with English Markdown, 31 web-only originals), 3 Korean
original posts excluded.
List the PyTorch sources as tutorials, hub and blog with their pinned
commits, scope and exclusion counts; describe the repository-root scope,
the pytorch-blog adapter and its linked-translation exception, and the two
hub repositories that carry no license file.
The term detail page embeds the Google Trends comparison for the same
candidate spellings the counter searched, most used first, up to the five
that one chart allows. The embed produces no statistics of its own and the
page cannot read the iframe, so the failure notice and an external link
stay next to it.
With several roots the empty-inventory guard only fired when every root was
wrong, so one typo silently dropped part of the corpus. Each translation and
original root is now checked on its own, and the blog slug index is built
only for the adapter that uses it.
Not all 14 english-missing tutorials are deleted originals: two have an
English counterpart under the other Sphinx extension, and one of those keeps
3,247 Korean characters in a file the site no longer builds. Also record how
much of the included corpus is still untranslated and make the Hugging Face
counts read in the same order as the table.
The intro no longer points at a usage table that is hidden on terms without
matches, the comparison link disappears instead of opening an empty query,
the checkbox group carries a name and a live selection count, and the iframe
sends no referrer to Google.
@9bow
9bow merged commit 1672881 into PyTorchKR:codex/hf-glossary-pr Sep 15, 2026
@9bow

9bow commented Sep 15, 2026

Copy link
Copy Markdown
Member

참고: 이 PR은 스택 base인 codex/hf-glossary-pr로 병합되어(1672881) main에는 반영되지 않았습니다. #56이 이미 squash로 main에 들어가 있어 이 브랜치를 그대로 main에 머지하면 src/pages/TermDetailPage.tsx에서 이력상 충돌이 발생하는 상태였습니다(내용 차이는 import 2줄 + 섹션 mount 2줄뿐).

그래서 revert 대신, 이 PR의 커밋 19개를 현재 main 위로 리베이스해 #59로 재적용했습니다.

이 PR의 작업 자체는 main에 모두 반영되었습니다. 기록용으로만 남겨 둡니다.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants