번역 문서 사용 통계와 출처 추가 가이드 (Step 2) - #57
Merged
9bow merged 19 commits intoSep 15, 2026
Merged
Conversation
8 tasks done
reST prose and sphinx-gallery text blocks are extracted in rst_source.py; usage_core keeps the shared matching rules and only takes a block extractor. Source roots may now be a list so one source can cover sibling document directories, each paired with one original root.
tutorials-kr@84b7db6e paired with pytorch/tutorials@c4d9d93 over the five *_source directories: 269 documents scanned, 250 included, 14 without an English counterpart and 5 code-only .py files excluded. Hugging Face snapshots and dictionary data are unchanged.
Record what paired-sphinx includes, why adding a format keeps the counting rule version, and the pinned commits, scope and exclusion counts of the PyTorch tutorials source.
Documents that live at a repository root are configured with '.', so hub-kr pairs with pytorch/hub by file name. The exclude list mirrors what pytorch.kr's _config.yml drops from the hub collection: 58 scanned, 46 included, 2 cards whose English originals upstream removed.
pytorch.kr posts name their original in frontmatter and quote it paragraph by paragraph, so org_link plus the translation category is the inclusion evidence. Posts pair by URL slug because the Korean and English date prefixes differ, and the reason records whether the English Markdown is still in the repository (paired-translation) or only on the web (linked-translation).
pytorch.kr@dbc281dc with pytorch/pytorch.github.io@9104164e, the last commit before the upstream blog left the repository: 48 posts scanned, 45 included (14 with English Markdown, 31 web-only originals), 3 Korean original posts excluded.
List the PyTorch sources as tutorials, hub and blog with their pinned commits, scope and exclusion counts; describe the repository-root scope, the pytorch-blog adapter and its linked-translation exception, and the two hub repositories that carry no license file.
The term detail page embeds the Google Trends comparison for the same candidate spellings the counter searched, most used first, up to the five that one chart allows. The embed produces no statistics of its own and the page cannot read the iframe, so the failure notice and an external link stay next to it.
With several roots the empty-inventory guard only fired when every root was wrong, so one typo silently dropped part of the corpus. Each translation and original root is now checked on its own, and the blog slug index is built only for the adapter that uses it.
Not all 14 english-missing tutorials are deleted originals: two have an English counterpart under the other Sphinx extension, and one of those keeps 3,247 Korean characters in a file the site no longer builds. Also record how much of the included corpus is still untranslated and make the Hugging Face counts read in the same order as the table.
The intro no longer points at a usage table that is hidden on terms without matches, the comparison link disappears instead of opening an empty query, the checkbox group carries a name and a live selection count, and the iframe sends no referrer to Google.
Member
|
참고: 이 PR은 스택 base인 그래서 revert 대신, 이 PR의 커밋 19개를 현재 main 위로 리베이스해 #59로 재적용했습니다.
이 PR의 작업 자체는 main에 모두 반영되었습니다. 기록용으로만 남겨 둡니다. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
번역 문서 사용 통계와 출처 추가 가이드
Step 1에 의존하는 후속 PR
이 PR은 Step 1 PR #56에 의존하는 stacked PR입니다. Base는 원본 저장소의
codex/hf-glossary-pr이며, #56의 HEAD와 동일한 커밋fe26403을 가리킵니다. 현재 diff에는 Step 2 전용 8개 커밋·24개 파일만 포함됩니다.Step 2만 비교하기에서 통계 기능과 가이드 변경만 볼 수 있습니다. 아래 커밋별 리뷰 순서도 Step 2 전용 커밋입니다.
#56을 먼저 병합하고, 이 PR의 base를
main으로 변경한 뒤 최종 병합해 주세요. 현재 base는 리뷰용 Step 1 브랜치이므로 지금 병합하면 main이 아닌 해당 브랜치에 반영됩니다. #56에 후속 커밋이 생기면 원본의 기준 브랜치도 명시적으로 동기화해야 합니다(포크와 자동 동기화되지 않습니다). Squash 병합처럼 커밋 이력이 달라지는 경우 Step 2 전용 커밋만 최신 main 위로 옮겨 중복 변경을 제거해야 합니다. 자동 병합은 설정하지 않았습니다.변경 사항
usage/sources.json에서 출처·커뮤니티·정확한 커밋·문서 경로·어댑터·제외 조건을 관리합니다. UI도 이 설정에서 생성한 출처 목록을 사용합니다.scripts/usage-statistics/에 모았습니다. 사전에 없는 미사용 검색 후보 3개는 제거했으며 기존 통계는 바뀌지 않습니다.--source로 선택한 출처만 갱신하며 다른 출처의 체크아웃 없이 기존 결과를 보존할 수 있습니다.null/—, 실제 미출현은 0, 한글 없는 후보는 집계 제외로 구분합니다.docs/usage-statistics/README.md에 설계·데이터 계약·집계 규칙을,adding-source.md에 출처 추가 절차와 검증 체크리스트를 작성했습니다.CONTRIBUTING.md에서 연결하며AGENTS.md는 변경하지 않습니다.범위와 해석
현재 데이터는 Transformers·smolagents·HF Blog 고정 커밋에서 나온 것입니다. 실제 PyTorch 문서는 아직 등록하거나 집계하지 않았습니다. 같은 상대 경로의 한국어·영문 Markdown을 연결하는 공통 어댑터를 제공하며, PyTorch의 실제 문서가 RST/MDX 등이라면 형식에 맞는 어댑터·테스트를 먼저 추가해야 합니다.
표기 횟수는 한국어 문자열의 출현 빈도이며 특정 영문 용어의 번역 대응·권장 번역·사용자 선호도를 뜻하지 않습니다. 영문 대응의 존재는 포함 범위를 정하기 위한 조건입니다.
현재 결과: 용어 263개 중 194개 출현 확인, 문서 254개 스캔 중 205개 포함. 출처 공통화 이전의 Step 2와 전체 용어별 횟수·문서 수·근거 문장이 동일함을 비교했습니다. 원문은 이전에 수집해 둔 커밋이므로 최신 원격 문서 전체의 통계라고 주장하지 않습니다.
커밋별 리뷰 순서
45c3016— 출처 설정·공통 집계·증분 갱신·검증·테스트99f6702— 출처별 HF 스냅샷과 공개 결과 (대량 생성 파일)ec5805c— 출처 목록에 따라 확장되는 상세 페이지 UIad3961e— 설계 및 출처 기여 가이드fa246b1— 미사용 후보 agentic·guardrail·open-vocabulary 제거aaca72b— 스크립트를 기능 폴더로 이동하고 실행·테스트·문서 경로 갱신f2821a8— 새 출처 온보딩·범위 검증·어댑터 구현 안내와 커뮤니티 표시 정책 보강 (문서만 변경)병합용 임시 도구·내부 검토 기록은 포함하지 않습니다. 통계 재현에 필요한 집계 도구와 상태 파일은 이번 기능의 유지보수 자료이므로 포함합니다.
검증
filesChanged: 0AGENTS.md가 Step 1과 동일함을 확인자동 브라우저 조작·스크린샷 기반 시각 검증은 수행하지 않았습니다. 새 서버·DB·스케줄러·배포 설정·Node 의존성은 추가하지 않습니다. 집계용 Python 파서만
scripts/usage-statistics/requirements.txt로 고정합니다.