Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
260 changes: 56 additions & 204 deletions skills/how-to-download-ref/SKILL.md

Large diffs are not rendered by default.

84 changes: 84 additions & 0 deletions skills/how-to-download-ref/references/acquisition.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,84 @@
# Optional acquisition paths

Use the installed `DOWNLOAD_REF_DIR` and resolved `KB` from SKILL.md.
Read the relevant section when handling APS/JATS details, arXiv source output,
missed DOI PDFs, or supplemental material. Dependency setup is in
[dependencies.md](dependencies.md).

**APS DOIs are handled automatically.** For any `10.1103/*` DOI the helper first
calls the [APS Harvest API](https://harvest.aps.org/docs/harvest-api), which serves
the *publisher's own* JATS XML — real sections, MathML3 equations, a structured
reference list — with **no API key and no institutional IP**. This is ground truth
and strictly beats parsing the PDF. Coverage is per *article*, not per journal: you
get `ok` for gold-OA titles (PRX, PRX Quantum, PRResearch, PRAB, PRPER), SCOAP3
titles (PRC, PRD), and any individually CC-licensed article in PRL/PRA/PRB;
`closed` (HTTP 401) falls through to the arXiv and PDF tiers below. Pass `--no-aps`
to skip. The same request both tests access and delivers the text, so there is no
separate open-access lookup to do.

Metadata uses cached JSON, then Semantic Scholar batches of at most 500,
then Crossref for missing DOIs, including deposited metadata normalized to usable
BibTeX. PDF acquisition tries S2's OA URL, Unpaywall repository copies, then the
arXiv preprint. Pass `--email <contact-address>` or set `SCIBRAIN_CONTACT_EMAIL`
to enable Unpaywall and Crossref's polite pool; without an email, Unpaywall is
skipped. API errors and HTML landing pages fall through to the next source.
PDFs must have both a `%PDF` header and `%%EOF` trailer. A DOI miss continues to
the DOI fallback below. Service contracts: [Crossref](https://www.crossref.org/documentation/retrieve-metadata/rest-api/),
[Unpaywall](https://unpaywall.org/products/api).


`--download-arxiv-source` additionally fetches each arXiv paper's e-print
LaTeX source, extracts it to `.raw/arxiv/<id>-src/`, flattens
`\input`/`\include` into `.raw/arxiv/<id>.tex`, and copies the source tree's
figure files into `.figures/arxiv__<id>/`. `src-miss` lines (PDF-only
submissions, withdrawn papers, fetch failures) are fine — those refs fall
back to PDF rendering in the rendering step. DOI entries whose Semantic Scholar record
names an arXiv preprint (`externalIds.ArXiv`) get the same treatment, into
`.raw/doi/<safe>.tex` and `.figures/doi__<safe>/`.

**Tip:** Set `SEMANTIC_SCHOLAR_API_KEY` in your environment to raise the Semantic Scholar rate limit from ~1 req/s to 100 req/s. Get a free key at https://www.semanticscholar.org/product/api#api-key-form.

## Sci-Hub fallback for paywalled PDFs (script)

If the fetch reports `miss` for any DOI (no open-access PDF and no arXiv preprint),
run the browser-based Sci-Hub helper. Pass the missed DOIs:

```sh
python3 "$DOWNLOAD_REF_DIR/helpers/scihub_download.py" --kb "$KB" \
--doi 10.1111/j.1467-9280.2006.01693.x \
--doi 10.3102/0034654316689306
```

It tries each mirror in `helpers/scihub_domains.toml` (in order) until one
serves the PDF, solving the mirrors' DDoS-Guard JavaScript challenge with a
headless browser, and saves to `$KB/.raw/doi/<safe>.pdf` (`<safe>` = DOI with
`/` → `-`) — the same place the fetch writes, so render.py picks it up. It
prints one `OK` / `MISS` / `SKIP` line per DOI.

- **Requires Playwright** (see dependencies.md). curl/urllib cannot pass DDoS-Guard.
- **Mirrors rotate.** If every DOI returns `MISS`, the domain list is likely
stale: web-search "working sci-hub mirror domains <year>" and edit
`helpers/scihub_domains.toml` (see its header), then re-run.
- If a stricter challenge blocks the headless browser, retry with `--headed`.

Skip this fallback when the requested full text is already available.

## APS extras (optional)

`aps_harvest.py` also runs standalone — useful for backfilling a KB built before
this path existed, or for pulling figures and supplemental material:

```sh
# probe one DOI without writing anything -> prints open | closed | notfound
python3 "$DOWNLOAD_REF_DIR/helpers/aps_harvest.py" --check 10.1103/PhysRevB.108.045101

# backfill JATS for every APS DOI already in the KB
python3 "$DOWNLOAD_REF_DIR/helpers/aps_harvest.py" --kb "$KB" --all

# ...and pull the BagIt package too: published PDF, figures, supplemental material
python3 "$DOWNLOAD_REF_DIR/helpers/aps_harvest.py" --kb "$KB" --all --bagit
```

`--bagit` is the only way to get **supplemental material**, which the arXiv
preprint route cannot provide. It is much heavier (tens of MB per article), so
use it per-DOI rather than across a whole KB.
56 changes: 56 additions & 0 deletions skills/how-to-download-ref/references/dependencies.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,56 @@
# Dependencies and rendering backends

Every helper runs under plain `python3`. Two of them want third-party packages:
`render.py` needs **pymupdf4llm** (highest-fidelity output, preserves figures) and
`scihub_download.py` needs **playwright**. Without them the renderer degrades to
`markitdown` → `pdftotext`, which is text-only — *figures missing, equations
mangled*. Check the backend needed for the current render before running it:

```sh
python3 -c "import pymupdf4llm; print('ok', pymupdf4llm.__version__)"
```

If that errors, install it for the **same** `python3` the helpers will use:

```sh
python3 -m pip install --user pymupdf4llm
# macOS / Homebrew, or any PEP 668 "externally managed" Python:
python3 -m pip install --user --break-system-packages pymupdf4llm
```

Both scripts also carry [PEP 723](https://peps.python.org/pep-0723/) inline
dependency metadata, so if you happen to have [uv](https://docs.astral.sh/uv/),
`uv run "$DOWNLOAD_REF_DIR/helpers/render.py" ...` resolves those deps on its own and you can skip the
install step entirely. That is an option, not a requirement — the metadata is
inert comments to a plain interpreter.

**Tesseract is not needed for normal papers.** arXiv and APS PDFs are born-digital,
so `render.py` runs `pymupdf4llm` with `use_ocr=NEVER` and only retries with OCR
when a PDF turns out to have no text layer at all — a scanned old paper, usually
from the Sci-Hub tier. Install a language pack only if you hit that:
`tesseract-data-eng` (Arch), `tesseract-ocr-eng` (Debian/Ubuntu), or
`brew install tesseract-lang` (macOS).

On Arch in particular, *any* `tesseract-data-*` satisfies the `tessdata`
dependency, so it is easy to have `tesseract` installed with `eng` absent.

The Sci-Hub fallback (the DOI fallback) additionally needs a Chromium for Playwright to
clear the mirrors' DDoS-Guard challenge. Only required if you expect to hit
paywalled DOIs:

```sh
python3 -m pip install --user playwright && python3 -m playwright install chromium
```

APS DOIs (`10.1103/*`) render from publisher JATS XML, which needs **pandoc**:

```sh
pandoc --version | head -1 # any 2.x/3.x works
```

If missing: `paru -S pandoc-cli` (Arch) / `apt install pandoc` / `brew install pandoc`.
Without it, APS refs silently fall back to the arXiv/PDF tiers.

For arXiv LaTeX sources (optional, when LaTeX sources are requested), `latexpand`
(ships with TeX Live) gives the cleanest flattening; if absent, a built-in Python
inliner is used — no action needed either way.
22 changes: 22 additions & 0 deletions skills/how-to-download-ref/references/maintenance.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
# Restore an existing KB

Use the installed `DOWNLOAD_REF_DIR` and resolved `KB` from SKILL.md.

## Regenerating a cloned KB

```sh
python3 "$DOWNLOAD_REF_DIR/helpers/kb_sync.py" --kb "$KB"
```

Requires `references.bib` and at least one rendered paper. This restores `.raw/`
and `.figures/` using each tracked entry's declared identifier namespace. Bib-only
references recover caches with a warning. It creates or rewrites no Markdown,
INDEX.md, or bibliography files. Complete caches need no network on repeat runs;
unavailable assets remain WARNs and can be retried. Invalid input or failed
restoration produces FAIL and a nonzero exit.

PDF figure restoration needs the same `pymupdf4llm` version used to render the
entry so filenames match tracked image links. Missing dependencies and mismatched
filenames are reported. LaTeX figures are restored from the cached source tree or
a new source download. Publisher JATS is restored for `full_text: jats` entries.
`--email` / `SCIBRAIN_CONTACT_EMAIL` enable Unpaywall here too.
20 changes: 20 additions & 0 deletions skills/how-to-download-ref/references/troubleshooting.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
# Acquisition and rendering troubleshooting

Read the matching row when a helper fails or gives unexpected output. Resource
paths use the installed `DOWNLOAD_REF_DIR` from SKILL.md.

## Common mistakes

| Mistake | Fix |
| --- | --- |
| Passing a relative `--kb` | Always absolute. Helpers don't `cd`; figures depend on absolute paths. |
| Forgetting `--download-arxiv-pdfs` in Step 4 | Without it, refs with no LaTeX source render `full_text: no` — the PDF is the only body for DOIs and PDF-only arXiv submissions. |
| Using `arXiv:XXXX` with prefix or `vN` suffix | Strip both — manifest takes bare ids: `1806.08734`. |
| Editing generated body text and losing it on re-render | Keep prose in NOTES.md. Human frontmatter `note`, `tags`, and `rating` survives re-rendering. |
| Cite-key collision with different content | `append` skips silently. Propose with `--bib` so the key is disambiguated up front (next content word of the title). |
| Drifting `--title` / `--source-note` between runs | `INDEX.md` regenerates wholesale; first-run values are canonical. Copy verbatim from existing `INDEX.md`. |
| Expecting `.figures/` images for `full_text: latex` refs to come from the PDF | They come from the source tarball; PDF image extraction runs only on the PDF path. |
| Rendered from PDF despite a `.tex` in `.raw/` | PDF is the default. To use LaTeX bodies, pass `--tex-source` in Step 5 (and `--download-arxiv-source` in Step 4). |
| APS paper rendered from PDF, math mangled | `pandoc` is missing, or the article is genuinely `closed`. Check with `aps_harvest.py --check <doi>`. |
| Reaching for MinerU/Marker on an APS DOI | Try Harvest first — a 401 is the only thing that justifies parsing a PDF at all. |
| APS DOI reported `notfound` | Harvest matches the DOI suffix case-sensitively; `aps_harvest.canonical_doi` restores APS's capitalisation before the request. Add the journal to `APS_JOURNAL_TOKENS` if a new title 404s. |
35 changes: 35 additions & 0 deletions tests/test_skill_relative_links.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
"""Every relative Markdown link inside a skill resolves within that skill's own directory.

Skills are installed one at a time (symlinked or copied), so a SKILL.md or any
reference it links to must not point outside the skill.
"""

import re
import shutil
from pathlib import Path

import pytest

ROOT = Path(__file__).resolve().parents[1]
SKILLS = sorted(p.parent.name for p in (ROOT / "skills").glob("*/SKILL.md"))


@pytest.mark.parametrize("skill", SKILLS)
def test_relative_links_resolve_inside_the_installed_skill(tmp_path, skill):
installed = (tmp_path / skill).resolve()
shutil.copytree(ROOT / "skills" / skill, installed)
pending = [installed / "SKILL.md"]
visited = set()
while pending:
source = pending.pop()
if source in visited:
continue
visited.add(source)
for link in re.findall(r"\]\(([^)\s]+)\)", source.read_text()):
if "://" in link or link.startswith("#") or link.startswith("mailto:"):
continue
target = (source.parent / link.split("#", 1)[0]).resolve()
assert target.is_relative_to(installed), (source, link)
assert target.exists(), (source, link)
if target.suffix == ".md":
pending.append(target)
Loading