Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 9 additions & 7 deletions .github/ISSUE_TEMPLATE/bug_report.yml
Original file line number Diff line number Diff line change
Expand Up @@ -12,12 +12,14 @@ body:

Two things answer most reports faster than we can:

- **Zero products?** Check you used a *filtered* category URL, not the
bare `/shopping/kids/items.aspx` hub — the hub carries no product
JSON-LD and correctly returns nothing. See
- **Zero products?** Open the `{out}_page1_debug.html` dump the run
wrote beside itself: a challenge, a sign-in wall or a real page
with no results answers most of these. See
[TROUBLESHOOTING.md](../blob/main/TROUBLESHOOTING.md).
- **Wrong currency or localised titles?** Amazon decides both from
your exit IP, not from the URL. See "Geo-redirect" in the README.
- **Wrong currency?** Amazon converts the price to the currency of
your exit IP's country, not the marketplace's — from a European
exit amazon.co.uk and amazon.co.jp both quote EUR. See "The
currency follows the exit IP, not the domain" in the README.

- type: dropdown
id: engine
Expand Down Expand Up @@ -66,7 +68,7 @@ body:
placeholder: |
python3 playwright_scraper.py \
--url "https://www.amazon.com/s?k=wireless+headphones" \
--pages 1 --out girls
--pages 1 --out headphones
validations:
required: true

Expand All @@ -86,7 +88,7 @@ body:
attributes:
label: What you expected instead
placeholder: >-
96 products, as the README says a category page yields.
16–30 organic tiles per search page, as the README's measured results say.
validations:
required: true

Expand Down
15 changes: 10 additions & 5 deletions .github/ISSUE_TEMPLATE/site_changed.yml
Original file line number Diff line number Diff line change
Expand Up @@ -9,8 +9,12 @@ body:
own report because the fix is different from a code bug: something on
amazon.com moved.

Useful to know before filing: the parser tries **JSON-LD first**, then a
CSS + URL-pattern fallback. Which one broke narrows the fix a lot.
Useful to know before filing: Amazon publishes no JSON-LD, so the
parser reads the page kind's own data-attribute anchor
(`[data-asin]` on search, `[id^="p13n-asin-index-"]` on best
sellers, `[data-hook="reviewContainer"]` on reviews) and falls back
to the `/dp/{ASIN}` URL pattern, logging a warning when it does.
Which one broke narrows the fix a lot.

- type: dropdown
id: what_broke
Expand Down Expand Up @@ -41,9 +45,10 @@ body:
description: |
Whichever of these you can get. A capture beats a description.

- JSON-LD as the page actually serves it:
`python3 -c "import json,sys,re;h=open('dump.html').read();print([m[:400] for m in re.findall(r'<script type=\"application/ld\\+json\">(.*?)</script>',h,re.S)][:2])"`
- Or open devtools and paste one `Product` node from the `ItemList`.
- The exact bytes the parser was given: rerun with
`--dump-html dump.html` (it writes on success too) and attach it,
with session ids and tokens removed.
- Whether the run log warned that the `/dp/` URL-pattern fallback ran.
- For a field problem: the value you got, and the value on the page.
render: text
validations:
Expand Down
16 changes: 16 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -126,6 +126,19 @@ one, and when it does the release notes say so first.
engine logs the status and does not classify on it). It now reads
`http_code`, falling back to `status` only if that is an integer. After the fix, one live call (`--wait-text results` on the README's search URL) answered HTTP 200, upstream 200, 16 rows, status complete.

- **farfetch-scraper leftovers removed from the issue templates, TROUBLESHOOTING,
SECURITY and `diff_runs.py`.** The bug-report tips sent a reader to
Farfetch's `/shopping/kids/items.aspx` hub "which carries no product
JSON-LD", the example wrote `--out girls` and expected "96 products"; the
site-change template said the parser "tries JSON-LD first" and asked for a
JSON-LD dump; TROUBLESHOOTING's 0-rows table named Farfetch's
`-item-<digits>.aspx` links and hub; SECURITY named Akamai; `diff_runs.py`
described `source_changed` as "DOM-corrected versus raw JSON-LD" and its
examples used `girls_clothing`. Amazon publishes no JSON-LD, so all of
these now describe the data-attribute anchors, the `/dp/{ASIN}` fallback,
AWS WAF, the `offscreen`/`split`/`detail` price sources and this README's
measured 16–30 tiles per search page.

- **The Scraper API's `x-debug` response header is redacted before it is
logged.** `SECURITY.md` names that header as one of three places
credentials reach a log unmasked, and the client logged it whole: the API
Expand Down Expand Up @@ -186,6 +199,9 @@ one, and when it does the release notes say so first.
with "Claude Code native binary not found". The channel is pinned rather
than a version, so a fixed upstream release needs no edit here.

- `SECURITY.md` said this project has no releases or version tags; it has
both. "Supported versions" now names the latest release and `main`.

## [0.1.3] — 2026-09-11

### Fixed
Expand Down
2 changes: 1 addition & 1 deletion CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -117,7 +117,7 @@ exit country, and what you
got. Product counts differ by country and by URL, so a bare "worked for me" is
not reproducible.

Do not add anything that submits the registration form. This project
Do not add anything that submits a registration or login form. This project
deliberately never does, and a captcha token proved valid by creating a real
account is not a result worth having.

Expand Down
8 changes: 4 additions & 4 deletions SECURITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,7 +55,7 @@ Not because these do not matter, but because they belong somewhere else:

- **Bypassing Amazon's bot protection.** This scraper drives an ordinary
browser and passes challenges the way a browser does. Anything about how
Akamai or reCAPTCHA behave is not a vulnerability in this repository.
AWS WAF or Amazon's own captcha behave is not a vulnerability in this repository.
- **The scraper stopped working.** Amazon changing its markup is expected —
file it as a normal issue, there is a template for exactly that.
- **Anything about 2Captcha's services** — the solver API, the Scraping Browser
Expand All @@ -79,9 +79,9 @@ Not because these do not matter, but because they belong somewhere else:

## Supported versions

`main` only. This project has no releases or version tags; fixes land on `main`
and you update by pulling. If you are running an old clone, update before
reporting.
The latest release and `main`. Fixes land on `main` first and ship in the
next tagged release (see the Releases page). If you are running an old clone,
update before reporting.

## If you have leaked a key

Expand Down
3 changes: 1 addition & 2 deletions TROUBLESHOOTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,9 +11,8 @@ answer this in one look.
| What the dump shows | Cause |
|---|---|
| A challenge or "verify you are human" page | Bot management. Use a browser engine (not a plain HTTP fetch), a residential IP, or a remote browser via `--cdp-endpoint`. |
| A real page, prices visible, still 0 rows | The JSON-LD path found nothing and the CSS fallback did not match. Check that product links still match `-item-<digits>.aspx`. |
| A real page, prices visible, still 0 rows | The page kind's data-attribute anchor (`[data-asin]`, `[id^="p13n-asin-index-"]`, `[data-hook="reviewContainer"]`) found nothing and the `/dp/{ASIN}` URL-pattern fallback did not match either. See "How it parses" in the README. |
| A real page in a different language, prices like `125 €` | Fine — that parses. If rows are still 0, it is not the locale. |
| A near-empty page | The hub URL. `/shopping/kids/items.aspx` has zero products in its JSON-LD; use a filtered category URL. |

## A local Selenium session will not start

Expand Down
16 changes: 8 additions & 8 deletions diff_runs.py
Original file line number Diff line number Diff line change
Expand Up @@ -8,15 +8,15 @@
monitoring and assortment tracking, but that nothing in this repo actually
computed.

python3 diff_runs.py --old girls_clothing.2026-09-01.json \\
--new girls_clothing.2026-09-07.json
python3 diff_runs.py --old headphones.2026-09-01.json \\
--new headphones.2026-09-07.json

Typical use is a scheduled re-run of one of the four scraper engines, kept
under a dated filename, diffed against the previous one:

python3 playwright_scraper.py --url "$URL" --out "girls_$(date +%F)"
python3 diff_runs.py --old "girls_$(ls -t girls_*.json | sed -n 2p)" \\
--new "girls_$(date +%F).json" --out diff.json
python3 playwright_scraper.py --url "$URL" --out "headphones_$(date +%F)"
python3 diff_runs.py --old "headphones_$(ls -t headphones_*.json | sed -n 2p)" \\
--new "headphones_$(date +%F).json" --out diff.json

Four buckets, each keyed on sku:

Expand All @@ -26,9 +26,9 @@
changed — sku present in both, with a different price,
original_price, discount_pct, currency or in_stock
source_changed — sku present in both with a different price, but also a
different price_source: one run got the DOM-corrected
figure and the other the raw JSON-LD one, so the two are
not comparable on price. Reported separately because this
different price_source: the two runs read the price from
different DOM nodes (offscreen / split / detail), so the
two are not comparable on price. Reported separately because this
says something about our own two snapshots, not about the
site — and --fail-on-change deliberately ignores it.

Expand Down
Loading