Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,4 +21,8 @@ Neither route needs the checksum or the size: an entry missing either one has th
downloaded on GitHub and the fields filled in for it, unless the pull request comes from
a fork, whose branch cannot be pushed to.

A dataset that is one file inside a zip on a record nobody can re-publish is written by
hand instead, with an `archive` block naming the zip and the member inside it.
`download()` then fetches only that member, and never the whole archive.

Full instructions, including what to do by hand and what CI checks: <https://electronmicroscopy.github.io/emdatabase/contributing.html>.
6 changes: 5 additions & 1 deletion docs/source/_build_docs.py
Original file line number Diff line number Diff line change
Expand Up @@ -88,7 +88,11 @@
link.href = pin.url;
link.target = "_blank";
link.rel = "noopener";
link.textContent = `⤓ Download ${pin.file}`;
// An archive entry's only link is the zip the file lives inside, so the
// label names the archive rather than a file the link does not serve.
link.textContent = item.archive
? `⤓ Download ${pin.url.split("/").pop()} (archive holding ${item.archive})`
: `⤓ Download ${pin.file}`;
const wrap = el("div", "emdb-dl-link");
wrap.appendChild(link);
detailsEl.appendChild(wrap);
Expand Down
41 changes: 39 additions & 2 deletions docs/source/contributing.rst
Original file line number Diff line number Diff line change
Expand Up @@ -82,6 +82,38 @@ on its format and whether it is required, and fill it in. Then check it:
That prints one line per problem and exits non-zero, or prints ``valid``.
:func:`emdatabase.metadata.validate_file` is the same check from Python.

A file inside an archive
------------------------

Some data worth shipping is one file inside a multi-gigabyte zip on a record
nobody can re-publish, where downloading all of it to get one file is not
reasonable. Such an entry adds an ``archive`` block naming the zip and the
member inside it:

.. code-block:: yaml

archive:
url: https://zenodo.org/records/0000000/files/Figures.zip
member: Figure_01/Panel_a/scan_x128_y128.raw

``download()`` then fetches only that member, over HTTP range requests: a few
requests read the zip's directory, and the member costs its own compressed bytes
rather than the whole archive. What comes back is a path, as for any other entry.

The entry's own ``file``, ``checksum`` and ``size_bytes`` go on describing the
member, which is the file you end up with; ``checksum`` and ``size_bytes``
inside the block describe the archive instead, and are optional. Give ``member``
as the complete path inside the zip - one file name can appear in several of its
directories, so a basename alone is ambiguous.

These entries are written by hand. The two fields CI otherwise fills in would be
taken from the archive rather than from the member, so a pull request leaving
them blank is refused rather than guessed at. ``download_url``, and the download
link on the docs site, point at the archive: the member has no link of its own.

A host that ignores ``Range`` and answers with the whole archive fails loudly,
naming the host, rather than quietly pulling gigabytes.

Contributing model weights
--------------------------

Expand Down Expand Up @@ -151,13 +183,18 @@ a dataset with one, fail as well.
files. It downloads the file behind each changed entry that is missing its
``checksum`` or ``size_bytes`` - a weights family's ``latest`` and each dated
version on their own links - fills the fields in and pushes the result back to
the branch. A fork's branch cannot be pushed to, so a pull request from one
the branch. An entry naming an ``archive`` is refused instead: its ``checksum``
and ``size_bytes`` describe the member inside the zip, and the only thing there
is to download is the whole archive, so those two are filled in by hand. A
fork's branch cannot be pushed to, so a pull request from one
fails instead and prints the values to paste in. An entry coming in through the
issue form is filled in the same way before its pull request is opened, so it
arrives complete.

``check_sources.yml`` runs weekly and asks each source server whether the file
is still there and still the size the entry claims.
is still there and still the size the entry claims. For an entry fetched out of
an archive it reads the archive's directory too, over a few range requests, and
checks the member is still there at the size the entry declares.

``check_weights.yml`` runs weekly as well, and what it does depends on where
the family is hosted.
Expand Down
156 changes: 156 additions & 0 deletions emdatabase/_archive.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,156 @@
"""Fetching one file out of a zip on someone else's server.

Some data worth shipping is a single member of a multi-gigabyte archive on a
record that cannot be re-published, where downloading all of it to get one file
is not reasonable. A zip's directory sits at its end and names the byte range of
every member, so with HTTP range requests the whole archive never has to move:
three requests find the directory, and the member costs its own stored bytes.

:class:`_HTTPRangeFile` is the seekable file ``zipfile`` reads that through, and
:class:`ArchiveMemberDownloader` is the pooch downloader
:meth:`~emdatabase.downloadable_dataset.DownloadableDataset._retrieve` hands to
:func:`pooch.retrieve` in place of :class:`pooch.HTTPDownloader` when an entry
names an ``archive``.
"""

from __future__ import annotations

import io
import urllib.parse
import urllib.request
import zipfile
from typing import Any

from emdatabase.downloadable_dataset import USER_AGENT, Progress

# The zip directory is read in many small seeks, so an unbuffered reader would
# cost hundreds of requests before a single byte of the member moved.
_BUFFER_SIZE = 1 << 20


class ArchiveError(Exception):
"""The host will not serve the archive in a way one member can be read out of.

Deliberately not an :class:`OSError`: ``zipfile`` turns any ``OSError``
raised while it is reading the directory into
``BadZipFile("File is not a zip file")``, which would replace the real
explanation with a wrong one.
"""


def _host(url: str) -> str:
return urllib.parse.urlsplit(url).netloc


class _HTTPRangeFile(io.RawIOBase):
"""A seekable, read-only file over HTTP range requests.

``zipfile`` only seeks and reads, so ranges stand in for a local copy.
"""

def __init__(self, url: str, timeout: float = 120) -> None:
self.url = url
self.timeout = timeout
self.pos = 0
request = urllib.request.Request(url, method="HEAD", headers={"User-Agent": USER_AGENT})
with urllib.request.urlopen(request, timeout=timeout) as response:
declared = response.headers["Content-Length"]
if declared is None:
raise ArchiveError(
f"{_host(url)} did not say how big {url} is, so the end of the archive - "
"where a zip keeps its directory - cannot be found."
)
self.size = int(declared)

def readable(self) -> bool:
return True

def seekable(self) -> bool:
return True

def tell(self) -> int:
return self.pos

def seek(self, offset: int, whence: int = io.SEEK_SET) -> int:
if whence == io.SEEK_SET:
self.pos = offset
elif whence == io.SEEK_CUR:
self.pos += offset
else:
self.pos = self.size + offset
return self.pos

def readinto(self, buffer) -> int: # pyright: ignore[reportMissingParameterType]
if self.pos >= self.size:
return 0
end = min(self.pos + len(buffer), self.size) - 1
request = urllib.request.Request(
self.url,
headers={"User-Agent": USER_AGENT, "Range": f"bytes={self.pos}-{end}"},
)
with urllib.request.urlopen(request, timeout=self.timeout) as response:
# A host that does not do ranges answers 200 with the whole body.
# Reading it would quietly pull the entire archive, which is the one
# thing this exists to avoid, so it is an error rather than a
# fallback.
if response.status != 206:
raise ArchiveError(
f"{_host(self.url)} ignored a Range request and answered "
f"{response.status}, so fetching one member would mean downloading "
f"all {self.size} bytes of {self.url}."
)
data = response.read()
buffer[: len(data)] = data
self.pos += len(data)
return len(data)


class ArchiveMemberDownloader:
"""A pooch downloader that pulls one member out of a remote zip.

pooch calls a downloader as ``(url, output_file, pooch)`` from inside
``pooch.core.stream_download``, which streams to a temporary file, checks it
against the entry's ``checksum`` and only then renames it into place,
deleting the temporary file on any failure. So this only has to produce the
member's bytes and drive ``progressbar`` the way :class:`pooch.HTTPDownloader`
does: ``total`` once before streaming, ``update(n)`` per chunk, then
``reset()``, ``update(total)``, ``close()``. The widgets' cancel works by
raising from ``update``, which aborts the stream and takes the temporary
file with it.
"""

def __init__(
self,
member: str,
progressbar: Progress | bool = False,
chunk_size: int = 4096,
) -> None:
self.member = member
# `True` means "build your own bar", which pooch's HTTPDownloader does
# and this does not: _retrieve has already swapped it for a Progress.
self.progressbar = None if isinstance(progressbar, bool) else progressbar
self.chunk_size = chunk_size

def __call__(self, url: str, output_file: str, _pooch: Any = None) -> None:
bar = self.progressbar
with io.BufferedReader(_HTTPRangeFile(url), buffer_size=_BUFFER_SIZE) as stream: # pyright: ignore[reportArgumentType]
with zipfile.ZipFile(stream) as archive:
try:
info = archive.getinfo(self.member)
except KeyError:
raise KeyError(
f"{url} holds no member {self.member!r}. Name the complete path "
"inside the archive: one file name can appear in several of its "
"directories."
) from None
if bar:
bar.total = info.file_size
with archive.open(info) as member, open(output_file, "wb") as out:
while chunk := member.read(self.chunk_size):
out.write(chunk)
if bar:
bar.update(len(chunk))
if bar:
bar.reset()
bar.update(info.file_size)
bar.close()
3 changes: 3 additions & 0 deletions emdatabase/catalogue.py
Original file line number Diff line number Diff line change
Expand Up @@ -113,6 +113,9 @@ def entry(name: str, ds: DownloadableDataset) -> dict:
"source": md.source,
"file": md.file,
"url": ds.download_url,
# For an entry fetched out of a zip, `url` is the archive rather than
# the file; this is the member inside it, and "" for everything else.
"archive": md.archive.member if md.archive else "",
"latest_checksum": ds.checksum or "",
"versions": _versions(ds),
"model_class": md.model.class_ if md.model else "",
Expand Down
17 changes: 16 additions & 1 deletion emdatabase/data/__init__.pyi
Original file line number Diff line number Diff line change
Expand Up @@ -326,6 +326,21 @@ class TutorialUNet(DownloadableDataset):
Model weights hosted at https://drive.google.com; see ``.versions`` for the dated snapshots.


"""
...

class TwistedBilayerWSe2Themis(DownloadableDataset):
"""
TwistedBilayerWSe2Themis

Electron ptychography of a twisted bilayer WSe2, acquired at 80 kV on an uncorrected Thermo Fisher Themis with an EMPAD. From the record for "Achieving sub-0.5-Angstrom resolution ptychography in an uncorrected electron microscope", which reaches 0.44 Angstrom without an aberration corrector (Science 384, adl2029). The file is one member of the record's 1.9 GB Fig_01.zip and is fetched out of it directly, without downloading the archive. It is a headerless raw: a 128 x 128 scan of 128 x 130 little-endian float32 frames, each frame the 128 x 128 detector followed by two rows of EMPAD metadata. Read it with numpy.fromfile(path, dtype="<f4").reshape(128, 128, 130, 128)[..., :128, :]. The same archive holds a Talos acquisition of the same scan under Fig_01/Panel_c-d_Talos/, so the member path matters.

DOI: 10.5281/zenodo.10431683

You can download this dataset here:
https://zenodo.org/records/10431683/files


"""
...

Expand All @@ -344,4 +359,4 @@ class ZrNbPrecipitate(DownloadableDataset):
"""
...

__all__ = ['AlNanocrystals', 'AmorphousFilm4nm4DSTEM', 'ApoferritinApollo15eps', 'BilayerWS2', 'CuZnEELSMapping', 'CuZnHAADF', 'FeAlStripes', 'HREBSDStrainPatterns', 'InSituElectrochemGrowth', 'LSMOLineScan', 'LSMOLineScanLowLoss', 'LSMOSTOLineScan', 'LSMOSTOLineScanLowLoss', 'LayeredCuNb4DSTEM', 'MgONanoCrystals', 'NiEBSDLarge', 'PdCuSiCrystallization', 'PdNiPGlass', 'PeakDetectionPolymers', 'SPEDAg', 'TutorialUNet', 'ZrNbPrecipitate']
__all__ = ['AlNanocrystals', 'AmorphousFilm4nm4DSTEM', 'ApoferritinApollo15eps', 'BilayerWS2', 'CuZnEELSMapping', 'CuZnHAADF', 'FeAlStripes', 'HREBSDStrainPatterns', 'InSituElectrochemGrowth', 'LSMOLineScan', 'LSMOLineScanLowLoss', 'LSMOSTOLineScan', 'LSMOSTOLineScanLowLoss', 'LayeredCuNb4DSTEM', 'MgONanoCrystals', 'NiEBSDLarge', 'PdCuSiCrystallization', 'PdNiPGlass', 'PeakDetectionPolymers', 'SPEDAg', 'TutorialUNet', 'TwistedBilayerWSe2Themis', 'ZrNbPrecipitate']
32 changes: 26 additions & 6 deletions emdatabase/downloadable_dataset.py
Original file line number Diff line number Diff line change
Expand Up @@ -94,13 +94,18 @@ class _Resolved:
``pinned`` is False only for a weights family's ``latest``, where the
checksum describes what the link served when the index was written rather
than what it has to serve now.

``member`` is set when ``url`` is an archive the file has to be taken out
of, rather than the file itself; ``checksum`` and ``size_bytes`` describe
the member either way, because that is what the caller ends up with.
"""

url: str
checksum: str | None
size_bytes: int | None
file: str
pinned: bool
member: str | None = None


class Progress(Protocol):
Expand Down Expand Up @@ -354,17 +359,21 @@ def _resolve(self, version: str | None = None) -> _Resolved:
A dataset is a single pinned file and takes no version. A weights entry
is a family: with no version it is the ``latest`` link, which is not
pinned, and with one it is that dated snapshot, which is.

A dataset naming an ``archive`` resolves to that archive's link and the
member inside it, since the archive is the only thing there is a link to.
"""
md = self.metadata
if md.kind != "weights":
if version is not None:
raise ValueError(f"{type(self).__name__} is a dataset and has no versions")
return _Resolved(
url=md.url or f"{md.source}/{md.file}",
url=md.archive.url if md.archive else (md.url or f"{md.source}/{md.file}"),
checksum=md.checksum,
size_bytes=md.size_bytes,
file=md.file,
pinned=True,
member=md.archive.member if md.archive else None,
)
if version is None:
if md.latest is None:
Expand Down Expand Up @@ -399,6 +408,9 @@ def download_url(self) -> str:
link. Where it is not - a Google Drive link, or anything else with a
query string - the entry gives the whole link as ``url`` instead, and a
weights entry gives one link per version.

For an entry whose file lives inside an archive this is the archive:
the member has no link of its own.
"""
return self._resolve(None).url

Expand Down Expand Up @@ -595,11 +607,19 @@ def _retrieve(
if shared is not None:
return shared
destination = self._resolve_destination(destination)
downloader = pooch.HTTPDownloader(
progressbar=progressbar, # pyright: ignore[reportArgumentType]
chunk_size=chunk_size,
headers={"User-Agent": USER_AGENT},
)
downloader: Any
if resolved.member is None:
downloader = pooch.HTTPDownloader(
progressbar=progressbar, # pyright: ignore[reportArgumentType]
chunk_size=chunk_size,
headers={"User-Agent": USER_AGENT},
)
else:
# Imported here, not at the top, because _archive imports this
# module for USER_AGENT and the progress protocol.
from emdatabase._archive import ArchiveMemberDownloader

downloader = ArchiveMemberDownloader(resolved.member, progressbar, chunk_size)
try:
if refresh:
# pooch keeps a file whose hash it was not given anything to
Expand Down
13 changes: 13 additions & 0 deletions emdatabase/index/TEMPLATE.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,19 @@ MyDatasetName:
file: MyDatasetName.zspy
# The file's Content-Length in bytes; the weekly CI check compares it to the server.
size_bytes: 1000000
# Only for a file that lives inside a zip on a record you cannot re-publish,
# where downloading the whole archive to get one file is not reasonable.
# `download()` fetches just this member, with HTTP range requests, and hands
# back a path like any other entry. `url`, `checksum` and `size_bytes` here
# describe the archive and are what `download_url` points at; `file`,
# `checksum` and `size_bytes` above go on describing the member, the file you
# end up with. Give the complete path inside the zip: the same file name can
# appear in several of its directories. Not allowed on a `kind: weights` entry.
# archive:
# url: https://zenodo.org/records/0000000/files/Figures.zip
# member: Figure_01/Panel_a/scan_x128_y128.raw
# checksum: md5:0123456789abcdef0123456789abcdef
# size_bytes: 2000000
# Who made the detector. See vendors.yaml for the names already in use.
detector_manufacturer: Direct Electron
# The detector model.
Expand Down
Loading
Loading