Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 26 additions & 5 deletions .github/scripts/check_latest_weights.py
Original file line number Diff line number Diff line change
Expand Up @@ -27,8 +27,10 @@
round. Those bytes are immutable and permanent, so nothing is downloaded and
nothing is copied to GitHub; retraining publishes a new record under the same
concept record instead, which the file URL cannot show. The script asks the
Zenodo API for the concept's newest record, and a new record id becomes a dated
version pointing at that record.
Zenodo API for the concept's newest record, and a new record serving a file
whose md5 differs becomes a dated version pointing at that record. One serving
the same md5 is not a new state of the weights - a record is versioned whole, so
anything published beside them mints a new id - so only the link moves.
"""

from __future__ import annotations
Expand Down Expand Up @@ -262,8 +264,13 @@ def _check_zenodo(
"""Ask the Zenodo API whether this concept has a newer record.

A record's files never change, so there is nothing to download and nothing
to archive: a new version of the model is a new record id, and the dated
version written for it points straight at that record.
to archive: the API's md5 is what the link serves, and a dated version
written for it points straight at that record.

A newer record is not by itself a newer file. Zenodo versions the whole
record, so a change to anything in it publishes a new id while the file this
entry names may be byte for byte what it was; only a checksum that differs
is a new state of the weights worth a date of its own.
"""
try:
record = fetch_json(f"{link.api}/{link.record_id}")
Expand Down Expand Up @@ -299,10 +306,24 @@ def _check_zenodo(
report.lines.append(f"- unchanged; Zenodo record {new_id} is still the latest version")
return

url = link.file_url(new_id, str(served.get("key") or link.key))
if latest.checksum and checksum == latest.checksum:
# A newer record is not a newer file. Zenodo versions the whole record,
# so anything else in it changing - a zip published beside the weights -
# is a new id while this file stays byte for byte what it was. Only the
# link moves: a dated version would pin a second date to bytes that
# already have one.
report.lines.append(
f"- unchanged; Zenodo record {new_id} is newer than {link.record_id} but serves "
f"the same `{checksum}`, so the link moved and no version was filed"
)
entry["latest"]["url"] = url
report.changed = True
return

date = published.strftime("%y%m%d")
if _date_taken(report, metadata, date, checksum):
return
url = link.file_url(new_id, str(served.get("key") or link.key))
report.lines.append(
f"- new Zenodo record {new_id}: the index has `{latest.checksum}` from record "
f"{link.record_id}, and the new record serves `{checksum}` "
Expand Down
178 changes: 9 additions & 169 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,168 +1,25 @@
emdatabase
----------

This is a simple project for aggregating different Electron Microscopy files which are hosted over different sources. It uses pooch to download datasets and should be
used as a way to host simple example datasets for method validation.
This is a project for aggregating different Electron Microscopy files which are hosted over different sources. It is intended to simplify downloading example datasets and trained machine learning model weights for tutorials and method validation.

Downloads go to `~/.cache/emdatabase` by default; shared read-only locations can be added with `emdatabase.add_location`.

List of datasets https://electronmicroscopy.github.io/emdatabase/datasets.html
A list of all datasets and model weights can be found in our [docs](https://electronmicroscopy.github.io/emdatabase/datasets.html).

## Installation

```bash
pip install emdatabase
```
You can install `emdatabase` via pip or as a local install in the usual way.

## Usage

Every dataset is a class under `emdatabase.data`. Calling `download()` fetches the
file to the data directory, verifies its checksum, and returns a path handle. Files
that are already present are not downloaded again.

```python
import emdatabase.data as data
import hyperspy.api as hs

path = data.LayeredCuNb4DSTEM().download()
s = hs.load(path, lazy=True)
```

By default the download runs on a background thread so a notebook cell returns
immediately. The handle it returns *is* the file path — a `pathlib.Path` subclass
pointing at the file's final location — so you can hand it straight to a loader as
above; it only blocks at the moment the file is actually opened. Read `path.done`
to check progress without blocking, or call `path.result()` to wait explicitly.
`download(background=False)` blocks instead, and returns the same type.

Any path pointing at the same file waits, however it was built, so a derived path
(`handle.parent / handle.name`, `handle.with_suffix(...)`) behaves too. The
exceptions are `str(handle)` and `Path(handle)`: both hand back an ordinary value
with no download attached, so `hs.load(str(handle))` will *not* wait. Keeping
`str()` non-blocking is deliberate — `repr()` needs it — so pass the handle itself.

## Finding a dataset

`search()` is the browser widget's search box, callable from Python; `filter()`
matches named fields. Both return dataset objects, so a result can be downloaded
directly.

```python
import emdatabase

emdatabase.list_datasets() # everything
emdatabase.search("amorphous") # any field
emdatabase.search("jeol eels") # all terms, any field
emdatabase.filter(technique="4D-STEM", tags="Strain") # exact, case-insensitive
emdatabase.filter(microscope_vendor=["JEOL", "Hitachi"]) # a list means any of
emdatabase.filter(downloaded=True) # what is already here
```

`technique`, `tags`, `authors` and `version` are several values per dataset, so they
test membership: a dataset that is both in-situ and 4D-STEM matches either.

An unknown field raises rather than being ignored, so a typo cannot quietly return
the whole index.

## Configuration

Data lives in named **locations**. `personal` is the one writable location,
where downloads go; every other one is read-only and searched first, so a copy
already on a group drive is used instead of refetched.

```python
from emdatabase import config

config.add_location("/group/example_data") # read-only
config.add_location("/big/disk/emdatabase", name="personal") # where downloads go
config.locations()
```

```
[Location(name='example_data', path=PosixPath('/group/example_data'), kind='shared'),
Location(name='personal', path=PosixPath('/big/disk/emdatabase'), kind='personal')]
```

`locations()` is the search order: the shared locations in the order they were
added, then `personal` last. A location is named after the last component of its
path unless you pass `name=`, and that name is the provenance — it is what
`catalogue.entry()["location"]`, `emdatabase.filter(location="example_data")` and
the browser widget report for a copy found there. Nothing is written to a shared
location unless you name it as a download's destination, which is how one is
seeded.

Removing one takes either the name or the path; `"personal"` is not deleted but
reset, putting downloads back in the default cache directory:

```python
config.remove_location("example_data")
config.remove_location("/group/example_data") # the same thing, by path
config.remove_location("personal")
```

Both functions persist to `~/.config/emdatabase/config.yaml`, which is read on
every import. Pass `persist=False` to change this process only, or use
`config.set` as a context manager for a change that lasts for a block:

```python
config.add_location("/scratch/em", name="personal", persist=False) # this process
with config.set({"locations.personal": "/scratch/em"}): # this block
...
```

The path does not have to exist when you add it — a share may be mounted later —
but you get a warning saying so.

### Seeding a shared location

`destination=` takes a location's name, which is how the copy gets onto the share
in the first place — run it once, from an account with write access:

```python
from emdatabase import data

data.CuZnHAADF().download(destination="example_data")
```

The file is written with your umask, so `chmod` it group-readable afterwards if
your umask is not; emdatabase does not set permissions for you.

### The key underneath

Configuration is dask-style: shipped defaults, then every `*.yaml` in
`~/.config/emdatabase/` (or wherever `EMDATABASE_CONFIG` points), then
environment variables, then `config.set` — each layer overriding the one before.
There are two keys, and `add_location` is a wrapper over writing the first one
yourself:

```yaml
# ~/.config/emdatabase/config.yaml
locations:
example_data: /group/example_data
cluster: /cluster/em_data
personal: /big/disk/emdatabase
check_updates: true
```

`personal: null` means pooch's cache directory (`~/.cache/emdatabase` on Linux),
and `config.data_dir()` reports whichever it resolves to.

`check_updates` is whether downloading a model's `latest` weights asks the index
on the project's `main` branch — kept current by a weekly job — whether newer
weights have been published, and warns if they have; `download(refresh=True)`
fetches them. Set it to `false` to skip the request.

On HPC, where a config file is often the wrong place to put a machine-specific
path, set the same key from the environment instead — prefix `EMDATABASE_`,
double underscore to nest — which needs no file and no write access:

```bash
export EMDATABASE_LOCATIONS__PERSONAL=/scratch/emdatabase
export EMDATABASE_LOCATIONS__GROUP=/group/example_data
```
Examples of how to download data and configure storage locations can be found in our [docs](https://electronmicroscopy.github.io/emdatabase/examples/index.html), as well as in the [quantem-tutorials](https://github.com/electronmicroscopy/quantem-tutorials/tree/main/tutorials/core) repository.

## Adding a dataset

We welcome contributions of new or existing data!

Datasets are described by a YAML file in `emdatabase/index/`, one entry per file,
validated against `emdatabase/index/json-schema.json`. The class name is generated
from the top-level key:
Expand All @@ -185,26 +42,9 @@ vocabulary they come from: `acquisition` (how the data was taken) and `ml_task`
(what a model does). A dataset declares acquisition techniques only; a
`kind: weights` entry declares one of those plus the ML tasks it performs.

`size_bytes` is the file's `Content-Length` in bytes; the test suite checks it against
the server on every run. `emdatabase/index/vendors.yaml` lists the microscope
vendors and detector manufacturers already in use - a new one is fine, but a name close
to one already on the list fails CI as a misspelling. A technique close to one in
`techniques.yaml` fails the same way.

Submissions go through an issue. Fill in the
[new dataset issue form](https://github.com/electronmicroscopy/emdatabase/issues/new?template=new_dataset.yaml)
and an action writes the file and opens the pull request; the
[Add Dataset page](https://electronmicroscopy.github.io/emdatabase/add_dataset.html)
says what to have ready first. Or run `python -m emdatabase.new_dataset <url>`, which
fetches the checksum and size, prompts for the rest and writes the file for you to open
a pull request with. See [CONTRIBUTING.md](CONTRIBUTING.md).

Both take one download link and split it into `source`, `file` and, when the file is
not served at `source/file`, `url`. A Google Drive share link - what the share button
copies - is rewritten to the `uc?export=download&id=<id>` link that serves the file.
Authors go in as one per line, `Name; Affiliation; ORCID`, with the ORCID optional.
and an action writes the file and opens the pull request. For very large datasets it might be easier to use the CLI
as described on the [Add Dataset page](https://electronmicroscopy.github.io/emdatabase/add_dataset.html).
See [CONTRIBUTING.md](CONTRIBUTING.md).

Neither route needs the checksum or the size. An entry that is missing either one has
the file downloaded on GitHub and the fields filled in and pushed back to the branch; a
pull request from a fork, whose branch cannot be pushed to, fails with the values to
paste in instead.
12 changes: 6 additions & 6 deletions docs/source/_build_docs.py
Original file line number Diff line number Diff line change
Expand Up @@ -350,8 +350,7 @@ def generate_weights_html() -> str:
"<p>Trained model checkpoints, one entry per model. Downloading an entry "
"follows its <code>latest</code> link, which serves whatever the current "
"weights are; every earlier state of that link is kept as a dated version, "
"pinned to its checksum. Pick one with the selector; the load snippet "
"opens it with <code>weights_only=True</code>.</p>"
"pinned to its checksum."
"<p><code>download()</code> warns when the index on the project's "
"<code>main</code> branch has newer weights than your installed release, "
"and <code>download(refresh=True)</code> fetches them; the "
Expand Down Expand Up @@ -379,12 +378,13 @@ def generate_add_dataset_html() -> str:
'<main class="app-main"><div class="explainer">'
'<div class="app-hero page">'
"<h1>Add a Dataset</h1>"
"<p>Submissions go through an issue. Fill in the new-dataset issue form and "
"<p>Submissions go through an issue on GitHub. Fill in the new-dataset issue form and "
"an action turns it into the entry’s YAML file, fills in whatever it can "
"work out for itself, and opens the pull request.</p>"
"The checksum and the size are filled in automatically by "
"downloading the file, so both can be left blank in most cases. If the file is very large,"
" you may want to provide the checksum manually.</p>"
"</div>"
'<p class="note">The checksum and the size are filled in automatically by '
"downloading the file, so both can be left blank.</p>"
'<div class="submit-row">'
'<a class="btn-primary" target="_blank" rel="noopener" href="'
+ _ISSUE_URL
Expand All @@ -393,7 +393,7 @@ def generate_add_dataset_html() -> str:
"<h2>From a terminal</h2>"
"<pre><code>" + escape(_CLI_COMMAND) + "</code></pre>"
"<p>It asks the same questions at the prompt and writes the YAML file, leaving "
"the pull request to you. A trained model checkpoint takes "
"the pull request to you. A model checkpoint takes "
"<code>--kind weights</code> as well.</p>"
"<p>It computes the checksum from the file on your own machine rather than on "
"GitHub, which is the route to take for a very large file.</p>"
Expand Down
Loading
Loading