From 51834f120fcb8172ea4acbb2926d13502ce5b88f Mon Sep 17 00:00:00 2001 From: Mohammed Minhajuddin <68331751+minhajuddin2510@users.noreply.github.com> Date: Wed, 23 Sep 2026 09:54:57 -0400 Subject: [PATCH 1/3] feat: publish qsv profiles in the Fair Store and deploy it to Compute Engine Mirror (scripts/populate_fairstore.py): - mirror every registered portal (CKAN and DCAT/Socrata) with --site all - publish each qsv profile as dataset qsv_* fields plus a data dictionary, stats and frequency DataStore tables (with table views) and the describegpt JSON; only output that matches the profiled file's header is published - fail closed: reads raise instead of returning {} (CKANClient.call), Socrata enrichment is required under --apply, and one portal's failure no longer stops the others - reruns write only what changed (mirror_digest / qsv_digest), apply source renames, and report records withdrawn at the source - DataStore loads are verified by row count and rebuilt, never doubled Deployment (deploy/fairstore): one VM running CKAN 2.11.6, Postgres, Solr, Redis and Caddy (automatic HTTPS), with no service account, the metadata server blocked from containers, security headers, Secure cookies, CSRF checks on extension endpoints, daily disk snapshots and deletion protection. --- .gitignore | 5 + README.md | 97 +- deploy/fairstore/caddy/Caddyfile | 48 + deploy/fairstore/deploy.sh | 174 ++ deploy/fairstore/docker-compose.yml | 117 + deploy/fairstore/remote-up.sh | 74 + deploy/fairstore/vm-startup.sh | 52 + docker-compose.yml | 49 +- docker/fairstore/Dockerfile | 10 +- docker/fairstore/init-db.sh | 16 +- docker/fairstore/pin-secrets.sh | 16 + docker/fairstore/secure-cookies.sh | 21 + docker/fairstore/templates.sh | 3 + docker/fairstore/templates/package/read.html | 24 + .../package/snippets/additional_info.html | 13 + scripts/populate_fairstore.py | 2284 +++++++++++++++-- .../data_layer/connectors/ckan.py | 164 +- .../data_layer/connectors/dcat.py | 82 +- .../data_layer/qsv_profiling.py | 4 +- tests/unit/test_fairstore_populate.py | 1690 +++++++++++- 20 files changed, 4630 insertions(+), 313 deletions(-) create mode 100644 deploy/fairstore/caddy/Caddyfile create mode 100755 deploy/fairstore/deploy.sh create mode 100644 deploy/fairstore/docker-compose.yml create mode 100755 deploy/fairstore/remote-up.sh create mode 100755 deploy/fairstore/vm-startup.sh create mode 100644 docker/fairstore/pin-secrets.sh create mode 100644 docker/fairstore/secure-cookies.sh create mode 100644 docker/fairstore/templates.sh create mode 100644 docker/fairstore/templates/package/read.html create mode 100644 docker/fairstore/templates/package/snippets/additional_info.html diff --git a/.gitignore b/.gitignore index ae0b915..6c7f146 100644 --- a/.gitignore +++ b/.gitignore @@ -158,6 +158,11 @@ logs/ *.log .secrets/ +# API token files (deploy/fairstore/deploy.sh writes the mirror's outside the +# repository; these catch one saved here by mistake) +fairstore-token* +*token.txt + # Runtime-seeded local storage (contains user data / secrets — never commit) /users.json /roles.json diff --git a/README.md b/README.md index af04a05..2a9c06a 100644 --- a/README.md +++ b/README.md @@ -360,36 +360,107 @@ examples/ # Sample generated notebooks ## CKAN Fair Store mirror -Start the local CKAN 2.11 Fair Store with organization hierarchy support: +Start the local CKAN 2.11 Fair Store with organization hierarchy support. +Set `FAIRSTORE_SECRET_KEY` (any long random string, e.g. in `.env`) so API +tokens and sessions survive the container being recreated: ```bash docker compose --profile fairstore up -d --build fairstore ``` -Preview a registered CKAN portal before writing anything: +On first run, create a sysadmin and an API token for the mirror (keep the token +out of the repository): ```bash -python -m scripts.populate_fairstore \ - --site wprdc \ - --target-url http://localhost:5001 +docker compose exec fairstore ckan -c /srv/app/ckan.ini user add fairadmin email=fairadmin@localhost.localdomain password= +docker compose exec fairstore ckan -c /srv/app/ckan.ini sysadmin add fairadmin +docker compose exec fairstore ckan -c /srv/app/ckan.ini user token add fairadmin mirror +``` + +Preview every registered portal (CKAN and DCAT) before writing anything, or +pass one `--site` ID: + +```bash +python -m scripts.populate_fairstore --site all --target-url http://localhost:5001 ``` -To apply the mirror, create a target CKAN sysadmin token, expose it through -`CKAN_API_KEY` (or a protected file), and add `--apply`. The command performs a -collision preflight first, then preserves source organization, dataset, and -resource names and UUIDs. It creates one parent organization for the source -portal, attaches source organizations beneath it, and adds any matching qsv -profile to the resource as separate metadata. Re-running patches the preserved -UUIDs, so it does not duplicate resources. +Add `--apply` with the token in `FAIRSTORE_API_KEY` or `--api-key-file` to +write (`FAIRSTORE_URL` sets the default `--target-url`; the app's own +`CKAN_URL`/`CKAN_API_KEY` are deliberately not used, and a portal that is the +target is never mirrored into itself). Every read fails closed — a portal that +errors mid-read is not mirrored from a partial snapshot — and with `--site all` +one portal's failure does not stop the others (the exit status is non-zero). +The command runs a collision preflight first, then: + +- creates one parent organization per source portal and attaches the portal's + publishers beneath it (ckanext-hierarchy); +- for a **CKAN** portal, preserves organization, group, dataset, and resource + names and UUIDs; +- for a **DCAT** portal (`/data.json` from Socrata, DKAN, ArcGIS Hub), turns + publishers into organizations, themes into groups, and distributions into + resources, with UUIDv5 identifiers derived from the catalog's own IDs. A + Socrata catalog is enriched from the Socrata Discovery API with the owning + agency (the hierarchy) and the column list (`source_data_dictionary`); pass + `--no-enrich` to skip it; +- keeps every source field: anything the Fair Store's default schema has no + column for (ckanext-scheming fields, DCAT-US fields) becomes an extra, and + links to files uploaded to the source keep pointing at the source; +- publishes each dataset's qsv profile (from `data/{ckan,dcat}_onboard/`): the + AI description (its prose; describegpt's provenance block stays in the raw + output), AI tags, row and column counts and column names become `qsv_*` + fields on the dataset, and four resources are added to it — a data + dictionary (types, AI labels and descriptions, key statistics), the full + `qsv stats` and `qsv frequency` tables (DataStore tables, shown as sortable + tables and downloadable as CSV/JSON), and the raw describegpt JSON. API keys + that describegpt writes into its attribution are redacted, and stats, + frequency or describegpt output that does not match the profiled file's + header (left behind by another file's profile in the same directory) is not + published. The source data itself is never copied; +- skips records the source itself mirrored from another portal, so each + dataset is mirrored once from its origin. + +Re-running patches the preserved UUIDs, so it does not duplicate anything, +and it only writes what changed: each mirrored record carries a digest of what +was last written (`mirror_digest`, `qsv_digest`), so an unchanged dataset, +resource or qsv table is left alone. A record the source renamed is renamed +(DCAT names are derived, so those keep the name they were published under); +records the source no longer publishes are reported (`withdrawn_at_source`), +never deleted. The dry run reports the same created/updated/unchanged plan. ```bash python -m scripts.populate_fairstore \ - --site wprdc \ + --site all \ --target-url http://localhost:5001 \ --api-key-file /path/to/protected-token \ --apply ``` +### Deploying the Fair Store on GCP + +`deploy/fairstore/deploy.sh` runs the same stack on one Compute Engine VM next +to Verikan (in `us-central1` by default; `PROJECT` names the GCP project): CKAN on uwsgi +(built from `ckan/ckan-base`), Postgres, Solr, Redis, and Caddy for automatic +HTTPS. It is idempotent — the first run creates the static IP, firewall rule, +VM and a daily snapshot schedule; later runs re-ship `docker/fairstore` and +restart the stack. Secrets are generated on the VM (`/opt/fairstore/.env`) and +never leave it. The host is `FAIRSTORE_HOST` if set, else the one already +saved on the VM, else `.sslip.io` (first deploy); point a DNS name at the IP +and rerun with `FAIRSTORE_HOST` set to move it (the search index is rebuilt +for the new URLs). + +The VM has no service account (the site calls no Google API; a VM that still +has one is stopped once to remove it), and containers are blocked from the +metadata server except for DNS. + +```bash +gcloud auth login +PROJECT= deploy/fairstore/deploy.sh +``` + +The script prints the command that mints the mirror's API token into +`~/.config/verikan/fairstore-token` (outside the repository); then run the +mirror above with `--target-url https://` and that `--api-key-file`. + ## Development ```bash diff --git a/deploy/fairstore/caddy/Caddyfile b/deploy/fairstore/caddy/Caddyfile new file mode 100644 index 0000000..59028a7 --- /dev/null +++ b/deploy/fairstore/caddy/Caddyfile @@ -0,0 +1,48 @@ +# HTTPS for the Fair Store; Caddy obtains and renews the certificate itself. +# +# Mounted as a directory (./caddy:/etc/caddy), not as a single file: a +# redeploy's tar replaces the file with a new inode, and a single-file bind +# mount keeps serving the old one. remote-up.sh runs `caddy reload` afterwards. +{ + servers { + # Only TCP 443 is published and allowed through the firewall, so + # don't advertise HTTP/3 (UDP) to clients that would try it first. + protocols h1 h2 + } +} + +{$FAIRSTORE_HOST} { + encode zstd gzip + + # Access log with the client's address; CKAN only ever sees Caddy's. + log { + output stdout + format json + } + + header { + Strict-Transport-Security "max-age=31536000" + X-Content-Type-Options nosniff + Referrer-Policy strict-origin-when-cross-origin + # Don't advertise the server software (Caddy adds both). + -Server + -Via + } + + # CKAN's "Embed" button hands out an iframe of a resource view page + # (/dataset//resource//view/), so only those stay frameable by + # other sites. + @not_view not path */view/* + header @not_view { + X-Frame-Options SAMEORIGIN + defer + } + + reverse_proxy fairstore:5000 { + # uwsgi's http router closes idle keep-alive connections; a POST sent on + # a pooled one fails with a 502, and Go never retries a non-idempotent request. + transport http { + keepalive off + } + } +} diff --git a/deploy/fairstore/deploy.sh b/deploy/fairstore/deploy.sh new file mode 100755 index 0000000..eff84a6 --- /dev/null +++ b/deploy/fairstore/deploy.sh @@ -0,0 +1,174 @@ +#!/usr/bin/env bash +# Deploy the Fair Store (CKAN 2.11 + ckanext-hierarchy) next to Verikan. +# +# One Compute Engine VM runs the same stack as the local `fairstore` compose +# profile — CKAN, Postgres, Solr, Redis — plus Caddy for automatic HTTPS. A VM +# rather than Cloud Run because Postgres and Solr need persistent disks. +# +# Idempotent: creates the static IP, firewall rule, VM and daily disk snapshots +# on the first run, then (re)ships docker/fairstore + this directory and +# restarts the stack. Secrets are generated on the VM and stay there. +# +# gcloud auth login # the session expires; re-auth is interactive +# PROJECT= deploy/fairstore/deploy.sh +# +# The VM runs without a service account: the site calls no Google API, and the +# default one is a project Editor. A VM that still has one is stopped, stripped +# of it and started again — a few minutes of downtime, once. +# +# The host is FAIRSTORE_HOST if set, else the one the VM already serves (its +# .env), else .sslip.io, a wildcard DNS name that resolves to the VM so +# Caddy can get a certificate before a real domain exists. Point a DNS name at +# the IP and rerun with FAIRSTORE_HOST set to move it. +set -euo pipefail + +PROJECT="${PROJECT:?set PROJECT to the GCP project id to deploy into}" +REGION="${REGION:-us-central1}" +ZONE="${ZONE:-us-central1-a}" +VM="${VM:-verikan-fairstore}" +MACHINE_TYPE="${MACHINE_TYPE:-e2-medium}" +DISK_SIZE="${DISK_SIZE:-50GB}" +NETWORK_TAG="${VM}-web" +REMOTE_DIR=/opt/fairstore +ROOT="$(cd "$(dirname "$0")/../.." && pwd)" +STARTUP_SCRIPT="$ROOT/deploy/fairstore/vm-startup.sh" +# Outside the repository, so the mirror's sysadmin token can never be committed. +TOKEN_FILE="${XDG_CONFIG_HOME:-$HOME/.config}/verikan/fairstore-token" + +gc() { gcloud --project "$PROJECT" --quiet "$@"; } +on_vm() { gc compute ssh "$VM" --zone "$ZONE" --command "$1"; } +vm_field() { gc compute instances describe "$VM" --zone "$ZONE" --format="value($1)"; } + +# wait_for +# Prints a dot per failed try and, on timeout, the last try's output (the why). +wait_for() { + local limit="$1" pause="$2" what="$3" deadline out + shift 3 + deadline=$((SECONDS + limit)) + printf 'Waiting for %s ' "$what" + until out="$("$@" 2>&1)"; do + if ((SECONDS >= deadline)); then + echo + echo "deploy: gave up waiting for $what after $((limit / 60)) min; last try said:" >&2 + printf '%s\n' "$out" | tail -n 20 >&2 + exit 1 + fi + printf . + sleep "$pause" + done + echo " ok" +} + +gc projects describe "$PROJECT" --format='value(projectId)' >/dev/null \ + || { echo "gcloud cannot reach $PROJECT; run: gcloud auth login" >&2; exit 1; } + +if ! gc compute addresses describe "${VM}-ip" --region "$REGION" >/dev/null 2>&1; then + gc compute addresses create "${VM}-ip" --region "$REGION" +fi +IP="$(gc compute addresses describe "${VM}-ip" --region "$REGION" --format='value(address)')" + +if ! gc compute firewall-rules describe "${VM}-web" >/dev/null 2>&1; then + gc compute firewall-rules create "${VM}-web" --network default --direction INGRESS \ + --allow tcp:80,tcp:443 --target-tags "$NETWORK_TAG" --source-ranges 0.0.0.0/0 +fi + +if ! gc compute instances describe "$VM" --zone "$ZONE" >/dev/null 2>&1; then + gc compute instances create "$VM" --zone "$ZONE" --machine-type "$MACHINE_TYPE" \ + --image-family debian-12 --image-project debian-cloud \ + --boot-disk-size "$DISK_SIZE" --boot-disk-type pd-balanced \ + --address "$IP" --tags "$NETWORK_TAG" \ + --no-service-account --no-scopes --deletion-protection \ + --labels app=verikan,component=fairstore \ + --metadata-from-file startup-script="$STARTUP_SCRIPT" +else + status="$(vm_field status)" + sa="$(vm_field 'serviceAccounts[].email')" + if [ -n "$sa" ]; then + echo "$VM runs as $sa, which the site never uses; removing it." + echo "That needs the VM stopped: the Fair Store is down for a few minutes." + [ "$status" = TERMINATED ] || gc compute instances stop "$VM" --zone "$ZONE" + gc compute instances set-service-account "$VM" --zone "$ZONE" \ + --no-service-account --no-scopes + status=TERMINATED + fi + # The boot disk holds every database: guard the VM against deletion. + gc compute instances update "$VM" --zone "$ZONE" --deletion-protection + # GCE runs the startup script stored in instance metadata, so ship the + # current one; it applies from the next boot (remote-up.sh covers this one). + gc compute instances add-metadata "$VM" --zone "$ZONE" \ + --metadata-from-file startup-script="$STARTUP_SCRIPT" + if [ "$status" = TERMINATED ]; then + echo "Starting $VM..." + gc compute instances start "$VM" --zone "$ZONE" + fi +fi + +# Postgres, Solr and uploads all live on the boot disk: keep a week of daily +# snapshots. +if ! gc compute resource-policies describe "${VM}-daily" --region "$REGION" >/dev/null 2>&1; then + gc compute resource-policies create snapshot-schedule "${VM}-daily" --region "$REGION" \ + --daily-schedule --start-time 07:00 --max-retention-days 7 \ + --on-source-disk-delete keep-auto-snapshots +fi +policies="$(gc compute disks describe "$VM" --zone "$ZONE" --format='value(resourcePolicies)')" +case ";${policies};" in + *"/resourcePolicies/${VM}-daily;"*) ;; + *) gc compute disks add-resource-policies "$VM" --zone "$ZONE" --resource-policies "${VM}-daily" ;; +esac + +# `compose ls` needs both the compose plugin and a running daemon. +wait_for 900 15 "Docker on $VM" on_vm "sudo docker compose ls" + +# Default to the host the VM already serves: falling back to sslip.io on every +# run would silently move a site that has been given a real domain. +if [ -n "${FAIRSTORE_HOST:-}" ]; then + HOST="$FAIRSTORE_HOST" +else + HOST="$(on_vm "if sudo test -f $REMOTE_DIR/.env; then sudo sed -n 's/^FAIRSTORE_HOST=//p' $REMOTE_DIR/.env; fi" \ + | tail -n 1 | tr -d '\r')" + HOST="${HOST:-${IP//./-}.sslip.io}" +fi +[[ "$HOST" =~ ^[A-Za-z0-9.-]+$ ]] || { echo "deploy: bad host name: '$HOST'" >&2; exit 1; } +echo "Deploying to https://$HOST" + +bundle="$(mktemp -d)" +trap 'rm -rf "$bundle"' EXIT +mkdir -p "$bundle/stack/ckan" +cp "$ROOT"/deploy/fairstore/{docker-compose.yml,remote-up.sh} "$bundle/stack/" +cp -R "$ROOT"/deploy/fairstore/caddy "$bundle/stack/" +cp -R "$ROOT"/docker/fairstore/. "$bundle/stack/ckan/" +# macOS tar would add AppleDouble ._* files and xattrs, and record the local +# user as owner. Members are named rather than ".", so extracting never resets +# the mode of /opt/fairstore itself. +COPYFILE_DISABLE=1 tar --no-xattrs --owner=root:0 --group=root:0 --exclude .DS_Store \ + -C "$bundle/stack" -czf "$bundle/fairstore.tgz" docker-compose.yml remote-up.sh caddy ckan +gc compute scp "$bundle/fairstore.tgz" "$VM:/tmp/fairstore.tgz" --zone "$ZONE" +# ckan/ is only a build context (its init-db.sh runs only when Postgres +# initialises an empty volume), so it is replaced whole and files deleted here +# disappear there. .env is never touched, and neither is the caddy/ directory +# itself: Caddy bind-mounts it, so only the files inside it are replaced. +# The first deploys' bundles left macOS ._* files and the operator's ownership +# at the top level; both are tidied here. +on_vm "sudo mkdir -p $REMOTE_DIR && sudo chown root:root $REMOTE_DIR \ + && sudo find $REMOTE_DIR -maxdepth 1 -name '._*' -delete \ + && sudo rm -rf $REMOTE_DIR/ckan \ + && sudo tar --no-same-owner -xzf /tmp/fairstore.tgz -C $REMOTE_DIR \ + && rm -f /tmp/fairstore.tgz \ + && sudo bash $REMOTE_DIR/remote-up.sh $HOST" + +wait_for 600 10 "https://$HOST" \ + curl -fsS -o /dev/null --max-time 20 "https://$HOST/api/3/action/status_show" +cat < "$TOKEN_FILE") +then load every registered portal: + python -m scripts.populate_fairstore --site all --apply \\ + --target-url https://$HOST --api-key-file "$TOKEN_FILE" + +The fairadmin password is FAIRSTORE_ADMIN_PASSWORD in $REMOTE_DIR/.env on the VM. +EOF diff --git a/deploy/fairstore/docker-compose.yml b/deploy/fairstore/docker-compose.yml new file mode 100644 index 0000000..fded497 --- /dev/null +++ b/deploy/fairstore/docker-compose.yml @@ -0,0 +1,117 @@ +# Production Fair Store: CKAN 2.11 (uwsgi) + ckanext-hierarchy behind Caddy, +# which provisions HTTPS automatically. deploy.sh ships this directory, with +# docker/fairstore as ./ckan, to one Compute Engine VM next to Verikan. Every +# secret comes from .env, generated on the VM on first deploy (remote-up.sh). +# +# Same services and names as the fairstore profile in the root +# docker-compose.yml, minus the dev server, plus Caddy and hardening. +name: fairstore + +# Container logs (Caddy's access log included) are rotated, not kept forever. +x-logging: &logging + driver: json-file + options: + max-size: "20m" + max-file: "5" + +services: + fairstore: + build: + context: ./ckan + args: + CKAN_BASE_IMAGE: ckan/ckan-base:2.11.6 + image: fairstore-ckan:2.11.6 + depends_on: + fairstore-db: + condition: service_healthy + fairstore-solr: + condition: service_started + fairstore-redis: + condition: service_started + environment: + - CKAN_SITE_URL=https://${FAIRSTORE_HOST:?} + - CKAN_SQLALCHEMY_URL=postgresql://ckan:${FAIRSTORE_DB_PASSWORD:?}@fairstore-db/ckan + - CKAN_DATASTORE_WRITE_URL=postgresql://ckan_datastore_write:${FAIRSTORE_DATASTORE_PASSWORD:?}@fairstore-db/datastore + - CKAN_DATASTORE_READ_URL=postgresql://ckan_datastore_read:${FAIRSTORE_DATASTORE_PASSWORD:?}@fairstore-db/datastore + - CKAN_SOLR_URL=http://fairstore-solr:8983/solr/ckan + - CKAN_REDIS_URL=redis://fairstore-redis:6379/1 + - CKAN_SECRET_KEY=${FAIRSTORE_SECRET_KEY:?} + # Created by the image's prerun on first start; skipped once it exists. + - CKAN_SYSADMIN_NAME=fairadmin + - CKAN_SYSADMIN_PASSWORD=${FAIRSTORE_ADMIN_PASSWORD:?} + - CKAN_SYSADMIN_EMAIL=fairadmin@${FAIRSTORE_HOST:?} + - CKAN__PLUGINS=envvars datastore datatables_view text_view image_view hierarchy_display hierarchy_form + - CKAN__VIEWS__DEFAULT_VIEWS=datatables_view text_view image_view + # The 2.11.6 image's ckan.ini sets ckan.uploads_enabled to an empty + # value, which turns file uploads off in the web forms. + - CKAN__UPLOADS_ENABLED=true + # CKAN exempts extension endpoints (the DataStore dictionary editor, the + # DataTables AJAX feed) from CSRF checks by default, and SameSite=Lax + # does not cover them here: sslip.io is not a public suffix, so every + # other *.sslip.io site counts as same-site. + - CKAN__CSRF_PROTECTION__IGNORE_EXTENSIONS=false + - CKAN__SITE_TITLE=Verikan Fair Store + # Records are written by the mirror and edited by admins; nobody signs + # up through the web, and the user list is not public. + - CKAN__AUTH__CREATE_USER_VIA_WEB=false + - CKAN__AUTH__PUBLIC_USER_DETAILS=false + volumes: + - fairstore_data:/var/lib/ckan + restart: unless-stopped + logging: *logging + + fairstore-db: + image: postgres:16-alpine + environment: + - POSTGRES_USER=postgres + - POSTGRES_PASSWORD=${FAIRSTORE_POSTGRES_PASSWORD:?} + - POSTGRES_DB=postgres + - CKAN_DB_PASSWORD=${FAIRSTORE_DB_PASSWORD:?} + - DATASTORE_DB_PASSWORD=${FAIRSTORE_DATASTORE_PASSWORD:?} + volumes: + - fairstore_db:/var/lib/postgresql/data + - ./ckan/init-db.sh:/docker-entrypoint-initdb.d/init-db.sh:ro + healthcheck: + test: ["CMD-SHELL", "pg_isready -U postgres"] + interval: 10s + timeout: 5s + retries: 10 + restart: unless-stopped + logging: *logging + + fairstore-solr: + image: ckan/ckan-solr:2.11-solr9 + volumes: + - fairstore_solr:/var/solr + restart: unless-stopped + logging: *logging + + fairstore-redis: + image: redis:7-alpine + restart: unless-stopped + logging: *logging + + caddy: + image: caddy:2-alpine + depends_on: + - fairstore + ports: + - "80:80" + - "443:443" + environment: + - FAIRSTORE_HOST=${FAIRSTORE_HOST:?} + volumes: + # The directory, not the file: a redeploy replaces the Caddyfile's inode, + # which a single-file bind mount would never see (caddy/Caddyfile). + - ./caddy:/etc/caddy:ro + - caddy_data:/data + - caddy_config:/config + restart: unless-stopped + logging: *logging + +volumes: + fairstore_data: + fairstore_db: + fairstore_solr: + caddy_data: + caddy_config: diff --git a/deploy/fairstore/remote-up.sh b/deploy/fairstore/remote-up.sh new file mode 100755 index 0000000..e212e11 --- /dev/null +++ b/deploy/fairstore/remote-up.sh @@ -0,0 +1,74 @@ +#!/usr/bin/env bash +# Runs on the Fair Store VM, as root, from deploy.sh: create the secrets on the +# first deploy, then build and (re)start the stack. +# +# remote-up.sh +# +# .env is generated once and never leaves the VM. Rotating a value means +# editing .env; note the database passwords are only applied when Postgres +# initialises an empty volume. +set -euo pipefail +cd "$(dirname "$0")" +host="${1:?usage: remote-up.sh }" +[[ "$host" =~ ^[A-Za-z0-9.-]+$ ]] || { echo "remote-up: bad hostname: $host" >&2; exit 1; } + +# retry +retry() { + local tries="$1" pause="$2" i + shift 2 + for ((i = 1; ; i++)); do + "$@" && return 0 + ((i < tries)) || return 1 + sleep "$pause" + done +} + +# Containers get no path to the metadata server, DNS (port 53) excepted; see +# vm-startup.sh, which re-adds these on every boot. Added here too so a deploy +# takes effect without a reboot. +for proto in tcp udp; do + iptables -C DOCKER-USER -d 169.254.169.254 -p "$proto" ! --dport 53 -j DROP 2>/dev/null \ + || iptables -I DOCKER-USER -d 169.254.169.254 -p "$proto" ! --dport 53 -j DROP +done + +if [ ! -f .env ]; then + umask 077 + secret() { openssl rand -hex 24; } + { + echo "FAIRSTORE_HOST=${host}" + echo "FAIRSTORE_SECRET_KEY=$(secret)" + echo "FAIRSTORE_POSTGRES_PASSWORD=$(secret)" + echo "FAIRSTORE_DB_PASSWORD=$(secret)" + echo "FAIRSTORE_DATASTORE_PASSWORD=$(secret)" + echo "FAIRSTORE_ADMIN_PASSWORD=$(secret)" + } > .env +fi +# Solr keeps each dataset as indexed, with absolute URLs built from the site +# URL, so a host change needs a reindex. The marker survives a failed run, so +# the next deploy still does it. +if [ "$(sed -n 's/^FAIRSTORE_HOST=//p' .env | tail -n 1)" != "$host" ]; then + touch .reindex-pending +fi +sed -i "s|^FAIRSTORE_HOST=.*|FAIRSTORE_HOST=${host}|" .env + +docker compose up -d --build --remove-orphans +# Caddy reads its config only at start; reload applies a changed Caddyfile +# without dropping connections. Retried because a just-created container may +# not have its admin endpoint up yet. +retry 10 3 docker compose exec -T caddy caddy reload --config /etc/caddy/Caddyfile --adapter caddyfile \ + || { echo "remote-up: caddy reload failed (the previous config is still live)" >&2; exit 1; } +# Before the caddy/ directory mount the Caddyfile sat here; nothing reads it now. +rm -f Caddyfile + +if [ -f .reindex-pending ]; then + ckan_up() { + docker compose exec -T fairstore python3 -c \ + 'import urllib.request as u; u.urlopen("http://127.0.0.1:5000/api/3/action/status_show", timeout=10)' \ + >/dev/null 2>&1 + } + echo "Site URL is now https://${host}: waiting for CKAN, then rebuilding the search index..." + retry 60 5 ckan_up \ + || { echo "remote-up: CKAN did not answer after 60 tries; search index NOT rebuilt (the next deploy retries)" >&2; exit 1; } + docker compose exec -T fairstore ckan -c /srv/app/ckan.ini search-index rebuild + rm -f .reindex-pending +fi diff --git a/deploy/fairstore/vm-startup.sh b/deploy/fairstore/vm-startup.sh new file mode 100755 index 0000000..25479b8 --- /dev/null +++ b/deploy/fairstore/vm-startup.sh @@ -0,0 +1,52 @@ +#!/bin/bash +# Compute Engine startup script for the Fair Store VM (runs as root on every +# boot). Installs Docker Engine and the compose plugin on first boot, then, on +# every boot, keeps containers away from the metadata server. The stack itself +# restarts on its own (restart: unless-stopped). +# +# GCE reads this from instance metadata, not from the repository: deploy.sh +# re-uploads it on each run, and a change takes effect at the next boot. +set -euo pipefail + +install_docker() { + # A first boot races GCE's own apt runs (unattended-upgrades, the guest + # agent), so wait for the lock rather than fail. + local apt=(apt-get -o DPkg::Lock::Timeout=600) + "${apt[@]}" update + "${apt[@]}" install -y ca-certificates curl + install -m 0755 -d /etc/apt/keyrings + curl -fsSL https://download.docker.com/linux/debian/gpg -o /etc/apt/keyrings/docker.asc + chmod a+r /etc/apt/keyrings/docker.asc + . /etc/os-release + echo "deb [arch=$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/docker.asc] https://download.docker.com/linux/debian ${VERSION_CODENAME} stable" \ + > /etc/apt/sources.list.d/docker.list + "${apt[@]}" update + "${apt[@]}" install -y docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin + systemctl enable --now docker +} + +# The compose plugin is the last package the stack needs: a first boot cut +# short after the docker CLI was unpacked is finished on the next boot. +docker compose version >/dev/null 2>&1 || install_docker + +# Nothing in the stack calls a Google API, so containers get no path to the +# metadata server (the VM has no service account either; this is the second +# layer). Port 53 stays open: on GCE that address is also the VM's DNS +# resolver, which the image build's RUN steps and any default-bridge +# container query directly (containers on the compose network go through +# Docker's embedded resolver), so a blanket DROP would break image builds. +# iptables rules do not survive a reboot, and DOCKER-USER only exists once +# dockerd has started. remote-up.sh adds the same rules on each deploy. +for _ in $(seq 60); do + iptables -n -L DOCKER-USER >/dev/null 2>&1 && break + sleep 5 +done +if ! iptables -n -L DOCKER-USER >/dev/null 2>&1; then + echo "vm-startup: no DOCKER-USER chain after 5 min (is dockerd running?);" \ + "containers can still reach the metadata server" >&2 + exit 1 +fi +for proto in tcp udp; do + iptables -C DOCKER-USER -d 169.254.169.254 -p "$proto" ! --dport 53 -j DROP 2>/dev/null \ + || iptables -I DOCKER-USER -d 169.254.169.254 -p "$proto" ! --dport 53 -j DROP +done diff --git a/docker-compose.yml b/docker-compose.yml index 8839665..707a888 100644 --- a/docker-compose.yml +++ b/docker-compose.yml @@ -76,8 +76,9 @@ services: # then open http://localhost:5001 (host 5001 -> container 5000; macOS # reserves 5000 for AirPlay). # - # First run: create an admin and an API token with - # docker compose exec fairstore ckan -c /srv/app/ckan.ini sysadmin add admin + # First run: create a sysadmin and an API token (commands under "CKAN Fair + # Store mirror" in README.md), load every registered portal with + # python -m scripts.populate_fairstore --site all --apply # then set CKAN_URL / CKAN_API_KEY for the app. # --------------------------------------------------------------------- fairstore: @@ -85,28 +86,48 @@ services: build: context: ./docker/fairstore dockerfile: Dockerfile - image: data-concierge-fairstore:2.11.3 + image: data-concierge-fairstore:2.11.6 platform: linux/amd64 depends_on: fairstore-db: condition: service_healthy fairstore-solr: condition: service_started + fairstore-redis: + condition: service_started ports: - "${FAIRSTORE_PORT:-5001}:5000" environment: - - CKAN_SITE_URL=http://localhost:5000 + # The URL browsers reach CKAN at. CKAN builds upload and DataStore dump + # links from it, so it must be the host port, not the container's 5000. + - CKAN_SITE_URL=http://localhost:${FAIRSTORE_PORT:-5001} # Credentials must match the users the postgres image's init scripts # create (see fairstore-db below): ckan / ckan_datastore_write / # ckan_datastore_read. - - CKAN_SQLALCHEMY_URL=postgresql://ckan:ckan@fairstore-db/ckan - - CKAN_DATASTORE_WRITE_URL=postgresql://ckan_datastore_write:datastore@fairstore-db/datastore - - CKAN_DATASTORE_READ_URL=postgresql://ckan_datastore_read:datastore@fairstore-db/datastore + - CKAN_SQLALCHEMY_URL=postgresql://ckan:${FAIRSTORE_DB_PASSWORD:-ckan}@fairstore-db/ckan + - CKAN_DATASTORE_WRITE_URL=postgresql://ckan_datastore_write:${FAIRSTORE_DATASTORE_PASSWORD:-datastore}@fairstore-db/datastore + - CKAN_DATASTORE_READ_URL=postgresql://ckan_datastore_read:${FAIRSTORE_DATASTORE_PASSWORD:-datastore}@fairstore-db/datastore - CKAN_SOLR_URL=http://fairstore-solr:8983/solr/ckan + - CKAN_REDIS_URL=redis://fairstore-redis:6379/1 + # Pins the session and API-token signing secrets (docker/fairstore/ + # pin-secrets.sh). Unset, each new container signs with fresh random + # secrets and every issued API token stops working. + - CKAN_SECRET_KEY=${FAIRSTORE_SECRET_KEY:-} # hierarchy_display must load before hierarchy_form. Organizations that # represent source portals can then be parents of the organizations # mirrored from those portals. - - CKAN__PLUGINS=envvars datastore hierarchy_display hierarchy_form + - CKAN__PLUGINS=envvars datastore datatables_view text_view image_view hierarchy_display hierarchy_form + # qsv profiles (stats, frequency, data dictionary) are DataStore tables, + # shown as sortable tables; describegpt's JSON output is shown inline. + - CKAN__VIEWS__DEFAULT_VIEWS=datatables_view text_view image_view + # The 2.11.6 image's ckan.ini sets ckan.uploads_enabled to an empty + # value, which turns file uploads off in the web forms. + - CKAN__UPLOADS_ENABLED=true + # CKAN exempts extension endpoints (the DataStore dictionary editor, the + # DataTables AJAX feed) from CSRF checks by default, and SameSite=Lax + # does not cover them here: sslip.io is not a public suffix, so every + # other *.sslip.io site counts as same-site. + - CKAN__CSRF_PROTECTION__IGNORE_EXTENSIONS=false volumes: - fairstore_data:/var/lib/ckan restart: unless-stopped @@ -119,8 +140,11 @@ services: image: postgres:16-alpine environment: - POSTGRES_USER=postgres - - POSTGRES_PASSWORD=postgres + - POSTGRES_PASSWORD=${FAIRSTORE_POSTGRES_PASSWORD:-postgres} - POSTGRES_DB=postgres + # Read by init-db.sh on first start; must match the URLs above. + - CKAN_DB_PASSWORD=${FAIRSTORE_DB_PASSWORD:-ckan} + - DATASTORE_DB_PASSWORD=${FAIRSTORE_DATASTORE_PASSWORD:-datastore} volumes: - fairstore_db:/var/lib/postgresql/data - ./docker/fairstore/init-db.sh:/docker-entrypoint-initdb.d/init-db.sh:ro @@ -138,6 +162,13 @@ services: - fairstore_solr:/var/solr restart: unless-stopped + # CKAN 2.11 logs a critical error on every start without Redis, and its + # background jobs queue lives there. + fairstore-redis: + profiles: ["fairstore"] + image: redis:7-alpine + restart: unless-stopped + # Optional: Redis for caching (uncomment if needed) # redis: # image: redis:7-alpine diff --git a/docker/fairstore/Dockerfile b/docker/fairstore/Dockerfile index 59a688b..3361ca3 100644 --- a/docker/fairstore/Dockerfile +++ b/docker/fairstore/Dockerfile @@ -1,4 +1,8 @@ -FROM ckan/ckan-dev:2.11.3 +# ckan-dev runs Flask's debug server (with the debug toolbar) for local work. +# A deployed Fair Store builds from ckan-base, the uwsgi production image; +# both share the same start scripts and /docker-entrypoint.d hooks. +ARG CKAN_BASE_IMAGE=ckan/ckan-dev:2.11.6 +FROM ${CKAN_BASE_IMAGE} # Keep the extension revision deterministic and aligned with the CKAN 2.11 # version used by the hosted Fair Store. This revision's CI covers CKAN 2.11. @@ -7,5 +11,9 @@ ARG CKANEXT_HIERARCHY_SHA=53c1ee74805d79aef7c9c01cc513eb881cb8928f USER root RUN pip install --no-cache-dir \ "git+https://github.com/ckan/ckanext-hierarchy.git@${CKANEXT_HIERARCHY_SHA}#egg=ckanext-hierarchy" +COPY pin-secrets.sh /docker-entrypoint.d/10-pin-secrets.sh +COPY templates.sh /docker-entrypoint.d/20-templates.sh +COPY secure-cookies.sh /docker-entrypoint.d/30-secure-cookies.sh +COPY templates /srv/app/fairstore_templates USER ckan diff --git a/docker/fairstore/init-db.sh b/docker/fairstore/init-db.sh index 80a01fc..3c23a06 100755 --- a/docker/fairstore/init-db.sh +++ b/docker/fairstore/init-db.sh @@ -6,16 +6,20 @@ # plain Postgres and creates exactly the users and databases CKAN needs here. # Runs once, on an empty data directory. # -# Local-only credentials. A deployed Fair Store needs real secrets (see the -# real secrets managed outside this repo). +# Passwords come from CKAN_DB_PASSWORD / DATASTORE_DB_PASSWORD. The defaults +# are local-only; a deployed Fair Store sets real ones (deploy/fairstore). +# They are passed as psql variables and quoted by psql (:'var'), not spliced +# into the SQL by the shell. set -e -psql -v ON_ERROR_STOP=1 --username "$POSTGRES_USER" <<-SQL - CREATE USER ckan WITH PASSWORD 'ckan'; +psql -v ON_ERROR_STOP=1 --username "$POSTGRES_USER" \ + -v ckan_pw="${CKAN_DB_PASSWORD:-ckan}" \ + -v datastore_pw="${DATASTORE_DB_PASSWORD:-datastore}" <<-SQL + CREATE USER ckan WITH PASSWORD :'ckan_pw'; CREATE DATABASE ckan OWNER ckan; - CREATE USER ckan_datastore_write WITH PASSWORD 'datastore'; - CREATE USER ckan_datastore_read WITH PASSWORD 'datastore'; + CREATE USER ckan_datastore_write WITH PASSWORD :'datastore_pw'; + CREATE USER ckan_datastore_read WITH PASSWORD :'datastore_pw'; CREATE DATABASE datastore OWNER ckan_datastore_write; SQL diff --git a/docker/fairstore/pin-secrets.sh b/docker/fairstore/pin-secrets.sh new file mode 100644 index 0000000..53d5b06 --- /dev/null +++ b/docker/fairstore/pin-secrets.sh @@ -0,0 +1,16 @@ +# Pin CKAN's signing secrets from the environment (sourced by the stock +# start scripts from /docker-entrypoint.d, just before the server starts). +# +# Those scripts write fresh random secrets into the container's ckan.ini +# whenever it is created, so recreating the container — a deploy, a config +# change — silently signed out every session and invalidated every API token, +# including the one the mirror uses. With CKAN_SECRET_KEY set, the secrets +# survive. Unset, the stock random-per-container behaviour is kept. +if [ -n "${CKAN_SECRET_KEY:-}" ]; then + ckan config-tool "$CKAN_INI" \ + "SECRET_KEY=${CKAN_SECRET_KEY}" \ + "WTF_CSRF_SECRET_KEY=${CKAN_SECRET_KEY}" \ + "api_token.jwt.encode.secret=string:${CKAN_SECRET_KEY}" \ + "api_token.jwt.decode.secret=string:${CKAN_SECRET_KEY}" + echo "pin-secrets: signing secrets taken from CKAN_SECRET_KEY" +fi diff --git a/docker/fairstore/secure-cookies.sh b/docker/fairstore/secure-cookies.sh new file mode 100644 index 0000000..bc1b9b9 --- /dev/null +++ b/docker/fairstore/secure-cookies.sh @@ -0,0 +1,21 @@ +# Mark CKAN's session and remember-me cookies Secure when the site is served +# over HTTPS (sourced by the stock start scripts from /docker-entrypoint.d). +# +# The image's ckan.ini ships both as false, and uppercase Flask keys cannot be +# set through the envvars plugin (it lowercases names), hence config-tool. +# Local dev on http://localhost keeps them false: a browser never sends a +# Secure cookie back over plain HTTP, so nobody could stay logged in. +# +# The image also ships REMEMBER_COOKIE_SAMESITE = None, which lets another site +# make a remembered sysadmin's browser send the one-year cookie with a +# cross-site POST (CKAN's DataStore dictionary form is CSRF-exempt); Lax +# works over plain HTTP too, so it is set everywhere. +ckan config-tool "$CKAN_INI" "REMEMBER_COOKIE_SAMESITE=Lax" +case "${CKAN_SITE_URL:-}" in + https://*) + ckan config-tool "$CKAN_INI" \ + "SESSION_COOKIE_SECURE=true" \ + "REMEMBER_COOKIE_SECURE=true" + echo "secure-cookies: session cookies marked Secure" + ;; +esac diff --git a/docker/fairstore/templates.sh b/docker/fairstore/templates.sh new file mode 100644 index 0000000..e10b968 --- /dev/null +++ b/docker/fairstore/templates.sh @@ -0,0 +1,3 @@ +# Load the Fair Store's template overrides (docker/fairstore/templates), which +# render the qsv AI summary on dataset pages. Sourced from /docker-entrypoint.d. +ckan config-tool "$CKAN_INI" "extra_template_paths = /srv/app/fairstore_templates" diff --git a/docker/fairstore/templates/package/read.html b/docker/fairstore/templates/package/read.html new file mode 100644 index 0000000..3b6df80 --- /dev/null +++ b/docker/fairstore/templates/package/read.html @@ -0,0 +1,24 @@ +{% ckan_extends %} + +{#- The dataset's qsv profile (scripts/populate_fairstore.py): the summary qsv + describegpt wrote, rendered as markdown under the source's own + description. The mirror stores only the prose; describegpt's provenance + block is carried by qsv_model / qsv_version and the describegpt resource. -#} +{% block package_notes %} + {{ super() }} + {% set qsv_text = pkg.extras | selectattr('key', 'equalto', 'qsv_description') | map(attribute='value') | first %} + {% if qsv_text %} + {% set qsv_model = pkg.extras | selectattr('key', 'equalto', 'qsv_model') | map(attribute='value') | first %} +
+

AI summary

+
+ {{ h.render_markdown(qsv_text) }} +
+

+ Written by qsv describegpt{% if qsv_model %} ({{ qsv_model }}){% endif %} from a profile of this + dataset's data; AI-written text may contain inaccuracies. The profile's data dictionary + and its other qsv resources are listed under Data and Resources. +

+
+ {% endif %} +{% endblock %} diff --git a/docker/fairstore/templates/package/snippets/additional_info.html b/docker/fairstore/templates/package/snippets/additional_info.html new file mode 100644 index 0000000..abd0663 --- /dev/null +++ b/docker/fairstore/templates/package/snippets/additional_info.html @@ -0,0 +1,13 @@ +{% ckan_extends %} + +{#- The qsv AI summary is rendered as markdown by package/read.html, so its raw + text is left out of the key/value table. -#} +{% block extras scoped %} + {% for extra in h.sorted_extras(pkg_dict.extras, exclude=['qsv_description']) %} + {% set key, value = extra %} + + {{ _(key|e) }} + {{ value }} + + {% endfor %} +{% endblock %} diff --git a/scripts/populate_fairstore.py b/scripts/populate_fairstore.py index 4283b9d..7285c81 100644 --- a/scripts/populate_fairstore.py +++ b/scripts/populate_fairstore.py @@ -1,22 +1,48 @@ #!/usr/bin/env python -"""Mirror a source CKAN catalog into the Fair Store. - -The mirror keeps source organization, group, dataset, and resource names and -UUIDs. A top-level organization represents the source portal; source root -organizations are attached beneath it through ckanext-hierarchy. Available -qsv profiling metadata from ``data/ckan_onboard//index.json`` is added to -the corresponding resource without replacing the source description. +"""Mirror source catalogs into the Fair Store. + +Both portal types in the registry are mirrored: + +``ckan`` + Source organization, group, dataset, and resource names and UUIDs are kept. +``dcat`` + A DCAT-US catalog (``/data.json`` from Socrata, DKAN, ArcGIS Hub, ...) has + no organization records or UUIDs, so they are derived deterministically: + publishers become organizations, themes become groups, distributions become + resources, and every UUID is a UUIDv5 of the source identifier (or the + source's own UUID, when it publishes one). A Socrata catalog is enriched + from the Socrata Discovery API, whose owning agency and column list the + catalog document omits. + +A top-level organization represents each source portal, and the portal's +publishers are attached beneath it through ckanext-hierarchy. Every field a +source publishes is kept: fields the Fair Store's default schema has no column +for are carried as extras instead of being silently dropped. Available qsv +profiling metadata from ``data/{ckan,dcat}_onboard//index.json`` is added +to the corresponding resource without replacing the source description. + +A source record that its portal itself mirrored from another portal (it +carries ``mirror_source_portal`` naming that portal) is skipped: mirror the +origin portal directly so provenance names the real source. The command is read-only unless ``--apply`` is supplied. Before any write it checks the whole source snapshot for name and UUID collisions. Re-running is -idempotent: existing objects are patched by their preserved source UUIDs. +idempotent: existing objects are patched by their preserved source UUIDs, and a +record the source renamed is renamed. Every read fails closed: a portal that +errors mid-read aborts that portal's run rather than being mirrored from a +partial snapshot, which would move the unseen records' datasets to the root. +Records the source has deleted are reported (``withdrawn_at_source``), never +deleted. """ from __future__ import annotations import argparse import asyncio +import csv +import hashlib import json +import logging import os import re import sys @@ -24,13 +50,29 @@ from collections.abc import Iterable from pathlib import Path from typing import Any +from urllib.parse import urlparse + +import httpx _PROJECT_ROOT = Path(__file__).resolve().parent.parent sys.path.insert(0, str(_PROJECT_ROOT / "src")) -from data_concierge.data_layer.connectors.ckan import CKANClient # noqa: E402 +from data_concierge.data_layer.connectors.ckan import CKANActionError, CKANClient # noqa: E402 +from data_concierge.data_layer.connectors.dcat import ( # noqa: E402 + DCATClient, + _as_list, + _text, + parse_dataset, + raw_dataset_nodes, +) from data_concierge.data_layer.onboard_index import _scrub_secrets # noqa: E402 -from data_concierge.gateway.ckan_sites import get_site # noqa: E402 +from data_concierge.data_layer.qsv_profiling import _extract_qsv_tags # noqa: E402 +from data_concierge.gateway.ckan_sites import ( # noqa: E402 + PORTAL_TYPE_DCAT, + get_site, + list_sites, + normalize_portal_type, +) _SLUG_RE = re.compile(r"[^a-z0-9-]+") _PROVENANCE_KEYS = { @@ -67,29 +109,144 @@ "private", "plugin_data", ) -_RESOURCE_FIELDS = ( - "id", +_CLEARABLE_PACKAGE_FIELDS = ( + "author", + "author_email", + "maintainer", + "maintainer_email", + "license_id", "url", - "description", - "format", - "hash", - "name", - "resource_type", - "url_type", - "mimetype", - "mimetype_inner", - "cache_url", - "size", - "created", - "last_modified", - "cache_last_updated", + "version", +) +# Package keys CKAN computes or the mirror sets itself. Anything else a source +# returns — ckanext-scheming fields such as a data steward or temporal +# coverage — is source metadata the Fair Store's default schema would drop, so +# it is carried as an extra. +_PACKAGE_MANAGED = frozenset( + { + *_PACKAGE_FIELDS, + "owner_org", + "organization", + "groups", + "tags", + "tag_string", + "extras", + "resources", + "license_title", + "license_url", + "isopen", + "num_resources", + "num_tags", + "metadata_created", + "metadata_modified", + "creator_user_id", + "relationships_as_object", + "relationships_as_subject", + "revision_id", + "tracking_summary", + } ) +# Resource keys describing the source CKAN's own bookkeeping rather than the +# resource. ``datastore_*`` keys are excluded separately: the Fair Store holds +# metadata only and must never claim that a DataStore table was copied. +_RESOURCE_MANAGED = frozenset( + {"package_id", "position", "state", "metadata_modified", "revision_id", "tracking_summary"} +) +# url_type values for which CKAN rewrites ``url`` to the *serving* site's own +# download route. Mirrored as-is, the link would point at a file the Fair Store +# never received, so the source's absolute URL is kept as a plain link. +_LOCAL_URL_TYPES = frozenset({"upload", "datastore"}) +# Longest term Solr will index (bytes, UTF-8). CKAN indexes every dataset extra +# as one such term; resource extras are not indexed, so they have no limit. +_SOLR_MAX_TERM_BYTES = 32766 + +_TAG_VALID_RE = re.compile(r"[\w \-.]{2,100}") +_TAG_INVALID_RE = re.compile(r"[^\w \-.]+") +_EMAIL_RE = re.compile(r"[^@\s]+@[^@\s]+\.[^@\s]+") + +_SOCRATA_VIEW_RE = re.compile(r"/api/views/([a-z0-9]{4}-[a-z0-9]{4})/?$") +_SOCRATA_DISCOVERY_URL = "https://api.us.socrata.com/api/catalog/v1" +# Socrata domains name their owning-agency field themselves, so match the +# usual spellings; ``attribution`` is the standard fallback. +_SOCRATA_OWNER_KEY_RE = re.compile(r"business[-_ ]?owner|agency|department|publisher", re.I) + +_MEDIA_TYPE_FORMATS = { + "text/csv": "CSV", + "application/csv": "CSV", + "text/tab-separated-values": "TSV", + "application/json": "JSON", + "application/geo+json": "GeoJSON", + "application/vnd.geo+json": "GeoJSON", + "application/xml": "XML", + "text/xml": "XML", + "application/rdf+xml": "RDF", + "application/vnd.google-earth.kml+xml": "KML", + "application/vnd.google-earth.kmz": "KMZ", + "application/vnd.ms-excel": "XLS", + "application/vnd.openxmlformats-officedocument.spreadsheetml.sheet": "XLSX", + "application/zip": "ZIP", + "application/pdf": "PDF", + "text/html": "HTML", +} + +# DCAT publishes licenses as URLs; CKAN's license register uses short IDs. +# Keys are normalized by _license_key. Unknown URLs are kept only in extras. +_LICENSE_IDS = { + "usa.gov/government-works": "other-pd", + "usa.gov/publicdomain/label/1.0": "other-pd", + "creativecommons.org/publicdomain/zero/1.0": "cc-zero", + "creativecommons.org/publicdomain/mark/1.0": "other-pd", + "creativecommons.org/licenses/by/4.0": "cc-by", + "creativecommons.org/licenses/by/3.0": "cc-by", + "creativecommons.org/licenses/by-sa/4.0": "cc-by-sa", + "creativecommons.org/licenses/by-sa/3.0": "cc-by-sa", + "creativecommons.org/licenses/by-nc/4.0": "cc-nc", + "opendatacommons.org/licenses/pddl/1.0": "odc-pddl", + "opendatacommons.org/licenses/by/1.0": "odc-by", + "opendatacommons.org/licenses/odbl/1.0": "odc-odbl", + "gnu.org/licenses/fdl-1.3": "gfdl", +} class MirrorError(RuntimeError): """Raised when a collision or failed CKAN action makes a mirror unsafe.""" +class DeletedInTarget(MirrorError): + """The object exists in the Fair Store but an admin deleted it there.""" + + +_ATTEMPTS = 3 + + +async def _read( + client: CKANClient, + action: str, + params: dict[str, Any], + *, + label: str, + missing_ok: bool = False, +) -> Any: + """A read that fails closed; ``None`` for a missing object when ``missing_ok``. + + ``CKANClient.action`` returns ``{}`` for any failure, which a mirror reads + as "nothing there": one 502 while paging a portal's organizations would + leave the unseen ones out of the snapshot, and ``--apply`` would then move + their datasets to the portal root. Transient failures are retried; anything + else, or a transient failure that persists, aborts the run. + """ + for attempt in range(_ATTEMPTS): + try: + return await client.call(action, params) + except CKANActionError as exc: + if missing_ok and exc.not_found: + return None + if not exc.transient or attempt == _ATTEMPTS - 1: + raise MirrorError(f"{action} failed for {label}: {exc}") from exc + await asyncio.sleep(2**attempt) + raise AssertionError("unreachable") + + def _slug(text: str, fallback: str) -> str: """Return a CKAN-safe slug while retaining the legacy helper API.""" slug = _SLUG_RE.sub("-", (text or "").lower()).strip("-") @@ -106,6 +263,7 @@ def _load_index(site: str, index_path: str | Path | None = None) -> dict[str, An if index_path else [ _PROJECT_ROOT / "data" / "ckan_onboard" / site / "index.json", + _PROJECT_ROOT / "data" / "dcat_onboard" / site / "index.json", _PROJECT_ROOT / "ckan_onboard" / site / "index.json", ] ) @@ -148,32 +306,196 @@ def _column_dictionary(resource: dict[str, Any]) -> list[dict[str, Any]]: def _qsv_by_resource(index: dict[str, Any]) -> dict[str, dict[str, Any]]: return { - str(resource["resource_id"]): resource + str(resource["resource_id"]): _vet_profile(resource) for dataset in index.get("datasets", []) or [] for resource in dataset.get("resources", []) or [] if resource.get("resource_id") } +def _raise_csv_field_limit() -> None: + # Python's csv module refuses fields over 128 KB, which qsv emits for + # geometry columns. + csv.field_size_limit(max(csv.field_size_limit(), 64 * 1024 * 1024)) + + +def _profiled_header(qsv: dict[str, Any]) -> list[str] | None: + """Column names of the file qsv profiled, while it is still on disk.""" + local_path = str(qsv.get("local_path") or "") + if not local_path: + return None + path = Path(local_path) + path = path if path.is_absolute() else _PROJECT_ROOT / path + if not path.is_file(): + return None + _raise_csv_field_limit() + with path.open(newline="", encoding="utf-8-sig", errors="replace") as handle: + return [name for name in next(csv.reader(handle), []) if name] + + +def _qsv_fields(path: Path) -> set[str] | None: + """The columns a qsv stats or frequency CSV reports on.""" + if not path.is_file(): + return None + _raise_csv_field_limit() + with path.open(newline="", encoding="utf-8", errors="replace") as handle: + return {row.get("field") or "" for row in csv.DictReader(handle)} - {""} + + +def _vet_profile(qsv: dict[str, Any]) -> dict[str, Any]: + """A profile checked against the file it names, ready to publish. + + onboard_*.py writes ``qsv_stats.csv`` and ``qsv_frequency.csv`` once per + dataset directory, so when a dataset has several CSVs and one profile + fails, the directory can hold another file's output — two WPRDC datasets + showed another file's columns and row counts under the profiled file's + name. Output that does not match the profiled file's header is left out, + as are the statistics merged from it. Tags are re-read from describegpt's + own output, whose key is sometimes capitalised. + """ + if "_vetted" in qsv: + return qsv + profile = dict(qsv, _vetted=True) + header = _profiled_header(qsv) + directory = _qsv_dir(qsv) + trusted = str(qsv.get("status") or "") == "ok" + + def describes(name: str, *, exact: bool) -> bool: + fields = _qsv_fields(directory / name) if directory else None + if fields is None: + return False + if header is None: + return trusted + return fields == set(header) if exact else fields <= set(header) + + profile["_stats_ok"] = describes("qsv_stats.csv", exact=True) + profile["_frequency_ok"] = describes("qsv_frequency.csv", exact=False) + profile["_header"] = header + if header is None and not trusted: + # A failed profile whose file is gone: its merged statistics cannot be + # checked against the file, so only names, labels and types remain. + profile["columns"] = [ + {**column, "stats": {}, "top_values": []} for column in qsv.get("columns") or [] + ] + elif header is not None and not (profile["_stats_ok"] and profile["_frequency_ok"]): + # Keep the file's own columns and the portal's DataStore fields; drop + # columns only the mismatched output contributed. + profile["columns"] = [ + { + **column, + "stats": column.get("stats") if profile["_stats_ok"] else {}, + "top_values": column.get("top_values") if profile["_frequency_ok"] else [], + } + for column in qsv.get("columns") or [] + if column.get("name") in header or column.get("ckan_info") or column.get("ckan_type") + ] + + # describegpt's own dictionary names the columns it described; like the + # stats, it can belong to another file in the same directory. + describegpt = directory / "qsv_dict.json" if directory else None + try: + output = ( + json.loads(describegpt.read_text(encoding="utf-8")) + if describegpt is not None and describegpt.is_file() + else None + ) + except (ValueError, OSError): + output = None + described = { + str(field.get("name")) + for field in ((output or {}).get("Dictionary") or {}).get("response", {}).get("fields") + or [] + if isinstance(field, dict) and field.get("name") + } + if output is None: + profile["_describegpt_ok"] = False + elif header is None or not described: + profile["_describegpt_ok"] = trusted + else: + profile["_describegpt_ok"] = described - {"_id"} <= set(header) + if profile["_describegpt_ok"]: + tags = _extract_qsv_tags(output or {}) + if tags: + profile["qsv_tags"] = tags + elif output is not None: + # Another file's AI output: its summary, tags and column labels go too. + profile["qsv_description"] = "" + profile["qsv_tags"] = [] + profile["columns"] = [ + { + key: value + for key, value in column.items() + if key not in ("qsv_label", "qsv_description") + } + for column in profile.get("columns") or [] + ] + return profile + + def _project(source: dict[str, Any], fields: Iterable[str]) -> dict[str, Any]: return { field: source[field] for field in fields if field in source and source[field] is not None } +def _is_empty(value: Any) -> bool: + return value is None or value == "" or value == [] or value == {} + + +def _extra_value(value: Any) -> str: + """Dataset extras are strings in CKAN; structured values are kept as JSON.""" + if isinstance(value, str): + return value + return json.dumps(value, sort_keys=True, ensure_ascii=False) + + +def _source_id(row: dict[str, Any]) -> str: + """The identifier the source itself uses for a record. + + For a CKAN source that is the preserved UUID. A DCAT snapshot records the + catalog's own identifier under ``_source_id`` because its UUIDs are derived. + """ + return str(row.get("_source_id") or row.get("id", "")) + + +def _with_image(payload: dict[str, Any], row: dict[str, Any]) -> dict[str, Any]: + """Point an organization or group image at a URL that resolves on the Fair Store. + + An uploaded image is stored as a bare filename that CKAN resolves against + the serving site's own ``/uploads``, where the mirror never put the file, + so the source's absolute ``image_display_url`` is used instead. + """ + image = str(row.get("image_url") or "") + if image and not image.startswith(("http://", "https://")): + image = str(row.get("image_display_url") or "") + if image: + payload["image_url"] = image + else: + payload.pop("image_url", None) + return payload + + def _provenance_extras( extras: list[dict[str, Any]] | None, *, site_id: str, source_url: str, source_id: str, + carried: dict[str, str] | None = None, ) -> list[dict[str, str]]: - """Preserve source extras and add stable mirror identity fields.""" - values = { - str(item.get("key")): str(item.get("value", "")) - for item in (extras or []) - if item.get("key") - } + """Preserve source extras and add stable mirror identity fields. + + ``carried`` holds source fields with no column in the target schema; a + source extra of the same name wins, and the identity fields win over both. + """ + values = dict(carried or {}) + values.update( + { + str(item.get("key")): str(item.get("value", "")) + for item in (extras or []) + if item.get("key") + } + ) values.update( { _PROVENANCE_KEYS["portal"]: site_id, @@ -184,6 +506,14 @@ def _provenance_extras( return [{"key": key, "value": value} for key, value in values.items()] +def _mirrored_from(row: dict[str, Any]) -> str: + """Name the portal a source record was itself mirrored from, if any.""" + for item in row.get("extras") or []: + if isinstance(item, dict) and item.get("key") == _PROVENANCE_KEYS["portal"]: + return str(item.get("value") or "").strip() + return "" + + def _group_refs(groups: list[dict[str, Any]] | None) -> list[dict[str, str]]: refs: list[dict[str, str]] = [] for group in groups or []: @@ -196,31 +526,81 @@ def _group_refs(groups: list[dict[str, Any]] | None) -> list[dict[str, str]]: return refs -async def _all_packages(client: CKANClient) -> list[dict[str, Any]]: +def _clean_tag(raw: str) -> str: + """Return ``raw`` as a valid CKAN tag, or ``""`` when nothing usable is left. + + CKAN accepts 2-100 word characters, spaces, hyphens and dots. Valid tags pass + through untouched; DCAT keywords such as ``l&i`` are rewritten (``l-i``). + """ + tag = raw.strip() + if _TAG_VALID_RE.fullmatch(tag): + return tag + tag = re.sub(r"-{2,}", "-", _TAG_INVALID_RE.sub("-", tag)).strip(" -") + tag = tag[:100].strip() + return tag if len(tag) >= 2 else "" + + +def _tag_refs(tags: list[dict[str, Any]] | None) -> list[dict[str, str]]: + refs: list[dict[str, str]] = [] + seen: set[str] = set() + for tag in tags or []: + name = _clean_tag(str(tag.get("name") or "")) + if not name or name in seen: + continue + seen.add(name) + refs.append( + { + "name": name, + **({"vocabulary_id": tag["vocabulary_id"]} if tag.get("vocabulary_id") else {}), + } + ) + return refs + + +async def _all_packages( + client: CKANClient, *, label: str = "portal", include_private: bool = False +) -> list[dict[str, Any]]: + """Every dataset the portal lists; raises rather than return part of them.""" + params: dict[str, Any] = {"q": "*:*", "rows": 1000, "sort": "name asc"} + if include_private: + # The Fair Store's own private and draft datasets still hold their + # UUIDs: left out, they would look new and never be updated. + params.update(include_private=True, include_drafts=True) packages: list[dict[str, Any]] = [] + seen_ids: set[str] = set() start = 0 while True: - page = await client.action( - "package_search", {"q": "*:*", "rows": 1000, "start": start, "sort": "name asc"} - ) - batch = page.get("results", []) if isinstance(page, dict) else [] - packages.extend(batch) + page = await _read(client, "package_search", {**params, "start": start}, label=label) + if not isinstance(page, dict): + raise MirrorError(f"package_search for {label} returned {type(page).__name__}") + batch = page.get("results") or [] + count = int(page.get("count") or 0) + packages.extend(row for row in batch if str(row.get("id")) not in seen_ids) + seen_ids.update(str(row.get("id")) for row in batch) start += len(batch) - if not batch or start >= int(page.get("count", 0)): + if start >= count: return packages + if not batch: + raise MirrorError(f"package_search for {label} stopped at {start} of {count} datasets") -async def _all_group_rows(client: CKANClient, action_name: str) -> list[dict[str, Any]]: +async def _all_group_rows( + client: CKANClient, action_name: str, *, label: str = "portal", include_extras: bool = False +) -> list[dict[str, Any]]: """Read every organization or group despite portal-side page-size caps.""" + # Sorted by name, which is unique: CKAN's default (title) order is not a + # total order, so offset paging over it can skip a row. + params: dict[str, Any] = {"all_fields": True, "limit": 1000, "sort": "name asc"} + if include_extras: + params["include_extras"] = True rows: list[dict[str, Any]] = [] seen_ids: set[str] = set() offset = 0 while True: - result = await client.action( - action_name, - {"all_fields": True, "limit": 1000, "offset": offset}, - ) - batch = list(result) if isinstance(result, list) else [] + result = await _read(client, action_name, {**params, "offset": offset}, label=label) + if not isinstance(result, list): + raise MirrorError(f"{action_name} for {label} returned {type(result).__name__}") + batch = list(result) if not batch: return rows new_rows = [row for row in batch if str(row.get("id")) not in seen_ids] @@ -231,15 +611,64 @@ async def _all_group_rows(client: CKANClient, action_name: str) -> list[dict[str offset += len(batch) +def _empty_copies() -> dict[str, Any]: + return {"organizations": 0, "groups": 0, "datasets": 0, "origins": []} + + +async def _check_complete( + client: CKANClient, action_name: str, rows: list[dict[str, Any]], *, label: str +) -> None: + """Cross-check paged ``all_fields`` rows against the plain name listing. + + The name listing is one unpaged call: WPRDC's CDN answers 403 to a + names-only listing that carries limit/offset. CKAN returns up to 1,000 + names that way; a portal at that cap is not checked. + """ + listed = await _read(client, action_name, {"sort": "name asc"}, label=label) + if not isinstance(listed, list): + raise MirrorError(f"{action_name} for {label} returned {type(listed).__name__}") + names = {str(item.get("name") if isinstance(item, dict) else item) for item in listed} + if len(listed) >= 1000: + print(f"warning: {action_name} for {label}: too many to cross-check", file=sys.stderr) + return + missing = names - {str(row.get("name")) for row in rows} + if missing: + raise MirrorError( + f"{action_name} for {label} paged {len(rows)} of {len(names)} records; " + f"missing e.g. {sorted(missing)[:5]}" + ) + + async def _source_snapshot( - client: CKANClient, *, organization: str | None = None, limit: int | None = None -) -> dict[str, list[dict[str, Any]]]: - organization_rows = await _all_group_rows(client, "organization_list") + client: CKANClient, + *, + site_id: str | None = None, + organization: str | None = None, + limit: int | None = None, +) -> dict[str, Any]: + """Read a CKAN source, leaving out records it mirrored from another portal.""" + copies = _empty_copies() + origins: set[str] = set() + + def is_copy(row: dict[str, Any], kind: str) -> bool: + origin = _mirrored_from(row) + if site_id and origin and origin != site_id: + copies[kind] += 1 + origins.add(origin) + return True + return False + + label = f"source {site_id or client.ckan_url}" + organization_rows = await _all_group_rows(client, "organization_list", label=label) + await _check_complete(client, "organization_list", organization_rows, label=label) organizations: list[dict[str, Any]] = [] - for row in organization_rows if isinstance(organization_rows, list) else []: + for row in organization_rows: if organization and row.get("name") != organization: continue - full = await client.action( + # The list row lacks the extras (including any mirror provenance) and + # the parent groups, so it is never a stand-in for the full record. + full = await _read( + client, "organization_show", { "id": row["id"], @@ -248,18 +677,29 @@ async def _source_snapshot( "include_groups": True, "include_tags": True, }, + label=f"{label} organization {row.get('name')}", ) - organizations.append(full or row) - - groups = await _all_group_rows(client, "group_list") + if not is_copy(full, "organizations"): + organizations.append(full) + + # With extras, so a group the source itself mirrored is seen as a copy. + groups = [ + group + for group in await _all_group_rows(client, "group_list", label=label, include_extras=True) + if not is_copy(group, "groups") + ] - packages = await _all_packages(client) + packages = [ + package + for package in await _all_packages(client, label=label) + if not is_copy(package, "datasets") + ] if organization: org_ids = {str(item.get("id")) for item in organizations} packages = [ package for package in packages - if package.get("organization", {}).get("name") == organization + if (package.get("organization") or {}).get("name") == organization or str(package.get("owner_org")) in org_ids ] if limit is not None: @@ -271,28 +711,542 @@ async def _source_snapshot( if group.get("name") } groups = [group for group in groups if group.get("name") in used_group_names] - return {"organizations": organizations, "groups": groups, "packages": packages} + copies["origins"] = sorted(origins) + return { + "organizations": organizations, + "groups": groups, + "packages": packages, + "copies": copies, + } -async def _target_snapshot(client: CKANClient) -> dict[str, list[dict[str, Any]]]: - organizations = await _all_group_rows(client, "organization_list") - groups = await _all_group_rows(client, "group_list") - packages = await _all_packages(client) +def _dcat_uuid(identifier: str, source_url: str, kind: str) -> str: + """Stable UUID for a DCAT record: the source's own UUID when it has one.""" + try: + return str(uuid.UUID(identifier)) + except (ValueError, TypeError): + pass + if urlparse(identifier).scheme in ("http", "https"): + return str(uuid.uuid5(uuid.NAMESPACE_URL, identifier)) + return str(uuid.uuid5(uuid.NAMESPACE_URL, f"{source_url.rstrip('/')}#{kind}/{identifier}")) + + +def _socrata_id(identifier: str) -> str: + match = _SOCRATA_VIEW_RE.search(urlparse(identifier or "").path) + return match.group(1) if match else "" + + +def _dcat_dataset_name(title: str, identifier: str, dataset_id: str) -> str: + """A unique, stable CKAN name: the title slug plus the source's short ID. + + CKAN names are site-wide, and portal titles repeat ("Crashes", "Budget"), so + every name carries the portal's own ID (a Socrata 4x4) or a UUID prefix. + """ + tail = urlparse(identifier).path.rstrip("/").rsplit("/", 1)[-1] if identifier else "" + suffix = tail.lower() if re.fullmatch(r"[a-z0-9][a-z0-9_-]{1,19}", tail.lower()) else "" + suffix = suffix or dataset_id[:8] + base = _slug(title, "dataset")[: 100 - len(suffix) - 1].rstrip("-") + return f"{base}-{suffix}" + + +def _publisher_chain(publisher: Any) -> list[str]: + """Publisher names from most specific to broadest (DCAT-US subOrganizationOf).""" + chain: list[str] = [] + node = publisher + while node and len(chain) < 10: + name = _text(node) + if name and name not in chain: + chain.append(name) + node = node.get("subOrganizationOf") if isinstance(node, dict) else None + return chain + + +def _lineage(parents: dict[str, str], name: str) -> set[str]: + """``name`` and every ancestor recorded for it in ``parents``.""" + seen: set[str] = set() + while name and name not in seen: + seen.add(name) + name = parents.get(name, "") + return seen + + +def _mailto(value: Any) -> str: + email = _text(value).removeprefix("mailto:").strip() + return email if _EMAIL_RE.fullmatch(email) else "" + + +def _license_key(url: str) -> str: + parsed = urlparse(url.strip().lower()) + host = parsed.netloc.removeprefix("www.") + return f"{host}{parsed.path.rstrip('/')}" if host else url.strip().lower() + + +def _dcat_license_id(url: str) -> str: + return _LICENSE_IDS.get(_license_key(url), "") if url else "" + + +def _media_format(media_type: str, declared: str) -> str: + return declared or _MEDIA_TYPE_FORMATS.get(media_type.lower(), "") + + +def _dcat_extras(raw: dict[str, Any]) -> list[dict[str, str]]: + """Every DCAT dataset field as a ``dcat_``-prefixed extra. + + Fields copied into a CKAN core field (title, description, landing page, + version) are not duplicated; keyword, theme, license and contactPoint are + kept verbatim because their CKAN counterparts are normalized. + """ + skipped = {"distribution", "title", "description", "landingPage", "version"} + return [ + {"key": f"dcat_{key.lstrip('@')}", "value": _extra_value(_stable(key, value))} + for key, value in raw.items() + if key not in skipped and not _is_empty(value) + ] + + +# Catalog fields that are sets. Socrata serves them in a different order on +# each fetch, so they are sorted to keep identical content identical. +_UNORDERED_FIELDS = frozenset({"keyword", "theme", "domain_tags"}) + + +def _stable(key: str, value: Any) -> Any: + if key in _UNORDERED_FIELDS and isinstance(value, list): + return sorted(value, key=lambda item: json.dumps(item, sort_keys=True)) + return value + + +def _socrata_owner(row: dict[str, Any]) -> tuple[str, str]: + """The owning agency Socrata records for a dataset, and the field it came from.""" + for item in (row.get("classification") or {}).get("domain_metadata") or []: + key, value = str(item.get("key") or ""), str(item.get("value") or "").strip() + if value and _SOCRATA_OWNER_KEY_RE.search(key): + return value, key + attribution = str((row.get("resource") or {}).get("attribution") or "").strip() + return (attribution, "attribution") if attribution else ("", "") + + +def _socrata_extras(row: dict[str, Any]) -> list[dict[str, str]]: + """Discovery API metadata the DCAT export omits, as ``socrata_`` extras. + + Volatile analytics (page views, downloads) and account names are left out. + """ + resource = row.get("resource") or {} + classification = row.get("classification") or {} + values: dict[str, Any] = { + "socrata_id": resource.get("id"), + "socrata_type": resource.get("type"), + "socrata_attribution": resource.get("attribution"), + "socrata_attribution_link": resource.get("attribution_link"), + "socrata_provenance": resource.get("provenance"), + "socrata_created_at": resource.get("createdAt"), + "socrata_updated_at": resource.get("updatedAt"), + "socrata_data_updated_at": resource.get("data_updated_at"), + "socrata_metadata_updated_at": resource.get("metadata_updated_at"), + "socrata_publication_date": resource.get("publication_date"), + "socrata_category": classification.get("domain_category"), + "socrata_domain_tags": _stable("domain_tags", classification.get("domain_tags")), + "socrata_license": (row.get("metadata") or {}).get("license"), + "socrata_permalink": row.get("permalink"), + } + for item in classification.get("domain_metadata") or []: + if item.get("key"): + values[f"socrata_{item['key']}"] = item.get("value") + return [ + {"key": key, "value": _extra_value(value)} + for key, value in values.items() + if not _is_empty(value) + ] + + +def _socrata_columns(row: dict[str, Any]) -> list[dict[str, str]]: + """The column list Socrata publishes for a dataset, as a data dictionary.""" + resource = row.get("resource") or {} + fields = resource.get("columns_field_name") or [] + + def at(key: str, index: int) -> str: + values = resource.get(key) or [] + return str(values[index] or "") if index < len(values) else "" + + # The column arrays agree with one another but not from fetch to fetch, + # so the dictionary is ordered by field name. + return sorted( + ( + { + "name": str(field), + "label": at("columns_name", index), + "type": at("columns_datatype", index), + "description": _scrub_secrets(at("columns_description", index)), + } + for index, field in enumerate(fields) + ), + key=lambda column: column["name"], + ) + + +def _dcat_resources(raw: dict[str, Any], *, dataset_id: str) -> list[dict[str, Any]]: + """CKAN resources for a dataset's DCAT distributions.""" + identifier = _text(raw.get("identifier") or raw.get("@id")) + resources: list[dict[str, Any]] = [] + used_keys: set[str] = set() + distributions = [d for d in raw.get("distribution") or [] if isinstance(d, dict)] + for index, dist in enumerate(distributions): + download_url = _text(dist.get("downloadURL")) + url = download_url or _text(dist.get("accessURL")) + media_type = _text(dist.get("mediaType")) + fmt = _media_format(media_type, _text(dist.get("format"))) + key = url or str(index) + if key in used_keys: + key = f"{key}#{index}" + used_keys.add(key) + resource: dict[str, Any] = { + "id": str(uuid.uuid5(uuid.NAMESPACE_URL, f"{dataset_id}/{key}")), + "_source_id": url or f"{identifier}#distribution-{index}", + "url": url, + "name": _text(dist.get("title")) or fmt or media_type or f"Distribution {index + 1}", + "description": _text(dist.get("description")), + "format": fmt, + "mimetype": media_type or None, + } + for field_name, value in dist.items(): + if field_name in {"title", "description", "mediaType", "format", "downloadURL"}: + continue + if field_name == "accessURL" and not download_url: + continue + if not _is_empty(value): + resource[f"dcat_{field_name.lstrip('@')}"] = value + resources.append(resource) + return resources + + +def _dcat_snapshot( + catalog: Any, + *, + site_title: str, + source_url: str, + socrata: dict[str, dict[str, Any]] | None = None, + limit: int | None = None, +) -> dict[str, Any]: + """Translate a DCAT catalog into the CKAN-shaped snapshot the mirror writes. + + Organizations come from each dataset's publisher chain, most specific first, + headed by the Socrata owning agency when the Discovery API names one. A + publisher that is the portal itself is represented by the portal's root + organization rather than duplicated beneath it. + """ + nodes = raw_dataset_nodes(catalog) + if limit is not None: + nodes = nodes[:limit] + host = (urlparse(source_url).hostname or "").lower() + portal_names = {host, host.removeprefix("www."), site_title.strip().lower()} - {""} + + org_parent: dict[str, str] = {} + org_field: dict[str, str] = {} + themes: dict[str, str] = {} + packages: list[dict[str, Any]] = [] + names: set[str] = set() + + def org_id(title: str) -> str: + return _dcat_uuid(title, source_url, "publisher") + + for raw in nodes: + identifier = _text(raw.get("identifier") or raw.get("@id")) + title = _text(raw.get("title")) + dataset_id = _dcat_uuid(identifier or title, source_url, "dataset") + row = (socrata or {}).get(_socrata_id(identifier)) + + chain = _publisher_chain(raw.get("publisher")) + fields = ["dcat:publisher"] * len(chain) + if row: + owner, owner_key = _socrata_owner(row) + if owner and owner not in chain: + chain.insert(0, owner) + fields.insert(0, f"socrata:{owner_key}") + kept = [(n, f) for n, f in zip(chain, fields, strict=True) if n.lower() not in portal_names] + for (child, _), (parent, _) in zip(kept, kept[1:], strict=False): + # A parent learned from any dataset fills in one not yet known, but + # a catalog that contradicts itself must not create a loop, which + # ckanext-hierarchy rejects mid-run. + if not org_parent.get(child) and child not in _lineage(org_parent, parent): + org_parent[child] = parent + for name_, field_ in kept: + org_parent.setdefault(name_, "") + org_field.setdefault(name_, field_) + + theme_slugs: list[str] = [] + for theme in _as_list(raw.get("theme")): + slug = _slug(theme, _dcat_uuid(theme, source_url, "theme")[:8]) + themes.setdefault(slug, theme) + if slug not in theme_slugs: + theme_slugs.append(slug) + + name = _dcat_dataset_name(title or identifier, identifier, dataset_id) + if name in names: + name = f"{name[:91]}-{dataset_id[:8]}" + names.add(name) + + contact = raw.get("contactPoint") if isinstance(raw.get("contactPoint"), dict) else {} + extras = _dcat_extras(raw) + (_socrata_extras(row) if row else []) + parsed = parse_dataset(raw) + tabular = parsed.tabular_distributions + first_tabular_url = tabular[0].best_url if tabular else "" + resources = _dcat_resources(raw, dataset_id=dataset_id) + columns = _socrata_columns(row) if row else [] + if columns: + # The column list goes on the file it describes, beside any qsv + # profile. As a dataset extra it would also be indexed as one Solr + # term, and a wide table's list overruns the term limit. + dictionary = { + "source_data_dictionary": json.dumps(columns), + "source_column_count": str(len(columns)), + } + described = next( + (r for r in resources if first_tabular_url and r["url"] == first_tabular_url), + resources[0] if resources else None, + ) + if described is not None: + described.update(dictionary) + else: + extras += [{"key": key, "value": value} for key, value in dictionary.items()] + + packages.append( + { + "id": dataset_id, + "_source_id": identifier or title, + # How onboard_dcat.py files this dataset's qsv profile. + "_short_id": parsed.id, + "_first_tabular_url": first_tabular_url, + "name": name, + "title": title or identifier, + "notes": _text(raw.get("description")), + "url": _text(raw.get("landingPage")) or None, + "version": _text(raw.get("version")) or None, + "maintainer": _text(contact.get("fn")) or None, + "maintainer_email": _mailto(contact.get("hasEmail")) or None, + "license_id": _dcat_license_id(_text(raw.get("license"))) or None, + "owner_org": org_id(kept[0][0]) if kept else "", + "tags": [ + {"name": keyword} for keyword in sorted(_as_list(raw.get("keyword")), key=str) + ], + "groups": [{"name": slug} for slug in theme_slugs], + "extras": extras, + "resources": resources, + } + ) + + # Names are assigned after every publisher is known so they do not depend on + # catalog order; two titles that slug alike are told apart by UUID prefix. + org_names: dict[str, str] = {} + for title in sorted(org_parent): + slug = _slug(title, org_id(title)[:8]) + org_names[title] = ( + slug if slug not in org_names.values() else f"{slug[:91]}-{org_id(title)[:8]}" + ) + organizations = [ + { + "id": org_id(title), + "_source_id": title, + "name": org_names[title], + "title": title, + "groups": ( + [{"name": org_names[org_parent[title]], "capacity": "parent"}] + if org_parent[title] + else [] + ), + "extras": [{"key": "publisher_source", "value": org_field[title]}], + } + for title in sorted(org_parent) + ] + groups = [ + { + "id": _dcat_uuid(theme, source_url, "theme"), + "_source_id": theme, + "name": slug, + "title": theme, + } + for slug, theme in sorted(themes.items()) + ] return { - "organizations": list(organizations) if isinstance(organizations, list) else [], - "groups": list(groups) if isinstance(groups, list) else [], + "organizations": organizations, + "groups": groups, "packages": packages, + "copies": _empty_copies(), + } + + +def _dcat_qsv_index(index: dict[str, Any], snapshot: dict[str, Any]) -> dict[str, Any]: + """Re-key a DCAT onboarding index to the mirrored resource UUIDs. + + ``onboard_dcat.py`` profiles each dataset's first tabular distribution and + files it under the dataset's short ID rather than a resource ID. + """ + by_short_id: dict[str, str] = {} + for package in snapshot["packages"]: + for resource in package.get("resources", []): + if ( + package.get("_first_tabular_url") + and resource.get("url") == package["_first_tabular_url"] + ): + by_short_id[str(package.get("_short_id"))] = str(resource["id"]) + break + rekeyed = [ + {**resource, "resource_id": by_short_id[str(resource.get("resource_id"))]} + for dataset in index.get("datasets", []) or [] + for resource in dataset.get("resources", []) or [] + if str(resource.get("resource_id")) in by_short_id + ] + return {"datasets": [{"resources": rekeyed}]} if rekeyed else {} + + +async def _socrata_page(client: httpx.AsyncClient, domain: str, offset: int) -> dict[str, Any]: + """One Discovery API page, retried when the failure is transient.""" + params = {"domains": domain, "search_context": domain, "limit": 1000, "offset": offset} + for attempt in range(_ATTEMPTS): + try: + resp = await client.get(_SOCRATA_DISCOVERY_URL, params=params) + if resp.status_code != 429 and resp.status_code < 500: + resp.raise_for_status() + body = resp.json() + if isinstance(body, dict): + return body + raise ValueError(f"unexpected {type(body).__name__} body") + error: Exception = httpx.HTTPStatusError( + f"HTTP {resp.status_code}", request=resp.request, response=resp + ) + except httpx.TransportError as exc: + error = exc + if attempt == _ATTEMPTS - 1: + raise error + await asyncio.sleep(2**attempt) + raise AssertionError("unreachable") + + +async def _socrata_metadata(catalog: Any, *, required: bool) -> dict[str, dict[str, Any]]: + """Discovery API records for a Socrata-hosted catalog, keyed by 4x4 ID. + + Returns ``{}`` for a catalog that is not Socrata's. The owning agency comes + only from this API — data.pa.gov's own catalog names one publisher for every + dataset — so when ``required`` (a run that writes) a failure raises instead + of mirroring a snapshot that would move every dataset to the portal root and + drop its ``socrata_*`` extras. Otherwise it warns and returns what it has. + """ + domains = { + (urlparse(_text(node.get("identifier"))).hostname or "") + for node in raw_dataset_nodes(catalog) + if _socrata_id(_text(node.get("identifier"))) + } - {""} + rows: dict[str, dict[str, Any]] = {} + async with httpx.AsyncClient(timeout=60, follow_redirects=True) as client: + for domain in sorted(domains): + offset = 0 + while True: + try: + body = await _socrata_page(client, domain, offset) + except (httpx.HTTPError, ValueError) as exc: + message = f"Socrata enrichment failed for {domain} at offset {offset}: {exc}" + if required: + raise MirrorError( + f"{message}. Retry later; --no-enrich mirrors without it, which " + "moves every dataset to the portal root organization." + ) from exc + print(f"warning: {message}", file=sys.stderr) + break + batch = body.get("results") or [] + for result in batch: + view_id = (result.get("resource") or {}).get("id") + if view_id: + rows[str(view_id)] = result + offset += len(batch) + total = int(body.get("resultSetSize") or 0) + if offset >= total: + break + if not batch: + message = f"Socrata enrichment for {domain} stopped at {offset} of {total}" + if required: + raise MirrorError(message) + print(f"warning: {message}", file=sys.stderr) + break + return rows + + +async def _target_snapshot(client: CKANClient) -> dict[str, list[dict[str, Any]]]: + label = f"target {client.ckan_url}" + # Extras carry the provenance that tells a rename of this portal's own + # record apart from a collision with someone else's. + return { + "organizations": await _all_group_rows( + client, "organization_list", label=label, include_extras=True + ), + "groups": await _all_group_rows(client, "group_list", label=label, include_extras=True), + "packages": await _all_packages(client, label=label, include_private=True), } +def _pin_target_names( + source: dict[str, Any], target: dict[str, list[dict[str, Any]]], *, site_id: str +) -> None: + """Keep the names already published for a DCAT portal's records. + + DCAT names are derived — from a title slug, and for organizations whose + titles slug alike, from sort order — so an edited title or a newly added + publisher would otherwise rename published datasets and break their URLs. + A record this portal already holds keeps its name; a new record whose name + is taken gets its UUID prefix appended. + """ + for kind in ("organizations", "packages"): + held = { + str(row.get("id")): str(row.get("name")) + for row in target[kind] + if _mirrored_from(row) == site_id and row.get("name") + } + rows = source[kind] + renamed: dict[str, str] = {} + for row in rows: + name = held.get(str(row.get("id"))) + if name and name != row.get("name"): + renamed[str(row.get("name"))] = name + row["name"] = name + # A new record may not take any name the Fair Store already uses, nor + # one given to an earlier record of this run. + taken = {str(row.get("name")) for row in target[kind] if row.get("name")} + taken |= {str(row.get("name")) for row in rows if str(row.get("id")) in held} + for row in rows: + row_id, name = str(row.get("id")), str(row.get("name")) + if row_id in held: + continue + if name in taken: + new_name = f"{name[:91]}-{row_id[:8]}" + renamed[name] = new_name + row["name"] = name = new_name + taken.add(name) + if kind == "organizations" and renamed: + # Parent references are by name. + for row in rows: + for parent in row.get("groups") or []: + if parent.get("name") in renamed: + parent["name"] = renamed[parent["name"]] + + def _assert_no_collisions( - source: dict[str, list[dict[str, Any]]], + source: dict[str, Any], target: dict[str, list[dict[str, Any]]], *, store_name: str, store_id: str, -) -> None: + site_id: str = "", + evictions: list[tuple[str, str, str]] | None = None, +) -> list[str]: + """Raise on anything a write would clobber; return the renames it will apply. + + The same UUID under a new name is a rename when the Fair Store's record is + this portal's own (its provenance names ``site_id``): the source renamed + it, and the patch carries the new name. Under any other provenance it is a + collision, as is a name held by a different UUID — unless the holder is this + portal's own record that its source no longer publishes and ``evictions`` + is given: the source reused the name, so the withdrawn record is to be + renamed out of the way, collected as ``(kind, id, new name)``. + """ collisions: list[str] = [] + renames: list[str] = [] def check( kind: str, @@ -303,12 +1257,37 @@ def check( ) -> None: by_name = {str(row.get("name")): row for row in existing if row.get("name")} by_id = {str(row.get("id")): row for row in existing if row.get("id")} + incoming_ids = {str(row.get("id", "")) for row in incoming} + seen_ids: set[str] = set() + seen_names: set[str] = set() for row in incoming: name, row_id = str(row.get("name", "")), str(row.get("id", "")) - if not allow_shared_name and name in by_name and str(by_name[name].get("id")) != row_id: - collisions.append(f"{kind} name {name!r} already has a different UUID") - if row_id in by_id and str(by_id[row_id].get("name")) != name: - collisions.append(f"{kind} UUID {row_id!r} already has a different name") + if row_id in seen_ids: + collisions.append(f"{kind} UUID {row_id!r} appears twice in the source") + if name in seen_names and not allow_shared_name: + collisions.append(f"{kind} name {name!r} appears twice in the source") + seen_ids.add(row_id) + seen_names.add(name) + holder = by_name.get(name) + if not allow_shared_name and holder is not None and str(holder.get("id")) != row_id: + holder_id = str(holder.get("id")) + if ( + evictions is not None + and site_id + and _mirrored_from(holder) == site_id + and holder_id not in incoming_ids + ): + evicted = f"{name[:80]}-withdrawn-{holder_id[:8]}" + evictions.append((kind, holder_id, evicted)) + renames.append(f"{kind} {name!r} (withdrawn at source) -> {evicted!r}") + else: + collisions.append(f"{kind} name {name!r} already has a different UUID") + held = by_id.get(row_id) + if held is not None and str(held.get("name")) != name: + if site_id and _mirrored_from(held) == site_id: + renames.append(f"{kind} {held.get('name')!r} -> {name!r}") + else: + collisions.append(f"{kind} UUID {row_id!r} already has a different name") root_existing = next( (row for row in target["organizations"] if row.get("name") == store_name), None @@ -352,35 +1331,643 @@ def check( if collisions: detail = "\n - ".join(collisions[:50]) raise MirrorError(f"Mirror preflight found {len(collisions)} collision(s):\n - {detail}") + return renames + + +_SHOW_ACTIONS = { + "organization_create": "organization_show", + "group_create": "group_show", + "package_create": "package_show", + "resource_create": "resource_show", +} async def _write_action( - client: CKANClient, action: str, payload: dict[str, Any], *, label: str + client: CKANClient, + action: str, + payload: dict[str, Any], + *, + label: str, + upload: tuple[str, bytes, str] | None = None, ) -> dict[str, Any]: - show_actions = { - "organization_create": "organization_show", - "group_create": "group_show", - "package_create": "package_show", - "resource_create": "resource_show", - } - for attempt in range(3): - result = await client.action(action, payload) - if result: - return result - - # A gateway can time out after CKAN commits the object. Resolve by the - # preserved UUID before retrying so a successful write is not treated - # as a failed duplicate create. - show_action = show_actions.get(action) + """Run a write, retrying transient failures; ``upload`` is (name, bytes, type). + + A create whose object turns out to exist is resolved by its preserved UUID: + a gateway can fail after CKAN commits, and the retry must not be reported + as a failure (or repeated). Writes without a UUID to check — DataStore + loads, view creation — are made safe by their callers instead. + """ + show_action = _SHOW_ACTIONS.get(action) + for attempt in range(_ATTEMPTS): + try: + if upload: + filename, content, content_type = upload + result = await client.upload_call( + action, payload, filename=filename, content=content, content_type=content_type + ) + else: + result = await client.call(action, payload) + return result if isinstance(result, dict) else {"result": result} + except CKANActionError as exc: + error = exc + if show_action and payload.get("id"): - existing = await client.action(show_action, {"id": payload["id"]}) + existing = await _read( + client, show_action, {"id": payload["id"]}, label=label, missing_ok=True + ) if existing: + if existing.get("state", "active") == "deleted": + raise DeletedInTarget(f"{label} was deleted in the Fair Store") + # Package, organization and group names are identifiers too; + # a resource name is only a label. + if ( + action != "resource_create" + and "name" in payload + and existing.get("name") != payload["name"] + ): + raise MirrorError( + f"{action} failed for {label}: UUID {payload['id']!r} is held by " + f"{existing.get('name')!r}: {error}" + ) return existing - if attempt < 2: - await asyncio.sleep(2**attempt) + if not error.transient or attempt == _ATTEMPTS - 1: + raise MirrorError(f"{action} failed for {label}: {error}") from error + await asyncio.sleep(2**attempt) + raise AssertionError("unreachable") + + +# -- qsv profiles as dataset metadata and resources --------------------------- +# +# onboard_ckan.py / onboard_dcat.py profile a dataset's first CSV with qsv: +# stats, frequency, and describegpt (AI descriptions, labels, tags). The mirror +# publishes that profile twice over: compact fields on the dataset, and the full +# outputs as resources on it. Tables go into the DataStore, so CKAN shows them +# as sortable tables and serves them as CSV/JSON downloads; describegpt's JSON +# is uploaded as a file. The profiled source data itself is never copied. + +_QSV_MODEL_RE = re.compile(r"^Model:\s*(\S+)", re.MULTILINE) +_QSV_VERSION_RE = re.compile(r"Generated by qsv v([\w.\-]+)") +# describegpt ends a description with a provenance block — "Generated by qsv +# vX describegpt", its command line, model, API URL, timestamp and an LLM +# warning — introduced however the model chose: "Attribution: Generated by…", +# an "## Attribution" heading, a bare line, a rule, or "@attribution". +_QSV_PROVENANCE_RE = re.compile( + r"^[^\n]*\bGenerated by qsv\b|^[ \t]*Command line:[ \t]*qsv describegpt\b", + re.IGNORECASE | re.MULTILINE, +) +_QSV_TRAILING_RE = re.compile( + r"(?:\n[ \t]*(?:(?:#{1,6}|\*\*||@)?[ \t]*attribution[ \t]*:?[ \t]*(?:\*\*|)?" + r"|-{3,}|\*{3,}|_{3,})?[ \t]*)+$", + re.IGNORECASE, +) +_DATASTORE_BATCH = 1000 +_MAX_SUMMARY_CELL = 10_000 + + +def _qsv_dir(qsv: dict[str, Any]) -> Path | None: + """The onboarding directory holding a profile's qsv output files.""" + local_path = str(qsv.get("local_path") or "") + if not local_path: + return None + path = Path(local_path) + directory = (path if path.is_absolute() else _PROJECT_ROOT / path).parent + return directory if directory.is_dir() else None + + +def _qsv_resource_id(profiled_resource_id: str, kind: str) -> str: + """Stable UUID for one qsv output of one profiled resource.""" + return str(uuid.uuid5(uuid.NAMESPACE_URL, f"qsv:{profiled_resource_id}:{kind}")) + + +def _profile_row_count(qsv: dict[str, Any]) -> int | None: + """The profiled file's row count, when it is one. + + A byte-capped DCAT download counts only the rows it kept, and a profile + whose download failed reports 0; neither is the file's size. + """ + count = qsv.get("row_count") + if isinstance(count, bool) or not isinstance(count, int) or qsv.get("truncated_download"): + return None + if count == 0 and str(qsv.get("status") or "") != "ok": + return None + return count + + +def _describegpt_prose(description: str) -> str: + """A describegpt description without its trailing provenance block.""" + description = description or "" + match = _QSV_PROVENANCE_RE.search(description) + prose = description[: match.start()] if match else description + return _QSV_TRAILING_RE.sub("", "\n" + prose).strip() + + +def _qsv_dataset_extras(qsv: dict[str, Any]) -> dict[str, str]: + """The profile's dataset-level summary, shown with the dataset's metadata. + + Column names are included so a search for a column finds its dataset. The + description keeps only describegpt's prose: the model and qsv version are + fields of their own, and the full output, provenance included, is the + describegpt resource. + """ + qsv = _vet_profile(qsv) + description = _scrub_secrets(str(qsv.get("qsv_description") or "")) + # The profiled file's own columns: the dictionary also lists the portal's + # DataStore fields, which a download can lack. + columns = qsv.get("_header") or [ + str(c["name"]) for c in qsv.get("columns") or [] if c.get("name") + ] + model = _QSV_MODEL_RE.search(description) + version = _QSV_VERSION_RE.search(description) + values = { + "qsv_description": _describegpt_prose(description), + "qsv_tags": json.dumps(qsv.get("qsv_tags") or []) if qsv.get("qsv_tags") else "", + "qsv_row_count": str( + _profile_row_count(qsv) if _profile_row_count(qsv) is not None else "" + ), + # Statistics from a byte-capped download describe the file's first rows. + "qsv_profile_truncated": "true" if qsv.get("truncated_download") else "", + "qsv_column_count": str(len(columns)) if columns else "", + # JSON, like qsv_tags: column names can contain commas. + "qsv_columns": json.dumps(columns, ensure_ascii=False) if columns else "", + "qsv_profiled_resource_id": str(qsv.get("resource_id") or ""), + "qsv_profiled_at": str(qsv.get("onboarded_at") or ""), + "qsv_status": str(qsv.get("status") or ""), + "qsv_model": model.group(1) if model else "", + "qsv_version": version.group(1) if version else "", + } + return {key: value for key, value in values.items() if value} + + +def _as_number(value: Any) -> float | int | None: + """A DataStore numeric cell: qsv writes blanks and text where numbers are absent.""" + try: + number = float(str(value).strip()) + except (TypeError, ValueError): + return None + return int(number) if number.is_integer() else number + + +def _top_values_text(top_values: list[dict[str, Any]] | None) -> str: + return "; ".join( + f"{item.get('value')} ({item.get('count')})" for item in (top_values or [])[:10] + ) + + +def _dictionary_table(qsv: dict[str, Any]) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]: + """One row per column: qsv type, AI label and description, key statistics.""" + fields = [ + { + "id": "column", + "type": "text", + "info": {"label": "Column", "notes": "Column name in the file"}, + }, + {"id": "label", "type": "text", "info": {"label": "Label", "notes": "AI-generated label"}}, + {"id": "type", "type": "text", "info": {"label": "Type", "notes": "Type inferred by qsv"}}, + { + "id": "description", + "type": "text", + "info": {"label": "Description", "notes": "AI-generated description"}, + }, + {"id": "null_count", "type": "numeric", "info": {"label": "Null count"}}, + {"id": "cardinality", "type": "numeric", "info": {"label": "Distinct values"}}, + {"id": "min", "type": "text", "info": {"label": "Minimum"}}, + {"id": "max", "type": "text", "info": {"label": "Maximum"}}, + {"id": "mean", "type": "numeric", "info": {"label": "Mean"}}, + {"id": "median", "type": "numeric", "info": {"label": "Median"}}, + {"id": "stddev", "type": "numeric", "info": {"label": "Standard deviation"}}, + {"id": "top_values", "type": "text", "info": {"label": "Most frequent values (count)"}}, + ] + records = [] + for column in _column_dictionary(qsv): + stats = column["stats"] + records.append( + { + "column": column["name"], + "label": column["label"], + "type": column["type"], + "description": column["description"], + "null_count": _as_number(stats.get("nullcount")), + "cardinality": _as_number(stats.get("cardinality")), + "min": str(stats.get("min") or ""), + "max": str(stats.get("max") or ""), + "mean": _as_number(stats.get("mean")), + "median": _as_number(stats.get("q2_median")), + "stddev": _as_number(stats.get("stddev")), + "top_values": _top_values_text(column["top_values"]), + } + ) + return fields, records + + +def _summary_cell(value: str) -> str: + """A text cell for a summary table, capped at :data:`_MAX_SUMMARY_CELL` chars. + + qsv reports a column's min/max/mode verbatim, so a geometry column yields + whole WKT polygons (one 311 mode is 157 KB) that would swamp the table view. + """ + value = _scrub_secrets(value) + if len(value) <= _MAX_SUMMARY_CELL: + return value + return f"{value[:_MAX_SUMMARY_CELL]}… [truncated: {len(value):,} characters]" + + +def _csv_table( + path: Path, numeric: frozenset[str] = frozenset() +) -> tuple[list[dict[str, Any]], list[dict[str, Any]]] | None: + """A qsv CSV as DataStore fields and records; ``numeric`` columns are typed.""" + if not path.is_file(): + return None + # Geometry cells are read whole, then capped. + _raise_csv_field_limit() + with path.open(newline="", encoding="utf-8", errors="replace") as handle: + reader = csv.DictReader(handle) + header = [name for name in reader.fieldnames or [] if name] + records = [ + { + name: ( + _as_number(row.get(name)) + if name in numeric + else _summary_cell(row.get(name) or "") + ) + for name in header + } + for row in reader + ] + fields = [{"id": name, "type": "numeric" if name in numeric else "text"} for name in header] + return fields, records + + +def _qsv_artifacts(qsv: dict[str, Any]) -> list[dict[str, Any]]: + """The resources to publish for one profile, in display order.""" + qsv = _vet_profile(qsv) + directory = _qsv_dir(qsv) + source = str(qsv.get("resource_name") or "the profiled file") + caveat = "Generated with qsv; AI-written text may contain inaccuracies." + artifacts: list[dict[str, Any]] = [] + + fields, records = _dictionary_table(qsv) + if records: + artifacts.append( + { + "kind": "dictionary", + "name": "Data dictionary (qsv)", + "description": ( + f"Column-level data dictionary for “{source}”: qsv-inferred types, " + f"AI-generated labels and descriptions (qsv describegpt), and key " + f"statistics. {caveat}" + ), + "table": (fields, records), + } + ) + if directory is None: + return artifacts + + stats = _csv_table(directory / "qsv_stats.csv") if qsv["_stats_ok"] else None + if stats: + artifacts.append( + { + "kind": "stats", + "name": "Summary statistics (qsv stats)", + "description": ( + f"qsv stats for “{source}”: one row per column with its type, " + "min/max, mean, quartiles, cardinality, null count and more." + ), + "table": stats, + } + ) + frequency = ( + _csv_table( + directory / "qsv_frequency.csv", numeric=frozenset({"count", "percentage", "rank"}) + ) + if qsv["_frequency_ok"] + else None + ) + if frequency: + artifacts.append( + { + "kind": "frequency", + "name": "Frequency distribution (qsv frequency)", + "description": ( + f"qsv frequency for “{source}”: the most common values in each " + "column, with counts, percentages and rank." + ), + "table": frequency, + } + ) + describegpt = directory / "qsv_dict.json" + if qsv["_describegpt_ok"] and describegpt.is_file(): + artifacts.append( + { + "kind": "describegpt", + "name": "AI descriptions (qsv describegpt)", + "description": ( + f"Raw qsv describegpt output for “{source}”: the AI-generated " + f"dataset description, tags, and column dictionary. {caveat}" + ), + "file": ( + "describegpt.json", + _scrub_secrets(describegpt.read_text(encoding="utf-8")).encode("utf-8"), + "application/json", + ), + } + ) + return artifacts + + +_QSV_KINDS = ("dictionary", "stats", "frequency", "describegpt") + + +async def _publish_qsv( + target: CKANClient, + *, + package_id: str, + qsv: dict[str, Any], + site_id: str, + existing_resources: dict[str, dict[str, Any]] | set[str], + apply: bool = True, +) -> dict[str, int]: + """Create, refresh or remove one profile's resources; counts by outcome. + + Each resource records a digest of what it was built from (``qsv_digest``), + and one whose digest matches — with its table still loaded — is left alone, + so a rerun rewrites only the profiles that changed. A kind the profile no + longer produces is removed: these resources are the Fair Store's own + derivation, and one left behind would be stale or, for output that turned + out to describe another file, wrong. ``apply=False`` only counts. + """ + held = ( + existing_resources + if isinstance(existing_resources, dict) + else {resource_id: {} for resource_id in existing_resources} + ) + counts = {"created": 0, "refreshed": 0, "unchanged": 0, "removed": 0} + profiled_id = str(qsv.get("resource_id") or "") + produced: set[str] = set() + for artifact in _qsv_artifacts(qsv): + produced.add(artifact["kind"]) + resource_id = _qsv_resource_id(profiled_id, artifact["kind"]) + resource: dict[str, Any] = { + "id": resource_id, + "package_id": package_id, + "name": artifact["name"], + "description": artifact["description"], + "qsv_artifact": artifact["kind"], + "qsv_profiled_resource_id": profiled_id, + "qsv_profiled_at": str(qsv.get("onboarded_at") or ""), + "qsv_profile_site": site_id, + } + if "file" in artifact: + resource["format"] = "JSON" + content = hashlib.sha256(artifact["file"][1]).hexdigest() + resource["qsv_digest"] = _digest({"resource": resource, "content": content}) + else: + fields, records = artifact["table"] + # What datastore_create does for an inline resource, done as a + # create of its own so a lost response is resolved by its UUID. + resource.update(format="CSV", url="_datastore_only_resource", url_type="datastore") + resource["qsv_digest"] = _digest( + {"resource": resource, "fields": fields, "records": records} + ) + existing = held.get(resource_id) + loaded = existing is not None and ( + existing.get("url_type") == "upload" + if "file" in artifact + else bool(existing.get("datastore_active")) + ) + if loaded and existing.get("qsv_digest") == resource["qsv_digest"]: + counts["unchanged"] += 1 + if apply and "table" in artifact: + await _ensure_table_view(target, resource_id, label=resource["name"]) + continue + counts["created" if existing is None else "refreshed"] += 1 + if not apply: + continue + + label = f"{artifact['kind']} for {profiled_id}" + action = "resource_create" if existing is None else "resource_patch" + if "file" in artifact: + # The file and its digest land in one call. + await _write_action(target, action, resource, label=label, upload=artifact["file"]) + continue + # The digest is recorded only once the rows are in and counted: a run + # that stops mid-load leaves a table that the next run must reload, + # not one it would take for unchanged. + digest = resource["qsv_digest"] + await _write_action(target, action, {**resource, "qsv_digest": ""}, label=label) + await _load_table(target, resource_id, fields, records, label=label) + await _write_action( + target, "resource_patch", {"id": resource_id, "qsv_digest": digest}, label=label + ) + await _ensure_table_view(target, resource_id, label=label) + + # Without the onboarding directory nothing shows the output is gone (the + # mirror may just be running from a checkout that lacks it), so nothing + # is removed. + removable = _qsv_dir(qsv) is not None + for kind in _QSV_KINDS: + resource_id = _qsv_resource_id(profiled_id, kind) + if not removable or kind in produced or resource_id not in held: + continue + counts["removed"] += 1 + if apply: + await _delete( + target, "resource_delete", {"id": resource_id}, label=f"{kind} for {profiled_id}" + ) + return counts + + +async def _delete(target: CKANClient, action: str, params: dict[str, Any], *, label: str) -> None: + """A delete; an object that is already gone is fine.""" + for attempt in range(_ATTEMPTS): + try: + await target.call(action, params) + return + except CKANActionError as exc: + if exc.not_found: + return + if not exc.transient or attempt == _ATTEMPTS - 1: + raise MirrorError(f"{action} failed for {label}: {exc}") from exc + await asyncio.sleep(2**attempt) + + +async def _drop_table(target: CKANClient, resource_id: str, *, label: str) -> None: + """Delete a resource's DataStore table; one that does not exist is fine.""" + await _delete( + target, "datastore_delete", {"resource_id": resource_id, "force": True}, label=label + ) + + +async def _load_table( + target: CKANClient, + resource_id: str, + fields: list[dict[str, Any]], + records: list[dict[str, Any]], + *, + label: str, +) -> None: + """Replace a resource's DataStore table with ``records``, verified by count. + + The table is rebuilt rather than appended to, so a rerun cannot duplicate + rows. An insert whose response is lost may still have committed, and its + retry then adds the batch twice, so the loaded row count is checked and the + table rebuilt until it matches. + """ + total: Any = None + for _ in range(_ATTEMPTS): + await _drop_table(target, resource_id, label=label) + await _write_action( + target, + "datastore_create", + { + "resource_id": resource_id, + "force": True, + "fields": fields, + "records": records[:_DATASTORE_BATCH], + }, + label=label, + ) + for start in range(_DATASTORE_BATCH, len(records), _DATASTORE_BATCH): + await _write_action( + target, + "datastore_upsert", + { + "resource_id": resource_id, + "force": True, + "method": "insert", + "records": records[start : start + _DATASTORE_BATCH], + }, + label=label, + ) + loaded = await _read( + target, "datastore_search", {"resource_id": resource_id, "limit": 0}, label=label + ) + total = loaded.get("total") if isinstance(loaded, dict) else None + if total == len(records): + return + print( + f"warning: {label} holds {total} rows, expected {len(records)}; reloading", + file=sys.stderr, + ) + raise MirrorError(f"{label}: DataStore holds {total} rows, expected {len(records)}") + + +async def _ensure_table_view(target: CKANClient, resource_id: str, *, label: str) -> None: + """Give a DataStore resource its table view. + + CKAN picks a resource's default views when the resource is created, before + its table exists, so the table view never qualifies and has to be added + once the rows are in. An existing one is left alone so reruns do not stack + duplicates — which is why a failed listing or create is re-listed, never + read as "no view yet". + """ + for attempt in range(_ATTEMPTS): + views = await _read( + target, "resource_view_list", {"id": resource_id}, label=f"views of {label}" + ) + if any( + isinstance(view, dict) and view.get("view_type") == "datatables_view" + for view in views or [] + ): + return + try: + await target.call( + "resource_view_create", + {"resource_id": resource_id, "view_type": "datatables_view", "title": "Table"}, + ) + return + except CKANActionError as exc: + if not exc.transient or attempt == _ATTEMPTS - 1: + raise MirrorError(f"table view failed for {label}: {exc}") from exc + await asyncio.sleep(2**attempt) + + +# The digest of what the mirror last wrote to a record: an unchanged record is +# skipped on the next run instead of being patched (and reindexed) again. +_DIGEST_KEY = "mirror_digest" + + +def _digest(value: Any) -> str: + encoded = json.dumps(value, sort_keys=True, ensure_ascii=False, default=str) + return hashlib.sha256(encoded.encode("utf-8")).hexdigest()[:20] + + +def _extra(row: dict[str, Any], key: str) -> str: + for item in row.get("extras") or []: + if isinstance(item, dict) and item.get("key") == key: + return str(item.get("value") or "") + return "" - raise MirrorError(f"{action} failed for {label}") + +def _stamp_package(payload: dict[str, Any]) -> str: + """Add the package's digest as an extra; extras are compared as a set.""" + extras = sorted(payload.get("extras") or [], key=lambda item: item["key"]) + digest = _digest({**payload, "extras": extras}) + payload["extras"] = [*payload.get("extras", []), {"key": _DIGEST_KEY, "value": digest}] + return digest + + +def _stamp_resource(payload: dict[str, Any]) -> str: + digest = _digest(payload) + payload[_DIGEST_KEY] = digest + return digest + + +def _package_payload( + package: dict[str, Any], + *, + site_id: str, + source_url: str, + store_id: str, + known_org_ids: set[str], + qsv: dict[str, Any] | None = None, +) -> dict[str, Any]: + payload = _project(package, _PACKAGE_FIELDS) + # A patch leaves out what it does not name, so a value the source cleared + # would survive in the Fair Store; name these as empty instead. + for field in _CLEARABLE_PACKAGE_FIELDS: + payload.setdefault(field, "") + owner_org = str(package.get("owner_org") or "") + payload["owner_org"] = owner_org if owner_org in known_org_ids else store_id + payload["tags"] = _tag_refs(package.get("tags")) + payload["groups"] = _group_refs(package.get("groups")) + carried = { + key: _extra_value(value) + for key, value in package.items() + if key not in _PACKAGE_MANAGED and not key.startswith("_") and not _is_empty(value) + } + if qsv: + carried.update(_qsv_dataset_extras(qsv)) + for key in ("metadata_created", "metadata_modified"): + if package.get(key): + carried[f"mirror_source_{key}"] = str(package[key]) + if not owner_org: + carried["mirror_source_owner_org"] = "" + extras = _provenance_extras( + package.get("extras"), + site_id=site_id, + source_url=source_url, + source_id=_source_id(package), + carried=carried, + ) + # CKAN also indexes each dataset extra as a single untokenized Solr term; + # one longer than Solr's limit fails the whole dataset with an HTTP 500. + # Such a value cannot be stored on the dataset, so it is named instead. + oversized = [ + item["key"] for item in extras if len(item["value"].encode("utf-8")) > _SOLR_MAX_TERM_BYTES + ] + if oversized: + print( + f"warning: {package.get('name')}: extras over Solr's term limit not stored: " + f"{', '.join(oversized)}", + file=sys.stderr, + ) + extras = [item for item in extras if item["key"] not in oversized] + extras.append({"key": "mirror_oversized_extras", "value": json.dumps(oversized)}) + payload["extras"] = extras + if not str(payload.get("notes") or "").strip(): + payload["notes"] = "No description was provided by the source catalog." + return payload def _resource_payload( @@ -391,38 +1978,83 @@ def _resource_payload( source_url: str, qsv: dict[str, Any] | None, ) -> dict[str, Any]: - payload = _project(resource, _RESOURCE_FIELDS) + """Every source resource field, minus the source CKAN's own bookkeeping.""" + payload = { + key: value + for key, value in resource.items() + if value is not None + and key not in _RESOURCE_MANAGED + and not key.startswith(("_", "datastore_")) + } payload["package_id"] = package_id + url_type = str(resource.get("url_type") or "") + if url_type in _LOCAL_URL_TYPES: + payload.pop("url_type", None) + payload["mirror_source_url_type"] = url_type + if resource.get("datastore_active"): + # Recorded so a consumer knows the source portal can be queried; the + # Fair Store itself has no DataStore table for this resource. + payload["mirror_source_datastore_active"] = True payload.update( { _PROVENANCE_KEYS["portal"]: site_id, _PROVENANCE_KEYS["url"]: source_url, - _PROVENANCE_KEYS["id"]: str(resource.get("id", "")), + _PROVENANCE_KEYS["id"]: _source_id(resource), } ) - # Preserve extension-owned spatial metadata when the destination has the - # same extension, but never claim that a DataStore table was copied. - for key, value in resource.items(): - if key.startswith("dataspatial_") and value is not None: - payload[key] = value if qsv: + qsv = _vet_profile(qsv) columns = _column_dictionary(qsv) payload.update( { - "qsv_description": _scrub_secrets(qsv.get("qsv_description", "")), + "qsv_description": _describegpt_prose( + _scrub_secrets(str(qsv.get("qsv_description") or "")) + ), "qsv_tags": json.dumps(qsv.get("qsv_tags") or []), "data_dictionary": json.dumps(columns), - "column_count": len(columns), - "row_count": qsv.get("row_count", 0), + "column_count": len(qsv.get("_header") or columns), + "row_count": _profile_row_count(qsv), "qsv_onboarded_at": qsv.get("onboarded_at", ""), } ) return payload +def _withdrawn( + source: dict[str, Any], target: dict[str, list[dict[str, Any]]], *, site_id: str +) -> dict[str, Any]: + """This portal's mirrored records that its source no longer publishes. + + Reported, never deleted: whether a record withdrawn upstream should leave + the Fair Store is the operator's call. qsv resources are the Fair Store's + own and are not counted. + """ + source_packages = {str(package.get("id")) for package in source["packages"]} + source_resources = { + str(resource.get("id")) + for package in source["packages"] + for resource in package.get("resources") or [] + } + ours = [package for package in target["packages"] if _mirrored_from(package) == site_id] + datasets = sorted( + str(package.get("name")) + for package in ours + if str(package.get("id")) not in source_packages + ) + resources = [ + str(resource.get("id")) + for package in ours + if str(package.get("id")) in source_packages + for resource in package.get("resources") or [] + if resource.get(_PROVENANCE_KEYS["portal"]) == site_id + and str(resource.get("id")) not in source_resources + ] + return {"datasets": len(datasets), "dataset_names": datasets[:50], "resources": len(resources)} + + async def mirror_catalog( *, - source: CKANClient, + source: CKANClient | None, target: CKANClient, site_id: str, site_title: str, @@ -431,15 +2063,40 @@ async def mirror_catalog( organization: str | None = None, qsv_index: dict[str, Any] | None = None, limit: int | None = None, + snapshot: dict[str, Any] | None = None, + pin_names: bool = False, ) -> dict[str, Any]: - """Plan or apply one complete, API-level CKAN metadata mirror.""" - source_data = await _source_snapshot(source, organization=organization, limit=limit) + """Plan or apply one complete, API-level metadata mirror. + + ``source`` is read as a CKAN portal unless a prepared ``snapshot`` (from + :func:`_dcat_snapshot`) is passed instead. ``pin_names`` keeps the names + already published for the portal's records (see :func:`_pin_target_names`); + without it a source rename is applied. + """ + if snapshot is not None: + source_data = snapshot + elif source is not None: + source_data = await _source_snapshot( + source, site_id=site_id, organization=organization, limit=limit + ) + else: + raise MirrorError("mirror_catalog needs a source client or a snapshot") target_data = await _target_snapshot(target) + if pin_names: + _pin_target_names(source_data, target_data, site_id=site_id) store_name = _slug(site_id, "source-store") if store_name in {str(row.get("name")) for row in source_data["organizations"]}: store_name = _slug(f"{store_name}-source", "source-store") store_id = str(uuid.uuid5(uuid.NAMESPACE_URL, source_url.rstrip("/"))) - _assert_no_collisions(source_data, target_data, store_name=store_name, store_id=store_id) + evictions: list[tuple[str, str, str]] = [] + renames = _assert_no_collisions( + source_data, + target_data, + store_name=store_name, + store_id=store_id, + site_id=site_id, + evictions=evictions, + ) qsv_resources = _qsv_by_resource(qsv_index or {}) resources = [ @@ -459,22 +2116,86 @@ async def mirror_catalog( "qsv_resources": sum( 1 for resource in resources if str(resource.get("id")) in qsv_resources ), + "skipped_copies": source_data.get("copies", _empty_copies()), + "renamed": renames, + # Planned in a dry run, done when applied. "created": 0, "updated": 0, + "unchanged": 0, + "qsv": {"created": 0, "refreshed": 0, "unchanged": 0, "removed": 0}, } - if not apply: - return summary + if organization is None and limit is None: + summary["withdrawn_at_source"] = _withdrawn(source_data, target_data, site_id=site_id) + + held_qsv = [ + resource + for package in target_data["packages"] + for resource in package.get("resources", []) or [] + if resource.get("qsv_profile_site") == site_id + ] + if held_qsv and not qsv_resources: + # Every profiled dataset would lose its qsv fields to the patch. + message = ( + f"the Fair Store holds {len(held_qsv)} qsv resources for {site_id}, but no qsv " + "profiles were loaded: is its onboarding index missing (see --index-path)?" + ) + if apply: + raise MirrorError(message) + print(f"warning: {message}", file=sys.stderr) + # Profiles no longer in the index; their resources are reported, not removed. + summary["qsv"]["orphaned_profiles"] = len( + {str(resource.get("qsv_profiled_resource_id")) for resource in held_qsv} + - set(qsv_resources) + ) - target_org_ids = {str(row.get("id")) for row in target_data["organizations"]} target_group_ids = {str(row.get("id")) for row in target_data["groups"]} target_group_names = {str(row.get("name")) for row in target_data["groups"]} - target_package_ids = {str(row.get("id")) for row in target_data["packages"]} - target_resource_ids = { - str(resource.get("id")) + target_packages = {str(row.get("id")): row for row in target_data["packages"]} + target_resources = { + str(resource.get("id")): resource for package in target_data["packages"] for resource in package.get("resources", []) or [] } + async def write(action: str, payload: dict[str, Any], *, label: str) -> None: + if apply: + await _write_action(target, action, payload, label=label) + + target_orgs = {str(row.get("id")): row for row in target_data["organizations"]} + deleted_orgs: set[str] = set() + + async def upsert_group( + kind: str, payload: dict[str, Any], held: dict[str, Any] | None, digest: str + ) -> bool: + """Create or patch an organization or group; False when unchanged or skipped.""" + if held is not None and _extra(held, _DIGEST_KEY) == digest: + summary["unchanged"] += 1 + return False + payload["extras"] = [*payload["extras"], {"key": _DIGEST_KEY, "value": digest}] + try: + await write( + f"{kind}_{'create' if held is None else 'patch'}", payload, label=payload["name"] + ) + except DeletedInTarget: + # An admin removed it from the Fair Store; that decision stands. + summary.setdefault("skipped_deleted_in_target", []).append(payload["name"]) + if kind == "organization": + deleted_orgs.add(str(payload["id"])) + return False + summary["created" if held is None else "updated"] += 1 + return True + + def group_digest(payload: dict[str, Any], parents: Any = None) -> str: + extras = sorted(payload.get("extras") or [], key=lambda item: item["key"]) + return _digest({**payload, "extras": extras, "parents": parents}) + + for kind, holder_id, evicted in evictions: + await write( + "organization_patch" if kind == "organization" else "package_patch", + {"id": holder_id, "name": evicted}, + label=f"{kind} {holder_id} (withdrawn at source)", + ) + store_payload = { "id": store_id, "name": store_name, @@ -484,126 +2205,236 @@ async def mirror_catalog( [], site_id=site_id, source_url=source_url, source_id=store_id ), } - store_exists = store_id in target_org_ids - await _write_action( - target, - "organization_patch" if store_exists else "organization_create", - store_payload, - label=store_name, + await upsert_group( + "organization", store_payload, target_orgs.get(store_id), group_digest(store_payload) ) - summary["updated" if store_exists else "created"] += 1 + # CKAN's no_loops validator resolves both child and parent rows; combining a + # caller-supplied child UUID with groups during create trips that validator, + # so the hierarchy is applied only after every source organization exists. + # A parent that is not being mirrored cannot be referenced, so an + # organization whose parents are all outside the snapshot hangs off the root + # — except in an --organization run, which leaves an existing one's parent. source_org_names = {str(row.get("name")) for row in source_data["organizations"]} + hierarchy: list[tuple[dict[str, Any], list[dict[str, str]]]] = [] for organization_row in source_data["organizations"]: - payload = _project(organization_row, _ORGANIZATION_FIELDS) - source_id = str(organization_row.get("id", "")) + org_id = str(organization_row.get("id", "")) + payload = _with_image(_project(organization_row, _ORGANIZATION_FIELDS), organization_row) payload["extras"] = _provenance_extras( organization_row.get("extras"), site_id=site_id, source_url=source_url, - source_id=source_id, + source_id=_source_id(organization_row), ) - exists = source_id in target_org_ids - await _write_action( - target, - "organization_patch" if exists else "organization_create", - payload, - label=str(organization_row.get("name")), - ) - summary["updated" if exists else "created"] += 1 - - # Apply hierarchy only after every source organization exists. CKAN's - # no_loops validator resolves both child and parent rows; combining a - # caller-supplied child UUID with groups during create trips that validator. - for organization_row in source_data["organizations"]: - parents = _group_refs(organization_row.get("groups")) - if not any(parent["name"] in source_org_names for parent in parents): - parents.append({"name": store_name, "capacity": "parent"}) - await _write_action( - target, + held = target_orgs.get(org_id) + parents: list[dict[str, str]] | None = [ + parent + for parent in _group_refs(organization_row.get("groups")) + if parent["name"] in source_org_names + ] or [{"name": store_name, "capacity": "parent"}] + if organization is not None and held is not None: + parents = None + if await upsert_group("organization", payload, held, group_digest(payload, parents)): + if parents is not None: + hierarchy.append((organization_row, parents)) + + deleted_names = { + str(row.get("name")) + for row in source_data["organizations"] + if str(row.get("id")) in deleted_orgs + } + for organization_row, parents in hierarchy: + parents = [parent for parent in parents if parent["name"] not in deleted_names] or [ + {"name": store_name, "capacity": "parent"} + ] + await write( "organization_patch", {"id": str(organization_row.get("id", "")), "groups": parents}, label=f"{organization_row.get('name')} hierarchy", ) + target_groups = {str(row.get("id")): row for row in target_data["groups"]} for group in source_data["groups"]: - payload = _project(group, _GROUP_FIELDS) - source_id = str(group.get("id", "")) - if str(group.get("name")) in target_group_names and source_id not in target_group_ids: + payload = _with_image(_project(group, _GROUP_FIELDS), group) + group_id = str(group.get("id", "")) + if str(group.get("name")) in target_group_names and group_id not in target_group_ids: continue payload["extras"] = _provenance_extras( - group.get("extras"), site_id=site_id, source_url=source_url, source_id=source_id + group.get("extras"), site_id=site_id, source_url=source_url, source_id=_source_id(group) ) - exists = source_id in target_group_ids - await _write_action( - target, - "group_patch" if exists else "group_create", - payload, - label=str(group.get("name")), - ) - summary["updated" if exists else "created"] += 1 - - known_org_ids = {str(row.get("id")) for row in source_data["organizations"]} - for package in source_data["packages"]: - payload = _project(package, _PACKAGE_FIELDS) - package_id = str(package.get("id", "")) - owner_org = str(package.get("owner_org") or "") - payload["owner_org"] = owner_org if owner_org in known_org_ids else store_id - payload["tags"] = [ - { - "name": str(tag["name"]), - **({"vocabulary_id": tag["vocabulary_id"]} if tag.get("vocabulary_id") else {}), - } - for tag in package.get("tags", []) or [] - if tag.get("name") + await upsert_group("group", payload, target_groups.get(group_id), group_digest(payload)) + + # Datasets of an organization an admin deleted here hang off the root. + known_org_ids = {str(row.get("id")) for row in source_data["organizations"]} - deleted_orgs + total = len(source_data["packages"]) + for position, package in enumerate(source_data["packages"], start=1): + profiles = [ + qsv_resources[str(resource.get("id"))] + for resource in package.get("resources", []) or [] + if str(resource.get("id")) in qsv_resources ] - payload["groups"] = _group_refs(package.get("groups")) - payload["extras"] = _provenance_extras( - package.get("extras"), + payload = _package_payload( + package, site_id=site_id, source_url=source_url, - source_id=package_id, + store_id=store_id, + known_org_ids=known_org_ids, + qsv=profiles[0] if profiles else None, ) - if not str(payload.get("notes") or "").strip(): - payload["notes"] = "No description was provided by the source catalog." - if not owner_org: - payload["extras"].append({"key": "mirror_source_owner_org", "value": ""}) - exists = package_id in target_package_ids - await _write_action( - target, - "package_patch" if exists else "package_create", - payload, - label=str(package.get("name")), - ) - summary["updated" if exists else "created"] += 1 - - for resource in package.get("resources", []) or []: - resource_id = str(resource.get("id", "")) - resource_payload = _resource_payload( + digest = _stamp_package(payload) + package_id = str(payload.get("id", "")) + resource_payloads = [ + _resource_payload( resource, package_id=package_id, site_id=site_id, source_url=source_url, - qsv=qsv_resources.get(resource_id), + qsv=qsv_resources.get(str(resource.get("id", ""))), ) - resource_exists = resource_id in target_resource_ids - await _write_action( + for resource in package.get("resources", []) or [] + ] + for resource_payload in resource_payloads: + _stamp_resource(resource_payload) + if apply and (position % 50 == 0 or position == total): + print(f"[{site_id}] {position}/{total} datasets", file=sys.stderr, flush=True) + + existing = target_packages.get(package_id) + if existing is None: + # One call creates the dataset and its resources with their + # preserved UUIDs; a separate create per resource costs a + # reindex each and made a portal-sized mirror several times slower. + try: + await write( + "package_create", + {**payload, "resources": resource_payloads}, + label=str(package.get("name")), + ) + except DeletedInTarget: + # An admin removed it from the Fair Store; that decision stands. + summary.setdefault("skipped_deleted_in_target", []).append(package.get("name")) + continue + summary["created"] += 1 + len(resource_payloads) + else: + if _extra(existing, _DIGEST_KEY) == digest: + summary["unchanged"] += 1 + else: + await write("package_patch", payload, label=str(package.get("name"))) + summary["updated"] += 1 + for resource_payload in resource_payloads: + held = target_resources.get(resource_payload["id"]) + if held is not None and held.get(_DIGEST_KEY) == resource_payload[_DIGEST_KEY]: + summary["unchanged"] += 1 + continue + # Replaced rather than patched: a field the source dropped + # must go here too (a patch would keep it). + await write( + "resource_create" if held is None else "resource_update", + resource_payload, + label=str(resource_payload["id"]), + ) + summary["created" if held is None else "updated"] += 1 + + for profile in profiles: + counts = await _publish_qsv( target, - "resource_patch" if resource_exists else "resource_create", - resource_payload, - label=resource_id, + package_id=package_id, + qsv=profile, + site_id=site_id, + existing_resources=target_resources, + apply=apply, ) - summary["updated" if resource_exists else "created"] += 1 + for outcome, count in counts.items(): + summary["qsv"][outcome] += count + qsv_counts = summary["qsv"] + summary["qsv_artifacts"] = ( + qsv_counts["created"] + qsv_counts["refreshed"] + qsv_counts["unchanged"] + ) return summary -async def _main(args: argparse.Namespace) -> dict[str, Any]: - site = get_site(args.site) - if not site: - raise MirrorError(f"Unknown site {args.site!r}; add it to ckan_sites.json first") - if site.get("portal_type", "ckan") != "ckan": - raise MirrorError("Exact organization mirroring currently requires a CKAN source portal") +async def _mirror_site( + site_id: str, site: dict[str, Any], args: argparse.Namespace, api_key: str +) -> dict[str, Any]: + # Neither client may fall back to CKAN_API_KEY: the target is written with + # the mirror's own token, and a source portal must never receive a key. + target = CKANClient(ckan_url=args.target_url, api_key=api_key or None, use_default_key=False) + try: + common = { + "target": target, + "site_id": site_id, + "site_title": site["name"], + "source_url": site["url"], + "apply": args.apply, + } + if normalize_portal_type(site.get("portal_type")) == PORTAL_TYPE_DCAT: + if args.organization: + raise MirrorError("--organization applies to CKAN source portals only") + dcat = DCATClient(site["url"], site.get("catalog_url")) + try: + catalog = await dcat.fetch_raw_catalog() + except (httpx.HTTPError, ValueError) as exc: + raise MirrorError(f"DCAT catalog for {site_id} could not be read: {exc}") from exc + finally: + await dcat.close() + socrata = ( + {} if args.no_enrich else await _socrata_metadata(catalog, required=args.apply) + ) + snapshot = _dcat_snapshot( + catalog, + site_title=site["name"], + source_url=site["url"], + socrata=socrata, + limit=args.limit, + ) + summary = await mirror_catalog( + source=None, + snapshot=snapshot, + pin_names=True, + qsv_index=_dcat_qsv_index(_load_index(site_id, args.index_path), snapshot), + **common, + ) + summary["socrata_enriched"] = sum( + 1 + for package in snapshot["packages"] + if any(extra["key"] == "socrata_id" for extra in package["extras"]) + ) + return summary + + source = CKANClient(ckan_url=site["url"], use_default_key=False) + try: + return await mirror_catalog( + source=source, + organization=args.organization, + qsv_index=_load_index(site_id, args.index_path), + limit=args.limit, + **common, + ) + finally: + await source.close() + finally: + await target.close() + + +def _same_portal(url: str, other: str) -> bool: + def key(value: str) -> tuple[str, str]: + parsed = urlparse(value.strip()) + host = (parsed.hostname or "").lower().removeprefix("www.") + return host, parsed.path.rstrip("/") + + return key(url) == key(other) + + +async def _main(args: argparse.Namespace) -> dict[str, Any] | list[dict[str, Any]]: + if args.site == "all": + if args.organization or args.index_path: + raise MirrorError("--organization and --index-path need a single --site") + sites = [(str(site["id"]), site) for site in list_sites()] + else: + site = get_site(args.site) + if not site: + raise MirrorError(f"Unknown site {args.site!r}; add it to ckan_sites.json first") + sites = [(args.site, site)] api_key = os.environ.get(args.api_key_env, "") if args.api_key_file: @@ -613,41 +2444,62 @@ async def _main(args: argparse.Namespace) -> dict[str, Any]: f"Set {args.api_key_env} or --api-key-file when using --apply; a sysadmin token is required" ) - source = CKANClient(ckan_url=site["url"]) - target = CKANClient(ckan_url=args.target_url, api_key=api_key or None) - try: - return await mirror_catalog( - source=source, - target=target, - site_id=args.site, - site_title=site["name"], - source_url=site["url"], - apply=args.apply, - organization=args.organization, - qsv_index=_load_index(args.site, args.index_path), - limit=args.limit, - ) - finally: - await source.close() - await target.close() + if args.site != "all" and _same_portal(sites[0][1]["url"], args.target_url): + raise MirrorError(f"{args.site} is the target portal; it cannot be mirrored into itself") + + summaries: list[dict[str, Any]] = [] + for site_id, site in sites: + if _same_portal(site["url"], args.target_url): + # data.dathere.com is both a registered portal and a Fair Store. + summaries.append({"site": site_id, "skipped": "this portal is the target"}) + continue + try: + summaries.append(await _mirror_site(site_id, site, args, api_key)) + except Exception as exc: + if args.site != "all": + raise + # One portal's outage must not stop the others from refreshing; + # the failure is reported and the exit status is non-zero. + print(f"error: {site_id}: {type(exc).__name__}: {exc}", file=sys.stderr) + summaries.append({"site": site_id, "error": f"{type(exc).__name__}: {exc}"}) + return summaries if args.site == "all" else summaries[0] def main() -> None: - parser = argparse.ArgumentParser(description="Mirror CKAN metadata and qsv results") - parser.add_argument("--site", default="wprdc", help="source site ID from ckan_sites.json") - parser.add_argument("--target-url", default=os.environ.get("CKAN_URL", "http://localhost:5001")) + parser = argparse.ArgumentParser(description="Mirror CKAN/DCAT metadata and qsv results") + parser.add_argument( + "--site", + default="wprdc", + help="source site ID from ckan_sites.json, or 'all' for every registered portal", + ) + # Deliberately not CKAN_URL/CKAN_API_KEY: those are the app's portal + # settings, and in the app container CKAN_URL is data.dathere.com. + parser.add_argument( + "--target-url", default=os.environ.get("FAIRSTORE_URL", "http://localhost:5001") + ) parser.add_argument("--organization", help="optionally mirror only one source organization") parser.add_argument("--index-path", help="override the qsv index.json path") parser.add_argument("--limit", type=int, help="limit datasets for a smoke test") parser.add_argument("--apply", action="store_true", help="write after collision preflight") - parser.add_argument("--api-key-env", default="CKAN_API_KEY") + parser.add_argument( + "--no-enrich", + action="store_true", + help="skip Socrata Discovery API enrichment (owning agency, column list) for DCAT portals", + ) + parser.add_argument("--api-key-env", default="FAIRSTORE_API_KEY") parser.add_argument("--api-key-file", help="read the target sysadmin token from a file") args = parser.parse_args() + # Development logging is DEBUG on stdout, where this command prints its + # JSON summary; a line per HTTP request would bury it. + for noisy in ("asyncio", "httpx", "httpcore"): + logging.getLogger(noisy).setLevel(logging.WARNING) try: summary = asyncio.run(_main(args)) except MirrorError as exc: parser.exit(2, f"error: {exc}\n") print(json.dumps(summary, indent=2, sort_keys=True)) + if isinstance(summary, list) and any("error" in item for item in summary): + sys.exit(1) if __name__ == "__main__": diff --git a/src/data_concierge/data_layer/connectors/ckan.py b/src/data_concierge/data_layer/connectors/ckan.py index eb4cbc4..3d5562d 100644 --- a/src/data_concierge/data_layer/connectors/ckan.py +++ b/src/data_concierge/data_layer/connectors/ckan.py @@ -4,6 +4,7 @@ search over pre-indexed CKAN resources. """ +import json from typing import Any import httpx @@ -16,6 +17,61 @@ logger = get_logger(__name__) +class CKANActionError(Exception): + """A CKAN action that did not succeed. + + ``status_code`` is ``None`` when no response arrived at all (connection + refused, reset, timed out). + """ + + def __init__( + self, + action: str, + message: str, + *, + status_code: int | None = None, + error_type: str = "", + ) -> None: + super().__init__(f"{action}: {message}") + self.action = action + self.status_code = status_code + self.error_type = error_type + + @property + def not_found(self) -> bool: + return self.status_code == 404 or self.error_type == "Not Found Error" + + @property + def transient(self) -> bool: + """Worth retrying: no response, a server or gateway error, or rate limiting.""" + return ( + self.status_code is None + or self.status_code in {408, 425, 429} + or self.status_code >= 500 + ) + + +def _action_result(action_name: str, response: httpx.Response) -> Any: + """The ``result`` of a CKAN action response, or :class:`CKANActionError`.""" + try: + body = response.json() + except ValueError: + body = None + if response.is_success and isinstance(body, dict) and body.get("success"): + return body.get("result") + error = body.get("error") if isinstance(body, dict) else None + error_type = str(error.get("__type", "")) if isinstance(error, dict) else "" + if isinstance(error, dict) and error.get("message"): + message = str(error["message"]) + elif error: + message = json.dumps(error, default=str)[:500] + else: + message = f"HTTP {response.status_code}" + raise CKANActionError( + action_name, message, status_code=response.status_code, error_type=error_type + ) + + class CKANClient: """Client for interacting with CKAN data portals. @@ -33,17 +89,22 @@ def __init__( self, ckan_url: str | None = None, api_key: str | None = None, + *, + use_default_key: bool = True, ) -> None: """Initialize the CKAN client. Args: ckan_url: Base URL for CKAN instance. Defaults to settings. api_key: Optional CKAN API key for authenticated access. + use_default_key: Fall back to ``CKAN_API_KEY`` when ``api_key`` is + not given. Pass ``False`` for a portal that key does not + belong to, so it is never sent to a third party. """ self.ckan_url = (ckan_url or settings.ckan_url).rstrip("/") self._api_key = api_key or ( settings.ckan_api_key.get_secret_value() - if hasattr(settings, 'ckan_api_key') and settings.ckan_api_key + if use_default_key and hasattr(settings, 'ckan_api_key') and settings.ckan_api_key else None ) self._client: httpx.AsyncClient | None = None @@ -71,6 +132,19 @@ async def close(self) -> None: await self._client.aclose() self._client = None + async def call(self, action_name: str, params: dict[str, Any] | None = None) -> Any: + """Execute a CKAN API action, raising :class:`CKANActionError` on failure. + + Unlike :meth:`action`, a failure is never mistaken for an empty result, + which is what a caller that writes based on what it read needs. + """ + client = await self._get_client() + try: + response = await client.post(f"/api/3/action/{action_name}", json=params or {}) + except httpx.HTTPError as e: + raise CKANActionError(action_name, str(e) or type(e).__name__) from e + return _action_result(action_name, response) + async def action(self, action_name: str, params: dict[str, Any] | None = None) -> dict[str, Any]: """Execute a CKAN API action. @@ -79,45 +153,89 @@ async def action(self, action_name: str, params: dict[str, Any] | None = None) - params: Action parameters Returns: - API response result + API response result, or ``{}`` on any failure (see :meth:`call`) """ - client = await self._get_client() - try: - response = await client.post( - f"/api/3/action/{action_name}", - json=params or {}, - ) - response.raise_for_status() - data = response.json() - - if not data.get("success"): - error_msg = data.get("error", {}).get("message", "Unknown error") - self.logger.error("CKAN action failed", action=action_name, error=error_msg) - return {} - - return data.get("result", {}) - - except httpx.HTTPStatusError as e: + result = await self.call(action_name, params) + except CKANActionError as e: # Log 404s at debug level since they're expected for invalid resource IDs - if e.response.status_code == 404: + if e.not_found: self.logger.debug( "CKAN resource not found", action=action_name, - status_code=404, + status_code=e.status_code, resource_id=params.get("resource_id") if params else None, ) + elif e.status_code is None: + self.logger.error("CKAN API error", action=action_name, error=str(e)) else: self.logger.error( - "CKAN HTTP error", + "CKAN action failed", action=action_name, - status_code=e.response.status_code, + status_code=e.status_code, error=str(e), ) return {} except Exception as e: self.logger.error("CKAN API error", action=action_name, error=str(e)) return {} + return {} if result is None else result + + async def upload_call( + self, + action_name: str, + fields: dict[str, Any], + *, + filename: str, + content: bytes, + content_type: str = "application/octet-stream", + ) -> Any: + """Execute a CKAN action that carries a file (``resource_create``/``_patch``). + + CKAN takes the file as the multipart ``upload`` field and every other + field as form text, so non-string values are sent JSON-encoded. A + separate client is used because the shared one pins a JSON + ``Content-Type`` that would override the multipart boundary. + + Raises :class:`CKANActionError` on failure, as :meth:`call` does. + """ + headers = {"Authorization": self._api_key} if self._api_key else {} + data = { + key: value if isinstance(value, str) else json.dumps(value) + for key, value in fields.items() + if value is not None + } + try: + async with httpx.AsyncClient( + base_url=self.ckan_url, timeout=120.0, headers=headers + ) as client: + response = await client.post( + f"/api/3/action/{action_name}", + data=data, + files={"upload": (filename, content, content_type)}, + ) + except httpx.HTTPError as e: + raise CKANActionError(action_name, str(e) or type(e).__name__) from e + return _action_result(action_name, response) + + async def upload_action( + self, + action_name: str, + fields: dict[str, Any], + *, + filename: str, + content: bytes, + content_type: str = "application/octet-stream", + ) -> dict[str, Any]: + """:meth:`upload_call`, returning ``{}`` on failure (as :meth:`action`).""" + try: + result = await self.upload_call( + action_name, fields, filename=filename, content=content, content_type=content_type + ) + except Exception as e: + self.logger.error("CKAN upload failed", action=action_name, error=str(e)[:500]) + return {} + return result or {} async def package_search( self, diff --git a/src/data_concierge/data_layer/connectors/dcat.py b/src/data_concierge/data_layer/connectors/dcat.py index 41a1071..631d500 100644 --- a/src/data_concierge/data_layer/connectors/dcat.py +++ b/src/data_concierge/data_layer/connectors/dcat.py @@ -322,38 +322,37 @@ def parse_dataset(raw: dict[str, Any]) -> DCATDataset: ) -def parse_catalog(body: Any, catalog_url: str) -> DCATCatalog: - """Parse a catalog document into :class:`DCATCatalog`. +def raw_dataset_nodes(body: Any) -> list[dict[str, Any]]: + """Return the unparsed dataset nodes of a catalog document. Handles the three shapes seen in the wild: DCAT-US 1.1 (``{"dataset": [...]}``), a bare JSON-LD array of nodes, and a JSON-LD document with ``@graph``. """ - raw_datasets: list[dict[str, Any]] = [] - title = "" - if isinstance(body, dict): - title = _text(body.get("title")) if isinstance(body.get("dataset"), list): - raw_datasets = [d for d in body["dataset"] if isinstance(d, dict)] - elif isinstance(body.get("@graph"), list): - raw_datasets = [ + return [d for d in body["dataset"] if isinstance(d, dict)] + if isinstance(body.get("@graph"), list): + return [ node for node in body["@graph"] if isinstance(node, dict) and "Dataset" in str(node.get("@type", "")) ] - elif isinstance(body, list): - raw_datasets = [ + return [] + if isinstance(body, list): + typed = [ node for node in body if isinstance(node, dict) and "Dataset" in str(node.get("@type", "")) ] - if not raw_datasets: - # A plain array of dataset objects with no @type annotation. - raw_datasets = [ - node for node in body if isinstance(node, dict) and node.get("title") - ] + # A plain array of dataset objects with no @type annotation. + return typed or [node for node in body if isinstance(node, dict) and node.get("title")] + return [] + - datasets = [parse_dataset(node) for node in raw_datasets] +def parse_catalog(body: Any, catalog_url: str) -> DCATCatalog: + """Parse a catalog document into :class:`DCATCatalog`.""" + title = _text(body.get("title")) if isinstance(body, dict) else "" + datasets = [parse_dataset(node) for node in raw_dataset_nodes(body)] # Drop entries with neither a title nor an identifier — nothing to cite. datasets = [d for d in datasets if d.title or d.identifier] return DCATCatalog( @@ -453,31 +452,46 @@ async def fetch_catalog(self, *, force: bool = False) -> DCATCatalog: if cached and not force and (time.time() - cached[0]) < CATALOG_TTL_SECONDS: return cached[1] - client = await self._http() - body_bytes = bytearray() - async with client.stream("GET", url) as resp: - resp.raise_for_status() - async for chunk in resp.aiter_bytes(): - body_bytes.extend(chunk) - if len(body_bytes) > MAX_CATALOG_BYTES: - raise ValueError( - f"DCAT catalog at {url} exceeds " - f"{MAX_CATALOG_BYTES // (1024 * 1024)} MB; " - "point catalog_url at a filtered catalog instead" - ) - - import json - - catalog = parse_catalog(json.loads(body_bytes.decode("utf-8")), url) + body, size = await self._download_catalog(url) + catalog = parse_catalog(body, url) _CATALOG_CACHE[url] = (time.time(), catalog) self.logger.info( "DCAT catalog loaded", url=url, datasets=len(catalog.datasets), - bytes=len(body_bytes), + bytes=size, ) return catalog + async def fetch_raw_catalog(self) -> Any: + """Fetch the catalog document unparsed and uncached. + + :meth:`fetch_catalog` keeps only the fields the agent searches on. A + caller that republishes the catalog (the Fair Store mirror) needs every + field the portal published, so it takes the document as-is. + """ + body, _size = await self._download_catalog(await self.resolve_catalog_url()) + return body + + async def _download_catalog(self, url: str) -> tuple[Any, int]: + """Stream the catalog under :data:`MAX_CATALOG_BYTES` and decode it.""" + client = await self._http() + body_bytes = bytearray() + async with client.stream("GET", url) as resp: + resp.raise_for_status() + async for chunk in resp.aiter_bytes(): + body_bytes.extend(chunk) + if len(body_bytes) > MAX_CATALOG_BYTES: + raise ValueError( + f"DCAT catalog at {url} exceeds " + f"{MAX_CATALOG_BYTES // (1024 * 1024)} MB; " + "point catalog_url at a filtered catalog instead" + ) + + import json + + return json.loads(body_bytes.decode("utf-8")), len(body_bytes) + async def search_datasets( self, query: str, diff --git a/src/data_concierge/data_layer/qsv_profiling.py b/src/data_concierge/data_layer/qsv_profiling.py index 457f129..6c82ca0 100644 --- a/src/data_concierge/data_layer/qsv_profiling.py +++ b/src/data_concierge/data_layer/qsv_profiling.py @@ -186,7 +186,9 @@ def _extract_qsv_tags(qsv_data: dict) -> list[str]: if isinstance(val, dict): resp = val.get("response", val) if isinstance(resp, dict): - raw = resp.get("tags", []) + # describegpt capitalises the key in some responses + # ({"Attribution": ..., "Tags": [...]}). + raw = next((v for k, v in resp.items() if str(k).lower() == "tags"), []) else: raw = resp if isinstance(raw, str): diff --git a/tests/unit/test_fairstore_populate.py b/tests/unit/test_fairstore_populate.py index 5601370..a91a935 100644 --- a/tests/unit/test_fairstore_populate.py +++ b/tests/unit/test_fairstore_populate.py @@ -12,6 +12,7 @@ populate_fairstore = importlib.util.module_from_spec(_spec) _spec.loader.exec_module(populate_fairstore) +CKANActionError = populate_fairstore.CKANActionError _slug = populate_fairstore._slug _column_dictionary = populate_fairstore._column_dictionary MirrorError = populate_fairstore.MirrorError @@ -71,14 +72,80 @@ def test_falls_back_to_ckan_label(self) -> None: class FakeCKAN: + """Programmed CKAN responses; a response that is (or returns) an exception is raised.""" + + ckan_url = "https://fake.example" + def __init__(self, responses): self.responses = responses self.calls = [] - async def action(self, name, payload=None): - self.calls.append((name, payload or {})) + def _respond(self, name, payload): response = self.responses.get(name, {}) - return response(payload or {}) if callable(response) else response + if callable(response): + response = response(payload) + if isinstance(response, Exception): + raise response + return response + + async def call(self, name, payload=None): + self.calls.append((name, payload or {})) + return self._respond(name, payload or {}) + + async def action(self, name, payload=None): + try: + return await self.call(name, payload) + except CKANActionError: + return {} + + async def upload_call(self, name, payload, *, filename, content, content_type): + self.calls.append((name, {**payload, "_upload": (filename, content, content_type)})) + return self._respond(name, payload) + + +def _http_error(action, status=502): + return CKANActionError(action, f"HTTP {status}", status_code=status) + + +def _datastore(): + """Handlers for a fake DataStore that counts the rows each table holds.""" + rows = {} + + def create(payload): + table = payload["resource_id"] + rows[table] = rows.get(table, 0) + len(payload.get("records") or []) + return {"resource_id": table} + + def upsert(payload): + rows[payload["resource_id"]] += len(payload["records"]) + return {"resource_id": payload["resource_id"]} + + def delete(payload): + if payload["resource_id"] not in rows: + return CKANActionError("datastore_delete", "Not found", status_code=404) + del rows[payload["resource_id"]] + return {} + + def search(payload): + return {"total": rows.get(payload["resource_id"], 0), "records": []} + + return { + "datastore_create": create, + "datastore_upsert": upsert, + "datastore_delete": delete, + "datastore_search": search, + "_rows": rows, + } + + +@pytest.fixture(autouse=True) +def _no_backoff(monkeypatch): + """Retries back off with asyncio.sleep; the tests need not wait.""" + + async def instant(_seconds): + return None + + monkeypatch.setattr(populate_fairstore.asyncio, "sleep", instant) @pytest.mark.asyncio @@ -100,7 +167,9 @@ def capped_page(payload): @pytest.mark.asyncio async def test_write_action_recovers_when_gateway_times_out_after_create(): resource = {"id": "resource-id", "name": "Created resource"} - client = FakeCKAN({"resource_create": {}, "resource_show": resource}) + client = FakeCKAN( + {"resource_create": _http_error("resource_create"), "resource_show": resource} + ) result = await populate_fairstore._write_action( client, @@ -168,6 +237,8 @@ async def test_mirror_preserves_source_identity_and_adds_hierarchy_and_qsv(): "organization_patch": lambda payload: payload, "package_create": lambda payload: payload, "resource_create": lambda payload: payload, + "resource_view_create": lambda payload: payload, + **_datastore(), } ) resource_id = package["resources"][0]["id"] @@ -199,6 +270,8 @@ async def test_mirror_preserves_source_identity_and_adds_hierarchy_and_qsv(): assert summary["datasets"] == 1 assert summary["qsv_resources"] == 1 + # No output files on disk for this profile, so only its dictionary table. + assert summary["qsv_artifacts"] == 1 child = next( payload for action, payload in target.calls @@ -220,7 +293,11 @@ async def test_mirror_preserves_source_identity_and_adds_hierarchy_and_qsv(): assert dataset["notes"] == "Original notes" assert dataset["tags"] == [{"name": "source tag"}] - resource = next(payload for action, payload in target.calls if action == "resource_create") + # A new dataset is created with its resources in the same call; only the + # qsv dictionary table is created on its own. + created = [payload for action, payload in target.calls if action == "resource_create"] + assert [payload.get("qsv_artifact") for payload in created] == ["dictionary"] + [resource] = dataset["resources"] assert resource["id"] == resource_id assert resource["url"] == "https://source.example/data.csv" assert resource["description"] == "Original resource description" @@ -262,9 +339,7 @@ async def test_mirror_supplies_description_when_source_notes_are_blank(): async def test_source_store_slug_does_not_replace_same_named_source_organization(): source, _organization, _package = _source_client() source_url = "https://source.example" - root_id = str( - populate_fairstore.uuid.uuid5(populate_fairstore.uuid.NAMESPACE_URL, source_url) - ) + root_id = str(populate_fairstore.uuid.uuid5(populate_fairstore.uuid.NAMESPACE_URL, source_url)) target = FakeCKAN( { "organization_list": [{"id": root_id, "name": "source-publisher"}], @@ -334,7 +409,7 @@ async def test_existing_source_ids_are_patched_instead_of_duplicated(): }, "organization_patch": lambda payload: payload, "package_patch": lambda payload: payload, - "resource_patch": lambda payload: payload, + "resource_update": lambda payload: payload, } ) @@ -353,4 +428,1599 @@ async def test_existing_source_ids_are_patched_instead_of_duplicated(): assert "resource_create" not in actions assert actions.count("organization_patch") == 3 assert actions.count("package_patch") == 1 - assert actions.count("resource_patch") == 1 + # A source resource is replaced whole, so a field the source dropped goes. + assert actions.count("resource_update") == 1 + + +def _mirrored_extras(portal): + return [ + {"key": "mirror_source_portal", "value": portal}, + {"key": "mirror_source_url", "value": f"https://{portal}.example"}, + ] + + +@pytest.mark.asyncio +async def test_source_snapshot_skips_records_the_portal_mirrored_from_elsewhere(): + """A Fair Store that already mirrors WPRDC must not re-publish WPRDC as its own.""" + native_org = {"id": "org-native", "name": "dathere", "extras": []} + copied_org = { + "id": "org-copy", + "name": "city-of-pittsburgh", + "extras": _mirrored_extras("wprdc"), + } + native = {"id": "pkg-native", "name": "native", "owner_org": "org-native", "resources": []} + copied = { + "id": "pkg-copy", + "name": "copied", + "owner_org": "org-copy", + "extras": _mirrored_extras("wprdc"), + "resources": [], + } + orgs = {row["id"]: row for row in (native_org, copied_org)} + client = FakeCKAN( + { + "organization_list": [native_org, copied_org], + "organization_show": lambda payload: orgs[payload["id"]], + "group_list": [], + "package_search": {"count": 2, "results": [native, copied]}, + } + ) + + snapshot = await populate_fairstore._source_snapshot(client, site_id="ckan") + + assert [row["name"] for row in snapshot["organizations"]] == ["dathere"] + assert [row["name"] for row in snapshot["packages"]] == ["native"] + assert snapshot["copies"] == { + "organizations": 1, + "groups": 0, + "datasets": 1, + "origins": ["wprdc"], + } + + # Mirroring the origin portal itself keeps its own records. + own = await populate_fairstore._source_snapshot(client, site_id="wprdc") + assert len(own["packages"]) == 2 + + +def test_package_payload_carries_schema_fields_the_target_would_drop(): + package = { + "id": "pkg", + "name": "pkg", + "title": "Package", + "notes": "Notes", + "owner_org": "org", + "data_steward_name": "Jane Steward", + "temporal_coverage": "2015-2024", + "related_documents": ["https://example.org/a.pdf"], + "num_resources": 3, + "tracking_summary": {"total": 9}, + "organization": {"name": "org"}, + "metadata_modified": "2026-09-01T00:00:00", + "extras": [{"key": "frequency", "value": "monthly"}], + } + + payload = populate_fairstore._package_payload( + package, + site_id="wprdc", + source_url="https://source.example", + store_id="root", + known_org_ids={"org"}, + ) + + extras = {item["key"]: item["value"] for item in payload["extras"]} + assert extras["data_steward_name"] == "Jane Steward" + assert extras["temporal_coverage"] == "2015-2024" + assert extras["related_documents"] == '["https://example.org/a.pdf"]' + assert extras["frequency"] == "monthly" + assert extras["mirror_source_metadata_modified"] == "2026-09-01T00:00:00" + assert extras["mirror_source_id"] == "pkg" + # Computed by CKAN, so neither copied as a field nor carried as an extra. + for computed in ("num_resources", "tracking_summary", "organization"): + assert computed not in extras + assert computed not in payload + assert payload["owner_org"] == "org" + + +def test_resource_payload_keeps_uploads_pointing_at_the_source_copy(): + """CKAN rewrites an upload's URL to the serving site, where no file exists.""" + resource = { + "id": "res", + "url": "https://source.example/dataset/pkg/resource/res/download/data.csv", + "url_type": "upload", + "name": "Data", + "format": "CSV", + "position": 0, + "tracking_summary": {"total": 1}, + "datastore_active": True, + "datastore_contains_all_records_of_source_file": True, + "dataspatial_wkt_field": "geom", + "preview_rows": [1, 2], + } + + payload = populate_fairstore._resource_payload( + resource, package_id="pkg", site_id="wprdc", source_url="https://source.example", qsv=None + ) + + assert payload["url"] == resource["url"] + assert "url_type" not in payload + assert payload["mirror_source_url_type"] == "upload" + assert payload["mirror_source_datastore_active"] is True + assert not any(key.startswith("datastore_") for key in payload) + assert payload["dataspatial_wkt_field"] == "geom" + assert payload["preview_rows"] == [1, 2] + for bookkeeping in ("position", "tracking_summary"): + assert bookkeeping not in payload + + +def test_resource_payload_leaves_plain_links_alone(): + payload = populate_fairstore._resource_payload( + {"id": "res", "url": "https://elsewhere.example/a.csv", "url_type": ""}, + package_id="pkg", + site_id="wprdc", + source_url="https://source.example", + qsv=None, + ) + assert payload["url_type"] == "" + assert "mirror_source_url_type" not in payload + + +def test_uploaded_images_use_the_source_absolute_url(): + uploaded = { + "image_url": "2015-10-09-seal.jpg", + "image_display_url": "https://source.example/uploads/group/2015-10-09-seal.jpg", + } + assert populate_fairstore._with_image({"image_url": "x"}, uploaded)["image_url"] == ( + "https://source.example/uploads/group/2015-10-09-seal.jpg" + ) + linked = {"image_url": "https://cdn.example/logo.png"} + assert populate_fairstore._with_image({}, linked)["image_url"] == "https://cdn.example/logo.png" + assert "image_url" not in populate_fairstore._with_image({"image_url": ""}, {}) + + +class TestCleanTag: + def test_valid_tags_pass_through_unchanged(self) -> None: + for tag in ("covid-19", "K-12 Education", "u.s.", "snake_case"): + assert populate_fairstore._clean_tag(tag) == tag + + def test_invalid_characters_become_hyphens(self) -> None: + assert populate_fairstore._clean_tag("l&i") == "l-i" + assert populate_fairstore._clean_tag("fw&gs") == "fw-gs" + + def test_unusable_tags_are_dropped(self) -> None: + assert populate_fairstore._clean_tag("&") == "" + assert populate_fairstore._clean_tag("a") == "" + + def test_duplicates_after_cleaning_collapse(self) -> None: + refs = populate_fairstore._tag_refs([{"name": "l&i"}, {"name": "l-i"}]) + assert refs == [{"name": "l-i"}] + + +_PORTAL = "https://data.example.gov" + + +def _dcat_dataset(view_id, title, **overrides): + dataset = { + "@type": "dcat:Dataset", + "identifier": f"{_PORTAL}/api/views/{view_id}", + "title": title, + "description": f"About {title}", + "landingPage": f"{_PORTAL}/d/{view_id}", + "modified": "2026-02-28", + "accessLevel": "public", + "keyword": ["health", "l&i"], + "theme": ["Health"], + "license": "https://www.usa.gov/government-works", + "contactPoint": {"fn": "Data Team", "hasEmail": "mailto:data@example.gov"}, + "publisher": {"@type": "org:Organization", "name": "data.example.gov"}, + "distribution": [ + { + "@type": "dcat:Distribution", + "downloadURL": f"{_PORTAL}/api/views/{view_id}/rows.csv", + "mediaType": "text/csv", + }, + { + "@type": "dcat:Distribution", + "downloadURL": f"{_PORTAL}/api/views/{view_id}/rows.json", + "mediaType": "application/json", + "describedBy": f"{_PORTAL}/api/views/{view_id}/columns.json", + }, + ], + } + dataset.update(overrides) + return dataset + + +def _socrata_row(view_id, owner): + return { + "resource": { + "id": view_id, + "type": "dataset", + "attribution": "Bureau of Records", + "columns_field_name": ["county", "deaths"], + "columns_name": ["County", "Deaths"], + "columns_datatype": ["Text", "Number"], + "columns_description": ["County name", "Count of deaths"], + "page_views": {"page_views_total": 10}, + }, + "classification": { + "domain_category": "Health", + "domain_metadata": [ + {"key": "Data-Management_Business-Owner", "value": owner}, + {"key": "Data-Management_Update-Frequency", "value": "Monthly"}, + ], + }, + "owner": {"display_name": "Someone"}, + } + + +def _snapshot(datasets, socrata=None): + return populate_fairstore._dcat_snapshot( + {"dataset": datasets}, + site_title="Example Open Data", + source_url=_PORTAL, + socrata=socrata, + ) + + +def test_dcat_snapshot_derives_stable_identity_and_keeps_every_field(): + snapshot = _snapshot([_dcat_dataset("abcd-1234", "Overdose Deaths")]) + [package] = snapshot["packages"] + + expected_id = str( + populate_fairstore.uuid.uuid5( + populate_fairstore.uuid.NAMESPACE_URL, f"{_PORTAL}/api/views/abcd-1234" + ) + ) + assert package["id"] == expected_id + assert package["name"] == "overdose-deaths-abcd-1234" + assert package["url"] == f"{_PORTAL}/d/abcd-1234" + assert package["license_id"] == "other-pd" + assert package["maintainer_email"] == "data@example.gov" + assert package["groups"] == [{"name": "health"}] + extras = {item["key"]: item["value"] for item in package["extras"]} + assert extras["dcat_identifier"] == f"{_PORTAL}/api/views/abcd-1234" + assert extras["dcat_accessLevel"] == "public" + assert extras["dcat_keyword"] == '["health", "l&i"]' + assert extras["dcat_license"] == "https://www.usa.gov/government-works" + assert "dcat_distribution" not in extras + + csv, json_dist = package["resources"] + assert csv["format"] == "CSV" and json_dist["format"] == "JSON" + assert json_dist["dcat_describedBy"] == f"{_PORTAL}/api/views/abcd-1234/columns.json" + assert csv["_source_id"] == f"{_PORTAL}/api/views/abcd-1234/rows.csv" + # Deterministic, so a rerun patches instead of duplicating. + assert ( + _snapshot([_dcat_dataset("abcd-1234", "Overdose Deaths")])["packages"][0]["resources"][0][ + "id" + ] + == csv["id"] + ) + + # The portal publishing under its own name is the root, not a child org. + assert snapshot["organizations"] == [] + assert package["owner_org"] == "" + assert snapshot["groups"][0]["title"] == "Health" + + +def test_dcat_snapshot_uses_socrata_owner_for_the_hierarchy_and_columns(): + socrata = { + "abcd-1234": _socrata_row("abcd-1234", "Department of Health (DOH)"), + "wxyz-9876": _socrata_row("wxyz-9876", "Department of Health (DOH)"), + } + snapshot = _snapshot( + [_dcat_dataset("abcd-1234", "Overdose Deaths"), _dcat_dataset("wxyz-9876", "Births")], + socrata=socrata, + ) + + [org] = snapshot["organizations"] + assert org["name"] == "department-of-health-doh" + assert org["title"] == "Department of Health (DOH)" + assert org["groups"] == [] # hangs off the portal root + assert {"key": "publisher_source", "value": "socrata:Data-Management_Business-Owner"} in org[ + "extras" + ] + assert all(package["owner_org"] == org["id"] for package in snapshot["packages"]) + + extras = {item["key"]: item["value"] for item in snapshot["packages"][0]["extras"]} + assert extras["socrata_id"] == "abcd-1234" + assert extras["socrata_attribution"] == "Bureau of Records" + assert extras["socrata_Data-Management_Update-Frequency"] == "Monthly" + # Volatile analytics and account names stay out. + assert not any("page_views" in key or "owner" == key for key in extras) + + # The column list sits on the CSV it describes, not on the dataset, where + # Solr would index it as a single (size-limited) term. + assert "source_data_dictionary" not in extras + csv, json_dist = snapshot["packages"][0]["resources"] + assert "source_data_dictionary" not in json_dist + assert csv["source_column_count"] == "2" + columns = populate_fairstore.json.loads(csv["source_data_dictionary"]) + assert columns[1] == { + "name": "deaths", + "label": "Deaths", + "type": "Number", + "description": "Count of deaths", + } + + +def test_dataset_extras_over_the_solr_term_limit_are_named_not_stored(): + """One oversized extra must not make CKAN reject the whole dataset.""" + package = { + "id": "pkg", + "name": "pkg", + "notes": "Notes", + "owner_org": "org", + "spatial": "x" * 40_000, + "extras": [{"key": "frequency", "value": "monthly"}], + } + + payload = populate_fairstore._package_payload( + package, + site_id="wprdc", + source_url="https://source.example", + store_id="root", + known_org_ids={"org"}, + ) + + extras = {item["key"]: item["value"] for item in payload["extras"]} + assert "spatial" not in extras + assert extras["mirror_oversized_extras"] == '["spatial"]' + assert extras["frequency"] == "monthly" + assert extras["mirror_source_id"] == "pkg" + + +def test_dcat_publisher_chain_becomes_nested_organizations(): + chained = _dcat_dataset( + "abcd-1234", + "Permits", + publisher={ + "name": "Office of Permits", + "subOrganizationOf": { + "name": "Department of State", + "subOrganizationOf": {"name": "data.example.gov"}, + }, + }, + ) + snapshot = _snapshot([chained]) + + orgs = {org["title"]: org for org in snapshot["organizations"]} + assert set(orgs) == {"Office of Permits", "Department of State"} + assert orgs["Office of Permits"]["groups"] == [ + {"name": "department-of-state", "capacity": "parent"} + ] + assert orgs["Department of State"]["groups"] == [] + assert snapshot["packages"][0]["owner_org"] == orgs["Office of Permits"]["id"] + + +def test_dcat_hierarchy_ignores_a_contradiction_that_would_loop(): + a_under_b = _dcat_dataset( + "abcd-1234", + "One", + publisher={"name": "Agency A", "subOrganizationOf": {"name": "Agency B"}}, + ) + b_under_a = _dcat_dataset( + "wxyz-9876", + "Two", + publisher={"name": "Agency B", "subOrganizationOf": {"name": "Agency A"}}, + ) + + orgs = {org["title"]: org for org in _snapshot([a_under_b, b_under_a])["organizations"]} + + assert orgs["Agency A"]["groups"] == [{"name": "agency-b", "capacity": "parent"}] + assert orgs["Agency B"]["groups"] == [] + + +def test_dcat_uuid_keeps_a_source_uuid_identifier(): + """DKAN catalogs identify datasets by UUID; the mirror keeps it, as for CKAN.""" + source_uuid = "0a1b2c3d-4e5f-4a6b-8c7d-9e0f1a2b3c4d" + assert populate_fairstore._dcat_uuid(source_uuid, _PORTAL, "dataset") == source_uuid + name = populate_fairstore._dcat_dataset_name("Water Quality", source_uuid, source_uuid) + assert name == "water-quality-0a1b2c3d" + + +def test_dcat_qsv_index_is_rekeyed_to_the_first_tabular_resource(): + snapshot = _snapshot([_dcat_dataset("abcd-1234", "Overdose Deaths")]) + index = {"datasets": [{"resources": [{"resource_id": "abcd-1234", "row_count": 5}]}]} + + rekeyed = populate_fairstore._dcat_qsv_index(index, snapshot) + + [resource] = rekeyed["datasets"][0]["resources"] + assert resource["resource_id"] == snapshot["packages"][0]["resources"][0]["id"] + assert resource["row_count"] == 5 + + +@pytest.mark.asyncio +async def test_dcat_mirror_creates_datasets_with_resources_under_the_portal_root(): + socrata = {"abcd-1234": _socrata_row("abcd-1234", "Department of Health (DOH)")} + snapshot = _snapshot([_dcat_dataset("abcd-1234", "Overdose Deaths")], socrata=socrata) + target = FakeCKAN( + { + "organization_list": [], + "group_list": [], + "package_search": {"count": 0, "results": []}, + "organization_create": lambda payload: payload, + "organization_patch": lambda payload: payload, + "group_create": lambda payload: payload, + "package_create": lambda payload: payload, + } + ) + + summary = await populate_fairstore.mirror_catalog( + source=None, + snapshot=snapshot, + target=target, + site_id="example", + site_title="Example Open Data", + source_url=_PORTAL, + apply=True, + ) + + assert summary["datasets"] == 1 and summary["resources"] == 2 + hierarchy = next( + payload + for action, payload in target.calls + if action == "organization_patch" and "groups" in payload + ) + assert hierarchy["groups"] == [{"name": "example", "capacity": "parent"}] + dataset = next(payload for action, payload in target.calls if action == "package_create") + assert dataset["tags"] == [{"name": "health"}, {"name": "l-i"}] + assert len(dataset["resources"]) == 2 + extras = {item["key"]: item["value"] for item in dataset["extras"]} + assert extras["mirror_source_id"] == f"{_PORTAL}/api/views/abcd-1234" + assert not any(key.startswith("_") for key in dataset) + assert not any(key.startswith("_") for resource in dataset["resources"] for key in resource) + + +_FAKE_KEY = "sk-or-v1-" + "0" * 64 +_ATTRIBUTION = ( + "\n\nAttribution:\nGenerated by qsv v20.0.0 describegpt\n" + f"Command line: qsv describegpt data.csv --all --api-key {_FAKE_KEY}\n" + "Model: google/gemini-2.5-flash-lite\n" + "WARNING: Description generated by an LLM and may contain inaccuracies." +) + + +def _qsv_profile(tmp_path, *, frequency_rows=2): + """An onboarded profile on disk, shaped like onboard_ckan.py's output.""" + directory = tmp_path / "data" / "ckan_onboard" / "wprdc" / "parking" + directory.mkdir(parents=True) + (directory / "qsv_stats.csv").write_text( + "field,type,min,max,mean,nullcount,cardinality\n" + "zone,String,A,Z,,0,26\n" + "spaces,Integer,1,400,52.5,3,120\n" + ) + rows = "".join( + f"zone,Z{index},{index + 1},1.5,{index + 1}\n" for index in range(frequency_rows) + ) + (directory / "qsv_frequency.csv").write_text("field,value,count,percentage,rank\n" + rows) + (directory / "qsv_dict.json").write_text( + populate_fairstore.json.dumps({"Description": {"response": "Parking." + _ATTRIBUTION}}) + ) + return { + "resource_id": "44444444-4444-4444-8444-444444444444", + "resource_name": "Parking CSV", + "local_path": str(directory / "data.csv"), + "onboarded_at": "2026-05-08T07:03:37+00:00", + "row_count": 138, + "status": "ok", + "qsv_description": "Parking spaces by zone." + _ATTRIBUTION, + "qsv_tags": ["parking", "zoning"], + "columns": [ + { + "name": "zone", + "qsv_label": "Zone", + "qsv_description": "Parking zone letter", + "qsv_type": "String", + "stats": {"nullcount": "0", "cardinality": "26", "min": "A", "max": "Z"}, + "top_values": [{"value": "A", "count": 9}], + }, + { + "name": "spaces", + "qsv_label": "Spaces", + "qsv_type": "Integer", + "stats": {"mean": "52.5", "q2_median": "40", "nullcount": "3"}, + }, + ], + } + + +def test_qsv_dataset_extras_summarise_the_profile_without_the_api_key(tmp_path): + extras = populate_fairstore._qsv_dataset_extras(_qsv_profile(tmp_path)) + + # Only the prose: the provenance block, key and command line included, + # stays in the describegpt resource. + assert extras["qsv_description"] == "Parking spaces by zone." + assert extras["qsv_tags"] == '["parking", "zoning"]' + assert extras["qsv_row_count"] == "138" + assert extras["qsv_column_count"] == "2" + assert extras["qsv_columns"] == '["zone", "spaces"]' + assert extras["qsv_model"] == "google/gemini-2.5-flash-lite" + assert extras["qsv_version"] == "20.0.0" + assert extras["qsv_profiled_resource_id"] == "44444444-4444-4444-8444-444444444444" + + +def test_qsv_artifacts_cover_dictionary_stats_frequency_and_describegpt(tmp_path): + artifacts = {a["kind"]: a for a in populate_fairstore._qsv_artifacts(_qsv_profile(tmp_path))} + + assert list(artifacts) == ["dictionary", "stats", "frequency", "describegpt"] + + fields, records = artifacts["dictionary"]["table"] + assert [f["id"] for f in fields][:4] == ["column", "label", "type", "description"] + assert records[0]["label"] == "Zone" and records[0]["cardinality"] == 26 + assert records[1]["median"] == 40 and records[1]["null_count"] == 3 + assert records[1]["description"] == "" + + fields, records = artifacts["frequency"]["table"] + types = {f["id"]: f["type"] for f in fields} + assert types == { + "field": "text", + "value": "text", + "count": "numeric", + "percentage": "numeric", + "rank": "numeric", + } + assert records[1] == {"field": "zone", "value": "Z1", "count": 2, "percentage": 1.5, "rank": 2} + + fields, records = artifacts["stats"]["table"] + assert records[1]["mean"] == "52.5" and records[0]["mean"] == "" + + filename, content, content_type = artifacts["describegpt"]["file"] + assert (filename, content_type) == ("describegpt.json", "application/json") + assert _FAKE_KEY.encode() not in content + + +def test_qsv_artifacts_without_output_files_still_publish_the_dictionary(tmp_path): + profile = _qsv_profile(tmp_path) + profile["local_path"] = str(tmp_path / "missing" / "data.csv") + assert [a["kind"] for a in populate_fairstore._qsv_artifacts(profile)] == ["dictionary"] + + +@pytest.mark.asyncio +async def test_publish_qsv_creates_datastore_tables_and_uploads_describegpt(tmp_path): + profile = _qsv_profile(tmp_path, frequency_rows=2500) + datastore = _datastore() + target = FakeCKAN( + { + "resource_create": lambda payload: payload, + "resource_view_create": lambda payload: payload, + **datastore, + } + ) + + counts = await populate_fairstore._publish_qsv( + target, + package_id="pkg", + qsv=profile, + site_id="wprdc", + existing_resources=set(), + ) + + assert counts == {"created": 4, "refreshed": 0, "unchanged": 0, "removed": 0} + expected_id = populate_fairstore._qsv_resource_id(profile["resource_id"], "frequency") + tables = [ + payload + for action, payload in target.calls + if action == "resource_create" and "_upload" not in payload + ] + assert [t["qsv_artifact"] for t in tables] == ["dictionary", "stats", "frequency"] + assert tables[2]["id"] == expected_id and tables[2]["package_id"] == "pkg" + # What datastore_create would set for an inline resource. + assert tables[2]["url_type"] == "datastore" + assert tables[2]["url"] == "_datastore_only_resource" + creates = [payload for action, payload in target.calls if action == "datastore_create"] + frequency = creates[2] + assert frequency["resource_id"] == expected_id + assert len(frequency["records"]) == 1000 + assert datastore["_rows"][expected_id] == 2500 + # The rest of the 2,500 rows follow in insert batches. + upserts = [payload for action, payload in target.calls if action == "datastore_upsert"] + assert [len(u["records"]) for u in upserts] == [1000, 500] + assert all(u["method"] == "insert" and u["resource_id"] == expected_id for u in upserts) + + [upload] = [ + payload + for action, payload in target.calls + if action == "resource_create" and "_upload" in payload + ] + assert upload["qsv_artifact"] == "describegpt" + assert upload["_upload"][0] == "describegpt.json" + + # Each table gets its sortable table view once its rows are loaded. + views = [payload for action, payload in target.calls if action == "resource_view_create"] + assert [v["view_type"] for v in views] == ["datatables_view"] * 3 + assert views[2]["resource_id"] == expected_id + + +@pytest.mark.asyncio +async def test_publish_qsv_rerun_replaces_tables_instead_of_appending(tmp_path): + profile = _qsv_profile(tmp_path) + existing = { + populate_fairstore._qsv_resource_id(profile["resource_id"], kind) + for kind in ("dictionary", "stats", "frequency", "describegpt") + } + datastore = _datastore() + # The tables hold the first run's rows. + for kind, count in (("dictionary", 2), ("stats", 2), ("frequency", 2)): + datastore["_rows"][populate_fairstore._qsv_resource_id(profile["resource_id"], kind)] = ( + count + ) + target = FakeCKAN( + { + "resource_patch": lambda payload: payload, + # The tables already have their view from the first run. + "resource_view_list": [{"view_type": "datatables_view"}], + **datastore, + } + ) + + await populate_fairstore._publish_qsv( + target, + package_id="pkg", + qsv=profile, + site_id="wprdc", + existing_resources=existing, + ) + + actions = [action for action, _ in target.calls] + assert "resource_create" not in actions + assert "resource_view_create" not in actions + assert actions.count("datastore_delete") == 3 + for index, action in enumerate(actions): + if action == "datastore_create": + assert actions[index - 1] == "datastore_delete" + creates = [payload for action, payload in target.calls if action == "datastore_create"] + assert all("resource" not in c and c["resource_id"] in existing for c in creates) + assert sorted(datastore["_rows"].values()) == [2, 2, 2] + + +@pytest.mark.asyncio +async def test_mirror_puts_the_qsv_profile_on_the_dataset(tmp_path): + source, _organization, package = _source_client() + profile = _qsv_profile(tmp_path) + profile["resource_id"] = package["resources"][0]["id"] + target = FakeCKAN( + { + "organization_list": [], + "group_list": [], + "package_search": {"count": 0, "results": []}, + "organization_create": lambda payload: payload, + "organization_patch": lambda payload: payload, + "package_create": lambda payload: payload, + "resource_create": lambda payload: payload, + "resource_view_create": lambda payload: payload, + **_datastore(), + } + ) + + summary = await populate_fairstore.mirror_catalog( + source=source, + target=target, + site_id="source-store", + site_title="Source Store", + source_url="https://source.example", + apply=True, + qsv_index={"datasets": [{"resources": [profile]}]}, + ) + + assert summary["qsv_artifacts"] == 4 + dataset = next(payload for action, payload in target.calls if action == "package_create") + extras = {item["key"]: item["value"] for item in dataset["extras"]} + assert extras["qsv_tags"] == '["parking", "zoning"]' + assert extras["qsv_columns"] == '["zone", "spaces"]' + assert extras["mirror_source_id"] == package["id"] + + +def test_source_clients_never_fall_back_to_the_configured_key(monkeypatch): + """The Fair Store token in CKAN_API_KEY must not be sent to a source portal.""" + from pydantic import SecretStr + + from data_concierge.core.config import settings + from data_concierge.data_layer.connectors.ckan import CKANClient + + monkeypatch.setattr(settings, "ckan_api_key", SecretStr("fairstore-sysadmin-token")) + + assert CKANClient("https://source.example", use_default_key=False)._api_key is None + assert CKANClient("https://target.example")._api_key == "fairstore-sysadmin-token" + + +def test_summary_tables_read_and_cap_geometry_sized_cells(tmp_path): + """qsv reports whole WKT polygons as a column's min/max/mode.""" + path = tmp_path / "qsv_stats.csv" + wkt = "POLYGON((" + "1 2," * 60_000 + "1 2))" + path.write_text(f'field,mode\ngeom,"{wkt}"\n') + + _fields, [record] = populate_fairstore._csv_table(path) + + assert len(record["mode"]) < 10_100 + assert record["mode"].endswith(f"[truncated: {len(wkt):,} characters]") + + +# -- reads fail closed (review findings: partial snapshots) ------------------- + + +@pytest.mark.asyncio +async def test_a_failed_page_aborts_instead_of_returning_part_of_the_organizations(): + rows = [{"id": str(index), "name": f"publisher-{index}"} for index in range(5)] + + def page(payload): + offset = payload.get("offset", 0) + return rows[offset : offset + 2] if offset == 0 else _http_error("organization_list") + + client = FakeCKAN({"organization_list": page}) + + with pytest.raises(MirrorError, match="organization_list failed"): + await populate_fairstore._all_group_rows(client, "organization_list") + # The second page was retried before giving up. + assert [payload["offset"] for _, payload in client.calls] == [0, 2, 2, 2] + + +@pytest.mark.asyncio +async def test_a_transient_failure_is_retried_and_the_read_completes(): + rows = [{"id": str(index), "name": f"publisher-{index}"} for index in range(3)] + failures = iter([True]) + + def page(payload): + offset = payload.get("offset", 0) + if offset == 2 and next(failures, False): + return _http_error("organization_list", 503) + return rows[offset : offset + 2] + + client = FakeCKAN({"organization_list": page}) + + assert await populate_fairstore._all_group_rows(client, "organization_list") == rows + + +@pytest.mark.asyncio +async def test_package_paging_that_stops_short_of_the_count_aborts(): + client = FakeCKAN( + { + "package_search": lambda payload: ( + {"count": 3, "results": [{"id": "a"}, {"id": "b"}]} + if payload["start"] == 0 + else {"count": 3, "results": []} + ) + } + ) + + with pytest.raises(MirrorError, match="stopped at 2 of 3"): + await populate_fairstore._all_packages(client) + + +@pytest.mark.asyncio +async def test_organization_show_failure_is_not_replaced_by_the_list_row(): + source, _organization, _package = _source_client() + source.responses["organization_show"] = _http_error("organization_show", 500) + target = FakeCKAN( + {"organization_list": [], "group_list": [], "package_search": {"count": 0, "results": []}} + ) + + with pytest.raises(MirrorError, match="organization_show failed"): + await populate_fairstore.mirror_catalog( + source=source, + target=target, + site_id="source-store", + site_title="Source Store", + source_url="https://source.example", + apply=True, + ) + assert not [action for action, _ in target.calls if action.endswith(("_create", "_patch"))] + + +@pytest.mark.asyncio +async def test_socrata_enrichment_failure_aborts_a_run_that_writes(monkeypatch): + catalog = {"dataset": [{"identifier": "https://data.example.gov/api/views/abcd-1234"}]} + + async def unavailable(_client, _domain, _offset): + raise populate_fairstore.httpx.ConnectError("down") + + monkeypatch.setattr(populate_fairstore, "_socrata_page", unavailable) + + with pytest.raises(MirrorError, match="--no-enrich"): + await populate_fairstore._socrata_metadata(catalog, required=True) + # A dry run only warns. + assert await populate_fairstore._socrata_metadata(catalog, required=False) == {} + + +@pytest.mark.asyncio +async def test_socrata_pages_are_retried_on_rate_limits(): + import httpx + + responses = iter( + [ + httpx.Response(429), + httpx.Response(200, json={"results": [{"resource": {"id": "abcd-1234"}}]}), + ] + ) + transport = httpx.MockTransport(lambda request: next(responses)) + async with httpx.AsyncClient(transport=transport) as client: + body = await populate_fairstore._socrata_page(client, "data.example.gov", 0) + + assert body["results"][0]["resource"]["id"] == "abcd-1234" + + +# -- renames and collisions ---------------------------------------------------- + + +def _target_row(row_id, name, portal): + return { + "id": row_id, + "name": name, + "extras": [{"key": "mirror_source_portal", "value": portal}], + } + + +def test_a_source_rename_of_this_portals_record_is_a_rename_not_a_collision(): + source = { + "organizations": [], + "groups": [], + "packages": [{"id": "pkg-1", "name": "budget-2024", "resources": []}], + } + target = { + "organizations": [], + "groups": [], + "packages": [_target_row("pkg-1", "budget-2023", "wprdc")], + } + + renames = populate_fairstore._assert_no_collisions( + source, target, store_name="wprdc", store_id="root", site_id="wprdc" + ) + + assert renames == ["dataset 'budget-2023' -> 'budget-2024'"] + + +def test_the_same_uuid_under_another_portals_provenance_is_still_a_collision(): + source = { + "organizations": [], + "groups": [], + "packages": [{"id": "pkg-1", "name": "budget-2024", "resources": []}], + } + target = { + "organizations": [], + "groups": [], + "packages": [_target_row("pkg-1", "budget-2023", "data-pa-gov")], + } + + with pytest.raises(MirrorError, match="already has a different name"): + populate_fairstore._assert_no_collisions( + source, target, store_name="wprdc", store_id="root", site_id="wprdc" + ) + + +def test_duplicate_uuids_within_one_snapshot_are_collisions(): + source = { + "organizations": [], + "groups": [], + "packages": [ + {"id": "same", "name": "first", "resources": []}, + {"id": "same", "name": "second", "resources": []}, + ], + } + empty = {"organizations": [], "groups": [], "packages": []} + + with pytest.raises(MirrorError, match="appears twice"): + populate_fairstore._assert_no_collisions( + source, empty, store_name="x", store_id="root", site_id="x" + ) + + +def test_dcat_names_already_published_are_kept_and_new_ones_avoid_them(): + source = { + "organizations": [ + # A new publisher now sorts first and took the bare slug. + {"id": "org-new", "name": "health", "groups": []}, + {"id": "org-old", "name": "health-0badc0de", "groups": []}, + {"id": "org-child", "name": "clinics", "groups": [{"name": "health-0badc0de"}]}, + ], + "groups": [], + "packages": [{"id": "pkg-1", "name": "retitled-abcd-1234", "resources": []}], + } + target = { + "organizations": [ + _target_row("org-old", "health", "data-pa-gov"), + _target_row("org-child", "clinics", "data-pa-gov"), + ], + "groups": [], + "packages": [_target_row("pkg-1", "original-title-abcd-1234", "data-pa-gov")], + } + + populate_fairstore._pin_target_names(source, target, site_id="data-pa-gov") + + names = {row["id"]: row["name"] for row in source["organizations"]} + assert names["org-old"] == "health" + assert names["org-new"] == "health-org-new" + assert source["organizations"][2]["groups"] == [{"name": "health"}] + assert source["packages"][0]["name"] == "original-title-abcd-1234" + # Pinned names pass the preflight as unchanged records. + assert ( + populate_fairstore._assert_no_collisions( + source, target, store_name="data-pa-gov", store_id="root", site_id="data-pa-gov" + ) + == [] + ) + + +# -- retried writes cannot duplicate or fake success --------------------------- + + +@pytest.mark.asyncio +async def test_an_insert_that_committed_before_its_502_is_rebuilt_not_doubled(): + datastore = _datastore() + lost = iter([True]) + + def upsert_then_lose_the_response(payload): + datastore["datastore_upsert"](payload) + return _http_error("datastore_upsert") if next(lost, False) else {} + + target = FakeCKAN({**datastore, "datastore_upsert": upsert_then_lose_the_response}) + records = [{"n": index} for index in range(2500)] + + await populate_fairstore._load_table(target, "table", [{"id": "n"}], records, label="t") + + assert datastore["_rows"]["table"] == 2500 + assert [action for action, _ in target.calls].count("datastore_delete") == 2 + + +@pytest.mark.asyncio +async def test_a_failed_view_listing_never_adds_a_second_table_view(): + target = FakeCKAN({"resource_view_list": _http_error("resource_view_list")}) + + with pytest.raises(MirrorError): + await populate_fairstore._ensure_table_view(target, "table", label="t") + assert "resource_view_create" not in [action for action, _ in target.calls] + + +@pytest.mark.asyncio +async def test_a_view_create_whose_response_was_lost_is_not_repeated(): + views = [] + + def create(payload): + views.append({"view_type": payload["view_type"]}) + return _http_error("resource_view_create") + + target = FakeCKAN( + {"resource_view_list": lambda _payload: list(views), "resource_view_create": create} + ) + + await populate_fairstore._ensure_table_view(target, "table", label="t") + + assert len(views) == 1 + + +@pytest.mark.asyncio +async def test_a_create_that_finds_an_admin_deleted_object_is_not_counted_as_created(): + client = FakeCKAN( + { + "package_create": _http_error("package_create", 409), + "package_show": {"id": "pkg", "name": "pkg", "state": "deleted"}, + } + ) + + with pytest.raises(populate_fairstore.DeletedInTarget): + await populate_fairstore._write_action( + client, "package_create", {"id": "pkg", "name": "pkg"}, label="pkg" + ) + + +@pytest.mark.asyncio +async def test_a_rejected_write_is_not_retried(): + client = FakeCKAN({"package_patch": _http_error("package_patch", 409)}) + + with pytest.raises(MirrorError, match="package_patch failed"): + await populate_fairstore._write_action(client, "package_patch", {"id": "pkg"}, label="pkg") + assert len(client.calls) == 1 + + +# -- reporting and run control ------------------------------------------------ + + +def test_records_withdrawn_at_the_source_are_reported(): + source = { + "packages": [{"id": "kept", "resources": [{"id": "r-kept"}]}], + } + kept = _target_row("kept", "kept", "wprdc") + kept["resources"] = [ + {"id": "r-kept", "mirror_source_portal": "wprdc"}, + {"id": "r-gone", "mirror_source_portal": "wprdc"}, + {"id": "qsv", "qsv_artifact": "stats"}, + ] + target = { + "packages": [ + kept, + _target_row("gone", "gone-dataset", "wprdc"), + _target_row("other", "other-portal", "data-pa-gov"), + ] + } + + report = populate_fairstore._withdrawn(source, target, site_id="wprdc") + + assert report == {"datasets": 1, "dataset_names": ["gone-dataset"], "resources": 1} + + +@pytest.mark.asyncio +async def test_all_sites_continue_past_a_failing_portal_and_skip_the_target(monkeypatch): + sites = [ + {"id": "broken", "url": "https://broken.example"}, + {"id": "fine", "url": "https://fine.example"}, + {"id": "store", "url": "https://www.store.example/"}, + ] + monkeypatch.setattr(populate_fairstore, "list_sites", lambda: sites) + + async def mirror_site(site_id, _site, _args, _api_key): + if site_id == "broken": + raise MirrorError("source down") + return {"site": site_id} + + monkeypatch.setattr(populate_fairstore, "_mirror_site", mirror_site) + args = populate_fairstore.argparse.Namespace( + site="all", + organization=None, + index_path=None, + api_key_env="FAIRSTORE_API_KEY", + api_key_file=None, + apply=False, + target_url="https://store.example", + ) + + summaries = await populate_fairstore._main(args) + + assert summaries == [ + {"site": "broken", "error": "MirrorError: source down"}, + {"site": "fine"}, + {"site": "store", "skipped": "this portal is the target"}, + ] + + +@pytest.mark.parametrize( + "footer", + [ + "\n\nAttribution: Generated by qsv v20.0.0 describegpt\nCommand line: qsv describegpt x", + "\n\nGenerated by qsv v20.0.0 describegpt\nCommand line: qsv describegpt x", + "\n\n## Attribution\nGenerated by qsv v20.0.0 describegpt\nModel: m", + "\n\n---\nGenerated by qsv v20.0.0 describegpt\nModel: m", + "\n\nGenerated by qsv v20.0.0 describegpt\nModel: m", + "\n\n@attribution Generated by qsv v20.0.0 describegpt\nModel: m", + "\n\n**Attribution**\n\nCommand line: qsv describegpt data.csv --all", + ], +) +def test_describegpt_prose_drops_every_provenance_style(footer): + prose = "Parking by zone.\n\n## Notable Characteristics\n\n* **Zones:** 26 of them." + assert populate_fairstore._describegpt_prose(prose + footer) == prose + + +def test_describegpt_prose_leaves_a_description_without_provenance_alone(): + assert populate_fairstore._describegpt_prose("Just prose.\n") == "Just prose." + + +# -- the client tells failures apart from empty results ----------------------- + + +@pytest.mark.asyncio +@pytest.mark.parametrize( + ("status", "body", "transient", "not_found"), + [ + (502, "Bad Gateway", True, False), + ( + 404, + {"success": False, "error": {"__type": "Not Found Error", "message": "x"}}, + False, + True, + ), + ( + 409, + {"success": False, "error": {"__type": "Validation Error", "name": ["taken"]}}, + False, + False, + ), + ], +) +async def test_ckan_client_call_raises_classified_errors(status, body, transient, not_found): + import httpx + + from data_concierge.data_layer.connectors.ckan import CKANClient + + def respond(request): + if isinstance(body, dict): + return httpx.Response(status, json=body) + return httpx.Response(status, text=body) + + client = CKANClient("https://ckan.example", use_default_key=False) + client._client = httpx.AsyncClient( + base_url="https://ckan.example", transport=httpx.MockTransport(respond) + ) + try: + with pytest.raises(CKANActionError) as caught: + await client.call("package_show", {"id": "x"}) + assert caught.value.transient is transient + assert caught.value.not_found is not_found + # The lenient wrapper keeps its contract for the app's callers. + assert await client.action("package_show", {"id": "x"}) == {} + finally: + await client.close() + + +# -- qsv output is published only for the file it describes ------------------- + + +def _profile_on_disk( + tmp_path, *, header, stats_fields, frequency_fields, status="qsv_failed", described=None +): + directory = tmp_path / "data" / "ckan_onboard" / "wprdc" / "blotter" + directory.mkdir(parents=True) + (directory / "historical.csv").write_text(",".join(header) + "\n") + (directory / "qsv_stats.csv").write_text( + "field,type\n" + "".join(f"{name},String\n" for name in stats_fields) + ) + (directory / "qsv_frequency.csv").write_text( + "field,value,count,percentage,rank\n" + + "".join(f"{name},x,1,100,1\n" for name in frequency_fields) + ) + fields = [{"name": name} for name in (described if described is not None else header)] + (directory / "qsv_dict.json").write_text( + populate_fairstore.json.dumps( + { + "Dictionary": {"response": {"fields": [{"name": "_id"}, *fields]}}, + "Tags": {"response": {"Attribution": "x", "Tags": ["crime"]}}, + } + ) + ) + + return { + "resource_id": "55555555-5555-4555-8555-555555555555", + "resource_name": "Historical Blotter Data", + "local_path": str(directory / "historical.csv"), + "status": status, + "qsv_tags": [], + "columns": [ + {"name": "ccr", "stats": {"cardinality": "9"}, "top_values": [{"value": "1"}]}, + # Contributed only by the other file's stats. + {"name": "COUNCIL_DISTRICT", "stats": {"cardinality": "9"}}, + # The portal's own DataStore field, absent from the download. + {"name": "_geom", "ckan_type": "text"}, + ], + } + + +def test_another_files_stats_and_frequency_are_not_published_under_this_files_name(tmp_path): + profile = _profile_on_disk( + tmp_path, + header=["ccr", "offense"], + stats_fields=["ccr", "offense", "COUNCIL_DISTRICT"], + frequency_fields=["ccr", "COUNCIL_DISTRICT"], + ) + + kinds = [artifact["kind"] for artifact in populate_fairstore._qsv_artifacts(profile)] + assert kinds == ["dictionary", "describegpt"] + + vetted = populate_fairstore._vet_profile(profile) + assert [column["name"] for column in vetted["columns"]] == ["ccr", "_geom"] + assert all(not column["stats"] for column in vetted["columns"]) + # The profiled file's columns, and describegpt's capitalised Tags key. + extras = populate_fairstore._qsv_dataset_extras(profile) + assert extras["qsv_columns"] == '["ccr", "offense"]' + assert extras["qsv_column_count"] == "2" + assert extras["qsv_tags"] == '["crime"]' + + +def test_matching_output_is_published_even_for_a_failed_describegpt_run(tmp_path): + profile = _profile_on_disk( + tmp_path, + header=["ccr", "offense"], + stats_fields=["ccr", "offense"], + frequency_fields=["offense"], + ) + + kinds = [artifact["kind"] for artifact in populate_fairstore._qsv_artifacts(profile)] + assert kinds == ["dictionary", "stats", "frequency", "describegpt"] + + +def test_extract_qsv_tags_reads_a_capitalised_tags_key(): + from data_concierge.data_layer.qsv_profiling import _extract_qsv_tags + + data = {"Tags": {"response": {"Attribution": "Generated by qsv", "Tags": ["a", "b"]}}} + assert _extract_qsv_tags(data) == ["a", "b"] + assert _extract_qsv_tags({"tags": {"response": {"tags": ["c"]}}}) == ["c"] + + +# -- reruns skip what has not changed ----------------------------------------- + + +@pytest.mark.asyncio +async def test_a_rerun_leaves_unchanged_qsv_resources_alone_and_removes_stale_ones(tmp_path): + profile = _qsv_profile(tmp_path) + first = FakeCKAN( + { + "resource_create": lambda payload: payload, + "resource_view_create": lambda payload: payload, + **_datastore(), + } + ) + await populate_fairstore._publish_qsv( + first, package_id="pkg", qsv=profile, site_id="wprdc", existing_resources=set() + ) + # What the Fair Store holds after the first run, as a snapshot reports it: + # each create, then the digest patched in once a table's rows are loaded. + held = {} + for action, payload in first.calls: + if action == "resource_create": + held[payload["id"]] = { + **{k: v for k, v in payload.items() if k != "_upload"}, + "datastore_active": "_upload" not in payload, + "url_type": "upload" if "_upload" in payload else "datastore", + } + elif action == "resource_patch": + held[payload["id"]].update(payload) + assert all(resource["qsv_digest"] for resource in held.values()) + # A kind this profile no longer produces, left by an earlier run. + stale = populate_fairstore._qsv_resource_id(profile["resource_id"], "frequency") + (Path(profile["local_path"]).parent / "qsv_frequency.csv").unlink() + + rerun = FakeCKAN({"resource_view_list": [{"view_type": "datatables_view"}], **_datastore()}) + counts = await populate_fairstore._publish_qsv( + rerun, package_id="pkg", qsv=profile, site_id="wprdc", existing_resources=held + ) + + assert counts == {"created": 0, "refreshed": 0, "unchanged": 3, "removed": 1} + writes = [ + (action, payload) for action, payload in rerun.calls if action != "resource_view_list" + ] + assert writes == [("resource_delete", {"id": stale})] + + +@pytest.mark.asyncio +async def test_a_changed_qsv_table_is_refreshed(tmp_path): + profile = _qsv_profile(tmp_path) + dictionary = populate_fairstore._qsv_resource_id(profile["resource_id"], "dictionary") + held = {dictionary: {"id": dictionary, "qsv_digest": "old", "datastore_active": True}} + target = FakeCKAN( + { + "resource_patch": lambda payload: payload, + "resource_create": lambda payload: payload, + "resource_view_create": lambda payload: payload, + **_datastore(), + } + ) + + counts = await populate_fairstore._publish_qsv( + target, package_id="pkg", qsv=profile, site_id="wprdc", existing_resources=held + ) + + assert counts["refreshed"] == 1 and counts["created"] == 3 + assert ("resource_patch", dictionary) in [(a, p.get("id")) for a, p in target.calls] + + +@pytest.mark.asyncio +async def test_an_unchanged_dataset_is_not_patched_and_a_dry_run_plans_it(tmp_path): + source, _organization, package = _source_client() + + def target_with(held_package): + return FakeCKAN( + { + "organization_list": [], + "group_list": [], + "package_search": {"count": 1, "results": [held_package]}, + "organization_create": lambda payload: payload, + "organization_patch": lambda payload: payload, + "package_patch": lambda payload: payload, + "resource_patch": lambda payload: payload, + } + ) + + run = { + "source": source, + "site_id": "source-store", + "site_title": "Source Store", + "source_url": "https://source.example", + } + # A first dry run against an empty-ish target computes the digests. + probe = target_with({"id": "other", "name": "other", "resources": []}) + plan = await populate_fairstore.mirror_catalog(target=probe, apply=False, **run) + assert plan["created"] >= 2 and not [a for a, _ in probe.calls if a.endswith("_create")] + + # Hold the dataset exactly as the mirror would write it. + payload = populate_fairstore._package_payload( + package, + site_id="source-store", + source_url="https://source.example", + store_id=str( + populate_fairstore.uuid.uuid5( + populate_fairstore.uuid.NAMESPACE_URL, "https://source.example" + ) + ), + known_org_ids={package["owner_org"]}, + ) + populate_fairstore._stamp_package(payload) + resource = populate_fairstore._resource_payload( + package["resources"][0], + package_id=package["id"], + site_id="source-store", + source_url="https://source.example", + qsv=None, + ) + populate_fairstore._stamp_resource(resource) + target = target_with({**payload, "resources": [resource]}) + + summary = await populate_fairstore.mirror_catalog(target=target, apply=True, **run) + + actions = [action for action, _ in target.calls] + assert "package_patch" not in actions and "resource_patch" not in actions + assert summary["unchanged"] == 2 + + +def test_a_failed_profile_without_a_description_still_builds_its_payloads(): + """index.json stores qsv_description as null for a failed describegpt run.""" + profile = {"resource_id": "r", "status": "qsv_failed", "qsv_description": None, "columns": []} + resource = populate_fairstore._resource_payload( + {"id": "r", "url": "https://source.example/r.csv"}, + package_id="p", + site_id="wprdc", + source_url="https://source.example", + qsv=profile, + ) + assert resource["qsv_description"] == "" + assert "qsv_description" not in populate_fairstore._qsv_dataset_extras(profile) + + +@pytest.mark.asyncio +async def test_a_table_whose_load_was_cut_short_is_reloaded_on_the_next_run(tmp_path): + profile = _qsv_profile(tmp_path, frequency_rows=2500) + frequency = populate_fairstore._qsv_resource_id(profile["resource_id"], "frequency") + held = {} + + def create(payload): + held[payload["id"]] = {**payload, "datastore_active": True} + return payload + + def patch(payload): + held[payload["id"]].update(payload) + return held[payload["id"]] + + datastore = _datastore() + target = FakeCKAN( + { + "resource_create": create, + "resource_patch": patch, + "resource_view_create": lambda payload: payload, + **datastore, + # Every insert batch fails: the frequency table stops at 1,000 rows. + "datastore_upsert": _http_error("datastore_upsert"), + } + ) + with pytest.raises(MirrorError): + await populate_fairstore._publish_qsv( + target, package_id="pkg", qsv=profile, site_id="wprdc", existing_resources=set() + ) + assert held[frequency]["qsv_digest"] == "" + + target.responses["datastore_upsert"] = datastore["datastore_upsert"] + counts = await populate_fairstore._publish_qsv( + target, package_id="pkg", qsv=profile, site_id="wprdc", existing_resources=held + ) + + assert counts["refreshed"] >= 1 + assert datastore["_rows"][frequency] == 2500 + assert held[frequency]["qsv_digest"] + + +@pytest.mark.asyncio +async def test_nothing_is_removed_when_the_onboarding_directory_is_missing(tmp_path): + profile = _qsv_profile(tmp_path) + profile["local_path"] = str(tmp_path / "elsewhere" / "data.csv") + held = { + populate_fairstore._qsv_resource_id(profile["resource_id"], kind): {} + for kind in ("dictionary", "stats", "frequency", "describegpt") + } + target = FakeCKAN( + { + "resource_patch": lambda payload: payload, + "resource_view_create": lambda p: p, + **_datastore(), + } + ) + + counts = await populate_fairstore._publish_qsv( + target, package_id="pkg", qsv=profile, site_id="wprdc", existing_resources=held + ) + + assert counts["removed"] == 0 + assert "resource_delete" not in [action for action, _ in target.calls] + + +# -- the second review's findings ---------------------------------------------- + + +def test_a_name_reused_by_the_source_evicts_the_withdrawn_record(): + source = { + "organizations": [], + "groups": [], + "packages": [{"id": "new-id", "name": "parking", "resources": []}], + } + target = { + "organizations": [], + "groups": [], + "packages": [_target_row("old-id-00000000", "parking", "wprdc")], + } + evictions = [] + + renames = populate_fairstore._assert_no_collisions( + source, target, store_name="wprdc", store_id="root", site_id="wprdc", evictions=evictions + ) + + assert evictions == [("dataset", "old-id-00000000", "parking-withdrawn-old-id-0")] + assert renames == ["dataset 'parking' (withdrawn at source) -> 'parking-withdrawn-old-id-0'"] + + +def test_a_name_held_by_another_portal_is_still_a_collision(): + source = {"organizations": [], "groups": [], "packages": [{"id": "new", "name": "parking"}]} + target = { + "organizations": [], + "groups": [], + "packages": [_target_row("old", "parking", "data-pa-gov")], + } + + with pytest.raises(MirrorError, match="already has a different UUID"): + populate_fairstore._assert_no_collisions( + source, target, store_name="wprdc", store_id="root", site_id="wprdc", evictions=[] + ) + + +def test_row_counts_are_published_only_when_they_are_the_files_own(): + count = populate_fairstore._profile_row_count + assert count({"row_count": 138, "status": "ok"}) == 138 + assert count({"row_count": 495251, "status": "qsv_failed"}) == 495251 + assert count({"row_count": 0, "status": "qsv_failed"}) is None + assert count({"row_count": 5000, "status": "ok", "truncated_download": True}) is None + + +def test_a_field_the_source_cleared_is_sent_as_empty(): + payload = populate_fairstore._package_payload( + {"id": "p", "name": "p", "title": "P", "notes": "n"}, + site_id="s", + source_url="https://s.example", + store_id="root", + known_org_ids=set(), + ) + assert payload["license_id"] == "" and payload["maintainer_email"] == "" + + +@pytest.mark.asyncio +async def test_an_empty_qsv_index_does_not_strip_published_profiles(): + source, _organization, package = _source_client() + held = { + **package, + "resources": [ + *({**resource, "package_id": package["id"]} for resource in package["resources"]), + { + "id": "q", + "package_id": package["id"], + "qsv_profile_site": "source-store", + "qsv_profiled_resource_id": "r", + }, + ], + } + target = FakeCKAN( + { + "organization_list": [], + "group_list": [], + "package_search": {"count": 1, "results": [held]}, + } + ) + + with pytest.raises(MirrorError, match="onboarding index missing"): + await populate_fairstore.mirror_catalog( + source=source, + target=target, + site_id="source-store", + site_title="Source Store", + source_url="https://source.example", + apply=True, + qsv_index={}, + ) + assert not [a for a, _ in target.calls if a.endswith(("_create", "_patch", "_update"))] + + +@pytest.mark.asyncio +async def test_an_organization_deleted_in_the_fair_store_is_skipped_not_fatal(): + source, organization, _package = _source_client() + target = FakeCKAN( + { + "organization_list": [], + "group_list": [], + "package_search": {"count": 0, "results": []}, + "organization_create": lambda payload: ( + _http_error("organization_create", 409) + if payload["id"] == organization["id"] + else payload + ), + "organization_show": {"id": organization["id"], "state": "deleted"}, + "organization_patch": lambda payload: payload, + "package_create": lambda payload: payload, + } + ) + + summary = await populate_fairstore.mirror_catalog( + source=source, + target=target, + site_id="source-store", + site_title="Source Store", + source_url="https://source.example", + apply=True, + ) + + assert summary["skipped_deleted_in_target"] == [organization["name"]] + dataset = next(payload for action, payload in target.calls if action == "package_create") + # Its dataset hangs off the portal root instead. + assert dataset["owner_org"] != organization["id"] + + +@pytest.mark.asyncio +async def test_unchanged_organizations_and_groups_are_not_patched_again(): + source, organization, package = _source_client() + first = FakeCKAN( + { + "organization_list": [], + "group_list": [], + "package_search": {"count": 0, "results": []}, + "organization_create": lambda payload: payload, + "organization_patch": lambda payload: payload, + "package_create": lambda payload: payload, + } + ) + run = { + "source": source, + "site_id": "source-store", + "site_title": "Source Store", + "source_url": "https://source.example", + "apply": True, + } + await populate_fairstore.mirror_catalog(target=first, **run) + created = [p for a, p in first.calls if a == "organization_create"] + + rerun = FakeCKAN( + { + "organization_list": created, + "group_list": [], + "package_search": {"count": 0, "results": []}, + "package_create": lambda payload: payload, + } + ) + summary = await populate_fairstore.mirror_catalog(target=rerun, **run) + + assert not [ + a + for a, _ in rerun.calls + if a.startswith("organization_") and a not in ("organization_list",) + ] + assert summary["unchanged"] >= 2 + + +@pytest.mark.asyncio +async def test_an_organization_the_paged_rows_missed_aborts_the_snapshot(): + client = FakeCKAN({"organization_list": ["a", "b", "c"]}) + + with pytest.raises(MirrorError, match="paged 2 of 3"): + await populate_fairstore._check_complete( + client, "organization_list", [{"name": "a"}, {"name": "b"}], label="portal" + ) + # One unpaged call: WPRDC's CDN rejects a names-only listing with limit/offset. + assert client.calls == [("organization_list", {"sort": "name asc"})] + + +def test_another_files_describegpt_output_is_not_published(tmp_path): + """uniform-crime-reporting-data: describegpt described blotter-data-ucr-coded.csv.""" + profile = _profile_on_disk( + tmp_path, + header=["ccr", "offense"], + stats_fields=["ccr", "offense"], + frequency_fields=["offense"], + described=["ccr", "offense", "COUNCIL_DISTRICT"], + ) + profile["qsv_description"] = "Crimes by council district." + profile["columns"][0]["qsv_label"] = "CCR number" + + kinds = [artifact["kind"] for artifact in populate_fairstore._qsv_artifacts(profile)] + extras = populate_fairstore._qsv_dataset_extras(profile) + vetted = populate_fairstore._vet_profile(profile) + + assert kinds == ["dictionary", "stats", "frequency"] + assert "qsv_description" not in extras and "qsv_tags" not in extras + assert "qsv_label" not in vetted["columns"][0] From f7617aab45d8fac7d013f133191a3dbe6663c2dc Mon Sep 17 00:00:00 2001 From: Mohammed Minhajuddin <68331751+minhajuddin2510@users.noreply.github.com> Date: Tue, 29 Sep 2026 10:27:38 -0400 Subject: [PATCH 2/3] feat: manage the Fair Store from the admin panel and offer it in chat Adds an Admin -> Fair Store page and the code behind it (gateway/fairstore.py): - Connection: the Fair Store URL and a sysadmin API token. The token is never returned to the browser and is bound to the origin it was saved for; a URL that is already a registered source portal is refused, so a mirror can never write into one. FAIRSTORE_URL / FAIRSTORE_API_KEY seed the settings. - Mirror runs: populate_fairstore runs under the onboarding job runner (one job at a time), with a live log, a per-portal summary and cancel. Running jobs heartbeat their record and use absolute log offsets, so any Cloud Run instance can show or cancel a run. - Runs on Cloud Run read qsv profiles from the storage backend (--qsv-source): sync_to_storage now writes sync_manifest.json (CSV headers, dataset directories, qsv outputs present) and stores JSON with API keys scrubbed, and staging from it produces the same writes as a local run. - Live status (reachability, token check, datasets per source portal) and the Fair Store's runtime site settings (title, about, logo, custom CSS). - Chat source: the agent searches the Fair Store's catalog and loads rows from the portal each dataset was mirrored from (CSV/TSV files only). Citations, agent-log sources and notebook code name the portal that served each call; the notebook editor shares the same per-call path. Also hardens the agent: numeric tool arguments sent as strings, a portal with DataStore SQL disabled, DCAT not-found answers counted as successes, and notebook cells for calls that fetched nothing. --- .env.example | 6 + .gitignore | 1 + README.md | 20 + scripts/populate_fairstore.py | 202 ++- src/data_concierge/agents/llm_agent.py | 435 ++++- src/data_concierge/agents/notebook_editor.py | 30 +- .../agents/notebook_generator.py | 13 +- src/data_concierge/core/config.py | 7 + .../data_layer/qsv_profiling.py | 96 +- src/data_concierge/gateway/ckan_sites.py | 5 + src/data_concierge/gateway/fairstore.py | 581 +++++++ src/data_concierge/gateway/onboarding_jobs.py | 531 +++++- src/data_concierge/gateway/router.py | 222 ++- src/data_concierge/ui/static/js/admin.js | 504 +++++- src/data_concierge/ui/static/js/dictionary.js | 14 + src/data_concierge/ui/templates/admin.html | 225 ++- .../ui/templates/dictionary.html | 10 +- src/data_concierge/ui/templates/index.html | 3 + src/data_concierge/ui/web.py | 34 +- tests/unit/test_fairstore_link.py | 1483 +++++++++++++++++ 20 files changed, 4236 insertions(+), 186 deletions(-) create mode 100644 src/data_concierge/gateway/fairstore.py create mode 100644 tests/unit/test_fairstore_link.py diff --git a/.env.example b/.env.example index b339b8c..cd6a17a 100644 --- a/.env.example +++ b/.env.example @@ -78,6 +78,12 @@ DATA_COMMONS_API_KEY= CKAN_URL=https://data.dathere.com CKAN_API_KEY= +# Fair Store — the CKAN that mirrors every registered portal. Seeds for the +# admin panel's Fair Store page (values saved there take precedence). The token +# is a CKAN sysadmin API token; it is only ever sent to FAIRSTORE_URL's host. +# FAIRSTORE_URL=https://fairstore.example.org +# FAIRSTORE_API_KEY= + # WPRDC (Western PA Regional Data Center) - City of Pittsburgh open data WPRDC_CKAN_URL=https://data.wprdc.org WPRDC_ORGANIZATION=city-of-pittsburgh diff --git a/.gitignore b/.gitignore index 6c7f146..f431d25 100644 --- a/.gitignore +++ b/.gitignore @@ -168,6 +168,7 @@ fairstore-token* /roles.json /notebook_verification/ /github_settings.json +/fairstore_settings.json /landing_settings.json /system_prompt_settings.json /runtime_settings.json diff --git a/README.md b/README.md index 2a9c06a..3cc060b 100644 --- a/README.md +++ b/README.md @@ -435,6 +435,26 @@ python -m scripts.populate_fairstore \ --apply ``` +### Managing the Fair Store from Verikan + +The admin panel's **Fair Store** page links Verikan to it: + +- **Connection** — the Fair Store URL and a sysadmin API token (`ckan user + token add verikan` on the Fair Store). `FAIRSTORE_URL` and + `FAIRSTORE_API_KEY` seed these; values saved on the page take precedence. + The token is never shown again and is only ever sent to the host it was saved + for: moving the URL to another host drops it. +- **Chat source** — optionally offers the Fair Store as a data source in chat. + The agent searches its catalog and loads each dataset's rows from the portal + it was mirrored from. +- **Mirror runs** — dry runs and writes of every portal or one, with a live log + and a per-portal summary. On Cloud Run the qsv profiles come from the storage + backend that onboarding syncs to (`--qsv-source`). +- **Site settings** — the Fair Store's title, description, home page and about + text, logo and custom CSS, read and saved live. + +The chat sidebar and the data dictionary link to the Fair Store once it is set. + ### Deploying the Fair Store on GCP `deploy/fairstore/deploy.sh` runs the same stack on one Compute Engine VM next diff --git a/scripts/populate_fairstore.py b/scripts/populate_fairstore.py index 7285c81..651754a 100644 --- a/scripts/populate_fairstore.py +++ b/scripts/populate_fairstore.py @@ -46,6 +46,7 @@ import os import re import sys +import tempfile import uuid from collections.abc import Iterable from pathlib import Path @@ -66,7 +67,10 @@ raw_dataset_nodes, ) from data_concierge.data_layer.onboard_index import _scrub_secrets # noqa: E402 -from data_concierge.data_layer.qsv_profiling import _extract_qsv_tags # noqa: E402 +from data_concierge.data_layer.qsv_profiling import ( # noqa: E402 + MIRRORED_QSV_OUTPUTS, + _extract_qsv_tags, +) from data_concierge.gateway.ckan_sites import ( # noqa: E402 PORTAL_TYPE_DCAT, get_site, @@ -256,21 +260,125 @@ def _slug(text: str, fallback: str) -> str: return slug[:100] -def _load_index(site: str, index_path: str | Path | None = None) -> dict[str, Any]: - """Load qsv output when present; mirroring itself does not require it.""" - candidates = ( - [Path(index_path)] - if index_path - else [ +_ONBOARD_PREFIXES = ("ckan_onboard", "dcat_onboard") +# What the mirror reads from an onboarding directory besides index.json. +_QSV_OUTPUTS = MIRRORED_QSV_OUTPUTS +# A dataset directory name as Path.name yields it: no separators, not . or .. +_SAFE_DIR_RE = re.compile(r"^(?!\.{1,2}$)[^/\\\x00-\x1f]+$") + + +def _load_index( + site: str, + index_path: str | Path | None = None, + *, + source: str = "local", + stage_root: Path | None = None, +) -> dict[str, Any]: + """Load qsv output when present; mirroring itself does not require it. + + ``source`` is ``local`` (this checkout's onboarding directories), ``storage`` + (the storage backend, which onboarding syncs to — GCS on Cloud Run), or + ``auto`` (local, else storage). Output read from storage is staged under + ``stage_root`` first, because the profile vetting reads files. + """ + if index_path: + return json.loads(Path(index_path).read_text(encoding="utf-8")) + if source in ("auto", "local"): + for path in ( _PROJECT_ROOT / "data" / "ckan_onboard" / site / "index.json", _PROJECT_ROOT / "data" / "dcat_onboard" / site / "index.json", _PROJECT_ROOT / "ckan_onboard" / site / "index.json", - ] + ): + if path.exists(): + return json.loads(path.read_text(encoding="utf-8")) + if source == "local": + return {} + root = stage_root or Path(tempfile.mkdtemp(prefix="fairstore-qsv-")) + return _stage_from_storage(site, root) + + +def _stage_from_storage(site: str, root: Path) -> dict[str, Any]: + """Copy one portal's onboarding output from the storage backend to ``root``. + + Returns its index with each ``local_path`` pointed at the staged directory. + Only what the mirror reads is copied (the profiled CSVs are not synced, and + are not needed). The ``sync_manifest.json`` that ``sync_to_storage`` writes + stands in for the local directory: each file's header (checked against its + qsv output), which dataset directories existed, and which qsv outputs they + held. A directory it does not list is not created, so, as locally, its qsv + resources are not removed; a stale object it does not list is not staged. + """ + from data_concierge.data_layer.storage import storage + + index: dict[str, Any] | None = None + prefix = "" + for prefix in _ONBOARD_PREFIXES: + index = storage.read_json(f"{prefix}/{site}/index.json") + if index: + break + if not index: + return {} + manifest = storage.read_json(f"{prefix}/{site}/sync_manifest.json") or {} + headers = manifest.get("headers") + dirs = manifest.get("dirs") + if not isinstance(headers, dict) or not isinstance(dirs, dict): + # Without it, output left by another file in the same directory passes + # the vetting and failed profiles lose their stats, so the run would + # rewrite (and delete) qsv resources a local run keeps. No profiles: + # the mirror's qsv guard then refuses to write. + print( + f"{site}: the onboarding output in the storage backend has no " + f"sync_manifest.json (it predates it); re-sync it with onboard_*.py " + f"--rebuild-index before mirroring its qsv profiles from storage", + file=sys.stderr, + ) + return {} + + staged: dict[str, Path] = {} + for dataset in index.get("datasets", []) or []: + for resource in dataset.get("resources", []) or []: + local_path = Path(str(resource.get("local_path") or "")) + # onboard_*.py write each resource to ///, + # which sync_to_storage stores under ///. + dataset_dir = local_path.parent.name + key_dir = f"{prefix}/{site}/{dataset_dir}" + listed = dirs.get(key_dir) + if ( + not local_path.name + or not _SAFE_DIR_RE.match(dataset_dir) + or not isinstance(listed, list) + ): + # Never fall back to a local path: point at a directory that + # does not exist, which is what the synced directory lacked. + resource["local_path"] = str(root / "_absent" / (local_path.name or "file")) + continue + if key_dir not in staged: + directory = root / prefix / site / dataset_dir + directory.mkdir(parents=True, exist_ok=True) + for name in _QSV_OUTPUTS: + if name not in listed: + continue + data = storage.read_bytes(f"{key_dir}/{name}") + if data is None: + # Listed but gone: staging less would remove its qsv + # resources, so no profiles (the qsv guard then refuses). + print( + f"{site}: {key_dir}/{name} is listed in sync_manifest.json but " + f"missing from the storage backend; re-sync the onboarding output", + file=sys.stderr, + ) + return {} + (directory / name).write_bytes(data) + staged[key_dir] = directory + resource["local_path"] = str(staged[key_dir] / local_path.name) + header = headers.get(f"{key_dir}/{local_path.name}") + if isinstance(header, list): + resource["profiled_header"] = [str(name) for name in header] + print( + f"{site}: staged qsv output for {len(staged)} datasets from the storage backend", + file=sys.stderr, ) - for path in candidates: - if path.exists(): - return json.loads(path.read_text(encoding="utf-8")) - return {} + return index def _column_dictionary(resource: dict[str, Any]) -> list[dict[str, Any]]: @@ -327,7 +435,9 @@ def _profiled_header(qsv: dict[str, Any]) -> list[str] | None: path = Path(local_path) path = path if path.is_absolute() else _PROJECT_ROOT / path if not path.is_file(): - return None + # Staged from the storage backend: the header recorded at sync time. + recorded = qsv.get("profiled_header") + return [str(name) for name in recorded if name] if isinstance(recorded, list) else None _raise_csv_field_limit() with path.open(newline="", encoding="utf-8-sig", errors="replace") as handle: return [name for name in next(csv.reader(handle), []) if name] @@ -2353,6 +2463,16 @@ def group_digest(payload: dict[str, Any], parents: Any = None) -> str: return summary +def _site_index(site_id: str, args: argparse.Namespace) -> dict[str, Any]: + stage_root = getattr(args, "stage_root", None) + return _load_index( + site_id, + args.index_path, + source=getattr(args, "qsv_source", "local"), + stage_root=Path(stage_root) / site_id if stage_root else None, + ) + + async def _mirror_site( site_id: str, site: dict[str, Any], args: argparse.Namespace, api_key: str ) -> dict[str, Any]: @@ -2391,7 +2511,7 @@ async def _mirror_site( source=None, snapshot=snapshot, pin_names=True, - qsv_index=_dcat_qsv_index(_load_index(site_id, args.index_path), snapshot), + qsv_index=_dcat_qsv_index(_site_index(site_id, args), snapshot), **common, ) summary["socrata_enriched"] = sum( @@ -2406,7 +2526,7 @@ async def _mirror_site( return await mirror_catalog( source=source, organization=args.organization, - qsv_index=_load_index(site_id, args.index_path), + qsv_index=_site_index(site_id, args), limit=args.limit, **common, ) @@ -2426,15 +2546,21 @@ def key(value: str) -> tuple[str, str]: async def _main(args: argparse.Namespace) -> dict[str, Any] | list[dict[str, Any]]: + requested = [part.strip() for part in args.site.split(",") if part.strip()] + many = args.site == "all" or len(requested) > 1 + if many and (args.organization or args.index_path): + raise MirrorError("--organization and --index-path need a single --site") if args.site == "all": - if args.organization or args.index_path: - raise MirrorError("--organization and --index-path need a single --site") sites = [(str(site["id"]), site) for site in list_sites()] else: - site = get_site(args.site) - if not site: - raise MirrorError(f"Unknown site {args.site!r}; add it to ckan_sites.json first") - sites = [(args.site, site)] + sites = [] + for site_id in dict.fromkeys(requested): + site = get_site(site_id) + if not site: + raise MirrorError(f"Unknown site {site_id!r}; add it to ckan_sites.json first") + sites.append((site_id, site)) + if not sites: + raise MirrorError("--site needs a portal ID, a comma-separated list, or 'all'") api_key = os.environ.get(args.api_key_env, "") if args.api_key_file: @@ -2444,7 +2570,7 @@ async def _main(args: argparse.Namespace) -> dict[str, Any] | list[dict[str, Any f"Set {args.api_key_env} or --api-key-file when using --apply; a sysadmin token is required" ) - if args.site != "all" and _same_portal(sites[0][1]["url"], args.target_url): + if not many and _same_portal(sites[0][1]["url"], args.target_url): raise MirrorError(f"{args.site} is the target portal; it cannot be mirrored into itself") summaries: list[dict[str, Any]] = [] @@ -2456,13 +2582,20 @@ async def _main(args: argparse.Namespace) -> dict[str, Any] | list[dict[str, Any try: summaries.append(await _mirror_site(site_id, site, args, api_key)) except Exception as exc: - if args.site != "all": + if not many: raise # One portal's outage must not stop the others from refreshing; # the failure is reported and the exit status is non-zero. print(f"error: {site_id}: {type(exc).__name__}: {exc}", file=sys.stderr) summaries.append({"site": site_id, "error": f"{type(exc).__name__}: {exc}"}) - return summaries if args.site == "all" else summaries[0] + return summaries if many else summaries[0] + + +def _write_summary(path: str | None, summary: Any) -> None: + """Write the run's summary where a supervisor (the admin panel) reads it.""" + if not path: + return + Path(path).write_text(json.dumps(summary, indent=2, sort_keys=True), encoding="utf-8") def main() -> None: @@ -2470,7 +2603,10 @@ def main() -> None: parser.add_argument( "--site", default="wprdc", - help="source site ID from ckan_sites.json, or 'all' for every registered portal", + help=( + "source site ID from ckan_sites.json, a comma-separated list of them, " + "or 'all' for every registered portal" + ), ) # Deliberately not CKAN_URL/CKAN_API_KEY: those are the app's portal # settings, and in the app container CKAN_URL is data.dathere.com. @@ -2488,15 +2624,29 @@ def main() -> None: ) parser.add_argument("--api-key-env", default="FAIRSTORE_API_KEY") parser.add_argument("--api-key-file", help="read the target sysadmin token from a file") + parser.add_argument( + "--qsv-source", + choices=("auto", "local", "storage"), + default="auto", + help=( + "where qsv profiles come from: this checkout's onboarding output (local), " + "the storage backend onboarding syncs to (storage), or local else storage (auto)" + ), + ) + parser.add_argument("--summary-file", help="also write the JSON summary to this file") args = parser.parse_args() # Development logging is DEBUG on stdout, where this command prints its # JSON summary; a line per HTTP request would bury it. for noisy in ("asyncio", "httpx", "httpcore"): logging.getLogger(noisy).setLevel(logging.WARNING) try: - summary = asyncio.run(_main(args)) + with tempfile.TemporaryDirectory(prefix="fairstore-qsv-") as stage_root: + args.stage_root = stage_root + summary = asyncio.run(_main(args)) except MirrorError as exc: + _write_summary(args.summary_file, {"error": str(exc)}) parser.exit(2, f"error: {exc}\n") + _write_summary(args.summary_file, summary) print(json.dumps(summary, indent=2, sort_keys=True)) if isinstance(summary, list) and any("error" in item for item in summary): sys.exit(1) diff --git a/src/data_concierge/agents/llm_agent.py b/src/data_concierge/agents/llm_agent.py index dd23882..0281deb 100644 --- a/src/data_concierge/agents/llm_agent.py +++ b/src/data_concierge/agents/llm_agent.py @@ -153,6 +153,16 @@ def _source_for_tool(tool_name: str, data_source: str) -> str: return f"ckan:{data_source}" +def _registered_config(site_id: str) -> dict[str, Any] | None: + """A registered portal's config, or ``None``.""" + try: + from data_concierge.gateway import ckan_sites + + return ckan_sites.get_portal_config(site_id) + except Exception: # pragma: no cover - registry failures are non-fatal + return None + + def _sources_from_trace( tool_calls_trace: list[dict[str, Any]], data_source: str, @@ -161,7 +171,9 @@ def _sources_from_trace( ) -> list[DataSource]: """The data sources a run actually touched, in first-use order. - A CKAN tool call attributes the primary portal; an ``mcp__{server}__*`` + A CKAN tool call attributes the portal that served it (``portal_id`` on the + trace entry: an override, or a Fair Store load routed to its origin), else + the primary portal; an ``mcp__{server}__*`` call attributes that server (named from ``_STATIC_PORTAL_CONFIGS`` when registered there). Falls back to the nominal portal when no tool ran, so citations are never empty. @@ -188,11 +200,15 @@ def add(source_id: str, name: str, url: str, quality: float) -> None: else: add(server_id, f"MCP server: {server_id}", "", 0.85) else: + served = str(entry.get("portal_id") or data_source) + cfg = portal_cfg if served == data_source else _registered_config(served) + if cfg is None: + served, cfg = data_source, portal_cfg add( - data_source, - portal_cfg.get("name", data_source), - portal_url, - portal_cfg.get("quality_score", 0.85), + served, + cfg.get("name", served), + portal_url if served == data_source else cfg.get("url", ""), + cfg.get("quality_score", 0.85), ) if not sources: @@ -283,6 +299,118 @@ def _all_portal_configs() -> dict[str, dict[str, Any]]: _HTML_TAG_RE = re.compile(r"<[^>]+>") +def _sql_action_missing(resp: httpx.Response) -> bool: + """CKAN's answer when DataStore SQL is disabled, not when a query is bad. + + A portal without ``ckan.datastore.sqlsearch.enabled`` (the Fair Store's + default) replies 400 "Action name not known: datastore_search_sql" — the + same status as a bad statement, so the status alone would keep it + retryable and the model would try again every turn. + """ + return resp.status_code == 400 and "Action name not known" in resp.text[:2000] + + +# Results that retrieved nothing (see the accounting in the tool loop). +_NOT_RETRIEVED_PREFIXES = ("Error", "HTTP ", "SQL error", "Tool unavailable:", "Rows elsewhere:") + + +def classify_tool_result(text: str) -> str: + """``retrieved``, ``error``, ``unavailable`` or ``redirected``. + + "Error" (not "Error:"): MCP failures read "Error calling ...". + notebook_generator._NOTHING_FETCHED uses the same prefixes. + """ + if text.startswith(LLMAnalysisAgent.TOOL_UNAVAILABLE_PREFIX): + return "unavailable" + if text.startswith(LLMAnalysisAgent.TOOL_REDIRECT_PREFIX): + return "redirected" + if text.startswith(("Error", "HTTP ", "SQL error")): + return "error" + return "retrieved" + + +def _normalize_load_args(tool_name: str, tool_input: Any) -> None: + """Repair ``fields`` sent as a JSON string (``'["NAME", "ACRES"]'``), which + DataStore rejects, before the call and the notebook both use it.""" + if tool_name != "load_resource_data" or not isinstance(tool_input, dict): + return + fields = tool_input.get("fields") + if isinstance(fields, str): + try: + parsed = json.loads(fields) + except ValueError: + parsed = [part.strip() for part in fields.split(",") if part.strip()] + tool_input["fields"] = [str(f) for f in parsed] if isinstance(parsed, list) else fields + + +_ID_ARGS = ("dataset_id", "resource_id") + + +def _is_tabular(resource: dict[str, Any]) -> bool: + """Whether a mirrored DCAT resource is a CSV/TSV file.""" + from data_concierge.data_layer.connectors.dcat import _TABULAR_MEDIA_TYPES + + mimetype = str(resource.get("mimetype") or "").split(";")[0].strip().lower() + fmt = str(resource.get("format") or "").strip().upper() + return mimetype in _TABULAR_MEDIA_TYPES or fmt in ("CSV", "TSV") + + +def _result_of(resp: httpx.Response) -> Any: + """An action response's ``result``, or ``None`` for anything unexpected.""" + if not resp.is_success: + return None + try: + body = resp.json() + except ValueError: + return None + return body.get("result") if isinstance(body, dict) else None + + +def _remember_portal(tool_input: Any, portal: str, portal_of_id: dict[str, str]) -> None: + """Record that ``portal`` answered a successful call about these IDs. + + ``portal`` is captured before the call: ``_execute_tool`` pops + ``portal_id`` from the input. + """ + if not isinstance(tool_input, dict): + return + for key in _ID_ARGS: + value = tool_input.get(key) + if isinstance(value, str) and value: + portal_of_id[value] = portal + + +def _stick_to_portal(tool_input: Any, portal_of_id: dict[str, str]) -> None: + """Send a call with no ``portal_id`` where its ID was last fetched from. + + An explicit ``portal_id`` always wins; this only fills one the model left + out for an ID it already used on another portal in this run. + """ + if not isinstance(tool_input, dict) or tool_input.get("portal_id"): + return + for key in _ID_ARGS: + portal = portal_of_id.get(str(tool_input.get(key) or "")) + if portal: + tool_input["portal_id"] = portal + return + + +def _int_param(params: dict, key: str, default: int, *, cap: int, floor: int = 1) -> int: + """A numeric tool argument, tolerating the model sending ``"100"``. + + A value below ``floor`` means "use the default", as ``int(x or default)`` + did before; CKAN loads pass ``floor=0`` because ``limit=0`` is a + count-only DataStore query. + """ + try: + value = int(params.get(key, default)) + except (TypeError, ValueError, OverflowError): + value = default + if value < floor: + value = default + return min(value, cap) + + def _summarize_error_body(text: str, limit: int = 200) -> str: """Condense an error body for the model. @@ -770,6 +898,49 @@ def _build_system_prompt( # -- tool execution ---------------------------------------------------- + async def _run_portal_tool( + self, + tool_name: str, + tool_input: Any, + portal_url: str, + *, + portal_of_id: dict[str, str] | None = None, + ) -> tuple[str, str, str | None]: + """One portal tool call, as both the analysis loop and the notebook editor run it. + + Repairs the model's arguments, fills a dropped ``portal_id`` from + ``portal_of_id`` (Fair Store chats), routes a Fair Store row load to + its origin, and resolves — before :meth:`_execute_tool` pops + ``portal_id`` — the portal the call reaches, so the notebook code and + the attribution name the portal that answered. Returns ``(result_text, + code, served_portal_id)``; the last is ``None`` for the primary portal. + """ + _normalize_load_args(tool_name, tool_input) + if portal_of_id is not None: + _stick_to_portal(tool_input, portal_of_id) + routed_from = await self._route_to_origin(tool_name, tool_input, portal_url) + code_url = self._tool_portal_url(tool_input, portal_url) + served: str | None = None + if isinstance(tool_input, dict) and tool_input.get("portal_id"): + # The registry ID of the URL actually called: canonical case, and an + # unknown portal_id names the portal _execute_tool fell back to. + served = _site_id_for_url(code_url) or str(tool_input["portal_id"]) + result_text = await self._execute_tool(tool_name, tool_input, portal_url) + if routed_from and classify_tool_result(result_text) == "retrieved": + # Only on success: a failed routed load keeps its error prefix. + result_text = f"{routed_from}\n{result_text}" + code = self._code_for_tool(tool_name, tool_input, code_url) + return result_text, code, served + + def _tool_portal_url(self, tool_input: Any, portal_url: str) -> str: + """The portal a tool call will reach: its ``portal_id``, else the primary.""" + override_id = tool_input.get("portal_id") if isinstance(tool_input, dict) else None + if override_id: + override_url = (self.get_portal_config(override_id) or {}).get("url") + if override_url: + return str(override_url) + return portal_url + async def _execute_tool( self, tool_name: str, @@ -856,7 +1027,7 @@ async def _tool_semantic_search(self, params: dict, site_id: str | None = None) querying. """ query = params.get("query", "") - n_results = min(params.get("n_results", 10), 20) + n_results = _int_param(params, "n_results", 10, cap=20) if not query: return "Error: query parameter is required" @@ -933,7 +1104,7 @@ def _get_dcat_client(self, portal_url: str, catalog_url: str | None = None) -> A async def _tool_dcat_search(self, dcat: Any, params: dict) -> str: query = params.get("query", "") - rows = min(int(params.get("rows", 10) or 10), 20) + rows = _int_param(params, "rows", 10, cap=20) matches = await dcat.search_datasets(query, limit=rows) if not matches: @@ -970,7 +1141,7 @@ async def _tool_dcat_dataset_info(self, dcat: Any, params: dict) -> str: ds = await dcat.get_dataset(dataset_id) if ds is None: return ( - f"Dataset '{dataset_id}' not found in the {dcat.portal_url} catalog. " + f"Error: dataset '{dataset_id}' not found in the {dcat.portal_url} catalog. " "Use search_datasets first and pass a Dataset ID from its results." ) @@ -1022,7 +1193,7 @@ async def _tool_dcat_load(self, dcat: Any, params: dict) -> str: get_dataset_info. """ resource_id = str(params.get("resource_id") or "").strip() - limit = min(int(params.get("limit", 100) or 100), 1000) + limit = _int_param(params, "limit", 100, cap=1000) if not resource_id: return "Error: resource_id is required (a dataset ID or a distribution URL)." @@ -1034,13 +1205,13 @@ async def _tool_dcat_load(self, dcat: Any, params: dict) -> str: ds = await dcat.get_dataset(resource_id) if ds is None: return ( - f"Dataset '{resource_id}' not found in the {dcat.portal_url} " + f"Error: dataset '{resource_id}' not found in the {dcat.portal_url} " "catalog. Use search_datasets to find a valid Dataset ID." ) tabular = ds.tabular_distributions if not tabular: return ( - f"Dataset '{ds.title}' has no tabular (CSV/TSV) distribution, " + f"Error: dataset '{ds.title}' has no tabular (CSV/TSV) distribution, " "so its rows cannot be loaded." ) target_url = tabular[0].best_url @@ -1053,7 +1224,7 @@ async def _tool_dcat_load(self, dcat: Any, params: dict) -> str: rows = result["rows"] if not rows: return ( - f"No rows returned from {target_url} " + f"Error: no rows returned from {target_url} " f"({result['bytes_read']:,} bytes read)." ) @@ -1083,7 +1254,7 @@ async def _tool_dcat_load(self, dcat: Any, params: dict) -> str: async def _tool_search(self, client: httpx.AsyncClient, params: dict) -> str: query = params.get("query", "") org = params.get("organization") - rows = min(params.get("rows", 10), 20) + rows = _int_param(params, "rows", 10, cap=20) search_params: dict[str, Any] = {"q": query, "rows": rows} if org: @@ -1141,6 +1312,20 @@ async def _tool_dataset_info(self, client: httpx.AsyncClient, params: dict) -> s if tags: lines.append(f"Tags: {', '.join(tags)}") + extras = { + str(e.get("key")): e.get("value") + for e in pkg.get("extras") or [] + if isinstance(e, dict) and e.get("key") + } + origin_id = str(extras.get("mirror_source_portal") or "") + origin = self._mirror_origin(origin_id) + if origin_id: + # A Fair Store record: metadata here, rows at the portal it mirrors. + label = origin["name"] if origin else extras.get("mirror_source_url") or origin_id + lines.append(f"Mirrored from: {label} (portal_id `{origin_id}`)") + if extras.get("qsv_description"): + lines.append(f"AI summary (qsv): {str(extras['qsv_description'])[:500]}") + resources = pkg.get("resources", []) lines.append(f"\nResources ({len(resources)}):") for i, res in enumerate(resources, 1): @@ -1150,15 +1335,71 @@ async def _tool_dataset_info(self, client: httpx.AsyncClient, params: dict) -> s lines.append(f" Format: {res.get('format', '?')} DataStore: {ds}") if res.get("description"): lines.append(f" {res['description'][:150]}") + route = self._origin_route(res, origin) + if route: + lines.append(f" Rows: {route}") return "\n".join(lines) + def _mirror_origin(self, origin_id: str) -> dict[str, Any] | None: + """The registered portal a Fair Store record was mirrored from.""" + if not origin_id: + return None + try: + from data_concierge.gateway import ckan_sites + + site = ckan_sites.get_site(origin_id) + except Exception: # pragma: no cover - registry failures are non-fatal + return None + if not site: + return None + return { + "id": origin_id, + "name": site.get("name") or origin_id, + "portal_type": _normalize_portal_type(site.get("portal_type")), + } + + @staticmethod + def _origin_route(resource: dict[str, Any], origin: dict[str, Any] | None) -> str: + """How to load a mirrored resource's rows from the portal that has them. + + A mirrored resource is a link: its rows are in the origin's DataStore + (CKAN, same resource UUID) or its downloadable file (DCAT), never in the + Fair Store's. The qsv tables the Fair Store generated itself carry no + ``mirror_source_portal`` and load here as usual. + """ + if not origin or resource.get("datastore_active"): + return "" + if str(resource.get("mirror_source_portal") or origin["id"]) != origin["id"]: + return "" + if origin["portal_type"] == _PORTAL_TYPE_DCAT: + url = str(resource.get("url") or "") + # Only a CSV/TSV file has rows to load: the mirror makes one resource + # per distribution, and a JSON, XML, KML or PDF one parsed as CSV + # "loads" garbage (DCATDistribution.is_tabular's rule). + if not url.lower().startswith(("http://", "https://")) or not _is_tabular(resource): + return "" + return ( + f"load_resource_data with portal_id=`{origin['id']}` and " + f"resource_id=`{url}`" + ) + if not resource.get("mirror_source_datastore_active"): + return "" + source_id = str(resource.get("mirror_source_id") or resource.get("id") or "") + return ( + f"load_resource_data with portal_id=`{origin['id']}` and " + f"resource_id=`{source_id}` (the origin's DataStore)" + ) + async def _tool_load(self, client: httpx.AsyncClient, params: dict) -> str: resource_id = params["resource_id"] - limit = min(params.get("limit", 100), 500) + limit = _int_param(params, "limit", 100, cap=500, floor=0) body: dict[str, Any] = {"resource_id": resource_id, "limit": limit} if params.get("offset"): - body["offset"] = params["offset"] + try: + body["offset"] = max(0, int(params["offset"])) + except (TypeError, ValueError): + pass if params.get("filters"): body["filters"] = params["filters"] if params.get("q"): @@ -1169,6 +1410,10 @@ async def _tool_load(self, client: httpx.AsyncClient, params: dict) -> str: body["fields"] = params["fields"] resp = await client.post("/api/3/action/datastore_search", json=body) + if resp.status_code == 404: + redirect = await self._mirrored_rows_hint(client, resource_id) + if redirect: + return redirect resp.raise_for_status() data = resp.json() @@ -1193,6 +1438,74 @@ async def _tool_load(self, client: httpx.AsyncClient, params: dict) -> str: lines.append(df.to_string(index=False, max_colwidth=40)) return "\n".join(lines) + async def _route_to_origin(self, tool_name: str, tool_input: Any, portal_url: str) -> str: + """Send a Fair Store row load to the portal the resource was mirrored from. + + The Fair Store holds metadata; a mirrored resource's rows are at its + origin (same UUID in a CKAN origin's DataStore, or the file a DCAT + origin publishes). Rewriting ``portal_id``/``resource_id`` here, before + the call, means the notebook records the portal that answered. Returns + a note for the model, or ``""`` when the load stays where it is. + """ + if tool_name != "load_resource_data" or not isinstance(tool_input, dict): + return "" + from data_concierge.gateway.fairstore import PORTAL_ID as FAIRSTORE_ID + + target_url = self._tool_portal_url(tool_input, portal_url) + if (tool_input.get("portal_id") or _site_id_for_url(target_url)) != FAIRSTORE_ID: + return "" + resource_id = str(tool_input.get("resource_id") or "") + if not resource_id or resource_id.lower().startswith(("http://", "https://")): + return "" + try: + client = await self._get_http_client(target_url) + resp = await client.get( + "/api/3/action/resource_show", params={"id": resource_id}, timeout=10.0 + ) + except httpx.HTTPError: + return "" + resource = _result_of(resp) + if not isinstance(resource, dict): + return "" + origin = self._mirror_origin(str(resource.get("mirror_source_portal") or "")) + if not self._origin_route(resource, origin) or origin is None: + return "" + if origin["portal_type"] == _PORTAL_TYPE_DCAT: + tool_input["resource_id"] = str(resource["url"]) + else: + tool_input["resource_id"] = str(resource.get("mirror_source_id") or resource_id) + tool_input["portal_id"] = origin["id"] + self.logger.info( + "Fair Store load routed to origin", resource_id=resource_id, origin=origin["id"] + ) + return ( + f"(The Fair Store mirrors this resource; its rows were loaded from " + f"{origin['name']}, portal_id `{origin['id']}`, resource " + f"`{tool_input['resource_id']}`.)" + ) + + async def _mirrored_rows_hint(self, client: httpx.AsyncClient, resource_id: str) -> str: + """Point a DataStore miss on a mirrored resource at the portal with the rows.""" + try: + # Short: this probe runs after every DataStore 404 on any CKAN portal. + resp = await client.get( + "/api/3/action/resource_show", params={"id": resource_id}, timeout=5.0 + ) + except httpx.HTTPError: + return "" + resource = _result_of(resp) + if not isinstance(resource, dict): + return "" + origin = self._mirror_origin(str(resource.get("mirror_source_portal") or "")) + route = self._origin_route(resource, origin) + if not route or not origin: + return "" + return ( + f"{self.TOOL_REDIRECT_PREFIX} resource `{resource_id}` is a mirror of " + f"{origin['name']}'s resource; its rows are not stored on this portal. " + f"Call {route}." + ) + def _sql_unavailable_status(self, portal_url: str) -> int | None: """Status that disabled SQL for this portal, or ``None`` if usable.""" entry = self._sql_disabled.get(portal_url) @@ -1223,6 +1536,10 @@ def _disable_sql(self, portal_url: str, status: int) -> None: # fetched nothing, and counting it as a failure would penalise the agent # for correctly routing around a portal-side outage. TOOL_UNAVAILABLE_PREFIX = "Tool unavailable:" + # A call that reached the right tool on the wrong portal: the rows exist, + # elsewhere. Excluded from the success rate like the above, but the tool is + # NOT withdrawn — the model needs it to follow the redirect. + TOOL_REDIRECT_PREFIX = "Rows elsewhere:" @staticmethod def _sql_unavailable_message(status: int, detail: str = "") -> str: @@ -1247,6 +1564,19 @@ def _sql_unavailable_message(status: int, detail: str = "") -> str: async def _tool_sql(self, client: httpx.AsyncClient, params: dict) -> str: portal_url = str(client.base_url).rstrip("/") + from data_concierge.gateway.fairstore import PORTAL_ID as FAIRSTORE_ID + + if _site_id_for_url(portal_url) == FAIRSTORE_ID: + # The Fair Store holds catalog metadata, no tables. Not a dead + # endpoint: tripping the breaker would withdraw SQL for the whole run + # (and 30 minutes of Fair Store chats) when the origin may serve it. + return ( + f"{self.TOOL_REDIRECT_PREFIX} the Fair Store holds catalog metadata " + "only, so run_sql_query has no tables here. Pass the portal_id the " + "dataset was mirrored from (get_dataset_info shows it as 'Mirrored " + "from'), or use load_resource_data, which loads from that portal." + ) + # Already known dead for this portal: answer without a round trip. known = self._sql_unavailable_status(portal_url) if known is not None: @@ -1270,7 +1600,7 @@ async def _tool_sql(self, client: httpx.AsyncClient, params: dict) -> str: "execution timeout. Narrow the query (add filters, " "aggregate, or reduce the result set)." ) - if resp.status_code in _SQL_ENDPOINT_DEAD_STATUSES: + if resp.status_code in _SQL_ENDPOINT_DEAD_STATUSES or _sql_action_missing(resp): # Endpoint-level failure: remember it so the rest of this query — # and the next one against the same portal — skips SQL entirely. self._disable_sql(portal_url, resp.status_code) @@ -1363,10 +1693,7 @@ def _code_for_dcat_tool(tool_name: str, tool_input: dict, portal_url: str) -> st # NOT CKAN's package_search, which does not exist here and would # 404 on every run. query = tool_input.get("query", "") - try: - rows = int(tool_input.get("n_results", 10)) - except (TypeError, ValueError): - rows = 10 + rows = _int_param(tool_input, "n_results", 10, cap=20) return ( "# Resource discovery. The concierge used a semantic (vector) search\n" "# over pre-indexed resource metadata here; the public equivalent for a\n" @@ -1378,7 +1705,7 @@ def _code_for_dcat_tool(tool_name: str, tool_input: dict, portal_url: str) -> st if tool_name == "search_datasets": query = tool_input.get("query", "") - rows = min(int(tool_input.get("rows", 10) or 10), 20) + rows = _int_param(tool_input, "rows", 10, cap=20) return "# Search the DCAT catalog (fetched once, searched locally)\n" + ( LLMAnalysisAgent._dcat_catalog_search_code(catalog_url, query, rows, helper) ) @@ -1410,7 +1737,7 @@ def _code_for_dcat_tool(tool_name: str, tool_input: dict, portal_url: str) -> st if tool_name == "load_resource_data": rid = str(tool_input.get("resource_id", "")) - limit = min(int(tool_input.get("limit", 100) or 100), 1000) + limit = _int_param(tool_input, "limit", 100, cap=1000) if rid.lower().startswith(("http://", "https://")): preamble = ( @@ -1465,10 +1792,7 @@ def _code_for_tool(tool_name: str, tool_input: dict, portal_url: str) -> str: if tool_name == "semantic_search_resources": q = tool_input.get("query", "") - try: - n = int(tool_input.get("n_results", 10)) - except (TypeError, ValueError): - n = 10 + n = _int_param(tool_input, "n_results", 10, cap=20) # The live tool searches a private Pinecone index, which the # published notebook cannot reach — importing it made every # notebook fail both in Colab and in the verifier's kernel. The @@ -1496,10 +1820,7 @@ def _code_for_tool(tool_name: str, tool_input: dict, portal_url: str) -> str: # the user later runs, so they cannot be interpolated raw: an org # like `x"}\nimport os\nos.system(...)` would inject code. repr the # Solr fq value and coerce rows to an int. - try: - rows = int(tool_input.get("rows", 10)) - except (TypeError, ValueError): - rows = 10 + rows = _int_param(tool_input, "rows", 10, cap=20) # repr the whole fq value so any quotes/newlines in org are escaped # into a proper Python string literal in the generated cell. fq = f', "fq": {f"organization:{org}"!r}' if org else "" @@ -1534,7 +1855,8 @@ def _code_for_tool(tool_name: str, tool_input: dict, portal_url: str) -> str: if tool_name == "load_resource_data": rid = tool_input["resource_id"] - limit = tool_input.get("limit", 100) + # What _tool_load sent, and never raw model text inside the code. + limit = _int_param(tool_input, "limit", 100, cap=500, floor=0) parts = [f'"resource_id": {rid!r}', f'"limit": {limit}'] if tool_input.get("filters"): parts.append(f'"filters": {json.dumps(tool_input["filters"])}') @@ -1839,6 +2161,12 @@ async def process(self, state: GraphState) -> GraphState: # noqa: C901 resource_record_counts: list[int] = [] all_tool_result_texts: list[str] = [] _prev_sql_errors: set[str] = set() # track SQL tool_use IDs that errored + # Fair Store chats reach origin portals through portal_id, and the + # model often drops it on the next call about the same dataset. + portal_of_id: dict[str, str] = {} + from data_concierge.gateway.fairstore import PORTAL_ID as _FAIRSTORE_ID + + sticky_portals = _FAIRSTORE_ID in (data_source, _site_id_for_url(portal_url)) for iteration in range(max_iterations): iter_start = time.monotonic() @@ -1930,14 +2258,17 @@ async def process(self, state: GraphState) -> GraphState: # noqa: C901 tool_start = time.monotonic() # Route to MCP or CKAN tool handler + called_portal: str | None = None if tool_name.startswith("mcp__"): result_text = await self._execute_mcp_tool(tool_name, tool_input) code = self._code_for_mcp_tool(tool_name, tool_input, result_text) else: - result_text = await self._execute_tool( - tool_name, tool_input, portal_url + result_text, code, called_portal = await self._run_portal_tool( + tool_name, + tool_input, + portal_url, + portal_of_id=portal_of_id if sticky_portals else None, ) - code = self._code_for_tool(tool_name, tool_input, portal_url) tool_elapsed_ms = round((time.monotonic() - tool_start) * 1000) @@ -1946,13 +2277,14 @@ async def process(self, state: GraphState) -> GraphState: # noqa: C901 # success nor a failure — it is excluded from the # success-rate denominator so the rate keeps measuring # only calls that actually attempted retrieval. - is_unavailable = result_text.startswith( - self.TOOL_UNAVAILABLE_PREFIX - ) - is_error = not is_unavailable and result_text.startswith( - ("Error:", "HTTP ", "SQL error") - ) - if is_unavailable: + outcome = classify_tool_result(result_text) + is_unavailable = outcome == "unavailable" + is_redirect = outcome == "redirected" + is_error = outcome == "error" + retrieved = outcome == "retrieved" + if is_redirect: + pass # fetched nothing, but the tool stays available + elif is_unavailable: # Withdraw the tool for the REST of this run, not # just the next one. Telling the model in prose not # to retry did not work — a measured trial had it @@ -1980,6 +2312,8 @@ async def process(self, state: GraphState) -> GraphState: # noqa: C901 successful_tool_calls += 1 all_tool_result_texts.append(result_text[:10000]) + if sticky_portals and called_portal and retrieved: + _remember_portal(tool_input, called_portal, portal_of_id) # Semantic search score if tool_name == "semantic_search_resources" and not is_error: @@ -1992,7 +2326,7 @@ async def process(self, state: GraphState) -> GraphState: # noqa: C901 max_semantic_score = score_val # Row/record counts from load_resource_data - if tool_name == "load_resource_data" and not is_error: + if tool_name == "load_resource_data" and retrieved: import re as _re total_match = _re.search(r"Total records:\s*([\d,]+)", result_text) @@ -2001,7 +2335,10 @@ async def process(self, state: GraphState) -> GraphState: # noqa: C901 try: count = int(total_match.group(1).replace(",", "")) resource_record_counts.append(count) - total_rows_loaded += min(count, tool_input.get("limit", 100)) + total_rows_loaded += min( + count, + _int_param(tool_input, "limit", 100, cap=500, floor=0), + ) except ValueError: pass else: @@ -2030,7 +2367,7 @@ async def process(self, state: GraphState) -> GraphState: # noqa: C901 resource_ids_seen.add(rid) # SQL queries also touch resources - if tool_name == "run_sql_query" and not is_error: + if tool_name == "run_sql_query" and retrieved: import re as _re uuids = _re.findall( @@ -2052,12 +2389,16 @@ async def process(self, state: GraphState) -> GraphState: # noqa: C901 ): resource_metadata_modified = mod_str + # The portal that served the call: a portal_id override or a + # Fair Store load routed to its origin, else the primary. + served_by = called_portal or data_source tool_calls_trace.append( { "agent": self.name, "action": tool_name, "tool_name": tool_name, "arguments": tool_input, + "portal_id": served_by, "result_preview": result_text[:800], "code": code, } @@ -2081,9 +2422,11 @@ async def process(self, state: GraphState) -> GraphState: # noqa: C901 "iteration": iteration + 1, "tool": tool_name, "tool_use_id": tool_id, - "source": _source_for_tool(tool_name, data_source), + "source": _source_for_tool(tool_name, served_by), "operation_type": _operation_type_for_tool(tool_name), - "status": "error" if is_error else "success", + "status": "success" if retrieved else "error", + # Why nothing was retrieved, when nothing was. + "outcome": outcome, "input": tool_input, "result": result_text[:10000], "result_chars": len(result_text), diff --git a/src/data_concierge/agents/notebook_editor.py b/src/data_concierge/agents/notebook_editor.py index 8a584e6..70ea768 100644 --- a/src/data_concierge/agents/notebook_editor.py +++ b/src/data_concierge/agents/notebook_editor.py @@ -32,8 +32,11 @@ AGENT_LOG_FORMAT_VERSION, TOOLS, _operation_type_for_tool, + _remember_portal, + _site_id_for_url, _source_for_tool, _utc_now, + classify_tool_result, get_llm_agent, ) from data_concierge.core.config import settings @@ -252,6 +255,13 @@ async def edit_notebook( # noqa: C901 - one linear tool loop, mirrors llm_agent portal_cfg = agent.get_portal_config(data_source) portal_url = portal_cfg.get("url", "") + # Fair Store chats: an ID fetched from an origin portal stays there (the + # model drops portal_id on follow-up calls), as in the analysis loop. + from data_concierge.gateway.fairstore import PORTAL_ID as FAIRSTORE_ID + + portal_of_id: dict[str, str] | None = ( + {} if FAIRSTORE_ID in (data_source, _site_id_for_url(portal_url)) else None + ) mcp_tools = agent._get_mcp_tools() from data_concierge.agents.llm_agent import _STATIC_PORTAL_CONFIGS @@ -408,16 +418,24 @@ async def edit_notebook( # noqa: C901 - one linear tool loop, mirrors llm_agent successful_calls += 1 result.tool_result_texts.append(result_text[:_MAX_TOOL_RESULT_CHARS]) else: - result_text = await agent._execute_tool(tool_name, tool_input, portal_url) - code = agent._code_for_tool(tool_name, tool_input, portal_url) - source = _source_for_tool(tool_name, data_source) + # The analysis loop's per-call steps: Fair Store routing, the + # served portal and the code URL resolved before the call. + result_text, code, served = await agent._run_portal_tool( + tool_name, tool_input, portal_url, portal_of_id=portal_of_id + ) + source = _source_for_tool(tool_name, served or data_source) operation_type = _operation_type_for_tool(tool_name) - is_error = result_text.startswith(("Error", "HTTP ", "SQL error")) - if is_error: + outcome = classify_tool_result(result_text) + is_error = outcome != "retrieved" + # A redirect or an absent capability fetched nothing and is + # neither a success nor a failure (as in the analysis loop). + if outcome == "error": failed_calls += 1 - else: + elif outcome == "retrieved": successful_calls += 1 result.tool_result_texts.append(result_text[:_MAX_TOOL_RESULT_CHARS]) + if portal_of_id is not None and served: + _remember_portal(tool_input, served, portal_of_id) result_for_model = result_text[:_MAX_TOOL_RESULT_CHARS] if code: diff --git a/src/data_concierge/agents/notebook_generator.py b/src/data_concierge/agents/notebook_generator.py index d16142d..cdc95e9 100644 --- a/src/data_concierge/agents/notebook_generator.py +++ b/src/data_concierge/agents/notebook_generator.py @@ -52,6 +52,12 @@ def _compile_safe_code_cell(code: str, label: str = "generated code") -> nbforma ) +# Tool results that fetched nothing, so their code would fail on replay: an +# error, a capability the portal lacks ("Tool unavailable:", e.g. DataStore SQL +# disabled) or rows that live on another portal ("Rows elsewhere:"). +_NOTHING_FETCHED = ("Error", "HTTP ", "SQL error", "Tool unavailable:", "Rows elsewhere:") + + class NotebookGeneratorAgent(BaseAgent): """Agent responsible for generating reproducible Google Colab notebooks. @@ -494,7 +500,7 @@ def _create_mcp_cells( # A failed exploratory call stays for provenance but must not be # an executable step (same rule as the CKAN trace cells). - failed = preview.startswith(("Error", "HTTP ", "SQL error")) + failed = preview.startswith(_NOTHING_FETCHED) if failed: md = md.replace( f"## Step {i}: {label}\n\n", @@ -584,8 +590,9 @@ def _create_cells_from_llm_trace( # column, bad SQL) is kept for provenance but must not be an # executable step: reproduced verbatim it fails again, so the # notebook never runs top to bottom and verification (#131) - # scores every answer as broken. - failed = preview.startswith(("Error", "HTTP ", "SQL error")) + # scores every answer as broken. So is a call that fetched nothing + # (see _NOTHING_FETCHED). + failed = preview.startswith(_NOTHING_FETCHED) # Markdown description label = self._format_action(action) diff --git a/src/data_concierge/core/config.py b/src/data_concierge/core/config.py index c6aa6b2..f41f786 100644 --- a/src/data_concierge/core/config.py +++ b/src/data_concierge/core/config.py @@ -155,6 +155,13 @@ class Settings(BaseSettings): # Provide via env / .env only — never commit a real key (issue #91). ckan_api_key: SecretStr = Field(default=SecretStr("")) + # Fair Store — the CKAN that mirrors every registered portal + # (scripts/populate_fairstore.py). Seeds for the admin panel's Fair Store + # settings; an admin's saved values (fairstore_settings.json) take precedence. + # Deliberately separate from CKAN_URL/CKAN_API_KEY, which name data.dathere.com. + fairstore_url: str = "" + fairstore_api_key: SecretStr = Field(default=SecretStr("")) + # WPRDC (Western PA Regional Data Center) - City of Pittsburgh open data wprdc_ckan_url: str = "https://data.wprdc.org" wprdc_organization: str = "city-of-pittsburgh" diff --git a/src/data_concierge/data_layer/qsv_profiling.py b/src/data_concierge/data_layer/qsv_profiling.py index 6c82ca0..161a840 100644 --- a/src/data_concierge/data_layer/qsv_profiling.py +++ b/src/data_concierge/data_layer/qsv_profiling.py @@ -33,6 +33,7 @@ from datetime import UTC, datetime from io import StringIO from pathlib import Path +from typing import Any import httpx @@ -308,24 +309,111 @@ def build_index( } +def csv_header(path: Path) -> list[str]: + """A CSV's column names, as the Fair Store mirror reads them.""" + csv.field_size_limit(max(csv.field_size_limit(), 64 * 1024 * 1024)) + with path.open(newline="", encoding="utf-8-sig", errors="replace") as handle: + return [name for name in next(csv.reader(handle), []) if name] + + +def scrubbed_json_bytes(path: Path) -> bytes: + """A JSON file's bytes with provider keys redacted, as the mirror publishes it. + + ``populate_fairstore`` scrubs the raw text of ``qsv_dict.json`` before + publishing it and hashes the result, so storing exactly that text keeps a + mirror run from storage and one from local files identical (a per-value + scrub keeps the word after the key, which the raw scrub removes, and every + describegpt resource would flip between the two). Should raw scrubbing + ever break the JSON — a key right before a closing quote — each string + value is scrubbed instead. + """ + from data_concierge.data_layer.onboard_index import _scrub_secrets + + text = path.read_text(encoding="utf-8") + scrubbed = _scrub_secrets(text) + try: + json.loads(scrubbed) + except ValueError: + scrubbed = json.dumps(_scrub_values(json.loads(text)), indent=2) + return scrubbed.encode("utf-8") + + +def _scrub_values(value: Any) -> Any: + from data_concierge.data_layer.onboard_index import _scrub_secrets + + if isinstance(value, str): + return _scrub_secrets(value) + if isinstance(value, list): + return [_scrub_values(item) for item in value] + if isinstance(value, dict): + return {key: _scrub_values(item) for key, item in value.items()} + return value + + +# The qsv outputs populate_fairstore reads from a dataset directory. +MIRRORED_QSV_OUTPUTS = ("qsv_dict.json", "qsv_stats.csv", "qsv_frequency.csv") + + +def recorded_headers(base_dir: Path, prefix: str) -> dict[str, list[str]]: + """Each downloaded CSV's header, keyed as its storage key would be.""" + headers: dict[str, list[str]] = {} + for csv_file in sorted(base_dir.rglob("*.csv")): + if csv_file.name.startswith("qsv_"): + continue + try: + header = csv_header(csv_file) + except (OSError, csv.Error): + continue # an unreadable download must not stop the sync + # Even an empty header: locally the mirror reads [] from such a file. + headers[f"{prefix}/{csv_file.relative_to(base_dir.parent)}"] = header + return headers + + +def sync_manifest(base_dir: Path, prefix: str) -> dict[str, Any]: + """What a mirror run from storage needs to know about the local directory. + + ``headers``: each downloaded CSV's header. ``dirs``: every dataset directory + and the qsv outputs it holds now — so a directory that is absent stays + absent, and an output deleted since an earlier sync is not staged. + """ + dirs = { + f"{prefix}/{directory.relative_to(base_dir.parent)}": sorted( + name for name in MIRRORED_QSV_OUTPUTS if (directory / name).is_file() + ) + for directory in sorted(p for p in base_dir.iterdir() if p.is_dir()) + } + return {"version": 1, "headers": recorded_headers(base_dir, prefix), "dirs": dirs} + + def sync_to_storage(base_dir: Path, site_id: str, prefix: str = "ckan_onboard") -> int: """Sync JSON and qsv CSV files from local disk to the unified storage backend. ``prefix`` is the storage key namespace. It defaults to ``ckan_onboard`` because ``onboard_index`` reads from there and existing deployments already have data under that prefix. + + The downloaded CSVs themselves are not synced; ``sync_manifest.json`` + records their headers and which dataset directories and qsv outputs exist + (:func:`sync_manifest`). ``populate_fairstore`` checks each qsv output + against the header of the file it describes, and runs from the storage + backend when the admin panel starts a mirror on Cloud Run. + + JSON is scrubbed of provider keys on the way: qsv describegpt records its + own command line, ``--api-key`` included, in ``qsv_dict.json``, + ``meta.json`` and ``index.json``. """ synced = 0 + manifest = sync_manifest(base_dir, prefix) for json_file in base_dir.rglob("*.json"): rel = json_file.relative_to(base_dir.parent) key = f"{prefix}/{rel}" - with open(json_file) as f: - data = json.load(f) - storage.write_json(key, data) + storage.write_bytes(key, scrubbed_json_bytes(json_file)) synced += 1 for csv_file in base_dir.rglob("qsv_*.csv"): rel = csv_file.relative_to(base_dir.parent) key = f"{prefix}/{rel}" storage.write_bytes(key, csv_file.read_bytes()) synced += 1 - return synced + # Last: a manifest must never list objects an interrupted sync did not write. + storage.write_json(f"{prefix}/{site_id}/sync_manifest.json", manifest) + return synced + 1 diff --git a/src/data_concierge/gateway/ckan_sites.py b/src/data_concierge/gateway/ckan_sites.py index d2d46db..a440e8a 100644 --- a/src/data_concierge/gateway/ckan_sites.py +++ b/src/data_concierge/gateway/ckan_sites.py @@ -230,12 +230,15 @@ def add_site( added_by: str = "admin", portal_type: str = PORTAL_TYPE_CKAN, catalog_url: str | None = None, + managed_by: str | None = None, ) -> dict[str, Any]: """Add a new portal. Auto-generates an ID from the name if not given. ``portal_type`` selects the access mechanism — ``ckan`` for a live action API, ``dcat`` for a catalog document. ``catalog_url`` is DCAT-only and optional: when omitted the standard catalog paths are probed under ``url``. + ``managed_by`` marks an entry another settings page owns (the Fair Store's + chat-source entry), which only that page changes or removes. Raises ``ValueError`` if ``url`` or ``name`` is empty, or if a site with the chosen ID already exists. @@ -270,6 +273,8 @@ def add_site( "added_by": added_by, "added_at": _now(), } + if managed_by: + entry["managed_by"] = managed_by sites.append(entry) _save(sites) logger.info( diff --git a/src/data_concierge/gateway/fairstore.py b/src/data_concierge/gateway/fairstore.py new file mode 100644 index 0000000..a576f54 --- /dev/null +++ b/src/data_concierge/gateway/fairstore.py @@ -0,0 +1,581 @@ +"""Verikan's link to the Fair Store. + +The Fair Store is a CKAN (2.11 + ckanext-hierarchy) that mirrors every +registered portal's catalog, with its organization hierarchy and the qsv +profiles onboarding produced (``scripts/populate_fairstore.py``). This module +is everything Verikan knows about it: + +**Connection settings.** The Fair Store's URL and a sysadmin API token, +seeded from ``FAIRSTORE_URL`` / ``FAIRSTORE_API_KEY`` and overridable from the +admin panel (``fairstore_settings.json`` through the storage backend — the same +pattern as the GitHub settings). The token is never returned to a client. + +**The token is bound to the origin it was saved for** (scheme, host and port). +Changing the URL to a different origin without supplying a new token drops the +saved token, and the environment token is used only while the URL is on the +environment URL's origin. A URL edit can therefore never send the sysadmin +token to another server, or another service on the same host. A URL that is +already a registered source portal is refused (:func:`conflicting_portal`). + +**Chat source.** With ``chat_source`` on, the Fair Store is registered in the +portal registry under :data:`PORTAL_ID`, so users can pick it in chat and the +agent can search every mirrored portal at once. + +**Status and site settings.** Live reads through the CKAN action API, and the +CKAN options a sysadmin can change at runtime (site title, about text, logo, +custom CSS) — the Fair Store's own ``/ckan-admin/config`` page. + +Mirror runs are started through :mod:`data_concierge.gateway.onboarding_jobs`, +which already runs one supervised job at a time. +""" + +from __future__ import annotations + +import asyncio +import ipaddress +from datetime import UTC, datetime +from typing import Any +from urllib.parse import urlparse + +import httpx + +from data_concierge.core.config import settings as app_settings +from data_concierge.core.logging import get_logger +from data_concierge.data_layer.connectors.ckan import CKANActionError, _action_result +from data_concierge.data_layer.storage import storage + +logger = get_logger(__name__) + +_KEY = "fairstore_settings.json" + +#: Registry ID of the Fair Store when it is offered as a chat source. +PORTAL_ID = "fairstore" + +DEFAULT_PORTAL_NAME = "Verikan Fair Store" +DEFAULT_PORTAL_DESCRIPTION = ( + "One catalog that mirrors every registered open data portal (WPRDC, " + "data.pa.gov, datHere) with each portal's organization hierarchy. Profiled " + "datasets carry qsv data dictionaries, summary statistics, frequency tables " + "and AI-written descriptions. Row data stays at each source portal." +) +DEFAULT_QUALITY_SCORE = 0.85 + +_TIMEOUT = httpx.Timeout(20.0, connect=10.0) +_MAX_TEXT = 20_000 + +# A token may travel over plain http only to this machine. +_LOCAL_HOSTS = frozenset({"localhost", "host.docker.internal"}) + +#: The CKAN options a sysadmin can change at runtime (``config_option_update``), +#: minus ``ckan.site_url`` (changing it breaks every link CKAN generates), +#: ``ckan.theme`` (2.11 ships one theme) and the logo-upload pseudo-options. +SITE_OPTIONS: dict[str, dict[str, str]] = { + "ckan.site_title": {"label": "Site title", "kind": "text"}, + "ckan.site_description": {"label": "Site description", "kind": "text"}, + "ckan.site_logo": {"label": "Logo URL", "kind": "text"}, + "ckan.site_intro_text": {"label": "Home page intro (Markdown)", "kind": "textarea"}, + "ckan.site_about": {"label": "About page (Markdown)", "kind": "textarea"}, + "ckan.site_custom_css": {"label": "Custom CSS", "kind": "code"}, +} + + +class FairStoreError(RuntimeError): + """The Fair Store could not be reached, or refused a call.""" + + +def _now() -> str: + return datetime.now(UTC).isoformat() + + +def _host(url: str) -> str: + return (urlparse(url).hostname or "").lower() + + +def _origin(url: str) -> str: + """``scheme://host:port`` — what a saved token is bound to (``""`` if unset).""" + if not url: + return "" + try: + parsed = httpx.URL(url) + except httpx.InvalidURL: + return "" + port = parsed.port or (443 if parsed.scheme == "https" else 80) + return f"{parsed.scheme}://{parsed.host.lower()}:{port}" + + +def conflicting_portal(url: str) -> str | None: + """A registered source portal already at this origin, if any. + + The Fair Store must be its own CKAN: pointed at a source portal (say + data.dathere.com), a mirror run would write every other portal into it. + """ + from data_concierge.gateway import ckan_sites + + origin = _origin(url) + if not origin: + return None + for site in ckan_sites.list_sites(): + if site.get("managed_by") == "fairstore": + continue + if _origin(str(site.get("url") or "")) == origin: + return str(site.get("id")) + return None + + +def normalize_url(raw: str) -> str: + """Validate and normalise a Fair Store base URL (``""`` stays unset). + + Raises ``ValueError`` for anything that is not a plain http(s) origin plus + optional path: credentials, a query or fragment, or plain http to a host + other than this machine (the sysadmin token would travel in clear text). + """ + url = (raw or "").strip().rstrip("/") + if not url: + return "" + if any(ord(ch) <= 32 or ord(ch) >= 127 for ch in url): + # urlparse silently drops tabs and newlines, and a non-ASCII host is a + # different string from the one the token is bound to: use punycode. + raise ValueError("The Fair Store URL may contain only printable ASCII (use punycode)") + try: + httpx.URL(url).port # noqa: B018 - validates the port as httpx will + except (httpx.InvalidURL, ValueError) as exc: + raise ValueError(f"The Fair Store URL is not valid: {exc}") from exc + parsed = urlparse(url) + if parsed.scheme not in ("http", "https") or not parsed.hostname: + raise ValueError("The Fair Store URL must start with https:// and name a host") + if parsed.username or parsed.password: + raise ValueError("Put the API token in the token field, not in the URL") + if parsed.query or parsed.fragment: + raise ValueError("The Fair Store URL must not carry a query or fragment") + if parsed.scheme == "http" and not _is_local(parsed.hostname): + raise ValueError( + "Use https:// for a Fair Store on another host — the sysadmin token " + "would otherwise be sent in clear text" + ) + return url + + +def _is_local(host: str) -> bool: + host = host.lower() + if host in _LOCAL_HOSTS or host.endswith(".localhost"): + return True + try: + return ipaddress.ip_address(host).is_loopback + except ValueError: + return False + + +def _env_url() -> str: + try: + return normalize_url(app_settings.fairstore_url) + except ValueError: + logger.warning("Ignoring an invalid FAIRSTORE_URL") + return "" + + +def _env_token() -> str: + return app_settings.fairstore_api_key.get_secret_value().strip() + + +class SettingsUnreadable(FairStoreError): + """The saved settings exist but could not be read (not the same as none).""" + + +def _read_saved(*, strict: bool = False) -> dict[str, Any]: + """The saved settings; ``{}`` when there are none. + + A read failure degrades to ``{}`` for display, but ``strict`` callers — a + save (which would otherwise overwrite the token) and the chat-source sync + (which would otherwise delete the portal entry) — get + :class:`SettingsUnreadable` instead. + """ + try: + saved = storage.read_json(_KEY) + except Exception as exc: # noqa: BLE001 - degrade to defaults + if strict: + raise SettingsUnreadable(f"Fair Store settings could not be read: {exc}") from exc + logger.warning("Failed to read Fair Store settings", error=str(exc)) + return {} + return saved if isinstance(saved, dict) else {} + + +def load_settings(*, strict: bool = False) -> dict[str, Any]: + """The effective settings, token included — never hand this to a client. + + ``token_source`` is ``"admin"``, ``"environment"`` or ``""``: the saved + token is used only for the origin (scheme, host and port) it was saved for, + and the environment's only for the environment URL's origin. + """ + saved = _read_saved(strict=strict) + url = saved.get("url") if isinstance(saved.get("url"), str) else None + url = url if url is not None else _env_url() + + token, source = "", "" + origin = _origin(url) + saved_token = str(saved.get("token") or "") + if saved_token and origin and saved.get("token_origin") == origin: + token, source = saved_token, "admin" + elif _env_token() and origin and origin == _origin(_env_url()): + token, source = _env_token(), "environment" + + mirror_sites = saved.get("mirror_sites") + return { + "url": url, + "url_source": "admin" if "url" in saved else ("environment" if url else ""), + "token": token, + "token_source": source, + "chat_source": bool(saved.get("chat_source", False)), + "portal_name": str(saved.get("portal_name") or DEFAULT_PORTAL_NAME), + "portal_description": str(saved.get("portal_description") or DEFAULT_PORTAL_DESCRIPTION), + "quality_score": float(saved.get("quality_score", DEFAULT_QUALITY_SCORE)), + "mirror_sites": [str(s) for s in mirror_sites] if isinstance(mirror_sites, list) else [], + "enrich": bool(saved.get("enrich", True)), + "updated_at": saved.get("updated_at"), + "updated_by": saved.get("updated_by"), + } + + +def public_settings(current: dict[str, Any] | None = None) -> dict[str, Any]: + """Settings safe to send to the admin UI: the token becomes set/unset flags. + + ``mirror_sites`` lists the chosen portals still registered; + ``mirror_sites_missing`` the ones since removed. + """ + from data_concierge.gateway import ckan_sites + + current = current if current is not None else load_settings() + token = current.get("token") or "" + known = set(ckan_sites.list_site_ids()) + chosen = current.get("mirror_sites") or [] + return { + **{k: v for k, v in current.items() if k != "token"}, + "mirror_sites": [s for s in chosen if s in known], + "mirror_sites_missing": [s for s in chosen if s not in known], + "conflicting_portal": conflicting_portal(current.get("url") or ""), + "token_set": bool(token), + "token_masked": f"••••{token[-4:]}" if len(token) > 12 else ("••••" if token else ""), + "configured": bool(current.get("url")), + "site_options": SITE_OPTIONS, + } + + +def save_settings(updates: dict[str, Any], *, updated_by: str = "admin") -> dict[str, Any]: + """Validate and persist an admin's changes; returns the public settings. + + ``token``: a non-empty string sets it, blank keeps it (the UI never sees + the current value), and ``clear_token`` removes it. A URL change to a + different origin (scheme, host or port) without a new token drops the saved + token — see the module docstring. Raises ``ValueError`` on invalid input, + including a URL that is already a registered source portal. + """ + saved = _read_saved(strict=True) + before = load_settings(strict=True) + + if "url" in updates and updates["url"] is not None: + saved["url"] = normalize_url(str(updates["url"])) + clash = conflicting_portal(saved["url"]) + if clash: + raise ValueError( + f"That URL is registered as the portal '{clash}' on the Data Portals page. " + "The Fair Store must be a separate CKAN, or a mirror run would write into " + f"that portal. (If '{clash}' is this Fair Store, delete that entry and use " + "'Offer the Fair Store as a data source in chat' instead.)" + ) + origin = _origin(saved.get("url") if "url" in saved else _env_url()) + + new_token = str(updates.get("token") or "").strip() + if new_token: + if any(ch.isspace() for ch in new_token) or len(new_token) > 4096: + raise ValueError("That does not look like a CKAN API token") + if not origin: + raise ValueError("Set the Fair Store URL before its API token") + saved["token"] = new_token + saved["token_origin"] = origin + elif updates.get("clear_token"): + saved.pop("token", None) + saved.pop("token_origin", None) + elif saved.get("token") and saved.get("token_origin") != origin: + saved.pop("token", None) + saved.pop("token_origin", None) + logger.info("Fair Store token dropped: the URL moved to another origin", origin=origin) + + if updates.get("chat_source") is not None: + saved["chat_source"] = bool(updates["chat_source"]) + for key in ("portal_name", "portal_description"): + if updates.get(key) is not None: + text = str(updates[key]).strip()[:2000] + if key == "portal_name" and not text: + raise ValueError("The chat source needs a name") + saved[key] = text + if updates.get("quality_score") is not None: + score = float(updates["quality_score"]) + if not 0.0 <= score <= 1.0: + raise ValueError("Quality score must be between 0 and 1") + saved["quality_score"] = score + if updates.get("mirror_sites") is not None: + from data_concierge.gateway import ckan_sites + + known = set(ckan_sites.list_site_ids()) + sites = [str(s).strip() for s in updates["mirror_sites"] if str(s).strip()] + unknown = [s for s in sites if s not in known] + if unknown: + raise ValueError(f"Unknown portal(s): {', '.join(sorted(unknown))}") + saved["mirror_sites"] = sorted(set(sites) - {PORTAL_ID}) + if updates.get("enrich") is not None: + saved["enrich"] = bool(updates["enrich"]) + + saved["updated_at"] = _now() + saved["updated_by"] = updated_by + storage.write_json(_KEY, saved) + + current = load_settings() + sync_chat_source(current) + logger.info( + "Fair Store settings saved", + updated_by=updated_by, + url_changed=before["url"] != current["url"], + token_changed=before["token"] != current["token"], + chat_source=current["chat_source"], + ) + return public_settings(current) + + +def sync_chat_source(current: dict[str, Any] | None = None) -> None: + """Make the portal registry's ``fairstore`` entry match the settings. + + Only an entry this module created (``managed_by == "fairstore"``) is ever + changed or removed, so a portal an admin registered by hand under the same + ID is left alone. + """ + from data_concierge.gateway import ckan_sites + + if current is None: + try: + current = load_settings(strict=True) + except SettingsUnreadable as exc: + # Unreadable is not "turned off": leave the registry as it is. + logger.warning("Skipping the Fair Store chat-source sync", error=str(exc)) + return + # Two instances syncing at once can both add the entry; add_site renames + # the second to fairstore-2, which the Data Portals page cannot delete. + for site in ckan_sites.list_sites(): + if site.get("managed_by") == "fairstore" and site.get("id") != PORTAL_ID: + ckan_sites.remove_site(str(site["id"])) + logger.info("Removed a duplicate Fair Store chat-source entry", site_id=site["id"]) + existing = ckan_sites.get_site(PORTAL_ID) + managed = bool(existing and existing.get("managed_by") == "fairstore") + wanted = bool(current.get("chat_source") and current.get("url")) + if wanted and conflicting_portal(current["url"]): + # An environment URL is not validated on save; never shadow a portal. + logger.warning("The Fair Store URL is a registered source portal; not offering it") + wanted = False + + if existing and not managed: + if wanted: + logger.warning( + "A hand-registered portal already uses the Fair Store's ID; leaving it", + site_id=PORTAL_ID, + ) + return + if not wanted: + if managed: + ckan_sites.remove_site(PORTAL_ID) + return + + fields = { + "url": current["url"], + "name": current["portal_name"], + "description": current["portal_description"], + "quality_score": current["quality_score"], + "portal_type": ckan_sites.PORTAL_TYPE_CKAN, + } + if managed: + if any(existing.get(key) != value for key, value in fields.items()): # type: ignore[union-attr] + ckan_sites.update_site(PORTAL_ID, fields) + else: + ckan_sites.add_site( + site_id=PORTAL_ID, + keywords=["fair store", "all portals", "catalog", "data dictionary", "qsv"], + added_by="fairstore", + managed_by="fairstore", + **fields, + ) + + +# ----------------------------------------------------------------------------- +# CKAN calls +# ----------------------------------------------------------------------------- + + +async def _call( + current: dict[str, Any], + action: str, + params: dict[str, Any] | None = None, + *, + auth: bool = False, +) -> Any: + """One CKAN action against the Fair Store (the token only when ``auth``).""" + url = current.get("url") or "" + if not url: + raise FairStoreError("The Fair Store URL is not configured") + headers = {"Content-Type": "application/json"} + if auth: + if not current.get("token"): + raise FairStoreError("No Fair Store API token is configured") + headers["Authorization"] = current["token"] + try: + async with httpx.AsyncClient(base_url=url, timeout=_TIMEOUT, headers=headers) as client: + response = await client.post(f"/api/3/action/{action}", json=params or {}) + except (httpx.HTTPError, httpx.InvalidURL) as exc: + raise FairStoreError(f"{action}: {type(exc).__name__}: {exc}") from exc + try: + return _action_result(action, response) + except CKANActionError as exc: + raise FairStoreError(str(exc)) from exc + + +async def _count(current: dict[str, Any], fq: str | None = None) -> int | None: + params: dict[str, Any] = {"rows": 0, "include_private": True} + if fq: + params["fq"] = fq + try: + result = await _call(current, "package_search", params, auth=bool(current.get("token"))) + except FairStoreError: + return None + return int(result.get("count", 0)) if isinstance(result, dict) else None + + +async def status(current: dict[str, Any] | None = None) -> dict[str, Any]: + """A live health and content summary for the admin panel. + + Never raises: an unreachable Fair Store, or a token CKAN refuses, is + reported in the result so the panel can say what is wrong. + """ + from data_concierge.gateway import ckan_sites + + current = current if current is not None else load_settings() + report: dict[str, Any] = { + "configured": bool(current.get("url")), + "url": current.get("url") or "", + "reachable": False, + "checked_at": _now(), + } + if not report["configured"]: + return report + + try: + info = await _call(current, "status_show") + except FairStoreError as exc: + report["error"] = str(exc) + return report + report["reachable"] = True + report["ckan_version"] = info.get("ckan_version") + report["site_title"] = info.get("site_title") + report["extensions"] = sorted(info.get("extensions") or []) + + # A sysadmin-only action tells a valid sysadmin token from any other. + if current.get("token"): + try: + await _call(current, "config_option_list", auth=True) + report["token"] = "sysadmin" + except FairStoreError as exc: + report["token"] = "rejected" + report["token_error"] = str(exc) + else: + report["token"] = "missing" + + sources = [s for s in ckan_sites.list_sites() if s.get("id") and s.get("id") != PORTAL_ID] + counts = await asyncio.gather( + _count(current), + *(_count(current, f'mirror_source_portal:"{s["id"]}"') for s in sources), + ) + report["datasets"] = counts[0] + report["by_portal"] = [ + {"site_id": s["id"], "name": s.get("name") or s["id"], "datasets": n} + for s, n in zip(sources, counts[1:], strict=True) + ] + for action, key in (("organization_list", "organizations"), ("group_list", "groups")): + try: + report[key] = len(await _call(current, action, {"all_fields": False})) + except FairStoreError: + report[key] = None + return report + + +def _refuse_conflict(current: dict[str, Any]) -> None: + clash = conflicting_portal(current.get("url") or "") + if clash: + raise FairStoreError( + f"The Fair Store URL is the registered portal '{clash}'; its site settings " + "are not edited from here." + ) + + +async def get_site_config(current: dict[str, Any] | None = None) -> dict[str, Any]: + """The Fair Store's runtime-editable site options (sysadmin token needed).""" + current = current if current is not None else load_settings() + _refuse_conflict(current) + editable = set(await _call(current, "config_option_list", auth=True) or []) + keys = [key for key in SITE_OPTIONS if key in editable] + values = await asyncio.gather( + *(_call(current, "config_option_show", {"key": key}, auth=True) for key in keys) + ) + return { + "options": [ + {**SITE_OPTIONS[key], "key": key, "value": "" if value is None else str(value)} + for key, value in zip(keys, values, strict=True) + ] + } + + +async def update_site_config( + values: dict[str, Any], current: dict[str, Any] | None = None +) -> dict[str, Any]: + """Change the Fair Store's site options; only :data:`SITE_OPTIONS` pass.""" + current = current if current is not None else load_settings() + _refuse_conflict(current) + unknown = sorted(set(values) - set(SITE_OPTIONS)) + if unknown: + raise ValueError(f"Not an editable Fair Store option: {', '.join(unknown)}") + payload = {} + for key, value in values.items(): + text = "" if value is None else str(value) + if len(text) > _MAX_TEXT: + raise ValueError(f"{SITE_OPTIONS[key]['label']} is longer than {_MAX_TEXT} characters") + payload[key] = text + if not payload: + raise ValueError("Nothing to update") + await _call(current, "config_option_update", payload, auth=True) + logger.info("Fair Store site options updated", keys=sorted(payload)) + return await get_site_config(current) + + +def public_info() -> dict[str, Any]: + """What any visitor may know: whether there is a Fair Store, and its address.""" + current = load_settings() + return { + "configured": bool(current.get("url")), + "url": current.get("url") or "", + "name": current.get("portal_name") or DEFAULT_PORTAL_NAME, + "chat_source": bool(current.get("chat_source") and current.get("url")), + "portal_id": PORTAL_ID, + } + + +__all__ = [ + "DEFAULT_PORTAL_DESCRIPTION", + "conflicting_portal", + "DEFAULT_PORTAL_NAME", + "PORTAL_ID", + "SITE_OPTIONS", + "FairStoreError", + "get_site_config", + "load_settings", + "normalize_url", + "public_info", + "public_settings", + "save_settings", + "status", + "sync_chat_source", + "update_site_config", +] diff --git a/src/data_concierge/gateway/onboarding_jobs.py b/src/data_concierge/gateway/onboarding_jobs.py index 6da6570..edba054 100644 --- a/src/data_concierge/gateway/onboarding_jobs.py +++ b/src/data_concierge/gateway/onboarding_jobs.py @@ -26,6 +26,24 @@ interleave their downloads. A second start is refused with the running job's ID rather than queued, so the caller always knows what is happening. +**Fair Store mirror runs share the runner.** ``populate_fairstore.py`` +(:func:`start_mirror_job`) runs under the same lock, log and records, tagged +``kind="fairstore_mirror"``: a mirror reads the onboarding output a concurrent +onboarding run would be rewriting. The Fair Store's sysadmin token reaches the +child through ``FAIRSTORE_API_KEY``, never argv. + +**Any instance can show a run.** Cloud Run sends each poll to whichever +instance it likes, and only one holds the child process. A running job +re-persists its record (log tail included) every :data:`HEARTBEAT_SECONDS`, so +another instance reports it as running from storage; a "running" record whose +heartbeat has gone stale is reported as interrupted. A cancel that lands on +another instance is written to ``.cancel.json`` and carried out by its +owner's next heartbeat. Log offsets are absolute line numbers +(``log_line_count``), so a poll answered by a lagging instance or after the +tail passes :data:`MAX_LOG_LINES` never repeats or stalls. Two starts landing +on different instances at the same moment can both run; the storage backend +has no create-only write to take a lease with. + **Cloud Run caveat.** A child process keeps running only while the instance is alive and scheduled. Without ``--no-cpu-throttling`` the CPU is throttled to near zero between requests and a background job crawls; with @@ -37,10 +55,14 @@ from __future__ import annotations import asyncio +import json import os import re import shutil import sys +import tempfile +import threading +import time import uuid from collections import deque from datetime import UTC, datetime @@ -61,6 +83,15 @@ MAX_LOG_LINES = 4000 MAX_JOB_RECORDS = 50 +KIND_ONBOARDING = "onboarding" +KIND_FAIRSTORE_MIRROR = "fairstore_mirror" + +# A running job's record is re-persisted this often (see the module docstring), +# and a stored "running" record whose heartbeat is older than STALE_AFTER has no +# live process anywhere. +HEARTBEAT_SECONDS = 10 +STALE_AFTER_SECONDS = 60 + STATUS_RUNNING = "running" STATUS_SUCCEEDED = "succeeded" STATUS_FAILED = "failed" @@ -72,6 +103,7 @@ "limit": (0, 5000), "concurrency": (1, 10), "max_mb": (1, 2000), + "mirror_limit": (1, 5000), } _NAME_RE = re.compile(r"^[A-Za-z0-9_.-]{1,64}$") _CONTROL_RE = re.compile(r"[\x00-\x1f\x7f]") @@ -113,6 +145,11 @@ def _is_noise(line: str) -> bool: _logs: dict[str, deque[str]] = {} _jobs: dict[str, dict[str, Any]] = {} _start_lock = asyncio.Lock() +# Orders a heartbeat's storage write against the final one: a heartbeat write +# still in flight in a worker thread when the job ends must not land after the +# final record (it would read "running" again, then "interrupted"). +_record_lock = threading.Lock() +_finished: set[str] = set() class JobError(RuntimeError): @@ -287,6 +324,72 @@ def build_command(site: dict[str, Any], options: dict[str, Any]) -> tuple[list[s return argv, script_name +def build_mirror_command( + options: dict[str, Any], + *, + target_url: str, + summary_path: str, +) -> tuple[list[str], list[dict[str, Any]]]: + """Build the validated argv for a Fair Store mirror run. + + ``options["site"]`` is ``"all"`` or a registered portal ID, or + ``options["sites"]`` a list of them — never the Fair Store's own + chat-source entry. Returns ``(argv, sites)``, ``sites`` empty for every + portal. The token is not an argument: :func:`start_mirror_job` passes it + through the environment. + """ + from data_concierge.gateway import ckan_sites + from data_concierge.gateway.fairstore import PORTAL_ID, normalize_url + + directory = scripts_dir() + if directory is None: + raise JobError("The mirror script is not available in this deployment") + script_path = directory / "populate_fairstore.py" + if not script_path.exists(): + raise JobError(f"populate_fairstore.py is missing from {directory}") + + try: + url = normalize_url(target_url) + except ValueError as exc: + raise JobError(str(exc)) from exc + if not url: + raise JobError("Set the Fair Store URL first") + + requested = options.get("sites") or [options.get("site") or "all"] + if not isinstance(requested, list) or len(requested) > 50: + raise JobError("Choose 'all' or up to 50 portals") + sites: list[dict[str, Any]] = [] + if requested != ["all"]: + for raw in dict.fromkeys(str(r).strip() for r in requested): + site = ckan_sites.get_site(raw) + if ( + site is None + or str(site.get("id", "")).lower() == PORTAL_ID + or site.get("managed_by") == "fairstore" + ): + raise JobError(f"'{raw}' is not a portal the Fair Store mirrors") + sites.append(site) + + argv: list[str] = [ + sys.executable, + str(script_path), + "--site", + ",".join(str(site["id"]) for site in sites) or "all", + "--target-url", + url, + "--summary-file", + summary_path, + ] + if options.get("apply"): + argv.append("--apply") + if options.get("enrich") is False: + argv.append("--no-enrich") + limit = _clean_int_option(options.get("limit"), "mirror_limit") + if limit is not None: + argv += ["--limit", str(limit)] + return argv, sites + + def _public(job: dict[str, Any]) -> dict[str, Any]: """Strip internal fields before a record leaves this module.""" return {k: v for k, v in job.items() if not k.startswith("_")} @@ -296,32 +399,77 @@ def _job_key(job_id: str) -> str: return f"{_STORAGE_PREFIX}/{job_id}.json" -def _persist(job: dict[str, Any]) -> None: - """Write a job record, and keep the index of recent jobs trimmed. +def _cancel_key(job_id: str) -> str: + # Separate from the record, which the owner's heartbeat keeps rewriting. + return f"{_STORAGE_PREFIX}/{job_id}.cancel.json" + - ``_public`` is not optional here: the live record carries the supervising - asyncio Task under ``_task``, which is not JSON-serializable, and writing - the raw dict silently loses every completed job. +def _write_record(record: dict[str, Any]) -> bool: + """Write a job record and the index; ``False`` if storage failed. + + Pass ``_public(job)``, never the live record: it carries the supervising + asyncio Task under ``_task``, which is not JSON-serializable, and writing it + silently lost every completed job. """ + job = record try: - storage.write_json(_job_key(job["id"]), _public(job)) + storage.write_json(_job_key(job["id"]), record) index = storage.read_json(_INDEX_KEY) or {} ids = [i for i in index.get("job_ids", []) if i != job["id"]] ids.insert(0, job["id"]) dropped = ids[MAX_JOB_RECORDS:] storage.write_json(_INDEX_KEY, {"job_ids": ids[:MAX_JOB_RECORDS]}) for old in dropped: - try: - storage.delete(_job_key(old)) - except Exception: # noqa: BLE001 - best-effort cleanup - pass + for key in (_job_key(old), _cancel_key(old)): + try: + storage.delete(key) + except Exception: # noqa: BLE001 - best-effort cleanup + pass except Exception as exc: # noqa: BLE001 - persistence must not kill a job logger.warning("Failed to persist onboarding job", job_id=job["id"], error=str(exc)) + return False + return True + + +FINAL_WRITE_ATTEMPTS = 5 + + +def _locked_write(record: dict[str, Any]) -> None: + with _record_lock: + _write_record(record) + + +def _final_write(record: dict[str, Any]) -> None: + """The terminal record: after any in-flight heartbeat, retried on failure. + + Other instances read a job only from storage, so a lost final write would + leave it "running" until the heartbeat went stale, then "interrupted" + without its result. Runs in a worker thread (blocking I/O and sleeps). + """ + with _record_lock: + _finished.add(record["id"]) + for attempt in range(FINAL_WRITE_ATTEMPTS): + if _write_record(record): + break + if attempt < FINAL_WRITE_ATTEMPTS - 1: + time.sleep(min(2**attempt, 10)) + else: + logger.error("Could not persist a finished job", job_id=record["id"]) + try: + storage.delete(_cancel_key(record["id"])) + except Exception: # noqa: BLE001 - best-effort cleanup + pass def _summary(job: dict[str, Any]) -> dict[str, Any]: - """A job record without its log, for list views.""" - return {k: v for k, v in job.items() if k != "log"} + """A job record without its log, for list views. + + Records written before mirror runs existed carry no ``kind``; they were + all onboarding runs. + """ + record = {k: v for k, v in job.items() if k != "log"} + record.setdefault("kind", KIND_ONBOARDING) + return record def running_job_id() -> str | None: @@ -331,6 +479,73 @@ def running_job_id() -> str | None: return None +def _seconds_since(timestamp: Any) -> float | None: + try: + then = datetime.fromisoformat(str(timestamp).replace("Z", "+00:00")) + except ValueError: + return None + if then.tzinfo is None: + then = then.replace(tzinfo=UTC) + return (datetime.now(UTC) - then).total_seconds() + + +def _settle_stored(stored: dict[str, Any]) -> dict[str, Any]: + """Decide what a stored record this instance holds no process for means. + + Running with a fresh heartbeat: another instance is running it. Running + without one: the instance that ran it is gone, so it was interrupted. + """ + if stored.get("status") == STATUS_RUNNING: + age = _seconds_since(stored.get("heartbeat_at")) + if age is not None and age < STALE_AFTER_SECONDS: + stored["running_elsewhere"] = True + else: + stored["status"] = STATUS_FAILED + stored["error"] = "interrupted (the server running it restarted)" + return stored + + +def _running_elsewhere() -> str | None: + """A job another instance is running, from its fresh heartbeat.""" + try: + index = storage.read_json(_INDEX_KEY) or {} + for job_id in index.get("job_ids", [])[:5]: + if job_id in _jobs: + continue + stored = storage.read_json(_job_key(job_id)) or {} + if _settle_stored(stored).get("running_elsewhere"): + return str(job_id) + except Exception as exc: # noqa: BLE001 - never block a start on a read error + logger.warning("Could not check for jobs on other instances", error=str(exc)) + return None + + +async def _heartbeat(job_id: str) -> None: + """Re-persist a running job, and honour a cancel recorded by another instance.""" + while True: + await asyncio.sleep(HEARTBEAT_SECONDS) + job = _jobs.get(job_id) + if job is None or job.get("status") != STATUS_RUNNING: + return + try: + cancel_requested = await asyncio.to_thread(storage.exists, _cancel_key(job_id)) + except Exception: # noqa: BLE001 + cancel_requested = False + if cancel_requested: + await cancel_job(job_id) + return + job["heartbeat_at"] = _now() + job["log"] = list(_logs.get(job_id, [])) + # Snapshot on the loop thread; only the storage write leaves it. + await asyncio.to_thread(_heartbeat_write, _public(job)) + + +def _heartbeat_write(record: dict[str, Any]) -> None: + with _record_lock: + if record["id"] not in _finished: + _write_record(record) + + async def _pump_output(job_id: str, process: asyncio.subprocess.Process) -> None: """Read the child's merged output into the job's log tail.""" log = _logs[job_id] @@ -348,9 +563,29 @@ async def _pump_output(job_id: str, process: asyncio.subprocess.Process) -> None job["log_line_count"] = job.get("log_line_count", 0) + 1 +def _collect_result(job: dict[str, Any]) -> None: + """Fold a run's JSON summary file into the record, then remove the file.""" + path = job.pop("_summary_path", None) + if not path: + return + try: + with open(path, encoding="utf-8") as handle: + job["result"] = json.load(handle) + except FileNotFoundError: + pass # the run failed before writing one; the log says why + except (OSError, ValueError) as exc: + logger.warning("Unreadable job summary", job_id=job.get("id"), error=str(exc)) + finally: + try: + os.unlink(path) + except OSError: + pass + + async def _supervise(job_id: str, process: asyncio.subprocess.Process) -> None: """Wait for the child, then record its outcome.""" job = _jobs[job_id] + beat = asyncio.create_task(_heartbeat(job_id)) try: await _pump_output(job_id, process) exit_code = await process.wait() @@ -369,20 +604,102 @@ async def _supervise(job_id: str, process: asyncio.subprocess.Process) -> None: except Exception as exc: # noqa: BLE001 job["status"] = STATUS_FAILED job["error"] = str(exc) - logger.error("Onboarding job crashed", job_id=job_id, error=str(exc)) + logger.error("Job crashed", job_id=job_id, error=str(exc)) finally: + beat.cancel() job["finished_at"] = _now() job["log"] = list(_logs.get(job_id, [])) + _collect_result(job) _processes.pop(job_id, None) - _persist(job) + # Snapshot on the loop thread; the write (and its lock wait) leaves it. + await asyncio.shield(asyncio.to_thread(_final_write, _public(job))) logger.info( - "Onboarding job finished", + "Job finished", job_id=job_id, status=job.get("status"), exit_code=job.get("exit_code"), ) +async def _spawn( + argv: list[str], + job: dict[str, Any], + *, + env_overrides: dict[str, str | None] | None = None, +) -> dict[str, Any]: + """Start ``argv`` as the one running job; the caller holds ``_start_lock``. + + ``env_overrides`` sets (or, with ``None``, removes) variables in the + child's copy of the server environment — how secrets reach it without + appearing in argv. + """ + job_id = job["id"] + env = dict(os.environ) + env["PYTHONUNBUFFERED"] = "1" + for key, value in (env_overrides or {}).items(): + if value is None: + env.pop(key, None) + else: + env[key] = value + + try: + process = await asyncio.create_subprocess_exec( + *argv, + stdout=asyncio.subprocess.PIPE, + stderr=asyncio.subprocess.STDOUT, + env=env, + cwd=str(scripts_dir().parent), # type: ignore[union-attr] + ) + except OSError as exc: + raise JobError(f"Could not start {job.get('script') or 'the job'}: {exc}") from exc + + _jobs[job_id] = job + _logs[job_id] = deque(maxlen=MAX_LOG_LINES) + _processes[job_id] = process + # Under the record lock: the previous job's final write may still be + # updating the index, and interleaving would drop one of the two. + await asyncio.to_thread(_locked_write, _public(job)) + + # Supervised in the background; the caller gets the job record now. + task = asyncio.create_task(_supervise(job_id, process)) + job["_task"] = task # not persisted (see _write_record / _public) + logger.info( + "Job started", + job_id=job_id, + kind=job.get("kind"), + site_id=job.get("site_id"), + script=job.get("script"), + started_by=job.get("started_by"), + ) + return _public(job) + + +def _refuse_if_running() -> None: + active = running_job_id() or _running_elsewhere() + if active: + raise JobError( + f"Job {active} is already running. Wait for it to finish or cancel it — " + "onboarding and Fair Store mirror runs share one slot, because a mirror " + "reads what an onboarding run writes." + ) + + +def _new_job(**fields: Any) -> dict[str, Any]: + now = _now() + return { + "id": uuid.uuid4().hex[:12], + "status": STATUS_RUNNING, + "started_at": now, + "heartbeat_at": now, + "finished_at": None, + "exit_code": None, + "error": None, + "log_line_count": 0, + "log": [], + **fields, + } + + async def start_job( site: dict[str, Any], options: dict[str, Any], @@ -391,71 +708,82 @@ async def start_job( ) -> dict[str, Any]: """Validate options, spawn the onboarding script, and return the job record.""" async with _start_lock: - active = running_job_id() - if active: - raise JobError( - f"Job {active} is already running. Wait for it to finish or cancel it " - "— concurrent runs would interleave downloads into the same directory." - ) - + _refuse_if_running() argv, script_name = build_command(site, options) - job_id = uuid.uuid4().hex[:12] - + job = _new_job( + kind=KIND_ONBOARDING, + site_id=site.get("id"), + site_name=site.get("name"), + portal_type=site.get("portal_type") or "ckan", + script=script_name, + # argv is safe to show: secrets go through the environment. + command=" ".join(argv[1:]), + options=dict(options), + started_by=started_by, + ) # The child inherits the server's environment so OPENROUTER_API_KEY and # PINECONE_API_KEY reach it without ever appearing in argv. - env = dict(os.environ) - env["PYTHONUNBUFFERED"] = "1" - - job: dict[str, Any] = { - "id": job_id, - "site_id": site.get("id"), - "site_name": site.get("name"), - "portal_type": site.get("portal_type") or "ckan", - "script": script_name, - # argv is safe to show: secrets go through the environment. - "command": " ".join(argv[1:]), - "options": dict(options), - "status": STATUS_RUNNING, - "started_by": started_by, - "started_at": _now(), - "finished_at": None, - "exit_code": None, - "error": None, - "log_line_count": 0, - "log": [], - } + return await _spawn(argv, job) - try: - process = await asyncio.create_subprocess_exec( - *argv, - stdout=asyncio.subprocess.PIPE, - stderr=asyncio.subprocess.STDOUT, - env=env, - cwd=str(scripts_dir().parent), # type: ignore[union-attr] - ) - except OSError as exc: - raise JobError(f"Could not start the onboarding script: {exc}") from exc - - _jobs[job_id] = job - _logs[job_id] = deque(maxlen=MAX_LOG_LINES) - _processes[job_id] = process - _persist(job) - - # Supervised in the background; the caller gets the job record now. - task = asyncio.create_task(_supervise(job_id, process)) - job["_task"] = task # not persisted (see _persist / _summary consumers) - logger.info( - "Onboarding job started", - job_id=job_id, - site_id=site.get("id"), - script=script_name, + +async def start_mirror_job( + options: dict[str, Any], + *, + target_url: str, + token: str, + started_by: str = "admin", +) -> dict[str, Any]: + """Start a Fair Store mirror run (``populate_fairstore.py``). + + ``options``: ``site`` (``"all"`` or a portal ID), ``apply`` (write; the + default is a dry run), ``enrich`` (Socrata enrichment, default on) and + ``limit`` (smoke test). Writing needs the sysadmin ``token``. + """ + if options.get("apply") and not token: + raise JobError("Writing to the Fair Store needs its sysadmin API token") + async with _start_lock: + _refuse_if_running() + fd, summary_path = tempfile.mkstemp(prefix="fairstore-mirror-", suffix=".json") + os.close(fd) + os.unlink(summary_path) # the script creates it; absent means no summary + argv, sites = build_mirror_command( + options, target_url=target_url, summary_path=summary_path + ) + if len(sites) == 1: + site_name = sites[0].get("name") or sites[0]["id"] + elif sites: + site_name = "Portals: " + ", ".join(str(site["id"]) for site in sites) + else: + site_name = "All portals" + job = _new_job( + kind=KIND_FAIRSTORE_MIRROR, + site_id=",".join(str(site["id"]) for site in sites) or "all", + site_name=site_name, + portal_type=(sites[0].get("portal_type") or "ckan") if len(sites) == 1 else None, + script="populate_fairstore.py", + command=" ".join(argv[1:]), + options={ + "site": ",".join(str(site["id"]) for site in sites) or "all", + "apply": bool(options.get("apply")), + "enrich": options.get("enrich") is not False, + "limit": options.get("limit") or None, + }, + target_url=target_url, started_by=started_by, + _summary_path=summary_path, + ) + return await _spawn( + argv, + job, + env_overrides={ + "FAIRSTORE_API_KEY": token or None, + "FAIRSTORE_URL": target_url, + }, ) - return _public(job) -def list_jobs(limit: int = 25) -> list[dict[str, Any]]: - """Recent jobs, newest first, without their logs.""" +def list_jobs(limit: int = 25, *, kind: str | None = None) -> list[dict[str, Any]]: + """Recent jobs, newest first, without their logs (optionally one ``kind``).""" records: list[dict[str, Any]] = [_summary(_public(j)) for j in _jobs.values()] seen = {r["id"] for r in records} @@ -467,16 +795,13 @@ def list_jobs(limit: int = 25) -> list[dict[str, Any]]: continue stored = storage.read_json(_job_key(job_id)) if stored: - # A job recorded as running that we have no process for did not - # survive a restart; reporting it as running would be a lie. - if stored.get("status") == STATUS_RUNNING: - stored["status"] = STATUS_FAILED - stored["error"] = "interrupted (server restarted during the run)" - records.append(_summary(stored)) + records.append(_summary(_settle_stored(stored))) seen.add(job_id) except Exception as exc: # noqa: BLE001 logger.warning("Failed to read onboarding job index", error=str(exc)) + if kind: + records = [r for r in records if r.get("kind") == kind] records.sort(key=lambda r: r.get("started_at") or "", reverse=True) return records[:limit] @@ -484,8 +809,9 @@ def list_jobs(limit: int = 25) -> list[dict[str, Any]]: def get_job(job_id: str, *, log_offset: int = 0) -> dict[str, Any] | None: """One job with a slice of its log. - ``log_offset`` is a line index into the retained tail, so a UI can poll for - just what it has not shown yet instead of refetching the whole log. + ``log_offset`` is an absolute line number (the previous response's + ``log_next_offset``), so a UI can poll for just what it has not shown yet, + from any instance, however long the log grows. """ job = _jobs.get(job_id) if job is not None: @@ -495,34 +821,53 @@ def get_job(job_id: str, *, log_offset: int = 0) -> dict[str, Any] | None: stored = storage.read_json(_job_key(job_id)) if not stored: return None - if stored.get("status") == STATUS_RUNNING: - stored["status"] = STATUS_FAILED - stored["error"] = "interrupted (server restarted during the run)" - record = _summary(stored) + record = _summary(_settle_stored(stored)) lines = list(stored.get("log", [])) - offset = max(0, min(int(log_offset or 0), len(lines))) - record["log"] = lines[offset:] + # Absolute line numbers: ``lines`` is the tail ending at log_line_count. + total = max(int(record.get("log_line_count") or 0), len(lines)) + first = total - len(lines) + offset = max(0, int(log_offset or 0)) + record["log"] = lines[max(0, offset - first) :] if offset < total else [] record["log_offset"] = offset - record["log_next_offset"] = len(lines) - record["log_truncated"] = record.get("log_line_count", 0) > len(lines) + # Never backwards: a poll answered from a lagging snapshot returns nothing. + record["log_next_offset"] = max(offset, total) + record["log_truncated"] = offset < first return record async def cancel_job(job_id: str) -> bool: - """Terminate a running job. Returns ``False`` if it was not running.""" + """Terminate a running job. Returns ``False`` if it was not running. + + For a job another instance is running, a ``.cancel.json`` key is + written; its owner's next heartbeat terminates it. + """ process = _processes.get(job_id) job = _jobs.get(job_id) - if process is None or job is None or job.get("status") != STATUS_RUNNING: + if process is None or job is None: + stored = storage.read_json(_job_key(job_id)) or {} + if job is None and _settle_stored(dict(stored)).get("running_elsewhere"): + storage.write_json(_cancel_key(job_id), {"requested_at": _now()}) + # The owner may have finished (and cleared cancels) meanwhile. + again = storage.read_json(_job_key(job_id)) or {} + if not _settle_stored(dict(again)).get("running_elsewhere"): + storage.delete(_cancel_key(job_id)) + return False + logger.info("Cancel requested for a job on another instance", job_id=job_id) + return True return False + if job.get("status") != STATUS_RUNNING or process.returncode is not None: + return False # already exited; _supervise records how # Mark first so _supervise does not overwrite the status with "failed" # when the child exits non-zero because we killed it. + previous = (job.get("status"), job.get("error")) job["status"] = STATUS_CANCELLED job["error"] = "cancelled by admin" try: process.terminate() except ProcessLookupError: + job["status"], job["error"] = previous return False try: @@ -532,11 +877,13 @@ async def cancel_job(job_id: str) -> bool: process.kill() except ProcessLookupError: pass - logger.info("Onboarding job cancelled", job_id=job_id) + logger.info("Job cancelled", job_id=job_id) return True __all__ = [ + "KIND_FAIRSTORE_MIRROR", + "KIND_ONBOARDING", "MAX_LOG_LINES", "STATUS_CANCELLED", "STATUS_FAILED", @@ -545,6 +892,7 @@ async def cancel_job(job_id: str) -> bool: "TERMINAL_STATUSES", "JobError", "build_command", + "build_mirror_command", "cancel_job", "get_job", "list_jobs", @@ -552,4 +900,5 @@ async def cancel_job(job_id: str) -> bool: "runtime_warnings", "scripts_dir", "start_job", + "start_mirror_job", ] diff --git a/src/data_concierge/gateway/router.py b/src/data_concierge/gateway/router.py index ca8aac9..653bbee 100644 --- a/src/data_concierge/gateway/router.py +++ b/src/data_concierge/gateway/router.py @@ -820,6 +820,19 @@ class UpdateCkanSiteRequest(BaseModel): keywords: list[str] | None = None +def _refuse_managed_site(site_id: str) -> None: + """Another settings page owns this entry (the Fair Store's chat source).""" + site = ckan_sites_store.get_site(site_id) + if site and site.get("managed_by") == "fairstore": + raise HTTPException( + status_code=status.HTTP_409_CONFLICT, + detail=( + "This entry follows the Fair Store settings. Change it, or turn off " + "'Offer the Fair Store as a data source in chat', on the Fair Store page." + ), + ) + + @router.get("/admin/ckan-sites") async def list_ckan_sites( _admin: dict[str, Any] = Depends(require_admin), @@ -860,6 +873,7 @@ async def update_ckan_site( _admin: dict[str, Any] = Depends(require_admin), ) -> dict[str, Any]: """Update an existing CKAN site entry.""" + _refuse_managed_site(site_id) updates = request.model_dump(exclude_none=True) entry = ckan_sites_store.update_site(site_id, updates) if not entry: @@ -876,6 +890,7 @@ async def delete_ckan_site( _admin: dict[str, Any] = Depends(require_admin), ) -> dict[str, Any]: """Remove a CKAN site from the registry.""" + _refuse_managed_site(site_id) if not ckan_sites_store.remove_site(site_id): raise HTTPException( status_code=status.HTTP_404_NOT_FOUND, @@ -921,7 +936,7 @@ async def list_onboarding_jobs( """Recent onboarding runs, plus anything about this host that would break one.""" from data_concierge.gateway import onboarding_jobs - jobs = onboarding_jobs.list_jobs() + jobs = onboarding_jobs.list_jobs(kind=onboarding_jobs.KIND_ONBOARDING) return { "count": len(jobs), "jobs": jobs, @@ -944,6 +959,14 @@ async def start_onboarding_job( status_code=status.HTTP_404_NOT_FOUND, detail=f"Portal '{request.site_id}' is not registered", ) + if site.get("managed_by") == "fairstore": + raise HTTPException( + status_code=status.HTTP_409_CONFLICT, + detail=( + "The Fair Store is not onboarded itself: its qsv profiles come from " + "onboarding the portals it mirrors." + ), + ) options = request.model_dump(exclude={"site_id"}) try: @@ -990,6 +1013,195 @@ async def cancel_onboarding_job( return {"message": f"Job '{job_id}' cancelled"} +# ============================================================================= +# Fair Store (admin-only) — the CKAN that mirrors every registered portal +# ============================================================================= + + +class FairStoreSettingsRequest(BaseModel): + """Changes to the Fair Store link. Omitted fields keep their value. + + ``token`` follows the GitHub settings' "blank = keep" rule, because the UI + never sees the current token; ``clear_token`` removes it. + """ + + url: str | None = Field(default=None, max_length=400) + token: str | None = Field(default=None, max_length=4096) + clear_token: bool = False + chat_source: bool | None = Field( + default=None, description="Offer the Fair Store as a data source in chat" + ) + portal_name: str | None = Field(default=None, max_length=200) + portal_description: str | None = Field(default=None, max_length=2000) + quality_score: float | None = Field(default=None, ge=0.0, le=1.0) + mirror_sites: list[str] | None = Field( + default=None, description="Portals a mirror run of 'all' covers; empty = every portal" + ) + enrich: bool | None = Field(default=None, description="Socrata enrichment for DCAT portals") + + +class FairStoreSiteConfigRequest(BaseModel): + """New values for the Fair Store's runtime site options (``ckan.site_*``).""" + + options: dict[str, str | None] + + +class StartMirrorJobRequest(BaseModel): + """A Fair Store mirror run. The default is a dry run of every portal.""" + + site: str = Field(default="all", min_length=1, max_length=80) + apply: bool = Field(default=False, description="Write to the Fair Store (else dry run)") + enrich: bool | None = Field(default=None, description="Defaults to the saved setting") + limit: int | None = Field(default=None, ge=1, le=5000, description="Smoke test only") + + +@router.get("/admin/fairstore") +async def get_fairstore_settings( + _admin: dict[str, Any] = Depends(require_admin), +) -> dict[str, Any]: + """The Fair Store link's settings (the token only as set/unset).""" + from data_concierge.gateway import fairstore + + return {"settings": fairstore.public_settings()} + + +@router.post("/admin/fairstore") +async def update_fairstore_settings( + request: FairStoreSettingsRequest, + admin_user: dict[str, Any] = Depends(require_admin), +) -> dict[str, Any]: + """Save the Fair Store link's settings (and its chat-source registration).""" + from data_concierge.gateway import fairstore + + try: + saved = fairstore.save_settings( + request.model_dump(exclude_none=True), + updated_by=admin_user.get("user", "admin"), + ) + except ValueError as exc: + raise HTTPException(status_code=status.HTTP_400_BAD_REQUEST, detail=str(exc)) from exc + except fairstore.SettingsUnreadable as exc: + raise HTTPException( + status_code=status.HTTP_503_SERVICE_UNAVAILABLE, detail=str(exc) + ) from exc + return {"message": "Fair Store settings saved", "settings": saved} + + +@router.get("/admin/fairstore/status") +async def get_fairstore_status( + _admin: dict[str, Any] = Depends(require_admin), +) -> dict[str, Any]: + """Live health, token check and dataset counts per source portal.""" + from data_concierge.gateway import fairstore, onboarding_jobs + + report = await fairstore.status() + mirror_runs = onboarding_jobs.list_jobs(limit=1, kind=onboarding_jobs.KIND_FAIRSTORE_MIRROR) + report["last_mirror_run"] = mirror_runs[0] if mirror_runs else None + return report + + +@router.get("/admin/fairstore/site-config") +async def get_fairstore_site_config( + _admin: dict[str, Any] = Depends(require_admin), +) -> dict[str, Any]: + """The Fair Store's runtime site options, read live (needs the token).""" + from data_concierge.gateway import fairstore + + try: + return await fairstore.get_site_config() + except fairstore.FairStoreError as exc: + raise HTTPException(status_code=status.HTTP_502_BAD_GATEWAY, detail=str(exc)) from exc + + +@router.post("/admin/fairstore/site-config") +async def update_fairstore_site_config( + request: FairStoreSiteConfigRequest, + admin_user: dict[str, Any] = Depends(require_admin), +) -> dict[str, Any]: + """Change the Fair Store's site title, about text, logo or custom CSS.""" + from data_concierge.gateway import fairstore + + try: + config = await fairstore.update_site_config(request.options) + except ValueError as exc: + raise HTTPException(status_code=status.HTTP_400_BAD_REQUEST, detail=str(exc)) from exc + except fairstore.FairStoreError as exc: + raise HTTPException(status_code=status.HTTP_502_BAD_GATEWAY, detail=str(exc)) from exc + logger.info( + "Fair Store site options changed from the admin panel", + admin=admin_user.get("user"), + keys=sorted(request.options), + ) + return {"message": "Fair Store site settings saved", **config} + + +@router.get("/admin/fairstore/mirror-jobs") +async def list_fairstore_mirror_jobs( + _admin: dict[str, Any] = Depends(require_admin), +) -> dict[str, Any]: + """Recent mirror runs. Logs and cancel use ``/admin/onboarding-jobs/{id}``.""" + from data_concierge.gateway import onboarding_jobs + + jobs = onboarding_jobs.list_jobs(kind=onboarding_jobs.KIND_FAIRSTORE_MIRROR) + warnings = [ + w for w in onboarding_jobs.runtime_warnings() if "qsv" not in w and "OPENROUTER" not in w + ] + return { + "count": len(jobs), + "jobs": jobs, + "running_job_id": onboarding_jobs.running_job_id(), + "warnings": warnings, + } + + +@router.post("/admin/fairstore/mirror-jobs") +async def start_fairstore_mirror_job( + request: StartMirrorJobRequest, + admin_user: dict[str, Any] = Depends(require_admin), +) -> dict[str, Any]: + """Start a mirror run into the Fair Store (a dry run unless ``apply``).""" + from data_concierge.gateway import fairstore, onboarding_jobs + + current = fairstore.load_settings() + options = request.model_dump() + if options.get("enrich") is None: + options["enrich"] = current["enrich"] + clash = fairstore.conflicting_portal(current["url"]) + if clash: + raise HTTPException( + status_code=status.HTTP_409_CONFLICT, + detail=( + f"The Fair Store URL is the source portal '{clash}'; a mirror would " + "write into it. Point the Fair Store at its own CKAN first." + ), + ) + chosen = [s for s in current["mirror_sites"] if ckan_sites_store.get_site(s)] + if request.site == "all" and current["mirror_sites"] and not chosen: + # Widening to every portal would not be what the saved subset asked for. + raise HTTPException( + status_code=status.HTTP_409_CONFLICT, + detail=( + "None of the portals chosen under Mirror defaults is registered any " + "more. Choose portals again (or none, for every portal) and save." + ), + ) + if request.site == "all" and chosen: + # "All" means the portals the settings choose, when they choose some + # (minus any since removed from the registry). + options["sites"] = chosen + try: + job = await onboarding_jobs.start_mirror_job( + options, + target_url=current["url"], + token=current["token"], + started_by=admin_user.get("user", "admin"), + ) + except onboarding_jobs.JobError as exc: + raise HTTPException(status_code=status.HTTP_409_CONFLICT, detail=str(exc)) from exc + mode = "Mirror" if request.apply else "Dry run" + return {"message": f"{mode} started for {job['site_name']}", "job": job} + + # Public read-only endpoint — lets the UI populate a picker without admin auth. @router.get("/ckan-sites") async def list_ckan_sites_public() -> dict[str, Any]: @@ -4432,6 +4644,14 @@ def _validate_site(site: str) -> str: return site +@router.get("/fairstore/info") +async def fairstore_info() -> dict[str, Any]: + """Whether a Fair Store is linked, and its public address (no secrets).""" + from data_concierge.gateway import fairstore + + return fairstore.public_info() + + @router.get("/fairstore/search") async def fairstore_search(q: str, site: str = "wprdc", limit: int = 10) -> dict[str, Any]: """Search the Fair Store data dictionary by keyword. diff --git a/src/data_concierge/ui/static/js/admin.js b/src/data_concierge/ui/static/js/admin.js index 07a4280..fc19fa3 100644 --- a/src/data_concierge/ui/static/js/admin.js +++ b/src/data_concierge/ui/static/js/admin.js @@ -197,6 +197,19 @@ function setupEventListeners() { }; wireRevealToggle('ghTokenToggle', 'ghToken', 'token'); wireRevealToggle('ghWebhookSecretToggle', 'ghWebhookSecret', 'webhook secret'); + wireRevealToggle('fsTokenToggle', 'fsToken', 'token'); + + // Fair Store + document.getElementById('fairstore-tab').addEventListener('shown.bs.tab', loadFairStore); + document.getElementById('fsRefreshBtn').addEventListener('click', loadFairStore); + document.getElementById('fsSettingsForm').addEventListener('submit', saveFairStoreSettings); + document.getElementById('fsMirrorForm').addEventListener('submit', startFairStoreMirror); + document.getElementById('fsMirrorApply').addEventListener('change', syncFairStoreMirrorLabel); + document.getElementById('fsChatSource').addEventListener('change', syncFairStoreChatFields); + document.getElementById('fsLogCloseBtn').addEventListener('click', closeFairStoreLog); + document.getElementById('fsCancelBtn').addEventListener('click', cancelFairStoreRun); + document.getElementById('fsSiteForm').addEventListener('submit', saveFairStoreSiteConfig); + document.getElementById('fsSiteReloadBtn').addEventListener('click', loadFairStoreSiteConfig); document.getElementById('ghPauseBtn').addEventListener('click', openPauseModal); document.getElementById('ghPauseConfirmBtn').addEventListener('click', confirmPauseToggle); } @@ -2690,6 +2703,9 @@ function renderCkanSiteCard(site) { ? `${escapeHtml(site.organization)}` : ''; const isDefault = site.added_by === 'default'; + // The Fair Store's chat-source entry follows its settings page; onboarding + // it would re-download files the source portals already profile. + const isManaged = site.managed_by === 'fairstore'; // Portal type drives which tools work against this site, so it is shown // on the card rather than hidden behind an edit form. const portalType = (site.portal_type === 'dcat') ? 'dcat' : 'ckan'; @@ -2710,6 +2726,7 @@ function renderCkanSiteCard(site) { ${typeBadge} ${orgBadge} ${isDefault ? 'built-in' : ''} + ${isManaged ? 'managed on the Fair Store page' : ''}

@@ -3023,6 +3040,7 @@ async function pollOnboardingLog() { try { const response = await adminFetch(`${ONBOARD_API}/${encodeURIComponent(jobId)}?log_offset=${_onboardLogOffset}`); const { job } = await response.json(); + if (jobId !== _onboardLogJobId) return; // switched jobs while waiting const pre = document.getElementById('onboardLog'); if (job.log && job.log.length) { @@ -3064,8 +3082,486 @@ async function cancelOnboardingJob() { try { await adminFetch(`${ONBOARD_API}/${encodeURIComponent(_onboardLogJobId)}/cancel`, { method: 'POST' }); showToast('Run cancelled', 'success'); + stopOnboardingPoll(); // one poll chain only, or log lines repeat pollOnboardingLog(); } catch (error) { showToast(`Could not cancel: ${error.message}`, 'danger'); } } + + +// ============================================================================= +// Fair Store — the CKAN that mirrors every registered portal +// ============================================================================= + +const FAIRSTORE_API = `${API_BASE}/admin/fairstore`; + +let _fsSettings = null; +let _fsLogJobId = null; +let _fsLogOffset = 0; +let _fsPollTimer = null; + +function _fsAlert(targetId, type, msg) { + const box = document.getElementById(targetId); + if (!box) return; + box.innerHTML = msg + ? `` + : ''; +} + +async function loadFairStore() { + try { + const [settingsRes, sitesRes] = await Promise.all([ + adminFetch(FAIRSTORE_API), + adminFetch(CKAN_SITES_API), + ]); + _fsSettings = (await settingsRes.json()).settings; + const sites = ((await sitesRes.json()).sites || []).filter(site => site.id !== 'fairstore'); + renderFairStoreSettings(_fsSettings, sites); + } catch (error) { + _fsAlert('fsAlert', 'danger', `Could not load the Fair Store settings: ${error.message}`); + return; + } + loadFairStoreStatus(); + loadFairStoreJobs(); + if (_fsSettings.configured && _fsSettings.token_set) { + loadFairStoreSiteConfig(); + } else { + document.getElementById('fsSiteFields').textContent = + 'Set the URL and a sysadmin token to edit these.'; + document.getElementById('fsSiteActions').classList.add('d-none'); + } +} + +function renderFairStoreSettings(settings, sites) { + document.getElementById('fsUrl').value = settings.url || ''; + document.getElementById('fsToken').value = ''; + document.getElementById('fsClearToken').checked = false; + document.getElementById('fsClearTokenWrap').classList.toggle('d-none', settings.token_source !== 'admin'); + const tokenHint = document.getElementById('fsTokenHint'); + if (settings.token_set) { + const from = settings.token_source === 'environment' ? ' (from the FAIRSTORE_API_KEY environment variable)' : ''; + tokenHint.textContent = `A token ending ${settings.token_masked.replace(/•/g, '')} is set${from}. Leave blank to keep it.`; + } else { + tokenHint.textContent = 'No token is set: mirror runs can only be dry runs, and site settings cannot be edited.'; + } + document.getElementById('fsUrlHint').textContent = settings.url_source === 'environment' + ? 'From the FAIRSTORE_URL environment variable. Saving here overrides it.' + : 'The address people and Verikan use. Must be https:// unless it is on this machine.'; + + document.getElementById('fsChatSource').checked = !!settings.chat_source; + document.getElementById('fsPortalName').value = settings.portal_name || ''; + document.getElementById('fsPortalDescription').value = settings.portal_description || ''; + document.getElementById('fsEnrichDefault').checked = settings.enrich !== false; + document.getElementById('fsMirrorEnrich').checked = settings.enrich !== false; + syncFairStoreChatFields(); + + const chosen = new Set(settings.mirror_sites || []); + const missing = settings.mirror_sites_missing || []; + document.getElementById('fsMirrorSites').innerHTML = (sites.map(site => ` +
+ + +
`).join('') || 'No portals are registered.') + + (missing.length + ? `
No longer registered: + ${escapeHtml(missing.join(', '))}. Save to update the defaults.
` + : ''); + if (settings.conflicting_portal) { + _fsAlert('fsAlert', 'danger', + `The Fair Store URL is the registered portal "${settings.conflicting_portal}". ` + + 'Mirror runs and site-settings edits are refused until it points at its own CKAN.'); + } else { + _fsAlert('fsAlert', '', ''); + } + + const select = document.getElementById('fsMirrorSite'); + const current = select.value; + const allLabel = chosen.size + ? `All portals (${[...chosen].join(', ')})` + : (missing.length ? 'Saved portals (none registered — update Mirror defaults)' : 'All portals'); + select.innerHTML = `` + sites.map(site => + ``).join(''); + if ([...select.options].some(o => o.value === current)) select.value = current; + + const open = document.getElementById('fsOpenLink'); + if (settings.url) { + open.href = settings.url; + open.classList.remove('d-none'); + } else { + open.classList.add('d-none'); + } + const updated = document.getElementById('fsUpdatedHint'); + updated.textContent = settings.updated_at + ? `Last saved ${new Date(settings.updated_at).toLocaleString()}${settings.updated_by ? ` by ${settings.updated_by}` : ''}.` + : ''; + syncFairStoreMirrorLabel(); +} + +function syncFairStoreChatFields() { + document.getElementById('fsChatFields').classList.toggle('d-none', !document.getElementById('fsChatSource').checked); +} + +function syncFairStoreMirrorLabel() { + const apply = document.getElementById('fsMirrorApply').checked; + document.getElementById('fsMirrorStartLabel').textContent = apply ? 'Start mirror' : 'Start dry run'; + const btn = document.getElementById('fsMirrorStartBtn'); + btn.classList.toggle('btn-primary', !apply); + btn.classList.toggle('btn-warning', apply); +} + +async function saveFairStoreSettings(e) { + e.preventDefault(); + const payload = { + url: document.getElementById('fsUrl').value.trim(), + chat_source: document.getElementById('fsChatSource').checked, + portal_name: document.getElementById('fsPortalName').value.trim(), + portal_description: document.getElementById('fsPortalDescription').value.trim(), + mirror_sites: [...document.querySelectorAll('.fs-mirror-site:checked')].map(el => el.value), + enrich: document.getElementById('fsEnrichDefault').checked, + }; + const token = document.getElementById('fsToken').value.trim(); + if (token) payload.token = token; + if (document.getElementById('fsClearToken').checked) payload.clear_token = true; + + const btn = document.getElementById('fsSaveBtn'); + btn.disabled = true; + try { + const response = await adminFetch(FAIRSTORE_API, { + method: 'POST', + headers: { 'Content-Type': 'application/json' }, + body: JSON.stringify(payload), + }); + const data = await response.json(); + const hadToken = !!(_fsSettings && _fsSettings.token_set); + _fsSettings = data.settings; + showToast(data.message || 'Fair Store settings saved', 'success'); + const tokenDropped = hadToken && !token && !payload.clear_token && !_fsSettings.token_set; + // After the reload: renderFairStoreSettings resets the alert box. + await loadFairStore(); + if (tokenDropped && !_fsSettings.conflicting_portal) { + _fsAlert('fsAlert', 'warning', + 'The saved token was removed because the URL now points at a different server ' + + '(host or port). Enter a token for the new Fair Store.'); + } + } catch (error) { + showToast(`Could not save: ${error.message}`, 'danger'); + } finally { + btn.disabled = false; + } +} + +function _fsNumber(n) { + return (n === null || n === undefined) ? '—' : Number(n).toLocaleString(); +} + +async function loadFairStoreStatus() { + const box = document.getElementById('fsStatus'); + try { + const response = await adminFetch(`${FAIRSTORE_API}/status`); + const st = await response.json(); + if (!st.configured) { + box.innerHTML = `
+ No Fair Store is linked yet. Enter its URL under + Connection below.
`; + return; + } + if (!st.reachable) { + box.innerHTML = `
+ Unreachable: ${escapeHtml(st.url)} +
${escapeHtml(st.error || '')}
`; + return; + } + const tokenBadge = { + sysadmin: 'Sysadmin token OK', + rejected: 'Token rejected', + missing: 'No token (read-only)', + }[st.token] || ''; + const byPortal = (st.by_portal || []).map(p => ` + ${escapeHtml(p.name)}${escapeHtml(p.site_id)} + ${_fsNumber(p.datasets)}`).join(''); + const last = st.last_mirror_run; + const lastLine = last + ? `Last mirror run: ${onboardStatusBadge(last.status)} ${escapeHtml(last.site_name || '')} + ${last.options && last.options.apply ? '' : 'dry run'} + ${last.started_at ? escapeHtml(new Date(last.started_at).toLocaleString()) : ''}` + : 'No mirror run from this panel yet.'; + box.innerHTML = ` +
+ ${st.token === 'rejected' ? `
${escapeHtml(st.token_error || 'The Fair Store refused the saved token.')}
` : ''} +
+
Datasets
${_fsNumber(st.datasets)}
+
Organizations
${_fsNumber(st.organizations)}
+
Groups
${_fsNumber(st.groups)}
+
+ ${byPortal ? `
+ + ${byPortal}
Mirrored fromPortal IDDatasets
` : ''} +
${lastLine}
`; + } catch (error) { + box.innerHTML = `
Could not check the Fair Store: ${escapeHtml(error.message)}
`; + } +} + +async function startFairStoreMirror(e) { + e.preventDefault(); + const apply = document.getElementById('fsMirrorApply').checked; + const siteSelect = document.getElementById('fsMirrorSite'); + const target = siteSelect.options[siteSelect.selectedIndex]?.text || 'All portals'; + if (apply && !confirm(`Write ${target} into the Fair Store?\n\nOnly records that changed at the source are written. ` + + 'Records the source withdrew are reported, not deleted.')) { + return; + } + const limitRaw = document.getElementById('fsMirrorLimit').value; + const payload = { + site: siteSelect.value || 'all', + apply, + enrich: document.getElementById('fsMirrorEnrich').checked, + limit: limitRaw ? Number(limitRaw) : null, + }; + const btn = document.getElementById('fsMirrorStartBtn'); + btn.disabled = true; + try { + const response = await adminFetch(`${FAIRSTORE_API}/mirror-jobs`, { + method: 'POST', + headers: { 'Content-Type': 'application/json' }, + body: JSON.stringify(payload), + }); + const data = await response.json(); + showToast(data.message || 'Mirror run started', 'success'); + await loadFairStoreJobs(); + if (data.job?.id) watchFairStoreJob(data.job.id); + } catch (error) { + showToast(`Could not start the run: ${error.message}`, 'danger'); + } finally { + btn.disabled = false; + } +} + +async function loadFairStoreJobs() { + const container = document.getElementById('fsJobList'); + try { + const response = await adminFetch(`${FAIRSTORE_API}/mirror-jobs`); + const data = await response.json(); + document.getElementById('fsMirrorWarnings').innerHTML = (data.warnings || []).map(w => + `
${escapeHtml(w)}
`).join(''); + if (!data.jobs || data.jobs.length === 0) { + container.innerHTML = '
No mirror runs yet.
'; + return; + } + container.innerHTML = ` +
+ + + + ${data.jobs.map(job => ` + + + + + + + + + `).join('')} + +
PortalsModeStatusStartedDurationStarted by
${escapeHtml(job.site_name || job.site_id || '')}${job.options && job.options.apply + ? 'write' + : 'dry run'}${onboardStatusBadge(job.status)} + ${job.error ? `
${escapeHtml(job.error)}
` : ''}
${job.started_at ? escapeHtml(new Date(job.started_at).toLocaleString()) : ''}${escapeHtml(_onboardDuration(job))}${escapeHtml(job.started_by || '')} + +
+
`; + bindClick(container, '.js-fs-log', (el) => watchFairStoreJob(el.dataset.jobId)); + } catch (error) { + container.innerHTML = `
Failed to load mirror runs: ${escapeHtml(error.message)}
`; + } +} + +function stopFairStorePoll() { + if (_fsPollTimer) { + clearTimeout(_fsPollTimer); + _fsPollTimer = null; + } +} + +function closeFairStoreLog() { + stopFairStorePoll(); + _fsLogJobId = null; + document.getElementById('fsLogPanel').classList.add('d-none'); +} + +function watchFairStoreJob(jobId) { + if (_fsLogJobId !== jobId) { + _fsLogOffset = 0; + document.getElementById('fsLog').textContent = ''; + document.getElementById('fsResult').innerHTML = ''; + } + _fsLogJobId = jobId; + document.getElementById('fsLogPanel').classList.remove('d-none'); + document.getElementById('fsLogTitle').textContent = `Mirror run ${jobId}`; + stopFairStorePoll(); + pollFairStoreLog(); +} + +function renderFairStoreResult(result) { + const box = document.getElementById('fsResult'); + if (!result) { + box.innerHTML = ''; + return; + } + if (!Array.isArray(result) && result.error) { + box.innerHTML = `
${escapeHtml(result.error)}
`; + return; + } + const rows = (Array.isArray(result) ? result : [result]).map(r => { + if (r.error) { + return `${escapeHtml(r.site)}${escapeHtml(r.error)}`; + } + if (r.skipped) { + return `${escapeHtml(r.site)}Skipped: ${escapeHtml(r.skipped)}`; + } + const q = r.qsv || {}; + const withdrawn = r.withdrawn_at_source ? (r.withdrawn_at_source.datasets || 0) : null; + return ` + ${escapeHtml(r.site)} + ${_fsNumber(r.datasets)} + ${_fsNumber(r.created)} + ${_fsNumber(r.updated)} + ${_fsNumber(r.unchanged)} + ${_fsNumber(q.created)} / ${_fsNumber(q.refreshed)} / ${_fsNumber(q.unchanged)} + ${withdrawn === null ? '—' : _fsNumber(withdrawn)} + `; + }).join(''); + const mode = (Array.isArray(result) ? result[0] : result)?.mode === 'apply' ? 'Written' : 'Planned (dry run)'; + box.innerHTML = ` +
${escapeHtml(mode)}
+
+ + + + ${rows}
PortalDatasetsCreatedUpdatedUnchangedqsv new / refreshed / sameWithdrawn at source
`; +} + +async function pollFairStoreLog() { + const jobId = _fsLogJobId; + if (!jobId) return; + try { + const response = await adminFetch(`${ONBOARD_API}/${encodeURIComponent(jobId)}?log_offset=${_fsLogOffset}`); + const { job } = await response.json(); + if (jobId !== _fsLogJobId) return; // switched jobs while waiting + + const pre = document.getElementById('fsLog'); + if (job.log && job.log.length) { + const atBottom = pre.scrollHeight - pre.scrollTop - pre.clientHeight < 40; + pre.textContent += job.log.join('\n') + '\n'; + if (atBottom) pre.scrollTop = pre.scrollHeight; + } + _fsLogOffset = job.log_next_offset ?? _fsLogOffset; + document.getElementById('fsLogStatus').innerHTML = onboardStatusBadge(job.status) + + (job.options && !job.options.apply ? ' dry run' : ''); + const isRunning = job.status === 'running'; + document.getElementById('fsCancelBtn').classList.toggle('d-none', !isRunning); + renderFairStoreResult(job.result); + + if (isRunning) { + _fsPollTimer = setTimeout(pollFairStoreLog, 2000); + } else { + stopFairStorePoll(); + loadFairStoreJobs(); + loadFairStoreStatus(); + } + } catch (error) { + stopFairStorePoll(); + document.getElementById('fsLogStatus').innerHTML = + `Lost contact: ${escapeHtml(error.message)}`; + } +} + +async function cancelFairStoreRun() { + if (!_fsLogJobId) return; + if (!confirm('Cancel this mirror run?\n\nWhat it already wrote stays; a rerun picks up the rest.')) return; + try { + await adminFetch(`${ONBOARD_API}/${encodeURIComponent(_fsLogJobId)}/cancel`, { method: 'POST' }); + showToast('Run cancelled', 'success'); + stopFairStorePoll(); // one poll chain only, or log lines repeat + pollFairStoreLog(); + } catch (error) { + showToast(`Could not cancel: ${error.message}`, 'danger'); + } +} + +async function loadFairStoreSiteConfig() { + const fields = document.getElementById('fsSiteFields'); + const actions = document.getElementById('fsSiteActions'); + fields.innerHTML = 'Reading from the Fair Store…'; + try { + const response = await adminFetch(`${FAIRSTORE_API}/site-config`); + renderFairStoreSiteConfig(await response.json()); + _fsAlert('fsSiteAlert', '', ''); + } catch (error) { + fields.innerHTML = ''; + actions.classList.add('d-none'); + _fsAlert('fsSiteAlert', 'warning', `Could not read the site settings: ${error.message}`); + } +} + +function renderFairStoreSiteConfig(config) { + const fields = document.getElementById('fsSiteFields'); + const options = config.options || []; + fields.classList.remove('text-muted', 'small'); + fields.innerHTML = options.map(opt => { + const id = `fsOpt-${opt.key.replace(/[^a-z0-9]/gi, '-')}`; + const value = escapeHtml(opt.value || ''); + const input = opt.kind === 'text' + ? `` + : ``; + return `
+ + ${input} +
`; + }).join('') || '
The Fair Store exposes no editable site settings.
'; + fields.querySelectorAll('.fs-site-opt').forEach(el => { el.dataset.original = el.value; }); + document.getElementById('fsSiteActions').classList.toggle('d-none', options.length === 0); +} + +async function saveFairStoreSiteConfig(e) { + e.preventDefault(); + const changed = {}; + document.querySelectorAll('.fs-site-opt').forEach(el => { + if (el.value !== el.dataset.original) changed[el.dataset.key] = el.value; + }); + if (Object.keys(changed).length === 0) { + showToast('Nothing changed', 'info'); + return; + } + const btn = document.getElementById('fsSiteSaveBtn'); + btn.disabled = true; + try { + const response = await adminFetch(`${FAIRSTORE_API}/site-config`, { + method: 'POST', + headers: { 'Content-Type': 'application/json' }, + body: JSON.stringify({ options: changed }), + }); + const data = await response.json(); + renderFairStoreSiteConfig(data); + showToast(data.message || 'Site settings saved', 'success'); + loadFairStoreStatus(); + } catch (error) { + _fsAlert('fsSiteAlert', 'danger', `Could not save: ${error.message}`); + } finally { + btn.disabled = false; + } +} diff --git a/src/data_concierge/ui/static/js/dictionary.js b/src/data_concierge/ui/static/js/dictionary.js index f486bae..20b80dc 100644 --- a/src/data_concierge/ui/static/js/dictionary.js +++ b/src/data_concierge/ui/static/js/dictionary.js @@ -144,6 +144,7 @@ async function openResource(resourceId, title) { ${d.description ? `

${escapeHtml(d.description)}

` : ''} + ${fairstoreDatasetLink(d.dataset_id)}
@@ -155,6 +156,19 @@ async function openResource(resourceId, title) { `; } +// The Fair Store mirrors these datasets under their source names, with the +// qsv dictionary, stats and frequency tables attached as resources. +function fairstoreDatasetLink(datasetId) { + const base = document.body.dataset.fairstoreUrl; + if (!base || !datasetId) return ''; + const href = `${base}/dataset/${encodeURIComponent(datasetId)}`; + return ``; +} + function backToList() { document.getElementById('dictDetailView').classList.add('d-none'); document.getElementById('dictListView').classList.remove('d-none'); diff --git a/src/data_concierge/ui/templates/admin.html b/src/data_concierge/ui/templates/admin.html index 7946acc..39fb889 100644 --- a/src/data_concierge/ui/templates/admin.html +++ b/src/data_concierge/ui/templates/admin.html @@ -103,6 +103,10 @@

Admin Dashboard

data-bs-target="#ckansites" type="button" role="tab" aria-controls="ckansites"> Data Portals +
Insights
+ + + +
+ + +
+
+ + Checking the Fair Store… +
+
+ +
+ + +
+
+

Mirror Runs

+

+ Copy the registered portals’ catalogs into the Fair Store. + Reruns only write what changed at the source, and records a portal + withdrew are reported, never deleted. Start with a dry run: it + reads everything and reports the plan without writing. +

+
+
+ +
+ + +
+ + +
+
+ + +
+
+
+ + +
+
+ + +
+
+
+ +
+
+ “Max datasets” is for smoke tests: a limited run skips the + withdrawn-at-source check. Socrata enrichment adds each data.pa.gov + dataset’s owning agency and column list. +
+ + +
+
+ + Loading mirror runs… +
+
+ +
+
+
+ Mirror run + +
+
+ + +
+
+
+

+                    
+ +
+ + +
+
+

Connection

+

+ Where the Fair Store is, the sysadmin API token Verikan uses to write + to it, and whether chat users can search it. +

+
+
+ + +
+ + +
+ The address people and Verikan use. Must be https:// unless it is on this machine. +
+
+
+ +
+ + +
+
Leave blank to keep the existing token.
+
+ + +
+
+ Create one on the Fair Store with + ckan user token add <sysadmin> verikan. Changing the URL to + another host drops the saved token, so it is never sent to a different server. +
+
+ +
+ + +
+
+
+ + +
+
+ + +
+
+ The agent searches the Fair Store’s catalog and loads rows from the + portal each dataset was mirrored from. +
+
+ +

Mirror defaults

+
Portals an “All portals” run covers (none checked = every registered portal):
+
+
+ + +
+ +
+ +
+
+ + +
+ + +
+
+

Fair Store Site Settings

+

+ The Fair Store’s own title, description, home page and about text, + logo and custom CSS — read from and saved to the Fair Store live. +

+
+
+ +
+
+
+ +
Set the URL and a sysadmin token to edit these.
+
+ +
+ + + +
@@ -1658,7 +1881,7 @@ - + diff --git a/src/data_concierge/ui/templates/dictionary.html b/src/data_concierge/ui/templates/dictionary.html index 0a87b61..a7305b1 100644 --- a/src/data_concierge/ui/templates/dictionary.html +++ b/src/data_concierge/ui/templates/dictionary.html @@ -42,7 +42,7 @@ })(); - +