diff --git a/.env.example b/.env.example index b339b8c..cd6a17a 100644 --- a/.env.example +++ b/.env.example @@ -78,6 +78,12 @@ DATA_COMMONS_API_KEY= CKAN_URL=https://data.dathere.com CKAN_API_KEY= +# Fair Store — the CKAN that mirrors every registered portal. Seeds for the +# admin panel's Fair Store page (values saved there take precedence). The token +# is a CKAN sysadmin API token; it is only ever sent to FAIRSTORE_URL's host. +# FAIRSTORE_URL=https://fairstore.example.org +# FAIRSTORE_API_KEY= + # WPRDC (Western PA Regional Data Center) - City of Pittsburgh open data WPRDC_CKAN_URL=https://data.wprdc.org WPRDC_ORGANIZATION=city-of-pittsburgh diff --git a/.gitignore b/.gitignore index ae0b915..f431d25 100644 --- a/.gitignore +++ b/.gitignore @@ -158,11 +158,17 @@ logs/ *.log .secrets/ +# API token files (deploy/fairstore/deploy.sh writes the mirror's outside the +# repository; these catch one saved here by mistake) +fairstore-token* +*token.txt + # Runtime-seeded local storage (contains user data / secrets — never commit) /users.json /roles.json /notebook_verification/ /github_settings.json +/fairstore_settings.json /landing_settings.json /system_prompt_settings.json /runtime_settings.json diff --git a/README.md b/README.md index af04a05..3cc060b 100644 --- a/README.md +++ b/README.md @@ -360,36 +360,127 @@ examples/ # Sample generated notebooks ## CKAN Fair Store mirror -Start the local CKAN 2.11 Fair Store with organization hierarchy support: +Start the local CKAN 2.11 Fair Store with organization hierarchy support. +Set `FAIRSTORE_SECRET_KEY` (any long random string, e.g. in `.env`) so API +tokens and sessions survive the container being recreated: ```bash docker compose --profile fairstore up -d --build fairstore ``` -Preview a registered CKAN portal before writing anything: +On first run, create a sysadmin and an API token for the mirror (keep the token +out of the repository): ```bash -python -m scripts.populate_fairstore \ - --site wprdc \ - --target-url http://localhost:5001 +docker compose exec fairstore ckan -c /srv/app/ckan.ini user add fairadmin email=fairadmin@localhost.localdomain password= +docker compose exec fairstore ckan -c /srv/app/ckan.ini sysadmin add fairadmin +docker compose exec fairstore ckan -c /srv/app/ckan.ini user token add fairadmin mirror +``` + +Preview every registered portal (CKAN and DCAT) before writing anything, or +pass one `--site` ID: + +```bash +python -m scripts.populate_fairstore --site all --target-url http://localhost:5001 ``` -To apply the mirror, create a target CKAN sysadmin token, expose it through -`CKAN_API_KEY` (or a protected file), and add `--apply`. The command performs a -collision preflight first, then preserves source organization, dataset, and -resource names and UUIDs. It creates one parent organization for the source -portal, attaches source organizations beneath it, and adds any matching qsv -profile to the resource as separate metadata. Re-running patches the preserved -UUIDs, so it does not duplicate resources. +Add `--apply` with the token in `FAIRSTORE_API_KEY` or `--api-key-file` to +write (`FAIRSTORE_URL` sets the default `--target-url`; the app's own +`CKAN_URL`/`CKAN_API_KEY` are deliberately not used, and a portal that is the +target is never mirrored into itself). Every read fails closed — a portal that +errors mid-read is not mirrored from a partial snapshot — and with `--site all` +one portal's failure does not stop the others (the exit status is non-zero). +The command runs a collision preflight first, then: + +- creates one parent organization per source portal and attaches the portal's + publishers beneath it (ckanext-hierarchy); +- for a **CKAN** portal, preserves organization, group, dataset, and resource + names and UUIDs; +- for a **DCAT** portal (`/data.json` from Socrata, DKAN, ArcGIS Hub), turns + publishers into organizations, themes into groups, and distributions into + resources, with UUIDv5 identifiers derived from the catalog's own IDs. A + Socrata catalog is enriched from the Socrata Discovery API with the owning + agency (the hierarchy) and the column list (`source_data_dictionary`); pass + `--no-enrich` to skip it; +- keeps every source field: anything the Fair Store's default schema has no + column for (ckanext-scheming fields, DCAT-US fields) becomes an extra, and + links to files uploaded to the source keep pointing at the source; +- publishes each dataset's qsv profile (from `data/{ckan,dcat}_onboard/`): the + AI description (its prose; describegpt's provenance block stays in the raw + output), AI tags, row and column counts and column names become `qsv_*` + fields on the dataset, and four resources are added to it — a data + dictionary (types, AI labels and descriptions, key statistics), the full + `qsv stats` and `qsv frequency` tables (DataStore tables, shown as sortable + tables and downloadable as CSV/JSON), and the raw describegpt JSON. API keys + that describegpt writes into its attribution are redacted, and stats, + frequency or describegpt output that does not match the profiled file's + header (left behind by another file's profile in the same directory) is not + published. The source data itself is never copied; +- skips records the source itself mirrored from another portal, so each + dataset is mirrored once from its origin. + +Re-running patches the preserved UUIDs, so it does not duplicate anything, +and it only writes what changed: each mirrored record carries a digest of what +was last written (`mirror_digest`, `qsv_digest`), so an unchanged dataset, +resource or qsv table is left alone. A record the source renamed is renamed +(DCAT names are derived, so those keep the name they were published under); +records the source no longer publishes are reported (`withdrawn_at_source`), +never deleted. The dry run reports the same created/updated/unchanged plan. ```bash python -m scripts.populate_fairstore \ - --site wprdc \ + --site all \ --target-url http://localhost:5001 \ --api-key-file /path/to/protected-token \ --apply ``` +### Managing the Fair Store from Verikan + +The admin panel's **Fair Store** page links Verikan to it: + +- **Connection** — the Fair Store URL and a sysadmin API token (`ckan user + token add verikan` on the Fair Store). `FAIRSTORE_URL` and + `FAIRSTORE_API_KEY` seed these; values saved on the page take precedence. + The token is never shown again and is only ever sent to the host it was saved + for: moving the URL to another host drops it. +- **Chat source** — optionally offers the Fair Store as a data source in chat. + The agent searches its catalog and loads each dataset's rows from the portal + it was mirrored from. +- **Mirror runs** — dry runs and writes of every portal or one, with a live log + and a per-portal summary. On Cloud Run the qsv profiles come from the storage + backend that onboarding syncs to (`--qsv-source`). +- **Site settings** — the Fair Store's title, description, home page and about + text, logo and custom CSS, read and saved live. + +The chat sidebar and the data dictionary link to the Fair Store once it is set. + +### Deploying the Fair Store on GCP + +`deploy/fairstore/deploy.sh` runs the same stack on one Compute Engine VM next +to Verikan (in `us-central1` by default; `PROJECT` names the GCP project): CKAN on uwsgi +(built from `ckan/ckan-base`), Postgres, Solr, Redis, and Caddy for automatic +HTTPS. It is idempotent — the first run creates the static IP, firewall rule, +VM and a daily snapshot schedule; later runs re-ship `docker/fairstore` and +restart the stack. Secrets are generated on the VM (`/opt/fairstore/.env`) and +never leave it. The host is `FAIRSTORE_HOST` if set, else the one already +saved on the VM, else `.sslip.io` (first deploy); point a DNS name at the IP +and rerun with `FAIRSTORE_HOST` set to move it (the search index is rebuilt +for the new URLs). + +The VM has no service account (the site calls no Google API; a VM that still +has one is stopped once to remove it), and containers are blocked from the +metadata server except for DNS. + +```bash +gcloud auth login +PROJECT= deploy/fairstore/deploy.sh +``` + +The script prints the command that mints the mirror's API token into +`~/.config/verikan/fairstore-token` (outside the repository); then run the +mirror above with `--target-url https://` and that `--api-key-file`. + ## Development ```bash diff --git a/deploy/fairstore/caddy/Caddyfile b/deploy/fairstore/caddy/Caddyfile new file mode 100644 index 0000000..59028a7 --- /dev/null +++ b/deploy/fairstore/caddy/Caddyfile @@ -0,0 +1,48 @@ +# HTTPS for the Fair Store; Caddy obtains and renews the certificate itself. +# +# Mounted as a directory (./caddy:/etc/caddy), not as a single file: a +# redeploy's tar replaces the file with a new inode, and a single-file bind +# mount keeps serving the old one. remote-up.sh runs `caddy reload` afterwards. +{ + servers { + # Only TCP 443 is published and allowed through the firewall, so + # don't advertise HTTP/3 (UDP) to clients that would try it first. + protocols h1 h2 + } +} + +{$FAIRSTORE_HOST} { + encode zstd gzip + + # Access log with the client's address; CKAN only ever sees Caddy's. + log { + output stdout + format json + } + + header { + Strict-Transport-Security "max-age=31536000" + X-Content-Type-Options nosniff + Referrer-Policy strict-origin-when-cross-origin + # Don't advertise the server software (Caddy adds both). + -Server + -Via + } + + # CKAN's "Embed" button hands out an iframe of a resource view page + # (/dataset//resource//view/), so only those stay frameable by + # other sites. + @not_view not path */view/* + header @not_view { + X-Frame-Options SAMEORIGIN + defer + } + + reverse_proxy fairstore:5000 { + # uwsgi's http router closes idle keep-alive connections; a POST sent on + # a pooled one fails with a 502, and Go never retries a non-idempotent request. + transport http { + keepalive off + } + } +} diff --git a/deploy/fairstore/deploy.sh b/deploy/fairstore/deploy.sh new file mode 100755 index 0000000..eff84a6 --- /dev/null +++ b/deploy/fairstore/deploy.sh @@ -0,0 +1,174 @@ +#!/usr/bin/env bash +# Deploy the Fair Store (CKAN 2.11 + ckanext-hierarchy) next to Verikan. +# +# One Compute Engine VM runs the same stack as the local `fairstore` compose +# profile — CKAN, Postgres, Solr, Redis — plus Caddy for automatic HTTPS. A VM +# rather than Cloud Run because Postgres and Solr need persistent disks. +# +# Idempotent: creates the static IP, firewall rule, VM and daily disk snapshots +# on the first run, then (re)ships docker/fairstore + this directory and +# restarts the stack. Secrets are generated on the VM and stay there. +# +# gcloud auth login # the session expires; re-auth is interactive +# PROJECT= deploy/fairstore/deploy.sh +# +# The VM runs without a service account: the site calls no Google API, and the +# default one is a project Editor. A VM that still has one is stopped, stripped +# of it and started again — a few minutes of downtime, once. +# +# The host is FAIRSTORE_HOST if set, else the one the VM already serves (its +# .env), else .sslip.io, a wildcard DNS name that resolves to the VM so +# Caddy can get a certificate before a real domain exists. Point a DNS name at +# the IP and rerun with FAIRSTORE_HOST set to move it. +set -euo pipefail + +PROJECT="${PROJECT:?set PROJECT to the GCP project id to deploy into}" +REGION="${REGION:-us-central1}" +ZONE="${ZONE:-us-central1-a}" +VM="${VM:-verikan-fairstore}" +MACHINE_TYPE="${MACHINE_TYPE:-e2-medium}" +DISK_SIZE="${DISK_SIZE:-50GB}" +NETWORK_TAG="${VM}-web" +REMOTE_DIR=/opt/fairstore +ROOT="$(cd "$(dirname "$0")/../.." && pwd)" +STARTUP_SCRIPT="$ROOT/deploy/fairstore/vm-startup.sh" +# Outside the repository, so the mirror's sysadmin token can never be committed. +TOKEN_FILE="${XDG_CONFIG_HOME:-$HOME/.config}/verikan/fairstore-token" + +gc() { gcloud --project "$PROJECT" --quiet "$@"; } +on_vm() { gc compute ssh "$VM" --zone "$ZONE" --command "$1"; } +vm_field() { gc compute instances describe "$VM" --zone "$ZONE" --format="value($1)"; } + +# wait_for +# Prints a dot per failed try and, on timeout, the last try's output (the why). +wait_for() { + local limit="$1" pause="$2" what="$3" deadline out + shift 3 + deadline=$((SECONDS + limit)) + printf 'Waiting for %s ' "$what" + until out="$("$@" 2>&1)"; do + if ((SECONDS >= deadline)); then + echo + echo "deploy: gave up waiting for $what after $((limit / 60)) min; last try said:" >&2 + printf '%s\n' "$out" | tail -n 20 >&2 + exit 1 + fi + printf . + sleep "$pause" + done + echo " ok" +} + +gc projects describe "$PROJECT" --format='value(projectId)' >/dev/null \ + || { echo "gcloud cannot reach $PROJECT; run: gcloud auth login" >&2; exit 1; } + +if ! gc compute addresses describe "${VM}-ip" --region "$REGION" >/dev/null 2>&1; then + gc compute addresses create "${VM}-ip" --region "$REGION" +fi +IP="$(gc compute addresses describe "${VM}-ip" --region "$REGION" --format='value(address)')" + +if ! gc compute firewall-rules describe "${VM}-web" >/dev/null 2>&1; then + gc compute firewall-rules create "${VM}-web" --network default --direction INGRESS \ + --allow tcp:80,tcp:443 --target-tags "$NETWORK_TAG" --source-ranges 0.0.0.0/0 +fi + +if ! gc compute instances describe "$VM" --zone "$ZONE" >/dev/null 2>&1; then + gc compute instances create "$VM" --zone "$ZONE" --machine-type "$MACHINE_TYPE" \ + --image-family debian-12 --image-project debian-cloud \ + --boot-disk-size "$DISK_SIZE" --boot-disk-type pd-balanced \ + --address "$IP" --tags "$NETWORK_TAG" \ + --no-service-account --no-scopes --deletion-protection \ + --labels app=verikan,component=fairstore \ + --metadata-from-file startup-script="$STARTUP_SCRIPT" +else + status="$(vm_field status)" + sa="$(vm_field 'serviceAccounts[].email')" + if [ -n "$sa" ]; then + echo "$VM runs as $sa, which the site never uses; removing it." + echo "That needs the VM stopped: the Fair Store is down for a few minutes." + [ "$status" = TERMINATED ] || gc compute instances stop "$VM" --zone "$ZONE" + gc compute instances set-service-account "$VM" --zone "$ZONE" \ + --no-service-account --no-scopes + status=TERMINATED + fi + # The boot disk holds every database: guard the VM against deletion. + gc compute instances update "$VM" --zone "$ZONE" --deletion-protection + # GCE runs the startup script stored in instance metadata, so ship the + # current one; it applies from the next boot (remote-up.sh covers this one). + gc compute instances add-metadata "$VM" --zone "$ZONE" \ + --metadata-from-file startup-script="$STARTUP_SCRIPT" + if [ "$status" = TERMINATED ]; then + echo "Starting $VM..." + gc compute instances start "$VM" --zone "$ZONE" + fi +fi + +# Postgres, Solr and uploads all live on the boot disk: keep a week of daily +# snapshots. +if ! gc compute resource-policies describe "${VM}-daily" --region "$REGION" >/dev/null 2>&1; then + gc compute resource-policies create snapshot-schedule "${VM}-daily" --region "$REGION" \ + --daily-schedule --start-time 07:00 --max-retention-days 7 \ + --on-source-disk-delete keep-auto-snapshots +fi +policies="$(gc compute disks describe "$VM" --zone "$ZONE" --format='value(resourcePolicies)')" +case ";${policies};" in + *"/resourcePolicies/${VM}-daily;"*) ;; + *) gc compute disks add-resource-policies "$VM" --zone "$ZONE" --resource-policies "${VM}-daily" ;; +esac + +# `compose ls` needs both the compose plugin and a running daemon. +wait_for 900 15 "Docker on $VM" on_vm "sudo docker compose ls" + +# Default to the host the VM already serves: falling back to sslip.io on every +# run would silently move a site that has been given a real domain. +if [ -n "${FAIRSTORE_HOST:-}" ]; then + HOST="$FAIRSTORE_HOST" +else + HOST="$(on_vm "if sudo test -f $REMOTE_DIR/.env; then sudo sed -n 's/^FAIRSTORE_HOST=//p' $REMOTE_DIR/.env; fi" \ + | tail -n 1 | tr -d '\r')" + HOST="${HOST:-${IP//./-}.sslip.io}" +fi +[[ "$HOST" =~ ^[A-Za-z0-9.-]+$ ]] || { echo "deploy: bad host name: '$HOST'" >&2; exit 1; } +echo "Deploying to https://$HOST" + +bundle="$(mktemp -d)" +trap 'rm -rf "$bundle"' EXIT +mkdir -p "$bundle/stack/ckan" +cp "$ROOT"/deploy/fairstore/{docker-compose.yml,remote-up.sh} "$bundle/stack/" +cp -R "$ROOT"/deploy/fairstore/caddy "$bundle/stack/" +cp -R "$ROOT"/docker/fairstore/. "$bundle/stack/ckan/" +# macOS tar would add AppleDouble ._* files and xattrs, and record the local +# user as owner. Members are named rather than ".", so extracting never resets +# the mode of /opt/fairstore itself. +COPYFILE_DISABLE=1 tar --no-xattrs --owner=root:0 --group=root:0 --exclude .DS_Store \ + -C "$bundle/stack" -czf "$bundle/fairstore.tgz" docker-compose.yml remote-up.sh caddy ckan +gc compute scp "$bundle/fairstore.tgz" "$VM:/tmp/fairstore.tgz" --zone "$ZONE" +# ckan/ is only a build context (its init-db.sh runs only when Postgres +# initialises an empty volume), so it is replaced whole and files deleted here +# disappear there. .env is never touched, and neither is the caddy/ directory +# itself: Caddy bind-mounts it, so only the files inside it are replaced. +# The first deploys' bundles left macOS ._* files and the operator's ownership +# at the top level; both are tidied here. +on_vm "sudo mkdir -p $REMOTE_DIR && sudo chown root:root $REMOTE_DIR \ + && sudo find $REMOTE_DIR -maxdepth 1 -name '._*' -delete \ + && sudo rm -rf $REMOTE_DIR/ckan \ + && sudo tar --no-same-owner -xzf /tmp/fairstore.tgz -C $REMOTE_DIR \ + && rm -f /tmp/fairstore.tgz \ + && sudo bash $REMOTE_DIR/remote-up.sh $HOST" + +wait_for 600 10 "https://$HOST" \ + curl -fsS -o /dev/null --max-time 20 "https://$HOST/api/3/action/status_show" +cat < "$TOKEN_FILE") +then load every registered portal: + python -m scripts.populate_fairstore --site all --apply \\ + --target-url https://$HOST --api-key-file "$TOKEN_FILE" + +The fairadmin password is FAIRSTORE_ADMIN_PASSWORD in $REMOTE_DIR/.env on the VM. +EOF diff --git a/deploy/fairstore/docker-compose.yml b/deploy/fairstore/docker-compose.yml new file mode 100644 index 0000000..fded497 --- /dev/null +++ b/deploy/fairstore/docker-compose.yml @@ -0,0 +1,117 @@ +# Production Fair Store: CKAN 2.11 (uwsgi) + ckanext-hierarchy behind Caddy, +# which provisions HTTPS automatically. deploy.sh ships this directory, with +# docker/fairstore as ./ckan, to one Compute Engine VM next to Verikan. Every +# secret comes from .env, generated on the VM on first deploy (remote-up.sh). +# +# Same services and names as the fairstore profile in the root +# docker-compose.yml, minus the dev server, plus Caddy and hardening. +name: fairstore + +# Container logs (Caddy's access log included) are rotated, not kept forever. +x-logging: &logging + driver: json-file + options: + max-size: "20m" + max-file: "5" + +services: + fairstore: + build: + context: ./ckan + args: + CKAN_BASE_IMAGE: ckan/ckan-base:2.11.6 + image: fairstore-ckan:2.11.6 + depends_on: + fairstore-db: + condition: service_healthy + fairstore-solr: + condition: service_started + fairstore-redis: + condition: service_started + environment: + - CKAN_SITE_URL=https://${FAIRSTORE_HOST:?} + - CKAN_SQLALCHEMY_URL=postgresql://ckan:${FAIRSTORE_DB_PASSWORD:?}@fairstore-db/ckan + - CKAN_DATASTORE_WRITE_URL=postgresql://ckan_datastore_write:${FAIRSTORE_DATASTORE_PASSWORD:?}@fairstore-db/datastore + - CKAN_DATASTORE_READ_URL=postgresql://ckan_datastore_read:${FAIRSTORE_DATASTORE_PASSWORD:?}@fairstore-db/datastore + - CKAN_SOLR_URL=http://fairstore-solr:8983/solr/ckan + - CKAN_REDIS_URL=redis://fairstore-redis:6379/1 + - CKAN_SECRET_KEY=${FAIRSTORE_SECRET_KEY:?} + # Created by the image's prerun on first start; skipped once it exists. + - CKAN_SYSADMIN_NAME=fairadmin + - CKAN_SYSADMIN_PASSWORD=${FAIRSTORE_ADMIN_PASSWORD:?} + - CKAN_SYSADMIN_EMAIL=fairadmin@${FAIRSTORE_HOST:?} + - CKAN__PLUGINS=envvars datastore datatables_view text_view image_view hierarchy_display hierarchy_form + - CKAN__VIEWS__DEFAULT_VIEWS=datatables_view text_view image_view + # The 2.11.6 image's ckan.ini sets ckan.uploads_enabled to an empty + # value, which turns file uploads off in the web forms. + - CKAN__UPLOADS_ENABLED=true + # CKAN exempts extension endpoints (the DataStore dictionary editor, the + # DataTables AJAX feed) from CSRF checks by default, and SameSite=Lax + # does not cover them here: sslip.io is not a public suffix, so every + # other *.sslip.io site counts as same-site. + - CKAN__CSRF_PROTECTION__IGNORE_EXTENSIONS=false + - CKAN__SITE_TITLE=Verikan Fair Store + # Records are written by the mirror and edited by admins; nobody signs + # up through the web, and the user list is not public. + - CKAN__AUTH__CREATE_USER_VIA_WEB=false + - CKAN__AUTH__PUBLIC_USER_DETAILS=false + volumes: + - fairstore_data:/var/lib/ckan + restart: unless-stopped + logging: *logging + + fairstore-db: + image: postgres:16-alpine + environment: + - POSTGRES_USER=postgres + - POSTGRES_PASSWORD=${FAIRSTORE_POSTGRES_PASSWORD:?} + - POSTGRES_DB=postgres + - CKAN_DB_PASSWORD=${FAIRSTORE_DB_PASSWORD:?} + - DATASTORE_DB_PASSWORD=${FAIRSTORE_DATASTORE_PASSWORD:?} + volumes: + - fairstore_db:/var/lib/postgresql/data + - ./ckan/init-db.sh:/docker-entrypoint-initdb.d/init-db.sh:ro + healthcheck: + test: ["CMD-SHELL", "pg_isready -U postgres"] + interval: 10s + timeout: 5s + retries: 10 + restart: unless-stopped + logging: *logging + + fairstore-solr: + image: ckan/ckan-solr:2.11-solr9 + volumes: + - fairstore_solr:/var/solr + restart: unless-stopped + logging: *logging + + fairstore-redis: + image: redis:7-alpine + restart: unless-stopped + logging: *logging + + caddy: + image: caddy:2-alpine + depends_on: + - fairstore + ports: + - "80:80" + - "443:443" + environment: + - FAIRSTORE_HOST=${FAIRSTORE_HOST:?} + volumes: + # The directory, not the file: a redeploy replaces the Caddyfile's inode, + # which a single-file bind mount would never see (caddy/Caddyfile). + - ./caddy:/etc/caddy:ro + - caddy_data:/data + - caddy_config:/config + restart: unless-stopped + logging: *logging + +volumes: + fairstore_data: + fairstore_db: + fairstore_solr: + caddy_data: + caddy_config: diff --git a/deploy/fairstore/remote-up.sh b/deploy/fairstore/remote-up.sh new file mode 100755 index 0000000..e212e11 --- /dev/null +++ b/deploy/fairstore/remote-up.sh @@ -0,0 +1,74 @@ +#!/usr/bin/env bash +# Runs on the Fair Store VM, as root, from deploy.sh: create the secrets on the +# first deploy, then build and (re)start the stack. +# +# remote-up.sh +# +# .env is generated once and never leaves the VM. Rotating a value means +# editing .env; note the database passwords are only applied when Postgres +# initialises an empty volume. +set -euo pipefail +cd "$(dirname "$0")" +host="${1:?usage: remote-up.sh }" +[[ "$host" =~ ^[A-Za-z0-9.-]+$ ]] || { echo "remote-up: bad hostname: $host" >&2; exit 1; } + +# retry +retry() { + local tries="$1" pause="$2" i + shift 2 + for ((i = 1; ; i++)); do + "$@" && return 0 + ((i < tries)) || return 1 + sleep "$pause" + done +} + +# Containers get no path to the metadata server, DNS (port 53) excepted; see +# vm-startup.sh, which re-adds these on every boot. Added here too so a deploy +# takes effect without a reboot. +for proto in tcp udp; do + iptables -C DOCKER-USER -d 169.254.169.254 -p "$proto" ! --dport 53 -j DROP 2>/dev/null \ + || iptables -I DOCKER-USER -d 169.254.169.254 -p "$proto" ! --dport 53 -j DROP +done + +if [ ! -f .env ]; then + umask 077 + secret() { openssl rand -hex 24; } + { + echo "FAIRSTORE_HOST=${host}" + echo "FAIRSTORE_SECRET_KEY=$(secret)" + echo "FAIRSTORE_POSTGRES_PASSWORD=$(secret)" + echo "FAIRSTORE_DB_PASSWORD=$(secret)" + echo "FAIRSTORE_DATASTORE_PASSWORD=$(secret)" + echo "FAIRSTORE_ADMIN_PASSWORD=$(secret)" + } > .env +fi +# Solr keeps each dataset as indexed, with absolute URLs built from the site +# URL, so a host change needs a reindex. The marker survives a failed run, so +# the next deploy still does it. +if [ "$(sed -n 's/^FAIRSTORE_HOST=//p' .env | tail -n 1)" != "$host" ]; then + touch .reindex-pending +fi +sed -i "s|^FAIRSTORE_HOST=.*|FAIRSTORE_HOST=${host}|" .env + +docker compose up -d --build --remove-orphans +# Caddy reads its config only at start; reload applies a changed Caddyfile +# without dropping connections. Retried because a just-created container may +# not have its admin endpoint up yet. +retry 10 3 docker compose exec -T caddy caddy reload --config /etc/caddy/Caddyfile --adapter caddyfile \ + || { echo "remote-up: caddy reload failed (the previous config is still live)" >&2; exit 1; } +# Before the caddy/ directory mount the Caddyfile sat here; nothing reads it now. +rm -f Caddyfile + +if [ -f .reindex-pending ]; then + ckan_up() { + docker compose exec -T fairstore python3 -c \ + 'import urllib.request as u; u.urlopen("http://127.0.0.1:5000/api/3/action/status_show", timeout=10)' \ + >/dev/null 2>&1 + } + echo "Site URL is now https://${host}: waiting for CKAN, then rebuilding the search index..." + retry 60 5 ckan_up \ + || { echo "remote-up: CKAN did not answer after 60 tries; search index NOT rebuilt (the next deploy retries)" >&2; exit 1; } + docker compose exec -T fairstore ckan -c /srv/app/ckan.ini search-index rebuild + rm -f .reindex-pending +fi diff --git a/deploy/fairstore/vm-startup.sh b/deploy/fairstore/vm-startup.sh new file mode 100755 index 0000000..25479b8 --- /dev/null +++ b/deploy/fairstore/vm-startup.sh @@ -0,0 +1,52 @@ +#!/bin/bash +# Compute Engine startup script for the Fair Store VM (runs as root on every +# boot). Installs Docker Engine and the compose plugin on first boot, then, on +# every boot, keeps containers away from the metadata server. The stack itself +# restarts on its own (restart: unless-stopped). +# +# GCE reads this from instance metadata, not from the repository: deploy.sh +# re-uploads it on each run, and a change takes effect at the next boot. +set -euo pipefail + +install_docker() { + # A first boot races GCE's own apt runs (unattended-upgrades, the guest + # agent), so wait for the lock rather than fail. + local apt=(apt-get -o DPkg::Lock::Timeout=600) + "${apt[@]}" update + "${apt[@]}" install -y ca-certificates curl + install -m 0755 -d /etc/apt/keyrings + curl -fsSL https://download.docker.com/linux/debian/gpg -o /etc/apt/keyrings/docker.asc + chmod a+r /etc/apt/keyrings/docker.asc + . /etc/os-release + echo "deb [arch=$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/docker.asc] https://download.docker.com/linux/debian ${VERSION_CODENAME} stable" \ + > /etc/apt/sources.list.d/docker.list + "${apt[@]}" update + "${apt[@]}" install -y docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin + systemctl enable --now docker +} + +# The compose plugin is the last package the stack needs: a first boot cut +# short after the docker CLI was unpacked is finished on the next boot. +docker compose version >/dev/null 2>&1 || install_docker + +# Nothing in the stack calls a Google API, so containers get no path to the +# metadata server (the VM has no service account either; this is the second +# layer). Port 53 stays open: on GCE that address is also the VM's DNS +# resolver, which the image build's RUN steps and any default-bridge +# container query directly (containers on the compose network go through +# Docker's embedded resolver), so a blanket DROP would break image builds. +# iptables rules do not survive a reboot, and DOCKER-USER only exists once +# dockerd has started. remote-up.sh adds the same rules on each deploy. +for _ in $(seq 60); do + iptables -n -L DOCKER-USER >/dev/null 2>&1 && break + sleep 5 +done +if ! iptables -n -L DOCKER-USER >/dev/null 2>&1; then + echo "vm-startup: no DOCKER-USER chain after 5 min (is dockerd running?);" \ + "containers can still reach the metadata server" >&2 + exit 1 +fi +for proto in tcp udp; do + iptables -C DOCKER-USER -d 169.254.169.254 -p "$proto" ! --dport 53 -j DROP 2>/dev/null \ + || iptables -I DOCKER-USER -d 169.254.169.254 -p "$proto" ! --dport 53 -j DROP +done diff --git a/docker-compose.yml b/docker-compose.yml index 8839665..707a888 100644 --- a/docker-compose.yml +++ b/docker-compose.yml @@ -76,8 +76,9 @@ services: # then open http://localhost:5001 (host 5001 -> container 5000; macOS # reserves 5000 for AirPlay). # - # First run: create an admin and an API token with - # docker compose exec fairstore ckan -c /srv/app/ckan.ini sysadmin add admin + # First run: create a sysadmin and an API token (commands under "CKAN Fair + # Store mirror" in README.md), load every registered portal with + # python -m scripts.populate_fairstore --site all --apply # then set CKAN_URL / CKAN_API_KEY for the app. # --------------------------------------------------------------------- fairstore: @@ -85,28 +86,48 @@ services: build: context: ./docker/fairstore dockerfile: Dockerfile - image: data-concierge-fairstore:2.11.3 + image: data-concierge-fairstore:2.11.6 platform: linux/amd64 depends_on: fairstore-db: condition: service_healthy fairstore-solr: condition: service_started + fairstore-redis: + condition: service_started ports: - "${FAIRSTORE_PORT:-5001}:5000" environment: - - CKAN_SITE_URL=http://localhost:5000 + # The URL browsers reach CKAN at. CKAN builds upload and DataStore dump + # links from it, so it must be the host port, not the container's 5000. + - CKAN_SITE_URL=http://localhost:${FAIRSTORE_PORT:-5001} # Credentials must match the users the postgres image's init scripts # create (see fairstore-db below): ckan / ckan_datastore_write / # ckan_datastore_read. - - CKAN_SQLALCHEMY_URL=postgresql://ckan:ckan@fairstore-db/ckan - - CKAN_DATASTORE_WRITE_URL=postgresql://ckan_datastore_write:datastore@fairstore-db/datastore - - CKAN_DATASTORE_READ_URL=postgresql://ckan_datastore_read:datastore@fairstore-db/datastore + - CKAN_SQLALCHEMY_URL=postgresql://ckan:${FAIRSTORE_DB_PASSWORD:-ckan}@fairstore-db/ckan + - CKAN_DATASTORE_WRITE_URL=postgresql://ckan_datastore_write:${FAIRSTORE_DATASTORE_PASSWORD:-datastore}@fairstore-db/datastore + - CKAN_DATASTORE_READ_URL=postgresql://ckan_datastore_read:${FAIRSTORE_DATASTORE_PASSWORD:-datastore}@fairstore-db/datastore - CKAN_SOLR_URL=http://fairstore-solr:8983/solr/ckan + - CKAN_REDIS_URL=redis://fairstore-redis:6379/1 + # Pins the session and API-token signing secrets (docker/fairstore/ + # pin-secrets.sh). Unset, each new container signs with fresh random + # secrets and every issued API token stops working. + - CKAN_SECRET_KEY=${FAIRSTORE_SECRET_KEY:-} # hierarchy_display must load before hierarchy_form. Organizations that # represent source portals can then be parents of the organizations # mirrored from those portals. - - CKAN__PLUGINS=envvars datastore hierarchy_display hierarchy_form + - CKAN__PLUGINS=envvars datastore datatables_view text_view image_view hierarchy_display hierarchy_form + # qsv profiles (stats, frequency, data dictionary) are DataStore tables, + # shown as sortable tables; describegpt's JSON output is shown inline. + - CKAN__VIEWS__DEFAULT_VIEWS=datatables_view text_view image_view + # The 2.11.6 image's ckan.ini sets ckan.uploads_enabled to an empty + # value, which turns file uploads off in the web forms. + - CKAN__UPLOADS_ENABLED=true + # CKAN exempts extension endpoints (the DataStore dictionary editor, the + # DataTables AJAX feed) from CSRF checks by default, and SameSite=Lax + # does not cover them here: sslip.io is not a public suffix, so every + # other *.sslip.io site counts as same-site. + - CKAN__CSRF_PROTECTION__IGNORE_EXTENSIONS=false volumes: - fairstore_data:/var/lib/ckan restart: unless-stopped @@ -119,8 +140,11 @@ services: image: postgres:16-alpine environment: - POSTGRES_USER=postgres - - POSTGRES_PASSWORD=postgres + - POSTGRES_PASSWORD=${FAIRSTORE_POSTGRES_PASSWORD:-postgres} - POSTGRES_DB=postgres + # Read by init-db.sh on first start; must match the URLs above. + - CKAN_DB_PASSWORD=${FAIRSTORE_DB_PASSWORD:-ckan} + - DATASTORE_DB_PASSWORD=${FAIRSTORE_DATASTORE_PASSWORD:-datastore} volumes: - fairstore_db:/var/lib/postgresql/data - ./docker/fairstore/init-db.sh:/docker-entrypoint-initdb.d/init-db.sh:ro @@ -138,6 +162,13 @@ services: - fairstore_solr:/var/solr restart: unless-stopped + # CKAN 2.11 logs a critical error on every start without Redis, and its + # background jobs queue lives there. + fairstore-redis: + profiles: ["fairstore"] + image: redis:7-alpine + restart: unless-stopped + # Optional: Redis for caching (uncomment if needed) # redis: # image: redis:7-alpine diff --git a/docker/fairstore/Dockerfile b/docker/fairstore/Dockerfile index 59a688b..3361ca3 100644 --- a/docker/fairstore/Dockerfile +++ b/docker/fairstore/Dockerfile @@ -1,4 +1,8 @@ -FROM ckan/ckan-dev:2.11.3 +# ckan-dev runs Flask's debug server (with the debug toolbar) for local work. +# A deployed Fair Store builds from ckan-base, the uwsgi production image; +# both share the same start scripts and /docker-entrypoint.d hooks. +ARG CKAN_BASE_IMAGE=ckan/ckan-dev:2.11.6 +FROM ${CKAN_BASE_IMAGE} # Keep the extension revision deterministic and aligned with the CKAN 2.11 # version used by the hosted Fair Store. This revision's CI covers CKAN 2.11. @@ -7,5 +11,9 @@ ARG CKANEXT_HIERARCHY_SHA=53c1ee74805d79aef7c9c01cc513eb881cb8928f USER root RUN pip install --no-cache-dir \ "git+https://github.com/ckan/ckanext-hierarchy.git@${CKANEXT_HIERARCHY_SHA}#egg=ckanext-hierarchy" +COPY pin-secrets.sh /docker-entrypoint.d/10-pin-secrets.sh +COPY templates.sh /docker-entrypoint.d/20-templates.sh +COPY secure-cookies.sh /docker-entrypoint.d/30-secure-cookies.sh +COPY templates /srv/app/fairstore_templates USER ckan diff --git a/docker/fairstore/init-db.sh b/docker/fairstore/init-db.sh index 80a01fc..3c23a06 100755 --- a/docker/fairstore/init-db.sh +++ b/docker/fairstore/init-db.sh @@ -6,16 +6,20 @@ # plain Postgres and creates exactly the users and databases CKAN needs here. # Runs once, on an empty data directory. # -# Local-only credentials. A deployed Fair Store needs real secrets (see the -# real secrets managed outside this repo). +# Passwords come from CKAN_DB_PASSWORD / DATASTORE_DB_PASSWORD. The defaults +# are local-only; a deployed Fair Store sets real ones (deploy/fairstore). +# They are passed as psql variables and quoted by psql (:'var'), not spliced +# into the SQL by the shell. set -e -psql -v ON_ERROR_STOP=1 --username "$POSTGRES_USER" <<-SQL - CREATE USER ckan WITH PASSWORD 'ckan'; +psql -v ON_ERROR_STOP=1 --username "$POSTGRES_USER" \ + -v ckan_pw="${CKAN_DB_PASSWORD:-ckan}" \ + -v datastore_pw="${DATASTORE_DB_PASSWORD:-datastore}" <<-SQL + CREATE USER ckan WITH PASSWORD :'ckan_pw'; CREATE DATABASE ckan OWNER ckan; - CREATE USER ckan_datastore_write WITH PASSWORD 'datastore'; - CREATE USER ckan_datastore_read WITH PASSWORD 'datastore'; + CREATE USER ckan_datastore_write WITH PASSWORD :'datastore_pw'; + CREATE USER ckan_datastore_read WITH PASSWORD :'datastore_pw'; CREATE DATABASE datastore OWNER ckan_datastore_write; SQL diff --git a/docker/fairstore/pin-secrets.sh b/docker/fairstore/pin-secrets.sh new file mode 100644 index 0000000..53d5b06 --- /dev/null +++ b/docker/fairstore/pin-secrets.sh @@ -0,0 +1,16 @@ +# Pin CKAN's signing secrets from the environment (sourced by the stock +# start scripts from /docker-entrypoint.d, just before the server starts). +# +# Those scripts write fresh random secrets into the container's ckan.ini +# whenever it is created, so recreating the container — a deploy, a config +# change — silently signed out every session and invalidated every API token, +# including the one the mirror uses. With CKAN_SECRET_KEY set, the secrets +# survive. Unset, the stock random-per-container behaviour is kept. +if [ -n "${CKAN_SECRET_KEY:-}" ]; then + ckan config-tool "$CKAN_INI" \ + "SECRET_KEY=${CKAN_SECRET_KEY}" \ + "WTF_CSRF_SECRET_KEY=${CKAN_SECRET_KEY}" \ + "api_token.jwt.encode.secret=string:${CKAN_SECRET_KEY}" \ + "api_token.jwt.decode.secret=string:${CKAN_SECRET_KEY}" + echo "pin-secrets: signing secrets taken from CKAN_SECRET_KEY" +fi diff --git a/docker/fairstore/secure-cookies.sh b/docker/fairstore/secure-cookies.sh new file mode 100644 index 0000000..bc1b9b9 --- /dev/null +++ b/docker/fairstore/secure-cookies.sh @@ -0,0 +1,21 @@ +# Mark CKAN's session and remember-me cookies Secure when the site is served +# over HTTPS (sourced by the stock start scripts from /docker-entrypoint.d). +# +# The image's ckan.ini ships both as false, and uppercase Flask keys cannot be +# set through the envvars plugin (it lowercases names), hence config-tool. +# Local dev on http://localhost keeps them false: a browser never sends a +# Secure cookie back over plain HTTP, so nobody could stay logged in. +# +# The image also ships REMEMBER_COOKIE_SAMESITE = None, which lets another site +# make a remembered sysadmin's browser send the one-year cookie with a +# cross-site POST (CKAN's DataStore dictionary form is CSRF-exempt); Lax +# works over plain HTTP too, so it is set everywhere. +ckan config-tool "$CKAN_INI" "REMEMBER_COOKIE_SAMESITE=Lax" +case "${CKAN_SITE_URL:-}" in + https://*) + ckan config-tool "$CKAN_INI" \ + "SESSION_COOKIE_SECURE=true" \ + "REMEMBER_COOKIE_SECURE=true" + echo "secure-cookies: session cookies marked Secure" + ;; +esac diff --git a/docker/fairstore/templates.sh b/docker/fairstore/templates.sh new file mode 100644 index 0000000..e10b968 --- /dev/null +++ b/docker/fairstore/templates.sh @@ -0,0 +1,3 @@ +# Load the Fair Store's template overrides (docker/fairstore/templates), which +# render the qsv AI summary on dataset pages. Sourced from /docker-entrypoint.d. +ckan config-tool "$CKAN_INI" "extra_template_paths = /srv/app/fairstore_templates" diff --git a/docker/fairstore/templates/package/read.html b/docker/fairstore/templates/package/read.html new file mode 100644 index 0000000..3b6df80 --- /dev/null +++ b/docker/fairstore/templates/package/read.html @@ -0,0 +1,24 @@ +{% ckan_extends %} + +{#- The dataset's qsv profile (scripts/populate_fairstore.py): the summary qsv + describegpt wrote, rendered as markdown under the source's own + description. The mirror stores only the prose; describegpt's provenance + block is carried by qsv_model / qsv_version and the describegpt resource. -#} +{% block package_notes %} + {{ super() }} + {% set qsv_text = pkg.extras | selectattr('key', 'equalto', 'qsv_description') | map(attribute='value') | first %} + {% if qsv_text %} + {% set qsv_model = pkg.extras | selectattr('key', 'equalto', 'qsv_model') | map(attribute='value') | first %} +
+

AI summary

+
+ {{ h.render_markdown(qsv_text) }} +
+

+ Written by qsv describegpt{% if qsv_model %} ({{ qsv_model }}){% endif %} from a profile of this + dataset's data; AI-written text may contain inaccuracies. The profile's data dictionary + and its other qsv resources are listed under Data and Resources. +

+
+ {% endif %} +{% endblock %} diff --git a/docker/fairstore/templates/package/snippets/additional_info.html b/docker/fairstore/templates/package/snippets/additional_info.html new file mode 100644 index 0000000..abd0663 --- /dev/null +++ b/docker/fairstore/templates/package/snippets/additional_info.html @@ -0,0 +1,13 @@ +{% ckan_extends %} + +{#- The qsv AI summary is rendered as markdown by package/read.html, so its raw + text is left out of the key/value table. -#} +{% block extras scoped %} + {% for extra in h.sorted_extras(pkg_dict.extras, exclude=['qsv_description']) %} + {% set key, value = extra %} + + {{ _(key|e) }} + {{ value }} + + {% endfor %} +{% endblock %} diff --git a/scripts/populate_fairstore.py b/scripts/populate_fairstore.py index 4283b9d..651754a 100644 --- a/scripts/populate_fairstore.py +++ b/scripts/populate_fairstore.py @@ -1,36 +1,82 @@ #!/usr/bin/env python -"""Mirror a source CKAN catalog into the Fair Store. - -The mirror keeps source organization, group, dataset, and resource names and -UUIDs. A top-level organization represents the source portal; source root -organizations are attached beneath it through ckanext-hierarchy. Available -qsv profiling metadata from ``data/ckan_onboard//index.json`` is added to -the corresponding resource without replacing the source description. +"""Mirror source catalogs into the Fair Store. + +Both portal types in the registry are mirrored: + +``ckan`` + Source organization, group, dataset, and resource names and UUIDs are kept. +``dcat`` + A DCAT-US catalog (``/data.json`` from Socrata, DKAN, ArcGIS Hub, ...) has + no organization records or UUIDs, so they are derived deterministically: + publishers become organizations, themes become groups, distributions become + resources, and every UUID is a UUIDv5 of the source identifier (or the + source's own UUID, when it publishes one). A Socrata catalog is enriched + from the Socrata Discovery API, whose owning agency and column list the + catalog document omits. + +A top-level organization represents each source portal, and the portal's +publishers are attached beneath it through ckanext-hierarchy. Every field a +source publishes is kept: fields the Fair Store's default schema has no column +for are carried as extras instead of being silently dropped. Available qsv +profiling metadata from ``data/{ckan,dcat}_onboard//index.json`` is added +to the corresponding resource without replacing the source description. + +A source record that its portal itself mirrored from another portal (it +carries ``mirror_source_portal`` naming that portal) is skipped: mirror the +origin portal directly so provenance names the real source. The command is read-only unless ``--apply`` is supplied. Before any write it checks the whole source snapshot for name and UUID collisions. Re-running is -idempotent: existing objects are patched by their preserved source UUIDs. +idempotent: existing objects are patched by their preserved source UUIDs, and a +record the source renamed is renamed. Every read fails closed: a portal that +errors mid-read aborts that portal's run rather than being mirrored from a +partial snapshot, which would move the unseen records' datasets to the root. +Records the source has deleted are reported (``withdrawn_at_source``), never +deleted. """ from __future__ import annotations import argparse import asyncio +import csv +import hashlib import json +import logging import os import re import sys +import tempfile import uuid from collections.abc import Iterable from pathlib import Path from typing import Any +from urllib.parse import urlparse + +import httpx _PROJECT_ROOT = Path(__file__).resolve().parent.parent sys.path.insert(0, str(_PROJECT_ROOT / "src")) -from data_concierge.data_layer.connectors.ckan import CKANClient # noqa: E402 +from data_concierge.data_layer.connectors.ckan import CKANActionError, CKANClient # noqa: E402 +from data_concierge.data_layer.connectors.dcat import ( # noqa: E402 + DCATClient, + _as_list, + _text, + parse_dataset, + raw_dataset_nodes, +) from data_concierge.data_layer.onboard_index import _scrub_secrets # noqa: E402 -from data_concierge.gateway.ckan_sites import get_site # noqa: E402 +from data_concierge.data_layer.qsv_profiling import ( # noqa: E402 + MIRRORED_QSV_OUTPUTS, + _extract_qsv_tags, +) +from data_concierge.gateway.ckan_sites import ( # noqa: E402 + PORTAL_TYPE_DCAT, + get_site, + list_sites, + normalize_portal_type, +) _SLUG_RE = re.compile(r"[^a-z0-9-]+") _PROVENANCE_KEYS = { @@ -67,29 +113,144 @@ "private", "plugin_data", ) -_RESOURCE_FIELDS = ( - "id", +_CLEARABLE_PACKAGE_FIELDS = ( + "author", + "author_email", + "maintainer", + "maintainer_email", + "license_id", "url", - "description", - "format", - "hash", - "name", - "resource_type", - "url_type", - "mimetype", - "mimetype_inner", - "cache_url", - "size", - "created", - "last_modified", - "cache_last_updated", + "version", +) +# Package keys CKAN computes or the mirror sets itself. Anything else a source +# returns — ckanext-scheming fields such as a data steward or temporal +# coverage — is source metadata the Fair Store's default schema would drop, so +# it is carried as an extra. +_PACKAGE_MANAGED = frozenset( + { + *_PACKAGE_FIELDS, + "owner_org", + "organization", + "groups", + "tags", + "tag_string", + "extras", + "resources", + "license_title", + "license_url", + "isopen", + "num_resources", + "num_tags", + "metadata_created", + "metadata_modified", + "creator_user_id", + "relationships_as_object", + "relationships_as_subject", + "revision_id", + "tracking_summary", + } +) +# Resource keys describing the source CKAN's own bookkeeping rather than the +# resource. ``datastore_*`` keys are excluded separately: the Fair Store holds +# metadata only and must never claim that a DataStore table was copied. +_RESOURCE_MANAGED = frozenset( + {"package_id", "position", "state", "metadata_modified", "revision_id", "tracking_summary"} ) +# url_type values for which CKAN rewrites ``url`` to the *serving* site's own +# download route. Mirrored as-is, the link would point at a file the Fair Store +# never received, so the source's absolute URL is kept as a plain link. +_LOCAL_URL_TYPES = frozenset({"upload", "datastore"}) +# Longest term Solr will index (bytes, UTF-8). CKAN indexes every dataset extra +# as one such term; resource extras are not indexed, so they have no limit. +_SOLR_MAX_TERM_BYTES = 32766 + +_TAG_VALID_RE = re.compile(r"[\w \-.]{2,100}") +_TAG_INVALID_RE = re.compile(r"[^\w \-.]+") +_EMAIL_RE = re.compile(r"[^@\s]+@[^@\s]+\.[^@\s]+") + +_SOCRATA_VIEW_RE = re.compile(r"/api/views/([a-z0-9]{4}-[a-z0-9]{4})/?$") +_SOCRATA_DISCOVERY_URL = "https://api.us.socrata.com/api/catalog/v1" +# Socrata domains name their owning-agency field themselves, so match the +# usual spellings; ``attribution`` is the standard fallback. +_SOCRATA_OWNER_KEY_RE = re.compile(r"business[-_ ]?owner|agency|department|publisher", re.I) + +_MEDIA_TYPE_FORMATS = { + "text/csv": "CSV", + "application/csv": "CSV", + "text/tab-separated-values": "TSV", + "application/json": "JSON", + "application/geo+json": "GeoJSON", + "application/vnd.geo+json": "GeoJSON", + "application/xml": "XML", + "text/xml": "XML", + "application/rdf+xml": "RDF", + "application/vnd.google-earth.kml+xml": "KML", + "application/vnd.google-earth.kmz": "KMZ", + "application/vnd.ms-excel": "XLS", + "application/vnd.openxmlformats-officedocument.spreadsheetml.sheet": "XLSX", + "application/zip": "ZIP", + "application/pdf": "PDF", + "text/html": "HTML", +} + +# DCAT publishes licenses as URLs; CKAN's license register uses short IDs. +# Keys are normalized by _license_key. Unknown URLs are kept only in extras. +_LICENSE_IDS = { + "usa.gov/government-works": "other-pd", + "usa.gov/publicdomain/label/1.0": "other-pd", + "creativecommons.org/publicdomain/zero/1.0": "cc-zero", + "creativecommons.org/publicdomain/mark/1.0": "other-pd", + "creativecommons.org/licenses/by/4.0": "cc-by", + "creativecommons.org/licenses/by/3.0": "cc-by", + "creativecommons.org/licenses/by-sa/4.0": "cc-by-sa", + "creativecommons.org/licenses/by-sa/3.0": "cc-by-sa", + "creativecommons.org/licenses/by-nc/4.0": "cc-nc", + "opendatacommons.org/licenses/pddl/1.0": "odc-pddl", + "opendatacommons.org/licenses/by/1.0": "odc-by", + "opendatacommons.org/licenses/odbl/1.0": "odc-odbl", + "gnu.org/licenses/fdl-1.3": "gfdl", +} class MirrorError(RuntimeError): """Raised when a collision or failed CKAN action makes a mirror unsafe.""" +class DeletedInTarget(MirrorError): + """The object exists in the Fair Store but an admin deleted it there.""" + + +_ATTEMPTS = 3 + + +async def _read( + client: CKANClient, + action: str, + params: dict[str, Any], + *, + label: str, + missing_ok: bool = False, +) -> Any: + """A read that fails closed; ``None`` for a missing object when ``missing_ok``. + + ``CKANClient.action`` returns ``{}`` for any failure, which a mirror reads + as "nothing there": one 502 while paging a portal's organizations would + leave the unseen ones out of the snapshot, and ``--apply`` would then move + their datasets to the portal root. Transient failures are retried; anything + else, or a transient failure that persists, aborts the run. + """ + for attempt in range(_ATTEMPTS): + try: + return await client.call(action, params) + except CKANActionError as exc: + if missing_ok and exc.not_found: + return None + if not exc.transient or attempt == _ATTEMPTS - 1: + raise MirrorError(f"{action} failed for {label}: {exc}") from exc + await asyncio.sleep(2**attempt) + raise AssertionError("unreachable") + + def _slug(text: str, fallback: str) -> str: """Return a CKAN-safe slug while retaining the legacy helper API.""" slug = _SLUG_RE.sub("-", (text or "").lower()).strip("-") @@ -99,20 +260,125 @@ def _slug(text: str, fallback: str) -> str: return slug[:100] -def _load_index(site: str, index_path: str | Path | None = None) -> dict[str, Any]: - """Load qsv output when present; mirroring itself does not require it.""" - candidates = ( - [Path(index_path)] - if index_path - else [ +_ONBOARD_PREFIXES = ("ckan_onboard", "dcat_onboard") +# What the mirror reads from an onboarding directory besides index.json. +_QSV_OUTPUTS = MIRRORED_QSV_OUTPUTS +# A dataset directory name as Path.name yields it: no separators, not . or .. +_SAFE_DIR_RE = re.compile(r"^(?!\.{1,2}$)[^/\\\x00-\x1f]+$") + + +def _load_index( + site: str, + index_path: str | Path | None = None, + *, + source: str = "local", + stage_root: Path | None = None, +) -> dict[str, Any]: + """Load qsv output when present; mirroring itself does not require it. + + ``source`` is ``local`` (this checkout's onboarding directories), ``storage`` + (the storage backend, which onboarding syncs to — GCS on Cloud Run), or + ``auto`` (local, else storage). Output read from storage is staged under + ``stage_root`` first, because the profile vetting reads files. + """ + if index_path: + return json.loads(Path(index_path).read_text(encoding="utf-8")) + if source in ("auto", "local"): + for path in ( _PROJECT_ROOT / "data" / "ckan_onboard" / site / "index.json", + _PROJECT_ROOT / "data" / "dcat_onboard" / site / "index.json", _PROJECT_ROOT / "ckan_onboard" / site / "index.json", - ] + ): + if path.exists(): + return json.loads(path.read_text(encoding="utf-8")) + if source == "local": + return {} + root = stage_root or Path(tempfile.mkdtemp(prefix="fairstore-qsv-")) + return _stage_from_storage(site, root) + + +def _stage_from_storage(site: str, root: Path) -> dict[str, Any]: + """Copy one portal's onboarding output from the storage backend to ``root``. + + Returns its index with each ``local_path`` pointed at the staged directory. + Only what the mirror reads is copied (the profiled CSVs are not synced, and + are not needed). The ``sync_manifest.json`` that ``sync_to_storage`` writes + stands in for the local directory: each file's header (checked against its + qsv output), which dataset directories existed, and which qsv outputs they + held. A directory it does not list is not created, so, as locally, its qsv + resources are not removed; a stale object it does not list is not staged. + """ + from data_concierge.data_layer.storage import storage + + index: dict[str, Any] | None = None + prefix = "" + for prefix in _ONBOARD_PREFIXES: + index = storage.read_json(f"{prefix}/{site}/index.json") + if index: + break + if not index: + return {} + manifest = storage.read_json(f"{prefix}/{site}/sync_manifest.json") or {} + headers = manifest.get("headers") + dirs = manifest.get("dirs") + if not isinstance(headers, dict) or not isinstance(dirs, dict): + # Without it, output left by another file in the same directory passes + # the vetting and failed profiles lose their stats, so the run would + # rewrite (and delete) qsv resources a local run keeps. No profiles: + # the mirror's qsv guard then refuses to write. + print( + f"{site}: the onboarding output in the storage backend has no " + f"sync_manifest.json (it predates it); re-sync it with onboard_*.py " + f"--rebuild-index before mirroring its qsv profiles from storage", + file=sys.stderr, + ) + return {} + + staged: dict[str, Path] = {} + for dataset in index.get("datasets", []) or []: + for resource in dataset.get("resources", []) or []: + local_path = Path(str(resource.get("local_path") or "")) + # onboard_*.py write each resource to ///, + # which sync_to_storage stores under ///. + dataset_dir = local_path.parent.name + key_dir = f"{prefix}/{site}/{dataset_dir}" + listed = dirs.get(key_dir) + if ( + not local_path.name + or not _SAFE_DIR_RE.match(dataset_dir) + or not isinstance(listed, list) + ): + # Never fall back to a local path: point at a directory that + # does not exist, which is what the synced directory lacked. + resource["local_path"] = str(root / "_absent" / (local_path.name or "file")) + continue + if key_dir not in staged: + directory = root / prefix / site / dataset_dir + directory.mkdir(parents=True, exist_ok=True) + for name in _QSV_OUTPUTS: + if name not in listed: + continue + data = storage.read_bytes(f"{key_dir}/{name}") + if data is None: + # Listed but gone: staging less would remove its qsv + # resources, so no profiles (the qsv guard then refuses). + print( + f"{site}: {key_dir}/{name} is listed in sync_manifest.json but " + f"missing from the storage backend; re-sync the onboarding output", + file=sys.stderr, + ) + return {} + (directory / name).write_bytes(data) + staged[key_dir] = directory + resource["local_path"] = str(staged[key_dir] / local_path.name) + header = headers.get(f"{key_dir}/{local_path.name}") + if isinstance(header, list): + resource["profiled_header"] = [str(name) for name in header] + print( + f"{site}: staged qsv output for {len(staged)} datasets from the storage backend", + file=sys.stderr, ) - for path in candidates: - if path.exists(): - return json.loads(path.read_text(encoding="utf-8")) - return {} + return index def _column_dictionary(resource: dict[str, Any]) -> list[dict[str, Any]]: @@ -148,32 +414,198 @@ def _column_dictionary(resource: dict[str, Any]) -> list[dict[str, Any]]: def _qsv_by_resource(index: dict[str, Any]) -> dict[str, dict[str, Any]]: return { - str(resource["resource_id"]): resource + str(resource["resource_id"]): _vet_profile(resource) for dataset in index.get("datasets", []) or [] for resource in dataset.get("resources", []) or [] if resource.get("resource_id") } +def _raise_csv_field_limit() -> None: + # Python's csv module refuses fields over 128 KB, which qsv emits for + # geometry columns. + csv.field_size_limit(max(csv.field_size_limit(), 64 * 1024 * 1024)) + + +def _profiled_header(qsv: dict[str, Any]) -> list[str] | None: + """Column names of the file qsv profiled, while it is still on disk.""" + local_path = str(qsv.get("local_path") or "") + if not local_path: + return None + path = Path(local_path) + path = path if path.is_absolute() else _PROJECT_ROOT / path + if not path.is_file(): + # Staged from the storage backend: the header recorded at sync time. + recorded = qsv.get("profiled_header") + return [str(name) for name in recorded if name] if isinstance(recorded, list) else None + _raise_csv_field_limit() + with path.open(newline="", encoding="utf-8-sig", errors="replace") as handle: + return [name for name in next(csv.reader(handle), []) if name] + + +def _qsv_fields(path: Path) -> set[str] | None: + """The columns a qsv stats or frequency CSV reports on.""" + if not path.is_file(): + return None + _raise_csv_field_limit() + with path.open(newline="", encoding="utf-8", errors="replace") as handle: + return {row.get("field") or "" for row in csv.DictReader(handle)} - {""} + + +def _vet_profile(qsv: dict[str, Any]) -> dict[str, Any]: + """A profile checked against the file it names, ready to publish. + + onboard_*.py writes ``qsv_stats.csv`` and ``qsv_frequency.csv`` once per + dataset directory, so when a dataset has several CSVs and one profile + fails, the directory can hold another file's output — two WPRDC datasets + showed another file's columns and row counts under the profiled file's + name. Output that does not match the profiled file's header is left out, + as are the statistics merged from it. Tags are re-read from describegpt's + own output, whose key is sometimes capitalised. + """ + if "_vetted" in qsv: + return qsv + profile = dict(qsv, _vetted=True) + header = _profiled_header(qsv) + directory = _qsv_dir(qsv) + trusted = str(qsv.get("status") or "") == "ok" + + def describes(name: str, *, exact: bool) -> bool: + fields = _qsv_fields(directory / name) if directory else None + if fields is None: + return False + if header is None: + return trusted + return fields == set(header) if exact else fields <= set(header) + + profile["_stats_ok"] = describes("qsv_stats.csv", exact=True) + profile["_frequency_ok"] = describes("qsv_frequency.csv", exact=False) + profile["_header"] = header + if header is None and not trusted: + # A failed profile whose file is gone: its merged statistics cannot be + # checked against the file, so only names, labels and types remain. + profile["columns"] = [ + {**column, "stats": {}, "top_values": []} for column in qsv.get("columns") or [] + ] + elif header is not None and not (profile["_stats_ok"] and profile["_frequency_ok"]): + # Keep the file's own columns and the portal's DataStore fields; drop + # columns only the mismatched output contributed. + profile["columns"] = [ + { + **column, + "stats": column.get("stats") if profile["_stats_ok"] else {}, + "top_values": column.get("top_values") if profile["_frequency_ok"] else [], + } + for column in qsv.get("columns") or [] + if column.get("name") in header or column.get("ckan_info") or column.get("ckan_type") + ] + + # describegpt's own dictionary names the columns it described; like the + # stats, it can belong to another file in the same directory. + describegpt = directory / "qsv_dict.json" if directory else None + try: + output = ( + json.loads(describegpt.read_text(encoding="utf-8")) + if describegpt is not None and describegpt.is_file() + else None + ) + except (ValueError, OSError): + output = None + described = { + str(field.get("name")) + for field in ((output or {}).get("Dictionary") or {}).get("response", {}).get("fields") + or [] + if isinstance(field, dict) and field.get("name") + } + if output is None: + profile["_describegpt_ok"] = False + elif header is None or not described: + profile["_describegpt_ok"] = trusted + else: + profile["_describegpt_ok"] = described - {"_id"} <= set(header) + if profile["_describegpt_ok"]: + tags = _extract_qsv_tags(output or {}) + if tags: + profile["qsv_tags"] = tags + elif output is not None: + # Another file's AI output: its summary, tags and column labels go too. + profile["qsv_description"] = "" + profile["qsv_tags"] = [] + profile["columns"] = [ + { + key: value + for key, value in column.items() + if key not in ("qsv_label", "qsv_description") + } + for column in profile.get("columns") or [] + ] + return profile + + def _project(source: dict[str, Any], fields: Iterable[str]) -> dict[str, Any]: return { field: source[field] for field in fields if field in source and source[field] is not None } +def _is_empty(value: Any) -> bool: + return value is None or value == "" or value == [] or value == {} + + +def _extra_value(value: Any) -> str: + """Dataset extras are strings in CKAN; structured values are kept as JSON.""" + if isinstance(value, str): + return value + return json.dumps(value, sort_keys=True, ensure_ascii=False) + + +def _source_id(row: dict[str, Any]) -> str: + """The identifier the source itself uses for a record. + + For a CKAN source that is the preserved UUID. A DCAT snapshot records the + catalog's own identifier under ``_source_id`` because its UUIDs are derived. + """ + return str(row.get("_source_id") or row.get("id", "")) + + +def _with_image(payload: dict[str, Any], row: dict[str, Any]) -> dict[str, Any]: + """Point an organization or group image at a URL that resolves on the Fair Store. + + An uploaded image is stored as a bare filename that CKAN resolves against + the serving site's own ``/uploads``, where the mirror never put the file, + so the source's absolute ``image_display_url`` is used instead. + """ + image = str(row.get("image_url") or "") + if image and not image.startswith(("http://", "https://")): + image = str(row.get("image_display_url") or "") + if image: + payload["image_url"] = image + else: + payload.pop("image_url", None) + return payload + + def _provenance_extras( extras: list[dict[str, Any]] | None, *, site_id: str, source_url: str, source_id: str, + carried: dict[str, str] | None = None, ) -> list[dict[str, str]]: - """Preserve source extras and add stable mirror identity fields.""" - values = { - str(item.get("key")): str(item.get("value", "")) - for item in (extras or []) - if item.get("key") - } + """Preserve source extras and add stable mirror identity fields. + + ``carried`` holds source fields with no column in the target schema; a + source extra of the same name wins, and the identity fields win over both. + """ + values = dict(carried or {}) + values.update( + { + str(item.get("key")): str(item.get("value", "")) + for item in (extras or []) + if item.get("key") + } + ) values.update( { _PROVENANCE_KEYS["portal"]: site_id, @@ -184,6 +616,14 @@ def _provenance_extras( return [{"key": key, "value": value} for key, value in values.items()] +def _mirrored_from(row: dict[str, Any]) -> str: + """Name the portal a source record was itself mirrored from, if any.""" + for item in row.get("extras") or []: + if isinstance(item, dict) and item.get("key") == _PROVENANCE_KEYS["portal"]: + return str(item.get("value") or "").strip() + return "" + + def _group_refs(groups: list[dict[str, Any]] | None) -> list[dict[str, str]]: refs: list[dict[str, str]] = [] for group in groups or []: @@ -196,31 +636,81 @@ def _group_refs(groups: list[dict[str, Any]] | None) -> list[dict[str, str]]: return refs -async def _all_packages(client: CKANClient) -> list[dict[str, Any]]: +def _clean_tag(raw: str) -> str: + """Return ``raw`` as a valid CKAN tag, or ``""`` when nothing usable is left. + + CKAN accepts 2-100 word characters, spaces, hyphens and dots. Valid tags pass + through untouched; DCAT keywords such as ``l&i`` are rewritten (``l-i``). + """ + tag = raw.strip() + if _TAG_VALID_RE.fullmatch(tag): + return tag + tag = re.sub(r"-{2,}", "-", _TAG_INVALID_RE.sub("-", tag)).strip(" -") + tag = tag[:100].strip() + return tag if len(tag) >= 2 else "" + + +def _tag_refs(tags: list[dict[str, Any]] | None) -> list[dict[str, str]]: + refs: list[dict[str, str]] = [] + seen: set[str] = set() + for tag in tags or []: + name = _clean_tag(str(tag.get("name") or "")) + if not name or name in seen: + continue + seen.add(name) + refs.append( + { + "name": name, + **({"vocabulary_id": tag["vocabulary_id"]} if tag.get("vocabulary_id") else {}), + } + ) + return refs + + +async def _all_packages( + client: CKANClient, *, label: str = "portal", include_private: bool = False +) -> list[dict[str, Any]]: + """Every dataset the portal lists; raises rather than return part of them.""" + params: dict[str, Any] = {"q": "*:*", "rows": 1000, "sort": "name asc"} + if include_private: + # The Fair Store's own private and draft datasets still hold their + # UUIDs: left out, they would look new and never be updated. + params.update(include_private=True, include_drafts=True) packages: list[dict[str, Any]] = [] + seen_ids: set[str] = set() start = 0 while True: - page = await client.action( - "package_search", {"q": "*:*", "rows": 1000, "start": start, "sort": "name asc"} - ) - batch = page.get("results", []) if isinstance(page, dict) else [] - packages.extend(batch) + page = await _read(client, "package_search", {**params, "start": start}, label=label) + if not isinstance(page, dict): + raise MirrorError(f"package_search for {label} returned {type(page).__name__}") + batch = page.get("results") or [] + count = int(page.get("count") or 0) + packages.extend(row for row in batch if str(row.get("id")) not in seen_ids) + seen_ids.update(str(row.get("id")) for row in batch) start += len(batch) - if not batch or start >= int(page.get("count", 0)): + if start >= count: return packages + if not batch: + raise MirrorError(f"package_search for {label} stopped at {start} of {count} datasets") -async def _all_group_rows(client: CKANClient, action_name: str) -> list[dict[str, Any]]: +async def _all_group_rows( + client: CKANClient, action_name: str, *, label: str = "portal", include_extras: bool = False +) -> list[dict[str, Any]]: """Read every organization or group despite portal-side page-size caps.""" + # Sorted by name, which is unique: CKAN's default (title) order is not a + # total order, so offset paging over it can skip a row. + params: dict[str, Any] = {"all_fields": True, "limit": 1000, "sort": "name asc"} + if include_extras: + params["include_extras"] = True rows: list[dict[str, Any]] = [] seen_ids: set[str] = set() offset = 0 while True: - result = await client.action( - action_name, - {"all_fields": True, "limit": 1000, "offset": offset}, - ) - batch = list(result) if isinstance(result, list) else [] + result = await _read(client, action_name, {**params, "offset": offset}, label=label) + if not isinstance(result, list): + raise MirrorError(f"{action_name} for {label} returned {type(result).__name__}") + batch = list(result) if not batch: return rows new_rows = [row for row in batch if str(row.get("id")) not in seen_ids] @@ -231,15 +721,64 @@ async def _all_group_rows(client: CKANClient, action_name: str) -> list[dict[str offset += len(batch) +def _empty_copies() -> dict[str, Any]: + return {"organizations": 0, "groups": 0, "datasets": 0, "origins": []} + + +async def _check_complete( + client: CKANClient, action_name: str, rows: list[dict[str, Any]], *, label: str +) -> None: + """Cross-check paged ``all_fields`` rows against the plain name listing. + + The name listing is one unpaged call: WPRDC's CDN answers 403 to a + names-only listing that carries limit/offset. CKAN returns up to 1,000 + names that way; a portal at that cap is not checked. + """ + listed = await _read(client, action_name, {"sort": "name asc"}, label=label) + if not isinstance(listed, list): + raise MirrorError(f"{action_name} for {label} returned {type(listed).__name__}") + names = {str(item.get("name") if isinstance(item, dict) else item) for item in listed} + if len(listed) >= 1000: + print(f"warning: {action_name} for {label}: too many to cross-check", file=sys.stderr) + return + missing = names - {str(row.get("name")) for row in rows} + if missing: + raise MirrorError( + f"{action_name} for {label} paged {len(rows)} of {len(names)} records; " + f"missing e.g. {sorted(missing)[:5]}" + ) + + async def _source_snapshot( - client: CKANClient, *, organization: str | None = None, limit: int | None = None -) -> dict[str, list[dict[str, Any]]]: - organization_rows = await _all_group_rows(client, "organization_list") + client: CKANClient, + *, + site_id: str | None = None, + organization: str | None = None, + limit: int | None = None, +) -> dict[str, Any]: + """Read a CKAN source, leaving out records it mirrored from another portal.""" + copies = _empty_copies() + origins: set[str] = set() + + def is_copy(row: dict[str, Any], kind: str) -> bool: + origin = _mirrored_from(row) + if site_id and origin and origin != site_id: + copies[kind] += 1 + origins.add(origin) + return True + return False + + label = f"source {site_id or client.ckan_url}" + organization_rows = await _all_group_rows(client, "organization_list", label=label) + await _check_complete(client, "organization_list", organization_rows, label=label) organizations: list[dict[str, Any]] = [] - for row in organization_rows if isinstance(organization_rows, list) else []: + for row in organization_rows: if organization and row.get("name") != organization: continue - full = await client.action( + # The list row lacks the extras (including any mirror provenance) and + # the parent groups, so it is never a stand-in for the full record. + full = await _read( + client, "organization_show", { "id": row["id"], @@ -248,18 +787,29 @@ async def _source_snapshot( "include_groups": True, "include_tags": True, }, + label=f"{label} organization {row.get('name')}", ) - organizations.append(full or row) - - groups = await _all_group_rows(client, "group_list") + if not is_copy(full, "organizations"): + organizations.append(full) + + # With extras, so a group the source itself mirrored is seen as a copy. + groups = [ + group + for group in await _all_group_rows(client, "group_list", label=label, include_extras=True) + if not is_copy(group, "groups") + ] - packages = await _all_packages(client) + packages = [ + package + for package in await _all_packages(client, label=label) + if not is_copy(package, "datasets") + ] if organization: org_ids = {str(item.get("id")) for item in organizations} packages = [ package for package in packages - if package.get("organization", {}).get("name") == organization + if (package.get("organization") or {}).get("name") == organization or str(package.get("owner_org")) in org_ids ] if limit is not None: @@ -271,28 +821,542 @@ async def _source_snapshot( if group.get("name") } groups = [group for group in groups if group.get("name") in used_group_names] - return {"organizations": organizations, "groups": groups, "packages": packages} + copies["origins"] = sorted(origins) + return { + "organizations": organizations, + "groups": groups, + "packages": packages, + "copies": copies, + } -async def _target_snapshot(client: CKANClient) -> dict[str, list[dict[str, Any]]]: - organizations = await _all_group_rows(client, "organization_list") - groups = await _all_group_rows(client, "group_list") - packages = await _all_packages(client) +def _dcat_uuid(identifier: str, source_url: str, kind: str) -> str: + """Stable UUID for a DCAT record: the source's own UUID when it has one.""" + try: + return str(uuid.UUID(identifier)) + except (ValueError, TypeError): + pass + if urlparse(identifier).scheme in ("http", "https"): + return str(uuid.uuid5(uuid.NAMESPACE_URL, identifier)) + return str(uuid.uuid5(uuid.NAMESPACE_URL, f"{source_url.rstrip('/')}#{kind}/{identifier}")) + + +def _socrata_id(identifier: str) -> str: + match = _SOCRATA_VIEW_RE.search(urlparse(identifier or "").path) + return match.group(1) if match else "" + + +def _dcat_dataset_name(title: str, identifier: str, dataset_id: str) -> str: + """A unique, stable CKAN name: the title slug plus the source's short ID. + + CKAN names are site-wide, and portal titles repeat ("Crashes", "Budget"), so + every name carries the portal's own ID (a Socrata 4x4) or a UUID prefix. + """ + tail = urlparse(identifier).path.rstrip("/").rsplit("/", 1)[-1] if identifier else "" + suffix = tail.lower() if re.fullmatch(r"[a-z0-9][a-z0-9_-]{1,19}", tail.lower()) else "" + suffix = suffix or dataset_id[:8] + base = _slug(title, "dataset")[: 100 - len(suffix) - 1].rstrip("-") + return f"{base}-{suffix}" + + +def _publisher_chain(publisher: Any) -> list[str]: + """Publisher names from most specific to broadest (DCAT-US subOrganizationOf).""" + chain: list[str] = [] + node = publisher + while node and len(chain) < 10: + name = _text(node) + if name and name not in chain: + chain.append(name) + node = node.get("subOrganizationOf") if isinstance(node, dict) else None + return chain + + +def _lineage(parents: dict[str, str], name: str) -> set[str]: + """``name`` and every ancestor recorded for it in ``parents``.""" + seen: set[str] = set() + while name and name not in seen: + seen.add(name) + name = parents.get(name, "") + return seen + + +def _mailto(value: Any) -> str: + email = _text(value).removeprefix("mailto:").strip() + return email if _EMAIL_RE.fullmatch(email) else "" + + +def _license_key(url: str) -> str: + parsed = urlparse(url.strip().lower()) + host = parsed.netloc.removeprefix("www.") + return f"{host}{parsed.path.rstrip('/')}" if host else url.strip().lower() + + +def _dcat_license_id(url: str) -> str: + return _LICENSE_IDS.get(_license_key(url), "") if url else "" + + +def _media_format(media_type: str, declared: str) -> str: + return declared or _MEDIA_TYPE_FORMATS.get(media_type.lower(), "") + + +def _dcat_extras(raw: dict[str, Any]) -> list[dict[str, str]]: + """Every DCAT dataset field as a ``dcat_``-prefixed extra. + + Fields copied into a CKAN core field (title, description, landing page, + version) are not duplicated; keyword, theme, license and contactPoint are + kept verbatim because their CKAN counterparts are normalized. + """ + skipped = {"distribution", "title", "description", "landingPage", "version"} + return [ + {"key": f"dcat_{key.lstrip('@')}", "value": _extra_value(_stable(key, value))} + for key, value in raw.items() + if key not in skipped and not _is_empty(value) + ] + + +# Catalog fields that are sets. Socrata serves them in a different order on +# each fetch, so they are sorted to keep identical content identical. +_UNORDERED_FIELDS = frozenset({"keyword", "theme", "domain_tags"}) + + +def _stable(key: str, value: Any) -> Any: + if key in _UNORDERED_FIELDS and isinstance(value, list): + return sorted(value, key=lambda item: json.dumps(item, sort_keys=True)) + return value + + +def _socrata_owner(row: dict[str, Any]) -> tuple[str, str]: + """The owning agency Socrata records for a dataset, and the field it came from.""" + for item in (row.get("classification") or {}).get("domain_metadata") or []: + key, value = str(item.get("key") or ""), str(item.get("value") or "").strip() + if value and _SOCRATA_OWNER_KEY_RE.search(key): + return value, key + attribution = str((row.get("resource") or {}).get("attribution") or "").strip() + return (attribution, "attribution") if attribution else ("", "") + + +def _socrata_extras(row: dict[str, Any]) -> list[dict[str, str]]: + """Discovery API metadata the DCAT export omits, as ``socrata_`` extras. + + Volatile analytics (page views, downloads) and account names are left out. + """ + resource = row.get("resource") or {} + classification = row.get("classification") or {} + values: dict[str, Any] = { + "socrata_id": resource.get("id"), + "socrata_type": resource.get("type"), + "socrata_attribution": resource.get("attribution"), + "socrata_attribution_link": resource.get("attribution_link"), + "socrata_provenance": resource.get("provenance"), + "socrata_created_at": resource.get("createdAt"), + "socrata_updated_at": resource.get("updatedAt"), + "socrata_data_updated_at": resource.get("data_updated_at"), + "socrata_metadata_updated_at": resource.get("metadata_updated_at"), + "socrata_publication_date": resource.get("publication_date"), + "socrata_category": classification.get("domain_category"), + "socrata_domain_tags": _stable("domain_tags", classification.get("domain_tags")), + "socrata_license": (row.get("metadata") or {}).get("license"), + "socrata_permalink": row.get("permalink"), + } + for item in classification.get("domain_metadata") or []: + if item.get("key"): + values[f"socrata_{item['key']}"] = item.get("value") + return [ + {"key": key, "value": _extra_value(value)} + for key, value in values.items() + if not _is_empty(value) + ] + + +def _socrata_columns(row: dict[str, Any]) -> list[dict[str, str]]: + """The column list Socrata publishes for a dataset, as a data dictionary.""" + resource = row.get("resource") or {} + fields = resource.get("columns_field_name") or [] + + def at(key: str, index: int) -> str: + values = resource.get(key) or [] + return str(values[index] or "") if index < len(values) else "" + + # The column arrays agree with one another but not from fetch to fetch, + # so the dictionary is ordered by field name. + return sorted( + ( + { + "name": str(field), + "label": at("columns_name", index), + "type": at("columns_datatype", index), + "description": _scrub_secrets(at("columns_description", index)), + } + for index, field in enumerate(fields) + ), + key=lambda column: column["name"], + ) + + +def _dcat_resources(raw: dict[str, Any], *, dataset_id: str) -> list[dict[str, Any]]: + """CKAN resources for a dataset's DCAT distributions.""" + identifier = _text(raw.get("identifier") or raw.get("@id")) + resources: list[dict[str, Any]] = [] + used_keys: set[str] = set() + distributions = [d for d in raw.get("distribution") or [] if isinstance(d, dict)] + for index, dist in enumerate(distributions): + download_url = _text(dist.get("downloadURL")) + url = download_url or _text(dist.get("accessURL")) + media_type = _text(dist.get("mediaType")) + fmt = _media_format(media_type, _text(dist.get("format"))) + key = url or str(index) + if key in used_keys: + key = f"{key}#{index}" + used_keys.add(key) + resource: dict[str, Any] = { + "id": str(uuid.uuid5(uuid.NAMESPACE_URL, f"{dataset_id}/{key}")), + "_source_id": url or f"{identifier}#distribution-{index}", + "url": url, + "name": _text(dist.get("title")) or fmt or media_type or f"Distribution {index + 1}", + "description": _text(dist.get("description")), + "format": fmt, + "mimetype": media_type or None, + } + for field_name, value in dist.items(): + if field_name in {"title", "description", "mediaType", "format", "downloadURL"}: + continue + if field_name == "accessURL" and not download_url: + continue + if not _is_empty(value): + resource[f"dcat_{field_name.lstrip('@')}"] = value + resources.append(resource) + return resources + + +def _dcat_snapshot( + catalog: Any, + *, + site_title: str, + source_url: str, + socrata: dict[str, dict[str, Any]] | None = None, + limit: int | None = None, +) -> dict[str, Any]: + """Translate a DCAT catalog into the CKAN-shaped snapshot the mirror writes. + + Organizations come from each dataset's publisher chain, most specific first, + headed by the Socrata owning agency when the Discovery API names one. A + publisher that is the portal itself is represented by the portal's root + organization rather than duplicated beneath it. + """ + nodes = raw_dataset_nodes(catalog) + if limit is not None: + nodes = nodes[:limit] + host = (urlparse(source_url).hostname or "").lower() + portal_names = {host, host.removeprefix("www."), site_title.strip().lower()} - {""} + + org_parent: dict[str, str] = {} + org_field: dict[str, str] = {} + themes: dict[str, str] = {} + packages: list[dict[str, Any]] = [] + names: set[str] = set() + + def org_id(title: str) -> str: + return _dcat_uuid(title, source_url, "publisher") + + for raw in nodes: + identifier = _text(raw.get("identifier") or raw.get("@id")) + title = _text(raw.get("title")) + dataset_id = _dcat_uuid(identifier or title, source_url, "dataset") + row = (socrata or {}).get(_socrata_id(identifier)) + + chain = _publisher_chain(raw.get("publisher")) + fields = ["dcat:publisher"] * len(chain) + if row: + owner, owner_key = _socrata_owner(row) + if owner and owner not in chain: + chain.insert(0, owner) + fields.insert(0, f"socrata:{owner_key}") + kept = [(n, f) for n, f in zip(chain, fields, strict=True) if n.lower() not in portal_names] + for (child, _), (parent, _) in zip(kept, kept[1:], strict=False): + # A parent learned from any dataset fills in one not yet known, but + # a catalog that contradicts itself must not create a loop, which + # ckanext-hierarchy rejects mid-run. + if not org_parent.get(child) and child not in _lineage(org_parent, parent): + org_parent[child] = parent + for name_, field_ in kept: + org_parent.setdefault(name_, "") + org_field.setdefault(name_, field_) + + theme_slugs: list[str] = [] + for theme in _as_list(raw.get("theme")): + slug = _slug(theme, _dcat_uuid(theme, source_url, "theme")[:8]) + themes.setdefault(slug, theme) + if slug not in theme_slugs: + theme_slugs.append(slug) + + name = _dcat_dataset_name(title or identifier, identifier, dataset_id) + if name in names: + name = f"{name[:91]}-{dataset_id[:8]}" + names.add(name) + + contact = raw.get("contactPoint") if isinstance(raw.get("contactPoint"), dict) else {} + extras = _dcat_extras(raw) + (_socrata_extras(row) if row else []) + parsed = parse_dataset(raw) + tabular = parsed.tabular_distributions + first_tabular_url = tabular[0].best_url if tabular else "" + resources = _dcat_resources(raw, dataset_id=dataset_id) + columns = _socrata_columns(row) if row else [] + if columns: + # The column list goes on the file it describes, beside any qsv + # profile. As a dataset extra it would also be indexed as one Solr + # term, and a wide table's list overruns the term limit. + dictionary = { + "source_data_dictionary": json.dumps(columns), + "source_column_count": str(len(columns)), + } + described = next( + (r for r in resources if first_tabular_url and r["url"] == first_tabular_url), + resources[0] if resources else None, + ) + if described is not None: + described.update(dictionary) + else: + extras += [{"key": key, "value": value} for key, value in dictionary.items()] + + packages.append( + { + "id": dataset_id, + "_source_id": identifier or title, + # How onboard_dcat.py files this dataset's qsv profile. + "_short_id": parsed.id, + "_first_tabular_url": first_tabular_url, + "name": name, + "title": title or identifier, + "notes": _text(raw.get("description")), + "url": _text(raw.get("landingPage")) or None, + "version": _text(raw.get("version")) or None, + "maintainer": _text(contact.get("fn")) or None, + "maintainer_email": _mailto(contact.get("hasEmail")) or None, + "license_id": _dcat_license_id(_text(raw.get("license"))) or None, + "owner_org": org_id(kept[0][0]) if kept else "", + "tags": [ + {"name": keyword} for keyword in sorted(_as_list(raw.get("keyword")), key=str) + ], + "groups": [{"name": slug} for slug in theme_slugs], + "extras": extras, + "resources": resources, + } + ) + + # Names are assigned after every publisher is known so they do not depend on + # catalog order; two titles that slug alike are told apart by UUID prefix. + org_names: dict[str, str] = {} + for title in sorted(org_parent): + slug = _slug(title, org_id(title)[:8]) + org_names[title] = ( + slug if slug not in org_names.values() else f"{slug[:91]}-{org_id(title)[:8]}" + ) + organizations = [ + { + "id": org_id(title), + "_source_id": title, + "name": org_names[title], + "title": title, + "groups": ( + [{"name": org_names[org_parent[title]], "capacity": "parent"}] + if org_parent[title] + else [] + ), + "extras": [{"key": "publisher_source", "value": org_field[title]}], + } + for title in sorted(org_parent) + ] + groups = [ + { + "id": _dcat_uuid(theme, source_url, "theme"), + "_source_id": theme, + "name": slug, + "title": theme, + } + for slug, theme in sorted(themes.items()) + ] return { - "organizations": list(organizations) if isinstance(organizations, list) else [], - "groups": list(groups) if isinstance(groups, list) else [], + "organizations": organizations, + "groups": groups, "packages": packages, + "copies": _empty_copies(), + } + + +def _dcat_qsv_index(index: dict[str, Any], snapshot: dict[str, Any]) -> dict[str, Any]: + """Re-key a DCAT onboarding index to the mirrored resource UUIDs. + + ``onboard_dcat.py`` profiles each dataset's first tabular distribution and + files it under the dataset's short ID rather than a resource ID. + """ + by_short_id: dict[str, str] = {} + for package in snapshot["packages"]: + for resource in package.get("resources", []): + if ( + package.get("_first_tabular_url") + and resource.get("url") == package["_first_tabular_url"] + ): + by_short_id[str(package.get("_short_id"))] = str(resource["id"]) + break + rekeyed = [ + {**resource, "resource_id": by_short_id[str(resource.get("resource_id"))]} + for dataset in index.get("datasets", []) or [] + for resource in dataset.get("resources", []) or [] + if str(resource.get("resource_id")) in by_short_id + ] + return {"datasets": [{"resources": rekeyed}]} if rekeyed else {} + + +async def _socrata_page(client: httpx.AsyncClient, domain: str, offset: int) -> dict[str, Any]: + """One Discovery API page, retried when the failure is transient.""" + params = {"domains": domain, "search_context": domain, "limit": 1000, "offset": offset} + for attempt in range(_ATTEMPTS): + try: + resp = await client.get(_SOCRATA_DISCOVERY_URL, params=params) + if resp.status_code != 429 and resp.status_code < 500: + resp.raise_for_status() + body = resp.json() + if isinstance(body, dict): + return body + raise ValueError(f"unexpected {type(body).__name__} body") + error: Exception = httpx.HTTPStatusError( + f"HTTP {resp.status_code}", request=resp.request, response=resp + ) + except httpx.TransportError as exc: + error = exc + if attempt == _ATTEMPTS - 1: + raise error + await asyncio.sleep(2**attempt) + raise AssertionError("unreachable") + + +async def _socrata_metadata(catalog: Any, *, required: bool) -> dict[str, dict[str, Any]]: + """Discovery API records for a Socrata-hosted catalog, keyed by 4x4 ID. + + Returns ``{}`` for a catalog that is not Socrata's. The owning agency comes + only from this API — data.pa.gov's own catalog names one publisher for every + dataset — so when ``required`` (a run that writes) a failure raises instead + of mirroring a snapshot that would move every dataset to the portal root and + drop its ``socrata_*`` extras. Otherwise it warns and returns what it has. + """ + domains = { + (urlparse(_text(node.get("identifier"))).hostname or "") + for node in raw_dataset_nodes(catalog) + if _socrata_id(_text(node.get("identifier"))) + } - {""} + rows: dict[str, dict[str, Any]] = {} + async with httpx.AsyncClient(timeout=60, follow_redirects=True) as client: + for domain in sorted(domains): + offset = 0 + while True: + try: + body = await _socrata_page(client, domain, offset) + except (httpx.HTTPError, ValueError) as exc: + message = f"Socrata enrichment failed for {domain} at offset {offset}: {exc}" + if required: + raise MirrorError( + f"{message}. Retry later; --no-enrich mirrors without it, which " + "moves every dataset to the portal root organization." + ) from exc + print(f"warning: {message}", file=sys.stderr) + break + batch = body.get("results") or [] + for result in batch: + view_id = (result.get("resource") or {}).get("id") + if view_id: + rows[str(view_id)] = result + offset += len(batch) + total = int(body.get("resultSetSize") or 0) + if offset >= total: + break + if not batch: + message = f"Socrata enrichment for {domain} stopped at {offset} of {total}" + if required: + raise MirrorError(message) + print(f"warning: {message}", file=sys.stderr) + break + return rows + + +async def _target_snapshot(client: CKANClient) -> dict[str, list[dict[str, Any]]]: + label = f"target {client.ckan_url}" + # Extras carry the provenance that tells a rename of this portal's own + # record apart from a collision with someone else's. + return { + "organizations": await _all_group_rows( + client, "organization_list", label=label, include_extras=True + ), + "groups": await _all_group_rows(client, "group_list", label=label, include_extras=True), + "packages": await _all_packages(client, label=label, include_private=True), } +def _pin_target_names( + source: dict[str, Any], target: dict[str, list[dict[str, Any]]], *, site_id: str +) -> None: + """Keep the names already published for a DCAT portal's records. + + DCAT names are derived — from a title slug, and for organizations whose + titles slug alike, from sort order — so an edited title or a newly added + publisher would otherwise rename published datasets and break their URLs. + A record this portal already holds keeps its name; a new record whose name + is taken gets its UUID prefix appended. + """ + for kind in ("organizations", "packages"): + held = { + str(row.get("id")): str(row.get("name")) + for row in target[kind] + if _mirrored_from(row) == site_id and row.get("name") + } + rows = source[kind] + renamed: dict[str, str] = {} + for row in rows: + name = held.get(str(row.get("id"))) + if name and name != row.get("name"): + renamed[str(row.get("name"))] = name + row["name"] = name + # A new record may not take any name the Fair Store already uses, nor + # one given to an earlier record of this run. + taken = {str(row.get("name")) for row in target[kind] if row.get("name")} + taken |= {str(row.get("name")) for row in rows if str(row.get("id")) in held} + for row in rows: + row_id, name = str(row.get("id")), str(row.get("name")) + if row_id in held: + continue + if name in taken: + new_name = f"{name[:91]}-{row_id[:8]}" + renamed[name] = new_name + row["name"] = name = new_name + taken.add(name) + if kind == "organizations" and renamed: + # Parent references are by name. + for row in rows: + for parent in row.get("groups") or []: + if parent.get("name") in renamed: + parent["name"] = renamed[parent["name"]] + + def _assert_no_collisions( - source: dict[str, list[dict[str, Any]]], + source: dict[str, Any], target: dict[str, list[dict[str, Any]]], *, store_name: str, store_id: str, -) -> None: + site_id: str = "", + evictions: list[tuple[str, str, str]] | None = None, +) -> list[str]: + """Raise on anything a write would clobber; return the renames it will apply. + + The same UUID under a new name is a rename when the Fair Store's record is + this portal's own (its provenance names ``site_id``): the source renamed + it, and the patch carries the new name. Under any other provenance it is a + collision, as is a name held by a different UUID — unless the holder is this + portal's own record that its source no longer publishes and ``evictions`` + is given: the source reused the name, so the withdrawn record is to be + renamed out of the way, collected as ``(kind, id, new name)``. + """ collisions: list[str] = [] + renames: list[str] = [] def check( kind: str, @@ -303,12 +1367,37 @@ def check( ) -> None: by_name = {str(row.get("name")): row for row in existing if row.get("name")} by_id = {str(row.get("id")): row for row in existing if row.get("id")} + incoming_ids = {str(row.get("id", "")) for row in incoming} + seen_ids: set[str] = set() + seen_names: set[str] = set() for row in incoming: name, row_id = str(row.get("name", "")), str(row.get("id", "")) - if not allow_shared_name and name in by_name and str(by_name[name].get("id")) != row_id: - collisions.append(f"{kind} name {name!r} already has a different UUID") - if row_id in by_id and str(by_id[row_id].get("name")) != name: - collisions.append(f"{kind} UUID {row_id!r} already has a different name") + if row_id in seen_ids: + collisions.append(f"{kind} UUID {row_id!r} appears twice in the source") + if name in seen_names and not allow_shared_name: + collisions.append(f"{kind} name {name!r} appears twice in the source") + seen_ids.add(row_id) + seen_names.add(name) + holder = by_name.get(name) + if not allow_shared_name and holder is not None and str(holder.get("id")) != row_id: + holder_id = str(holder.get("id")) + if ( + evictions is not None + and site_id + and _mirrored_from(holder) == site_id + and holder_id not in incoming_ids + ): + evicted = f"{name[:80]}-withdrawn-{holder_id[:8]}" + evictions.append((kind, holder_id, evicted)) + renames.append(f"{kind} {name!r} (withdrawn at source) -> {evicted!r}") + else: + collisions.append(f"{kind} name {name!r} already has a different UUID") + held = by_id.get(row_id) + if held is not None and str(held.get("name")) != name: + if site_id and _mirrored_from(held) == site_id: + renames.append(f"{kind} {held.get('name')!r} -> {name!r}") + else: + collisions.append(f"{kind} UUID {row_id!r} already has a different name") root_existing = next( (row for row in target["organizations"] if row.get("name") == store_name), None @@ -352,35 +1441,643 @@ def check( if collisions: detail = "\n - ".join(collisions[:50]) raise MirrorError(f"Mirror preflight found {len(collisions)} collision(s):\n - {detail}") + return renames + + +_SHOW_ACTIONS = { + "organization_create": "organization_show", + "group_create": "group_show", + "package_create": "package_show", + "resource_create": "resource_show", +} async def _write_action( - client: CKANClient, action: str, payload: dict[str, Any], *, label: str + client: CKANClient, + action: str, + payload: dict[str, Any], + *, + label: str, + upload: tuple[str, bytes, str] | None = None, ) -> dict[str, Any]: - show_actions = { - "organization_create": "organization_show", - "group_create": "group_show", - "package_create": "package_show", - "resource_create": "resource_show", - } - for attempt in range(3): - result = await client.action(action, payload) - if result: - return result - - # A gateway can time out after CKAN commits the object. Resolve by the - # preserved UUID before retrying so a successful write is not treated - # as a failed duplicate create. - show_action = show_actions.get(action) + """Run a write, retrying transient failures; ``upload`` is (name, bytes, type). + + A create whose object turns out to exist is resolved by its preserved UUID: + a gateway can fail after CKAN commits, and the retry must not be reported + as a failure (or repeated). Writes without a UUID to check — DataStore + loads, view creation — are made safe by their callers instead. + """ + show_action = _SHOW_ACTIONS.get(action) + for attempt in range(_ATTEMPTS): + try: + if upload: + filename, content, content_type = upload + result = await client.upload_call( + action, payload, filename=filename, content=content, content_type=content_type + ) + else: + result = await client.call(action, payload) + return result if isinstance(result, dict) else {"result": result} + except CKANActionError as exc: + error = exc + if show_action and payload.get("id"): - existing = await client.action(show_action, {"id": payload["id"]}) + existing = await _read( + client, show_action, {"id": payload["id"]}, label=label, missing_ok=True + ) if existing: + if existing.get("state", "active") == "deleted": + raise DeletedInTarget(f"{label} was deleted in the Fair Store") + # Package, organization and group names are identifiers too; + # a resource name is only a label. + if ( + action != "resource_create" + and "name" in payload + and existing.get("name") != payload["name"] + ): + raise MirrorError( + f"{action} failed for {label}: UUID {payload['id']!r} is held by " + f"{existing.get('name')!r}: {error}" + ) return existing - if attempt < 2: - await asyncio.sleep(2**attempt) + if not error.transient or attempt == _ATTEMPTS - 1: + raise MirrorError(f"{action} failed for {label}: {error}") from error + await asyncio.sleep(2**attempt) + raise AssertionError("unreachable") + + +# -- qsv profiles as dataset metadata and resources --------------------------- +# +# onboard_ckan.py / onboard_dcat.py profile a dataset's first CSV with qsv: +# stats, frequency, and describegpt (AI descriptions, labels, tags). The mirror +# publishes that profile twice over: compact fields on the dataset, and the full +# outputs as resources on it. Tables go into the DataStore, so CKAN shows them +# as sortable tables and serves them as CSV/JSON downloads; describegpt's JSON +# is uploaded as a file. The profiled source data itself is never copied. + +_QSV_MODEL_RE = re.compile(r"^Model:\s*(\S+)", re.MULTILINE) +_QSV_VERSION_RE = re.compile(r"Generated by qsv v([\w.\-]+)") +# describegpt ends a description with a provenance block — "Generated by qsv +# vX describegpt", its command line, model, API URL, timestamp and an LLM +# warning — introduced however the model chose: "Attribution: Generated by…", +# an "## Attribution" heading, a bare line, a rule, or "@attribution". +_QSV_PROVENANCE_RE = re.compile( + r"^[^\n]*\bGenerated by qsv\b|^[ \t]*Command line:[ \t]*qsv describegpt\b", + re.IGNORECASE | re.MULTILINE, +) +_QSV_TRAILING_RE = re.compile( + r"(?:\n[ \t]*(?:(?:#{1,6}|\*\*||@)?[ \t]*attribution[ \t]*:?[ \t]*(?:\*\*|)?" + r"|-{3,}|\*{3,}|_{3,})?[ \t]*)+$", + re.IGNORECASE, +) +_DATASTORE_BATCH = 1000 +_MAX_SUMMARY_CELL = 10_000 + + +def _qsv_dir(qsv: dict[str, Any]) -> Path | None: + """The onboarding directory holding a profile's qsv output files.""" + local_path = str(qsv.get("local_path") or "") + if not local_path: + return None + path = Path(local_path) + directory = (path if path.is_absolute() else _PROJECT_ROOT / path).parent + return directory if directory.is_dir() else None + + +def _qsv_resource_id(profiled_resource_id: str, kind: str) -> str: + """Stable UUID for one qsv output of one profiled resource.""" + return str(uuid.uuid5(uuid.NAMESPACE_URL, f"qsv:{profiled_resource_id}:{kind}")) + + +def _profile_row_count(qsv: dict[str, Any]) -> int | None: + """The profiled file's row count, when it is one. + + A byte-capped DCAT download counts only the rows it kept, and a profile + whose download failed reports 0; neither is the file's size. + """ + count = qsv.get("row_count") + if isinstance(count, bool) or not isinstance(count, int) or qsv.get("truncated_download"): + return None + if count == 0 and str(qsv.get("status") or "") != "ok": + return None + return count + + +def _describegpt_prose(description: str) -> str: + """A describegpt description without its trailing provenance block.""" + description = description or "" + match = _QSV_PROVENANCE_RE.search(description) + prose = description[: match.start()] if match else description + return _QSV_TRAILING_RE.sub("", "\n" + prose).strip() + + +def _qsv_dataset_extras(qsv: dict[str, Any]) -> dict[str, str]: + """The profile's dataset-level summary, shown with the dataset's metadata. + + Column names are included so a search for a column finds its dataset. The + description keeps only describegpt's prose: the model and qsv version are + fields of their own, and the full output, provenance included, is the + describegpt resource. + """ + qsv = _vet_profile(qsv) + description = _scrub_secrets(str(qsv.get("qsv_description") or "")) + # The profiled file's own columns: the dictionary also lists the portal's + # DataStore fields, which a download can lack. + columns = qsv.get("_header") or [ + str(c["name"]) for c in qsv.get("columns") or [] if c.get("name") + ] + model = _QSV_MODEL_RE.search(description) + version = _QSV_VERSION_RE.search(description) + values = { + "qsv_description": _describegpt_prose(description), + "qsv_tags": json.dumps(qsv.get("qsv_tags") or []) if qsv.get("qsv_tags") else "", + "qsv_row_count": str( + _profile_row_count(qsv) if _profile_row_count(qsv) is not None else "" + ), + # Statistics from a byte-capped download describe the file's first rows. + "qsv_profile_truncated": "true" if qsv.get("truncated_download") else "", + "qsv_column_count": str(len(columns)) if columns else "", + # JSON, like qsv_tags: column names can contain commas. + "qsv_columns": json.dumps(columns, ensure_ascii=False) if columns else "", + "qsv_profiled_resource_id": str(qsv.get("resource_id") or ""), + "qsv_profiled_at": str(qsv.get("onboarded_at") or ""), + "qsv_status": str(qsv.get("status") or ""), + "qsv_model": model.group(1) if model else "", + "qsv_version": version.group(1) if version else "", + } + return {key: value for key, value in values.items() if value} + + +def _as_number(value: Any) -> float | int | None: + """A DataStore numeric cell: qsv writes blanks and text where numbers are absent.""" + try: + number = float(str(value).strip()) + except (TypeError, ValueError): + return None + return int(number) if number.is_integer() else number + + +def _top_values_text(top_values: list[dict[str, Any]] | None) -> str: + return "; ".join( + f"{item.get('value')} ({item.get('count')})" for item in (top_values or [])[:10] + ) + + +def _dictionary_table(qsv: dict[str, Any]) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]: + """One row per column: qsv type, AI label and description, key statistics.""" + fields = [ + { + "id": "column", + "type": "text", + "info": {"label": "Column", "notes": "Column name in the file"}, + }, + {"id": "label", "type": "text", "info": {"label": "Label", "notes": "AI-generated label"}}, + {"id": "type", "type": "text", "info": {"label": "Type", "notes": "Type inferred by qsv"}}, + { + "id": "description", + "type": "text", + "info": {"label": "Description", "notes": "AI-generated description"}, + }, + {"id": "null_count", "type": "numeric", "info": {"label": "Null count"}}, + {"id": "cardinality", "type": "numeric", "info": {"label": "Distinct values"}}, + {"id": "min", "type": "text", "info": {"label": "Minimum"}}, + {"id": "max", "type": "text", "info": {"label": "Maximum"}}, + {"id": "mean", "type": "numeric", "info": {"label": "Mean"}}, + {"id": "median", "type": "numeric", "info": {"label": "Median"}}, + {"id": "stddev", "type": "numeric", "info": {"label": "Standard deviation"}}, + {"id": "top_values", "type": "text", "info": {"label": "Most frequent values (count)"}}, + ] + records = [] + for column in _column_dictionary(qsv): + stats = column["stats"] + records.append( + { + "column": column["name"], + "label": column["label"], + "type": column["type"], + "description": column["description"], + "null_count": _as_number(stats.get("nullcount")), + "cardinality": _as_number(stats.get("cardinality")), + "min": str(stats.get("min") or ""), + "max": str(stats.get("max") or ""), + "mean": _as_number(stats.get("mean")), + "median": _as_number(stats.get("q2_median")), + "stddev": _as_number(stats.get("stddev")), + "top_values": _top_values_text(column["top_values"]), + } + ) + return fields, records + + +def _summary_cell(value: str) -> str: + """A text cell for a summary table, capped at :data:`_MAX_SUMMARY_CELL` chars. + + qsv reports a column's min/max/mode verbatim, so a geometry column yields + whole WKT polygons (one 311 mode is 157 KB) that would swamp the table view. + """ + value = _scrub_secrets(value) + if len(value) <= _MAX_SUMMARY_CELL: + return value + return f"{value[:_MAX_SUMMARY_CELL]}… [truncated: {len(value):,} characters]" + + +def _csv_table( + path: Path, numeric: frozenset[str] = frozenset() +) -> tuple[list[dict[str, Any]], list[dict[str, Any]]] | None: + """A qsv CSV as DataStore fields and records; ``numeric`` columns are typed.""" + if not path.is_file(): + return None + # Geometry cells are read whole, then capped. + _raise_csv_field_limit() + with path.open(newline="", encoding="utf-8", errors="replace") as handle: + reader = csv.DictReader(handle) + header = [name for name in reader.fieldnames or [] if name] + records = [ + { + name: ( + _as_number(row.get(name)) + if name in numeric + else _summary_cell(row.get(name) or "") + ) + for name in header + } + for row in reader + ] + fields = [{"id": name, "type": "numeric" if name in numeric else "text"} for name in header] + return fields, records + + +def _qsv_artifacts(qsv: dict[str, Any]) -> list[dict[str, Any]]: + """The resources to publish for one profile, in display order.""" + qsv = _vet_profile(qsv) + directory = _qsv_dir(qsv) + source = str(qsv.get("resource_name") or "the profiled file") + caveat = "Generated with qsv; AI-written text may contain inaccuracies." + artifacts: list[dict[str, Any]] = [] + + fields, records = _dictionary_table(qsv) + if records: + artifacts.append( + { + "kind": "dictionary", + "name": "Data dictionary (qsv)", + "description": ( + f"Column-level data dictionary for “{source}”: qsv-inferred types, " + f"AI-generated labels and descriptions (qsv describegpt), and key " + f"statistics. {caveat}" + ), + "table": (fields, records), + } + ) + if directory is None: + return artifacts + + stats = _csv_table(directory / "qsv_stats.csv") if qsv["_stats_ok"] else None + if stats: + artifacts.append( + { + "kind": "stats", + "name": "Summary statistics (qsv stats)", + "description": ( + f"qsv stats for “{source}”: one row per column with its type, " + "min/max, mean, quartiles, cardinality, null count and more." + ), + "table": stats, + } + ) + frequency = ( + _csv_table( + directory / "qsv_frequency.csv", numeric=frozenset({"count", "percentage", "rank"}) + ) + if qsv["_frequency_ok"] + else None + ) + if frequency: + artifacts.append( + { + "kind": "frequency", + "name": "Frequency distribution (qsv frequency)", + "description": ( + f"qsv frequency for “{source}”: the most common values in each " + "column, with counts, percentages and rank." + ), + "table": frequency, + } + ) + describegpt = directory / "qsv_dict.json" + if qsv["_describegpt_ok"] and describegpt.is_file(): + artifacts.append( + { + "kind": "describegpt", + "name": "AI descriptions (qsv describegpt)", + "description": ( + f"Raw qsv describegpt output for “{source}”: the AI-generated " + f"dataset description, tags, and column dictionary. {caveat}" + ), + "file": ( + "describegpt.json", + _scrub_secrets(describegpt.read_text(encoding="utf-8")).encode("utf-8"), + "application/json", + ), + } + ) + return artifacts + + +_QSV_KINDS = ("dictionary", "stats", "frequency", "describegpt") + + +async def _publish_qsv( + target: CKANClient, + *, + package_id: str, + qsv: dict[str, Any], + site_id: str, + existing_resources: dict[str, dict[str, Any]] | set[str], + apply: bool = True, +) -> dict[str, int]: + """Create, refresh or remove one profile's resources; counts by outcome. + + Each resource records a digest of what it was built from (``qsv_digest``), + and one whose digest matches — with its table still loaded — is left alone, + so a rerun rewrites only the profiles that changed. A kind the profile no + longer produces is removed: these resources are the Fair Store's own + derivation, and one left behind would be stale or, for output that turned + out to describe another file, wrong. ``apply=False`` only counts. + """ + held = ( + existing_resources + if isinstance(existing_resources, dict) + else {resource_id: {} for resource_id in existing_resources} + ) + counts = {"created": 0, "refreshed": 0, "unchanged": 0, "removed": 0} + profiled_id = str(qsv.get("resource_id") or "") + produced: set[str] = set() + for artifact in _qsv_artifacts(qsv): + produced.add(artifact["kind"]) + resource_id = _qsv_resource_id(profiled_id, artifact["kind"]) + resource: dict[str, Any] = { + "id": resource_id, + "package_id": package_id, + "name": artifact["name"], + "description": artifact["description"], + "qsv_artifact": artifact["kind"], + "qsv_profiled_resource_id": profiled_id, + "qsv_profiled_at": str(qsv.get("onboarded_at") or ""), + "qsv_profile_site": site_id, + } + if "file" in artifact: + resource["format"] = "JSON" + content = hashlib.sha256(artifact["file"][1]).hexdigest() + resource["qsv_digest"] = _digest({"resource": resource, "content": content}) + else: + fields, records = artifact["table"] + # What datastore_create does for an inline resource, done as a + # create of its own so a lost response is resolved by its UUID. + resource.update(format="CSV", url="_datastore_only_resource", url_type="datastore") + resource["qsv_digest"] = _digest( + {"resource": resource, "fields": fields, "records": records} + ) + existing = held.get(resource_id) + loaded = existing is not None and ( + existing.get("url_type") == "upload" + if "file" in artifact + else bool(existing.get("datastore_active")) + ) + if loaded and existing.get("qsv_digest") == resource["qsv_digest"]: + counts["unchanged"] += 1 + if apply and "table" in artifact: + await _ensure_table_view(target, resource_id, label=resource["name"]) + continue + counts["created" if existing is None else "refreshed"] += 1 + if not apply: + continue + + label = f"{artifact['kind']} for {profiled_id}" + action = "resource_create" if existing is None else "resource_patch" + if "file" in artifact: + # The file and its digest land in one call. + await _write_action(target, action, resource, label=label, upload=artifact["file"]) + continue + # The digest is recorded only once the rows are in and counted: a run + # that stops mid-load leaves a table that the next run must reload, + # not one it would take for unchanged. + digest = resource["qsv_digest"] + await _write_action(target, action, {**resource, "qsv_digest": ""}, label=label) + await _load_table(target, resource_id, fields, records, label=label) + await _write_action( + target, "resource_patch", {"id": resource_id, "qsv_digest": digest}, label=label + ) + await _ensure_table_view(target, resource_id, label=label) + + # Without the onboarding directory nothing shows the output is gone (the + # mirror may just be running from a checkout that lacks it), so nothing + # is removed. + removable = _qsv_dir(qsv) is not None + for kind in _QSV_KINDS: + resource_id = _qsv_resource_id(profiled_id, kind) + if not removable or kind in produced or resource_id not in held: + continue + counts["removed"] += 1 + if apply: + await _delete( + target, "resource_delete", {"id": resource_id}, label=f"{kind} for {profiled_id}" + ) + return counts + + +async def _delete(target: CKANClient, action: str, params: dict[str, Any], *, label: str) -> None: + """A delete; an object that is already gone is fine.""" + for attempt in range(_ATTEMPTS): + try: + await target.call(action, params) + return + except CKANActionError as exc: + if exc.not_found: + return + if not exc.transient or attempt == _ATTEMPTS - 1: + raise MirrorError(f"{action} failed for {label}: {exc}") from exc + await asyncio.sleep(2**attempt) + + +async def _drop_table(target: CKANClient, resource_id: str, *, label: str) -> None: + """Delete a resource's DataStore table; one that does not exist is fine.""" + await _delete( + target, "datastore_delete", {"resource_id": resource_id, "force": True}, label=label + ) + + +async def _load_table( + target: CKANClient, + resource_id: str, + fields: list[dict[str, Any]], + records: list[dict[str, Any]], + *, + label: str, +) -> None: + """Replace a resource's DataStore table with ``records``, verified by count. + + The table is rebuilt rather than appended to, so a rerun cannot duplicate + rows. An insert whose response is lost may still have committed, and its + retry then adds the batch twice, so the loaded row count is checked and the + table rebuilt until it matches. + """ + total: Any = None + for _ in range(_ATTEMPTS): + await _drop_table(target, resource_id, label=label) + await _write_action( + target, + "datastore_create", + { + "resource_id": resource_id, + "force": True, + "fields": fields, + "records": records[:_DATASTORE_BATCH], + }, + label=label, + ) + for start in range(_DATASTORE_BATCH, len(records), _DATASTORE_BATCH): + await _write_action( + target, + "datastore_upsert", + { + "resource_id": resource_id, + "force": True, + "method": "insert", + "records": records[start : start + _DATASTORE_BATCH], + }, + label=label, + ) + loaded = await _read( + target, "datastore_search", {"resource_id": resource_id, "limit": 0}, label=label + ) + total = loaded.get("total") if isinstance(loaded, dict) else None + if total == len(records): + return + print( + f"warning: {label} holds {total} rows, expected {len(records)}; reloading", + file=sys.stderr, + ) + raise MirrorError(f"{label}: DataStore holds {total} rows, expected {len(records)}") + + +async def _ensure_table_view(target: CKANClient, resource_id: str, *, label: str) -> None: + """Give a DataStore resource its table view. + + CKAN picks a resource's default views when the resource is created, before + its table exists, so the table view never qualifies and has to be added + once the rows are in. An existing one is left alone so reruns do not stack + duplicates — which is why a failed listing or create is re-listed, never + read as "no view yet". + """ + for attempt in range(_ATTEMPTS): + views = await _read( + target, "resource_view_list", {"id": resource_id}, label=f"views of {label}" + ) + if any( + isinstance(view, dict) and view.get("view_type") == "datatables_view" + for view in views or [] + ): + return + try: + await target.call( + "resource_view_create", + {"resource_id": resource_id, "view_type": "datatables_view", "title": "Table"}, + ) + return + except CKANActionError as exc: + if not exc.transient or attempt == _ATTEMPTS - 1: + raise MirrorError(f"table view failed for {label}: {exc}") from exc + await asyncio.sleep(2**attempt) + + +# The digest of what the mirror last wrote to a record: an unchanged record is +# skipped on the next run instead of being patched (and reindexed) again. +_DIGEST_KEY = "mirror_digest" + + +def _digest(value: Any) -> str: + encoded = json.dumps(value, sort_keys=True, ensure_ascii=False, default=str) + return hashlib.sha256(encoded.encode("utf-8")).hexdigest()[:20] + + +def _extra(row: dict[str, Any], key: str) -> str: + for item in row.get("extras") or []: + if isinstance(item, dict) and item.get("key") == key: + return str(item.get("value") or "") + return "" + + +def _stamp_package(payload: dict[str, Any]) -> str: + """Add the package's digest as an extra; extras are compared as a set.""" + extras = sorted(payload.get("extras") or [], key=lambda item: item["key"]) + digest = _digest({**payload, "extras": extras}) + payload["extras"] = [*payload.get("extras", []), {"key": _DIGEST_KEY, "value": digest}] + return digest - raise MirrorError(f"{action} failed for {label}") + +def _stamp_resource(payload: dict[str, Any]) -> str: + digest = _digest(payload) + payload[_DIGEST_KEY] = digest + return digest + + +def _package_payload( + package: dict[str, Any], + *, + site_id: str, + source_url: str, + store_id: str, + known_org_ids: set[str], + qsv: dict[str, Any] | None = None, +) -> dict[str, Any]: + payload = _project(package, _PACKAGE_FIELDS) + # A patch leaves out what it does not name, so a value the source cleared + # would survive in the Fair Store; name these as empty instead. + for field in _CLEARABLE_PACKAGE_FIELDS: + payload.setdefault(field, "") + owner_org = str(package.get("owner_org") or "") + payload["owner_org"] = owner_org if owner_org in known_org_ids else store_id + payload["tags"] = _tag_refs(package.get("tags")) + payload["groups"] = _group_refs(package.get("groups")) + carried = { + key: _extra_value(value) + for key, value in package.items() + if key not in _PACKAGE_MANAGED and not key.startswith("_") and not _is_empty(value) + } + if qsv: + carried.update(_qsv_dataset_extras(qsv)) + for key in ("metadata_created", "metadata_modified"): + if package.get(key): + carried[f"mirror_source_{key}"] = str(package[key]) + if not owner_org: + carried["mirror_source_owner_org"] = "" + extras = _provenance_extras( + package.get("extras"), + site_id=site_id, + source_url=source_url, + source_id=_source_id(package), + carried=carried, + ) + # CKAN also indexes each dataset extra as a single untokenized Solr term; + # one longer than Solr's limit fails the whole dataset with an HTTP 500. + # Such a value cannot be stored on the dataset, so it is named instead. + oversized = [ + item["key"] for item in extras if len(item["value"].encode("utf-8")) > _SOLR_MAX_TERM_BYTES + ] + if oversized: + print( + f"warning: {package.get('name')}: extras over Solr's term limit not stored: " + f"{', '.join(oversized)}", + file=sys.stderr, + ) + extras = [item for item in extras if item["key"] not in oversized] + extras.append({"key": "mirror_oversized_extras", "value": json.dumps(oversized)}) + payload["extras"] = extras + if not str(payload.get("notes") or "").strip(): + payload["notes"] = "No description was provided by the source catalog." + return payload def _resource_payload( @@ -391,38 +2088,83 @@ def _resource_payload( source_url: str, qsv: dict[str, Any] | None, ) -> dict[str, Any]: - payload = _project(resource, _RESOURCE_FIELDS) + """Every source resource field, minus the source CKAN's own bookkeeping.""" + payload = { + key: value + for key, value in resource.items() + if value is not None + and key not in _RESOURCE_MANAGED + and not key.startswith(("_", "datastore_")) + } payload["package_id"] = package_id + url_type = str(resource.get("url_type") or "") + if url_type in _LOCAL_URL_TYPES: + payload.pop("url_type", None) + payload["mirror_source_url_type"] = url_type + if resource.get("datastore_active"): + # Recorded so a consumer knows the source portal can be queried; the + # Fair Store itself has no DataStore table for this resource. + payload["mirror_source_datastore_active"] = True payload.update( { _PROVENANCE_KEYS["portal"]: site_id, _PROVENANCE_KEYS["url"]: source_url, - _PROVENANCE_KEYS["id"]: str(resource.get("id", "")), + _PROVENANCE_KEYS["id"]: _source_id(resource), } ) - # Preserve extension-owned spatial metadata when the destination has the - # same extension, but never claim that a DataStore table was copied. - for key, value in resource.items(): - if key.startswith("dataspatial_") and value is not None: - payload[key] = value if qsv: + qsv = _vet_profile(qsv) columns = _column_dictionary(qsv) payload.update( { - "qsv_description": _scrub_secrets(qsv.get("qsv_description", "")), + "qsv_description": _describegpt_prose( + _scrub_secrets(str(qsv.get("qsv_description") or "")) + ), "qsv_tags": json.dumps(qsv.get("qsv_tags") or []), "data_dictionary": json.dumps(columns), - "column_count": len(columns), - "row_count": qsv.get("row_count", 0), + "column_count": len(qsv.get("_header") or columns), + "row_count": _profile_row_count(qsv), "qsv_onboarded_at": qsv.get("onboarded_at", ""), } ) return payload +def _withdrawn( + source: dict[str, Any], target: dict[str, list[dict[str, Any]]], *, site_id: str +) -> dict[str, Any]: + """This portal's mirrored records that its source no longer publishes. + + Reported, never deleted: whether a record withdrawn upstream should leave + the Fair Store is the operator's call. qsv resources are the Fair Store's + own and are not counted. + """ + source_packages = {str(package.get("id")) for package in source["packages"]} + source_resources = { + str(resource.get("id")) + for package in source["packages"] + for resource in package.get("resources") or [] + } + ours = [package for package in target["packages"] if _mirrored_from(package) == site_id] + datasets = sorted( + str(package.get("name")) + for package in ours + if str(package.get("id")) not in source_packages + ) + resources = [ + str(resource.get("id")) + for package in ours + if str(package.get("id")) in source_packages + for resource in package.get("resources") or [] + if resource.get(_PROVENANCE_KEYS["portal"]) == site_id + and str(resource.get("id")) not in source_resources + ] + return {"datasets": len(datasets), "dataset_names": datasets[:50], "resources": len(resources)} + + async def mirror_catalog( *, - source: CKANClient, + source: CKANClient | None, target: CKANClient, site_id: str, site_title: str, @@ -431,15 +2173,40 @@ async def mirror_catalog( organization: str | None = None, qsv_index: dict[str, Any] | None = None, limit: int | None = None, + snapshot: dict[str, Any] | None = None, + pin_names: bool = False, ) -> dict[str, Any]: - """Plan or apply one complete, API-level CKAN metadata mirror.""" - source_data = await _source_snapshot(source, organization=organization, limit=limit) + """Plan or apply one complete, API-level metadata mirror. + + ``source`` is read as a CKAN portal unless a prepared ``snapshot`` (from + :func:`_dcat_snapshot`) is passed instead. ``pin_names`` keeps the names + already published for the portal's records (see :func:`_pin_target_names`); + without it a source rename is applied. + """ + if snapshot is not None: + source_data = snapshot + elif source is not None: + source_data = await _source_snapshot( + source, site_id=site_id, organization=organization, limit=limit + ) + else: + raise MirrorError("mirror_catalog needs a source client or a snapshot") target_data = await _target_snapshot(target) + if pin_names: + _pin_target_names(source_data, target_data, site_id=site_id) store_name = _slug(site_id, "source-store") if store_name in {str(row.get("name")) for row in source_data["organizations"]}: store_name = _slug(f"{store_name}-source", "source-store") store_id = str(uuid.uuid5(uuid.NAMESPACE_URL, source_url.rstrip("/"))) - _assert_no_collisions(source_data, target_data, store_name=store_name, store_id=store_id) + evictions: list[tuple[str, str, str]] = [] + renames = _assert_no_collisions( + source_data, + target_data, + store_name=store_name, + store_id=store_id, + site_id=site_id, + evictions=evictions, + ) qsv_resources = _qsv_by_resource(qsv_index or {}) resources = [ @@ -459,22 +2226,86 @@ async def mirror_catalog( "qsv_resources": sum( 1 for resource in resources if str(resource.get("id")) in qsv_resources ), + "skipped_copies": source_data.get("copies", _empty_copies()), + "renamed": renames, + # Planned in a dry run, done when applied. "created": 0, "updated": 0, + "unchanged": 0, + "qsv": {"created": 0, "refreshed": 0, "unchanged": 0, "removed": 0}, } - if not apply: - return summary + if organization is None and limit is None: + summary["withdrawn_at_source"] = _withdrawn(source_data, target_data, site_id=site_id) + + held_qsv = [ + resource + for package in target_data["packages"] + for resource in package.get("resources", []) or [] + if resource.get("qsv_profile_site") == site_id + ] + if held_qsv and not qsv_resources: + # Every profiled dataset would lose its qsv fields to the patch. + message = ( + f"the Fair Store holds {len(held_qsv)} qsv resources for {site_id}, but no qsv " + "profiles were loaded: is its onboarding index missing (see --index-path)?" + ) + if apply: + raise MirrorError(message) + print(f"warning: {message}", file=sys.stderr) + # Profiles no longer in the index; their resources are reported, not removed. + summary["qsv"]["orphaned_profiles"] = len( + {str(resource.get("qsv_profiled_resource_id")) for resource in held_qsv} + - set(qsv_resources) + ) - target_org_ids = {str(row.get("id")) for row in target_data["organizations"]} target_group_ids = {str(row.get("id")) for row in target_data["groups"]} target_group_names = {str(row.get("name")) for row in target_data["groups"]} - target_package_ids = {str(row.get("id")) for row in target_data["packages"]} - target_resource_ids = { - str(resource.get("id")) + target_packages = {str(row.get("id")): row for row in target_data["packages"]} + target_resources = { + str(resource.get("id")): resource for package in target_data["packages"] for resource in package.get("resources", []) or [] } + async def write(action: str, payload: dict[str, Any], *, label: str) -> None: + if apply: + await _write_action(target, action, payload, label=label) + + target_orgs = {str(row.get("id")): row for row in target_data["organizations"]} + deleted_orgs: set[str] = set() + + async def upsert_group( + kind: str, payload: dict[str, Any], held: dict[str, Any] | None, digest: str + ) -> bool: + """Create or patch an organization or group; False when unchanged or skipped.""" + if held is not None and _extra(held, _DIGEST_KEY) == digest: + summary["unchanged"] += 1 + return False + payload["extras"] = [*payload["extras"], {"key": _DIGEST_KEY, "value": digest}] + try: + await write( + f"{kind}_{'create' if held is None else 'patch'}", payload, label=payload["name"] + ) + except DeletedInTarget: + # An admin removed it from the Fair Store; that decision stands. + summary.setdefault("skipped_deleted_in_target", []).append(payload["name"]) + if kind == "organization": + deleted_orgs.add(str(payload["id"])) + return False + summary["created" if held is None else "updated"] += 1 + return True + + def group_digest(payload: dict[str, Any], parents: Any = None) -> str: + extras = sorted(payload.get("extras") or [], key=lambda item: item["key"]) + return _digest({**payload, "extras": extras, "parents": parents}) + + for kind, holder_id, evicted in evictions: + await write( + "organization_patch" if kind == "organization" else "package_patch", + {"id": holder_id, "name": evicted}, + label=f"{kind} {holder_id} (withdrawn at source)", + ) + store_payload = { "id": store_id, "name": store_name, @@ -484,126 +2315,252 @@ async def mirror_catalog( [], site_id=site_id, source_url=source_url, source_id=store_id ), } - store_exists = store_id in target_org_ids - await _write_action( - target, - "organization_patch" if store_exists else "organization_create", - store_payload, - label=store_name, + await upsert_group( + "organization", store_payload, target_orgs.get(store_id), group_digest(store_payload) ) - summary["updated" if store_exists else "created"] += 1 + # CKAN's no_loops validator resolves both child and parent rows; combining a + # caller-supplied child UUID with groups during create trips that validator, + # so the hierarchy is applied only after every source organization exists. + # A parent that is not being mirrored cannot be referenced, so an + # organization whose parents are all outside the snapshot hangs off the root + # — except in an --organization run, which leaves an existing one's parent. source_org_names = {str(row.get("name")) for row in source_data["organizations"]} + hierarchy: list[tuple[dict[str, Any], list[dict[str, str]]]] = [] for organization_row in source_data["organizations"]: - payload = _project(organization_row, _ORGANIZATION_FIELDS) - source_id = str(organization_row.get("id", "")) + org_id = str(organization_row.get("id", "")) + payload = _with_image(_project(organization_row, _ORGANIZATION_FIELDS), organization_row) payload["extras"] = _provenance_extras( organization_row.get("extras"), site_id=site_id, source_url=source_url, - source_id=source_id, - ) - exists = source_id in target_org_ids - await _write_action( - target, - "organization_patch" if exists else "organization_create", - payload, - label=str(organization_row.get("name")), + source_id=_source_id(organization_row), ) - summary["updated" if exists else "created"] += 1 - - # Apply hierarchy only after every source organization exists. CKAN's - # no_loops validator resolves both child and parent rows; combining a - # caller-supplied child UUID with groups during create trips that validator. - for organization_row in source_data["organizations"]: - parents = _group_refs(organization_row.get("groups")) - if not any(parent["name"] in source_org_names for parent in parents): - parents.append({"name": store_name, "capacity": "parent"}) - await _write_action( - target, + held = target_orgs.get(org_id) + parents: list[dict[str, str]] | None = [ + parent + for parent in _group_refs(organization_row.get("groups")) + if parent["name"] in source_org_names + ] or [{"name": store_name, "capacity": "parent"}] + if organization is not None and held is not None: + parents = None + if await upsert_group("organization", payload, held, group_digest(payload, parents)): + if parents is not None: + hierarchy.append((organization_row, parents)) + + deleted_names = { + str(row.get("name")) + for row in source_data["organizations"] + if str(row.get("id")) in deleted_orgs + } + for organization_row, parents in hierarchy: + parents = [parent for parent in parents if parent["name"] not in deleted_names] or [ + {"name": store_name, "capacity": "parent"} + ] + await write( "organization_patch", {"id": str(organization_row.get("id", "")), "groups": parents}, label=f"{organization_row.get('name')} hierarchy", ) + target_groups = {str(row.get("id")): row for row in target_data["groups"]} for group in source_data["groups"]: - payload = _project(group, _GROUP_FIELDS) - source_id = str(group.get("id", "")) - if str(group.get("name")) in target_group_names and source_id not in target_group_ids: + payload = _with_image(_project(group, _GROUP_FIELDS), group) + group_id = str(group.get("id", "")) + if str(group.get("name")) in target_group_names and group_id not in target_group_ids: continue payload["extras"] = _provenance_extras( - group.get("extras"), site_id=site_id, source_url=source_url, source_id=source_id - ) - exists = source_id in target_group_ids - await _write_action( - target, - "group_patch" if exists else "group_create", - payload, - label=str(group.get("name")), + group.get("extras"), site_id=site_id, source_url=source_url, source_id=_source_id(group) ) - summary["updated" if exists else "created"] += 1 - - known_org_ids = {str(row.get("id")) for row in source_data["organizations"]} - for package in source_data["packages"]: - payload = _project(package, _PACKAGE_FIELDS) - package_id = str(package.get("id", "")) - owner_org = str(package.get("owner_org") or "") - payload["owner_org"] = owner_org if owner_org in known_org_ids else store_id - payload["tags"] = [ - { - "name": str(tag["name"]), - **({"vocabulary_id": tag["vocabulary_id"]} if tag.get("vocabulary_id") else {}), - } - for tag in package.get("tags", []) or [] - if tag.get("name") + await upsert_group("group", payload, target_groups.get(group_id), group_digest(payload)) + + # Datasets of an organization an admin deleted here hang off the root. + known_org_ids = {str(row.get("id")) for row in source_data["organizations"]} - deleted_orgs + total = len(source_data["packages"]) + for position, package in enumerate(source_data["packages"], start=1): + profiles = [ + qsv_resources[str(resource.get("id"))] + for resource in package.get("resources", []) or [] + if str(resource.get("id")) in qsv_resources ] - payload["groups"] = _group_refs(package.get("groups")) - payload["extras"] = _provenance_extras( - package.get("extras"), + payload = _package_payload( + package, site_id=site_id, source_url=source_url, - source_id=package_id, - ) - if not str(payload.get("notes") or "").strip(): - payload["notes"] = "No description was provided by the source catalog." - if not owner_org: - payload["extras"].append({"key": "mirror_source_owner_org", "value": ""}) - exists = package_id in target_package_ids - await _write_action( - target, - "package_patch" if exists else "package_create", - payload, - label=str(package.get("name")), + store_id=store_id, + known_org_ids=known_org_ids, + qsv=profiles[0] if profiles else None, ) - summary["updated" if exists else "created"] += 1 - - for resource in package.get("resources", []) or []: - resource_id = str(resource.get("id", "")) - resource_payload = _resource_payload( + digest = _stamp_package(payload) + package_id = str(payload.get("id", "")) + resource_payloads = [ + _resource_payload( resource, package_id=package_id, site_id=site_id, source_url=source_url, - qsv=qsv_resources.get(resource_id), + qsv=qsv_resources.get(str(resource.get("id", ""))), ) - resource_exists = resource_id in target_resource_ids - await _write_action( + for resource in package.get("resources", []) or [] + ] + for resource_payload in resource_payloads: + _stamp_resource(resource_payload) + if apply and (position % 50 == 0 or position == total): + print(f"[{site_id}] {position}/{total} datasets", file=sys.stderr, flush=True) + + existing = target_packages.get(package_id) + if existing is None: + # One call creates the dataset and its resources with their + # preserved UUIDs; a separate create per resource costs a + # reindex each and made a portal-sized mirror several times slower. + try: + await write( + "package_create", + {**payload, "resources": resource_payloads}, + label=str(package.get("name")), + ) + except DeletedInTarget: + # An admin removed it from the Fair Store; that decision stands. + summary.setdefault("skipped_deleted_in_target", []).append(package.get("name")) + continue + summary["created"] += 1 + len(resource_payloads) + else: + if _extra(existing, _DIGEST_KEY) == digest: + summary["unchanged"] += 1 + else: + await write("package_patch", payload, label=str(package.get("name"))) + summary["updated"] += 1 + for resource_payload in resource_payloads: + held = target_resources.get(resource_payload["id"]) + if held is not None and held.get(_DIGEST_KEY) == resource_payload[_DIGEST_KEY]: + summary["unchanged"] += 1 + continue + # Replaced rather than patched: a field the source dropped + # must go here too (a patch would keep it). + await write( + "resource_create" if held is None else "resource_update", + resource_payload, + label=str(resource_payload["id"]), + ) + summary["created" if held is None else "updated"] += 1 + + for profile in profiles: + counts = await _publish_qsv( target, - "resource_patch" if resource_exists else "resource_create", - resource_payload, - label=resource_id, + package_id=package_id, + qsv=profile, + site_id=site_id, + existing_resources=target_resources, + apply=apply, ) - summary["updated" if resource_exists else "created"] += 1 + for outcome, count in counts.items(): + summary["qsv"][outcome] += count + qsv_counts = summary["qsv"] + summary["qsv_artifacts"] = ( + qsv_counts["created"] + qsv_counts["refreshed"] + qsv_counts["unchanged"] + ) return summary -async def _main(args: argparse.Namespace) -> dict[str, Any]: - site = get_site(args.site) - if not site: - raise MirrorError(f"Unknown site {args.site!r}; add it to ckan_sites.json first") - if site.get("portal_type", "ckan") != "ckan": - raise MirrorError("Exact organization mirroring currently requires a CKAN source portal") +def _site_index(site_id: str, args: argparse.Namespace) -> dict[str, Any]: + stage_root = getattr(args, "stage_root", None) + return _load_index( + site_id, + args.index_path, + source=getattr(args, "qsv_source", "local"), + stage_root=Path(stage_root) / site_id if stage_root else None, + ) + + +async def _mirror_site( + site_id: str, site: dict[str, Any], args: argparse.Namespace, api_key: str +) -> dict[str, Any]: + # Neither client may fall back to CKAN_API_KEY: the target is written with + # the mirror's own token, and a source portal must never receive a key. + target = CKANClient(ckan_url=args.target_url, api_key=api_key or None, use_default_key=False) + try: + common = { + "target": target, + "site_id": site_id, + "site_title": site["name"], + "source_url": site["url"], + "apply": args.apply, + } + if normalize_portal_type(site.get("portal_type")) == PORTAL_TYPE_DCAT: + if args.organization: + raise MirrorError("--organization applies to CKAN source portals only") + dcat = DCATClient(site["url"], site.get("catalog_url")) + try: + catalog = await dcat.fetch_raw_catalog() + except (httpx.HTTPError, ValueError) as exc: + raise MirrorError(f"DCAT catalog for {site_id} could not be read: {exc}") from exc + finally: + await dcat.close() + socrata = ( + {} if args.no_enrich else await _socrata_metadata(catalog, required=args.apply) + ) + snapshot = _dcat_snapshot( + catalog, + site_title=site["name"], + source_url=site["url"], + socrata=socrata, + limit=args.limit, + ) + summary = await mirror_catalog( + source=None, + snapshot=snapshot, + pin_names=True, + qsv_index=_dcat_qsv_index(_site_index(site_id, args), snapshot), + **common, + ) + summary["socrata_enriched"] = sum( + 1 + for package in snapshot["packages"] + if any(extra["key"] == "socrata_id" for extra in package["extras"]) + ) + return summary + + source = CKANClient(ckan_url=site["url"], use_default_key=False) + try: + return await mirror_catalog( + source=source, + organization=args.organization, + qsv_index=_site_index(site_id, args), + limit=args.limit, + **common, + ) + finally: + await source.close() + finally: + await target.close() + + +def _same_portal(url: str, other: str) -> bool: + def key(value: str) -> tuple[str, str]: + parsed = urlparse(value.strip()) + host = (parsed.hostname or "").lower().removeprefix("www.") + return host, parsed.path.rstrip("/") + + return key(url) == key(other) + + +async def _main(args: argparse.Namespace) -> dict[str, Any] | list[dict[str, Any]]: + requested = [part.strip() for part in args.site.split(",") if part.strip()] + many = args.site == "all" or len(requested) > 1 + if many and (args.organization or args.index_path): + raise MirrorError("--organization and --index-path need a single --site") + if args.site == "all": + sites = [(str(site["id"]), site) for site in list_sites()] + else: + sites = [] + for site_id in dict.fromkeys(requested): + site = get_site(site_id) + if not site: + raise MirrorError(f"Unknown site {site_id!r}; add it to ckan_sites.json first") + sites.append((site_id, site)) + if not sites: + raise MirrorError("--site needs a portal ID, a comma-separated list, or 'all'") api_key = os.environ.get(args.api_key_env, "") if args.api_key_file: @@ -613,41 +2570,86 @@ async def _main(args: argparse.Namespace) -> dict[str, Any]: f"Set {args.api_key_env} or --api-key-file when using --apply; a sysadmin token is required" ) - source = CKANClient(ckan_url=site["url"]) - target = CKANClient(ckan_url=args.target_url, api_key=api_key or None) - try: - return await mirror_catalog( - source=source, - target=target, - site_id=args.site, - site_title=site["name"], - source_url=site["url"], - apply=args.apply, - organization=args.organization, - qsv_index=_load_index(args.site, args.index_path), - limit=args.limit, - ) - finally: - await source.close() - await target.close() + if not many and _same_portal(sites[0][1]["url"], args.target_url): + raise MirrorError(f"{args.site} is the target portal; it cannot be mirrored into itself") + + summaries: list[dict[str, Any]] = [] + for site_id, site in sites: + if _same_portal(site["url"], args.target_url): + # data.dathere.com is both a registered portal and a Fair Store. + summaries.append({"site": site_id, "skipped": "this portal is the target"}) + continue + try: + summaries.append(await _mirror_site(site_id, site, args, api_key)) + except Exception as exc: + if not many: + raise + # One portal's outage must not stop the others from refreshing; + # the failure is reported and the exit status is non-zero. + print(f"error: {site_id}: {type(exc).__name__}: {exc}", file=sys.stderr) + summaries.append({"site": site_id, "error": f"{type(exc).__name__}: {exc}"}) + return summaries if many else summaries[0] + + +def _write_summary(path: str | None, summary: Any) -> None: + """Write the run's summary where a supervisor (the admin panel) reads it.""" + if not path: + return + Path(path).write_text(json.dumps(summary, indent=2, sort_keys=True), encoding="utf-8") def main() -> None: - parser = argparse.ArgumentParser(description="Mirror CKAN metadata and qsv results") - parser.add_argument("--site", default="wprdc", help="source site ID from ckan_sites.json") - parser.add_argument("--target-url", default=os.environ.get("CKAN_URL", "http://localhost:5001")) + parser = argparse.ArgumentParser(description="Mirror CKAN/DCAT metadata and qsv results") + parser.add_argument( + "--site", + default="wprdc", + help=( + "source site ID from ckan_sites.json, a comma-separated list of them, " + "or 'all' for every registered portal" + ), + ) + # Deliberately not CKAN_URL/CKAN_API_KEY: those are the app's portal + # settings, and in the app container CKAN_URL is data.dathere.com. + parser.add_argument( + "--target-url", default=os.environ.get("FAIRSTORE_URL", "http://localhost:5001") + ) parser.add_argument("--organization", help="optionally mirror only one source organization") parser.add_argument("--index-path", help="override the qsv index.json path") parser.add_argument("--limit", type=int, help="limit datasets for a smoke test") parser.add_argument("--apply", action="store_true", help="write after collision preflight") - parser.add_argument("--api-key-env", default="CKAN_API_KEY") + parser.add_argument( + "--no-enrich", + action="store_true", + help="skip Socrata Discovery API enrichment (owning agency, column list) for DCAT portals", + ) + parser.add_argument("--api-key-env", default="FAIRSTORE_API_KEY") parser.add_argument("--api-key-file", help="read the target sysadmin token from a file") + parser.add_argument( + "--qsv-source", + choices=("auto", "local", "storage"), + default="auto", + help=( + "where qsv profiles come from: this checkout's onboarding output (local), " + "the storage backend onboarding syncs to (storage), or local else storage (auto)" + ), + ) + parser.add_argument("--summary-file", help="also write the JSON summary to this file") args = parser.parse_args() + # Development logging is DEBUG on stdout, where this command prints its + # JSON summary; a line per HTTP request would bury it. + for noisy in ("asyncio", "httpx", "httpcore"): + logging.getLogger(noisy).setLevel(logging.WARNING) try: - summary = asyncio.run(_main(args)) + with tempfile.TemporaryDirectory(prefix="fairstore-qsv-") as stage_root: + args.stage_root = stage_root + summary = asyncio.run(_main(args)) except MirrorError as exc: + _write_summary(args.summary_file, {"error": str(exc)}) parser.exit(2, f"error: {exc}\n") + _write_summary(args.summary_file, summary) print(json.dumps(summary, indent=2, sort_keys=True)) + if isinstance(summary, list) and any("error" in item for item in summary): + sys.exit(1) if __name__ == "__main__": diff --git a/src/data_concierge/agents/llm_agent.py b/src/data_concierge/agents/llm_agent.py index dd23882..0281deb 100644 --- a/src/data_concierge/agents/llm_agent.py +++ b/src/data_concierge/agents/llm_agent.py @@ -153,6 +153,16 @@ def _source_for_tool(tool_name: str, data_source: str) -> str: return f"ckan:{data_source}" +def _registered_config(site_id: str) -> dict[str, Any] | None: + """A registered portal's config, or ``None``.""" + try: + from data_concierge.gateway import ckan_sites + + return ckan_sites.get_portal_config(site_id) + except Exception: # pragma: no cover - registry failures are non-fatal + return None + + def _sources_from_trace( tool_calls_trace: list[dict[str, Any]], data_source: str, @@ -161,7 +171,9 @@ def _sources_from_trace( ) -> list[DataSource]: """The data sources a run actually touched, in first-use order. - A CKAN tool call attributes the primary portal; an ``mcp__{server}__*`` + A CKAN tool call attributes the portal that served it (``portal_id`` on the + trace entry: an override, or a Fair Store load routed to its origin), else + the primary portal; an ``mcp__{server}__*`` call attributes that server (named from ``_STATIC_PORTAL_CONFIGS`` when registered there). Falls back to the nominal portal when no tool ran, so citations are never empty. @@ -188,11 +200,15 @@ def add(source_id: str, name: str, url: str, quality: float) -> None: else: add(server_id, f"MCP server: {server_id}", "", 0.85) else: + served = str(entry.get("portal_id") or data_source) + cfg = portal_cfg if served == data_source else _registered_config(served) + if cfg is None: + served, cfg = data_source, portal_cfg add( - data_source, - portal_cfg.get("name", data_source), - portal_url, - portal_cfg.get("quality_score", 0.85), + served, + cfg.get("name", served), + portal_url if served == data_source else cfg.get("url", ""), + cfg.get("quality_score", 0.85), ) if not sources: @@ -283,6 +299,118 @@ def _all_portal_configs() -> dict[str, dict[str, Any]]: _HTML_TAG_RE = re.compile(r"<[^>]+>") +def _sql_action_missing(resp: httpx.Response) -> bool: + """CKAN's answer when DataStore SQL is disabled, not when a query is bad. + + A portal without ``ckan.datastore.sqlsearch.enabled`` (the Fair Store's + default) replies 400 "Action name not known: datastore_search_sql" — the + same status as a bad statement, so the status alone would keep it + retryable and the model would try again every turn. + """ + return resp.status_code == 400 and "Action name not known" in resp.text[:2000] + + +# Results that retrieved nothing (see the accounting in the tool loop). +_NOT_RETRIEVED_PREFIXES = ("Error", "HTTP ", "SQL error", "Tool unavailable:", "Rows elsewhere:") + + +def classify_tool_result(text: str) -> str: + """``retrieved``, ``error``, ``unavailable`` or ``redirected``. + + "Error" (not "Error:"): MCP failures read "Error calling ...". + notebook_generator._NOTHING_FETCHED uses the same prefixes. + """ + if text.startswith(LLMAnalysisAgent.TOOL_UNAVAILABLE_PREFIX): + return "unavailable" + if text.startswith(LLMAnalysisAgent.TOOL_REDIRECT_PREFIX): + return "redirected" + if text.startswith(("Error", "HTTP ", "SQL error")): + return "error" + return "retrieved" + + +def _normalize_load_args(tool_name: str, tool_input: Any) -> None: + """Repair ``fields`` sent as a JSON string (``'["NAME", "ACRES"]'``), which + DataStore rejects, before the call and the notebook both use it.""" + if tool_name != "load_resource_data" or not isinstance(tool_input, dict): + return + fields = tool_input.get("fields") + if isinstance(fields, str): + try: + parsed = json.loads(fields) + except ValueError: + parsed = [part.strip() for part in fields.split(",") if part.strip()] + tool_input["fields"] = [str(f) for f in parsed] if isinstance(parsed, list) else fields + + +_ID_ARGS = ("dataset_id", "resource_id") + + +def _is_tabular(resource: dict[str, Any]) -> bool: + """Whether a mirrored DCAT resource is a CSV/TSV file.""" + from data_concierge.data_layer.connectors.dcat import _TABULAR_MEDIA_TYPES + + mimetype = str(resource.get("mimetype") or "").split(";")[0].strip().lower() + fmt = str(resource.get("format") or "").strip().upper() + return mimetype in _TABULAR_MEDIA_TYPES or fmt in ("CSV", "TSV") + + +def _result_of(resp: httpx.Response) -> Any: + """An action response's ``result``, or ``None`` for anything unexpected.""" + if not resp.is_success: + return None + try: + body = resp.json() + except ValueError: + return None + return body.get("result") if isinstance(body, dict) else None + + +def _remember_portal(tool_input: Any, portal: str, portal_of_id: dict[str, str]) -> None: + """Record that ``portal`` answered a successful call about these IDs. + + ``portal`` is captured before the call: ``_execute_tool`` pops + ``portal_id`` from the input. + """ + if not isinstance(tool_input, dict): + return + for key in _ID_ARGS: + value = tool_input.get(key) + if isinstance(value, str) and value: + portal_of_id[value] = portal + + +def _stick_to_portal(tool_input: Any, portal_of_id: dict[str, str]) -> None: + """Send a call with no ``portal_id`` where its ID was last fetched from. + + An explicit ``portal_id`` always wins; this only fills one the model left + out for an ID it already used on another portal in this run. + """ + if not isinstance(tool_input, dict) or tool_input.get("portal_id"): + return + for key in _ID_ARGS: + portal = portal_of_id.get(str(tool_input.get(key) or "")) + if portal: + tool_input["portal_id"] = portal + return + + +def _int_param(params: dict, key: str, default: int, *, cap: int, floor: int = 1) -> int: + """A numeric tool argument, tolerating the model sending ``"100"``. + + A value below ``floor`` means "use the default", as ``int(x or default)`` + did before; CKAN loads pass ``floor=0`` because ``limit=0`` is a + count-only DataStore query. + """ + try: + value = int(params.get(key, default)) + except (TypeError, ValueError, OverflowError): + value = default + if value < floor: + value = default + return min(value, cap) + + def _summarize_error_body(text: str, limit: int = 200) -> str: """Condense an error body for the model. @@ -770,6 +898,49 @@ def _build_system_prompt( # -- tool execution ---------------------------------------------------- + async def _run_portal_tool( + self, + tool_name: str, + tool_input: Any, + portal_url: str, + *, + portal_of_id: dict[str, str] | None = None, + ) -> tuple[str, str, str | None]: + """One portal tool call, as both the analysis loop and the notebook editor run it. + + Repairs the model's arguments, fills a dropped ``portal_id`` from + ``portal_of_id`` (Fair Store chats), routes a Fair Store row load to + its origin, and resolves — before :meth:`_execute_tool` pops + ``portal_id`` — the portal the call reaches, so the notebook code and + the attribution name the portal that answered. Returns ``(result_text, + code, served_portal_id)``; the last is ``None`` for the primary portal. + """ + _normalize_load_args(tool_name, tool_input) + if portal_of_id is not None: + _stick_to_portal(tool_input, portal_of_id) + routed_from = await self._route_to_origin(tool_name, tool_input, portal_url) + code_url = self._tool_portal_url(tool_input, portal_url) + served: str | None = None + if isinstance(tool_input, dict) and tool_input.get("portal_id"): + # The registry ID of the URL actually called: canonical case, and an + # unknown portal_id names the portal _execute_tool fell back to. + served = _site_id_for_url(code_url) or str(tool_input["portal_id"]) + result_text = await self._execute_tool(tool_name, tool_input, portal_url) + if routed_from and classify_tool_result(result_text) == "retrieved": + # Only on success: a failed routed load keeps its error prefix. + result_text = f"{routed_from}\n{result_text}" + code = self._code_for_tool(tool_name, tool_input, code_url) + return result_text, code, served + + def _tool_portal_url(self, tool_input: Any, portal_url: str) -> str: + """The portal a tool call will reach: its ``portal_id``, else the primary.""" + override_id = tool_input.get("portal_id") if isinstance(tool_input, dict) else None + if override_id: + override_url = (self.get_portal_config(override_id) or {}).get("url") + if override_url: + return str(override_url) + return portal_url + async def _execute_tool( self, tool_name: str, @@ -856,7 +1027,7 @@ async def _tool_semantic_search(self, params: dict, site_id: str | None = None) querying. """ query = params.get("query", "") - n_results = min(params.get("n_results", 10), 20) + n_results = _int_param(params, "n_results", 10, cap=20) if not query: return "Error: query parameter is required" @@ -933,7 +1104,7 @@ def _get_dcat_client(self, portal_url: str, catalog_url: str | None = None) -> A async def _tool_dcat_search(self, dcat: Any, params: dict) -> str: query = params.get("query", "") - rows = min(int(params.get("rows", 10) or 10), 20) + rows = _int_param(params, "rows", 10, cap=20) matches = await dcat.search_datasets(query, limit=rows) if not matches: @@ -970,7 +1141,7 @@ async def _tool_dcat_dataset_info(self, dcat: Any, params: dict) -> str: ds = await dcat.get_dataset(dataset_id) if ds is None: return ( - f"Dataset '{dataset_id}' not found in the {dcat.portal_url} catalog. " + f"Error: dataset '{dataset_id}' not found in the {dcat.portal_url} catalog. " "Use search_datasets first and pass a Dataset ID from its results." ) @@ -1022,7 +1193,7 @@ async def _tool_dcat_load(self, dcat: Any, params: dict) -> str: get_dataset_info. """ resource_id = str(params.get("resource_id") or "").strip() - limit = min(int(params.get("limit", 100) or 100), 1000) + limit = _int_param(params, "limit", 100, cap=1000) if not resource_id: return "Error: resource_id is required (a dataset ID or a distribution URL)." @@ -1034,13 +1205,13 @@ async def _tool_dcat_load(self, dcat: Any, params: dict) -> str: ds = await dcat.get_dataset(resource_id) if ds is None: return ( - f"Dataset '{resource_id}' not found in the {dcat.portal_url} " + f"Error: dataset '{resource_id}' not found in the {dcat.portal_url} " "catalog. Use search_datasets to find a valid Dataset ID." ) tabular = ds.tabular_distributions if not tabular: return ( - f"Dataset '{ds.title}' has no tabular (CSV/TSV) distribution, " + f"Error: dataset '{ds.title}' has no tabular (CSV/TSV) distribution, " "so its rows cannot be loaded." ) target_url = tabular[0].best_url @@ -1053,7 +1224,7 @@ async def _tool_dcat_load(self, dcat: Any, params: dict) -> str: rows = result["rows"] if not rows: return ( - f"No rows returned from {target_url} " + f"Error: no rows returned from {target_url} " f"({result['bytes_read']:,} bytes read)." ) @@ -1083,7 +1254,7 @@ async def _tool_dcat_load(self, dcat: Any, params: dict) -> str: async def _tool_search(self, client: httpx.AsyncClient, params: dict) -> str: query = params.get("query", "") org = params.get("organization") - rows = min(params.get("rows", 10), 20) + rows = _int_param(params, "rows", 10, cap=20) search_params: dict[str, Any] = {"q": query, "rows": rows} if org: @@ -1141,6 +1312,20 @@ async def _tool_dataset_info(self, client: httpx.AsyncClient, params: dict) -> s if tags: lines.append(f"Tags: {', '.join(tags)}") + extras = { + str(e.get("key")): e.get("value") + for e in pkg.get("extras") or [] + if isinstance(e, dict) and e.get("key") + } + origin_id = str(extras.get("mirror_source_portal") or "") + origin = self._mirror_origin(origin_id) + if origin_id: + # A Fair Store record: metadata here, rows at the portal it mirrors. + label = origin["name"] if origin else extras.get("mirror_source_url") or origin_id + lines.append(f"Mirrored from: {label} (portal_id `{origin_id}`)") + if extras.get("qsv_description"): + lines.append(f"AI summary (qsv): {str(extras['qsv_description'])[:500]}") + resources = pkg.get("resources", []) lines.append(f"\nResources ({len(resources)}):") for i, res in enumerate(resources, 1): @@ -1150,15 +1335,71 @@ async def _tool_dataset_info(self, client: httpx.AsyncClient, params: dict) -> s lines.append(f" Format: {res.get('format', '?')} DataStore: {ds}") if res.get("description"): lines.append(f" {res['description'][:150]}") + route = self._origin_route(res, origin) + if route: + lines.append(f" Rows: {route}") return "\n".join(lines) + def _mirror_origin(self, origin_id: str) -> dict[str, Any] | None: + """The registered portal a Fair Store record was mirrored from.""" + if not origin_id: + return None + try: + from data_concierge.gateway import ckan_sites + + site = ckan_sites.get_site(origin_id) + except Exception: # pragma: no cover - registry failures are non-fatal + return None + if not site: + return None + return { + "id": origin_id, + "name": site.get("name") or origin_id, + "portal_type": _normalize_portal_type(site.get("portal_type")), + } + + @staticmethod + def _origin_route(resource: dict[str, Any], origin: dict[str, Any] | None) -> str: + """How to load a mirrored resource's rows from the portal that has them. + + A mirrored resource is a link: its rows are in the origin's DataStore + (CKAN, same resource UUID) or its downloadable file (DCAT), never in the + Fair Store's. The qsv tables the Fair Store generated itself carry no + ``mirror_source_portal`` and load here as usual. + """ + if not origin or resource.get("datastore_active"): + return "" + if str(resource.get("mirror_source_portal") or origin["id"]) != origin["id"]: + return "" + if origin["portal_type"] == _PORTAL_TYPE_DCAT: + url = str(resource.get("url") or "") + # Only a CSV/TSV file has rows to load: the mirror makes one resource + # per distribution, and a JSON, XML, KML or PDF one parsed as CSV + # "loads" garbage (DCATDistribution.is_tabular's rule). + if not url.lower().startswith(("http://", "https://")) or not _is_tabular(resource): + return "" + return ( + f"load_resource_data with portal_id=`{origin['id']}` and " + f"resource_id=`{url}`" + ) + if not resource.get("mirror_source_datastore_active"): + return "" + source_id = str(resource.get("mirror_source_id") or resource.get("id") or "") + return ( + f"load_resource_data with portal_id=`{origin['id']}` and " + f"resource_id=`{source_id}` (the origin's DataStore)" + ) + async def _tool_load(self, client: httpx.AsyncClient, params: dict) -> str: resource_id = params["resource_id"] - limit = min(params.get("limit", 100), 500) + limit = _int_param(params, "limit", 100, cap=500, floor=0) body: dict[str, Any] = {"resource_id": resource_id, "limit": limit} if params.get("offset"): - body["offset"] = params["offset"] + try: + body["offset"] = max(0, int(params["offset"])) + except (TypeError, ValueError): + pass if params.get("filters"): body["filters"] = params["filters"] if params.get("q"): @@ -1169,6 +1410,10 @@ async def _tool_load(self, client: httpx.AsyncClient, params: dict) -> str: body["fields"] = params["fields"] resp = await client.post("/api/3/action/datastore_search", json=body) + if resp.status_code == 404: + redirect = await self._mirrored_rows_hint(client, resource_id) + if redirect: + return redirect resp.raise_for_status() data = resp.json() @@ -1193,6 +1438,74 @@ async def _tool_load(self, client: httpx.AsyncClient, params: dict) -> str: lines.append(df.to_string(index=False, max_colwidth=40)) return "\n".join(lines) + async def _route_to_origin(self, tool_name: str, tool_input: Any, portal_url: str) -> str: + """Send a Fair Store row load to the portal the resource was mirrored from. + + The Fair Store holds metadata; a mirrored resource's rows are at its + origin (same UUID in a CKAN origin's DataStore, or the file a DCAT + origin publishes). Rewriting ``portal_id``/``resource_id`` here, before + the call, means the notebook records the portal that answered. Returns + a note for the model, or ``""`` when the load stays where it is. + """ + if tool_name != "load_resource_data" or not isinstance(tool_input, dict): + return "" + from data_concierge.gateway.fairstore import PORTAL_ID as FAIRSTORE_ID + + target_url = self._tool_portal_url(tool_input, portal_url) + if (tool_input.get("portal_id") or _site_id_for_url(target_url)) != FAIRSTORE_ID: + return "" + resource_id = str(tool_input.get("resource_id") or "") + if not resource_id or resource_id.lower().startswith(("http://", "https://")): + return "" + try: + client = await self._get_http_client(target_url) + resp = await client.get( + "/api/3/action/resource_show", params={"id": resource_id}, timeout=10.0 + ) + except httpx.HTTPError: + return "" + resource = _result_of(resp) + if not isinstance(resource, dict): + return "" + origin = self._mirror_origin(str(resource.get("mirror_source_portal") or "")) + if not self._origin_route(resource, origin) or origin is None: + return "" + if origin["portal_type"] == _PORTAL_TYPE_DCAT: + tool_input["resource_id"] = str(resource["url"]) + else: + tool_input["resource_id"] = str(resource.get("mirror_source_id") or resource_id) + tool_input["portal_id"] = origin["id"] + self.logger.info( + "Fair Store load routed to origin", resource_id=resource_id, origin=origin["id"] + ) + return ( + f"(The Fair Store mirrors this resource; its rows were loaded from " + f"{origin['name']}, portal_id `{origin['id']}`, resource " + f"`{tool_input['resource_id']}`.)" + ) + + async def _mirrored_rows_hint(self, client: httpx.AsyncClient, resource_id: str) -> str: + """Point a DataStore miss on a mirrored resource at the portal with the rows.""" + try: + # Short: this probe runs after every DataStore 404 on any CKAN portal. + resp = await client.get( + "/api/3/action/resource_show", params={"id": resource_id}, timeout=5.0 + ) + except httpx.HTTPError: + return "" + resource = _result_of(resp) + if not isinstance(resource, dict): + return "" + origin = self._mirror_origin(str(resource.get("mirror_source_portal") or "")) + route = self._origin_route(resource, origin) + if not route or not origin: + return "" + return ( + f"{self.TOOL_REDIRECT_PREFIX} resource `{resource_id}` is a mirror of " + f"{origin['name']}'s resource; its rows are not stored on this portal. " + f"Call {route}." + ) + def _sql_unavailable_status(self, portal_url: str) -> int | None: """Status that disabled SQL for this portal, or ``None`` if usable.""" entry = self._sql_disabled.get(portal_url) @@ -1223,6 +1536,10 @@ def _disable_sql(self, portal_url: str, status: int) -> None: # fetched nothing, and counting it as a failure would penalise the agent # for correctly routing around a portal-side outage. TOOL_UNAVAILABLE_PREFIX = "Tool unavailable:" + # A call that reached the right tool on the wrong portal: the rows exist, + # elsewhere. Excluded from the success rate like the above, but the tool is + # NOT withdrawn — the model needs it to follow the redirect. + TOOL_REDIRECT_PREFIX = "Rows elsewhere:" @staticmethod def _sql_unavailable_message(status: int, detail: str = "") -> str: @@ -1247,6 +1564,19 @@ def _sql_unavailable_message(status: int, detail: str = "") -> str: async def _tool_sql(self, client: httpx.AsyncClient, params: dict) -> str: portal_url = str(client.base_url).rstrip("/") + from data_concierge.gateway.fairstore import PORTAL_ID as FAIRSTORE_ID + + if _site_id_for_url(portal_url) == FAIRSTORE_ID: + # The Fair Store holds catalog metadata, no tables. Not a dead + # endpoint: tripping the breaker would withdraw SQL for the whole run + # (and 30 minutes of Fair Store chats) when the origin may serve it. + return ( + f"{self.TOOL_REDIRECT_PREFIX} the Fair Store holds catalog metadata " + "only, so run_sql_query has no tables here. Pass the portal_id the " + "dataset was mirrored from (get_dataset_info shows it as 'Mirrored " + "from'), or use load_resource_data, which loads from that portal." + ) + # Already known dead for this portal: answer without a round trip. known = self._sql_unavailable_status(portal_url) if known is not None: @@ -1270,7 +1600,7 @@ async def _tool_sql(self, client: httpx.AsyncClient, params: dict) -> str: "execution timeout. Narrow the query (add filters, " "aggregate, or reduce the result set)." ) - if resp.status_code in _SQL_ENDPOINT_DEAD_STATUSES: + if resp.status_code in _SQL_ENDPOINT_DEAD_STATUSES or _sql_action_missing(resp): # Endpoint-level failure: remember it so the rest of this query — # and the next one against the same portal — skips SQL entirely. self._disable_sql(portal_url, resp.status_code) @@ -1363,10 +1693,7 @@ def _code_for_dcat_tool(tool_name: str, tool_input: dict, portal_url: str) -> st # NOT CKAN's package_search, which does not exist here and would # 404 on every run. query = tool_input.get("query", "") - try: - rows = int(tool_input.get("n_results", 10)) - except (TypeError, ValueError): - rows = 10 + rows = _int_param(tool_input, "n_results", 10, cap=20) return ( "# Resource discovery. The concierge used a semantic (vector) search\n" "# over pre-indexed resource metadata here; the public equivalent for a\n" @@ -1378,7 +1705,7 @@ def _code_for_dcat_tool(tool_name: str, tool_input: dict, portal_url: str) -> st if tool_name == "search_datasets": query = tool_input.get("query", "") - rows = min(int(tool_input.get("rows", 10) or 10), 20) + rows = _int_param(tool_input, "rows", 10, cap=20) return "# Search the DCAT catalog (fetched once, searched locally)\n" + ( LLMAnalysisAgent._dcat_catalog_search_code(catalog_url, query, rows, helper) ) @@ -1410,7 +1737,7 @@ def _code_for_dcat_tool(tool_name: str, tool_input: dict, portal_url: str) -> st if tool_name == "load_resource_data": rid = str(tool_input.get("resource_id", "")) - limit = min(int(tool_input.get("limit", 100) or 100), 1000) + limit = _int_param(tool_input, "limit", 100, cap=1000) if rid.lower().startswith(("http://", "https://")): preamble = ( @@ -1465,10 +1792,7 @@ def _code_for_tool(tool_name: str, tool_input: dict, portal_url: str) -> str: if tool_name == "semantic_search_resources": q = tool_input.get("query", "") - try: - n = int(tool_input.get("n_results", 10)) - except (TypeError, ValueError): - n = 10 + n = _int_param(tool_input, "n_results", 10, cap=20) # The live tool searches a private Pinecone index, which the # published notebook cannot reach — importing it made every # notebook fail both in Colab and in the verifier's kernel. The @@ -1496,10 +1820,7 @@ def _code_for_tool(tool_name: str, tool_input: dict, portal_url: str) -> str: # the user later runs, so they cannot be interpolated raw: an org # like `x"}\nimport os\nos.system(...)` would inject code. repr the # Solr fq value and coerce rows to an int. - try: - rows = int(tool_input.get("rows", 10)) - except (TypeError, ValueError): - rows = 10 + rows = _int_param(tool_input, "rows", 10, cap=20) # repr the whole fq value so any quotes/newlines in org are escaped # into a proper Python string literal in the generated cell. fq = f', "fq": {f"organization:{org}"!r}' if org else "" @@ -1534,7 +1855,8 @@ def _code_for_tool(tool_name: str, tool_input: dict, portal_url: str) -> str: if tool_name == "load_resource_data": rid = tool_input["resource_id"] - limit = tool_input.get("limit", 100) + # What _tool_load sent, and never raw model text inside the code. + limit = _int_param(tool_input, "limit", 100, cap=500, floor=0) parts = [f'"resource_id": {rid!r}', f'"limit": {limit}'] if tool_input.get("filters"): parts.append(f'"filters": {json.dumps(tool_input["filters"])}') @@ -1839,6 +2161,12 @@ async def process(self, state: GraphState) -> GraphState: # noqa: C901 resource_record_counts: list[int] = [] all_tool_result_texts: list[str] = [] _prev_sql_errors: set[str] = set() # track SQL tool_use IDs that errored + # Fair Store chats reach origin portals through portal_id, and the + # model often drops it on the next call about the same dataset. + portal_of_id: dict[str, str] = {} + from data_concierge.gateway.fairstore import PORTAL_ID as _FAIRSTORE_ID + + sticky_portals = _FAIRSTORE_ID in (data_source, _site_id_for_url(portal_url)) for iteration in range(max_iterations): iter_start = time.monotonic() @@ -1930,14 +2258,17 @@ async def process(self, state: GraphState) -> GraphState: # noqa: C901 tool_start = time.monotonic() # Route to MCP or CKAN tool handler + called_portal: str | None = None if tool_name.startswith("mcp__"): result_text = await self._execute_mcp_tool(tool_name, tool_input) code = self._code_for_mcp_tool(tool_name, tool_input, result_text) else: - result_text = await self._execute_tool( - tool_name, tool_input, portal_url + result_text, code, called_portal = await self._run_portal_tool( + tool_name, + tool_input, + portal_url, + portal_of_id=portal_of_id if sticky_portals else None, ) - code = self._code_for_tool(tool_name, tool_input, portal_url) tool_elapsed_ms = round((time.monotonic() - tool_start) * 1000) @@ -1946,13 +2277,14 @@ async def process(self, state: GraphState) -> GraphState: # noqa: C901 # success nor a failure — it is excluded from the # success-rate denominator so the rate keeps measuring # only calls that actually attempted retrieval. - is_unavailable = result_text.startswith( - self.TOOL_UNAVAILABLE_PREFIX - ) - is_error = not is_unavailable and result_text.startswith( - ("Error:", "HTTP ", "SQL error") - ) - if is_unavailable: + outcome = classify_tool_result(result_text) + is_unavailable = outcome == "unavailable" + is_redirect = outcome == "redirected" + is_error = outcome == "error" + retrieved = outcome == "retrieved" + if is_redirect: + pass # fetched nothing, but the tool stays available + elif is_unavailable: # Withdraw the tool for the REST of this run, not # just the next one. Telling the model in prose not # to retry did not work — a measured trial had it @@ -1980,6 +2312,8 @@ async def process(self, state: GraphState) -> GraphState: # noqa: C901 successful_tool_calls += 1 all_tool_result_texts.append(result_text[:10000]) + if sticky_portals and called_portal and retrieved: + _remember_portal(tool_input, called_portal, portal_of_id) # Semantic search score if tool_name == "semantic_search_resources" and not is_error: @@ -1992,7 +2326,7 @@ async def process(self, state: GraphState) -> GraphState: # noqa: C901 max_semantic_score = score_val # Row/record counts from load_resource_data - if tool_name == "load_resource_data" and not is_error: + if tool_name == "load_resource_data" and retrieved: import re as _re total_match = _re.search(r"Total records:\s*([\d,]+)", result_text) @@ -2001,7 +2335,10 @@ async def process(self, state: GraphState) -> GraphState: # noqa: C901 try: count = int(total_match.group(1).replace(",", "")) resource_record_counts.append(count) - total_rows_loaded += min(count, tool_input.get("limit", 100)) + total_rows_loaded += min( + count, + _int_param(tool_input, "limit", 100, cap=500, floor=0), + ) except ValueError: pass else: @@ -2030,7 +2367,7 @@ async def process(self, state: GraphState) -> GraphState: # noqa: C901 resource_ids_seen.add(rid) # SQL queries also touch resources - if tool_name == "run_sql_query" and not is_error: + if tool_name == "run_sql_query" and retrieved: import re as _re uuids = _re.findall( @@ -2052,12 +2389,16 @@ async def process(self, state: GraphState) -> GraphState: # noqa: C901 ): resource_metadata_modified = mod_str + # The portal that served the call: a portal_id override or a + # Fair Store load routed to its origin, else the primary. + served_by = called_portal or data_source tool_calls_trace.append( { "agent": self.name, "action": tool_name, "tool_name": tool_name, "arguments": tool_input, + "portal_id": served_by, "result_preview": result_text[:800], "code": code, } @@ -2081,9 +2422,11 @@ async def process(self, state: GraphState) -> GraphState: # noqa: C901 "iteration": iteration + 1, "tool": tool_name, "tool_use_id": tool_id, - "source": _source_for_tool(tool_name, data_source), + "source": _source_for_tool(tool_name, served_by), "operation_type": _operation_type_for_tool(tool_name), - "status": "error" if is_error else "success", + "status": "success" if retrieved else "error", + # Why nothing was retrieved, when nothing was. + "outcome": outcome, "input": tool_input, "result": result_text[:10000], "result_chars": len(result_text), diff --git a/src/data_concierge/agents/notebook_editor.py b/src/data_concierge/agents/notebook_editor.py index 8a584e6..70ea768 100644 --- a/src/data_concierge/agents/notebook_editor.py +++ b/src/data_concierge/agents/notebook_editor.py @@ -32,8 +32,11 @@ AGENT_LOG_FORMAT_VERSION, TOOLS, _operation_type_for_tool, + _remember_portal, + _site_id_for_url, _source_for_tool, _utc_now, + classify_tool_result, get_llm_agent, ) from data_concierge.core.config import settings @@ -252,6 +255,13 @@ async def edit_notebook( # noqa: C901 - one linear tool loop, mirrors llm_agent portal_cfg = agent.get_portal_config(data_source) portal_url = portal_cfg.get("url", "") + # Fair Store chats: an ID fetched from an origin portal stays there (the + # model drops portal_id on follow-up calls), as in the analysis loop. + from data_concierge.gateway.fairstore import PORTAL_ID as FAIRSTORE_ID + + portal_of_id: dict[str, str] | None = ( + {} if FAIRSTORE_ID in (data_source, _site_id_for_url(portal_url)) else None + ) mcp_tools = agent._get_mcp_tools() from data_concierge.agents.llm_agent import _STATIC_PORTAL_CONFIGS @@ -408,16 +418,24 @@ async def edit_notebook( # noqa: C901 - one linear tool loop, mirrors llm_agent successful_calls += 1 result.tool_result_texts.append(result_text[:_MAX_TOOL_RESULT_CHARS]) else: - result_text = await agent._execute_tool(tool_name, tool_input, portal_url) - code = agent._code_for_tool(tool_name, tool_input, portal_url) - source = _source_for_tool(tool_name, data_source) + # The analysis loop's per-call steps: Fair Store routing, the + # served portal and the code URL resolved before the call. + result_text, code, served = await agent._run_portal_tool( + tool_name, tool_input, portal_url, portal_of_id=portal_of_id + ) + source = _source_for_tool(tool_name, served or data_source) operation_type = _operation_type_for_tool(tool_name) - is_error = result_text.startswith(("Error", "HTTP ", "SQL error")) - if is_error: + outcome = classify_tool_result(result_text) + is_error = outcome != "retrieved" + # A redirect or an absent capability fetched nothing and is + # neither a success nor a failure (as in the analysis loop). + if outcome == "error": failed_calls += 1 - else: + elif outcome == "retrieved": successful_calls += 1 result.tool_result_texts.append(result_text[:_MAX_TOOL_RESULT_CHARS]) + if portal_of_id is not None and served: + _remember_portal(tool_input, served, portal_of_id) result_for_model = result_text[:_MAX_TOOL_RESULT_CHARS] if code: diff --git a/src/data_concierge/agents/notebook_generator.py b/src/data_concierge/agents/notebook_generator.py index d16142d..cdc95e9 100644 --- a/src/data_concierge/agents/notebook_generator.py +++ b/src/data_concierge/agents/notebook_generator.py @@ -52,6 +52,12 @@ def _compile_safe_code_cell(code: str, label: str = "generated code") -> nbforma ) +# Tool results that fetched nothing, so their code would fail on replay: an +# error, a capability the portal lacks ("Tool unavailable:", e.g. DataStore SQL +# disabled) or rows that live on another portal ("Rows elsewhere:"). +_NOTHING_FETCHED = ("Error", "HTTP ", "SQL error", "Tool unavailable:", "Rows elsewhere:") + + class NotebookGeneratorAgent(BaseAgent): """Agent responsible for generating reproducible Google Colab notebooks. @@ -494,7 +500,7 @@ def _create_mcp_cells( # A failed exploratory call stays for provenance but must not be # an executable step (same rule as the CKAN trace cells). - failed = preview.startswith(("Error", "HTTP ", "SQL error")) + failed = preview.startswith(_NOTHING_FETCHED) if failed: md = md.replace( f"## Step {i}: {label}\n\n", @@ -584,8 +590,9 @@ def _create_cells_from_llm_trace( # column, bad SQL) is kept for provenance but must not be an # executable step: reproduced verbatim it fails again, so the # notebook never runs top to bottom and verification (#131) - # scores every answer as broken. - failed = preview.startswith(("Error", "HTTP ", "SQL error")) + # scores every answer as broken. So is a call that fetched nothing + # (see _NOTHING_FETCHED). + failed = preview.startswith(_NOTHING_FETCHED) # Markdown description label = self._format_action(action) diff --git a/src/data_concierge/core/config.py b/src/data_concierge/core/config.py index c6aa6b2..f41f786 100644 --- a/src/data_concierge/core/config.py +++ b/src/data_concierge/core/config.py @@ -155,6 +155,13 @@ class Settings(BaseSettings): # Provide via env / .env only — never commit a real key (issue #91). ckan_api_key: SecretStr = Field(default=SecretStr("")) + # Fair Store — the CKAN that mirrors every registered portal + # (scripts/populate_fairstore.py). Seeds for the admin panel's Fair Store + # settings; an admin's saved values (fairstore_settings.json) take precedence. + # Deliberately separate from CKAN_URL/CKAN_API_KEY, which name data.dathere.com. + fairstore_url: str = "" + fairstore_api_key: SecretStr = Field(default=SecretStr("")) + # WPRDC (Western PA Regional Data Center) - City of Pittsburgh open data wprdc_ckan_url: str = "https://data.wprdc.org" wprdc_organization: str = "city-of-pittsburgh" diff --git a/src/data_concierge/data_layer/connectors/ckan.py b/src/data_concierge/data_layer/connectors/ckan.py index eb4cbc4..3d5562d 100644 --- a/src/data_concierge/data_layer/connectors/ckan.py +++ b/src/data_concierge/data_layer/connectors/ckan.py @@ -4,6 +4,7 @@ search over pre-indexed CKAN resources. """ +import json from typing import Any import httpx @@ -16,6 +17,61 @@ logger = get_logger(__name__) +class CKANActionError(Exception): + """A CKAN action that did not succeed. + + ``status_code`` is ``None`` when no response arrived at all (connection + refused, reset, timed out). + """ + + def __init__( + self, + action: str, + message: str, + *, + status_code: int | None = None, + error_type: str = "", + ) -> None: + super().__init__(f"{action}: {message}") + self.action = action + self.status_code = status_code + self.error_type = error_type + + @property + def not_found(self) -> bool: + return self.status_code == 404 or self.error_type == "Not Found Error" + + @property + def transient(self) -> bool: + """Worth retrying: no response, a server or gateway error, or rate limiting.""" + return ( + self.status_code is None + or self.status_code in {408, 425, 429} + or self.status_code >= 500 + ) + + +def _action_result(action_name: str, response: httpx.Response) -> Any: + """The ``result`` of a CKAN action response, or :class:`CKANActionError`.""" + try: + body = response.json() + except ValueError: + body = None + if response.is_success and isinstance(body, dict) and body.get("success"): + return body.get("result") + error = body.get("error") if isinstance(body, dict) else None + error_type = str(error.get("__type", "")) if isinstance(error, dict) else "" + if isinstance(error, dict) and error.get("message"): + message = str(error["message"]) + elif error: + message = json.dumps(error, default=str)[:500] + else: + message = f"HTTP {response.status_code}" + raise CKANActionError( + action_name, message, status_code=response.status_code, error_type=error_type + ) + + class CKANClient: """Client for interacting with CKAN data portals. @@ -33,17 +89,22 @@ def __init__( self, ckan_url: str | None = None, api_key: str | None = None, + *, + use_default_key: bool = True, ) -> None: """Initialize the CKAN client. Args: ckan_url: Base URL for CKAN instance. Defaults to settings. api_key: Optional CKAN API key for authenticated access. + use_default_key: Fall back to ``CKAN_API_KEY`` when ``api_key`` is + not given. Pass ``False`` for a portal that key does not + belong to, so it is never sent to a third party. """ self.ckan_url = (ckan_url or settings.ckan_url).rstrip("/") self._api_key = api_key or ( settings.ckan_api_key.get_secret_value() - if hasattr(settings, 'ckan_api_key') and settings.ckan_api_key + if use_default_key and hasattr(settings, 'ckan_api_key') and settings.ckan_api_key else None ) self._client: httpx.AsyncClient | None = None @@ -71,6 +132,19 @@ async def close(self) -> None: await self._client.aclose() self._client = None + async def call(self, action_name: str, params: dict[str, Any] | None = None) -> Any: + """Execute a CKAN API action, raising :class:`CKANActionError` on failure. + + Unlike :meth:`action`, a failure is never mistaken for an empty result, + which is what a caller that writes based on what it read needs. + """ + client = await self._get_client() + try: + response = await client.post(f"/api/3/action/{action_name}", json=params or {}) + except httpx.HTTPError as e: + raise CKANActionError(action_name, str(e) or type(e).__name__) from e + return _action_result(action_name, response) + async def action(self, action_name: str, params: dict[str, Any] | None = None) -> dict[str, Any]: """Execute a CKAN API action. @@ -79,45 +153,89 @@ async def action(self, action_name: str, params: dict[str, Any] | None = None) - params: Action parameters Returns: - API response result + API response result, or ``{}`` on any failure (see :meth:`call`) """ - client = await self._get_client() - try: - response = await client.post( - f"/api/3/action/{action_name}", - json=params or {}, - ) - response.raise_for_status() - data = response.json() - - if not data.get("success"): - error_msg = data.get("error", {}).get("message", "Unknown error") - self.logger.error("CKAN action failed", action=action_name, error=error_msg) - return {} - - return data.get("result", {}) - - except httpx.HTTPStatusError as e: + result = await self.call(action_name, params) + except CKANActionError as e: # Log 404s at debug level since they're expected for invalid resource IDs - if e.response.status_code == 404: + if e.not_found: self.logger.debug( "CKAN resource not found", action=action_name, - status_code=404, + status_code=e.status_code, resource_id=params.get("resource_id") if params else None, ) + elif e.status_code is None: + self.logger.error("CKAN API error", action=action_name, error=str(e)) else: self.logger.error( - "CKAN HTTP error", + "CKAN action failed", action=action_name, - status_code=e.response.status_code, + status_code=e.status_code, error=str(e), ) return {} except Exception as e: self.logger.error("CKAN API error", action=action_name, error=str(e)) return {} + return {} if result is None else result + + async def upload_call( + self, + action_name: str, + fields: dict[str, Any], + *, + filename: str, + content: bytes, + content_type: str = "application/octet-stream", + ) -> Any: + """Execute a CKAN action that carries a file (``resource_create``/``_patch``). + + CKAN takes the file as the multipart ``upload`` field and every other + field as form text, so non-string values are sent JSON-encoded. A + separate client is used because the shared one pins a JSON + ``Content-Type`` that would override the multipart boundary. + + Raises :class:`CKANActionError` on failure, as :meth:`call` does. + """ + headers = {"Authorization": self._api_key} if self._api_key else {} + data = { + key: value if isinstance(value, str) else json.dumps(value) + for key, value in fields.items() + if value is not None + } + try: + async with httpx.AsyncClient( + base_url=self.ckan_url, timeout=120.0, headers=headers + ) as client: + response = await client.post( + f"/api/3/action/{action_name}", + data=data, + files={"upload": (filename, content, content_type)}, + ) + except httpx.HTTPError as e: + raise CKANActionError(action_name, str(e) or type(e).__name__) from e + return _action_result(action_name, response) + + async def upload_action( + self, + action_name: str, + fields: dict[str, Any], + *, + filename: str, + content: bytes, + content_type: str = "application/octet-stream", + ) -> dict[str, Any]: + """:meth:`upload_call`, returning ``{}`` on failure (as :meth:`action`).""" + try: + result = await self.upload_call( + action_name, fields, filename=filename, content=content, content_type=content_type + ) + except Exception as e: + self.logger.error("CKAN upload failed", action=action_name, error=str(e)[:500]) + return {} + return result or {} async def package_search( self, diff --git a/src/data_concierge/data_layer/connectors/dcat.py b/src/data_concierge/data_layer/connectors/dcat.py index 41a1071..631d500 100644 --- a/src/data_concierge/data_layer/connectors/dcat.py +++ b/src/data_concierge/data_layer/connectors/dcat.py @@ -322,38 +322,37 @@ def parse_dataset(raw: dict[str, Any]) -> DCATDataset: ) -def parse_catalog(body: Any, catalog_url: str) -> DCATCatalog: - """Parse a catalog document into :class:`DCATCatalog`. +def raw_dataset_nodes(body: Any) -> list[dict[str, Any]]: + """Return the unparsed dataset nodes of a catalog document. Handles the three shapes seen in the wild: DCAT-US 1.1 (``{"dataset": [...]}``), a bare JSON-LD array of nodes, and a JSON-LD document with ``@graph``. """ - raw_datasets: list[dict[str, Any]] = [] - title = "" - if isinstance(body, dict): - title = _text(body.get("title")) if isinstance(body.get("dataset"), list): - raw_datasets = [d for d in body["dataset"] if isinstance(d, dict)] - elif isinstance(body.get("@graph"), list): - raw_datasets = [ + return [d for d in body["dataset"] if isinstance(d, dict)] + if isinstance(body.get("@graph"), list): + return [ node for node in body["@graph"] if isinstance(node, dict) and "Dataset" in str(node.get("@type", "")) ] - elif isinstance(body, list): - raw_datasets = [ + return [] + if isinstance(body, list): + typed = [ node for node in body if isinstance(node, dict) and "Dataset" in str(node.get("@type", "")) ] - if not raw_datasets: - # A plain array of dataset objects with no @type annotation. - raw_datasets = [ - node for node in body if isinstance(node, dict) and node.get("title") - ] + # A plain array of dataset objects with no @type annotation. + return typed or [node for node in body if isinstance(node, dict) and node.get("title")] + return [] + - datasets = [parse_dataset(node) for node in raw_datasets] +def parse_catalog(body: Any, catalog_url: str) -> DCATCatalog: + """Parse a catalog document into :class:`DCATCatalog`.""" + title = _text(body.get("title")) if isinstance(body, dict) else "" + datasets = [parse_dataset(node) for node in raw_dataset_nodes(body)] # Drop entries with neither a title nor an identifier — nothing to cite. datasets = [d for d in datasets if d.title or d.identifier] return DCATCatalog( @@ -453,31 +452,46 @@ async def fetch_catalog(self, *, force: bool = False) -> DCATCatalog: if cached and not force and (time.time() - cached[0]) < CATALOG_TTL_SECONDS: return cached[1] - client = await self._http() - body_bytes = bytearray() - async with client.stream("GET", url) as resp: - resp.raise_for_status() - async for chunk in resp.aiter_bytes(): - body_bytes.extend(chunk) - if len(body_bytes) > MAX_CATALOG_BYTES: - raise ValueError( - f"DCAT catalog at {url} exceeds " - f"{MAX_CATALOG_BYTES // (1024 * 1024)} MB; " - "point catalog_url at a filtered catalog instead" - ) - - import json - - catalog = parse_catalog(json.loads(body_bytes.decode("utf-8")), url) + body, size = await self._download_catalog(url) + catalog = parse_catalog(body, url) _CATALOG_CACHE[url] = (time.time(), catalog) self.logger.info( "DCAT catalog loaded", url=url, datasets=len(catalog.datasets), - bytes=len(body_bytes), + bytes=size, ) return catalog + async def fetch_raw_catalog(self) -> Any: + """Fetch the catalog document unparsed and uncached. + + :meth:`fetch_catalog` keeps only the fields the agent searches on. A + caller that republishes the catalog (the Fair Store mirror) needs every + field the portal published, so it takes the document as-is. + """ + body, _size = await self._download_catalog(await self.resolve_catalog_url()) + return body + + async def _download_catalog(self, url: str) -> tuple[Any, int]: + """Stream the catalog under :data:`MAX_CATALOG_BYTES` and decode it.""" + client = await self._http() + body_bytes = bytearray() + async with client.stream("GET", url) as resp: + resp.raise_for_status() + async for chunk in resp.aiter_bytes(): + body_bytes.extend(chunk) + if len(body_bytes) > MAX_CATALOG_BYTES: + raise ValueError( + f"DCAT catalog at {url} exceeds " + f"{MAX_CATALOG_BYTES // (1024 * 1024)} MB; " + "point catalog_url at a filtered catalog instead" + ) + + import json + + return json.loads(body_bytes.decode("utf-8")), len(body_bytes) + async def search_datasets( self, query: str, diff --git a/src/data_concierge/data_layer/qsv_profiling.py b/src/data_concierge/data_layer/qsv_profiling.py index 457f129..161a840 100644 --- a/src/data_concierge/data_layer/qsv_profiling.py +++ b/src/data_concierge/data_layer/qsv_profiling.py @@ -33,6 +33,7 @@ from datetime import UTC, datetime from io import StringIO from pathlib import Path +from typing import Any import httpx @@ -186,7 +187,9 @@ def _extract_qsv_tags(qsv_data: dict) -> list[str]: if isinstance(val, dict): resp = val.get("response", val) if isinstance(resp, dict): - raw = resp.get("tags", []) + # describegpt capitalises the key in some responses + # ({"Attribution": ..., "Tags": [...]}). + raw = next((v for k, v in resp.items() if str(k).lower() == "tags"), []) else: raw = resp if isinstance(raw, str): @@ -306,24 +309,111 @@ def build_index( } +def csv_header(path: Path) -> list[str]: + """A CSV's column names, as the Fair Store mirror reads them.""" + csv.field_size_limit(max(csv.field_size_limit(), 64 * 1024 * 1024)) + with path.open(newline="", encoding="utf-8-sig", errors="replace") as handle: + return [name for name in next(csv.reader(handle), []) if name] + + +def scrubbed_json_bytes(path: Path) -> bytes: + """A JSON file's bytes with provider keys redacted, as the mirror publishes it. + + ``populate_fairstore`` scrubs the raw text of ``qsv_dict.json`` before + publishing it and hashes the result, so storing exactly that text keeps a + mirror run from storage and one from local files identical (a per-value + scrub keeps the word after the key, which the raw scrub removes, and every + describegpt resource would flip between the two). Should raw scrubbing + ever break the JSON — a key right before a closing quote — each string + value is scrubbed instead. + """ + from data_concierge.data_layer.onboard_index import _scrub_secrets + + text = path.read_text(encoding="utf-8") + scrubbed = _scrub_secrets(text) + try: + json.loads(scrubbed) + except ValueError: + scrubbed = json.dumps(_scrub_values(json.loads(text)), indent=2) + return scrubbed.encode("utf-8") + + +def _scrub_values(value: Any) -> Any: + from data_concierge.data_layer.onboard_index import _scrub_secrets + + if isinstance(value, str): + return _scrub_secrets(value) + if isinstance(value, list): + return [_scrub_values(item) for item in value] + if isinstance(value, dict): + return {key: _scrub_values(item) for key, item in value.items()} + return value + + +# The qsv outputs populate_fairstore reads from a dataset directory. +MIRRORED_QSV_OUTPUTS = ("qsv_dict.json", "qsv_stats.csv", "qsv_frequency.csv") + + +def recorded_headers(base_dir: Path, prefix: str) -> dict[str, list[str]]: + """Each downloaded CSV's header, keyed as its storage key would be.""" + headers: dict[str, list[str]] = {} + for csv_file in sorted(base_dir.rglob("*.csv")): + if csv_file.name.startswith("qsv_"): + continue + try: + header = csv_header(csv_file) + except (OSError, csv.Error): + continue # an unreadable download must not stop the sync + # Even an empty header: locally the mirror reads [] from such a file. + headers[f"{prefix}/{csv_file.relative_to(base_dir.parent)}"] = header + return headers + + +def sync_manifest(base_dir: Path, prefix: str) -> dict[str, Any]: + """What a mirror run from storage needs to know about the local directory. + + ``headers``: each downloaded CSV's header. ``dirs``: every dataset directory + and the qsv outputs it holds now — so a directory that is absent stays + absent, and an output deleted since an earlier sync is not staged. + """ + dirs = { + f"{prefix}/{directory.relative_to(base_dir.parent)}": sorted( + name for name in MIRRORED_QSV_OUTPUTS if (directory / name).is_file() + ) + for directory in sorted(p for p in base_dir.iterdir() if p.is_dir()) + } + return {"version": 1, "headers": recorded_headers(base_dir, prefix), "dirs": dirs} + + def sync_to_storage(base_dir: Path, site_id: str, prefix: str = "ckan_onboard") -> int: """Sync JSON and qsv CSV files from local disk to the unified storage backend. ``prefix`` is the storage key namespace. It defaults to ``ckan_onboard`` because ``onboard_index`` reads from there and existing deployments already have data under that prefix. + + The downloaded CSVs themselves are not synced; ``sync_manifest.json`` + records their headers and which dataset directories and qsv outputs exist + (:func:`sync_manifest`). ``populate_fairstore`` checks each qsv output + against the header of the file it describes, and runs from the storage + backend when the admin panel starts a mirror on Cloud Run. + + JSON is scrubbed of provider keys on the way: qsv describegpt records its + own command line, ``--api-key`` included, in ``qsv_dict.json``, + ``meta.json`` and ``index.json``. """ synced = 0 + manifest = sync_manifest(base_dir, prefix) for json_file in base_dir.rglob("*.json"): rel = json_file.relative_to(base_dir.parent) key = f"{prefix}/{rel}" - with open(json_file) as f: - data = json.load(f) - storage.write_json(key, data) + storage.write_bytes(key, scrubbed_json_bytes(json_file)) synced += 1 for csv_file in base_dir.rglob("qsv_*.csv"): rel = csv_file.relative_to(base_dir.parent) key = f"{prefix}/{rel}" storage.write_bytes(key, csv_file.read_bytes()) synced += 1 - return synced + # Last: a manifest must never list objects an interrupted sync did not write. + storage.write_json(f"{prefix}/{site_id}/sync_manifest.json", manifest) + return synced + 1 diff --git a/src/data_concierge/gateway/ckan_sites.py b/src/data_concierge/gateway/ckan_sites.py index d2d46db..a440e8a 100644 --- a/src/data_concierge/gateway/ckan_sites.py +++ b/src/data_concierge/gateway/ckan_sites.py @@ -230,12 +230,15 @@ def add_site( added_by: str = "admin", portal_type: str = PORTAL_TYPE_CKAN, catalog_url: str | None = None, + managed_by: str | None = None, ) -> dict[str, Any]: """Add a new portal. Auto-generates an ID from the name if not given. ``portal_type`` selects the access mechanism — ``ckan`` for a live action API, ``dcat`` for a catalog document. ``catalog_url`` is DCAT-only and optional: when omitted the standard catalog paths are probed under ``url``. + ``managed_by`` marks an entry another settings page owns (the Fair Store's + chat-source entry), which only that page changes or removes. Raises ``ValueError`` if ``url`` or ``name`` is empty, or if a site with the chosen ID already exists. @@ -270,6 +273,8 @@ def add_site( "added_by": added_by, "added_at": _now(), } + if managed_by: + entry["managed_by"] = managed_by sites.append(entry) _save(sites) logger.info( diff --git a/src/data_concierge/gateway/fairstore.py b/src/data_concierge/gateway/fairstore.py new file mode 100644 index 0000000..a576f54 --- /dev/null +++ b/src/data_concierge/gateway/fairstore.py @@ -0,0 +1,581 @@ +"""Verikan's link to the Fair Store. + +The Fair Store is a CKAN (2.11 + ckanext-hierarchy) that mirrors every +registered portal's catalog, with its organization hierarchy and the qsv +profiles onboarding produced (``scripts/populate_fairstore.py``). This module +is everything Verikan knows about it: + +**Connection settings.** The Fair Store's URL and a sysadmin API token, +seeded from ``FAIRSTORE_URL`` / ``FAIRSTORE_API_KEY`` and overridable from the +admin panel (``fairstore_settings.json`` through the storage backend — the same +pattern as the GitHub settings). The token is never returned to a client. + +**The token is bound to the origin it was saved for** (scheme, host and port). +Changing the URL to a different origin without supplying a new token drops the +saved token, and the environment token is used only while the URL is on the +environment URL's origin. A URL edit can therefore never send the sysadmin +token to another server, or another service on the same host. A URL that is +already a registered source portal is refused (:func:`conflicting_portal`). + +**Chat source.** With ``chat_source`` on, the Fair Store is registered in the +portal registry under :data:`PORTAL_ID`, so users can pick it in chat and the +agent can search every mirrored portal at once. + +**Status and site settings.** Live reads through the CKAN action API, and the +CKAN options a sysadmin can change at runtime (site title, about text, logo, +custom CSS) — the Fair Store's own ``/ckan-admin/config`` page. + +Mirror runs are started through :mod:`data_concierge.gateway.onboarding_jobs`, +which already runs one supervised job at a time. +""" + +from __future__ import annotations + +import asyncio +import ipaddress +from datetime import UTC, datetime +from typing import Any +from urllib.parse import urlparse + +import httpx + +from data_concierge.core.config import settings as app_settings +from data_concierge.core.logging import get_logger +from data_concierge.data_layer.connectors.ckan import CKANActionError, _action_result +from data_concierge.data_layer.storage import storage + +logger = get_logger(__name__) + +_KEY = "fairstore_settings.json" + +#: Registry ID of the Fair Store when it is offered as a chat source. +PORTAL_ID = "fairstore" + +DEFAULT_PORTAL_NAME = "Verikan Fair Store" +DEFAULT_PORTAL_DESCRIPTION = ( + "One catalog that mirrors every registered open data portal (WPRDC, " + "data.pa.gov, datHere) with each portal's organization hierarchy. Profiled " + "datasets carry qsv data dictionaries, summary statistics, frequency tables " + "and AI-written descriptions. Row data stays at each source portal." +) +DEFAULT_QUALITY_SCORE = 0.85 + +_TIMEOUT = httpx.Timeout(20.0, connect=10.0) +_MAX_TEXT = 20_000 + +# A token may travel over plain http only to this machine. +_LOCAL_HOSTS = frozenset({"localhost", "host.docker.internal"}) + +#: The CKAN options a sysadmin can change at runtime (``config_option_update``), +#: minus ``ckan.site_url`` (changing it breaks every link CKAN generates), +#: ``ckan.theme`` (2.11 ships one theme) and the logo-upload pseudo-options. +SITE_OPTIONS: dict[str, dict[str, str]] = { + "ckan.site_title": {"label": "Site title", "kind": "text"}, + "ckan.site_description": {"label": "Site description", "kind": "text"}, + "ckan.site_logo": {"label": "Logo URL", "kind": "text"}, + "ckan.site_intro_text": {"label": "Home page intro (Markdown)", "kind": "textarea"}, + "ckan.site_about": {"label": "About page (Markdown)", "kind": "textarea"}, + "ckan.site_custom_css": {"label": "Custom CSS", "kind": "code"}, +} + + +class FairStoreError(RuntimeError): + """The Fair Store could not be reached, or refused a call.""" + + +def _now() -> str: + return datetime.now(UTC).isoformat() + + +def _host(url: str) -> str: + return (urlparse(url).hostname or "").lower() + + +def _origin(url: str) -> str: + """``scheme://host:port`` — what a saved token is bound to (``""`` if unset).""" + if not url: + return "" + try: + parsed = httpx.URL(url) + except httpx.InvalidURL: + return "" + port = parsed.port or (443 if parsed.scheme == "https" else 80) + return f"{parsed.scheme}://{parsed.host.lower()}:{port}" + + +def conflicting_portal(url: str) -> str | None: + """A registered source portal already at this origin, if any. + + The Fair Store must be its own CKAN: pointed at a source portal (say + data.dathere.com), a mirror run would write every other portal into it. + """ + from data_concierge.gateway import ckan_sites + + origin = _origin(url) + if not origin: + return None + for site in ckan_sites.list_sites(): + if site.get("managed_by") == "fairstore": + continue + if _origin(str(site.get("url") or "")) == origin: + return str(site.get("id")) + return None + + +def normalize_url(raw: str) -> str: + """Validate and normalise a Fair Store base URL (``""`` stays unset). + + Raises ``ValueError`` for anything that is not a plain http(s) origin plus + optional path: credentials, a query or fragment, or plain http to a host + other than this machine (the sysadmin token would travel in clear text). + """ + url = (raw or "").strip().rstrip("/") + if not url: + return "" + if any(ord(ch) <= 32 or ord(ch) >= 127 for ch in url): + # urlparse silently drops tabs and newlines, and a non-ASCII host is a + # different string from the one the token is bound to: use punycode. + raise ValueError("The Fair Store URL may contain only printable ASCII (use punycode)") + try: + httpx.URL(url).port # noqa: B018 - validates the port as httpx will + except (httpx.InvalidURL, ValueError) as exc: + raise ValueError(f"The Fair Store URL is not valid: {exc}") from exc + parsed = urlparse(url) + if parsed.scheme not in ("http", "https") or not parsed.hostname: + raise ValueError("The Fair Store URL must start with https:// and name a host") + if parsed.username or parsed.password: + raise ValueError("Put the API token in the token field, not in the URL") + if parsed.query or parsed.fragment: + raise ValueError("The Fair Store URL must not carry a query or fragment") + if parsed.scheme == "http" and not _is_local(parsed.hostname): + raise ValueError( + "Use https:// for a Fair Store on another host — the sysadmin token " + "would otherwise be sent in clear text" + ) + return url + + +def _is_local(host: str) -> bool: + host = host.lower() + if host in _LOCAL_HOSTS or host.endswith(".localhost"): + return True + try: + return ipaddress.ip_address(host).is_loopback + except ValueError: + return False + + +def _env_url() -> str: + try: + return normalize_url(app_settings.fairstore_url) + except ValueError: + logger.warning("Ignoring an invalid FAIRSTORE_URL") + return "" + + +def _env_token() -> str: + return app_settings.fairstore_api_key.get_secret_value().strip() + + +class SettingsUnreadable(FairStoreError): + """The saved settings exist but could not be read (not the same as none).""" + + +def _read_saved(*, strict: bool = False) -> dict[str, Any]: + """The saved settings; ``{}`` when there are none. + + A read failure degrades to ``{}`` for display, but ``strict`` callers — a + save (which would otherwise overwrite the token) and the chat-source sync + (which would otherwise delete the portal entry) — get + :class:`SettingsUnreadable` instead. + """ + try: + saved = storage.read_json(_KEY) + except Exception as exc: # noqa: BLE001 - degrade to defaults + if strict: + raise SettingsUnreadable(f"Fair Store settings could not be read: {exc}") from exc + logger.warning("Failed to read Fair Store settings", error=str(exc)) + return {} + return saved if isinstance(saved, dict) else {} + + +def load_settings(*, strict: bool = False) -> dict[str, Any]: + """The effective settings, token included — never hand this to a client. + + ``token_source`` is ``"admin"``, ``"environment"`` or ``""``: the saved + token is used only for the origin (scheme, host and port) it was saved for, + and the environment's only for the environment URL's origin. + """ + saved = _read_saved(strict=strict) + url = saved.get("url") if isinstance(saved.get("url"), str) else None + url = url if url is not None else _env_url() + + token, source = "", "" + origin = _origin(url) + saved_token = str(saved.get("token") or "") + if saved_token and origin and saved.get("token_origin") == origin: + token, source = saved_token, "admin" + elif _env_token() and origin and origin == _origin(_env_url()): + token, source = _env_token(), "environment" + + mirror_sites = saved.get("mirror_sites") + return { + "url": url, + "url_source": "admin" if "url" in saved else ("environment" if url else ""), + "token": token, + "token_source": source, + "chat_source": bool(saved.get("chat_source", False)), + "portal_name": str(saved.get("portal_name") or DEFAULT_PORTAL_NAME), + "portal_description": str(saved.get("portal_description") or DEFAULT_PORTAL_DESCRIPTION), + "quality_score": float(saved.get("quality_score", DEFAULT_QUALITY_SCORE)), + "mirror_sites": [str(s) for s in mirror_sites] if isinstance(mirror_sites, list) else [], + "enrich": bool(saved.get("enrich", True)), + "updated_at": saved.get("updated_at"), + "updated_by": saved.get("updated_by"), + } + + +def public_settings(current: dict[str, Any] | None = None) -> dict[str, Any]: + """Settings safe to send to the admin UI: the token becomes set/unset flags. + + ``mirror_sites`` lists the chosen portals still registered; + ``mirror_sites_missing`` the ones since removed. + """ + from data_concierge.gateway import ckan_sites + + current = current if current is not None else load_settings() + token = current.get("token") or "" + known = set(ckan_sites.list_site_ids()) + chosen = current.get("mirror_sites") or [] + return { + **{k: v for k, v in current.items() if k != "token"}, + "mirror_sites": [s for s in chosen if s in known], + "mirror_sites_missing": [s for s in chosen if s not in known], + "conflicting_portal": conflicting_portal(current.get("url") or ""), + "token_set": bool(token), + "token_masked": f"••••{token[-4:]}" if len(token) > 12 else ("••••" if token else ""), + "configured": bool(current.get("url")), + "site_options": SITE_OPTIONS, + } + + +def save_settings(updates: dict[str, Any], *, updated_by: str = "admin") -> dict[str, Any]: + """Validate and persist an admin's changes; returns the public settings. + + ``token``: a non-empty string sets it, blank keeps it (the UI never sees + the current value), and ``clear_token`` removes it. A URL change to a + different origin (scheme, host or port) without a new token drops the saved + token — see the module docstring. Raises ``ValueError`` on invalid input, + including a URL that is already a registered source portal. + """ + saved = _read_saved(strict=True) + before = load_settings(strict=True) + + if "url" in updates and updates["url"] is not None: + saved["url"] = normalize_url(str(updates["url"])) + clash = conflicting_portal(saved["url"]) + if clash: + raise ValueError( + f"That URL is registered as the portal '{clash}' on the Data Portals page. " + "The Fair Store must be a separate CKAN, or a mirror run would write into " + f"that portal. (If '{clash}' is this Fair Store, delete that entry and use " + "'Offer the Fair Store as a data source in chat' instead.)" + ) + origin = _origin(saved.get("url") if "url" in saved else _env_url()) + + new_token = str(updates.get("token") or "").strip() + if new_token: + if any(ch.isspace() for ch in new_token) or len(new_token) > 4096: + raise ValueError("That does not look like a CKAN API token") + if not origin: + raise ValueError("Set the Fair Store URL before its API token") + saved["token"] = new_token + saved["token_origin"] = origin + elif updates.get("clear_token"): + saved.pop("token", None) + saved.pop("token_origin", None) + elif saved.get("token") and saved.get("token_origin") != origin: + saved.pop("token", None) + saved.pop("token_origin", None) + logger.info("Fair Store token dropped: the URL moved to another origin", origin=origin) + + if updates.get("chat_source") is not None: + saved["chat_source"] = bool(updates["chat_source"]) + for key in ("portal_name", "portal_description"): + if updates.get(key) is not None: + text = str(updates[key]).strip()[:2000] + if key == "portal_name" and not text: + raise ValueError("The chat source needs a name") + saved[key] = text + if updates.get("quality_score") is not None: + score = float(updates["quality_score"]) + if not 0.0 <= score <= 1.0: + raise ValueError("Quality score must be between 0 and 1") + saved["quality_score"] = score + if updates.get("mirror_sites") is not None: + from data_concierge.gateway import ckan_sites + + known = set(ckan_sites.list_site_ids()) + sites = [str(s).strip() for s in updates["mirror_sites"] if str(s).strip()] + unknown = [s for s in sites if s not in known] + if unknown: + raise ValueError(f"Unknown portal(s): {', '.join(sorted(unknown))}") + saved["mirror_sites"] = sorted(set(sites) - {PORTAL_ID}) + if updates.get("enrich") is not None: + saved["enrich"] = bool(updates["enrich"]) + + saved["updated_at"] = _now() + saved["updated_by"] = updated_by + storage.write_json(_KEY, saved) + + current = load_settings() + sync_chat_source(current) + logger.info( + "Fair Store settings saved", + updated_by=updated_by, + url_changed=before["url"] != current["url"], + token_changed=before["token"] != current["token"], + chat_source=current["chat_source"], + ) + return public_settings(current) + + +def sync_chat_source(current: dict[str, Any] | None = None) -> None: + """Make the portal registry's ``fairstore`` entry match the settings. + + Only an entry this module created (``managed_by == "fairstore"``) is ever + changed or removed, so a portal an admin registered by hand under the same + ID is left alone. + """ + from data_concierge.gateway import ckan_sites + + if current is None: + try: + current = load_settings(strict=True) + except SettingsUnreadable as exc: + # Unreadable is not "turned off": leave the registry as it is. + logger.warning("Skipping the Fair Store chat-source sync", error=str(exc)) + return + # Two instances syncing at once can both add the entry; add_site renames + # the second to fairstore-2, which the Data Portals page cannot delete. + for site in ckan_sites.list_sites(): + if site.get("managed_by") == "fairstore" and site.get("id") != PORTAL_ID: + ckan_sites.remove_site(str(site["id"])) + logger.info("Removed a duplicate Fair Store chat-source entry", site_id=site["id"]) + existing = ckan_sites.get_site(PORTAL_ID) + managed = bool(existing and existing.get("managed_by") == "fairstore") + wanted = bool(current.get("chat_source") and current.get("url")) + if wanted and conflicting_portal(current["url"]): + # An environment URL is not validated on save; never shadow a portal. + logger.warning("The Fair Store URL is a registered source portal; not offering it") + wanted = False + + if existing and not managed: + if wanted: + logger.warning( + "A hand-registered portal already uses the Fair Store's ID; leaving it", + site_id=PORTAL_ID, + ) + return + if not wanted: + if managed: + ckan_sites.remove_site(PORTAL_ID) + return + + fields = { + "url": current["url"], + "name": current["portal_name"], + "description": current["portal_description"], + "quality_score": current["quality_score"], + "portal_type": ckan_sites.PORTAL_TYPE_CKAN, + } + if managed: + if any(existing.get(key) != value for key, value in fields.items()): # type: ignore[union-attr] + ckan_sites.update_site(PORTAL_ID, fields) + else: + ckan_sites.add_site( + site_id=PORTAL_ID, + keywords=["fair store", "all portals", "catalog", "data dictionary", "qsv"], + added_by="fairstore", + managed_by="fairstore", + **fields, + ) + + +# ----------------------------------------------------------------------------- +# CKAN calls +# ----------------------------------------------------------------------------- + + +async def _call( + current: dict[str, Any], + action: str, + params: dict[str, Any] | None = None, + *, + auth: bool = False, +) -> Any: + """One CKAN action against the Fair Store (the token only when ``auth``).""" + url = current.get("url") or "" + if not url: + raise FairStoreError("The Fair Store URL is not configured") + headers = {"Content-Type": "application/json"} + if auth: + if not current.get("token"): + raise FairStoreError("No Fair Store API token is configured") + headers["Authorization"] = current["token"] + try: + async with httpx.AsyncClient(base_url=url, timeout=_TIMEOUT, headers=headers) as client: + response = await client.post(f"/api/3/action/{action}", json=params or {}) + except (httpx.HTTPError, httpx.InvalidURL) as exc: + raise FairStoreError(f"{action}: {type(exc).__name__}: {exc}") from exc + try: + return _action_result(action, response) + except CKANActionError as exc: + raise FairStoreError(str(exc)) from exc + + +async def _count(current: dict[str, Any], fq: str | None = None) -> int | None: + params: dict[str, Any] = {"rows": 0, "include_private": True} + if fq: + params["fq"] = fq + try: + result = await _call(current, "package_search", params, auth=bool(current.get("token"))) + except FairStoreError: + return None + return int(result.get("count", 0)) if isinstance(result, dict) else None + + +async def status(current: dict[str, Any] | None = None) -> dict[str, Any]: + """A live health and content summary for the admin panel. + + Never raises: an unreachable Fair Store, or a token CKAN refuses, is + reported in the result so the panel can say what is wrong. + """ + from data_concierge.gateway import ckan_sites + + current = current if current is not None else load_settings() + report: dict[str, Any] = { + "configured": bool(current.get("url")), + "url": current.get("url") or "", + "reachable": False, + "checked_at": _now(), + } + if not report["configured"]: + return report + + try: + info = await _call(current, "status_show") + except FairStoreError as exc: + report["error"] = str(exc) + return report + report["reachable"] = True + report["ckan_version"] = info.get("ckan_version") + report["site_title"] = info.get("site_title") + report["extensions"] = sorted(info.get("extensions") or []) + + # A sysadmin-only action tells a valid sysadmin token from any other. + if current.get("token"): + try: + await _call(current, "config_option_list", auth=True) + report["token"] = "sysadmin" + except FairStoreError as exc: + report["token"] = "rejected" + report["token_error"] = str(exc) + else: + report["token"] = "missing" + + sources = [s for s in ckan_sites.list_sites() if s.get("id") and s.get("id") != PORTAL_ID] + counts = await asyncio.gather( + _count(current), + *(_count(current, f'mirror_source_portal:"{s["id"]}"') for s in sources), + ) + report["datasets"] = counts[0] + report["by_portal"] = [ + {"site_id": s["id"], "name": s.get("name") or s["id"], "datasets": n} + for s, n in zip(sources, counts[1:], strict=True) + ] + for action, key in (("organization_list", "organizations"), ("group_list", "groups")): + try: + report[key] = len(await _call(current, action, {"all_fields": False})) + except FairStoreError: + report[key] = None + return report + + +def _refuse_conflict(current: dict[str, Any]) -> None: + clash = conflicting_portal(current.get("url") or "") + if clash: + raise FairStoreError( + f"The Fair Store URL is the registered portal '{clash}'; its site settings " + "are not edited from here." + ) + + +async def get_site_config(current: dict[str, Any] | None = None) -> dict[str, Any]: + """The Fair Store's runtime-editable site options (sysadmin token needed).""" + current = current if current is not None else load_settings() + _refuse_conflict(current) + editable = set(await _call(current, "config_option_list", auth=True) or []) + keys = [key for key in SITE_OPTIONS if key in editable] + values = await asyncio.gather( + *(_call(current, "config_option_show", {"key": key}, auth=True) for key in keys) + ) + return { + "options": [ + {**SITE_OPTIONS[key], "key": key, "value": "" if value is None else str(value)} + for key, value in zip(keys, values, strict=True) + ] + } + + +async def update_site_config( + values: dict[str, Any], current: dict[str, Any] | None = None +) -> dict[str, Any]: + """Change the Fair Store's site options; only :data:`SITE_OPTIONS` pass.""" + current = current if current is not None else load_settings() + _refuse_conflict(current) + unknown = sorted(set(values) - set(SITE_OPTIONS)) + if unknown: + raise ValueError(f"Not an editable Fair Store option: {', '.join(unknown)}") + payload = {} + for key, value in values.items(): + text = "" if value is None else str(value) + if len(text) > _MAX_TEXT: + raise ValueError(f"{SITE_OPTIONS[key]['label']} is longer than {_MAX_TEXT} characters") + payload[key] = text + if not payload: + raise ValueError("Nothing to update") + await _call(current, "config_option_update", payload, auth=True) + logger.info("Fair Store site options updated", keys=sorted(payload)) + return await get_site_config(current) + + +def public_info() -> dict[str, Any]: + """What any visitor may know: whether there is a Fair Store, and its address.""" + current = load_settings() + return { + "configured": bool(current.get("url")), + "url": current.get("url") or "", + "name": current.get("portal_name") or DEFAULT_PORTAL_NAME, + "chat_source": bool(current.get("chat_source") and current.get("url")), + "portal_id": PORTAL_ID, + } + + +__all__ = [ + "DEFAULT_PORTAL_DESCRIPTION", + "conflicting_portal", + "DEFAULT_PORTAL_NAME", + "PORTAL_ID", + "SITE_OPTIONS", + "FairStoreError", + "get_site_config", + "load_settings", + "normalize_url", + "public_info", + "public_settings", + "save_settings", + "status", + "sync_chat_source", + "update_site_config", +] diff --git a/src/data_concierge/gateway/onboarding_jobs.py b/src/data_concierge/gateway/onboarding_jobs.py index 6da6570..edba054 100644 --- a/src/data_concierge/gateway/onboarding_jobs.py +++ b/src/data_concierge/gateway/onboarding_jobs.py @@ -26,6 +26,24 @@ interleave their downloads. A second start is refused with the running job's ID rather than queued, so the caller always knows what is happening. +**Fair Store mirror runs share the runner.** ``populate_fairstore.py`` +(:func:`start_mirror_job`) runs under the same lock, log and records, tagged +``kind="fairstore_mirror"``: a mirror reads the onboarding output a concurrent +onboarding run would be rewriting. The Fair Store's sysadmin token reaches the +child through ``FAIRSTORE_API_KEY``, never argv. + +**Any instance can show a run.** Cloud Run sends each poll to whichever +instance it likes, and only one holds the child process. A running job +re-persists its record (log tail included) every :data:`HEARTBEAT_SECONDS`, so +another instance reports it as running from storage; a "running" record whose +heartbeat has gone stale is reported as interrupted. A cancel that lands on +another instance is written to ``.cancel.json`` and carried out by its +owner's next heartbeat. Log offsets are absolute line numbers +(``log_line_count``), so a poll answered by a lagging instance or after the +tail passes :data:`MAX_LOG_LINES` never repeats or stalls. Two starts landing +on different instances at the same moment can both run; the storage backend +has no create-only write to take a lease with. + **Cloud Run caveat.** A child process keeps running only while the instance is alive and scheduled. Without ``--no-cpu-throttling`` the CPU is throttled to near zero between requests and a background job crawls; with @@ -37,10 +55,14 @@ from __future__ import annotations import asyncio +import json import os import re import shutil import sys +import tempfile +import threading +import time import uuid from collections import deque from datetime import UTC, datetime @@ -61,6 +83,15 @@ MAX_LOG_LINES = 4000 MAX_JOB_RECORDS = 50 +KIND_ONBOARDING = "onboarding" +KIND_FAIRSTORE_MIRROR = "fairstore_mirror" + +# A running job's record is re-persisted this often (see the module docstring), +# and a stored "running" record whose heartbeat is older than STALE_AFTER has no +# live process anywhere. +HEARTBEAT_SECONDS = 10 +STALE_AFTER_SECONDS = 60 + STATUS_RUNNING = "running" STATUS_SUCCEEDED = "succeeded" STATUS_FAILED = "failed" @@ -72,6 +103,7 @@ "limit": (0, 5000), "concurrency": (1, 10), "max_mb": (1, 2000), + "mirror_limit": (1, 5000), } _NAME_RE = re.compile(r"^[A-Za-z0-9_.-]{1,64}$") _CONTROL_RE = re.compile(r"[\x00-\x1f\x7f]") @@ -113,6 +145,11 @@ def _is_noise(line: str) -> bool: _logs: dict[str, deque[str]] = {} _jobs: dict[str, dict[str, Any]] = {} _start_lock = asyncio.Lock() +# Orders a heartbeat's storage write against the final one: a heartbeat write +# still in flight in a worker thread when the job ends must not land after the +# final record (it would read "running" again, then "interrupted"). +_record_lock = threading.Lock() +_finished: set[str] = set() class JobError(RuntimeError): @@ -287,6 +324,72 @@ def build_command(site: dict[str, Any], options: dict[str, Any]) -> tuple[list[s return argv, script_name +def build_mirror_command( + options: dict[str, Any], + *, + target_url: str, + summary_path: str, +) -> tuple[list[str], list[dict[str, Any]]]: + """Build the validated argv for a Fair Store mirror run. + + ``options["site"]`` is ``"all"`` or a registered portal ID, or + ``options["sites"]`` a list of them — never the Fair Store's own + chat-source entry. Returns ``(argv, sites)``, ``sites`` empty for every + portal. The token is not an argument: :func:`start_mirror_job` passes it + through the environment. + """ + from data_concierge.gateway import ckan_sites + from data_concierge.gateway.fairstore import PORTAL_ID, normalize_url + + directory = scripts_dir() + if directory is None: + raise JobError("The mirror script is not available in this deployment") + script_path = directory / "populate_fairstore.py" + if not script_path.exists(): + raise JobError(f"populate_fairstore.py is missing from {directory}") + + try: + url = normalize_url(target_url) + except ValueError as exc: + raise JobError(str(exc)) from exc + if not url: + raise JobError("Set the Fair Store URL first") + + requested = options.get("sites") or [options.get("site") or "all"] + if not isinstance(requested, list) or len(requested) > 50: + raise JobError("Choose 'all' or up to 50 portals") + sites: list[dict[str, Any]] = [] + if requested != ["all"]: + for raw in dict.fromkeys(str(r).strip() for r in requested): + site = ckan_sites.get_site(raw) + if ( + site is None + or str(site.get("id", "")).lower() == PORTAL_ID + or site.get("managed_by") == "fairstore" + ): + raise JobError(f"'{raw}' is not a portal the Fair Store mirrors") + sites.append(site) + + argv: list[str] = [ + sys.executable, + str(script_path), + "--site", + ",".join(str(site["id"]) for site in sites) or "all", + "--target-url", + url, + "--summary-file", + summary_path, + ] + if options.get("apply"): + argv.append("--apply") + if options.get("enrich") is False: + argv.append("--no-enrich") + limit = _clean_int_option(options.get("limit"), "mirror_limit") + if limit is not None: + argv += ["--limit", str(limit)] + return argv, sites + + def _public(job: dict[str, Any]) -> dict[str, Any]: """Strip internal fields before a record leaves this module.""" return {k: v for k, v in job.items() if not k.startswith("_")} @@ -296,32 +399,77 @@ def _job_key(job_id: str) -> str: return f"{_STORAGE_PREFIX}/{job_id}.json" -def _persist(job: dict[str, Any]) -> None: - """Write a job record, and keep the index of recent jobs trimmed. +def _cancel_key(job_id: str) -> str: + # Separate from the record, which the owner's heartbeat keeps rewriting. + return f"{_STORAGE_PREFIX}/{job_id}.cancel.json" + - ``_public`` is not optional here: the live record carries the supervising - asyncio Task under ``_task``, which is not JSON-serializable, and writing - the raw dict silently loses every completed job. +def _write_record(record: dict[str, Any]) -> bool: + """Write a job record and the index; ``False`` if storage failed. + + Pass ``_public(job)``, never the live record: it carries the supervising + asyncio Task under ``_task``, which is not JSON-serializable, and writing it + silently lost every completed job. """ + job = record try: - storage.write_json(_job_key(job["id"]), _public(job)) + storage.write_json(_job_key(job["id"]), record) index = storage.read_json(_INDEX_KEY) or {} ids = [i for i in index.get("job_ids", []) if i != job["id"]] ids.insert(0, job["id"]) dropped = ids[MAX_JOB_RECORDS:] storage.write_json(_INDEX_KEY, {"job_ids": ids[:MAX_JOB_RECORDS]}) for old in dropped: - try: - storage.delete(_job_key(old)) - except Exception: # noqa: BLE001 - best-effort cleanup - pass + for key in (_job_key(old), _cancel_key(old)): + try: + storage.delete(key) + except Exception: # noqa: BLE001 - best-effort cleanup + pass except Exception as exc: # noqa: BLE001 - persistence must not kill a job logger.warning("Failed to persist onboarding job", job_id=job["id"], error=str(exc)) + return False + return True + + +FINAL_WRITE_ATTEMPTS = 5 + + +def _locked_write(record: dict[str, Any]) -> None: + with _record_lock: + _write_record(record) + + +def _final_write(record: dict[str, Any]) -> None: + """The terminal record: after any in-flight heartbeat, retried on failure. + + Other instances read a job only from storage, so a lost final write would + leave it "running" until the heartbeat went stale, then "interrupted" + without its result. Runs in a worker thread (blocking I/O and sleeps). + """ + with _record_lock: + _finished.add(record["id"]) + for attempt in range(FINAL_WRITE_ATTEMPTS): + if _write_record(record): + break + if attempt < FINAL_WRITE_ATTEMPTS - 1: + time.sleep(min(2**attempt, 10)) + else: + logger.error("Could not persist a finished job", job_id=record["id"]) + try: + storage.delete(_cancel_key(record["id"])) + except Exception: # noqa: BLE001 - best-effort cleanup + pass def _summary(job: dict[str, Any]) -> dict[str, Any]: - """A job record without its log, for list views.""" - return {k: v for k, v in job.items() if k != "log"} + """A job record without its log, for list views. + + Records written before mirror runs existed carry no ``kind``; they were + all onboarding runs. + """ + record = {k: v for k, v in job.items() if k != "log"} + record.setdefault("kind", KIND_ONBOARDING) + return record def running_job_id() -> str | None: @@ -331,6 +479,73 @@ def running_job_id() -> str | None: return None +def _seconds_since(timestamp: Any) -> float | None: + try: + then = datetime.fromisoformat(str(timestamp).replace("Z", "+00:00")) + except ValueError: + return None + if then.tzinfo is None: + then = then.replace(tzinfo=UTC) + return (datetime.now(UTC) - then).total_seconds() + + +def _settle_stored(stored: dict[str, Any]) -> dict[str, Any]: + """Decide what a stored record this instance holds no process for means. + + Running with a fresh heartbeat: another instance is running it. Running + without one: the instance that ran it is gone, so it was interrupted. + """ + if stored.get("status") == STATUS_RUNNING: + age = _seconds_since(stored.get("heartbeat_at")) + if age is not None and age < STALE_AFTER_SECONDS: + stored["running_elsewhere"] = True + else: + stored["status"] = STATUS_FAILED + stored["error"] = "interrupted (the server running it restarted)" + return stored + + +def _running_elsewhere() -> str | None: + """A job another instance is running, from its fresh heartbeat.""" + try: + index = storage.read_json(_INDEX_KEY) or {} + for job_id in index.get("job_ids", [])[:5]: + if job_id in _jobs: + continue + stored = storage.read_json(_job_key(job_id)) or {} + if _settle_stored(stored).get("running_elsewhere"): + return str(job_id) + except Exception as exc: # noqa: BLE001 - never block a start on a read error + logger.warning("Could not check for jobs on other instances", error=str(exc)) + return None + + +async def _heartbeat(job_id: str) -> None: + """Re-persist a running job, and honour a cancel recorded by another instance.""" + while True: + await asyncio.sleep(HEARTBEAT_SECONDS) + job = _jobs.get(job_id) + if job is None or job.get("status") != STATUS_RUNNING: + return + try: + cancel_requested = await asyncio.to_thread(storage.exists, _cancel_key(job_id)) + except Exception: # noqa: BLE001 + cancel_requested = False + if cancel_requested: + await cancel_job(job_id) + return + job["heartbeat_at"] = _now() + job["log"] = list(_logs.get(job_id, [])) + # Snapshot on the loop thread; only the storage write leaves it. + await asyncio.to_thread(_heartbeat_write, _public(job)) + + +def _heartbeat_write(record: dict[str, Any]) -> None: + with _record_lock: + if record["id"] not in _finished: + _write_record(record) + + async def _pump_output(job_id: str, process: asyncio.subprocess.Process) -> None: """Read the child's merged output into the job's log tail.""" log = _logs[job_id] @@ -348,9 +563,29 @@ async def _pump_output(job_id: str, process: asyncio.subprocess.Process) -> None job["log_line_count"] = job.get("log_line_count", 0) + 1 +def _collect_result(job: dict[str, Any]) -> None: + """Fold a run's JSON summary file into the record, then remove the file.""" + path = job.pop("_summary_path", None) + if not path: + return + try: + with open(path, encoding="utf-8") as handle: + job["result"] = json.load(handle) + except FileNotFoundError: + pass # the run failed before writing one; the log says why + except (OSError, ValueError) as exc: + logger.warning("Unreadable job summary", job_id=job.get("id"), error=str(exc)) + finally: + try: + os.unlink(path) + except OSError: + pass + + async def _supervise(job_id: str, process: asyncio.subprocess.Process) -> None: """Wait for the child, then record its outcome.""" job = _jobs[job_id] + beat = asyncio.create_task(_heartbeat(job_id)) try: await _pump_output(job_id, process) exit_code = await process.wait() @@ -369,20 +604,102 @@ async def _supervise(job_id: str, process: asyncio.subprocess.Process) -> None: except Exception as exc: # noqa: BLE001 job["status"] = STATUS_FAILED job["error"] = str(exc) - logger.error("Onboarding job crashed", job_id=job_id, error=str(exc)) + logger.error("Job crashed", job_id=job_id, error=str(exc)) finally: + beat.cancel() job["finished_at"] = _now() job["log"] = list(_logs.get(job_id, [])) + _collect_result(job) _processes.pop(job_id, None) - _persist(job) + # Snapshot on the loop thread; the write (and its lock wait) leaves it. + await asyncio.shield(asyncio.to_thread(_final_write, _public(job))) logger.info( - "Onboarding job finished", + "Job finished", job_id=job_id, status=job.get("status"), exit_code=job.get("exit_code"), ) +async def _spawn( + argv: list[str], + job: dict[str, Any], + *, + env_overrides: dict[str, str | None] | None = None, +) -> dict[str, Any]: + """Start ``argv`` as the one running job; the caller holds ``_start_lock``. + + ``env_overrides`` sets (or, with ``None``, removes) variables in the + child's copy of the server environment — how secrets reach it without + appearing in argv. + """ + job_id = job["id"] + env = dict(os.environ) + env["PYTHONUNBUFFERED"] = "1" + for key, value in (env_overrides or {}).items(): + if value is None: + env.pop(key, None) + else: + env[key] = value + + try: + process = await asyncio.create_subprocess_exec( + *argv, + stdout=asyncio.subprocess.PIPE, + stderr=asyncio.subprocess.STDOUT, + env=env, + cwd=str(scripts_dir().parent), # type: ignore[union-attr] + ) + except OSError as exc: + raise JobError(f"Could not start {job.get('script') or 'the job'}: {exc}") from exc + + _jobs[job_id] = job + _logs[job_id] = deque(maxlen=MAX_LOG_LINES) + _processes[job_id] = process + # Under the record lock: the previous job's final write may still be + # updating the index, and interleaving would drop one of the two. + await asyncio.to_thread(_locked_write, _public(job)) + + # Supervised in the background; the caller gets the job record now. + task = asyncio.create_task(_supervise(job_id, process)) + job["_task"] = task # not persisted (see _write_record / _public) + logger.info( + "Job started", + job_id=job_id, + kind=job.get("kind"), + site_id=job.get("site_id"), + script=job.get("script"), + started_by=job.get("started_by"), + ) + return _public(job) + + +def _refuse_if_running() -> None: + active = running_job_id() or _running_elsewhere() + if active: + raise JobError( + f"Job {active} is already running. Wait for it to finish or cancel it — " + "onboarding and Fair Store mirror runs share one slot, because a mirror " + "reads what an onboarding run writes." + ) + + +def _new_job(**fields: Any) -> dict[str, Any]: + now = _now() + return { + "id": uuid.uuid4().hex[:12], + "status": STATUS_RUNNING, + "started_at": now, + "heartbeat_at": now, + "finished_at": None, + "exit_code": None, + "error": None, + "log_line_count": 0, + "log": [], + **fields, + } + + async def start_job( site: dict[str, Any], options: dict[str, Any], @@ -391,71 +708,82 @@ async def start_job( ) -> dict[str, Any]: """Validate options, spawn the onboarding script, and return the job record.""" async with _start_lock: - active = running_job_id() - if active: - raise JobError( - f"Job {active} is already running. Wait for it to finish or cancel it " - "— concurrent runs would interleave downloads into the same directory." - ) - + _refuse_if_running() argv, script_name = build_command(site, options) - job_id = uuid.uuid4().hex[:12] - + job = _new_job( + kind=KIND_ONBOARDING, + site_id=site.get("id"), + site_name=site.get("name"), + portal_type=site.get("portal_type") or "ckan", + script=script_name, + # argv is safe to show: secrets go through the environment. + command=" ".join(argv[1:]), + options=dict(options), + started_by=started_by, + ) # The child inherits the server's environment so OPENROUTER_API_KEY and # PINECONE_API_KEY reach it without ever appearing in argv. - env = dict(os.environ) - env["PYTHONUNBUFFERED"] = "1" - - job: dict[str, Any] = { - "id": job_id, - "site_id": site.get("id"), - "site_name": site.get("name"), - "portal_type": site.get("portal_type") or "ckan", - "script": script_name, - # argv is safe to show: secrets go through the environment. - "command": " ".join(argv[1:]), - "options": dict(options), - "status": STATUS_RUNNING, - "started_by": started_by, - "started_at": _now(), - "finished_at": None, - "exit_code": None, - "error": None, - "log_line_count": 0, - "log": [], - } + return await _spawn(argv, job) - try: - process = await asyncio.create_subprocess_exec( - *argv, - stdout=asyncio.subprocess.PIPE, - stderr=asyncio.subprocess.STDOUT, - env=env, - cwd=str(scripts_dir().parent), # type: ignore[union-attr] - ) - except OSError as exc: - raise JobError(f"Could not start the onboarding script: {exc}") from exc - - _jobs[job_id] = job - _logs[job_id] = deque(maxlen=MAX_LOG_LINES) - _processes[job_id] = process - _persist(job) - - # Supervised in the background; the caller gets the job record now. - task = asyncio.create_task(_supervise(job_id, process)) - job["_task"] = task # not persisted (see _persist / _summary consumers) - logger.info( - "Onboarding job started", - job_id=job_id, - site_id=site.get("id"), - script=script_name, + +async def start_mirror_job( + options: dict[str, Any], + *, + target_url: str, + token: str, + started_by: str = "admin", +) -> dict[str, Any]: + """Start a Fair Store mirror run (``populate_fairstore.py``). + + ``options``: ``site`` (``"all"`` or a portal ID), ``apply`` (write; the + default is a dry run), ``enrich`` (Socrata enrichment, default on) and + ``limit`` (smoke test). Writing needs the sysadmin ``token``. + """ + if options.get("apply") and not token: + raise JobError("Writing to the Fair Store needs its sysadmin API token") + async with _start_lock: + _refuse_if_running() + fd, summary_path = tempfile.mkstemp(prefix="fairstore-mirror-", suffix=".json") + os.close(fd) + os.unlink(summary_path) # the script creates it; absent means no summary + argv, sites = build_mirror_command( + options, target_url=target_url, summary_path=summary_path + ) + if len(sites) == 1: + site_name = sites[0].get("name") or sites[0]["id"] + elif sites: + site_name = "Portals: " + ", ".join(str(site["id"]) for site in sites) + else: + site_name = "All portals" + job = _new_job( + kind=KIND_FAIRSTORE_MIRROR, + site_id=",".join(str(site["id"]) for site in sites) or "all", + site_name=site_name, + portal_type=(sites[0].get("portal_type") or "ckan") if len(sites) == 1 else None, + script="populate_fairstore.py", + command=" ".join(argv[1:]), + options={ + "site": ",".join(str(site["id"]) for site in sites) or "all", + "apply": bool(options.get("apply")), + "enrich": options.get("enrich") is not False, + "limit": options.get("limit") or None, + }, + target_url=target_url, started_by=started_by, + _summary_path=summary_path, + ) + return await _spawn( + argv, + job, + env_overrides={ + "FAIRSTORE_API_KEY": token or None, + "FAIRSTORE_URL": target_url, + }, ) - return _public(job) -def list_jobs(limit: int = 25) -> list[dict[str, Any]]: - """Recent jobs, newest first, without their logs.""" +def list_jobs(limit: int = 25, *, kind: str | None = None) -> list[dict[str, Any]]: + """Recent jobs, newest first, without their logs (optionally one ``kind``).""" records: list[dict[str, Any]] = [_summary(_public(j)) for j in _jobs.values()] seen = {r["id"] for r in records} @@ -467,16 +795,13 @@ def list_jobs(limit: int = 25) -> list[dict[str, Any]]: continue stored = storage.read_json(_job_key(job_id)) if stored: - # A job recorded as running that we have no process for did not - # survive a restart; reporting it as running would be a lie. - if stored.get("status") == STATUS_RUNNING: - stored["status"] = STATUS_FAILED - stored["error"] = "interrupted (server restarted during the run)" - records.append(_summary(stored)) + records.append(_summary(_settle_stored(stored))) seen.add(job_id) except Exception as exc: # noqa: BLE001 logger.warning("Failed to read onboarding job index", error=str(exc)) + if kind: + records = [r for r in records if r.get("kind") == kind] records.sort(key=lambda r: r.get("started_at") or "", reverse=True) return records[:limit] @@ -484,8 +809,9 @@ def list_jobs(limit: int = 25) -> list[dict[str, Any]]: def get_job(job_id: str, *, log_offset: int = 0) -> dict[str, Any] | None: """One job with a slice of its log. - ``log_offset`` is a line index into the retained tail, so a UI can poll for - just what it has not shown yet instead of refetching the whole log. + ``log_offset`` is an absolute line number (the previous response's + ``log_next_offset``), so a UI can poll for just what it has not shown yet, + from any instance, however long the log grows. """ job = _jobs.get(job_id) if job is not None: @@ -495,34 +821,53 @@ def get_job(job_id: str, *, log_offset: int = 0) -> dict[str, Any] | None: stored = storage.read_json(_job_key(job_id)) if not stored: return None - if stored.get("status") == STATUS_RUNNING: - stored["status"] = STATUS_FAILED - stored["error"] = "interrupted (server restarted during the run)" - record = _summary(stored) + record = _summary(_settle_stored(stored)) lines = list(stored.get("log", [])) - offset = max(0, min(int(log_offset or 0), len(lines))) - record["log"] = lines[offset:] + # Absolute line numbers: ``lines`` is the tail ending at log_line_count. + total = max(int(record.get("log_line_count") or 0), len(lines)) + first = total - len(lines) + offset = max(0, int(log_offset or 0)) + record["log"] = lines[max(0, offset - first) :] if offset < total else [] record["log_offset"] = offset - record["log_next_offset"] = len(lines) - record["log_truncated"] = record.get("log_line_count", 0) > len(lines) + # Never backwards: a poll answered from a lagging snapshot returns nothing. + record["log_next_offset"] = max(offset, total) + record["log_truncated"] = offset < first return record async def cancel_job(job_id: str) -> bool: - """Terminate a running job. Returns ``False`` if it was not running.""" + """Terminate a running job. Returns ``False`` if it was not running. + + For a job another instance is running, a ``.cancel.json`` key is + written; its owner's next heartbeat terminates it. + """ process = _processes.get(job_id) job = _jobs.get(job_id) - if process is None or job is None or job.get("status") != STATUS_RUNNING: + if process is None or job is None: + stored = storage.read_json(_job_key(job_id)) or {} + if job is None and _settle_stored(dict(stored)).get("running_elsewhere"): + storage.write_json(_cancel_key(job_id), {"requested_at": _now()}) + # The owner may have finished (and cleared cancels) meanwhile. + again = storage.read_json(_job_key(job_id)) or {} + if not _settle_stored(dict(again)).get("running_elsewhere"): + storage.delete(_cancel_key(job_id)) + return False + logger.info("Cancel requested for a job on another instance", job_id=job_id) + return True return False + if job.get("status") != STATUS_RUNNING or process.returncode is not None: + return False # already exited; _supervise records how # Mark first so _supervise does not overwrite the status with "failed" # when the child exits non-zero because we killed it. + previous = (job.get("status"), job.get("error")) job["status"] = STATUS_CANCELLED job["error"] = "cancelled by admin" try: process.terminate() except ProcessLookupError: + job["status"], job["error"] = previous return False try: @@ -532,11 +877,13 @@ async def cancel_job(job_id: str) -> bool: process.kill() except ProcessLookupError: pass - logger.info("Onboarding job cancelled", job_id=job_id) + logger.info("Job cancelled", job_id=job_id) return True __all__ = [ + "KIND_FAIRSTORE_MIRROR", + "KIND_ONBOARDING", "MAX_LOG_LINES", "STATUS_CANCELLED", "STATUS_FAILED", @@ -545,6 +892,7 @@ async def cancel_job(job_id: str) -> bool: "TERMINAL_STATUSES", "JobError", "build_command", + "build_mirror_command", "cancel_job", "get_job", "list_jobs", @@ -552,4 +900,5 @@ async def cancel_job(job_id: str) -> bool: "runtime_warnings", "scripts_dir", "start_job", + "start_mirror_job", ] diff --git a/src/data_concierge/gateway/router.py b/src/data_concierge/gateway/router.py index ca8aac9..653bbee 100644 --- a/src/data_concierge/gateway/router.py +++ b/src/data_concierge/gateway/router.py @@ -820,6 +820,19 @@ class UpdateCkanSiteRequest(BaseModel): keywords: list[str] | None = None +def _refuse_managed_site(site_id: str) -> None: + """Another settings page owns this entry (the Fair Store's chat source).""" + site = ckan_sites_store.get_site(site_id) + if site and site.get("managed_by") == "fairstore": + raise HTTPException( + status_code=status.HTTP_409_CONFLICT, + detail=( + "This entry follows the Fair Store settings. Change it, or turn off " + "'Offer the Fair Store as a data source in chat', on the Fair Store page." + ), + ) + + @router.get("/admin/ckan-sites") async def list_ckan_sites( _admin: dict[str, Any] = Depends(require_admin), @@ -860,6 +873,7 @@ async def update_ckan_site( _admin: dict[str, Any] = Depends(require_admin), ) -> dict[str, Any]: """Update an existing CKAN site entry.""" + _refuse_managed_site(site_id) updates = request.model_dump(exclude_none=True) entry = ckan_sites_store.update_site(site_id, updates) if not entry: @@ -876,6 +890,7 @@ async def delete_ckan_site( _admin: dict[str, Any] = Depends(require_admin), ) -> dict[str, Any]: """Remove a CKAN site from the registry.""" + _refuse_managed_site(site_id) if not ckan_sites_store.remove_site(site_id): raise HTTPException( status_code=status.HTTP_404_NOT_FOUND, @@ -921,7 +936,7 @@ async def list_onboarding_jobs( """Recent onboarding runs, plus anything about this host that would break one.""" from data_concierge.gateway import onboarding_jobs - jobs = onboarding_jobs.list_jobs() + jobs = onboarding_jobs.list_jobs(kind=onboarding_jobs.KIND_ONBOARDING) return { "count": len(jobs), "jobs": jobs, @@ -944,6 +959,14 @@ async def start_onboarding_job( status_code=status.HTTP_404_NOT_FOUND, detail=f"Portal '{request.site_id}' is not registered", ) + if site.get("managed_by") == "fairstore": + raise HTTPException( + status_code=status.HTTP_409_CONFLICT, + detail=( + "The Fair Store is not onboarded itself: its qsv profiles come from " + "onboarding the portals it mirrors." + ), + ) options = request.model_dump(exclude={"site_id"}) try: @@ -990,6 +1013,195 @@ async def cancel_onboarding_job( return {"message": f"Job '{job_id}' cancelled"} +# ============================================================================= +# Fair Store (admin-only) — the CKAN that mirrors every registered portal +# ============================================================================= + + +class FairStoreSettingsRequest(BaseModel): + """Changes to the Fair Store link. Omitted fields keep their value. + + ``token`` follows the GitHub settings' "blank = keep" rule, because the UI + never sees the current token; ``clear_token`` removes it. + """ + + url: str | None = Field(default=None, max_length=400) + token: str | None = Field(default=None, max_length=4096) + clear_token: bool = False + chat_source: bool | None = Field( + default=None, description="Offer the Fair Store as a data source in chat" + ) + portal_name: str | None = Field(default=None, max_length=200) + portal_description: str | None = Field(default=None, max_length=2000) + quality_score: float | None = Field(default=None, ge=0.0, le=1.0) + mirror_sites: list[str] | None = Field( + default=None, description="Portals a mirror run of 'all' covers; empty = every portal" + ) + enrich: bool | None = Field(default=None, description="Socrata enrichment for DCAT portals") + + +class FairStoreSiteConfigRequest(BaseModel): + """New values for the Fair Store's runtime site options (``ckan.site_*``).""" + + options: dict[str, str | None] + + +class StartMirrorJobRequest(BaseModel): + """A Fair Store mirror run. The default is a dry run of every portal.""" + + site: str = Field(default="all", min_length=1, max_length=80) + apply: bool = Field(default=False, description="Write to the Fair Store (else dry run)") + enrich: bool | None = Field(default=None, description="Defaults to the saved setting") + limit: int | None = Field(default=None, ge=1, le=5000, description="Smoke test only") + + +@router.get("/admin/fairstore") +async def get_fairstore_settings( + _admin: dict[str, Any] = Depends(require_admin), +) -> dict[str, Any]: + """The Fair Store link's settings (the token only as set/unset).""" + from data_concierge.gateway import fairstore + + return {"settings": fairstore.public_settings()} + + +@router.post("/admin/fairstore") +async def update_fairstore_settings( + request: FairStoreSettingsRequest, + admin_user: dict[str, Any] = Depends(require_admin), +) -> dict[str, Any]: + """Save the Fair Store link's settings (and its chat-source registration).""" + from data_concierge.gateway import fairstore + + try: + saved = fairstore.save_settings( + request.model_dump(exclude_none=True), + updated_by=admin_user.get("user", "admin"), + ) + except ValueError as exc: + raise HTTPException(status_code=status.HTTP_400_BAD_REQUEST, detail=str(exc)) from exc + except fairstore.SettingsUnreadable as exc: + raise HTTPException( + status_code=status.HTTP_503_SERVICE_UNAVAILABLE, detail=str(exc) + ) from exc + return {"message": "Fair Store settings saved", "settings": saved} + + +@router.get("/admin/fairstore/status") +async def get_fairstore_status( + _admin: dict[str, Any] = Depends(require_admin), +) -> dict[str, Any]: + """Live health, token check and dataset counts per source portal.""" + from data_concierge.gateway import fairstore, onboarding_jobs + + report = await fairstore.status() + mirror_runs = onboarding_jobs.list_jobs(limit=1, kind=onboarding_jobs.KIND_FAIRSTORE_MIRROR) + report["last_mirror_run"] = mirror_runs[0] if mirror_runs else None + return report + + +@router.get("/admin/fairstore/site-config") +async def get_fairstore_site_config( + _admin: dict[str, Any] = Depends(require_admin), +) -> dict[str, Any]: + """The Fair Store's runtime site options, read live (needs the token).""" + from data_concierge.gateway import fairstore + + try: + return await fairstore.get_site_config() + except fairstore.FairStoreError as exc: + raise HTTPException(status_code=status.HTTP_502_BAD_GATEWAY, detail=str(exc)) from exc + + +@router.post("/admin/fairstore/site-config") +async def update_fairstore_site_config( + request: FairStoreSiteConfigRequest, + admin_user: dict[str, Any] = Depends(require_admin), +) -> dict[str, Any]: + """Change the Fair Store's site title, about text, logo or custom CSS.""" + from data_concierge.gateway import fairstore + + try: + config = await fairstore.update_site_config(request.options) + except ValueError as exc: + raise HTTPException(status_code=status.HTTP_400_BAD_REQUEST, detail=str(exc)) from exc + except fairstore.FairStoreError as exc: + raise HTTPException(status_code=status.HTTP_502_BAD_GATEWAY, detail=str(exc)) from exc + logger.info( + "Fair Store site options changed from the admin panel", + admin=admin_user.get("user"), + keys=sorted(request.options), + ) + return {"message": "Fair Store site settings saved", **config} + + +@router.get("/admin/fairstore/mirror-jobs") +async def list_fairstore_mirror_jobs( + _admin: dict[str, Any] = Depends(require_admin), +) -> dict[str, Any]: + """Recent mirror runs. Logs and cancel use ``/admin/onboarding-jobs/{id}``.""" + from data_concierge.gateway import onboarding_jobs + + jobs = onboarding_jobs.list_jobs(kind=onboarding_jobs.KIND_FAIRSTORE_MIRROR) + warnings = [ + w for w in onboarding_jobs.runtime_warnings() if "qsv" not in w and "OPENROUTER" not in w + ] + return { + "count": len(jobs), + "jobs": jobs, + "running_job_id": onboarding_jobs.running_job_id(), + "warnings": warnings, + } + + +@router.post("/admin/fairstore/mirror-jobs") +async def start_fairstore_mirror_job( + request: StartMirrorJobRequest, + admin_user: dict[str, Any] = Depends(require_admin), +) -> dict[str, Any]: + """Start a mirror run into the Fair Store (a dry run unless ``apply``).""" + from data_concierge.gateway import fairstore, onboarding_jobs + + current = fairstore.load_settings() + options = request.model_dump() + if options.get("enrich") is None: + options["enrich"] = current["enrich"] + clash = fairstore.conflicting_portal(current["url"]) + if clash: + raise HTTPException( + status_code=status.HTTP_409_CONFLICT, + detail=( + f"The Fair Store URL is the source portal '{clash}'; a mirror would " + "write into it. Point the Fair Store at its own CKAN first." + ), + ) + chosen = [s for s in current["mirror_sites"] if ckan_sites_store.get_site(s)] + if request.site == "all" and current["mirror_sites"] and not chosen: + # Widening to every portal would not be what the saved subset asked for. + raise HTTPException( + status_code=status.HTTP_409_CONFLICT, + detail=( + "None of the portals chosen under Mirror defaults is registered any " + "more. Choose portals again (or none, for every portal) and save." + ), + ) + if request.site == "all" and chosen: + # "All" means the portals the settings choose, when they choose some + # (minus any since removed from the registry). + options["sites"] = chosen + try: + job = await onboarding_jobs.start_mirror_job( + options, + target_url=current["url"], + token=current["token"], + started_by=admin_user.get("user", "admin"), + ) + except onboarding_jobs.JobError as exc: + raise HTTPException(status_code=status.HTTP_409_CONFLICT, detail=str(exc)) from exc + mode = "Mirror" if request.apply else "Dry run" + return {"message": f"{mode} started for {job['site_name']}", "job": job} + + # Public read-only endpoint — lets the UI populate a picker without admin auth. @router.get("/ckan-sites") async def list_ckan_sites_public() -> dict[str, Any]: @@ -4432,6 +4644,14 @@ def _validate_site(site: str) -> str: return site +@router.get("/fairstore/info") +async def fairstore_info() -> dict[str, Any]: + """Whether a Fair Store is linked, and its public address (no secrets).""" + from data_concierge.gateway import fairstore + + return fairstore.public_info() + + @router.get("/fairstore/search") async def fairstore_search(q: str, site: str = "wprdc", limit: int = 10) -> dict[str, Any]: """Search the Fair Store data dictionary by keyword. diff --git a/src/data_concierge/ui/static/js/admin.js b/src/data_concierge/ui/static/js/admin.js index 07a4280..fc19fa3 100644 --- a/src/data_concierge/ui/static/js/admin.js +++ b/src/data_concierge/ui/static/js/admin.js @@ -197,6 +197,19 @@ function setupEventListeners() { }; wireRevealToggle('ghTokenToggle', 'ghToken', 'token'); wireRevealToggle('ghWebhookSecretToggle', 'ghWebhookSecret', 'webhook secret'); + wireRevealToggle('fsTokenToggle', 'fsToken', 'token'); + + // Fair Store + document.getElementById('fairstore-tab').addEventListener('shown.bs.tab', loadFairStore); + document.getElementById('fsRefreshBtn').addEventListener('click', loadFairStore); + document.getElementById('fsSettingsForm').addEventListener('submit', saveFairStoreSettings); + document.getElementById('fsMirrorForm').addEventListener('submit', startFairStoreMirror); + document.getElementById('fsMirrorApply').addEventListener('change', syncFairStoreMirrorLabel); + document.getElementById('fsChatSource').addEventListener('change', syncFairStoreChatFields); + document.getElementById('fsLogCloseBtn').addEventListener('click', closeFairStoreLog); + document.getElementById('fsCancelBtn').addEventListener('click', cancelFairStoreRun); + document.getElementById('fsSiteForm').addEventListener('submit', saveFairStoreSiteConfig); + document.getElementById('fsSiteReloadBtn').addEventListener('click', loadFairStoreSiteConfig); document.getElementById('ghPauseBtn').addEventListener('click', openPauseModal); document.getElementById('ghPauseConfirmBtn').addEventListener('click', confirmPauseToggle); } @@ -2690,6 +2703,9 @@ function renderCkanSiteCard(site) { ? `${escapeHtml(site.organization)}` : ''; const isDefault = site.added_by === 'default'; + // The Fair Store's chat-source entry follows its settings page; onboarding + // it would re-download files the source portals already profile. + const isManaged = site.managed_by === 'fairstore'; // Portal type drives which tools work against this site, so it is shown // on the card rather than hidden behind an edit form. const portalType = (site.portal_type === 'dcat') ? 'dcat' : 'ckan'; @@ -2710,6 +2726,7 @@ function renderCkanSiteCard(site) { ${typeBadge} ${orgBadge} ${isDefault ? 'built-in' : ''} + ${isManaged ? 'managed on the Fair Store page' : ''}

@@ -3023,6 +3040,7 @@ async function pollOnboardingLog() { try { const response = await adminFetch(`${ONBOARD_API}/${encodeURIComponent(jobId)}?log_offset=${_onboardLogOffset}`); const { job } = await response.json(); + if (jobId !== _onboardLogJobId) return; // switched jobs while waiting const pre = document.getElementById('onboardLog'); if (job.log && job.log.length) { @@ -3064,8 +3082,486 @@ async function cancelOnboardingJob() { try { await adminFetch(`${ONBOARD_API}/${encodeURIComponent(_onboardLogJobId)}/cancel`, { method: 'POST' }); showToast('Run cancelled', 'success'); + stopOnboardingPoll(); // one poll chain only, or log lines repeat pollOnboardingLog(); } catch (error) { showToast(`Could not cancel: ${error.message}`, 'danger'); } } + + +// ============================================================================= +// Fair Store — the CKAN that mirrors every registered portal +// ============================================================================= + +const FAIRSTORE_API = `${API_BASE}/admin/fairstore`; + +let _fsSettings = null; +let _fsLogJobId = null; +let _fsLogOffset = 0; +let _fsPollTimer = null; + +function _fsAlert(targetId, type, msg) { + const box = document.getElementById(targetId); + if (!box) return; + box.innerHTML = msg + ? `` + : ''; +} + +async function loadFairStore() { + try { + const [settingsRes, sitesRes] = await Promise.all([ + adminFetch(FAIRSTORE_API), + adminFetch(CKAN_SITES_API), + ]); + _fsSettings = (await settingsRes.json()).settings; + const sites = ((await sitesRes.json()).sites || []).filter(site => site.id !== 'fairstore'); + renderFairStoreSettings(_fsSettings, sites); + } catch (error) { + _fsAlert('fsAlert', 'danger', `Could not load the Fair Store settings: ${error.message}`); + return; + } + loadFairStoreStatus(); + loadFairStoreJobs(); + if (_fsSettings.configured && _fsSettings.token_set) { + loadFairStoreSiteConfig(); + } else { + document.getElementById('fsSiteFields').textContent = + 'Set the URL and a sysadmin token to edit these.'; + document.getElementById('fsSiteActions').classList.add('d-none'); + } +} + +function renderFairStoreSettings(settings, sites) { + document.getElementById('fsUrl').value = settings.url || ''; + document.getElementById('fsToken').value = ''; + document.getElementById('fsClearToken').checked = false; + document.getElementById('fsClearTokenWrap').classList.toggle('d-none', settings.token_source !== 'admin'); + const tokenHint = document.getElementById('fsTokenHint'); + if (settings.token_set) { + const from = settings.token_source === 'environment' ? ' (from the FAIRSTORE_API_KEY environment variable)' : ''; + tokenHint.textContent = `A token ending ${settings.token_masked.replace(/•/g, '')} is set${from}. Leave blank to keep it.`; + } else { + tokenHint.textContent = 'No token is set: mirror runs can only be dry runs, and site settings cannot be edited.'; + } + document.getElementById('fsUrlHint').textContent = settings.url_source === 'environment' + ? 'From the FAIRSTORE_URL environment variable. Saving here overrides it.' + : 'The address people and Verikan use. Must be https:// unless it is on this machine.'; + + document.getElementById('fsChatSource').checked = !!settings.chat_source; + document.getElementById('fsPortalName').value = settings.portal_name || ''; + document.getElementById('fsPortalDescription').value = settings.portal_description || ''; + document.getElementById('fsEnrichDefault').checked = settings.enrich !== false; + document.getElementById('fsMirrorEnrich').checked = settings.enrich !== false; + syncFairStoreChatFields(); + + const chosen = new Set(settings.mirror_sites || []); + const missing = settings.mirror_sites_missing || []; + document.getElementById('fsMirrorSites').innerHTML = (sites.map(site => ` +
+ + +
`).join('') || 'No portals are registered.') + + (missing.length + ? `
No longer registered: + ${escapeHtml(missing.join(', '))}. Save to update the defaults.
` + : ''); + if (settings.conflicting_portal) { + _fsAlert('fsAlert', 'danger', + `The Fair Store URL is the registered portal "${settings.conflicting_portal}". ` + + 'Mirror runs and site-settings edits are refused until it points at its own CKAN.'); + } else { + _fsAlert('fsAlert', '', ''); + } + + const select = document.getElementById('fsMirrorSite'); + const current = select.value; + const allLabel = chosen.size + ? `All portals (${[...chosen].join(', ')})` + : (missing.length ? 'Saved portals (none registered — update Mirror defaults)' : 'All portals'); + select.innerHTML = `` + sites.map(site => + ``).join(''); + if ([...select.options].some(o => o.value === current)) select.value = current; + + const open = document.getElementById('fsOpenLink'); + if (settings.url) { + open.href = settings.url; + open.classList.remove('d-none'); + } else { + open.classList.add('d-none'); + } + const updated = document.getElementById('fsUpdatedHint'); + updated.textContent = settings.updated_at + ? `Last saved ${new Date(settings.updated_at).toLocaleString()}${settings.updated_by ? ` by ${settings.updated_by}` : ''}.` + : ''; + syncFairStoreMirrorLabel(); +} + +function syncFairStoreChatFields() { + document.getElementById('fsChatFields').classList.toggle('d-none', !document.getElementById('fsChatSource').checked); +} + +function syncFairStoreMirrorLabel() { + const apply = document.getElementById('fsMirrorApply').checked; + document.getElementById('fsMirrorStartLabel').textContent = apply ? 'Start mirror' : 'Start dry run'; + const btn = document.getElementById('fsMirrorStartBtn'); + btn.classList.toggle('btn-primary', !apply); + btn.classList.toggle('btn-warning', apply); +} + +async function saveFairStoreSettings(e) { + e.preventDefault(); + const payload = { + url: document.getElementById('fsUrl').value.trim(), + chat_source: document.getElementById('fsChatSource').checked, + portal_name: document.getElementById('fsPortalName').value.trim(), + portal_description: document.getElementById('fsPortalDescription').value.trim(), + mirror_sites: [...document.querySelectorAll('.fs-mirror-site:checked')].map(el => el.value), + enrich: document.getElementById('fsEnrichDefault').checked, + }; + const token = document.getElementById('fsToken').value.trim(); + if (token) payload.token = token; + if (document.getElementById('fsClearToken').checked) payload.clear_token = true; + + const btn = document.getElementById('fsSaveBtn'); + btn.disabled = true; + try { + const response = await adminFetch(FAIRSTORE_API, { + method: 'POST', + headers: { 'Content-Type': 'application/json' }, + body: JSON.stringify(payload), + }); + const data = await response.json(); + const hadToken = !!(_fsSettings && _fsSettings.token_set); + _fsSettings = data.settings; + showToast(data.message || 'Fair Store settings saved', 'success'); + const tokenDropped = hadToken && !token && !payload.clear_token && !_fsSettings.token_set; + // After the reload: renderFairStoreSettings resets the alert box. + await loadFairStore(); + if (tokenDropped && !_fsSettings.conflicting_portal) { + _fsAlert('fsAlert', 'warning', + 'The saved token was removed because the URL now points at a different server ' + + '(host or port). Enter a token for the new Fair Store.'); + } + } catch (error) { + showToast(`Could not save: ${error.message}`, 'danger'); + } finally { + btn.disabled = false; + } +} + +function _fsNumber(n) { + return (n === null || n === undefined) ? '—' : Number(n).toLocaleString(); +} + +async function loadFairStoreStatus() { + const box = document.getElementById('fsStatus'); + try { + const response = await adminFetch(`${FAIRSTORE_API}/status`); + const st = await response.json(); + if (!st.configured) { + box.innerHTML = `
+ No Fair Store is linked yet. Enter its URL under + Connection below.
`; + return; + } + if (!st.reachable) { + box.innerHTML = `
+ Unreachable: ${escapeHtml(st.url)} +
${escapeHtml(st.error || '')}
`; + return; + } + const tokenBadge = { + sysadmin: 'Sysadmin token OK', + rejected: 'Token rejected', + missing: 'No token (read-only)', + }[st.token] || ''; + const byPortal = (st.by_portal || []).map(p => ` + ${escapeHtml(p.name)}${escapeHtml(p.site_id)} + ${_fsNumber(p.datasets)}`).join(''); + const last = st.last_mirror_run; + const lastLine = last + ? `Last mirror run: ${onboardStatusBadge(last.status)} ${escapeHtml(last.site_name || '')} + ${last.options && last.options.apply ? '' : 'dry run'} + ${last.started_at ? escapeHtml(new Date(last.started_at).toLocaleString()) : ''}` + : 'No mirror run from this panel yet.'; + box.innerHTML = ` +
+ ${st.token === 'rejected' ? `
${escapeHtml(st.token_error || 'The Fair Store refused the saved token.')}
` : ''} +
+
Datasets
${_fsNumber(st.datasets)}
+
Organizations
${_fsNumber(st.organizations)}
+
Groups
${_fsNumber(st.groups)}
+
+ ${byPortal ? `
+ + ${byPortal}
Mirrored fromPortal IDDatasets
` : ''} +
${lastLine}
`; + } catch (error) { + box.innerHTML = `
Could not check the Fair Store: ${escapeHtml(error.message)}
`; + } +} + +async function startFairStoreMirror(e) { + e.preventDefault(); + const apply = document.getElementById('fsMirrorApply').checked; + const siteSelect = document.getElementById('fsMirrorSite'); + const target = siteSelect.options[siteSelect.selectedIndex]?.text || 'All portals'; + if (apply && !confirm(`Write ${target} into the Fair Store?\n\nOnly records that changed at the source are written. ` + + 'Records the source withdrew are reported, not deleted.')) { + return; + } + const limitRaw = document.getElementById('fsMirrorLimit').value; + const payload = { + site: siteSelect.value || 'all', + apply, + enrich: document.getElementById('fsMirrorEnrich').checked, + limit: limitRaw ? Number(limitRaw) : null, + }; + const btn = document.getElementById('fsMirrorStartBtn'); + btn.disabled = true; + try { + const response = await adminFetch(`${FAIRSTORE_API}/mirror-jobs`, { + method: 'POST', + headers: { 'Content-Type': 'application/json' }, + body: JSON.stringify(payload), + }); + const data = await response.json(); + showToast(data.message || 'Mirror run started', 'success'); + await loadFairStoreJobs(); + if (data.job?.id) watchFairStoreJob(data.job.id); + } catch (error) { + showToast(`Could not start the run: ${error.message}`, 'danger'); + } finally { + btn.disabled = false; + } +} + +async function loadFairStoreJobs() { + const container = document.getElementById('fsJobList'); + try { + const response = await adminFetch(`${FAIRSTORE_API}/mirror-jobs`); + const data = await response.json(); + document.getElementById('fsMirrorWarnings').innerHTML = (data.warnings || []).map(w => + `
${escapeHtml(w)}
`).join(''); + if (!data.jobs || data.jobs.length === 0) { + container.innerHTML = '
No mirror runs yet.
'; + return; + } + container.innerHTML = ` +
+ + + + ${data.jobs.map(job => ` + + + + + + + + + `).join('')} + +
PortalsModeStatusStartedDurationStarted by
${escapeHtml(job.site_name || job.site_id || '')}${job.options && job.options.apply + ? 'write' + : 'dry run'}${onboardStatusBadge(job.status)} + ${job.error ? `
${escapeHtml(job.error)}
` : ''}
${job.started_at ? escapeHtml(new Date(job.started_at).toLocaleString()) : ''}${escapeHtml(_onboardDuration(job))}${escapeHtml(job.started_by || '')} + +
+
`; + bindClick(container, '.js-fs-log', (el) => watchFairStoreJob(el.dataset.jobId)); + } catch (error) { + container.innerHTML = `
Failed to load mirror runs: ${escapeHtml(error.message)}
`; + } +} + +function stopFairStorePoll() { + if (_fsPollTimer) { + clearTimeout(_fsPollTimer); + _fsPollTimer = null; + } +} + +function closeFairStoreLog() { + stopFairStorePoll(); + _fsLogJobId = null; + document.getElementById('fsLogPanel').classList.add('d-none'); +} + +function watchFairStoreJob(jobId) { + if (_fsLogJobId !== jobId) { + _fsLogOffset = 0; + document.getElementById('fsLog').textContent = ''; + document.getElementById('fsResult').innerHTML = ''; + } + _fsLogJobId = jobId; + document.getElementById('fsLogPanel').classList.remove('d-none'); + document.getElementById('fsLogTitle').textContent = `Mirror run ${jobId}`; + stopFairStorePoll(); + pollFairStoreLog(); +} + +function renderFairStoreResult(result) { + const box = document.getElementById('fsResult'); + if (!result) { + box.innerHTML = ''; + return; + } + if (!Array.isArray(result) && result.error) { + box.innerHTML = `
${escapeHtml(result.error)}
`; + return; + } + const rows = (Array.isArray(result) ? result : [result]).map(r => { + if (r.error) { + return `${escapeHtml(r.site)}${escapeHtml(r.error)}`; + } + if (r.skipped) { + return `${escapeHtml(r.site)}Skipped: ${escapeHtml(r.skipped)}`; + } + const q = r.qsv || {}; + const withdrawn = r.withdrawn_at_source ? (r.withdrawn_at_source.datasets || 0) : null; + return ` + ${escapeHtml(r.site)} + ${_fsNumber(r.datasets)} + ${_fsNumber(r.created)} + ${_fsNumber(r.updated)} + ${_fsNumber(r.unchanged)} + ${_fsNumber(q.created)} / ${_fsNumber(q.refreshed)} / ${_fsNumber(q.unchanged)} + ${withdrawn === null ? '—' : _fsNumber(withdrawn)} + `; + }).join(''); + const mode = (Array.isArray(result) ? result[0] : result)?.mode === 'apply' ? 'Written' : 'Planned (dry run)'; + box.innerHTML = ` +
${escapeHtml(mode)}
+
+ + + + ${rows}
PortalDatasetsCreatedUpdatedUnchangedqsv new / refreshed / sameWithdrawn at source
`; +} + +async function pollFairStoreLog() { + const jobId = _fsLogJobId; + if (!jobId) return; + try { + const response = await adminFetch(`${ONBOARD_API}/${encodeURIComponent(jobId)}?log_offset=${_fsLogOffset}`); + const { job } = await response.json(); + if (jobId !== _fsLogJobId) return; // switched jobs while waiting + + const pre = document.getElementById('fsLog'); + if (job.log && job.log.length) { + const atBottom = pre.scrollHeight - pre.scrollTop - pre.clientHeight < 40; + pre.textContent += job.log.join('\n') + '\n'; + if (atBottom) pre.scrollTop = pre.scrollHeight; + } + _fsLogOffset = job.log_next_offset ?? _fsLogOffset; + document.getElementById('fsLogStatus').innerHTML = onboardStatusBadge(job.status) + + (job.options && !job.options.apply ? ' dry run' : ''); + const isRunning = job.status === 'running'; + document.getElementById('fsCancelBtn').classList.toggle('d-none', !isRunning); + renderFairStoreResult(job.result); + + if (isRunning) { + _fsPollTimer = setTimeout(pollFairStoreLog, 2000); + } else { + stopFairStorePoll(); + loadFairStoreJobs(); + loadFairStoreStatus(); + } + } catch (error) { + stopFairStorePoll(); + document.getElementById('fsLogStatus').innerHTML = + `Lost contact: ${escapeHtml(error.message)}`; + } +} + +async function cancelFairStoreRun() { + if (!_fsLogJobId) return; + if (!confirm('Cancel this mirror run?\n\nWhat it already wrote stays; a rerun picks up the rest.')) return; + try { + await adminFetch(`${ONBOARD_API}/${encodeURIComponent(_fsLogJobId)}/cancel`, { method: 'POST' }); + showToast('Run cancelled', 'success'); + stopFairStorePoll(); // one poll chain only, or log lines repeat + pollFairStoreLog(); + } catch (error) { + showToast(`Could not cancel: ${error.message}`, 'danger'); + } +} + +async function loadFairStoreSiteConfig() { + const fields = document.getElementById('fsSiteFields'); + const actions = document.getElementById('fsSiteActions'); + fields.innerHTML = 'Reading from the Fair Store…'; + try { + const response = await adminFetch(`${FAIRSTORE_API}/site-config`); + renderFairStoreSiteConfig(await response.json()); + _fsAlert('fsSiteAlert', '', ''); + } catch (error) { + fields.innerHTML = ''; + actions.classList.add('d-none'); + _fsAlert('fsSiteAlert', 'warning', `Could not read the site settings: ${error.message}`); + } +} + +function renderFairStoreSiteConfig(config) { + const fields = document.getElementById('fsSiteFields'); + const options = config.options || []; + fields.classList.remove('text-muted', 'small'); + fields.innerHTML = options.map(opt => { + const id = `fsOpt-${opt.key.replace(/[^a-z0-9]/gi, '-')}`; + const value = escapeHtml(opt.value || ''); + const input = opt.kind === 'text' + ? `` + : ``; + return `
+ + ${input} +
`; + }).join('') || '
The Fair Store exposes no editable site settings.
'; + fields.querySelectorAll('.fs-site-opt').forEach(el => { el.dataset.original = el.value; }); + document.getElementById('fsSiteActions').classList.toggle('d-none', options.length === 0); +} + +async function saveFairStoreSiteConfig(e) { + e.preventDefault(); + const changed = {}; + document.querySelectorAll('.fs-site-opt').forEach(el => { + if (el.value !== el.dataset.original) changed[el.dataset.key] = el.value; + }); + if (Object.keys(changed).length === 0) { + showToast('Nothing changed', 'info'); + return; + } + const btn = document.getElementById('fsSiteSaveBtn'); + btn.disabled = true; + try { + const response = await adminFetch(`${FAIRSTORE_API}/site-config`, { + method: 'POST', + headers: { 'Content-Type': 'application/json' }, + body: JSON.stringify({ options: changed }), + }); + const data = await response.json(); + renderFairStoreSiteConfig(data); + showToast(data.message || 'Site settings saved', 'success'); + loadFairStoreStatus(); + } catch (error) { + _fsAlert('fsSiteAlert', 'danger', `Could not save: ${error.message}`); + } finally { + btn.disabled = false; + } +} diff --git a/src/data_concierge/ui/static/js/app.js b/src/data_concierge/ui/static/js/app.js index 99ed9c3..6fc662b 100644 --- a/src/data_concierge/ui/static/js/app.js +++ b/src/data_concierge/ui/static/js/app.js @@ -534,8 +534,11 @@ function updateAuthUI() { logoutBtn.classList.add('d-none'); authStatus.classList.add('d-none'); } - // Admin shortcut in the navbar — only for signed-in admins. - if (adminBtn) adminBtn.classList.toggle('d-none', !(isAuthenticated && _isAdmin)); + // Admin shortcuts — the top bar button and the sidebar link — only for + // signed-in admins. + const showAdmin = isAuthenticated && _isAdmin; + if (adminBtn) adminBtn.classList.toggle('d-none', !showAdmin); + document.getElementById('adminTopBtn')?.classList.toggle('d-none', !showAdmin); } function showLoginModal(pendingQuery = null) { diff --git a/src/data_concierge/ui/static/js/dictionary.js b/src/data_concierge/ui/static/js/dictionary.js index f486bae..20b80dc 100644 --- a/src/data_concierge/ui/static/js/dictionary.js +++ b/src/data_concierge/ui/static/js/dictionary.js @@ -144,6 +144,7 @@ async function openResource(resourceId, title) { ${d.description ? `

${escapeHtml(d.description)}

` : ''} + ${fairstoreDatasetLink(d.dataset_id)}
@@ -155,6 +156,19 @@ async function openResource(resourceId, title) { `; } +// The Fair Store mirrors these datasets under their source names, with the +// qsv dictionary, stats and frequency tables attached as resources. +function fairstoreDatasetLink(datasetId) { + const base = document.body.dataset.fairstoreUrl; + if (!base || !datasetId) return ''; + const href = `${base}/dataset/${encodeURIComponent(datasetId)}`; + return ``; +} + function backToList() { document.getElementById('dictDetailView').classList.add('d-none'); document.getElementById('dictListView').classList.remove('d-none'); diff --git a/src/data_concierge/ui/templates/admin.html b/src/data_concierge/ui/templates/admin.html index 7946acc..39fb889 100644 --- a/src/data_concierge/ui/templates/admin.html +++ b/src/data_concierge/ui/templates/admin.html @@ -103,6 +103,10 @@

Admin Dashboard

data-bs-target="#ckansites" type="button" role="tab" aria-controls="ckansites"> Data Portals +
Insights
+ + + +
+ + +
+
+ + Checking the Fair Store… +
+
+ +
+ + +
+
+

Mirror Runs

+

+ Copy the registered portals’ catalogs into the Fair Store. + Reruns only write what changed at the source, and records a portal + withdrew are reported, never deleted. Start with a dry run: it + reads everything and reports the plan without writing. +

+
+
+ +
+ + +
+ + +
+
+ + +
+
+
+ + +
+
+ + +
+
+
+ +
+
+ “Max datasets” is for smoke tests: a limited run skips the + withdrawn-at-source check. Socrata enrichment adds each data.pa.gov + dataset’s owning agency and column list. +
+ + +
+
+ + Loading mirror runs… +
+
+ +
+
+
+ Mirror run + +
+
+ + +
+
+
+

+                    
+ +
+ + +
+
+

Connection

+

+ Where the Fair Store is, the sysadmin API token Verikan uses to write + to it, and whether chat users can search it. +

+
+
+ + +
+ + +
+ The address people and Verikan use. Must be https:// unless it is on this machine. +
+
+
+ +
+ + +
+
Leave blank to keep the existing token.
+
+ + +
+
+ Create one on the Fair Store with + ckan user token add <sysadmin> verikan. Changing the URL to + another host drops the saved token, so it is never sent to a different server. +
+
+ +
+ + +
+
+
+ + +
+
+ + +
+
+ The agent searches the Fair Store’s catalog and loads rows from the + portal each dataset was mirrored from. +
+
+ +

Mirror defaults

+
Portals an “All portals” run covers (none checked = every registered portal):
+
+
+ + +
+ +
+ +
+
+ + +
+ + +
+
+

Fair Store Site Settings

+

+ The Fair Store’s own title, description, home page and about text, + logo and custom CSS — read from and saved to the Fair Store live. +

+
+
+ +
+
+
+ +
Set the URL and a sysadmin token to edit these.
+
+ +
+ + + +
@@ -1658,7 +1881,7 @@ - + diff --git a/src/data_concierge/ui/templates/dictionary.html b/src/data_concierge/ui/templates/dictionary.html index 0a87b61..a7305b1 100644 --- a/src/data_concierge/ui/templates/dictionary.html +++ b/src/data_concierge/ui/templates/dictionary.html @@ -42,7 +42,7 @@ })(); - +