diff --git a/nim-skills/proteinmpnn-nim/SKILL.md b/nim-skills/proteinmpnn-nim/SKILL.md
index 900b559..6d6568c 100644
--- a/nim-skills/proteinmpnn-nim/SKILL.md
+++ b/nim-skills/proteinmpnn-nim/SKILL.md
@@ -1,15 +1,20 @@
---
name: proteinmpnn-nim
description: >
- Run ProteinMPNN inverse folding via NVIDIA NIM to design protein sequences for a target backbone. Use for ProteinMPNN, inverse folding, sequence design, backbone redesign, fixed chains/residues, omit_AAs, sampling temperature, soluble model, hosted NVIDIA API, local Docker, PDB input, and multi-FASTA output.
+ Run ProteinMPNN inverse folding via NVIDIA NIM to design protein sequences for a target backbone. Sends user-provided PDB files and design parameters to NVIDIA's hosted API, authenticated with an environment API key, or to a user-selected local NIM. Use for sequence design, backbone redesign, fixed chains and residues, omit_AAs, sampling temperature, soluble model, local Docker, and multi-FASTA output.
license: Apache-2.0 AND CC-BY-4.0
-compatibility: "requests>=2.28"
+compatibility: "Python >=3.10; requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
+permissions:
+ - network
+ - env
---
# ProteinMPNN NIM
-Design protein sequences for a supplied backbone PDB. Use this `SKILL.md` for
+
+
+Design protein sequences for a supplied backbone PDB. Use this guide for
first-pass hosted/local usage; load supplemental files only when needed:
- `references/api.md`: exact endpoints, schemas, Docker flags, response fields.
@@ -20,7 +25,7 @@ first-pass hosted/local usage; load supplemental files only when needed:
## Choose Mode
-Honor an explicitly configured runtime before asking. `NIM_API_MODE=local` selects
+Honor the user's explicit mode; otherwise use the configured runtime. `NIM_API_MODE=local` selects
the local service at `PROTEINMPNN_NIM_URL`; the URL defaults to
`http://localhost:8000` for a NIM running in the same host or container. Ask only
when neither the environment nor the user's request makes the mode clear:
@@ -28,7 +33,8 @@ when neither the environment nor the user's request makes the mode clear:
> Hosted NVIDIA API or local Docker NIM?
- Hosted: `https://health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict`
-- Local: `${PROTEINMPNN_NIM_URL:-http://localhost:8000}/biology/ipd/proteinmpnn/predict`
+- Local: append `/biology/ipd/proteinmpnn/predict` to `PROTEINMPNN_NIM_URL`
+ (default base URL: `http://localhost:8000`).
Local inference paths do not include `/v1/`. Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Supported local Docker
startup uses `NGC_API_KEY` (or `NVIDIA_API_KEY` via the preflight) for
@@ -37,6 +43,25 @@ into the container with `-e NGC_API_KEY`. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.
+## Data Transfer and Authorization
+
+Before a hosted request, tell the user that the **entire PDB file and design
+parameters will be uploaded to NVIDIA's hosted API** at the endpoint above.
+Proceed if the user has explicitly requested hosted processing of that PDB or
+already approved the transfer; otherwise ask for confirmation before submitting.
+For confidential structures, recommend a local NIM in the user's approved
+environment. A configured local URL may point to another machine; use only the
+configured or user-selected destination. Do not switch from local to hosted processing without
+the user's authorization.
+
+The client reads `NIM_API_MODE`, `PROTEINMPNN_NIM_URL`, and, for hosted mode,
+`NGC_API_KEY` from the environment. It sends the key only in the HTTPS
+Authorization header to the hosted endpoint; local inference sends no key.
+Keep credentials out of logs and saved artifacts. The output directory contains
+the full input PDB in `request.json` and the returned sequences and scores, so
+use a location appropriate for the input's sensitivity. See
+[`references/api.md`](references/api.md) for endpoint and data-handling details.
+
## Local Docker
For local setup, run the full sequence — env preflight, `docker login`,
@@ -56,53 +81,68 @@ proteinmpnn_nim_url="${PROTEINMPNN_NIM_URL:-http://localhost:8000}"
until curl -sf "${proteinmpnn_nim_url%/}/v1/health/ready"; do sleep 5; done
```
-## Request Pattern
-
-Read PDB content inline; do not send only a file path.
-
-```python
-import os
-from pathlib import Path
-import requests
-
-HOSTED = os.getenv("NIM_API_MODE", "hosted").strip().lower() != "local"
-pdb_content = Path("1R42.pdb").read_text()
-nim_url = os.getenv("PROTEINMPNN_NIM_URL", "http://localhost:8000").rstrip("/")
-url = (
- "https://health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict"
- if HOSTED else f"{nim_url}/biology/ipd/proteinmpnn/predict"
-)
-headers = {"Content-Type": "application/json"}
-if HOSTED:
- headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"
-
-payload = {
- "input_pdb": pdb_content,
- "num_seq_per_target": 10,
- "sampling_temp": [0.1],
- "use_soluble_model": False,
- "ca_only": False,
-}
-response = requests.post(url, headers=headers, json=payload, timeout=300)
-response.raise_for_status()
-result = response.json()
-```
+## Instructions
+
+For a request to execute a design, run [`scripts/design.py`](scripts/design.py)
+and inspect its results. Writing a request script alone does not complete an
+execution request. If the user asks only for code or setup instructions, provide
+those without submitting an inference request.
+
+1. Use the user's PDB path and requested sequence count. The client reads the
+ entire PDB into `input_pdb`; do not replace or truncate the supplied backbone.
+2. Select `--mode hosted` or `--mode local` and follow **Data Transfer and
+ Authorization** above before submitting. Hosted mode uploads the PDB to the
+ documented NVIDIA endpoint and requires `NGC_API_KEY` in the environment.
+ Check only whether the key is set; do not print it, dump the environment, or
+ save authentication headers. Local inference sends no authorization header.
+3. Choose a new `--output-dir` for each request. The client reserves it before
+ submitting, preserves the raw response for diagnostics, and validates the
+ designed sequence count and score alignment before reporting completion.
+4. Read `summary.json` and report the actual results described below. If the
+ request or validation fails, report the failure and diagnostic path; do not
+ substitute example sequences or repeatedly resubmit the same request.
-Common controls:
+## Examples
+
+Run from this skill's directory, or use an absolute path to `scripts/design.py`.
+Substitute the user's input path and a new output directory:
+
+```bash
+python scripts/design.py --mode hosted \
+ --pdb /path/to/backbone.pdb --num-sequences 10 \
+ --temperature 0.1 --output-dir /path/to/new-design-run
+```
-- Redesign only chain A: `"input_pdb_chains": ["A"]`.
-- Exclude amino acids: `"omit_AAs": ["C"]` or `"omit_AAs": ["M"]`.
-- Diversity: `"sampling_temp": [0.1, 0.3, 0.5]` (always a list).
-- Solubility bias: `"use_soluble_model": True`.
-- Candidate count: `num_seq_per_target` is 1-100.
+For a running local NIM, use `--mode local`; the client honors
+`PROTEINMPNN_NIM_URL`. To design only chain A, exclude cysteine, or request the
+soluble model, add `--chains A`, `--omit-aas C`, or `--soluble` respectively.
+`--seed` sets `random_seed`; `--ca-only` selects the CA-only model. The helper
+uses one temperature per request; run separate output directories for a
+temperature sweep. For advanced JSONL controls or a custom batch request,
+use [`references/api.md`](references/api.md) and the post-response example in
+[`references/examples.md`](references/examples.md).
## Save And Report Output
-Save the returned `mfasta` and pair scores only with designed (non-native/WT)
-rows, using the snippet in [`references/examples.md`](references/examples.md)
-under **Save Multi-FASTA**. Validate promising designs by predicting structures
-with Boltz2 or OpenFold3 and comparing them to the target backbone. For
-FASTA/score sanity checks, read `references/validation.md`.
+The client writes `request.json`, `response.raw`, `response.json`,
+`designed_sequences.fa`, and `summary.json` into the requested output directory.
+The FASTA preserves the complete returned `mfasta`, including a native/WT entry
+when present. The summary contains only designed sequences, each paired with
+its actual score, and records whether scores came from the JSON array or FASTA
+headers. It is also printed after the artifacts are saved and checked.
+
+In the final response, report:
+
+- The number of **designed** sequences, excluding the native/WT reference.
+- Each design's identifier and actual returned score, plus its sequence (for
+ long sequences, give a clearly labelled preview and link to the full FASTA).
+- The saved FASTA and summary paths, and the raw response path for provenance.
+- That these are inverse-folding candidates, with no fold-back validation
+ performed unless it was actually requested and run.
+
+Do not treat a score as proof that a sequence folds or binds. Further validation
+with Boltz2 or OpenFold3 is an optional next step. For FASTA/score sanity checks,
+read [`references/validation.md`](references/validation.md).
## Limits And Troubleshooting
diff --git a/nim-skills/proteinmpnn-nim/config/skillspector-baseline.yml b/nim-skills/proteinmpnn-nim/config/skillspector-baseline.yml
index bae7f22..8affdd5 100644
--- a/nim-skills/proteinmpnn-nim/config/skillspector-baseline.yml
+++ b/nim-skills/proteinmpnn-nim/config/skillspector-baseline.yml
@@ -3,7 +3,7 @@
# Audited false-positive suppression, auto-applied by the NVSkills Tier 1
# runner via config/skillspector-baseline.yml. Suppressed findings remain in
# the report JSON marked `suppressed: true` with the reason below.
-version: 1
+version: 2
rules:
- id: "PE3"
@@ -16,3 +16,15 @@ rules:
`docker login nvcr.io`. This is first-party, user-facing setup guidance for
the user's own credential file — not credential theft. No SSH keys, cloud
credential stores, or third-party secret files are accessed.
+ - id: "PE3"
+ path: "evals/evals.json"
+ # Match only the literal dotenv token, including PE3's trailing space.
+ message: ".env "
+ reason: >-
+ Reviewed 2026-09-30. The two matches are in deferred Docker evaluation
+ 3's expected output and setup assertion. They describe optional loading
+ of the user's own repo-root dotenv file for NGC authentication, the same
+ setup documented in references/api.md. This JSON is evaluation text and
+ does not read or transmit credentials. The active hosted evaluation and
+ design client read NGC_API_KEY from the environment. This exception is
+ limited to the literal dotenv matches in this evaluation file.
diff --git a/nim-skills/proteinmpnn-nim/evals/README.md b/nim-skills/proteinmpnn-nim/evals/README.md
new file mode 100644
index 0000000..77cd589
--- /dev/null
+++ b/nim-skills/proteinmpnn-nim/evals/README.md
@@ -0,0 +1,37 @@
+# ProteinMPNN evaluations
+
+The default `config.yml` selects the hosted request in `evals.json`. It uses
+`NGC_API_KEY`, stages `files/1R42.pdb`, and requires an executed inference request,
+saved designed sequences, and response-derived scores.
+
+## Local GPU task
+
+`harbor/proteinmpnn-local-design` is a separate native Harbor task with a custom
+grader. It requires the ProteinMPNN NIM image, a GPU, its model download
+credentials, and a running NIM server in the task container. A shared Astra
+sandbox using the generic agent template ignores the task image and cannot
+satisfy its loopback health check.
+
+For a dedicated local GPU evaluation, use a separate checkout and replace
+`evals/config.yml` with:
+
+```yaml
+schema_version: 1
+harbor:
+ task_source: native_harbor
+ custom_dockerfile_mode: preserve
+ base_image_mode: disabled
+ n_attempts: 1
+ pass_threshold: 0.8
+ stop_on_pass: false
+ n_concurrent: 1
+ runtime_env:
+ - NGC_API_KEY
+grading:
+ mode: custom_only
+```
+
+Run on a GPU-capable Docker host with `--env-mode docker`, or a dedicated sandbox
+template configured with the NIM image, GPU, credentials, and server startup.
+The native task's readiness check must pass before an agent starts. Keep the
+hosted configuration as the shared CI default.
diff --git a/nim-skills/proteinmpnn-nim/evals/config.yml b/nim-skills/proteinmpnn-nim/evals/config.yml
index dac0144..6ed0cda 100644
--- a/nim-skills/proteinmpnn-nim/evals/config.yml
+++ b/nim-skills/proteinmpnn-nim/evals/config.yml
@@ -1,13 +1,11 @@
schema_version: 1
-# ProteinMPNN is the repository's local-GPU BYOT example. ACES supports one
-# task source per skill, so this selects the native task below. The hosted
-# evals.json dataset remains available if this is switched back to evals_json
-# with aces_default grading.
+# The shared Astra sandbox uses a fixed agent image and does not launch the
+# native task's GPU NIM image. Use the hosted execution dataset for CI.
+# The native local-GPU task is retained under harbor/; see README.md to run it
+# in a runtime configured for that image and GPU.
harbor:
- task_source: native_harbor
- custom_dockerfile_mode: preserve
- base_image_mode: disabled
+ task_source: evals_json
n_attempts: 1
pass_threshold: 0.8
stop_on_pass: false
@@ -16,4 +14,4 @@ harbor:
- NGC_API_KEY
grading:
- mode: custom_only
+ mode: aces_default
diff --git a/nim-skills/proteinmpnn-nim/evals/evals.json b/nim-skills/proteinmpnn-nim/evals/evals.json
index a0239ec..924fb07 100644
--- a/nim-skills/proteinmpnn-nim/evals/evals.json
+++ b/nim-skills/proteinmpnn-nim/evals/evals.json
@@ -9,41 +9,13 @@
"evals/files/1R42.pdb"
],
"assertions": [
- {
- "id": "hosted-request-executed",
- "description": "Executes the hosted request instead of only writing code",
- "check": "Trajectory shows successful execution of the hosted request, and the final response reports actual response-derived designed sequences, scores, and the saved FASTA path"
- },
- {
- "id": "hosted-endpoint-url",
- "description": "Uses the correct hosted ProteinMPNN endpoint URL",
- "check": "Script contains 'health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict'"
- },
- {
- "id": "bearer-auth-header",
- "description": "Sets Authorization header with Bearer token from NGC_API_KEY",
- "check": "Script contains 'Authorization' and 'Bearer' and 'NGC_API_KEY'"
- },
- {
- "id": "input-pdb-field",
- "description": "Sends PDB file content as 'input_pdb' field (inline string, not file path)",
- "check": "Script reads 1R42.pdb and passes its content to 'input_pdb' field in the payload"
- },
- {
- "id": "num-seq-per-target",
- "description": "num_seq_per_target is set to 10",
- "check": "Script contains 'num_seq_per_target' and '10'"
- },
- {
- "id": "saves-mfasta-output",
- "description": "Saves the mfasta string from the response to a .fa file",
- "check": "Script accesses 'mfasta' from response and writes it to a file with .fa or .fasta extension"
- },
- {
- "id": "scores-reported",
- "description": "Reports scores only for designed sequences",
- "check": "Script accounts for the native/WT FASTA row when present and pairs scores with designed sequences, not the WT row"
- }
+ "[hosted-request-executed] Executes the hosted request instead of only writing code: Trajectory shows successful execution of the hosted request, and the final response reports actual response-derived designed sequences, scores, and the saved FASTA path",
+ "[hosted-endpoint-url] Uses the correct hosted ProteinMPNN endpoint URL: Script contains 'health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict'",
+ "[bearer-auth-header] Sets Authorization header with Bearer token from NGC_API_KEY: Script contains 'Authorization' and 'Bearer' and 'NGC_API_KEY'",
+ "[input-pdb-field] Sends PDB file content as 'input_pdb' field (inline string, not file path): Script reads 1R42.pdb and passes its content to 'input_pdb' field in the payload",
+ "[num-seq-per-target] num_seq_per_target is set to 10: Script contains 'num_seq_per_target' and '10'",
+ "[saves-mfasta-output] Saves the mfasta string from the response to a .fa file: Script accesses 'mfasta' from response and writes it to a file with .fa or .fasta extension",
+ "[scores-reported] Reports scores only for designed sequences: Script accounts for the native/WT FASTA row when present and pairs scores with designed sequences, not the WT row"
]
}
],
@@ -54,36 +26,12 @@
"expected_output": "A Python script targeting the hosted endpoint with input_pdb_chains=['A'], num_seq_per_target=5, sampling_temp=[0.2], and omit_AAs=['C'], reading structure.pdb and saving the designed sequences.",
"files": [],
"assertions": [
- {
- "id": "hosted-endpoint-url",
- "description": "Uses the correct hosted endpoint URL",
- "check": "Script contains 'health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict'"
- },
- {
- "id": "input-pdb-chains",
- "description": "Uses input_pdb_chains to limit design to chain A",
- "check": "Script contains 'input_pdb_chains' and 'A'"
- },
- {
- "id": "sampling-temp-array",
- "description": "sampling_temp is passed as an array (even for a single value)",
- "check": "Script contains 'sampling_temp' with value [0.2] or as a list containing 0.2"
- },
- {
- "id": "omit-cysteines",
- "description": "omit_AAs excludes cysteine ('C')",
- "check": "Script contains 'omit_AAs' and 'C'"
- },
- {
- "id": "num-seq-5",
- "description": "num_seq_per_target is set to 5",
- "check": "Script contains 'num_seq_per_target' and '5'"
- },
- {
- "id": "saves-output",
- "description": "Saves designed sequences to a FASTA file",
- "check": "Script writes the mfasta response to a file"
- }
+ "[hosted-endpoint-url] Uses the correct hosted endpoint URL: Script contains 'health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict'",
+ "[input-pdb-chains] Uses input_pdb_chains to limit design to chain A: Script contains 'input_pdb_chains' and 'A'",
+ "[sampling-temp-array] sampling_temp is passed as an array (even for a single value): Script contains 'sampling_temp' with value [0.2] or as a list containing 0.2",
+ "[omit-cysteines] omit_AAs excludes cysteine ('C'): Script contains 'omit_AAs' and 'C'",
+ "[num-seq-5] num_seq_per_target is set to 5: Script contains 'num_seq_per_target' and '5'",
+ "[saves-output] Saves designed sequences to a FASTA file: Script writes the mfasta response to a file"
]
},
{
@@ -93,36 +41,12 @@
"expected_output": "Docker setup commands using shell env first and optional repo-root .env overrides, requiring NGC_API_KEY or NVIDIA_API_KEY fallback plus LOCAL_NIM_CACHE, using the correct image tag, single GPU flag, and the unique cache mount path /home/nvs/.cache/nim (not /opt/nim/.cache), then a health check and no-auth request script to localhost:8000.",
"files": [],
"assertions": [
- {
- "id": "docker-image-tag",
- "description": "References the correct ProteinMPNN container image",
- "check": "Output contains 'nvcr.io/nim/ipd/proteinmpnn'"
- },
- {
- "id": "env-contract-and-cache",
- "description": "Local setup uses the repo env contract and LOCAL_NIM_CACHE",
- "check": "Output sources repo-root .env only if present, supports NVIDIA_API_KEY fallback to NGC_API_KEY, requires LOCAL_NIM_CACHE, and mounts LOCAL_NIM_CACHE to /home/nvs/.cache/nim"
- },
- {
- "id": "cache-mount-path",
- "description": "Volume mount uses /home/nvs/.cache/nim (not /opt/nim/.cache)",
- "check": "Output contains '/home/nvs/.cache/nim' as the container-side cache path in the -v mount"
- },
- {
- "id": "health-check",
- "description": "Includes health check before sending request",
- "check": "Output contains health check against localhost:8000/v1/health/ready"
- },
- {
- "id": "local-endpoint-no-v1",
- "description": "Local request uses path without /v1/ prefix",
- "check": "Script contains 'localhost:8000/biology/ipd/proteinmpnn/predict'"
- },
- {
- "id": "input-pdb-inline",
- "description": "PDB content is sent inline in input_pdb field, not as a file path",
- "check": "Script reads backbone.pdb file content and passes the string content to input_pdb"
- }
+ "[docker-image-tag] References the correct ProteinMPNN container image: Output contains 'nvcr.io/nim/ipd/proteinmpnn'",
+ "[env-contract-and-cache] Local setup uses the repo env contract and LOCAL_NIM_CACHE: Output sources repo-root .env only if present, supports NVIDIA_API_KEY fallback to NGC_API_KEY, requires LOCAL_NIM_CACHE, and mounts LOCAL_NIM_CACHE to /home/nvs/.cache/nim",
+ "[cache-mount-path] Volume mount uses /home/nvs/.cache/nim (not /opt/nim/.cache): Output contains '/home/nvs/.cache/nim' as the container-side cache path in the -v mount",
+ "[health-check] Includes health check before sending request: Output contains health check against localhost:8000/v1/health/ready",
+ "[local-endpoint-no-v1] Local request uses path without /v1/ prefix: Script contains 'localhost:8000/biology/ipd/proteinmpnn/predict'",
+ "[input-pdb-inline] PDB content is sent inline in input_pdb field, not as a file path: Script reads backbone.pdb file content and passes the string content to input_pdb"
]
},
{
@@ -131,36 +55,12 @@
"expected_output": "A Python script with use_soluble_model=True, sampling_temp=[0.1, 0.3, 0.5], and omit_AAs=['M'], calling the hosted endpoint and saving the multi-FASTA output.",
"files": [],
"assertions": [
- {
- "id": "hosted-endpoint-url",
- "description": "Uses the correct hosted endpoint URL",
- "check": "Script contains 'health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict'"
- },
- {
- "id": "soluble-model",
- "description": "use_soluble_model is set to True",
- "check": "Script contains 'use_soluble_model' and 'True'"
- },
- {
- "id": "multiple-temperatures",
- "description": "sampling_temp contains multiple temperature values",
- "check": "Script contains 'sampling_temp' with a list containing 0.1, 0.3, and 0.5 (or similar multiple values)"
- },
- {
- "id": "omit-methionine",
- "description": "omit_AAs excludes methionine ('M')",
- "check": "Script contains 'omit_AAs' and 'M'"
- },
- {
- "id": "saves-mfasta",
- "description": "Saves multi-FASTA output to a file",
- "check": "Script writes mfasta from response to a .fa or .fasta file"
- },
- {
- "id": "pdb-read-inline",
- "description": "Reads PDB file and passes content inline (not path)",
- "check": "Script reads protein.pdb and passes its text content as input_pdb string"
- }
+ "[hosted-endpoint-url] Uses the correct hosted endpoint URL: Script contains 'health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict'",
+ "[soluble-model] use_soluble_model is set to True: Script contains 'use_soluble_model' and 'True'",
+ "[multiple-temperatures] sampling_temp contains multiple temperature values: Script contains 'sampling_temp' with a list containing 0.1, 0.3, and 0.5 (or similar multiple values)",
+ "[omit-methionine] omit_AAs excludes methionine ('M'): Script contains 'omit_AAs' and 'M'",
+ "[saves-mfasta] Saves multi-FASTA output to a file: Script writes mfasta from response to a .fa or .fasta file",
+ "[pdb-read-inline] Reads PDB file and passes content inline (not path): Script reads protein.pdb and passes its text content as input_pdb string"
]
}
]
diff --git a/nim-skills/proteinmpnn-nim/references/api.md b/nim-skills/proteinmpnn-nim/references/api.md
index 201b5c8..16c42f8 100644
--- a/nim-skills/proteinmpnn-nim/references/api.md
+++ b/nim-skills/proteinmpnn-nim/references/api.md
@@ -5,16 +5,41 @@
| Mode | Method | URL |
|---|---|---|
| Hosted | POST | `https://health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict` |
-| Local Docker | POST | `${PROTEINMPNN_NIM_URL:-http://localhost:8000}/biology/ipd/proteinmpnn/predict` |
-| Health (local) | GET | `${PROTEINMPNN_NIM_URL:-http://localhost:8000}/v1/health/ready` |
+| Local Docker | POST | `http://localhost:8000/biology/ipd/proteinmpnn/predict` |
+| Health (local) | GET | `http://localhost:8000/v1/health/ready` |
**IMPORTANT**: Local path has no `/v1/` prefix.
-Set `NIM_API_MODE=local` and, when the client is not in the NIM container,
-set `PROTEINMPNN_NIM_URL` to the container-reachable base URL. Do not use the
+The local URLs above use the default base URL. Set `NIM_API_MODE=local` and,
+when the client is not in the NIM container, set `PROTEINMPNN_NIM_URL` to the
+container-reachable base URL and append the same endpoint paths. Do not use the
client container's `localhost` for a NIM running in a separate container. Local
inference requests do not use an authorization header.
+## Data Handling and Permissions
+
+`SKILL.md` declares `network` for inference HTTP requests and `env` for reading
+`NIM_API_MODE`, `PROTEINMPNN_NIM_URL`, and the hosted `NGC_API_KEY`. The existing
+`Read` and `Write` tool declarations cover the user's PDB and saved artifacts.
+
+- **Hosted:** the full PDB content and design parameters leave the user's
+ environment in a JSON POST to
+ `https://health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict`. The API key
+ is sent only as an HTTPS Bearer authorization header, not in the JSON body or
+ saved request. The client does not follow redirects.
+- **Local:** the same input is sent to the user-selected `PROTEINMPNN_NIM_URL`
+ (default `http://localhost:8000`) with no authorization header. Use an approved
+ NIM deployment for confidential structures; setting a remote URL still sends
+ the structure to that machine. Registry authentication and model downloads
+ during Docker setup are separate from inference.
+- **Authorization:** disclose the hosted upload before execution. An explicit
+ request to process the PDB with the hosted API or prior approval authorizes
+ that transfer; otherwise obtain confirmation first. Never silently fall back
+ from local to hosted processing.
+- **Artifacts:** `request.json` retains the full input PDB; response, FASTA,
+ and summary files retain the returned sequences and scores. Choose an output
+ location suitable for this data, and never save or print API credentials.
+
---
## Request Body Schema
@@ -52,9 +77,17 @@ All fields are optional (minimum: provide `input_pdb`).
| Field | Type | Description |
|---|---|---|
| `mfasta` | string | Multi-FASTA string with all designed sequences |
-| `scores` | array[float] | Log-probabilities per designed sequence (higher = more confident) |
+| `scores` | array[float] | Returned sequence scores; match to designed records and preserve the values |
| `probs` | array | Per-position amino acid probabilities |
+The [NIM endpoint documentation](https://docs.nvidia.com/nim/bionemo/proteinmpnn/latest/endpoints.html)
+describes the JSON scores as log-probabilities. The original ProteinMPNN FASTA
+[`score` and `global_score` fields](https://github.com/dauparas/ProteinMPNN#readme)
+are negative log-probabilities (lower is better); `score` covers designed
+residues, while `global_score` covers all residues. Keep the score source explicit
+and verify the served version's convention before ranking across these fields.
+Scores are not calibrated folding or binding probabilities.
+
### Example mfasta output
```
diff --git a/nim-skills/proteinmpnn-nim/references/examples.md b/nim-skills/proteinmpnn-nim/references/examples.md
index df3e94c..d357b54 100644
--- a/nim-skills/proteinmpnn-nim/references/examples.md
+++ b/nim-skills/proteinmpnn-nim/references/examples.md
@@ -36,10 +36,37 @@ payload = {
}
```
-## Save Multi-FASTA
+## Save Multi-FASTA and Report Scores
+
+For common requests, use `scripts/design.py`; it performs the request, saves the
+raw response and FASTA, and prints every designed sequence with its score and
+artifact paths. For custom request code, apply the same result parser after a
+successful response. Run this example from the skill directory and set
+`expected_count` to the requested number of designs for that request:
```python
-mfasta = result["mfasta"]
-with open("designed_sequences.fa", "w", encoding="utf-8") as handle:
- handle.write(mfasta)
+import json
+from pathlib import Path
+from scripts.design import design_results
+
+summary = design_results(result, expected_count)
+output = Path("design-run") # choose a new directory for each run
+output.mkdir(parents=True, exist_ok=False)
+(output / "response.json").write_text(json.dumps(result, indent=2, allow_nan=False) + "\n")
+fasta_path = output / "designed_sequences.fa"
+fasta_path.write_text(result["mfasta"])
+summary["fasta_path"] = str(fasta_path.resolve())
+(output / "summary.json").write_text(json.dumps(summary, indent=2, allow_nan=False) + "\n")
+print(json.dumps(summary, indent=2, allow_nan=False))
```
+
+Keep the native/WT row in the saved raw FASTA, but exclude it from the design
+count and score table. An array with one score per design maps to designed rows
+only; an array that includes the native row must lose that row's score too.
+Never silently truncate mismatched arrays with `zip`. If JSON scores are absent,
+the helper can use each design's exact `score=` header field; it does not confuse
+`global_score=` with `score=` or assign the native header score to a design.
+
+Report the generated count, design identifiers, returned scores, sequence text
+or labelled previews, and the saved paths. Report failure if the request or
+result validation failed; example sequences are not execution evidence.
diff --git a/nim-skills/proteinmpnn-nim/references/validation.md b/nim-skills/proteinmpnn-nim/references/validation.md
index 8c68723..029a46d 100644
--- a/nim-skills/proteinmpnn-nim/references/validation.md
+++ b/nim-skills/proteinmpnn-nim/references/validation.md
@@ -10,13 +10,21 @@ designed sequences as useful.
- Designed sequence count matches `num_seq_per_target` after accounting for any
native/WT row.
- `scores`, when present, are reported for designed sequences only.
+- If the leading FASTA record is native/WT, determine whether the JSON score
+ array includes it before matching scores. Reject unexplained count mismatches.
+- When JSON scores are absent, preserve scores from the corresponding design
+ headers and record that source. Do not substitute `global_score` or a WT score.
## Artifact Checks
- Save `mfasta` as `.fa` or `.fasta`.
+- Preserve the raw HTTP body and parsed response; keep a summary of the actual
+ design count, identifiers, sequences, matched scores, and artifact paths.
- Keep request metadata including input PDB name, chains, temperatures, omitted
residues, and soluble-model flag.
- Do not overwrite outputs from multiple temperatures.
+- Confirm the files were written before saying the task is complete. An HTTP
+ error, a pending response, or malformed output is not a completed design.
## Scientific Checks
diff --git a/nim-skills/proteinmpnn-nim/scripts/design.py b/nim-skills/proteinmpnn-nim/scripts/design.py
new file mode 100644
index 0000000..77aa05b
--- /dev/null
+++ b/nim-skills/proteinmpnn-nim/scripts/design.py
@@ -0,0 +1,202 @@
+#!/usr/bin/env python3
+"""Execute one ProteinMPNN request and save sequences with response-derived scores."""
+
+from __future__ import annotations
+
+import argparse
+import json
+import math
+import os
+from pathlib import Path
+import re
+import sys
+from urllib.parse import urlsplit
+
+import requests
+
+
+HOSTED_URL = "https://health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict"
+AMINO_ACIDS = set("ACDEFGHIKLMNPQRSTVWYX")
+DESIGN_HEADER = re.compile(r"(?:^|,)\s*(?:T|sample|seq)\s*=")
+NATIVE_HEADER = re.compile(r"(?:^|[\s,|])(?:native|wt|wild[-_ ]type)(?:$|[\s,|])", re.I)
+
+
+def finite_number(value: object, label: str) -> float:
+ if isinstance(value, bool) or not isinstance(value, (int, float)) or not math.isfinite(value):
+ raise ValueError(f"{label} must be a finite number")
+ return float(value)
+
+
+def parse_fasta(text: object) -> list[dict]:
+ if not isinstance(text, str) or not text.strip():
+ raise ValueError("Response mfasta must be a nonempty string")
+ records: list[dict] = []
+ for line in text.splitlines():
+ line = line.strip()
+ if not line:
+ continue
+ if line.startswith(">"):
+ if not line[1:].strip():
+ raise ValueError("FASTA header must not be empty")
+ records.append({"header": line[1:].strip(), "sequence": ""})
+ elif not records:
+ raise ValueError("FASTA sequence appears before its header")
+ else:
+ records[-1]["sequence"] += line
+ if not records:
+ raise ValueError("Response contains no FASTA records")
+ for record in records:
+ # ProteinMPNN separates chains with '/'; retain that representation.
+ if any(not chain or set(chain.upper()) - AMINO_ACIDS for chain in record["sequence"].split("/")):
+ raise ValueError("FASTA contains an empty chain or invalid amino-acid sequence")
+ return records
+
+
+def design_results(result: object, expected_count: int) -> dict:
+ """Align scores without assigning the first design's score to a native row."""
+ if not isinstance(result, dict):
+ raise ValueError("ProteinMPNN must return a JSON object")
+ records = parse_fasta(result.get("mfasta"))
+ first_header = records[0]["header"]
+ first_is_native = bool(NATIVE_HEADER.search(first_header)) or (
+ not DESIGN_HEADER.search(first_header)
+ and (
+ bool(re.search(r"(?:^|,)\s*(?:fixed_chains|designed_chains)\s*=", first_header))
+ or (len(records) == expected_count + 1
+ and all(DESIGN_HEADER.search(row["header"]) for row in records[1:]))
+ )
+ )
+ native_count = int(first_is_native)
+ designs = records[native_count:]
+ if len(designs) != expected_count or any(NATIVE_HEADER.search(row["header"]) for row in designs):
+ raise ValueError(f"Expected {expected_count} designed sequences, plus an optional leading native/WT record")
+
+ scores = result.get("scores")
+ score_source = "response.scores"
+ if scores is None:
+ # Some responses carry scores in FASTA headers instead of a JSON array.
+ scores = []
+ for row in designs:
+ match = re.search(r"(?:^|,)\s*score\s*=\s*([^,\s]+)", row["header"])
+ if match is None:
+ raise ValueError("Missing designed-sequence scores in both scores and FASTA headers")
+ scores.append(float(match.group(1)))
+ score_source = "mfasta.header.score"
+ elif not isinstance(scores, list):
+ raise ValueError("Response scores must be an array")
+ elif native_count and len(scores) == len(records):
+ # When the array includes the native record, remove its corresponding score.
+ scores = scores[1:]
+ if len(scores) != len(designs):
+ raise ValueError("Score count does not match the designed FASTA records; inspect response.json")
+ for index, (row, score) in enumerate(zip(designs, scores), start=1):
+ finite_number(score, "Designed-sequence score")
+ row.update({"design_index": index, "score": score})
+ return {
+ "generated_count": len(designs), "native_count": native_count,
+ "score_source": score_source, "scores": scores, "sequences": designs,
+ }
+
+
+def design(args: argparse.Namespace) -> dict:
+ if not 1 <= args.num_sequences <= 100:
+ raise ValueError("--num-sequences must be between 1 and 100")
+ if not 0 <= finite_number(args.temperature, "Temperature") <= 1:
+ raise ValueError("--temperature must be between 0 and 1")
+ if finite_number(args.timeout, "Timeout") <= 0:
+ raise ValueError("--timeout must be positive")
+ if args.omit_aas and any(aa not in AMINO_ACIDS - {"X"} for aa in args.omit_aas):
+ raise ValueError("--omit-aas must contain standard one-letter amino-acid codes")
+ pdb = args.pdb.resolve()
+ pdb_content = pdb.read_text(encoding="utf-8")
+ if not any(line.startswith("ATOM ") for line in pdb_content.splitlines()):
+ raise ValueError("Input PDB must contain ATOM records")
+
+ mode = args.mode or os.getenv("NIM_API_MODE") or ("local" if os.getenv("PROTEINMPNN_NIM_URL") else None)
+ headers = {"Content-Type": "application/json"}
+ if mode == "hosted":
+ key = os.getenv("NGC_API_KEY")
+ if not key:
+ raise ValueError("Set NGC_API_KEY in the environment for hosted ProteinMPNN")
+ headers["Authorization"] = f"Bearer {key}"
+ url = HOSTED_URL
+ elif mode == "local":
+ base = os.getenv("PROTEINMPNN_NIM_URL", "http://localhost:8000").rstrip("/")
+ parsed = urlsplit(base)
+ if (parsed.scheme not in {"http", "https"} or not parsed.netloc or parsed.username
+ or parsed.password or parsed.query or parsed.fragment):
+ raise ValueError("PROTEINMPNN_NIM_URL must be an HTTP(S) base URL without credentials, query, or fragment")
+ url = f"{base}/biology/ipd/proteinmpnn/predict"
+ else:
+ raise ValueError("Choose --mode hosted or local, or set NIM_API_MODE")
+
+ payload = {
+ "input_pdb": pdb_content, "num_seq_per_target": args.num_sequences,
+ "sampling_temp": [args.temperature], "use_soluble_model": args.soluble,
+ "ca_only": args.ca_only,
+ }
+ for name, value in (("input_pdb_chains", args.chains), ("omit_AAs", args.omit_aas), ("random_seed", args.seed)):
+ if value is not None:
+ payload[name] = value
+
+ output = args.output_dir.resolve()
+ # One atomic reservation prevents concurrent runs from sharing artifacts.
+ try:
+ output.mkdir(parents=True, exist_ok=False)
+ except FileExistsError as exc:
+ raise FileExistsError("Choose a new --output-dir; this directory already exists") from exc
+ paths = {name: output / filename for name, filename in {
+ "request": "request.json", "raw_response": "response.raw", "response": "response.json",
+ "fasta": "designed_sequences.fa", "summary": "summary.json",
+ }.items()}
+ paths["request"].write_text(json.dumps(payload, indent=2) + "\n", encoding="utf-8")
+ # Print metadata, never credentials or a potentially large input structure.
+ print(json.dumps({"event": "request", "method": "POST", "endpoint": url, "mode": mode,
+ "input_pdb_path": str(pdb), "num_seq_per_target": args.num_sequences,
+ "sampling_temp": [args.temperature], "request_path": str(paths["request"])}), flush=True)
+ response = requests.post(url, headers=headers, json=payload, timeout=(10, args.timeout), allow_redirects=False)
+ paths["raw_response"].write_bytes(response.content)
+ if response.status_code != 200:
+ raise RuntimeError(f"ProteinMPNN returned HTTP {response.status_code}; inspect response.raw; design is not complete")
+ result = response.json()
+ paths["response"].write_text(json.dumps(result, indent=2, allow_nan=False) + "\n", encoding="utf-8")
+ summary = design_results(result, args.num_sequences)
+ # Preserve the complete returned FASTA, including any native reference header.
+ paths["fasta"].write_bytes(result["mfasta"].encode("utf-8"))
+ summary.update({"status": "completed", "http_status": response.status_code,
+ "mode": mode, "endpoint": url, "input_pdb_path": str(pdb),
+ "artifacts": {name: str(path) for name, path in paths.items()}})
+ paths["summary"].write_text(json.dumps(summary, indent=2, allow_nan=False) + "\n", encoding="utf-8")
+ if paths["fasta"].read_bytes() != result["mfasta"].encode("utf-8") or json.loads(paths["summary"].read_text()) != summary:
+ raise RuntimeError("Saved artifacts do not match the response")
+ print(json.dumps(summary, indent=2, allow_nan=False), flush=True)
+ return summary
+
+
+def main() -> int:
+ parser = argparse.ArgumentParser(description=__doc__)
+ parser.add_argument("--pdb", type=Path, required=True, help="Use the user's supplied PDB path")
+ parser.add_argument("--mode", choices=["hosted", "local"], help="Explicit mode overrides environment configuration")
+ parser.add_argument("--num-sequences", type=int, default=10)
+ parser.add_argument("--temperature", type=float, default=0.1, help="One temperature per request")
+ parser.add_argument("--chains", nargs="+", help="Chains to design; omit to design all chains")
+ parser.add_argument("--omit-aas", nargs="+", help="One-letter amino-acid codes to exclude")
+ parser.add_argument("--soluble", action="store_true")
+ parser.add_argument("--ca-only", action="store_true")
+ parser.add_argument("--seed", type=int)
+ parser.add_argument("--timeout", type=float, default=300, help="Response read timeout in seconds")
+ parser.add_argument("--output-dir", type=Path, required=True, help="New directory reserved for this run")
+ args = parser.parse_args()
+ try:
+ design(args)
+ except requests.RequestException as exc:
+ print(f"ProteinMPNN request failed ({type(exc).__name__}); no completed design to report", file=sys.stderr)
+ return 1
+ except (ValueError, RuntimeError, OSError) as exc:
+ print(str(exc), file=sys.stderr)
+ return 1
+ return 0
+
+
+if __name__ == "__main__":
+ raise SystemExit(main())
diff --git a/pyproject.toml b/pyproject.toml
index 540a060..672be38 100644
--- a/pyproject.toml
+++ b/pyproject.toml
@@ -26,7 +26,7 @@ requires-python = ">=3.10"
# biotite -> */complexa-binder-design/scripts/{pdb_interface,pipeline,preflight_design,pdb_to_boltz_template_cif}.py
# numpy -> */complexa-binder-design/scripts/validate_binders.py, */protein-binder-design/scripts/metrics.py
# pyyaml -> */complexa-binder-design/scripts/pipeline.py
-# requests -> */evo2-nim/scripts/generate.py
+# requests -> */evo2-nim/scripts/generate.py, */proteinmpnn-nim/scripts/design.py
# Everything else the scripts import is Python standard library.
dependencies = [
"biotite>=1.0",
diff --git a/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/BENCHMARK.md b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/BENCHMARK.md
new file mode 100644
index 0000000..e971a36
--- /dev/null
+++ b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/BENCHMARK.md
@@ -0,0 +1,119 @@
+# Skill Benchmark: proteinmpnn-nim
+
+> ✅ **Overall verdict: PASS — Recommended for publication**
+
+## Publication Recommendation
+
+Recommended for publication based on the completed evaluation evidence in this report.
+
+## Evaluation Metadata
+
+- Skill: `proteinmpnn-nim`
+- Evaluation date: 2026-09-30
+- Evaluator version: `1.5.6`
+- Agents: Claude Code (`aws/anthropic/bedrock-claude-opus-4-8`), Codex (`openai/openai/gpt-5.5`)
+- Tasks: 1 evaluation tasks (1 positive)
+- Dataset digest: `sha256:52236518d474794d5676f553c4879b1270b621cbf3f0a5dcf19ee85143e8cc9c` (skill-evaluator-dataset-snapshot/1)
+- Attempts per task: 3
+- Environment: `k8s-sandbox`
+- Tier 2 evidence: required for publication
+- Tier 3 evidence: required for publication
+
+Each task attempt ran in its own isolated sandbox pod.
+
+## What This Report Answers
+
+The three-tier evaluation checks whether the skill:
+
+- is safe to use;
+- produces correct answers;
+- is discovered and activated when needed;
+- helps the agent complete the user's goal and expected workflow; and
+- avoids wasted skill and tool usage.
+
+## Results at a Glance
+
+| Measure | Claude Code (Baseline → Skill Uplift) | Codex (Baseline → Skill Uplift) |
+|---|---:|---:|
+| Overall | 98.0% — baseline ran, but no comparable score was available; uplift unavailable | 92.9% — baseline ran, but no comparable score was available; uplift unavailable |
+| Security | 100.0% → 100.0% (±0.0 points) | 50.0% → 100.0% (+50.0 points) |
+| Correctness | 100.0% → 100.0% (±0.0 points) | 100.0% → 100.0% (±0.0 points) |
+| Discoverability | 95.0% — baseline ran, but no comparable score was available; uplift unavailable | 85.0% — baseline ran, but no comparable score was available; uplift unavailable |
+| Effectiveness | 65.0% → 100.0% (+35.0 points) | 57.9% → 100.0% (+42.1 points) |
+| Efficiency | 95.0% — baseline ran, but no comparable score was available; uplift unavailable | 79.4% — baseline ran, but no comparable score was available; uplift unavailable |
+
+**How to read this table:** baseline is the same task attempted without the target skill. Scores are rounded to one decimal; threshold-adjacent values use additional precision so their displayed band matches the verdict. Uplift is derived from those displayed scores and shown in percentage points.
+
+Example: `47.0% → 92.0% (+45.0 points)` means the skill-assisted run scored 92.0%, 45.0 percentage points above its 47.0% no-skill baseline.
+
+## Token Usage
+
+Actual Tier 3 execution usage is reported for every observed agent/case pair and both conditions.
+
+| Agent | Dataset case | With skill | Without skill | Delta | Change | Coverage |
+|---|---|---:|---:|---:|---:|---|
+| claude-code | All cases | 273,504 | 721,343 | -447,839 | -62.08% | skill 1/1; base 1/1 |
+| claude-code | 1 | 273,504 | 721,343 | -447,839 | -62.08% | skill 1/1; base 1/1 |
+| codex | All cases | 122,990 | 281,801 | -158,811 | -56.36% | skill 1/1; base 1/1 |
+| codex | 1 | 122,990 | 281,801 | -158,811 | -56.36% | skill 1/1; base 1/1 |
+| ALL AGENTS | Dataset aggregate | 396,494 | 1,003,144 | -606,650 | -60.47% | skill 2/2; base 2/2 |
+
+Prompt tokens include cached reads, so total tokens are `prompt + completion` (cached is not added twice). The Efficiency score uses `(prompt - cached) + completion`. N/A means the relevant trajectory counters were not available; coverage is never estimated.
+
+## Tier Status
+
+| Tier | Purpose | Status | Evidence |
+|---|---|---|---|
+| Tier 1 | Static validation | **PASSED WITH OBSERVATIONS** | 11 validator(s); 20 finding(s) |
+| Tier 2 | Semantic deduplication | **PASSED** | 2 validator(s); 0 finding(s) |
+| Tier 3 | Live agent evaluation | **PASS** | 2 agent(s); 1 task(s) |
+
+## Findings and Observations
+
+
+Show detailed findings and successful checks
+
+- **MEDIUM** QUALITY/quality_correctness: No documented scripts in table format (`skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/SKILL.md`)
+- **MEDIUM** QUALITY/quality_correctness: Instructions don't mention 'run_script' (`skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/SKILL.md`)
+- **MEDIUM** QUALITY/quality_correctness: SKILL_SPEC recommended field missing: 'metadata.author' (`skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/SKILL.md`)
+- **MEDIUM** QUALITY/quality_correctness: SKILL_SPEC recommended field missing: 'metadata.tags' (`skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/SKILL.md`)
+- **MEDIUM** SCHEMA/folder_hierarchy: Unexpected nesting depth for general skill (`skills/bionemo-agent-toolkit/skills/proteinmpnn-nim`)
+- 15 additional finding(s) are available in the full evaluation artifacts.
+
+
+
+## Scoring Methodology
+
+
+Show dimension definitions, source signals, and thresholds
+
+| Dimension | Question | Scored signals |
+|---|---|---|
+| Security | Is it safe to use? | `security` (100%) |
+| Correctness | Is the answer correct? | `accuracy` (100%) |
+| Discoverability | Was the right skill loaded when needed? | `skill_execution` (100%) |
+| Effectiveness | Did the skill help complete the task? | `goal_accuracy` (50%) + `behavior_check` (50%) |
+| Efficiency | Did it avoid wasted tool calls and token usage? | `skill_efficiency` (50%) + `token_efficiency` (50%) |
+
+- Dimension bands: PASS at 50% or above; NEUTRAL from 40% to below 50%; FAIL below 40%.
+- Overall Tier 3 lift: PASS at +5 points or more; FAIL at -10 points or less; values between those bands are NEUTRAL.
+- Overall verdict: PASS only when every configured dimension passes for at least one supported agent. Lift is reported as diagnostic evidence and does not override this gate.
+- The 50% attempt pass threshold is a separate per-task gate; it is not the dimension pass threshold.
+- Effectiveness is the equal-weight mean of goal completion (`goal_accuracy`) and expected workflow adherence (`behavior_check`).
+- Efficiency is 50% tool-call productivity (the backward-compatible `skill_efficiency` wire id) and 50% `token_efficiency`. Positive-case skill routing is scored under Discoverability, not Efficiency; a negative case without a routing target is N/A. N/A sources are omitted, remaining weights are renormalized, and the dimension is marked partial.
+
+Signals present in this run:
+
+- `security` (Security): unsafe operations, secret leakage, and unauthorized access.
+- `skill_execution` (Skill Execution): whether the expected skill was selected, decoys were avoided, and the workflow executed.
+- `skill_efficiency` (Tool Productivity): tool-call productivity (legacy wire id; routing is scored under Discoverability).
+- `accuracy` (Accuracy): final-answer correctness against the reference answer.
+- `goal_accuracy` (Goal Accuracy): whether the user's goal was achieved.
+- `behavior_check` (Behavior Check): whether the expected workflow behavior was followed.
+- `token_efficiency` (Token Efficiency): actual uncached prompt plus completion usage (50% of Efficiency).
+
+
+
+## Freshness
+
+Regenerate this benchmark when the skill, evaluation dataset, target agent/model, evaluator version, environment, or scoring policy changes.
diff --git a/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/SKILL.md b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/SKILL.md
index 900b559..6d6568c 100644
--- a/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/SKILL.md
+++ b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/SKILL.md
@@ -1,15 +1,20 @@
---
name: proteinmpnn-nim
description: >
- Run ProteinMPNN inverse folding via NVIDIA NIM to design protein sequences for a target backbone. Use for ProteinMPNN, inverse folding, sequence design, backbone redesign, fixed chains/residues, omit_AAs, sampling temperature, soluble model, hosted NVIDIA API, local Docker, PDB input, and multi-FASTA output.
+ Run ProteinMPNN inverse folding via NVIDIA NIM to design protein sequences for a target backbone. Sends user-provided PDB files and design parameters to NVIDIA's hosted API, authenticated with an environment API key, or to a user-selected local NIM. Use for sequence design, backbone redesign, fixed chains and residues, omit_AAs, sampling temperature, soluble model, local Docker, and multi-FASTA output.
license: Apache-2.0 AND CC-BY-4.0
-compatibility: "requests>=2.28"
+compatibility: "Python >=3.10; requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
+permissions:
+ - network
+ - env
---
# ProteinMPNN NIM
-Design protein sequences for a supplied backbone PDB. Use this `SKILL.md` for
+
+
+Design protein sequences for a supplied backbone PDB. Use this guide for
first-pass hosted/local usage; load supplemental files only when needed:
- `references/api.md`: exact endpoints, schemas, Docker flags, response fields.
@@ -20,7 +25,7 @@ first-pass hosted/local usage; load supplemental files only when needed:
## Choose Mode
-Honor an explicitly configured runtime before asking. `NIM_API_MODE=local` selects
+Honor the user's explicit mode; otherwise use the configured runtime. `NIM_API_MODE=local` selects
the local service at `PROTEINMPNN_NIM_URL`; the URL defaults to
`http://localhost:8000` for a NIM running in the same host or container. Ask only
when neither the environment nor the user's request makes the mode clear:
@@ -28,7 +33,8 @@ when neither the environment nor the user's request makes the mode clear:
> Hosted NVIDIA API or local Docker NIM?
- Hosted: `https://health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict`
-- Local: `${PROTEINMPNN_NIM_URL:-http://localhost:8000}/biology/ipd/proteinmpnn/predict`
+- Local: append `/biology/ipd/proteinmpnn/predict` to `PROTEINMPNN_NIM_URL`
+ (default base URL: `http://localhost:8000`).
Local inference paths do not include `/v1/`. Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Supported local Docker
startup uses `NGC_API_KEY` (or `NVIDIA_API_KEY` via the preflight) for
@@ -37,6 +43,25 @@ into the container with `-e NGC_API_KEY`. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.
+## Data Transfer and Authorization
+
+Before a hosted request, tell the user that the **entire PDB file and design
+parameters will be uploaded to NVIDIA's hosted API** at the endpoint above.
+Proceed if the user has explicitly requested hosted processing of that PDB or
+already approved the transfer; otherwise ask for confirmation before submitting.
+For confidential structures, recommend a local NIM in the user's approved
+environment. A configured local URL may point to another machine; use only the
+configured or user-selected destination. Do not switch from local to hosted processing without
+the user's authorization.
+
+The client reads `NIM_API_MODE`, `PROTEINMPNN_NIM_URL`, and, for hosted mode,
+`NGC_API_KEY` from the environment. It sends the key only in the HTTPS
+Authorization header to the hosted endpoint; local inference sends no key.
+Keep credentials out of logs and saved artifacts. The output directory contains
+the full input PDB in `request.json` and the returned sequences and scores, so
+use a location appropriate for the input's sensitivity. See
+[`references/api.md`](references/api.md) for endpoint and data-handling details.
+
## Local Docker
For local setup, run the full sequence — env preflight, `docker login`,
@@ -56,53 +81,68 @@ proteinmpnn_nim_url="${PROTEINMPNN_NIM_URL:-http://localhost:8000}"
until curl -sf "${proteinmpnn_nim_url%/}/v1/health/ready"; do sleep 5; done
```
-## Request Pattern
-
-Read PDB content inline; do not send only a file path.
-
-```python
-import os
-from pathlib import Path
-import requests
-
-HOSTED = os.getenv("NIM_API_MODE", "hosted").strip().lower() != "local"
-pdb_content = Path("1R42.pdb").read_text()
-nim_url = os.getenv("PROTEINMPNN_NIM_URL", "http://localhost:8000").rstrip("/")
-url = (
- "https://health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict"
- if HOSTED else f"{nim_url}/biology/ipd/proteinmpnn/predict"
-)
-headers = {"Content-Type": "application/json"}
-if HOSTED:
- headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"
-
-payload = {
- "input_pdb": pdb_content,
- "num_seq_per_target": 10,
- "sampling_temp": [0.1],
- "use_soluble_model": False,
- "ca_only": False,
-}
-response = requests.post(url, headers=headers, json=payload, timeout=300)
-response.raise_for_status()
-result = response.json()
-```
+## Instructions
+
+For a request to execute a design, run [`scripts/design.py`](scripts/design.py)
+and inspect its results. Writing a request script alone does not complete an
+execution request. If the user asks only for code or setup instructions, provide
+those without submitting an inference request.
+
+1. Use the user's PDB path and requested sequence count. The client reads the
+ entire PDB into `input_pdb`; do not replace or truncate the supplied backbone.
+2. Select `--mode hosted` or `--mode local` and follow **Data Transfer and
+ Authorization** above before submitting. Hosted mode uploads the PDB to the
+ documented NVIDIA endpoint and requires `NGC_API_KEY` in the environment.
+ Check only whether the key is set; do not print it, dump the environment, or
+ save authentication headers. Local inference sends no authorization header.
+3. Choose a new `--output-dir` for each request. The client reserves it before
+ submitting, preserves the raw response for diagnostics, and validates the
+ designed sequence count and score alignment before reporting completion.
+4. Read `summary.json` and report the actual results described below. If the
+ request or validation fails, report the failure and diagnostic path; do not
+ substitute example sequences or repeatedly resubmit the same request.
-Common controls:
+## Examples
+
+Run from this skill's directory, or use an absolute path to `scripts/design.py`.
+Substitute the user's input path and a new output directory:
+
+```bash
+python scripts/design.py --mode hosted \
+ --pdb /path/to/backbone.pdb --num-sequences 10 \
+ --temperature 0.1 --output-dir /path/to/new-design-run
+```
-- Redesign only chain A: `"input_pdb_chains": ["A"]`.
-- Exclude amino acids: `"omit_AAs": ["C"]` or `"omit_AAs": ["M"]`.
-- Diversity: `"sampling_temp": [0.1, 0.3, 0.5]` (always a list).
-- Solubility bias: `"use_soluble_model": True`.
-- Candidate count: `num_seq_per_target` is 1-100.
+For a running local NIM, use `--mode local`; the client honors
+`PROTEINMPNN_NIM_URL`. To design only chain A, exclude cysteine, or request the
+soluble model, add `--chains A`, `--omit-aas C`, or `--soluble` respectively.
+`--seed` sets `random_seed`; `--ca-only` selects the CA-only model. The helper
+uses one temperature per request; run separate output directories for a
+temperature sweep. For advanced JSONL controls or a custom batch request,
+use [`references/api.md`](references/api.md) and the post-response example in
+[`references/examples.md`](references/examples.md).
## Save And Report Output
-Save the returned `mfasta` and pair scores only with designed (non-native/WT)
-rows, using the snippet in [`references/examples.md`](references/examples.md)
-under **Save Multi-FASTA**. Validate promising designs by predicting structures
-with Boltz2 or OpenFold3 and comparing them to the target backbone. For
-FASTA/score sanity checks, read `references/validation.md`.
+The client writes `request.json`, `response.raw`, `response.json`,
+`designed_sequences.fa`, and `summary.json` into the requested output directory.
+The FASTA preserves the complete returned `mfasta`, including a native/WT entry
+when present. The summary contains only designed sequences, each paired with
+its actual score, and records whether scores came from the JSON array or FASTA
+headers. It is also printed after the artifacts are saved and checked.
+
+In the final response, report:
+
+- The number of **designed** sequences, excluding the native/WT reference.
+- Each design's identifier and actual returned score, plus its sequence (for
+ long sequences, give a clearly labelled preview and link to the full FASTA).
+- The saved FASTA and summary paths, and the raw response path for provenance.
+- That these are inverse-folding candidates, with no fold-back validation
+ performed unless it was actually requested and run.
+
+Do not treat a score as proof that a sequence folds or binds. Further validation
+with Boltz2 or OpenFold3 is an optional next step. For FASTA/score sanity checks,
+read [`references/validation.md`](references/validation.md).
## Limits And Troubleshooting
diff --git a/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/config/skillspector-baseline.yml b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/config/skillspector-baseline.yml
index bae7f22..8affdd5 100644
--- a/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/config/skillspector-baseline.yml
+++ b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/config/skillspector-baseline.yml
@@ -3,7 +3,7 @@
# Audited false-positive suppression, auto-applied by the NVSkills Tier 1
# runner via config/skillspector-baseline.yml. Suppressed findings remain in
# the report JSON marked `suppressed: true` with the reason below.
-version: 1
+version: 2
rules:
- id: "PE3"
@@ -16,3 +16,15 @@ rules:
`docker login nvcr.io`. This is first-party, user-facing setup guidance for
the user's own credential file — not credential theft. No SSH keys, cloud
credential stores, or third-party secret files are accessed.
+ - id: "PE3"
+ path: "evals/evals.json"
+ # Match only the literal dotenv token, including PE3's trailing space.
+ message: ".env "
+ reason: >-
+ Reviewed 2026-09-30. The two matches are in deferred Docker evaluation
+ 3's expected output and setup assertion. They describe optional loading
+ of the user's own repo-root dotenv file for NGC authentication, the same
+ setup documented in references/api.md. This JSON is evaluation text and
+ does not read or transmit credentials. The active hosted evaluation and
+ design client read NGC_API_KEY from the environment. This exception is
+ limited to the literal dotenv matches in this evaluation file.
diff --git a/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/evals/README.md b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/evals/README.md
new file mode 100644
index 0000000..77cd589
--- /dev/null
+++ b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/evals/README.md
@@ -0,0 +1,37 @@
+# ProteinMPNN evaluations
+
+The default `config.yml` selects the hosted request in `evals.json`. It uses
+`NGC_API_KEY`, stages `files/1R42.pdb`, and requires an executed inference request,
+saved designed sequences, and response-derived scores.
+
+## Local GPU task
+
+`harbor/proteinmpnn-local-design` is a separate native Harbor task with a custom
+grader. It requires the ProteinMPNN NIM image, a GPU, its model download
+credentials, and a running NIM server in the task container. A shared Astra
+sandbox using the generic agent template ignores the task image and cannot
+satisfy its loopback health check.
+
+For a dedicated local GPU evaluation, use a separate checkout and replace
+`evals/config.yml` with:
+
+```yaml
+schema_version: 1
+harbor:
+ task_source: native_harbor
+ custom_dockerfile_mode: preserve
+ base_image_mode: disabled
+ n_attempts: 1
+ pass_threshold: 0.8
+ stop_on_pass: false
+ n_concurrent: 1
+ runtime_env:
+ - NGC_API_KEY
+grading:
+ mode: custom_only
+```
+
+Run on a GPU-capable Docker host with `--env-mode docker`, or a dedicated sandbox
+template configured with the NIM image, GPU, credentials, and server startup.
+The native task's readiness check must pass before an agent starts. Keep the
+hosted configuration as the shared CI default.
diff --git a/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/evals/config.yml b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/evals/config.yml
index dac0144..6ed0cda 100644
--- a/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/evals/config.yml
+++ b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/evals/config.yml
@@ -1,13 +1,11 @@
schema_version: 1
-# ProteinMPNN is the repository's local-GPU BYOT example. ACES supports one
-# task source per skill, so this selects the native task below. The hosted
-# evals.json dataset remains available if this is switched back to evals_json
-# with aces_default grading.
+# The shared Astra sandbox uses a fixed agent image and does not launch the
+# native task's GPU NIM image. Use the hosted execution dataset for CI.
+# The native local-GPU task is retained under harbor/; see README.md to run it
+# in a runtime configured for that image and GPU.
harbor:
- task_source: native_harbor
- custom_dockerfile_mode: preserve
- base_image_mode: disabled
+ task_source: evals_json
n_attempts: 1
pass_threshold: 0.8
stop_on_pass: false
@@ -16,4 +14,4 @@ harbor:
- NGC_API_KEY
grading:
- mode: custom_only
+ mode: aces_default
diff --git a/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/evals/evals.json b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/evals/evals.json
index a0239ec..924fb07 100644
--- a/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/evals/evals.json
+++ b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/evals/evals.json
@@ -9,41 +9,13 @@
"evals/files/1R42.pdb"
],
"assertions": [
- {
- "id": "hosted-request-executed",
- "description": "Executes the hosted request instead of only writing code",
- "check": "Trajectory shows successful execution of the hosted request, and the final response reports actual response-derived designed sequences, scores, and the saved FASTA path"
- },
- {
- "id": "hosted-endpoint-url",
- "description": "Uses the correct hosted ProteinMPNN endpoint URL",
- "check": "Script contains 'health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict'"
- },
- {
- "id": "bearer-auth-header",
- "description": "Sets Authorization header with Bearer token from NGC_API_KEY",
- "check": "Script contains 'Authorization' and 'Bearer' and 'NGC_API_KEY'"
- },
- {
- "id": "input-pdb-field",
- "description": "Sends PDB file content as 'input_pdb' field (inline string, not file path)",
- "check": "Script reads 1R42.pdb and passes its content to 'input_pdb' field in the payload"
- },
- {
- "id": "num-seq-per-target",
- "description": "num_seq_per_target is set to 10",
- "check": "Script contains 'num_seq_per_target' and '10'"
- },
- {
- "id": "saves-mfasta-output",
- "description": "Saves the mfasta string from the response to a .fa file",
- "check": "Script accesses 'mfasta' from response and writes it to a file with .fa or .fasta extension"
- },
- {
- "id": "scores-reported",
- "description": "Reports scores only for designed sequences",
- "check": "Script accounts for the native/WT FASTA row when present and pairs scores with designed sequences, not the WT row"
- }
+ "[hosted-request-executed] Executes the hosted request instead of only writing code: Trajectory shows successful execution of the hosted request, and the final response reports actual response-derived designed sequences, scores, and the saved FASTA path",
+ "[hosted-endpoint-url] Uses the correct hosted ProteinMPNN endpoint URL: Script contains 'health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict'",
+ "[bearer-auth-header] Sets Authorization header with Bearer token from NGC_API_KEY: Script contains 'Authorization' and 'Bearer' and 'NGC_API_KEY'",
+ "[input-pdb-field] Sends PDB file content as 'input_pdb' field (inline string, not file path): Script reads 1R42.pdb and passes its content to 'input_pdb' field in the payload",
+ "[num-seq-per-target] num_seq_per_target is set to 10: Script contains 'num_seq_per_target' and '10'",
+ "[saves-mfasta-output] Saves the mfasta string from the response to a .fa file: Script accesses 'mfasta' from response and writes it to a file with .fa or .fasta extension",
+ "[scores-reported] Reports scores only for designed sequences: Script accounts for the native/WT FASTA row when present and pairs scores with designed sequences, not the WT row"
]
}
],
@@ -54,36 +26,12 @@
"expected_output": "A Python script targeting the hosted endpoint with input_pdb_chains=['A'], num_seq_per_target=5, sampling_temp=[0.2], and omit_AAs=['C'], reading structure.pdb and saving the designed sequences.",
"files": [],
"assertions": [
- {
- "id": "hosted-endpoint-url",
- "description": "Uses the correct hosted endpoint URL",
- "check": "Script contains 'health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict'"
- },
- {
- "id": "input-pdb-chains",
- "description": "Uses input_pdb_chains to limit design to chain A",
- "check": "Script contains 'input_pdb_chains' and 'A'"
- },
- {
- "id": "sampling-temp-array",
- "description": "sampling_temp is passed as an array (even for a single value)",
- "check": "Script contains 'sampling_temp' with value [0.2] or as a list containing 0.2"
- },
- {
- "id": "omit-cysteines",
- "description": "omit_AAs excludes cysteine ('C')",
- "check": "Script contains 'omit_AAs' and 'C'"
- },
- {
- "id": "num-seq-5",
- "description": "num_seq_per_target is set to 5",
- "check": "Script contains 'num_seq_per_target' and '5'"
- },
- {
- "id": "saves-output",
- "description": "Saves designed sequences to a FASTA file",
- "check": "Script writes the mfasta response to a file"
- }
+ "[hosted-endpoint-url] Uses the correct hosted endpoint URL: Script contains 'health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict'",
+ "[input-pdb-chains] Uses input_pdb_chains to limit design to chain A: Script contains 'input_pdb_chains' and 'A'",
+ "[sampling-temp-array] sampling_temp is passed as an array (even for a single value): Script contains 'sampling_temp' with value [0.2] or as a list containing 0.2",
+ "[omit-cysteines] omit_AAs excludes cysteine ('C'): Script contains 'omit_AAs' and 'C'",
+ "[num-seq-5] num_seq_per_target is set to 5: Script contains 'num_seq_per_target' and '5'",
+ "[saves-output] Saves designed sequences to a FASTA file: Script writes the mfasta response to a file"
]
},
{
@@ -93,36 +41,12 @@
"expected_output": "Docker setup commands using shell env first and optional repo-root .env overrides, requiring NGC_API_KEY or NVIDIA_API_KEY fallback plus LOCAL_NIM_CACHE, using the correct image tag, single GPU flag, and the unique cache mount path /home/nvs/.cache/nim (not /opt/nim/.cache), then a health check and no-auth request script to localhost:8000.",
"files": [],
"assertions": [
- {
- "id": "docker-image-tag",
- "description": "References the correct ProteinMPNN container image",
- "check": "Output contains 'nvcr.io/nim/ipd/proteinmpnn'"
- },
- {
- "id": "env-contract-and-cache",
- "description": "Local setup uses the repo env contract and LOCAL_NIM_CACHE",
- "check": "Output sources repo-root .env only if present, supports NVIDIA_API_KEY fallback to NGC_API_KEY, requires LOCAL_NIM_CACHE, and mounts LOCAL_NIM_CACHE to /home/nvs/.cache/nim"
- },
- {
- "id": "cache-mount-path",
- "description": "Volume mount uses /home/nvs/.cache/nim (not /opt/nim/.cache)",
- "check": "Output contains '/home/nvs/.cache/nim' as the container-side cache path in the -v mount"
- },
- {
- "id": "health-check",
- "description": "Includes health check before sending request",
- "check": "Output contains health check against localhost:8000/v1/health/ready"
- },
- {
- "id": "local-endpoint-no-v1",
- "description": "Local request uses path without /v1/ prefix",
- "check": "Script contains 'localhost:8000/biology/ipd/proteinmpnn/predict'"
- },
- {
- "id": "input-pdb-inline",
- "description": "PDB content is sent inline in input_pdb field, not as a file path",
- "check": "Script reads backbone.pdb file content and passes the string content to input_pdb"
- }
+ "[docker-image-tag] References the correct ProteinMPNN container image: Output contains 'nvcr.io/nim/ipd/proteinmpnn'",
+ "[env-contract-and-cache] Local setup uses the repo env contract and LOCAL_NIM_CACHE: Output sources repo-root .env only if present, supports NVIDIA_API_KEY fallback to NGC_API_KEY, requires LOCAL_NIM_CACHE, and mounts LOCAL_NIM_CACHE to /home/nvs/.cache/nim",
+ "[cache-mount-path] Volume mount uses /home/nvs/.cache/nim (not /opt/nim/.cache): Output contains '/home/nvs/.cache/nim' as the container-side cache path in the -v mount",
+ "[health-check] Includes health check before sending request: Output contains health check against localhost:8000/v1/health/ready",
+ "[local-endpoint-no-v1] Local request uses path without /v1/ prefix: Script contains 'localhost:8000/biology/ipd/proteinmpnn/predict'",
+ "[input-pdb-inline] PDB content is sent inline in input_pdb field, not as a file path: Script reads backbone.pdb file content and passes the string content to input_pdb"
]
},
{
@@ -131,36 +55,12 @@
"expected_output": "A Python script with use_soluble_model=True, sampling_temp=[0.1, 0.3, 0.5], and omit_AAs=['M'], calling the hosted endpoint and saving the multi-FASTA output.",
"files": [],
"assertions": [
- {
- "id": "hosted-endpoint-url",
- "description": "Uses the correct hosted endpoint URL",
- "check": "Script contains 'health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict'"
- },
- {
- "id": "soluble-model",
- "description": "use_soluble_model is set to True",
- "check": "Script contains 'use_soluble_model' and 'True'"
- },
- {
- "id": "multiple-temperatures",
- "description": "sampling_temp contains multiple temperature values",
- "check": "Script contains 'sampling_temp' with a list containing 0.1, 0.3, and 0.5 (or similar multiple values)"
- },
- {
- "id": "omit-methionine",
- "description": "omit_AAs excludes methionine ('M')",
- "check": "Script contains 'omit_AAs' and 'M'"
- },
- {
- "id": "saves-mfasta",
- "description": "Saves multi-FASTA output to a file",
- "check": "Script writes mfasta from response to a .fa or .fasta file"
- },
- {
- "id": "pdb-read-inline",
- "description": "Reads PDB file and passes content inline (not path)",
- "check": "Script reads protein.pdb and passes its text content as input_pdb string"
- }
+ "[hosted-endpoint-url] Uses the correct hosted endpoint URL: Script contains 'health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict'",
+ "[soluble-model] use_soluble_model is set to True: Script contains 'use_soluble_model' and 'True'",
+ "[multiple-temperatures] sampling_temp contains multiple temperature values: Script contains 'sampling_temp' with a list containing 0.1, 0.3, and 0.5 (or similar multiple values)",
+ "[omit-methionine] omit_AAs excludes methionine ('M'): Script contains 'omit_AAs' and 'M'",
+ "[saves-mfasta] Saves multi-FASTA output to a file: Script writes mfasta from response to a .fa or .fasta file",
+ "[pdb-read-inline] Reads PDB file and passes content inline (not path): Script reads protein.pdb and passes its text content as input_pdb string"
]
}
]
diff --git a/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/references/api.md b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/references/api.md
index 201b5c8..16c42f8 100644
--- a/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/references/api.md
+++ b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/references/api.md
@@ -5,16 +5,41 @@
| Mode | Method | URL |
|---|---|---|
| Hosted | POST | `https://health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict` |
-| Local Docker | POST | `${PROTEINMPNN_NIM_URL:-http://localhost:8000}/biology/ipd/proteinmpnn/predict` |
-| Health (local) | GET | `${PROTEINMPNN_NIM_URL:-http://localhost:8000}/v1/health/ready` |
+| Local Docker | POST | `http://localhost:8000/biology/ipd/proteinmpnn/predict` |
+| Health (local) | GET | `http://localhost:8000/v1/health/ready` |
**IMPORTANT**: Local path has no `/v1/` prefix.
-Set `NIM_API_MODE=local` and, when the client is not in the NIM container,
-set `PROTEINMPNN_NIM_URL` to the container-reachable base URL. Do not use the
+The local URLs above use the default base URL. Set `NIM_API_MODE=local` and,
+when the client is not in the NIM container, set `PROTEINMPNN_NIM_URL` to the
+container-reachable base URL and append the same endpoint paths. Do not use the
client container's `localhost` for a NIM running in a separate container. Local
inference requests do not use an authorization header.
+## Data Handling and Permissions
+
+`SKILL.md` declares `network` for inference HTTP requests and `env` for reading
+`NIM_API_MODE`, `PROTEINMPNN_NIM_URL`, and the hosted `NGC_API_KEY`. The existing
+`Read` and `Write` tool declarations cover the user's PDB and saved artifacts.
+
+- **Hosted:** the full PDB content and design parameters leave the user's
+ environment in a JSON POST to
+ `https://health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict`. The API key
+ is sent only as an HTTPS Bearer authorization header, not in the JSON body or
+ saved request. The client does not follow redirects.
+- **Local:** the same input is sent to the user-selected `PROTEINMPNN_NIM_URL`
+ (default `http://localhost:8000`) with no authorization header. Use an approved
+ NIM deployment for confidential structures; setting a remote URL still sends
+ the structure to that machine. Registry authentication and model downloads
+ during Docker setup are separate from inference.
+- **Authorization:** disclose the hosted upload before execution. An explicit
+ request to process the PDB with the hosted API or prior approval authorizes
+ that transfer; otherwise obtain confirmation first. Never silently fall back
+ from local to hosted processing.
+- **Artifacts:** `request.json` retains the full input PDB; response, FASTA,
+ and summary files retain the returned sequences and scores. Choose an output
+ location suitable for this data, and never save or print API credentials.
+
---
## Request Body Schema
@@ -52,9 +77,17 @@ All fields are optional (minimum: provide `input_pdb`).
| Field | Type | Description |
|---|---|---|
| `mfasta` | string | Multi-FASTA string with all designed sequences |
-| `scores` | array[float] | Log-probabilities per designed sequence (higher = more confident) |
+| `scores` | array[float] | Returned sequence scores; match to designed records and preserve the values |
| `probs` | array | Per-position amino acid probabilities |
+The [NIM endpoint documentation](https://docs.nvidia.com/nim/bionemo/proteinmpnn/latest/endpoints.html)
+describes the JSON scores as log-probabilities. The original ProteinMPNN FASTA
+[`score` and `global_score` fields](https://github.com/dauparas/ProteinMPNN#readme)
+are negative log-probabilities (lower is better); `score` covers designed
+residues, while `global_score` covers all residues. Keep the score source explicit
+and verify the served version's convention before ranking across these fields.
+Scores are not calibrated folding or binding probabilities.
+
### Example mfasta output
```
diff --git a/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/references/examples.md b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/references/examples.md
index df3e94c..d357b54 100644
--- a/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/references/examples.md
+++ b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/references/examples.md
@@ -36,10 +36,37 @@ payload = {
}
```
-## Save Multi-FASTA
+## Save Multi-FASTA and Report Scores
+
+For common requests, use `scripts/design.py`; it performs the request, saves the
+raw response and FASTA, and prints every designed sequence with its score and
+artifact paths. For custom request code, apply the same result parser after a
+successful response. Run this example from the skill directory and set
+`expected_count` to the requested number of designs for that request:
```python
-mfasta = result["mfasta"]
-with open("designed_sequences.fa", "w", encoding="utf-8") as handle:
- handle.write(mfasta)
+import json
+from pathlib import Path
+from scripts.design import design_results
+
+summary = design_results(result, expected_count)
+output = Path("design-run") # choose a new directory for each run
+output.mkdir(parents=True, exist_ok=False)
+(output / "response.json").write_text(json.dumps(result, indent=2, allow_nan=False) + "\n")
+fasta_path = output / "designed_sequences.fa"
+fasta_path.write_text(result["mfasta"])
+summary["fasta_path"] = str(fasta_path.resolve())
+(output / "summary.json").write_text(json.dumps(summary, indent=2, allow_nan=False) + "\n")
+print(json.dumps(summary, indent=2, allow_nan=False))
```
+
+Keep the native/WT row in the saved raw FASTA, but exclude it from the design
+count and score table. An array with one score per design maps to designed rows
+only; an array that includes the native row must lose that row's score too.
+Never silently truncate mismatched arrays with `zip`. If JSON scores are absent,
+the helper can use each design's exact `score=` header field; it does not confuse
+`global_score=` with `score=` or assign the native header score to a design.
+
+Report the generated count, design identifiers, returned scores, sequence text
+or labelled previews, and the saved paths. Report failure if the request or
+result validation failed; example sequences are not execution evidence.
diff --git a/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/references/validation.md b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/references/validation.md
index 8c68723..029a46d 100644
--- a/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/references/validation.md
+++ b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/references/validation.md
@@ -10,13 +10,21 @@ designed sequences as useful.
- Designed sequence count matches `num_seq_per_target` after accounting for any
native/WT row.
- `scores`, when present, are reported for designed sequences only.
+- If the leading FASTA record is native/WT, determine whether the JSON score
+ array includes it before matching scores. Reject unexplained count mismatches.
+- When JSON scores are absent, preserve scores from the corresponding design
+ headers and record that source. Do not substitute `global_score` or a WT score.
## Artifact Checks
- Save `mfasta` as `.fa` or `.fasta`.
+- Preserve the raw HTTP body and parsed response; keep a summary of the actual
+ design count, identifiers, sequences, matched scores, and artifact paths.
- Keep request metadata including input PDB name, chains, temperatures, omitted
residues, and soluble-model flag.
- Do not overwrite outputs from multiple temperatures.
+- Confirm the files were written before saying the task is complete. An HTTP
+ error, a pending response, or malformed output is not a completed design.
## Scientific Checks
diff --git a/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/scripts/design.py b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/scripts/design.py
new file mode 100644
index 0000000..77aa05b
--- /dev/null
+++ b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/scripts/design.py
@@ -0,0 +1,202 @@
+#!/usr/bin/env python3
+"""Execute one ProteinMPNN request and save sequences with response-derived scores."""
+
+from __future__ import annotations
+
+import argparse
+import json
+import math
+import os
+from pathlib import Path
+import re
+import sys
+from urllib.parse import urlsplit
+
+import requests
+
+
+HOSTED_URL = "https://health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict"
+AMINO_ACIDS = set("ACDEFGHIKLMNPQRSTVWYX")
+DESIGN_HEADER = re.compile(r"(?:^|,)\s*(?:T|sample|seq)\s*=")
+NATIVE_HEADER = re.compile(r"(?:^|[\s,|])(?:native|wt|wild[-_ ]type)(?:$|[\s,|])", re.I)
+
+
+def finite_number(value: object, label: str) -> float:
+ if isinstance(value, bool) or not isinstance(value, (int, float)) or not math.isfinite(value):
+ raise ValueError(f"{label} must be a finite number")
+ return float(value)
+
+
+def parse_fasta(text: object) -> list[dict]:
+ if not isinstance(text, str) or not text.strip():
+ raise ValueError("Response mfasta must be a nonempty string")
+ records: list[dict] = []
+ for line in text.splitlines():
+ line = line.strip()
+ if not line:
+ continue
+ if line.startswith(">"):
+ if not line[1:].strip():
+ raise ValueError("FASTA header must not be empty")
+ records.append({"header": line[1:].strip(), "sequence": ""})
+ elif not records:
+ raise ValueError("FASTA sequence appears before its header")
+ else:
+ records[-1]["sequence"] += line
+ if not records:
+ raise ValueError("Response contains no FASTA records")
+ for record in records:
+ # ProteinMPNN separates chains with '/'; retain that representation.
+ if any(not chain or set(chain.upper()) - AMINO_ACIDS for chain in record["sequence"].split("/")):
+ raise ValueError("FASTA contains an empty chain or invalid amino-acid sequence")
+ return records
+
+
+def design_results(result: object, expected_count: int) -> dict:
+ """Align scores without assigning the first design's score to a native row."""
+ if not isinstance(result, dict):
+ raise ValueError("ProteinMPNN must return a JSON object")
+ records = parse_fasta(result.get("mfasta"))
+ first_header = records[0]["header"]
+ first_is_native = bool(NATIVE_HEADER.search(first_header)) or (
+ not DESIGN_HEADER.search(first_header)
+ and (
+ bool(re.search(r"(?:^|,)\s*(?:fixed_chains|designed_chains)\s*=", first_header))
+ or (len(records) == expected_count + 1
+ and all(DESIGN_HEADER.search(row["header"]) for row in records[1:]))
+ )
+ )
+ native_count = int(first_is_native)
+ designs = records[native_count:]
+ if len(designs) != expected_count or any(NATIVE_HEADER.search(row["header"]) for row in designs):
+ raise ValueError(f"Expected {expected_count} designed sequences, plus an optional leading native/WT record")
+
+ scores = result.get("scores")
+ score_source = "response.scores"
+ if scores is None:
+ # Some responses carry scores in FASTA headers instead of a JSON array.
+ scores = []
+ for row in designs:
+ match = re.search(r"(?:^|,)\s*score\s*=\s*([^,\s]+)", row["header"])
+ if match is None:
+ raise ValueError("Missing designed-sequence scores in both scores and FASTA headers")
+ scores.append(float(match.group(1)))
+ score_source = "mfasta.header.score"
+ elif not isinstance(scores, list):
+ raise ValueError("Response scores must be an array")
+ elif native_count and len(scores) == len(records):
+ # When the array includes the native record, remove its corresponding score.
+ scores = scores[1:]
+ if len(scores) != len(designs):
+ raise ValueError("Score count does not match the designed FASTA records; inspect response.json")
+ for index, (row, score) in enumerate(zip(designs, scores), start=1):
+ finite_number(score, "Designed-sequence score")
+ row.update({"design_index": index, "score": score})
+ return {
+ "generated_count": len(designs), "native_count": native_count,
+ "score_source": score_source, "scores": scores, "sequences": designs,
+ }
+
+
+def design(args: argparse.Namespace) -> dict:
+ if not 1 <= args.num_sequences <= 100:
+ raise ValueError("--num-sequences must be between 1 and 100")
+ if not 0 <= finite_number(args.temperature, "Temperature") <= 1:
+ raise ValueError("--temperature must be between 0 and 1")
+ if finite_number(args.timeout, "Timeout") <= 0:
+ raise ValueError("--timeout must be positive")
+ if args.omit_aas and any(aa not in AMINO_ACIDS - {"X"} for aa in args.omit_aas):
+ raise ValueError("--omit-aas must contain standard one-letter amino-acid codes")
+ pdb = args.pdb.resolve()
+ pdb_content = pdb.read_text(encoding="utf-8")
+ if not any(line.startswith("ATOM ") for line in pdb_content.splitlines()):
+ raise ValueError("Input PDB must contain ATOM records")
+
+ mode = args.mode or os.getenv("NIM_API_MODE") or ("local" if os.getenv("PROTEINMPNN_NIM_URL") else None)
+ headers = {"Content-Type": "application/json"}
+ if mode == "hosted":
+ key = os.getenv("NGC_API_KEY")
+ if not key:
+ raise ValueError("Set NGC_API_KEY in the environment for hosted ProteinMPNN")
+ headers["Authorization"] = f"Bearer {key}"
+ url = HOSTED_URL
+ elif mode == "local":
+ base = os.getenv("PROTEINMPNN_NIM_URL", "http://localhost:8000").rstrip("/")
+ parsed = urlsplit(base)
+ if (parsed.scheme not in {"http", "https"} or not parsed.netloc or parsed.username
+ or parsed.password or parsed.query or parsed.fragment):
+ raise ValueError("PROTEINMPNN_NIM_URL must be an HTTP(S) base URL without credentials, query, or fragment")
+ url = f"{base}/biology/ipd/proteinmpnn/predict"
+ else:
+ raise ValueError("Choose --mode hosted or local, or set NIM_API_MODE")
+
+ payload = {
+ "input_pdb": pdb_content, "num_seq_per_target": args.num_sequences,
+ "sampling_temp": [args.temperature], "use_soluble_model": args.soluble,
+ "ca_only": args.ca_only,
+ }
+ for name, value in (("input_pdb_chains", args.chains), ("omit_AAs", args.omit_aas), ("random_seed", args.seed)):
+ if value is not None:
+ payload[name] = value
+
+ output = args.output_dir.resolve()
+ # One atomic reservation prevents concurrent runs from sharing artifacts.
+ try:
+ output.mkdir(parents=True, exist_ok=False)
+ except FileExistsError as exc:
+ raise FileExistsError("Choose a new --output-dir; this directory already exists") from exc
+ paths = {name: output / filename for name, filename in {
+ "request": "request.json", "raw_response": "response.raw", "response": "response.json",
+ "fasta": "designed_sequences.fa", "summary": "summary.json",
+ }.items()}
+ paths["request"].write_text(json.dumps(payload, indent=2) + "\n", encoding="utf-8")
+ # Print metadata, never credentials or a potentially large input structure.
+ print(json.dumps({"event": "request", "method": "POST", "endpoint": url, "mode": mode,
+ "input_pdb_path": str(pdb), "num_seq_per_target": args.num_sequences,
+ "sampling_temp": [args.temperature], "request_path": str(paths["request"])}), flush=True)
+ response = requests.post(url, headers=headers, json=payload, timeout=(10, args.timeout), allow_redirects=False)
+ paths["raw_response"].write_bytes(response.content)
+ if response.status_code != 200:
+ raise RuntimeError(f"ProteinMPNN returned HTTP {response.status_code}; inspect response.raw; design is not complete")
+ result = response.json()
+ paths["response"].write_text(json.dumps(result, indent=2, allow_nan=False) + "\n", encoding="utf-8")
+ summary = design_results(result, args.num_sequences)
+ # Preserve the complete returned FASTA, including any native reference header.
+ paths["fasta"].write_bytes(result["mfasta"].encode("utf-8"))
+ summary.update({"status": "completed", "http_status": response.status_code,
+ "mode": mode, "endpoint": url, "input_pdb_path": str(pdb),
+ "artifacts": {name: str(path) for name, path in paths.items()}})
+ paths["summary"].write_text(json.dumps(summary, indent=2, allow_nan=False) + "\n", encoding="utf-8")
+ if paths["fasta"].read_bytes() != result["mfasta"].encode("utf-8") or json.loads(paths["summary"].read_text()) != summary:
+ raise RuntimeError("Saved artifacts do not match the response")
+ print(json.dumps(summary, indent=2, allow_nan=False), flush=True)
+ return summary
+
+
+def main() -> int:
+ parser = argparse.ArgumentParser(description=__doc__)
+ parser.add_argument("--pdb", type=Path, required=True, help="Use the user's supplied PDB path")
+ parser.add_argument("--mode", choices=["hosted", "local"], help="Explicit mode overrides environment configuration")
+ parser.add_argument("--num-sequences", type=int, default=10)
+ parser.add_argument("--temperature", type=float, default=0.1, help="One temperature per request")
+ parser.add_argument("--chains", nargs="+", help="Chains to design; omit to design all chains")
+ parser.add_argument("--omit-aas", nargs="+", help="One-letter amino-acid codes to exclude")
+ parser.add_argument("--soluble", action="store_true")
+ parser.add_argument("--ca-only", action="store_true")
+ parser.add_argument("--seed", type=int)
+ parser.add_argument("--timeout", type=float, default=300, help="Response read timeout in seconds")
+ parser.add_argument("--output-dir", type=Path, required=True, help="New directory reserved for this run")
+ args = parser.parse_args()
+ try:
+ design(args)
+ except requests.RequestException as exc:
+ print(f"ProteinMPNN request failed ({type(exc).__name__}); no completed design to report", file=sys.stderr)
+ return 1
+ except (ValueError, RuntimeError, OSError) as exc:
+ print(str(exc), file=sys.stderr)
+ return 1
+ return 0
+
+
+if __name__ == "__main__":
+ raise SystemExit(main())
diff --git a/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/skill-card.md b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/skill-card.md
new file mode 100644
index 0000000..ae7680f
--- /dev/null
+++ b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/skill-card.md
@@ -0,0 +1,86 @@
+## Description:
+Run ProteinMPNN inverse folding via NVIDIA NIM to design protein sequences for a target backbone.
+
+This skill is ready for commercial/non-commercial use.
+
+## Owner
+NVIDIA
+
+### License/Terms of Use:
+Apache-2.0 AND CC-BY-4.0
+## Use Case:
+Developers and engineers designing protein sequences for target backbone structures using ProteinMPNN inverse folding through NVIDIA's hosted API or a local NIM deployment.
+
+### Deployment Geography for Use:
+Global
+
+## Requirements / Dependencies:
+**Requires API Key or External Credential:** [Yes]
+**Credential Type(s):** [API key]
+
+Do not include secrets in prompts/logs/output; use least-privilege credentials; rotate keys as appropriate.
+
+## Known Risks and Mitigations:
+Risk: Review before execution as proposals could introduce incorrect or misleading guidance into skills.
+Mitigation: Review and scan skill before deployment.
+
+## Reference(s):
+- [ProteinMPNN NIM — API Reference](references/api.md)
+- [ProteinMPNN Science Notes](references/science.md)
+- [ProteinMPNN Parameter Guidance](references/parameters.md)
+- [ProteinMPNN Validation](references/validation.md)
+- [ProteinMPNN Examples](references/examples.md)
+
+
+## Skill Output:
+**Output Type(s):** [Shell commands, Code, Files]
+**Output Format:** [Markdown with inline bash code blocks, FASTA files, and JSON artifacts]
+**Output Parameters:** [1D]
+**Other Properties Related to Output:** [None]
+
+## Evaluation Agents Used:
+- Claude Code (`aws/anthropic/bedrock-claude-opus-4-8`)
+- Codex (`openai/openai/gpt-5.5`)
+
+
+
+## Evaluation Tasks:
+1 evaluation task (1 positive) across 3 attempts per task in isolated k8s-sandbox pods.
+
+## Evaluation Metrics Used:
+Reported benchmark dimensions:
+- Security: Whether the skill avoids unsafe operations, secret leakage, and unauthorized access.
+- Correctness: Final-answer correctness against the reference answer.
+- Discoverability: Whether the expected skill was selected, decoys were avoided, and the workflow executed.
+- Effectiveness: Whether the skill helped complete the user's goal (50% goal completion + 50% expected workflow adherence).
+- Efficiency: Tool-call productivity (50%) and token efficiency (50%), measuring avoidance of wasted tool calls and token usage.
+
+Underlying evaluation signals used in this run:
+- `security`: Checks for unsafe operations, secret leakage, and unauthorized access.
+- `accuracy`: Final-answer correctness against the reference answer.
+- `skill_execution`: Whether the expected skill was selected and the workflow executed.
+- `goal_accuracy`: Whether the user's goal was achieved.
+- `behavior_check`: Whether the expected workflow behavior was followed.
+- `skill_efficiency`: Tool-call productivity, measuring avoidance of wasted tool calls.
+- `token_efficiency`: Actual uncached prompt plus completion token usage.
+
+
+
+## Evaluation Results:
+| Measure | Claude Code (Baseline → Skill Uplift) | Codex (Baseline → Skill Uplift) |
+|---|---:|---:|
+| Overall | 98.0% | 92.9% |
+| Security | 100.0% → 100.0% (±0.0 pts) | 50.0% → 100.0% (+50.0 pts) |
+| Correctness | 100.0% → 100.0% (±0.0 pts) | 100.0% → 100.0% (±0.0 pts) |
+| Discoverability | 95.0% | 85.0% |
+| Effectiveness | 65.0% → 100.0% (+35.0 pts) | 57.9% → 100.0% (+42.1 pts) |
+| Efficiency | 95.0% | 79.4% |
+
+## Skill Version(s):
+0.1.0 (source: pyproject.toml)
+
+## Ethical Considerations:
+NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal team to ensure this skill meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
+
+(For Release on NVIDIA Platforms Only)
+Please report quality, risk, security vulnerabilities or NVIDIA AI Concerns [here](https://app.intigriti.com/programs/nvidia/nvidiavdp/detail).
diff --git a/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/skill.oms.sig b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/skill.oms.sig
new file mode 100644
index 0000000..8f68112
--- /dev/null
+++ b/skills/bionemo-agent-toolkit/skills/proteinmpnn-nim/skill.oms.sig
@@ -0,0 +1 @@
+{"mediaType":"application/vnd.dev.sigstore.bundle.v0.3+json","verificationMaterial":{"x509CertificateChain":{"certificates":[{"rawBytes":"MIICgzCCAgmgAwIBAgIUKIyS7SxNteQIiWzK1dWj85E6520wCgYIKoZIzj0EAwMwVTELMAkGA1UEBhMCVVMxGzAZBgNVBAoMEk5WSURJQSBDb3Jwb3JhdGlvbjEpMCcGA1UEAwwgTlZJRElBIEFnZW50IENhcGFiaWxpdGllcyBJQ0EgMDEwHhcNMjYwNDAxMDAwMDAwWhcNMjgwNDIyMTUzMzA5WjBUMQswCQYDVQQGEwJVUzEbMBkGA1UECgwSTlZJRElBIENvcnBvcmF0aW9uMSgwJgYDVQQDDB9OVklESUEgQWdlbnQgU2tpbGxzIFNpZ25pbmcgMDAxMHYwEAYHKoZIzj0CAQYFK4EEACIDYgAEYoRM9bQl/dGlwSRNi6bTpIJUXH8Nv9GciP6LSflJYYMLCc296kpyuTSsk5ddbAWiDcFX3C/ydX3jwc+qCLYP6uHy9XphyLjOQ27Yb2J6rBLVtRBS1mgGco/Gr7fL6ODco4GaMIGXMB0GA1UdDgQWBBRQ/5ZW3nJ6lmo9SVk7I15o7UGmpTAfBgNVHSMEGDAWgBRPGpILxMBBleJSsBGjrMKsby1CgjAMBgNVHRMBAf8EAjAAMA4GA1UdDwEB/wQEAwIHgDA3BggrBgEFBQcBAQQrMCkwJwYIKwYBBQUHMAGGG2h0dHA6Ly9vY3NwLm5kaXMubnZpZGlhLmNvbTAKBggqhkjOPQQDAwNoADBlAjAUygu/GiOCIXrgGr4SmLgeEVDcEitfFUv7ALbvLVGVyMysB3mxmO/uInZfXzWcJZsCMQDxuoxj4ZmO30jhkPIcCxGFCOvnUsnfU3TfGcouYm4M6iRpbKvtVnHPiy4bi6pcKf0="},{"rawBytes":"MIICiDCCAg6gAwIBAgIUZsIuSv9NkpJCNqtYEfCouVv5BzowCgYIKoZIzj0EAwMwUTELMAkGA1UEBhMCVVMxGzAZBgNVBAoMEk5WSURJQSBDb3Jwb3JhdGlvbjElMCMGA1UEAwwcTlZJRElBIEFnZW50IENhcGFiaWxpdGllcyBDQTAgFw0yNjA0MDEwMDAwMDBaGA85OTk5MTIzMTIzNTk1OVowVTELMAkGA1UEBhMCVVMxGzAZBgNVBAoMEk5WSURJQSBDb3Jwb3JhdGlvbjEpMCcGA1UEAwwgTlZJRElBIEFnZW50IENhcGFiaWxpdGllcyBJQ0EgMDEwdjAQBgcqhkjOPQIBBgUrgQQAIgNiAASI72cR3ctKGg4VWnB3bNja6g1Z2PnOmFEopkPof+QeIcPk9rT+g9MjJnq51EQXL93a7C2GJ9J985G4o2V85VD7wJ1RaXhluHW2rf3y8bQGeAYaKMr5s/hUgn+M3/9WlWejgaAwgZ0wHQYDVR0OBBYEFE8akgvEwEGV4lKwEaOswqxvLUKCMB8GA1UdIwQYMBaAFItnoAjjfuCEUvzyvWyI2vOGvwPjMBIGA1UdEwEB/wQIMAYBAf8CAQAwDgYDVR0PAQH/BAQDAgEGMDcGCCsGAQUFBwEBBCswKTAnBggrBgEFBQcwAYYbaHR0cDovL29jc3AubmRpcy5udmlkaWEuY29tMAoGCCqGSM49BAMDA2gAMGUCMQCeIMMfAbyzPDacw2MxG+Yt1cikrJX/DVxiGfXuHmkkXn6VgSzE79+lkqDErpVO2gYCMCNEColOyvUvkzZGUEI1hQ3PfMgi3FIo9tHoBKMw4/wGBLFpu/0ubtmbBXM6/UMOEw=="},{"rawBytes":"MIICRTCCAcygAwIBAgIUeJdY3rV86EdvFmG7L8LJBsyQFYkwCgYIKoZIzj0EAwMwUTELMAkGA1UEBhMCVVMxGzAZBgNVBAoMEk5WSURJQSBDb3Jwb3JhdGlvbjElMCMGA1UEAwwcTlZJRElBIEFnZW50IENhcGFiaWxpdGllcyBDQTAgFw0yNjA0MDEwMDAwMDBaGA85OTk5MTIzMTIzNTk1OVowUTELMAkGA1UEBhMCVVMxGzAZBgNVBAoMEk5WSURJQSBDb3Jwb3JhdGlvbjElMCMGA1UEAwwcTlZJRElBIEFnZW50IENhcGFiaWxpdGllcyBDQTB2MBAGByqGSM49AgEGBSuBBAAiA2IABAYpiXCDjJ9NT2eSDhyHJVSw1Tbze18cGG2F/578oWvHxg23eQAhNRYdq88i1iOshZSO6C29doKui5Xpmo/7Ctw9Sx4PP2RzOmIuOLCuTdNtKcTRwi4GEsd5BAFvWj42M6NjMGEwHQYDVR0OBBYEFItnoAjjfuCEUvzyvWyI2vOGvwPjMB8GA1UdIwQYMBaAFItnoAjjfuCEUvzyvWyI2vOGvwPjMA8GA1UdEwEB/wQFMAMBAf8wDgYDVR0PAQH/BAQDAgEGMAoGCCqGSM49BAMDA2cAMGQCMCwtAjWLaNwgGWNCgdyNoTyvNhqWRECRJV2r3+7w8g0PL6NHLOsbkgE09BH95h8XlgIwTaQmbbUh2ChAJ5TA1wRiVDnCcvbzHlZl2jM2FcwQQZlk19LOAbyGMRixbu2Ww/rj"}]},"tlogEntries":[]},"dsseEnvelope":{"payload":"ewogICJfdHlwZSI6ICJodHRwczovL2luLXRvdG8uaW8vU3RhdGVtZW50L3YxIiwKICAic3ViamVjdCI6IFsKICAgIHsKICAgICAgIm5hbWUiOiAicHJvdGVpbm1wbm4tbmltIiwKICAgICAgImRpZ2VzdCI6IHsKICAgICAgICAic2hhMjU2IjogIjYxMjk1M2YxNmY2Zjc5Y2Y0ZjQxMmU4NzIyNDg1NjE1ODU2MWQ2ZmQzYWJiNjQ3NDZjZjA4NDdiNDM5MDc0YWMiCiAgICAgIH0KICAgIH0KICBdLAogICJwcmVkaWNhdGVUeXBlIjogImh0dHBzOi8vbW9kZWxfc2lnbmluZy9zaWduYXR1cmUvdjEuMCIsCiAgInByZWRpY2F0ZSI6IHsKICAgICJzZXJpYWxpemF0aW9uIjogewogICAgICAiaGFzaF90eXBlIjogInNoYTI1NiIsCiAgICAgICJpZ25vcmVfcGF0aHMiOiBbCiAgICAgICAgIi5naXRhdHRyaWJ1dGVzIiwKICAgICAgICAiLmdpdGlnbm9yZSIsCiAgICAgICAgIi5naXRodWIiLAogICAgICAgICIuZ2l0IgogICAgICBdLAogICAgICAibWV0aG9kIjogImZpbGVzIiwKICAgICAgImFsbG93X3N5bWxpbmtzIjogZmFsc2UKICAgIH0sCiAgICAicmVzb3VyY2VzIjogWwogICAgICB7CiAgICAgICAgIm5hbWUiOiAiQkVOQ0hNQVJLLm1kIiwKICAgICAgICAiYWxnb3JpdGhtIjogInNoYTI1NiIsCiAgICAgICAgImRpZ2VzdCI6ICIzZTY3YTQwMTFiM2YxNzQyZDRiYzRkNzA5NWJkYzYyNjJkMGY0MGVhZGQzZjBiOGU0ZGQ1ODE3YWFkM2RlZmY2IgogICAgICB9LAogICAgICB7CiAgICAgICAgIm5hbWUiOiAiU0tJTEwubWQiLAogICAgICAgICJhbGdvcml0aG0iOiAic2hhMjU2IiwKICAgICAgICAiZGlnZXN0IjogIjE3MjllODYyZjVjZWU5YjcwODQxNjBlMGIzZGE5ZGRkMzhmNTI2MzUzMmE1ZGQxNjVjZDJmMzQyYTY2OWU1YTUiCiAgICAgIH0sCiAgICAgIHsKICAgICAgICAibmFtZSI6ICJjb25maWcvc2tpbGxzcGVjdG9yLWJhc2VsaW5lLnltbCIsCiAgICAgICAgImFsZ29yaXRobSI6ICJzaGEyNTYiLAogICAgICAgICJkaWdlc3QiOiAiMDZkMzRjN2U2MTkzZWQxNmRmOGI3NmVmN2MxZmRlNDlmYTQyMDkzOGJiYTU5MWU0MDM2MzgxODZlODQ0YzFmNyIKICAgICAgfSwKICAgICAgewogICAgICAgICJuYW1lIjogImV2YWxzL1JFQURNRS5tZCIsCiAgICAgICAgImFsZ29yaXRobSI6ICJzaGEyNTYiLAogICAgICAgICJkaWdlc3QiOiAiMjEzMTBiYTgyZjQ0MDg5MTc1NzgwMmVmOWNjZTRkNmU3NDVlNDVjZjNhNDhlMTMxOWYzYTQxNjM2NDU4ZWM3MiIKICAgICAgfSwKICAgICAgewogICAgICAgICJuYW1lIjogImV2YWxzL2NvbmZpZy55bWwiLAogICAgICAgICJhbGdvcml0aG0iOiAic2hhMjU2IiwKICAgICAgICAiZGlnZXN0IjogImMzNjE1Njc3MDYzMzNmNjg3MWY3OGQ2MzNjZDU0YzI1MzZmOTQwZmM3NzI1NGJjZGUyOWRlZjZmMWRjM2E4ZjkiCiAgICAgIH0sCiAgICAgIHsKICAgICAgICAibmFtZSI6ICJldmFscy9ldmFscy5qc29uIiwKICAgICAgICAiYWxnb3JpdGhtIjogInNoYTI1NiIsCiAgICAgICAgImRpZ2VzdCI6ICI1ZTZmY2U1MzYyMDA4NGVlZGM3NTRmNWU1NzVkNWRhZDM5ZjYzMDBiNDdlZGM1YWRiY2EzNjNhYzAzYTY5NTc1IgogICAgICB9LAogICAgICB7CiAgICAgICAgIm5hbWUiOiAiZXZhbHMvZmlsZXMvMVI0Mi5wZGIiLAogICAgICAgICJhbGdvcml0aG0iOiAic2hhMjU2IiwKICAgICAgICAiZGlnZXN0IjogImY1Mzc3MWJhNjU0ODJhYzc2ZGU3MjFhZDdkMDJmZTc3MTA4NzU0NTU3MzEwMGFhNmU5YTc0NGI0Y2M3YmM5YTciCiAgICAgIH0sCiAgICAgIHsKICAgICAgICAibmFtZSI6ICJldmFscy9oYXJib3IvZGF0YXNldC50b21sIiwKICAgICAgICAiYWxnb3JpdGhtIjogInNoYTI1NiIsCiAgICAgICAgImRpZ2VzdCI6ICIxNDNiOThkMWQxMjQ2NGJlOWI4NzBmNTdmYjE4NjRkZjM5MzI2NjIwMmZhZDViMmUxN2ZhNmQ2Y2UwYzA0ZmRmIgogICAgICB9LAogICAgICB7CiAgICAgICAgIm5hbWUiOiAiZXZhbHMvaGFyYm9yL3Byb3RlaW5tcG5uLWxvY2FsLWRlc2lnbi9lbnZpcm9ubWVudC9Eb2NrZXJmaWxlIiwKICAgICAgICAiYWxnb3JpdGhtIjogInNoYTI1NiIsCiAgICAgICAgImRpZ2VzdCI6ICJiYmIwYTM4MjFmYjEzYzk5OTI3ZmU1YTRhNTRiMDAxODlmMWI4ZTdjYzRmYWUzODIyMmZjMDc5NjZkY2Y2YWZkIgogICAgICB9LAogICAgICB7CiAgICAgICAgIm5hbWUiOiAiZXZhbHMvaGFyYm9yL3Byb3RlaW5tcG5uLWxvY2FsLWRlc2lnbi9lbnZpcm9ubWVudC9pbnB1dC8xUjQyLnBkYiIsCiAgICAgICAgImFsZ29yaXRobSI6ICJzaGEyNTYiLAogICAgICAgICJkaWdlc3QiOiAiZjUzNzcxYmE2NTQ4MmFjNzZkZTcyMWFkN2QwMmZlNzcxMDg3NTQ1NTczMTAwYWE2ZTlhNzQ0YjRjYzdiYzlhNyIKICAgICAgfSwKICAgICAgewogICAgICAgICJuYW1lIjogImV2YWxzL2hhcmJvci9wcm90ZWlubXBubi1sb2NhbC1kZXNpZ24vaW5zdHJ1Y3Rpb24ubWQiLAogICAgICAgICJhbGdvcml0aG0iOiAic2hhMjU2IiwKICAgICAgICAiZGlnZXN0IjogIjcwZTQ4MDYwNDEwNTRlNjI5NzY0MDExNmJkMWFlM2UxYjY1OGQyZmI3OTZhMzU3ZTZiNWU3MzM2NzkyNGEyMmQiCiAgICAgIH0sCiAgICAgIHsKICAgICAgICAibmFtZSI6ICJldmFscy9oYXJib3IvcHJvdGVpbm1wbm4tbG9jYWwtZGVzaWduL3Rhc2sudG9tbCIsCiAgICAgICAgImFsZ29yaXRobSI6ICJzaGEyNTYiLAogICAgICAgICJkaWdlc3QiOiAiNTY3ZmEyMTI3NDdiMmMyZGI2ZGMxZjU2MTdkMzdiYjk4MTg1ZmU1Y2U3YjM2MjA5MWIwMzNkYTI5ZDJjN2Q2MSIKICAgICAgfSwKICAgICAgewogICAgICAgICJuYW1lIjogImV2YWxzL2hhcmJvci9wcm90ZWlubXBubi1sb2NhbC1kZXNpZ24vdGVzdHMvZ3JhZGVyLnB5IiwKICAgICAgICAiYWxnb3JpdGhtIjogInNoYTI1NiIsCiAgICAgICAgImRpZ2VzdCI6ICJjYWYyMTJjZmZlYWE3ODUzMjg0YjFhZTkzZTRhYWI2YmZhZDZlODczMTEwNTllY2UxMGE1ZWNkMzYzNTUzYjA3IgogICAgICB9LAogICAgICB7CiAgICAgICAgIm5hbWUiOiAiZXZhbHMvaGFyYm9yL3Byb3RlaW5tcG5uLWxvY2FsLWRlc2lnbi90ZXN0cy90ZXN0LnNoIiwKICAgICAgICAiYWxnb3JpdGhtIjogInNoYTI1NiIsCiAgICAgICAgImRpZ2VzdCI6ICJhZTFmMzVhNGVkMDExMTNjMGM1YTcwZDRkOTE3NTIwZGVmZTBjZGI1OTFiY2NhODc1ZDg0ZGFmMTcxNWMwYzZmIgogICAgICB9LAogICAgICB7CiAgICAgICAgIm5hbWUiOiAiZXZhbHMvdHJpZ2dlcl9ldmFscy5qc29uIiwKICAgICAgICAiYWxnb3JpdGhtIjogInNoYTI1NiIsCiAgICAgICAgImRpZ2VzdCI6ICJlY2QyMzMwYzRmMDUzYjNjNDdjZmJkMjIxNzUzMmNkNThhMDczNGVhYjY5OWM1NzQzZmE2NmE2ZjBmY2IwNWVlIgogICAgICB9LAogICAgICB7CiAgICAgICAgIm5hbWUiOiAicmVmZXJlbmNlcy9hcGkubWQiLAogICAgICAgICJhbGdvcml0aG0iOiAic2hhMjU2IiwKICAgICAgICAiZGlnZXN0IjogIjRjNzkwMWI1OTIzNTMxMWYxNzRlZmYyM2E4OWZjZmZkMmRmOGJlMDhlYWYzMWM5NGJlNjE4OTExNzRlYjI4ODEiCiAgICAgIH0sCiAgICAgIHsKICAgICAgICAibmFtZSI6ICJyZWZlcmVuY2VzL2V4YW1wbGVzLm1kIiwKICAgICAgICAiYWxnb3JpdGhtIjogInNoYTI1NiIsCiAgICAgICAgImRpZ2VzdCI6ICIzYjg4YjdjNTFjY2M5ZDdjZDY3NzgxNGI3ZWU0ZjYyZTFmMTdmMTRiNDc4YzliYTU3N2M4ZWJiMzIwN2I5ZjgwIgogICAgICB9LAogICAgICB7CiAgICAgICAgIm5hbWUiOiAicmVmZXJlbmNlcy9wYXJhbWV0ZXJzLm1kIiwKICAgICAgICAiYWxnb3JpdGhtIjogInNoYTI1NiIsCiAgICAgICAgImRpZ2VzdCI6ICI2NTI4ZmRjZWQ3N2I3NWRkYjA2ZjNkNDM0Y2MzMDY1NTgzNjgyYjllMmQ4ZDQ4MGU0MjFmYmRiMWZhNDlkZWI3IgogICAgICB9LAogICAgICB7CiAgICAgICAgIm5hbWUiOiAicmVmZXJlbmNlcy9zY2llbmNlLm1kIiwKICAgICAgICAiYWxnb3JpdGhtIjogInNoYTI1NiIsCiAgICAgICAgImRpZ2VzdCI6ICI3NDBmZDQ1YzUyNmY1YTFhYmZmZmIzMzdhM2NjMzBkM2ZiYjRmN2MwMDE1NmY4ZTVmMjI0YzlmZWRiNTVhNWQyIgogICAgICB9LAogICAgICB7CiAgICAgICAgIm5hbWUiOiAicmVmZXJlbmNlcy92YWxpZGF0aW9uLm1kIiwKICAgICAgICAiYWxnb3JpdGhtIjogInNoYTI1NiIsCiAgICAgICAgImRpZ2VzdCI6ICI3ZDZlNjhhZjhjZDBjYzQwMjYzYjNlMGY5Y2E0NThjZmRiNzRlYmU1M2MzZmE5MjI3Y2Q1YTk0ZDlmM2M0OWE3IgogICAgICB9LAogICAgICB7CiAgICAgICAgIm5hbWUiOiAic2NyaXB0cy9kZXNpZ24ucHkiLAogICAgICAgICJhbGdvcml0aG0iOiAic2hhMjU2IiwKICAgICAgICAiZGlnZXN0IjogIjJiMDdhM2RmNWY0OWM5Y2NlYzdjYThjYzU1ODQ0YWNkMmMyZjVlZjRiM2MxM2I3M2RhZWJlODZkYWVlMTM0NDkiCiAgICAgIH0sCiAgICAgIHsKICAgICAgICAibmFtZSI6ICJza2lsbC1jYXJkLm1kIiwKICAgICAgICAiYWxnb3JpdGhtIjogInNoYTI1NiIsCiAgICAgICAgImRpZ2VzdCI6ICIyZDZiZTk3M2RhNDBmN2YwMDRjMWIxZGZlZmM3ZDRjZmJiMzRjYzMxMTU5MTcwNTAyNDk0YWJhYzcxMDU3ZTlkIgogICAgICB9CiAgICBdCiAgfQp9","payloadType":"application/vnd.in-toto+json","signatures":[{"sig":"MGUCMDT8u1PUcyHc09la3sbmmGvCQjzyEFk8zosX0elBaspBP2jDf3XAEW18oEmSbUd6uwIxANEp+wtFOmWbpLNNaSlPCUeZm0yb7M4aRlTfC8mHrfIuyTRvKfb12gc3jv0qvtAY9Q==","keyid":""}]}}
\ No newline at end of file
diff --git a/tests/test_proteinmpnn_design.py b/tests/test_proteinmpnn_design.py
new file mode 100644
index 0000000..1ae71fc
--- /dev/null
+++ b/tests/test_proteinmpnn_design.py
@@ -0,0 +1,245 @@
+"""ProteinMPNN client contracts using synthetic data; no GPU or API key required."""
+
+import argparse
+from concurrent.futures import ThreadPoolExecutor
+from contextlib import redirect_stdout
+from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
+import importlib.util
+import io
+import json
+import os
+from pathlib import Path
+import subprocess
+import sys
+import tempfile
+import threading
+import unittest
+from unittest.mock import patch
+
+import requests
+
+
+SCRIPT = Path(__file__).resolve().parents[1] / "nim-skills" / "proteinmpnn-nim" / "scripts" / "design.py"
+spec = importlib.util.spec_from_file_location("proteinmpnn_design", SCRIPT)
+client = importlib.util.module_from_spec(spec)
+spec.loader.exec_module(client)
+
+
+def response_data(native=True):
+ prefix = ">input, score=9.1, fixed_chains=[], designed_chains=['A']\nAAA\n" if native else ""
+ return {"mfasta": prefix + ">T=0.1, sample=1, score=1.2\nAC\nD\n"
+ ">T=0.1, sample=2, score=0.8\nEFG\n", "scores": [1.2, 0.8]}
+
+
+def response(data=None, status=200):
+ value = requests.Response()
+ value.status_code = status
+ value._content = json.dumps(response_data() if data is None else data).encode()
+ return value
+
+
+class ResultTests(unittest.TestCase):
+ def test_native_is_excluded_without_shifting_design_scores(self):
+ summary = client.design_results(response_data(), 2)
+ self.assertEqual((summary["generated_count"], summary["native_count"]), (2, 1))
+ self.assertEqual([(row["sequence"], row["score"]) for row in summary["sequences"]],
+ [("ACD", 1.2), ("EFG", 0.8)])
+
+ def test_score_array_can_include_the_native_record(self):
+ data = response_data()
+ data["scores"] = [9.1, 1.2, 0.8]
+ summary = client.design_results(data, 2)
+ self.assertEqual(summary["scores"], [1.2, 0.8])
+
+ def test_native_only_response_cannot_count_as_one_design(self):
+ data = {"mfasta": ">input, score=9.1, fixed_chains=[], designed_chains=['A']\nAAA\n", "scores": [9.1]}
+ with self.assertRaises(ValueError):
+ client.design_results(data, 1)
+
+ def test_no_native_record_keeps_every_design_and_chain_separator(self):
+ data = response_data(native=False)
+ data["mfasta"] = data["mfasta"].replace("EFG", "EFG/ACD")
+ summary = client.design_results(data, 2)
+ self.assertEqual(summary["native_count"], 0)
+ self.assertEqual(summary["sequences"][1]["sequence"], "EFG/ACD")
+ self.assertEqual(summary["scores"], [1.2, 0.8])
+
+ def test_header_fallback_uses_design_score_not_global_or_native_score(self):
+ data = response_data()
+ del data["scores"]
+ data["mfasta"] = data["mfasta"].replace("sample=1, score=1.2", "sample=1, global_score=7.3, score=1.2")
+ summary = client.design_results(data, 2)
+ self.assertEqual(summary["scores"], [1.2, 0.8])
+ self.assertEqual(summary["score_source"], "mfasta.header.score")
+
+ def test_invalid_responses_never_invent_or_silently_truncate_results(self):
+ bad = [{}, [], {"mfasta": ""}, {"mfasta": "ACD"},
+ {"mfasta": ">design\nACD\n", "scores": [1, 2]},
+ {**response_data(), "scores": [1]},
+ {**response_data(), "scores": [1, 2, 3, 4]},
+ {**response_data(), "scores": [True, 0.8]},
+ {**response_data(), "scores": [float("nan"), 0.8]},
+ {**response_data(), "scores": "1.2,0.8"},
+ {"mfasta": ">sample=1\nACD\n>sample=2\nEFG\n"},
+ {"mfasta": ">sample=1\nAC?\n>sample=2\nEFG\n", "scores": [1, 2]},
+ {"mfasta": ">sample=1\n\n>sample=2\nEFG\n", "scores": [1, 2]},
+ {"mfasta": ">sample=1\nACD\n>sample=2\nEFG\n>sample=3\nHIK\n", "scores": [1, 2, 3]}]
+ for data in bad:
+ with self.subTest(data=data), self.assertRaises(ValueError):
+ client.design_results(data, 2)
+
+
+class ClientTests(unittest.TestCase):
+ def setUp(self):
+ self.temp = tempfile.TemporaryDirectory()
+ self.addCleanup(self.temp.cleanup)
+ self.root = Path(self.temp.name)
+ self.output = self.root / "run"
+ self.pdb = self.root / "user-backbone.pdb"
+ self.pdb.write_text("HEADER synthetic test fixture\nATOM 1 CA ALA A 1 0.000 0.000 0.000\nEND\n")
+ self.args = argparse.Namespace(pdb=self.pdb, mode="hosted", num_sequences=2,
+ temperature=0.1, chains=None, omit_aas=None, soluble=False,
+ ca_only=False, seed=None, timeout=30, output_dir=self.output)
+
+ def test_hosted_execution_saves_and_prints_real_results_without_key(self):
+ self.args.chains = ["A"]
+ self.args.omit_aas = ["C"]
+ self.args.seed = 7
+ stdout = io.StringIO()
+ with patch.dict(os.environ, {"NGC_API_KEY": "test-secret", "NIM_API_MODE": "local"}, clear=True), \
+ patch.object(client.requests, "post", return_value=response()) as post, redirect_stdout(stdout):
+ summary = client.design(self.args)
+ self.assertEqual(post.call_count, 1)
+ self.assertEqual(post.call_args.args[0], client.HOSTED_URL)
+ self.assertEqual(post.call_args.kwargs["headers"]["Authorization"], "Bearer test-secret")
+ self.assertFalse(post.call_args.kwargs["allow_redirects"])
+ self.assertEqual(post.call_args.kwargs["timeout"], (10, 30))
+ saved = json.loads((self.output / "request.json").read_text())
+ self.assertEqual(saved, post.call_args.kwargs["json"])
+ self.assertEqual(saved["input_pdb"], self.pdb.read_text())
+ self.assertEqual((saved["num_seq_per_target"], saved["sampling_temp"], saved["random_seed"]), (2, [0.1], 7))
+ self.assertEqual((saved["input_pdb_chains"], saved["omit_AAs"]), (["A"], ["C"]))
+ self.assertEqual((self.output / "designed_sequences.fa").read_text(), response_data()["mfasta"])
+ self.assertEqual(json.loads((self.output / "response.json").read_text()), response_data())
+ self.assertEqual(json.loads((self.output / "summary.json").read_text()), summary)
+ self.assertIn(json.dumps(summary, indent=2), stdout.getvalue())
+ self.assertNotIn(self.pdb.read_text(), stdout.getvalue())
+ self.assertNotIn("test-secret", stdout.getvalue())
+ for path in self.output.iterdir():
+ self.assertNotIn(b"test-secret", path.read_bytes())
+
+ def test_failure_keeps_exact_body_without_success_artifacts(self):
+ cases = [response(status=code) for code in (202, 302, 401, 503)]
+ malformed = response()
+ malformed._content = b'{"mfasta":'
+ cases.extend([malformed, response({**response_data(), "scores": [float("nan"), 0.8]}),
+ response({**response_data(), "scores": [1]})])
+ for index, result in enumerate(cases):
+ with self.subTest(index=index):
+ self.args.output_dir = self.root / str(index)
+ stdout = io.StringIO()
+ with patch.dict(os.environ, {"NGC_API_KEY": "test-secret"}, clear=True), \
+ patch.object(client.requests, "post", return_value=result) as post, redirect_stdout(stdout), \
+ self.assertRaises((ValueError, RuntimeError)):
+ client.design(self.args)
+ self.assertEqual(post.call_count, 1)
+ self.assertEqual((self.args.output_dir / "response.raw").read_bytes(), result.content)
+ self.assertFalse((self.args.output_dir / "designed_sequences.fa").exists())
+ self.assertFalse((self.args.output_dir / "summary.json").exists())
+ self.assertNotIn('"status": "completed"', stdout.getvalue())
+
+ def test_crlf_fasta_is_preserved_exactly(self):
+ data = response_data()
+ data["mfasta"] = data["mfasta"].replace("\n", "\r\n")
+ with patch.dict(os.environ, {"NGC_API_KEY": "test-key"}), \
+ patch.object(client.requests, "post", return_value=response(data)), redirect_stdout(io.StringIO()):
+ summary = client.design(self.args)
+ self.assertEqual(summary["scores"], [1.2, 0.8])
+ self.assertEqual((self.output / "designed_sequences.fa").read_bytes(), data["mfasta"].encode())
+
+ def test_requires_mode_and_hosted_key_before_request(self):
+ for mode in [None, "hosted"]:
+ self.args.mode = mode
+ with self.subTest(mode=mode), patch.dict(os.environ, {}, clear=True), \
+ patch.object(client.requests, "post") as post, self.assertRaises(ValueError):
+ client.design(self.args)
+ post.assert_not_called()
+ self.assertFalse(self.output.exists())
+
+ def test_existing_directory_is_not_overwritten_or_resubmitted(self):
+ self.output.mkdir()
+ for existing_file in [False, True]:
+ if existing_file:
+ (self.output / "summary.json").write_text("existing")
+ with self.subTest(existing_file=existing_file), patch.dict(os.environ, {"NGC_API_KEY": "test-key"}), \
+ patch.object(client.requests, "post") as post, self.assertRaises(FileExistsError):
+ client.design(self.args)
+ post.assert_not_called()
+ self.assertEqual((self.output / "summary.json").read_text(), "existing")
+
+ def test_concurrent_runs_make_only_one_request(self):
+ ready = threading.Barrier(2)
+ mkdir = Path.mkdir
+
+ def synchronized_mkdir(path, *args, **kwargs):
+ if path == self.output:
+ ready.wait(timeout=5)
+ return mkdir(path, *args, **kwargs)
+
+ def run(seed):
+ try:
+ return client.design(argparse.Namespace(**{**vars(self.args), "seed": seed}))
+ except FileExistsError:
+ return None
+
+ with patch.dict(os.environ, {"NGC_API_KEY": "test-key"}), patch.object(Path, "mkdir", synchronized_mkdir), \
+ patch.object(client.requests, "post", return_value=response()) as post, redirect_stdout(io.StringIO()):
+ with ThreadPoolExecutor(max_workers=2) as pool:
+ results = list(pool.map(run, [10, 20]))
+ self.assertEqual(sum(result is not None for result in results), 1)
+ self.assertEqual(post.call_count, 1)
+ self.assertEqual(json.loads((self.output / "request.json").read_text()), post.call_args.kwargs["json"])
+
+ def test_cli_local_execution_preserves_actual_scores_and_omits_credentials(self):
+ received = []
+
+ class Handler(BaseHTTPRequestHandler):
+ def do_POST(self):
+ body = self.rfile.read(int(self.headers["Content-Length"]))
+ received.append((self.path, dict(self.headers), json.loads(body)))
+ data = json.dumps(response_data()).encode()
+ self.send_response(200)
+ self.send_header("Content-Type", "application/json")
+ self.send_header("Content-Length", str(len(data)))
+ self.end_headers()
+ self.wfile.write(data)
+
+ def log_message(self, *_):
+ pass
+
+ server = ThreadingHTTPServer(("127.0.0.1", 0), Handler)
+ thread = threading.Thread(target=server.serve_forever, daemon=True)
+ thread.start()
+ try:
+ env = {"PROTEINMPNN_NIM_URL": f"http://127.0.0.1:{server.server_port}", "NGC_API_KEY": "must-not-send"}
+ completed = subprocess.run([sys.executable, str(SCRIPT), "--mode", "local", "--pdb", str(self.pdb),
+ "--num-sequences", "2", "--output-dir", str(self.output)],
+ env=env, capture_output=True, text=True, timeout=15)
+ finally:
+ server.shutdown()
+ server.server_close()
+ thread.join(timeout=2)
+ self.assertEqual(completed.returncode, 0, completed.stderr)
+ self.assertEqual(len(received), 1)
+ self.assertEqual(received[0][0], "/biology/ipd/proteinmpnn/predict")
+ self.assertNotIn("Authorization", received[0][1])
+ self.assertEqual(received[0][2]["input_pdb"], self.pdb.read_text())
+ summary = json.loads((self.output / "summary.json").read_text())
+ self.assertEqual(summary["scores"], [1.2, 0.8])
+ self.assertEqual(summary["generated_count"], 2)
+ self.assertIn('"sequence": "ACD"', completed.stdout)
+ self.assertIn(str(self.output / "designed_sequences.fa"), completed.stdout)
+
+
+if __name__ == "__main__":
+ unittest.main()