Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
134 changes: 87 additions & 47 deletions nim-skills/proteinmpnn-nim/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,15 +1,20 @@
---
name: proteinmpnn-nim
description: >
Run ProteinMPNN inverse folding via NVIDIA NIM to design protein sequences for a target backbone. Use for ProteinMPNN, inverse folding, sequence design, backbone redesign, fixed chains/residues, omit_AAs, sampling temperature, soluble model, hosted NVIDIA API, local Docker, PDB input, and multi-FASTA output.
Run ProteinMPNN inverse folding via NVIDIA NIM to design protein sequences for a target backbone. Sends user-provided PDB files and design parameters to NVIDIA's hosted API, authenticated with an environment API key, or to a user-selected local NIM. Use for sequence design, backbone redesign, fixed chains and residues, omit_AAs, sampling temperature, soluble model, local Docker, and multi-FASTA output.
license: Apache-2.0 AND CC-BY-4.0
compatibility: "requests>=2.28"
compatibility: "Python >=3.10; requests>=2.28"
allowed-tools: Bash, Read, Write, AskUserQuestion
permissions:
- network
- env
---

# ProteinMPNN NIM

Design protein sequences for a supplied backbone PDB. Use this `SKILL.md` for
<!-- nv-carps: dummy edit to trigger NIM skill validation. -->

Design protein sequences for a supplied backbone PDB. Use this guide for
first-pass hosted/local usage; load supplemental files only when needed:

- `references/api.md`: exact endpoints, schemas, Docker flags, response fields.
Expand All @@ -20,15 +25,16 @@ first-pass hosted/local usage; load supplemental files only when needed:

## Choose Mode

Honor an explicitly configured runtime before asking. `NIM_API_MODE=local` selects
Honor the user's explicit mode; otherwise use the configured runtime. `NIM_API_MODE=local` selects
the local service at `PROTEINMPNN_NIM_URL`; the URL defaults to
`http://localhost:8000` for a NIM running in the same host or container. Ask only
when neither the environment nor the user's request makes the mode clear:

> Hosted NVIDIA API or local Docker NIM?

- Hosted: `https://health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict`
- Local: `${PROTEINMPNN_NIM_URL:-http://localhost:8000}/biology/ipd/proteinmpnn/predict`
- Local: append `/biology/ipd/proteinmpnn/predict` to `PROTEINMPNN_NIM_URL`
(default base URL: `http://localhost:8000`).

Local inference paths do not include `/v1/`. Hosted requests use `Authorization: Bearer $NGC_API_KEY`. Supported local Docker
startup uses `NGC_API_KEY` (or `NVIDIA_API_KEY` via the preflight) for
Expand All @@ -37,6 +43,25 @@ into the container with `-e NGC_API_KEY`. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.

## Data Transfer and Authorization

Before a hosted request, tell the user that the **entire PDB file and design
parameters will be uploaded to NVIDIA's hosted API** at the endpoint above.
Proceed if the user has explicitly requested hosted processing of that PDB or
already approved the transfer; otherwise ask for confirmation before submitting.
For confidential structures, recommend a local NIM in the user's approved
environment. A configured local URL may point to another machine; use only the
configured or user-selected destination. Do not switch from local to hosted processing without
the user's authorization.

The client reads `NIM_API_MODE`, `PROTEINMPNN_NIM_URL`, and, for hosted mode,
`NGC_API_KEY` from the environment. It sends the key only in the HTTPS
Authorization header to the hosted endpoint; local inference sends no key.
Keep credentials out of logs and saved artifacts. The output directory contains
the full input PDB in `request.json` and the returned sequences and scores, so
use a location appropriate for the input's sensitivity. See
[`references/api.md`](references/api.md) for endpoint and data-handling details.

## Local Docker

For local setup, run the full sequence — env preflight, `docker login`,
Expand All @@ -56,53 +81,68 @@ proteinmpnn_nim_url="${PROTEINMPNN_NIM_URL:-http://localhost:8000}"
until curl -sf "${proteinmpnn_nim_url%/}/v1/health/ready"; do sleep 5; done
```

## Request Pattern

Read PDB content inline; do not send only a file path.

```python
import os
from pathlib import Path
import requests

HOSTED = os.getenv("NIM_API_MODE", "hosted").strip().lower() != "local"
pdb_content = Path("1R42.pdb").read_text()
nim_url = os.getenv("PROTEINMPNN_NIM_URL", "http://localhost:8000").rstrip("/")
url = (
"https://health.api.nvidia.com/v1/biology/ipd/proteinmpnn/predict"
if HOSTED else f"{nim_url}/biology/ipd/proteinmpnn/predict"
)
headers = {"Content-Type": "application/json"}
if HOSTED:
headers["Authorization"] = f"Bearer {os.environ['NGC_API_KEY']}"

payload = {
"input_pdb": pdb_content,
"num_seq_per_target": 10,
"sampling_temp": [0.1],
"use_soluble_model": False,
"ca_only": False,
}
response = requests.post(url, headers=headers, json=payload, timeout=300)
response.raise_for_status()
result = response.json()
```
## Instructions

For a request to execute a design, run [`scripts/design.py`](scripts/design.py)
and inspect its results. Writing a request script alone does not complete an
execution request. If the user asks only for code or setup instructions, provide
those without submitting an inference request.

1. Use the user's PDB path and requested sequence count. The client reads the
entire PDB into `input_pdb`; do not replace or truncate the supplied backbone.
2. Select `--mode hosted` or `--mode local` and follow **Data Transfer and
Authorization** above before submitting. Hosted mode uploads the PDB to the
documented NVIDIA endpoint and requires `NGC_API_KEY` in the environment.
Check only whether the key is set; do not print it, dump the environment, or
save authentication headers. Local inference sends no authorization header.
3. Choose a new `--output-dir` for each request. The client reserves it before
submitting, preserves the raw response for diagnostics, and validates the
designed sequence count and score alignment before reporting completion.
4. Read `summary.json` and report the actual results described below. If the
request or validation fails, report the failure and diagnostic path; do not
substitute example sequences or repeatedly resubmit the same request.

Common controls:
## Examples

Run from this skill's directory, or use an absolute path to `scripts/design.py`.
Substitute the user's input path and a new output directory:

```bash
python scripts/design.py --mode hosted \
--pdb /path/to/backbone.pdb --num-sequences 10 \
--temperature 0.1 --output-dir /path/to/new-design-run
```

- Redesign only chain A: `"input_pdb_chains": ["A"]`.
- Exclude amino acids: `"omit_AAs": ["C"]` or `"omit_AAs": ["M"]`.
- Diversity: `"sampling_temp": [0.1, 0.3, 0.5]` (always a list).
- Solubility bias: `"use_soluble_model": True`.
- Candidate count: `num_seq_per_target` is 1-100.
For a running local NIM, use `--mode local`; the client honors
`PROTEINMPNN_NIM_URL`. To design only chain A, exclude cysteine, or request the
soluble model, add `--chains A`, `--omit-aas C`, or `--soluble` respectively.
`--seed` sets `random_seed`; `--ca-only` selects the CA-only model. The helper
uses one temperature per request; run separate output directories for a
temperature sweep. For advanced JSONL controls or a custom batch request,
use [`references/api.md`](references/api.md) and the post-response example in
[`references/examples.md`](references/examples.md).

## Save And Report Output

Save the returned `mfasta` and pair scores only with designed (non-native/WT)
rows, using the snippet in [`references/examples.md`](references/examples.md)
under **Save Multi-FASTA**. Validate promising designs by predicting structures
with Boltz2 or OpenFold3 and comparing them to the target backbone. For
FASTA/score sanity checks, read `references/validation.md`.
The client writes `request.json`, `response.raw`, `response.json`,
`designed_sequences.fa`, and `summary.json` into the requested output directory.
The FASTA preserves the complete returned `mfasta`, including a native/WT entry
when present. The summary contains only designed sequences, each paired with
its actual score, and records whether scores came from the JSON array or FASTA
headers. It is also printed after the artifacts are saved and checked.

In the final response, report:

- The number of **designed** sequences, excluding the native/WT reference.
- Each design's identifier and actual returned score, plus its sequence (for
long sequences, give a clearly labelled preview and link to the full FASTA).
- The saved FASTA and summary paths, and the raw response path for provenance.
- That these are inverse-folding candidates, with no fold-back validation
performed unless it was actually requested and run.

Do not treat a score as proof that a sequence folds or binds. Further validation
with Boltz2 or OpenFold3 is an optional next step. For FASTA/score sanity checks,
read [`references/validation.md`](references/validation.md).

## Limits And Troubleshooting

Expand Down
14 changes: 13 additions & 1 deletion nim-skills/proteinmpnn-nim/config/skillspector-baseline.yml
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
# Audited false-positive suppression, auto-applied by the NVSkills Tier 1
# runner via config/skillspector-baseline.yml. Suppressed findings remain in
# the report JSON marked `suppressed: true` with the reason below.
version: 1
version: 2

rules:
- id: "PE3"
Expand All @@ -16,3 +16,15 @@ rules:
`docker login nvcr.io`. This is first-party, user-facing setup guidance for
the user's own credential file — not credential theft. No SSH keys, cloud
credential stores, or third-party secret files are accessed.
- id: "PE3"
path: "evals/evals.json"
# Match only the literal dotenv token, including PE3's trailing space.
message: ".env "
reason: >-
Reviewed 2026-09-30. The two matches are in deferred Docker evaluation
3's expected output and setup assertion. They describe optional loading
of the user's own repo-root dotenv file for NGC authentication, the same
setup documented in references/api.md. This JSON is evaluation text and
does not read or transmit credentials. The active hosted evaluation and
design client read NGC_API_KEY from the environment. This exception is
limited to the literal dotenv matches in this evaluation file.
37 changes: 37 additions & 0 deletions nim-skills/proteinmpnn-nim/evals/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
# ProteinMPNN evaluations

The default `config.yml` selects the hosted request in `evals.json`. It uses
`NGC_API_KEY`, stages `files/1R42.pdb`, and requires an executed inference request,
saved designed sequences, and response-derived scores.

## Local GPU task

`harbor/proteinmpnn-local-design` is a separate native Harbor task with a custom
grader. It requires the ProteinMPNN NIM image, a GPU, its model download
credentials, and a running NIM server in the task container. A shared Astra
sandbox using the generic agent template ignores the task image and cannot
satisfy its loopback health check.

For a dedicated local GPU evaluation, use a separate checkout and replace
`evals/config.yml` with:

```yaml
schema_version: 1
harbor:
task_source: native_harbor
custom_dockerfile_mode: preserve
base_image_mode: disabled
n_attempts: 1
pass_threshold: 0.8
stop_on_pass: false
n_concurrent: 1
runtime_env:
- NGC_API_KEY
grading:
mode: custom_only
```

Run on a GPU-capable Docker host with `--env-mode docker`, or a dedicated sandbox
template configured with the NIM image, GPU, credentials, and server startup.
The native task's readiness check must pass before an agent starts. Keep the
hosted configuration as the shared CI default.
14 changes: 6 additions & 8 deletions nim-skills/proteinmpnn-nim/evals/config.yml
Original file line number Diff line number Diff line change
@@ -1,13 +1,11 @@
schema_version: 1

# ProteinMPNN is the repository's local-GPU BYOT example. ACES supports one
# task source per skill, so this selects the native task below. The hosted
# evals.json dataset remains available if this is switched back to evals_json
# with aces_default grading.
# The shared Astra sandbox uses a fixed agent image and does not launch the
# native task's GPU NIM image. Use the hosted execution dataset for CI.
# The native local-GPU task is retained under harbor/; see README.md to run it
# in a runtime configured for that image and GPU.
harbor:
task_source: native_harbor
custom_dockerfile_mode: preserve
base_image_mode: disabled
task_source: evals_json
n_attempts: 1
pass_threshold: 0.8
stop_on_pass: false
Expand All @@ -16,4 +14,4 @@ harbor:
- NGC_API_KEY

grading:
mode: custom_only
mode: aces_default
Loading
Loading