Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
132 changes: 67 additions & 65 deletions nim-skills/evo2-nim/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,15 +9,35 @@ allowed-tools: Bash, Read, Write, AskUserQuestion

# Evo 2 NIM

Use Evo 2 for DNA generation and, locally, layer-output extraction. Use this
`SKILL.md` for basic hosted/local use; load supplemental files only when needed:
Use Evo 2 for DNA generation and, locally, layer-output extraction. Load
supplemental files only when needed:

- `references/api.md`: exact schemas, layer names, Docker flags, hardware notes.
- `references/science.md`: genomic use cases, limits, and interpretation.
- `references/parameters.md`: generation/forward parameter effects.
- `references/validation.md`: DNA, probability, timing, and tensor checks.
- `references/examples.md`: compact hosted/local request patterns.

## Instructions

For generation, use `scripts/generate.py` to execute the request, validate the
response, and save its artifacts. Resolve the script path relative to this
skill's directory and choose an output directory in the user's workspace.
Use the user's sequence and requested parameters; the example below is only
a smoke test.

1. Select the requested mode. For hosted generation, go directly to the
generation example; Docker setup and local forward passes are separate tasks.
2. When the user asks to run generation, execute the client and inspect its
exit status and result. Writing a script alone does not complete that request.
3. Report the generated DNA (or its file for long sequences), actual
`elapsed_ms`, sampled-probability summary, seed, and artifact paths from the
successful run. Read the saved response or metrics if any result is unclear.

If the request or validation fails, report the actual failure and any diagnostic
files. Do not replace an unavailable API response with example values. For a
code-only request, provide the command without making an inference call.

## Choose Mode

Honor `NIM_API_MODE` when it is set. Accepted values are `hosted` and `local`.
Expand Down Expand Up @@ -45,6 +65,46 @@ into the container with `-e NGC_API_KEY`. Local inference requests use no
auth header after readiness. Warm-cache key-free startup varies by
image/version and should not be assumed.

## Examples

Normalize prompts before sending. Use A/C/G/T unless ambiguous bases are a
deliberate modeling choice and clearly reported.

For a hosted generation request, run the bundled client with the user's inputs
(the script path below is relative to the skill directory):

```bash
python scripts/generate.py \
--mode hosted \
--sequence ACTGACTGACTGACTG \
--num-tokens 64 --seed 1 \
--temperature 0.7 --top-k 3 --top-p 0.0 \
--output-dir /path/to/workspace/evo2-output
```

For an already-ready local NIM, use `--mode local`; the client resolves
`EVO2_NIM_URL` and sends no Authorization header. It never switches endpoints
after a failed request. Set `--timeout` for a longer read if the user requests
a larger generation; failed requests are not automatically resubmitted.

The client saves `request.json`, the actual `response.json`, `generated.fasta`,
and `metrics.json` in the chosen output directory. It also saves the exact
response body in `response.raw` before checking HTTP status or parsing JSON,
so diagnostics survive malformed JSON and non-finite probability/timing values.
It validates the requested number of generated bases, A/C/G/T alphabet, finite sampled probabilities in
`[0, 1]`, and nonnegative timing before printing a successful summary. Existing
directories are never reused, even if empty. Choose an output directory that
does not exist; the client creates it atomically so concurrent runs cannot
overwrite each other's artifacts.
The FASTA contains generated bases only, not the input prompt prepended again.

`sampled_probs` is requested by the client and summarized with count/min/max/mean;
the full values stay in the saved response. A missing or malformed probability
array is a validation failure, not permission to invent confidence values.
Only request `enable_logits` in a custom request when needed; logits can make
responses large. See `references/api.md` for custom payloads.
`random_seed` supports development reproducibility, not biological certainty.

## Local Docker Requirements

Evo 2 local deployment requires FP8-capable GPUs. Do not present A100 as
Expand Down Expand Up @@ -100,68 +160,6 @@ If RTX PRO 6000 Blackwell Workstation fails with no Transformer Engine
attention backend, treat it as outside the current validated matrix and rerun
on a documented GPU/runtime.

## DNA Generation

Normalize prompts before sending. Use A/C/G/T unless ambiguous bases are a
deliberate modeling choice and clearly reported.

```python
import json
import os
from pathlib import Path
import requests

def clean_dna(value: str) -> str:
seq = "".join(value.upper().split())
invalid = sorted(set(seq) - set("ACGT"))
if invalid:
raise ValueError(f"Unexpected DNA characters: {''.join(invalid)}")
return seq

prompt = clean_dna("ACTGACTGACTGACTG")
mode = os.getenv("NIM_API_MODE")
if mode is None:
mode = "local" if os.getenv("EVO2_NIM_URL") else "hosted"
if mode not in {"hosted", "local"}:
raise ValueError("NIM_API_MODE must be 'hosted' or 'local'")

nim_url = os.getenv("EVO2_NIM_URL", "http://localhost:8000").rstrip("/")
url = (
"https://health.api.nvidia.com/v1/biology/arc/evo2-40b/generate"
if mode == "hosted" else f"{nim_url}/biology/arc/evo2/generate"
)
headers = {"Content-Type": "application/json"}
if mode == "hosted":
api_key = os.getenv("NGC_API_KEY")
if not api_key:
raise RuntimeError("Set NGC_API_KEY for hosted Evo 2")
headers["Authorization"] = f"Bearer {api_key}"

payload = {
"sequence": prompt,
"num_tokens": 64,
"temperature": 0.7,
"top_k": 3,
"top_p": 0.0,
"random_seed": 1,
"enable_sampled_probs": True,
"enable_elapsed_ms_per_token": True,
}
response = requests.post(url, headers=headers, json=payload, timeout=180)
response.raise_for_status()
result = response.json()
seq = result["sequence"]
if sorted(set(seq.upper()) - set("ACGT")):
raise ValueError("Generated sequence contains unexpected non-ACGT bases")

Path("evo2_generation.json").write_text(json.dumps(result, indent=2) + "\n")
Path("evo2_generated.fa").write_text(f">evo2_generated\n{seq}\n")
print(f"Generated {len(seq)} bases in {result.get('elapsed_ms')} ms")
```

Only request `enable_logits` when needed; logits can make responses large.
`random_seed` supports development reproducibility, not biological certainty.

## Local Forward Pass

Forward returns base64-encoded NPZ tensors.
Expand All @@ -177,8 +175,12 @@ mode = os.getenv("NIM_API_MODE", "local")
if mode != "local":
raise RuntimeError("Evo 2 /forward is available only in local mode")
nim_url = os.getenv("EVO2_NIM_URL", "http://localhost:8000").rstrip("/")
sequence = "ACTGACTGACTG" # Replace with the user's DNA sequence.
sequence = "".join(sequence.upper().split())
if not sequence or set(sequence) - set("ACGT"):
raise ValueError("Expected nonempty A/C/G/T DNA")
payload = {
"sequence": clean_dna("ACTGACTGACTG"),
"sequence": sequence,
"output_layers": ["output_layer", "decoder.layers.3.self_attention"],
}
response = requests.post(
Expand Down
31 changes: 31 additions & 0 deletions nim-skills/evo2-nim/config/skillspector-baseline.yml
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,37 @@
version: 1

rules:
- id: "TT3"
path: "*scripts/generate.py"
reason: >-
Reviewed 2026-09-29 against the client and its request-contract tests.
The reported tainted URL is EVO2_NIM_URL, the user-selected local NIM
address, not a credential. It is validated as an HTTP(S) URL without
embedded credentials, query, or fragment. Local requests have no
Authorization header. Hosted requests use the fixed health.api.nvidia.com
endpoint and send NGC_API_KEY only as its required Bearer authentication;
redirects are disabled. Neither request artifacts nor summaries contain
the key. This exception covers the documented inference client only.
- id: "LP1"
path: "*scripts/generate.py"
reason: >-
Reviewed 2026-09-29 against the client and its request-contract tests.
The agent invokes this client through the declared Bash tool. Its
documented capabilities are reading NGC_API_KEY/NIM_API_MODE/EVO2_NIM_URL,
submitting the user's DNA to the chosen Evo 2 endpoint, and saving outputs
through the declared filesystem tools. The scanner maps Bash only to
shell and requires separate Env/WebFetch tool names even though this
Python client does not call either tool. The environment and network
operations are the explicitly requested inference workflow.
- id: "PE3"
path: "*evals/evals.json"
reason: >-
Reviewed 2026-09-29. The finding is the literal .env filename in a
deferred local-Docker evaluation assertion. It describes the same
user-owned repo-root dotenv setup already reviewed in SKILL.md and
references/api.md; this JSON is test data and does not load credentials.
The active hosted evaluation reads NGC_API_KEY from the environment.
No credential-file access is added to the executable generation client.
- id: "PE3"
path: "*SKILL.md"
reason: >-
Expand Down
114 changes: 19 additions & 95 deletions nim-skills/evo2-nim/evals/evals.json
Original file line number Diff line number Diff line change
Expand Up @@ -7,41 +7,13 @@
"expected_output": "A successfully executed hosted Evo 2 generation request with Bearer auth and exact request fields, plus actual response-derived DNA, sampled probabilities, elapsed timing, and saved JSON and FASTA artifacts.",
"files": [],
"assertions": [
{
"id": "hosted-request-executed",
"description": "Executes the hosted request instead of only writing code",
"check": "Trajectory shows successful execution of the hosted request, and the final response reports actual response-derived DNA, elapsed timing, and saved artifact paths"
},
{
"id": "hosted-endpoint-url",
"description": "Uses the correct hosted Evo2 generation endpoint",
"check": "Script contains 'https://health.api.nvidia.com/v1/biology/arc/evo2-40b/generate'"
},
{
"id": "bearer-auth-header",
"description": "Sets hosted Authorization header from NGC_API_KEY",
"check": "Script contains 'Authorization', 'Bearer', and 'NGC_API_KEY'"
},
{
"id": "generate-fields",
"description": "Uses exact Evo2 generation request field names",
"check": "Script contains 'sequence', 'num_tokens', 'temperature', 'top_k', 'top_p', and 'random_seed'"
},
{
"id": "sampled-probs-field",
"description": "Requests or handles sampled probabilities using the correct field",
"check": "Script contains 'enable_sampled_probs' and handles 'sampled_probs'"
},
{
"id": "dna-validation",
"description": "Validates generated DNA alphabet before claiming success",
"check": "Script checks generated sequence characters against A/C/G/T or an explicitly documented DNA alphabet"
},
{
"id": "save-json-fasta",
"description": "Saves response JSON and generated FASTA output",
"check": "Script writes a .json file and a .fa or .fasta file"
}
"[hosted-request-executed] Executes the hosted request instead of only writing code: Trajectory shows successful execution of the hosted request, and the final response reports actual response-derived DNA, elapsed timing, and saved artifact paths",
"[hosted-endpoint-url] Uses the correct hosted Evo2 generation endpoint: Script contains 'https://health.api.nvidia.com/v1/biology/arc/evo2-40b/generate'",
"[bearer-auth-header] Sets hosted Authorization header from NGC_API_KEY: Script contains 'Authorization', 'Bearer', and 'NGC_API_KEY'",
"[generate-fields] Uses exact Evo2 generation request field names: Script contains 'sequence', 'num_tokens', 'temperature', 'top_k', 'top_p', and 'random_seed'",
"[sampled-probs-field] Requests or handles sampled probabilities using the correct field: Script contains 'enable_sampled_probs' and handles 'sampled_probs'",
"[dna-validation] Validates generated DNA alphabet before claiming success: Script checks generated sequence characters against A/C/G/T or an explicitly documented DNA alphabet",
"[save-json-fasta] Saves response JSON and generated FASTA output: Script writes a .json file and a .fa or .fasta file"
]
}
],
Expand All @@ -53,36 +25,12 @@
"expected_output": "Docker startup and health-check commands using the Evo2 image, repo env contract, FP8-capable GPU guidance, optional NIM_VARIANT, local no-auth inference, and an EVO2_NIM_URL-aware generation request with a localhost fallback.",
"files": [],
"assertions": [
{
"id": "env-contract",
"description": "Uses repo env contract including .env, NGC_API_KEY/NVIDIA_API_KEY fallback, and LOCAL_NIM_CACHE",
"check": "Output mentions '.env', 'NGC_API_KEY', 'NVIDIA_API_KEY', and 'LOCAL_NIM_CACHE'"
},
{
"id": "docker-image",
"description": "Uses the correct Evo2 Docker image",
"check": "Output contains 'nvcr.io/nim/arc/evo2:2'"
},
{
"id": "cache-mount",
"description": "Mounts local cache to the documented container cache target",
"check": "Output contains '/opt/nim/.cache' and 'LOCAL_NIM_CACHE'"
},
{
"id": "variant-gpu-guidance",
"description": "Documents optional NIM_VARIANT=7b, GPU selection, FP8-compatible hardware, and 40B memory requirements",
"check": "Output contains 'NIM_VARIANT', 'NIM_TEST_GPUS', 'FP8', 40B guidance for 2x H100 80GB or 1x H200 141GB, and 7B fallback guidance such as H100, H200, RTX 6000 Ada, or L40S; it must not present A100 as compatible for local Evo2"
},
{
"id": "health-check",
"description": "Polls local readiness before inference",
"check": "Output builds the readiness URL from EVO2_NIM_URL with localhost:8000 only as a fallback"
},
{
"id": "local-no-auth-endpoint",
"description": "Uses local generation endpoint with no Authorization header",
"check": "Script builds the local generation endpoint from EVO2_NIM_URL and does not send 'Authorization' to local inference"
}
"[env-contract] Uses repo env contract including .env, NGC_API_KEY/NVIDIA_API_KEY fallback, and LOCAL_NIM_CACHE: Output mentions '.env', 'NGC_API_KEY', 'NVIDIA_API_KEY', and 'LOCAL_NIM_CACHE'",
"[docker-image] Uses the correct Evo2 Docker image: Output contains 'nvcr.io/nim/arc/evo2:2'",
"[cache-mount] Mounts local cache to the documented container cache target: Output contains '/opt/nim/.cache' and 'LOCAL_NIM_CACHE'",
"[variant-gpu-guidance] Documents optional NIM_VARIANT=7b, GPU selection, FP8-compatible hardware, and 40B memory requirements: Output contains 'NIM_VARIANT', 'NIM_TEST_GPUS', 'FP8', 40B guidance for 2x H100 80GB or 1x H200 141GB, and 7B fallback guidance such as H100, H200, RTX 6000 Ada, or L40S; it must not present A100 as compatible for local Evo2",
"[health-check] Polls local readiness before inference: Output builds the readiness URL from EVO2_NIM_URL with localhost:8000 only as a fallback",
"[local-no-auth-endpoint] Uses local generation endpoint with no Authorization header: Script builds the local generation endpoint from EVO2_NIM_URL and does not send 'Authorization' to local inference"
]
},
{
Expand All @@ -91,36 +39,12 @@
"expected_output": "A local Python script that calls the Evo2 forward endpoint, decodes the base64 NPZ response with numpy, saves it, and prints tensor names, shapes, dtypes, and finite-value summaries.",
"files": [],
"assertions": [
{
"id": "local-forward-endpoint",
"description": "Uses the correct local forward endpoint",
"check": "Script builds the local forward endpoint from EVO2_NIM_URL with localhost:8000 only as a fallback"
},
{
"id": "forward-fields",
"description": "Uses exact forward request fields",
"check": "Script contains 'sequence' and 'output_layers'"
},
{
"id": "requested-layers",
"description": "Requests the user-specified Evo2 layer names",
"check": "Script contains 'output_layer' and 'decoder.layers.3.self_attention'"
},
{
"id": "base64-npz-decode",
"description": "Decodes base64 NPZ response data",
"check": "Script contains 'base64' and 'np.load' or 'numpy.load'"
},
{
"id": "save-npz",
"description": "Saves decoded tensor artifacts as NPZ",
"check": "Script writes a .npz file"
},
{
"id": "finite-summary",
"description": "Checks or reports tensor shape, dtype, and finite numeric values",
"check": "Script reports 'shape' and 'dtype' and checks 'isfinite' or prints numeric summaries"
}
"[local-forward-endpoint] Uses the correct local forward endpoint: Script builds the local forward endpoint from EVO2_NIM_URL with localhost:8000 only as a fallback",
"[forward-fields] Uses exact forward request fields: Script contains 'sequence' and 'output_layers'",
"[requested-layers] Requests the user-specified Evo2 layer names: Script contains 'output_layer' and 'decoder.layers.3.self_attention'",
"[base64-npz-decode] Decodes base64 NPZ response data: Script contains 'base64' and 'np.load' or 'numpy.load'",
"[save-npz] Saves decoded tensor artifacts as NPZ: Script writes a .npz file",
"[finite-summary] Checks or reports tensor shape, dtype, and finite numeric values: Script reports 'shape' and 'dtype' and checks 'isfinite' or prints numeric summaries"
]
}
]
Expand Down
Loading
Loading