Skip to content

Data model for verification methods (follow-up to #101 / #186) #191

Description

@a-moskvin

Context

#101 identified a gap: control tables say what a control requires, not how to check it is implemented.

#186 piloted this as a free-text verification_method column on Agentic_AIUC1.md. Review on #186 found three problems:

  1. Method text is stored per row. The same method repeats across entries and will drift.
  2. Review state and citation sit inside prose (DRAFT — … (NPW C02)), so tools can't read them, and no field records who reviewed a method.
  3. The pilot used a legacy table: the pilot framework AIUC-1 already publishes their own evidence layer.

Goal

This issue proposes a structured data model for verification methods:

  • Hold each verification method once and reuse it across controls and frameworks.
  • Store method–control links, with their own review state.
  • Track provenance, licensing, and review state.
  • Allow both methods from third-party catalogues and methods created in the project.
  • Keep Markdown tables readable.
  • Non-breaking: files and consumers that don't use verification data are unaffected.

Overview

Two new source files, hand-edited through PRs similar to incidents.json, each with a companion schema:

File Holds Schema
data/verification-methods.json Method file: each method once (section 1) data/verification-methods-schema.json
data/verification-links.json Links between methods and controls, at control or row level (section 2) data/verification-links-schema.json

1. Method file: data/verification-methods.json

1.1 File structure

The file holds the method objects (1.2) in a methods array:

{
  "version": "1.0",
  "description": "Verification methods: how to check a control is implemented. Methods with a `source` are adapted from the cited source; the original wording is in `source.text`.",
  "methods": [ { "...": "method objects, see 1.2" } ]
}

1.2 Method object

Field Type Required Description
id string yes Project ID, pattern ^VM-\d{4}$.
name string yes Short title, max 80 characters.
method_type enum yes One of: examine, interview, test (see 1.2).
procedure string yes What the assessor does.
expected_result string yes What a pass looks like, must be observable.
source object / null yes Source the method is derived from; null for a method authored in the project (see 1.4, 1.5).
status enum yes One of: draft, reviewed, deprecated (see 1.6).
reviewed_by string[] yes GitHub handles of SME reviewers, empty for draft.
frequency object optional Recommended minimum frequency (see 1.3).
evidence string optional Artefact(s) the check produces or consumes.
review_date string (date) optional Date of the latest review.
status_note string optional Free-text status context. Required when status is deprecated: the reason, and the replacement ID if there is one.

1.3 method_type values

Based on the assessment methods in NIST SP 800-53A.

Value Definition Example
examine Review, inspect or analyse artefacts: documents, configurations, logs, records Diff the tool registry against the per-agent allowlist
interview Discuss with people responsible for the control to confirm understanding or practice Confirm with reviewers that approvals are made with sufficient context
test Exercise the control under defined conditions and compare actual with expected behaviour Attempt an out-of-scope tool call and confirm denial

1.4 frequency (optional)

A recommended minimum, not a requirement.

"frequency": { "mode": "continuous" }
mode value Definition Example
continuous Automated and always on; results available at any time Behavioural deviation alerting on agent actions
periodic Performed on a fixed schedule Review of the tool registry against allowlists
event_driven Performed when a defined change or event occurs Injection test on model or tool change

frequency is an object, not a plain string, so that details such as interval or triggers can be added later without breaking existing data (section 5).

1.5 Provenance

Provenance is expressed by source:

  • source present: the method is derived from the cited source. The source's original wording is kept in source.text.
  • source: null: the method was originated in the project.

1.5 source object

Field Required Description
name yes Catalogue or standard name
version yes Version the method was taken from
id yes Identifier within the source (e.g. C02)
url yes Public URL of the source
license yes Source licence. Must be compatible with CC BY-SA 4.0.
text yes The source's original wording, kept unchanged

1.6 status

Value Definition Required fields
draft Proposed, not yet reviewed by an SME —
reviewed Accepted by at least one SME non-empty reviewed_by
deprecated Retired: obsolete, rejected in review, replaced, or withdrawn by the source. status_note

Deprecated methods stay in the methods file for consistency, and ID never used for something else.

1.7 Change rules

  • Editorial change (clarity, typo; what is tested and the pass criteria are unchanged): same ID. unmodified becomes adapted.
  • Substantive change (what is tested, or the pass criteria): a new method with a new ID. The old method becomes deprecated, with status_note naming the replacement (e.g. "Replaced by VM-0031"). Its links should be re-pointed in the same PR (2.1).

1.8 Examples

Adapted from an external catalogue:

{
  "id": "VM-0001",
  "name": "Injection test across ingestion paths",
  "method_type": "test",
  "procedure": "Submit a maintained injection payload set through each distinct ingestion path (user input, retrieved documents, tool output).",
  "expected_result": "No payload alters the agent goal or triggers an unauthorised action; all attempts are logged.",
  "source": {
    "name": "NPW Agentic AI Control Catalogue",
    "version": "2.4.0",
    "id": "C02",
    "url": "https://www.newpacificway.com/ai-controls",
    "license": "CC BY-SA 4.0",
    "text": "Injection test executed through each distinct ingestion path."
  },
  "status": "draft",
  "reviewed_by": [],
  "frequency": { "mode": "event_driven" },
  "evidence": "Test report with payload set version and per-path results."
}

Originated in the project:

{
  "id": "VM-0002",
  "name": "Supply-chain scope of third-party adversarial testing",
  "method_type": "examine",
  "procedure": "Review the scope section of the current third-party adversarial test report.",
  "expected_result": "Tools, connectors and MCP servers used by the agent are explicitly in scope.",
  "source": null,
  "status": "draft",
  "reviewed_by": []
}

2. Links file: data/verification-links.json

2.1 Why a separate links file

Links are stored once, in their own file, rather than inside method records or framework registries for the following reasons:

  1. A link needs its own review. "Is this method well written?" and "Does this method verify control X?" are separate judgments.
  2. Separate change lifecycles. Editing a method returns it to draft. Linking a reviewed method to a new control does not touch the method record, so its review stands.
  3. One place for both levels. Control-level and row-level links live in the same table; a row-level link is a control-level link with the risk entry set.
  4. Fewer merge conflicts. Contributors linking the same method to different controls edit independent records, not one shared method object.
  5. Registries stay owned by ingest tooling. scripts/ingest-framework.mjs writes a framework registry whole on re-ingest, so links stored on registry controls would be lost on the next sync.

2.2 Control level and row level

A link is applied either to a control (control level) or to risk–control pair (row level). Both levels are needed:

  • One control often serves several risks. 94% of 3,706 mapping rows use a control that is also mapped to other risks of the same framework.
  • Some methods depend on the risk. Controls that state an outcome ("limit agent actions", "apply a cybersecurity baseline") are verified differently depending on the risk they address. For example, CoSAI WS4-AGT-2.2 "Agent cybersecurity baseline" is mapped to ASI03 (identity: verify authentication), ASI05 (code execution: verify input validation) and ASI07 (inter-agent communication: verify secure channels). These three threats require different verification methods for the same control.
  • Some methods don't depend on the risk. Controls that names a specific mechanism ("sign deployed artefacts") can be verified in the same way for every risk.

Rule for contributors: does the method verify the control the same way regardless of the risk?

  • Yes: a control-level link, entry_id: null.
  • No: a row-level link, entry_id set to the risk entry.

2.3 File structure

The file holds the link objects (2.4) in a links array:

{
  "version": "1.0",
  "links": [
    {
      ...
    },
    {
      ...
    }
  ]
}

2.4 Link object

Field Type Required Description
method_id string yes Must exist in the methods file verification-methods.json
framework string yes Mapped to framework's id in data/frameworks/<id>.json
control_id string yes Mapped to control_id in data/frameworks/<id>.json
entry_id string / null yes null: the link applies to the control. An entry ID from a risk list (LLM01, ASI04, DSGAI07, AST01): the link applies only to that risk–control row.
reviewed_by string[] yes GitHub handles of SMEs who confirmed the method verifies this control or row. Empty means draft.

Combination (method_id, framework, control_id, entry_id) is unique. No separate link ID or status field is needed.

2.5 Link lifecycle

  • A new link starts with empty reviewed_by (draft). It counts as reviewed once an SME adds their handle.
  • A wrong or obsolete link is deleted; Git keeps the history.
  • When a method is deprecated with a replacement, its links are re-pointed to the replacement in the same PR, and their reviewed_by is cleared.
  • When a method is deprecated without a replacement, its links are deleted.
  • Editing a method does not change its links: the method's review covers its text, and the link's review covers its applicability.

3. Validation rules (scripts/validate.js)

Methods file

  1. The methods file validates against verification-methods-schema.json, including enum values, additionalProperties: false and the 80-character name limit.
  2. IDs are unique and match ^VM-\d{4}$.
  3. reviewed requires a non-empty reviewed_by.
  4. deprecated requires status_note.
  5. When source is present, all its fields are required.

Links file

  1. The links file validates against verification-links-schema.json.
  2. The key (method_id, framework, control_id, entry_id) is unique.
  3. method_id exists in the methods file. A link to a deprecated method is a warning, not an error.
  4. (framework, control_id) resolves to a control in data/frameworks/.
  5. When entry_id is set, a mapping row exists for that entry, framework and control IDs.

4. Considerations for future development

elements that can be added later without impacting existing data:

Item What it would add Add when
frequency.interval ISO 8601-type coded duration for periodic (e.g. P3M) Users need the cadence, not just the mode
frequency.triggers Events for event_driven (e.g. model_change, tool_change) Users need to filter by trigger
superseded status + superseded_by Machine-readable replacement chain Replacements become common enough that status_note is insufficient
Link status and review_date Richer link lifecycle Link review needs dating or states beyond reviewed/not reviewed

Disclosure: NPW (New Pacific Way Ltd.) is my company, and the NPW Agentic AI Control Catalogue is its catalogue. Methods derived from it will say so in their source field and in the PR.

Activity

  1. emmanuelgjr commented on Oct 5, 2026

    @emmanuelgjr
    Contributor

    Thanks, @a-moskvin. This is a well-reasoned proposal, and it answers all three problems from #186. I checked its claims against current main before replying.

    What holds up

    • Separating method review from link review is the right split. "Is this method well written?" and "does it verify this control?" are different judgments. Keeping them apart is what lets a reviewed method be reused without re-review.
    • Both link levels are justified by the data. Re-measured on main: 93.8% of the 3,771 mapping rows use a control that's also mapped to another risk in the same framework. Your CoSAI WS4-AGT-2.2 example is real: it's mapped to ASI03, ASI05 and ASI07, and also to LLM06 and four DSGAI entries.
    • You're right about the registries, and it corrects my earlier suggestion. On feat(schema): optional verification_method field + Agentic_AIUC1 pilot #186 I proposed holding methods on the registry controls. ingest-framework.mjs writes the whole registry on re-ingest, so they'd be wiped on the next sync. A separate links file is better.
    • Deprecating rather than reusing method IDs matches a lesson this repo learned the hard way: several retired ASVS and ATLAS identifiers now name different requirements upstream.
    • examine / interview / test from SP 800-53A is a good, recognisable vocabulary. Thanks for the disclosure, too.

    Before implementation

    1. Allow citing a source whose licence isn't CC BY-SA-compatible. As written, source.license must be compatible and source.text is required verbatim. That rules out the case from feat(schema): optional verification_method field + Agentic_AIUC1 pilot #186: pointing at AIUC-1's own evidence (its terms are proprietary) or at ISO. Proposal: make source.text optional, and forbidden when the licence isn't compatible, so a method can cite a source by id + url without reproducing it. A validator rule enforces that. It also lets a link point at a framework's canonical verification instead of an alternative to it.

    2. Make the framework key explicit. §2.4 says the registry id (aiuc-1), but mapping rows are keyed by display name (AIUC-1). The id is the better key, because display names change: ASVS was just renamed from 4.0.3 to 5.0.0. Rule 10 then needs to state that it resolves id → registry name before matching the row.

    3. Keep review states from overstating each other. Two validator warnings:

      • a link with a non-empty reviewed_by whose method is still draft;
      • a row-level link on a mapping row that is itself unreviewed.

      Otherwise a consumer could show "reviewed" on top of unreviewed data. Proposal: reviewed_by follows the existing schema v2 convention (named humans, per docs/SCHEMA_V2_MIGRATION.md) rather than introducing GitHub handles, so all three review fields mean the same thing.

    4. Scope the consumers. The proposal says what's stored but not what reads it. Suggest v1 = data, schemas and validator only: no generate.js, export or webapp changes. The webapp is frozen except by maintainer approval, and OLIR has no field for this anyway. Consumers can follow once the data exists.

    5. Merge conflicts. Parallel PRs that each append to one links array will conflict at its end. Proposal: one links file per framework (data/verification-links/<framework-id>.json), or a single file with a sort order the validator enforces.

    6. Small consistency fixes in the text:

      • there are two §1.5 headings;
      • §1.2 points to "see 1.2" / "see 1.3" for method_type and frequency, which live in §1.3 and §1.4;
      • §1.7 changes unmodified to adapted, but neither value is defined anywhere.

    Suggested first PR

    Once this settles:

    • the two schemas, the validator rules (with tests) and the link files;
    • a pilot of up to 10 methods, all draft, on one framework with no verification layer of its own. Your choice of framework; name it here first.

    #186 can then be closed as superseded when that PR lands.

  2. a-moskvin commented on Oct 6, 2026

    @a-moskvin
    Author

    Thanks @emmanuelgjr .

    All six points accepted:

    1. Licences: source.text is optional, so a source under a non-compatible licence (e.g. AIUC-1, ISO) is cited by id + url only. I've left licence compatibility to contributor and reviewer judgment rather than a validator rule: a fixed list of compatible licences would need maintaining across licences and versions. license stays required, so the licence is always visible in review.
    2. Framework key: the registry id. Row-level links resolve id → registry name before matching the mapping row.
    3. Review states: two warnings, for a reviewed link on a draft method and for a reviewed row-level link on an unreviewed row. reviewed_by holds named humans, per SCHEMA_V2_MIGRATION.md.
    4. Consumers: v1 is data, schemas and validator only.
    5. Merge conflicts: one links file per framework (data/verification-links/<framework-id>.json, with framework declared once per file), plus an enforced sort order.
    6. Text fixes: done in docs/VERIFICATION_METHODS.md, which becomes the spec.

    I'd like to split delivery:

    • PR 1 with the schemas, validator rules, tests and docs, and no methods;
    • PR 2 with the pilot of up to 10 draft methods.

    That keeps the model reviewable on its own. I'll name the pilot framework here before PR 2.

  3. a-moskvin commented on Oct 7, 2026

    @a-moskvin
    Author

    Opened PR 1 with the schemas, validator rules, tests and docs, and no methods: #215

    Opened PR 2 #216 : pilot framework: NIST AI RMF 1.0.

    Scope: the pilot links only to the 6 subcategories whose registry titles match AI RMF 1.0 (GV-1.6, GV-1.7, MP-2.3, MP-4.1, MP-5.1, MS-2.6). The other 5 mapped subcategories have confusing labelling in the registry; I'll open a separate issue with the details.

    10 draft methods: 6 control-level and 4 row-level links (ASI01, ASI02, ASI03, ASI10); 8 adapted from the NPW catalogue v2.4.0 and 2 project-authored (source: null). All draft, none reviewed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions