AI-powered synthetic data generation pipeline built on the Strands Agents SDK. Produces realistic PDF documents from JSON schemas, validates them through multi-stage critique loops, and optionally applies image augmentation to simulate real-world scanning/faxing artifacts. Each document is paired with a ground-truth JSON label, so the output is a ready-made benchmark set. SEED also generates structured (tabular) data — CSV, Parquet, Excel, or JSON, one file per entity — from the same planned schema.
Designed for building evaluation datasets for document understanding systems (OCR, Key Information Extraction (KIE), document classification) and for tabular datasets to exercise data pipelines, analytics, and ML workflows. Use it from the command line (seed-data) or the typed Python API (seed_data.Generator).
Not a production-ready solution. This asset represents a proof-of-value for the services included and is not intended as a production-ready solution. You must determine how the AWS Shared Responsibility Model applies to your specific use case and implement the controls needed to achieve your desired security outcomes. AWS offers a broad set of security tools and configurations to enable our customers.
Ultimately it is your responsibility as the developer of a full-stack application to ensure all of its aspects are secure. We provide security best practices in the repository documentation and a secure baseline, but Amazon holds no responsibility for the security of applications built from this tool.
⚡ New here? Read the Quickstart — install from PyPI and generate your first document in a few minutes.
Install from PyPI (the default renderer is pure Python — nothing else to set up):
pip install seed-data
# Configure AWS credentials with Bedrock access
export AWS_PROFILE=your-profile-name
# Copy the built-in schema library into a local, editable folder
seed-data clone-schema-library ./schemas
# Generate a single document (PDF + ground-truth JSON) into ./output
seed-data --schema-dir fcc-invoice --output ./output--schema-dir accepts a bundled schema name (like fcc-invoice) or a path
to a schema directory, so the command above works with or without the clone step.
Browse the schema library on GitHub: awslabs/…/schemas. See the full Quickstart for batches, augmentation, and model selection.
Working from a repo clone (uv)?
# Install (creates the venv and installs from uv.lock)
uv sync
export AWS_PROFILE=your-profile-name
# Generate a single FCC invoice
uv run seed-data --schema-dir fcc-invoice
# Generate a diverse batch of 4 invoices with augmentation
uv run seed-data --schema-dir fcc-invoice \
--count 4 --augment \
--scenario "FCC broadcast invoices for packaged food companies in the midwest"Everything is available from both the command line and Python; they run the same pipeline and produce the same artifacts.
Command line:
seed-data --schema-dir invoice --scenario "Midwest food-distributor invoice" # single
seed-data --schema-dir fcc-invoice --count 10 --scenario "Local TV stations" # batch
seed-data packet lending-package --scenario "First-time homebuyer in Portland" # packet
seed-data infer-schema ./samples/*.pdf --name invoice --output ./schemas/invoice # schema from real docs
seed-data plan "Customers and their orders" --output ./schema.json # any input -> schema
seed-data generate-structured ./schema.json --rows 500 --format parquet # structured (tabular)
seed-data plan-and-generate "Customers and their orders" --output structured --rows 500 # plan + generatePython:
from seed_data import Generator, ModelConfig
gen = Generator(models=ModelConfig(doc="gpt-oss", critic="haiku"), threshold=5)
doc = gen.generate("invoice", scenario="Midwest food-distributor invoice")
batch = gen.generate_batch("fcc-invoice", count=10, scenario="Local TV stations")
packet = gen.generate_packet("lending-package", scenario="First-time homebuyer in Portland")
schema = gen.infer_schema("./samples/*.pdf", name="invoice") # reverse-engineer a schema from real docs
ing = gen.plan("Customers and their orders") # any input -> InferredSchema
table = gen.generate_structured(ing, rows=500) # structured (tabular) data
result = gen.plan_and_generate("Customers and their orders") # plan + generate, one callSchema from documents — the inverse of generation. Point SEED at real example
documents (PDF/PNG/JPEG, local or s3://) and a vision model reverse-engineers a
schema you can generate unlimited synthetic look-alikes from. Works for a single
document type, or splits one concatenated multi-document PDF into a per-type packet.
See Schema from Documents.
Full docs: CLI Usage · Python API Usage.
Alongside documents, SEED generates structured (tabular) datasets. A run writes
one file per entity into the output directory, named from the lowercased
entity name with spaces replaced by underscores: <entity>.csv, .json,
.xlsx, or .parquet. A schema with entities Customer and Order at
--format csv produces output/customer.csv and output/order.csv.
Structured generation is an opt-in extra — pandas and the file-format engines
(openpyxl for .xlsx, pyarrow for .parquet) are kept out of the base install so
pip install seed-data stays lean for document-only users:
pip install "seed-data[structured]"The document pipeline never needs the extra. Structured commands raise a clear
ImportError pointing at it if it is missing. Every output format works with the
extra installed — no separate engine install.
Planning is the shared front door. seed-data plan takes free text, example
data files (CSV/JSON/XLSX), formal schemas (JSON Schema, SQL DDL), documents
(PDF/PNG/JPEG, read with a vision model), ERD diagrams — or several of those
together — and writes one unified InferredSchema JSON. The same planned
schema can drive either modality: pass it to generate-structured for tables,
or to generate-documents for PDFs.
# 1. Plan a schema from anything
seed-data plan "Customers with orders and line items" ./samples/orders.csv \
--name retail --output ./schema.json
# 2. Generate tables from that schema
seed-data generate-structured ./schema.json --rows 500 --format csv --output ./output
# ...or generate documents from the very same schema
seed-data generate-documents ./schema.json --entity Order --count 3seed-data plan-and-generate does planning + generation in one shot, with no intermediate schema
file. Note that on plan-and-generate — and only on plan-and-generate — --output selects the modality
and --output-dir selects the path:
seed-data plan-and-generate "Customers with orders and line items" ./samples/orders.csv \
--output structured --rows 500 --format parquet \
--output-dir ./output --save-schema ./schema.jsonFrom Python, the same two steps with Generator:
from seed_data import Generator
gen = Generator(output_dir="./output")
schema = gen.plan("Customers with orders and line items", "./samples/orders.csv",
name="retail")
result = gen.generate_structured(schema, rows=500, format="csv")
print(result.success) # True when files were written
print(result.output_paths) # ['./output/customer.csv', './output/order.csv', ...]
print(result.format) # 'csv'
print(result.row_counts) # {'Customer': 500, 'Order': 500}
print(result.evaluation) # metric name -> score
print(result.token_usage) # {'inputTokens': ..., 'outputTokens': ..., 'totalTokens': ...}
print(result.error) # None on successGenerator.available_input_types() lists what planning can classify:
free_text, example_data, schema, document, erd.
A structured run writes a flat directory of per-entity files:
output/
├── customer.csv # one file per entity, lowercased, spaces -> _
├── order.csv
└── order_line_item.csv
free text ─┐
example data ─┤
schema ─┼─→ plan ─→ InferredSchema ─┬─→ structured pipeline ─→ CSV/Parquet/
documents ─┤ │ Excel/JSON
ERD ─┘ └─→ document pipeline ─→ PDF + JSON label
plan classifies each input and normalizes everything into one
InferredSchema, so the choice of modality is made after planning, not
before it. One schema can therefore drive tables, documents, or both. The
document pipeline is unchanged — --schema-dir with a legacy schema directory
still enters it directly, without planning.
data_generator → data_critic → [doc_generator → doc_critic loop] → (augmentor → aug_critic)
↑ (reject) ↑ (reject) ↑ (reject)
└─────────┘ └──────────┘ └──────────┘
The data generator invents schema-conforming JSON; the data critic validates it
(deterministic JSON-Schema check + LLM domain review); the doc generator writes
HTML/CSS and renders a PDF; the vision doc critic reviews the render and loops
until it passes the threshold. With --augment, an augmentor applies aging
effects and an aug critic checks legibility.
A batch turns one high-level --scenario into N distinct scenarios, then runs a
self-contained pipeline graph for each as sibling nodes — the Strands graph
executes them concurrently (no thread pools or worker flags). A specific scenario
produces far more varied output than a generic one:
# Generic — less diverse
--scenario "Generate diverse FCC invoices"
# Specific — much more diverse
--scenario "FCC broadcast invoices for packaged food companies advertising in the midwest"
--scenario "Local car dealership TV ads across small-market stations in the southeast"A packet is a set of different document types that share context — like a loan application with a credit report, pay stubs, and bank statements, all for the same applicant — merged into one multi-page PDF. See the Packets guide.
| Agent | Role | CLI default model |
|---|---|---|
| Scenario Planner | Creates diverse scenarios from a scenario brief + guidance (batch) | nova2-lite |
| Data Generator | Produces JSON data from schema. Has a calculator tool for math. | gpt-oss |
| Data Critic | Validates data against schema and domain rules. Structured output. | sonnet |
| Doc Generator | Writes HTML/CSS and renders to PDF. Has an editor for targeted fixes. | gpt-oss |
| Doc Critic | Vision model evaluates PDF quality — layout, typography, truncation, math. | sonnet |
| Augmentor | Picks an augraphy config and applies document aging effects. | gpt-oss |
| Aug Critic | Evaluates the augmented doc for legibility and realism. Structured output. | sonnet |
This project uses uv and ships a committed
uv.lock for reproducible installs. uv is the recommended workflow.
# Install uv (see https://docs.astral.sh/uv/getting-started/installation/)
curl -LsSf https://astral.sh/uv/install.sh | sh
# Create the virtual environment and install all dependencies from uv.lock,
# including the dev group (pytest, ruff, docs). Requires Python 3.12+.
uv sync
# Run any command inside the managed environment with `uv run`, e.g.
uv run seed-data --helpuv sync installs the dev dependency group by default (configured via
[tool.uv] default-groups). Use uv sync --no-dev for a runtime-only install.
The dev group already includes the structured stack, so a uv sync checkout
can generate both modalities with no extra step.
Alternative: pip or conda
# pip (add the [dev] extra for pytest/ruff/docs tooling)
pip install -e ".[dev]"
# or conda for the environment, then pip for the package
conda create -n seed python=3.12 -y
conda activate seed
pip install -e ".[dev]"Note: pip resolves dependencies fresh and does not use uv.lock.
| Install | Adds | Use for |
|---|---|---|
pip install seed-data |
base only | Documents. Stays lean — no pandas |
pip install "seed-data[structured]" |
pandas, openpyxl, pyarrow | Structured (tabular) generation |
pip install "seed-data[all]" |
every optional feature | Both modalities |
pip install -e ".[dev]" |
[all] + pytest, ruff, mkdocs |
Contributing |
The base install is CI-enforced to import and pass its test suite with no pandas
present, so the published document-only offering keeps working exactly as
before. Structured commands raise a clear ImportError telling you to install
the extra; the document pipeline never needs it. All four output formats —
including --format parquet — work with [structured] alone.
The extra also pins numpy and scipy, but those are not what makes it heavy: both
already arrive in the base install as transitive dependencies of augraphy. They
are named in the extra only to declare the versions this code uses directly.
pandas is the dependency the base install genuinely omits.
The default renderer is xhtml2pdf (pure Python, no system libraries) — a fresh
pip install seed-data renders PDFs out of the box. Select a renderer with --renderer:
--renderer |
Backend | System libraries | Notes |
|---|---|---|---|
xhtml2pdf (default) |
ReportLab | none | Pure Python; works everywhere |
weasyprint |
WeasyPrint | Pango, Cairo, GDK-PixBuf | Richer CSS; needs the libraries below |
reportlab |
ReportLab (Python script) | none | LLM writes a ReportLab script instead of HTML |
WeasyPrint only — install its system libraries first:
# macOS
brew install pango gdk-pixbuf libffi
# Ubuntu/Debian
apt-get install libpango-1.0-0 libgdk-pixbuf2.0-0seed-data --schema-dir fcc-invoice \
--scenario "Local TV station in Portland, Oregon"seed-data --schema-dir fcc-invoice \
--scenario "Local TV station in Portland, Oregon" \
--augment--count > 1 switches into batch mode and fans out concurrent pipeline graphs:
seed-data --schema-dir fcc-invoice \
--count 10 \
--scenario "CPG brands on local TV stations in the American southwest" \
--augmentseed-data packet lending-package \
--count 3 --doc-workers 2 --augment \
--scenario "First-time homebuyers in the Pacific Northwest, 30-year fixed mortgage"Common flags (single-document and batch):
| Flag | Default | Description |
|---|---|---|
--schema-dir |
required | Bundled schema name, or a directory with schema.json |
--output |
./output |
Output root directory |
--scenario |
What to generate this run (for --count > 1, the theme to diversify) |
|
--count |
1 |
Number of documents; > 1 enables batch mode |
--data-model |
gpt-oss |
Model for data generation |
--doc-model |
gpt-oss |
Model for PDF generation |
--critic-model |
sonnet |
Model for all critics |
--batch-model |
nova2-lite |
Model for batch scenario planning |
--aug-model |
gpt-oss |
Model for augmentation decisions |
--renderer |
xhtml2pdf |
xhtml2pdf (pure Python), weasyprint, or reportlab |
--augment |
off | Enable augraphy image augmentation |
--no-critic-samples |
off | Disable reference sample PDFs in the doc critic |
--threshold |
5 |
Acceptance score, 1–10 |
--max-attempts |
5 |
Max critic-retry cycles |
--seed |
Seed for batch scenario planning (regression-stable sets) | |
--timeout |
3600 |
Safety timeout in seconds |
--quiet |
off | Suppress stage progress output |
Packet subcommand (seed-data packet <name|path>):
| Flag | Default | Description |
|---|---|---|
packet |
required | Bundled packet name, or a packet directory path |
--count |
1 |
Number of packets to generate |
--doc-workers |
3 |
Parallel sub-documents within each packet |
--shuffle |
off | Randomize sub-document order in the merged PDF |
--context-model |
nova2-lite |
Model for shared-context resolution |
Plan subcommand (seed-data plan <inputs...>) — any inputs to one
InferredSchema JSON:
| Flag | Default | Description |
|---|---|---|
inputs |
required | One or more free-text descriptions, file paths/globs, and/or s3:// URIs |
--name |
dataset |
Logical dataset name |
--output |
./schema.json |
Path to write the InferredSchema JSON |
--quiet |
off | Suppress stage progress output |
Structured subcommand (seed-data generate-structured <schema>):
| Flag | Default | Description |
|---|---|---|
schema |
required | Bundled schema name, InferredSchema JSON path, or schema directory |
--rows |
100 |
Target records per entity |
--format |
csv |
csv, parquet, excel, or json |
--output |
./output |
Output directory |
--quiet |
off | Suppress stage progress output |
Document subcommand (seed-data generate-documents <schema>) — the modern
spelling of the default mode, and what lets a planned schema drive the
document pipeline:
| Flag | Default | Description |
|---|---|---|
schema |
required | InferredSchema JSON path, schema directory, or bundled name |
--entity |
For a multi-entity InferredSchema: which entity to render | |
--count |
1 |
Number of documents; > 1 plans diverse scenarios and fans out |
--scenario |
What to generate this run | |
--output |
./output |
Output directory |
--data-model |
gpt-oss |
Model for data generation |
--doc-model |
gpt-oss |
Model for PDF generation |
--critic-model |
sonnet |
Model for all critics |
--batch-model |
nova2-lite |
Model for batch scenario planning |
--aug-model |
gpt-oss |
Model for augmentation decisions |
--renderer |
xhtml2pdf |
xhtml2pdf (pure Python), weasyprint, or reportlab |
--augment |
off | Enable augraphy image augmentation |
--no-critic-samples |
off | Disable reference sample PDFs in the doc critic |
--threshold |
5 |
Acceptance score, 1–10 |
--max-attempts |
5 |
Max critic-retry cycles |
--timeout |
3600 |
Safety timeout in seconds |
--seed |
Seed for batch scenario planning | |
--quiet |
off | Suppress stage progress output |
End-to-end subcommand (seed-data plan-and-generate <inputs...>) — plan, then generate, in
one shot with no intermediate schema file.
On plan-and-generate, --output is the modality, not a path. The path is --output-dir.
Every other command uses --output for the path.
| Flag | Default | Description |
|---|---|---|
inputs |
required | Same as plan: free-text descriptions, paths/globs, s3:// URIs |
--output |
structured |
Which modality to generate: structured or documents |
--output-dir |
./output |
Directory to write artifacts to |
--name |
dataset |
Logical dataset name |
--save-schema |
Also write the planned InferredSchema JSON to this path | |
--rows |
100 |
structured only: target records per entity |
--format |
csv |
structured only: csv, parquet, excel, or json |
--count |
1 |
documents only: how many to generate |
--scenario |
documents only: what to generate this run | |
--entity |
documents only: which entity of a multi-entity schema to render | |
--augment |
off | documents only: apply image augmentation |
--data-model |
gpt-oss |
Model for data generation |
--doc-model |
gpt-oss |
Model for PDF generation |
--critic-model |
sonnet |
Model for all critics |
--batch-model |
nova2-lite |
Model for batch scenario planning |
--aug-model |
gpt-oss |
Model for augmentation decisions |
--renderer |
xhtml2pdf |
xhtml2pdf (pure Python), weasyprint, or reportlab |
--threshold |
5 |
Acceptance score, 1–10 |
--timeout |
3600 |
Safety timeout in seconds |
--quiet |
off | Suppress stage progress output |
Utility subcommand — copy the bundled schema library out to edit locally:
seed-data clone-schema-library ./schemasRun seed-data --help, seed-data packet --help,
seed-data infer-schema --help, seed-data plan --help,
seed-data generate-structured --help, seed-data generate-documents --help,
and seed-data plan-and-generate --help for the complete list.
- A single document takes roughly 30–100 seconds end-to-end.
- All workers launch concurrently. The Python Strands graph has no built-in concurrency cap, so large counts may hit Bedrock rate limits (adaptive retries absorb this, but wall-clock grows). Start with modest counts.
- Partial failures are normal. Failed documents come back with
success=Falseandverdict="error"; the rest still complete. - Use a specific
--scenario. Vague scenarios produce same-y output.
| Key | Model | Best For |
|---|---|---|
gpt-oss |
GPT-OSS 120B | Doc/data generation, augmentation (fast tool calling) |
sonnet |
Claude Sonnet | Critics (respects scope rules, good at evaluation) |
haiku |
Claude Haiku 4.5 | Fast critic (cheaper; less strict than sonnet) |
nova2-lite |
Amazon Nova 2 Lite | Scenario/context planning (fast, cheap) |
nemotron-super |
Nemotron Super 120B | Doc generation (slow but high first-pass quality) |
seed-data --help lists every available model key.
Each document type lives in its own directory:
schemas/fcc-invoice/
├── schema.json # JSON Schema — REQUIRED
├── generation_guidance.md # Visual style, data ranges, math rules — optional
└── samples/ # Reference PDFs for the doc critic — optional, local-only
├── README.md # What samples are for + handling rules
└── *.pdf # git-ignored — never committed
The optional samples/ directory holds real example PDFs used only by the
document critic as a visual style reference — they are never seen by the
generation stages and no content from them is reproduced in the output. See
critique.py and any
schemas/<type>/samples/README.md for the full explanation.
Sample PDFs are local-only and git-ignored. Do not commit them: real
documents may contain third-party names or PII you don't have the right to
redistribute, which can trigger data-distribution policies and block
open-source release. You are responsible for having the rights to any sample
you add. Use --no-critic-samples to skip samples entirely. The repository
ships with no sample documents.
This is the key file for controlling document quality and diversity. It tells agents:
- Visual style: layout, fonts, density, colors
- Data realism: value ranges, naming conventions, line item counts
- Math rules: arithmetic constraints (e.g., "GrossTotal = sum of LineItemRate")
- Common issues to avoid: known pitfalls for this document type
- Generation Choices: optional fields with
x-probabilityand table presentation variations (see below)
The scenario planner also reads this file to create diverse scenarios that respect the document type's constraints.
Document diversity comes from two sources of per-run variation:
- Optional fields with
x-probability— schema fields marked with anx-probabilityannotation are independently included or omitted per document based on their probability. - Table presentation variations — sections that can render either as tables
or inline fields, declared in
generation_guidance.mdwith ~50/50 probabilities.
See docs/docs/Guides/generation-choices.md for the full guide: conventions, examples for nested and array fields, and how to add it to a new document type.
Single-document and batch runs write:
output/
├── pdfs/ # Clean PDFs
├── data/ # Ground-truth JSON label per document
├── generation_scripts/ # HTML files used to render PDFs
├── augmented/ # Augmented PDFs (when --augment)
└── config/ # Copy of schema + guidance for reproducibility
Packet runs use the evaluation-dataset layout (merged PDFs in input/,
per-section labels in baseline/); see the Packets guide.
Structured runs (generate-structured, or plan-and-generate --output structured) write one
file per entity directly into the output directory — no subdirectories:
output/
├── customer.csv # lowercased entity name, spaces -> underscores
├── order.csv
└── order_line_item.csv
The extension follows --format: .csv, .json, .xlsx (excel), or
.parquet. StructuredResult.output_paths lists the files that were written.
# Run the full unit-test suite (no LLM calls). Integration tests are excluded
# by default (see addopts in pyproject.toml).
uv run pytest
# Or run a single module
uv run pytest tests/test_unit.py -v
# CLI smoke tests (no Bedrock) — verify the command works out of the box
uv run pytest tests/test_cli_smoke.py -vPacket configs live in src/seed_data/packets/<packet-name>/packet.json.
See the Packets guide for the full guide including config fields,
shared context modes, label format, and architecture.
{
"name": "insurance-claim-packet",
"description": "Property insurance claim submission packet.",
"documents": [
{
"document_class": "Insurance Claim",
"schema_dir": "insurance-claim",
"required": true,
"min_instances": 1,
"max_instances": 1
},
{
"document_class": "Invoice",
"schema_dir": "invoice",
"required": true,
"min_instances": 1,
"max_instances": 3
}
],
"shared_context": {
"claimant_name": "Full legal name of the person filing the claim",
"claimant_address": "Street address of the claimant"
},
"shared_context_hint": "All documents relate to the same insurance claim."
}seed-data runs AI agents that write and execute code and read/write files on the host in order to render documents. Read this section before running it on anything you don't fully trust. See also the disclaimer at the top of this README — you are responsible for the security of anything you build with this tool.
| Aspect | Detail |
|---|---|
| Real data in generation? | No. The generation stages produce entirely fictional content — all names, addresses, and financial figures are invented. |
| PII in output? | No. Output documents and ground-truth JSON contain no real personal or customer data. |
| Document templates? | No. Each PDF is rendered from freshly generated HTML/CSS (or ReportLab) code; no templates or real documents are copied. |
| Reference samples? | Optional, local-only. Real PDFs in a schema's samples/ directory are sent only to the quality-review critic as a visual style reference. The generation stages never see them and no content from them is reproduced in output. They are git-ignored — see Schema Directories. You are responsible for having the rights to any sample you provide. |
| Reproducibility | Batch/packet runs write a manifest recording the command, model versions, git hash, and config. |
The document-generation agent uses tools that act on the host with the privileges of the user running the CLI:
- Code execution — with the
reportlabrenderer the agent runs Python via ashelltool; with the HTML renderers (xhtml2pdf/weasyprint) it renders agent-authored HTML. Treat all agent-generated code inoutput/generation_scripts/as untrusted model output — do not blindly re-run it later. - Filesystem access — file read/write tools operate on model-supplied paths with no built-in sandbox or path allowlist.
- Tool consent — the Strands Agents SDK can prompt for human approval before
executing tools. Setting
BYPASS_TOOL_CONSENT=truedisables that prompt for autonomous runs; leave it unset when running on untrusted input so you can approve tool calls.
Recommendation: run the pipeline in an isolated environment (container or VM) with no sensitive data on disk, restricted network egress, and dedicated, least-privilege AWS credentials — not your personal developer profile.
schema.json, generation_guidance.md, sample PDFs, and the --scenario text
are all injected into agent prompts. Treat content from untrusted sources as a
prompt-injection vector: a malicious schema or guidance file could attempt to
steer an agent into misusing the code/file tools. Only run schemas you trust, or
run inside the isolated environment described above.
Model calls go to Amazon Bedrock using the credentials resolved from your
AWS_PROFILE (or an explicit boto3 session passed to Generator). Any code the
agent executes inherits that environment, so those credentials — and whatever
they can access — are reachable by agent-run code. Use a scoped, Bedrock-only
profile and avoid broad/admin credentials.
Generation runs make repeated Bedrock calls and use critic-driven retry loops.
Large batch counts or high --max-attempts increase token spend. Keep
--max-attempts/--timeout conservative and monitor Bedrock usage.