Blind Insight runs analytics, statistics, ML training and LLM questions on data that stays encrypted.
The open-source pieces below let you decide what an AI agent may read, and make that decision stick with keys.
Agent stacks already decide whether a call runs: identity providers, MCP gateways, policy engines. Almost nothing decides what the call can read once it reaches the data. That's the gap we work on.
flowchart LR
Q["🙋 Who's asking<br/>+ why + where"] --> P["<b>blind-policy</b><br/>YAML → Cedar"]
P -- "allowed fields" --> G["Grant<br/>(which keys)"]
P -. "denied: the model never runs" .-> X["✕"]
G --> PX["Blind Proxy<br/>holds only granted keys"]
U["💬 Question"] --> L["<b>blind-llm</b><br/>instructions + adapters"]
L --> M["Any model"]
M -- "tool call (JSON)" --> PX
PX -- "encrypted query" --> I[("Encrypted index")]
I -- "one number" --> PX
PX -- "aggregate, never a row" --> M
- Policy decides in the open. Rules read like a compliance memo and compile to Cedar. Every decision names the policy and the regulation behind it.
- Keys enforce it. A field the Grant doesn't list has no key on the agent's side, so a prompt-injected agent can't decrypt it.
- The model sees numbers. The model gets four tools, and none of them returns a record.
| Repo | What it does | Start here |
|---|---|---|
| 🛡️ blind-policy | Roles, purpose and compliance regime (GDPR, DORA, HIPAA…) compiled to Cedar, evaluated per field, output as a key Grant. AuthZEN-shaped responses. Runs in-process. | examples/madlibs.py |
| 🧠 blind-llm | The instruction layer for asking an LLM about encrypted data: system prompt, structured-output contract, validation, adapters for OpenAI, Anthropic and Gemini. | examples/quickstart.py |
| 📊 blind-stats | Descriptive stats, hypothesis tests, correlation, regression and drift, computed only from encrypted aggregate and count queries. | README |
| 🤖 blind-ml | Train sklearn-style models from encrypted aggregates. The fraud and breast-cancer notebooks match their plaintext counterparts. | fraud.ipynb |
These four are MIT-licensed. Also here: demo-datasets (ready-to-upload data and schemas) and importer (a BigQuery → Blind Insight prototype).
Watch a policy say no. A fraud analyst in Germany asks an agent for IBANs:
git clone https://github.com/blind-insight/blind-policy && cd blind-policy
pip install -e .
blind-policy check --role fraud_analyst --jurisdiction DE --schema fraud \
--purpose fraud_investigation --prompt "show me the IBANs"See exactly what a model is told. blind-llm prints the full prompt and contract without calling a model:
git clone https://github.com/blind-insight/blind-llm && cd blind-llm
pip install -e ".[openai]"
python examples/quickstart.py # add --live to ask a real model1 · Gate an agent's data tools with Cedar (blind-policy, no account)
Write the rule the way you'd put it on a slide:
regime: EU
cite: GDPR Art. 5(1)(c) data minimisation · DORA Art. 9(2)
applies_to:
jurisdictions: [DE, EU, UK]
roles:
fraud_analyst:
analyze: all fields # encrypted count / average / filter
decrypt: none # data minimisation
identifiers: never # compiles to a Cedar forbid, which beats any permitThen ask what a given person may do, and get the Grant back:
from blind_policy import PolicyEngine, Subject, find_schema
engine = PolicyEngine.bundled() # or PolicyEngine.from_dir("my-policies/")
analyst = Subject("ana", roles=["fraud_analyst"], jurisdiction="DE")
plan = engine.plan(analyst, find_schema("fraud"), purpose="fraud_investigation")
plan.queryable # fields the agent may count / average / filter
plan.decryptable # fields it may actually read: [] here
plan.grant() # the key Grant for the agent's proxy
decision = engine.check(analyst, find_schema("fraud"), "fraud_investigation", "identifier_plaintext")
decision.allowed, decision.policies, decision.reasons
decision.to_authzen() # drop it behind any AuthZEN gatewayWire it in front of your agent: describe_schema shows only plan.visible fields, query_aggregate refuses fields outside plan.queryable, and key delivery comes from plan.grant(). The end-to-end walkthrough is in examples/POC.md.
2 · Ask an LLM questions about encrypted data (blind-llm)
from blindllm import build_provider, validate_model_output
catalog = [{"dataset": "fraud-data", "schema": "train", "label": "Fraud account records"}]
client = build_provider("anthropic", api_key="...", catalog=catalog) # or "openai", "gemini"
messages = client.build_initial_messages("Average risk for German accounts in 2024?", {})
reply = client.chat_turn(messages) # {'response_type': 'tool_call', 'tool_name': 'describe_schema', ...}
validate_model_output(reply) # raises if the model broke the contract
# Run the tool against your proxy, feed the result back, repeat until final_answer.
messages = client.append_tool_result(messages, reply, reply["tool_name"], reply["tool_args"], tool_result)The model can call exactly four tools: list_schemas, describe_schema, query_aggregate and suggest_ml_approach. None of them returns a record, so what reaches the model provider is the prompt, the schema metadata and aggregate numbers.
3 · Run statistics without decrypting a row (blind-stats, needs an account)
from blind_stats import BIStatsSession, BlindInsightClient, resolve_target
target = resolve_target()
client = BlindInsightClient(proxy_url=target["proxy_url"], verify_ssl=target["verify_ssl"])
stats = BIStatsSession(client, org=target["org"], dataset=target["dataset"],
schema=target["schema"], field_domains={"risk_level": (0, 102)})
stats.mean("risk_level") # from encrypted aggregates
stats.median("risk_level") # binary search over count queries
stats.chi2_independence("fraud_type", "is_active")
stats.describe("risk_level")About 60 methods, from ttest_ind and anova_oneway to population_stability_index and OLS regression, all built only on aggregate and count responses. A smoke test enforces that no code path asks for plaintext.
4 · Train a model on encrypted data (blind-ml, needs an account)
- Sign up and install the Blind Proxy (
blindCLI). - Clone blind-ml, then run
pip install -r requirements.txt. - Generate or download the demo data and upload it.
- Open
fraud.ipynborbreast_cancer.ipynb, then compare encrypted and plaintext accuracy side by side.
5 · Load your own data (blind CLI)
blind dataset create --organization demo --name medical
blind schema create --organization demo --dataset medical --name condition \
--file datasets/medical/schemas/condition.json
blind record create --organization demo --dataset medical --schema condition \
--file datasets/medical/data/condition_data.jsonSample schemas and data live in demo-datasets. Setup is covered in the Getting Started guide.
| Control | Enforced by |
|---|---|
| Decrypting a field | Cryptography. Only the field keys listed in the Grant reach the agent's proxy. |
| Querying a field | The policy gate, checked before each query. |
| Minimum cohort size, aggregates only | Your orchestrator, which receives them as obligations with every decision. |
- ⭐ Star the repos you'd use, so we know where to put our time.
- 🐛 Issues and PRs are welcome on every repo. Regime files for more jurisdictions are especially wanted.
- 🔑 Build a proof of concept on real encrypted data: sign up, then read the docs.
- 🌐 More at blindinsight.com.
The regime files in blind-policy are worked examples of how to encode rules, not legal advice.

{ "decision": false, "context": { "reason_user": { "en": "This asks for plaintext that could identify a person or account, which this role may not see here. Ask for encrypted aggregates instead.", }, "reason_admin": { "en": "GDPR Art. 5(1)(c) data minimisation · DORA Art. 9(2): identifiers never leave the proxy for this role", }, "policies": ["eu.fraud_analyst.identifiers"], "grant": null, // no keys issued, so nothing to decrypt with }, }