Applied AI engineer building LLM systems for healthcare: verification, regulatory review and patient communication. Founder of PharmaTools.AI, a suite of production AI tools used by clinicians, medical writers and patients.
- Retrieve evidence rather than invent it
- Expose uncertainty rather than conceal it
- Constrain capability where consequences are high
- Help humans audit reasoning, not replace judgement
The thread through my work: don't ask the model to police itself. My published position argues for constraining healthcare LLMs to translation rather than interpretation. OpenGATE replaces the LLM judge with deterministic checks, RSI Loop shows an auditor rejecting reward-hacked self-improvements, and Redacta keeps re-identification maps out of the model's reach by construction.
| Tool | What it does | Adoption |
|---|---|---|
| OpenGATE | Deterministic grounding verification with no LLM judge. Runs as a CI release gate on the tools below | |
| PubCrawl | MCP server for PubMed, Europe PMC, ClinicalTrials.gov and US/UK drug labels (14 tools) | |
| Redacta | De-identifies clinical text before it reaches an AI, then re-identifies locally. Nine surfaces, including a self-hosted Kubernetes service | |
| StudyDiff | Explains why two studies disagree, grounding every statement in the source (live demo) |
Reference implementations: RSI Loop (validated self-improvement with an auditor that blocks specification gaming) · LitRAG (RAG with a built-in citation-faithfulness eval)
- Observer Zero: Do LLM Agents Form Epistemic Communities? Preprint, 2026. DOI · code
Across 235 runs (85 pre-registered), LLM agents detected a hidden change of physical law in 90–100% of worlds but correctly diagnosed it in 0 of 40 opportunities. - Translation, not Interpretation: Rethinking Language Model Design for Healthcare. SN Comprehensive Clinical Medicine, 2026. DOI
Argues for constraining healthcare LLMs to translational tasks: a narrower surface and a lower harm ceiling. - Validation of an AI-powered mobile application for personalizing medical note explanations. Frontiers in Digital Health, 2026. DOI
Patiently AI: 87.3% of outputs rated clinically safe by experts, 70% patient preference, reading level down ~3 grades. - A Day in the Life of an MSL Powered by AI. JNGR 5.0, 2025. PDF
A design proposal for an AI-powered Medical Science Liaison (MSL) training platform, combining RAG for personalised content, multimodal mechanism-of-action visualisation and explainable recommendations.
| Product | What it does | Of note |
|---|---|---|
| Patiently AI | Turns medical notes into patient-friendly language, without crossing into diagnosis | 5× award winner · peer-reviewed validation · iOS / Android / web |
| RefCheckr | Checks clinical claims against references; rejects any quote it can't find in the source | Web app · Word add-in |
| MedCheckr | ABPI Code review with clause-level citations | Code Clarity Awards Winner 2024 |
| PosterLens | Structured extraction from scientific posters | Presented at ESMO AI & Digital Oncology 2025 |
| BiomarkerFinder | Biomarker associations in plain language, with provenance | Open Targets Hackathon winner |
| HushMap | Sensory-friendly place finder for neurodivergent users | Community-sourced · Apple Watch |
LLM systems: Claude API · MCP · RAG (Pinecone, FAISS) · structured outputs · evals
Application: TypeScript · Python · Swift · Postgres · Docker · Kubernetes




