I build AI systems that hold up beyond the demo: agentic workflows, context engineering, typed tool use, structured outputs, hard validation, bounded repair, human approval, and observable delivery.
My work spans shipped open-source AI products, LLM evaluation infrastructure, and industrial/scientific machine learning. I work end-to-end across Python/FastAPI, TypeScript/React, Tauri/Rust, SQLite, CI/CD, and observability.
Explore Remis · AI Engineering · LinkedIn
Remis — Agentic AI product and LLM workflow system
Designed, built, and maintain an AI-native desktop product that turns model calls into a governed localisation workflow.
- Shipped 29 public releases, 500+ installer downloads, and localisation releases reaching 8,000+ Steam Workshop users.
- Built a localhost Agent API and repository-bundled operator skill, plus approval-gated Copilot architecture with model-selected tools, typed plans, persistent sessions, and controlled execution.
- Orchestrated cloud and local models behind context assembly, structured outputs, deterministic validators, bounded repair, checkpoint recovery, and human review.
- Delivered the complete Windows product across Python/FastAPI, React, Tauri/Rust, and SQLite, supported by regression coverage across 120+ test files.
Product · Use with an AI Agent · Architecture
Aventine — Reproducible LLM evaluation infrastructure
Created a public benchmark for complete translation recipes—including models, prompts, context, repair, and validators—across frozen tasks, hard validation, calibrated multi-provider judging, and reproducible result contracts.
- Built multilingual MQM and ACES calibration packs, bounded judge runners, MetricX/xCOMET baselines, and human-gold/judge/metric alignment analysis.
- Designed the benchmark so structurally unsafe output cannot win on style scores and inconsistent judgments remain visible rather than being averaged away.
- Built LAVA, an engineer-in-the-loop corrosion intelligence workflow selected for the AGS NSW Generative AI Showcase in Geotechnical Engineering and Engineering Geology Practice.
- ARC Industrial Transformation Training Centre Scholar in a BlueScope-linked research collaboration, building reproducible modelling workflows across 50+ electrochemical datasets and 100+ microscopy images.
- Translate noisy evidence, model uncertainty, and engineering assumptions into auditable decision support rather than opaque predictions.
AI Agents · Agentic Workflows · RAG / Context Engineering · Tool Calling · Structured Outputs · Human-in-the-Loop · LLM Evaluation · LLM-as-a-Judge · LLMOps · Industrial AI
I am open to Applied AI, AI Agent, LLM Evaluation, and LLMOps roles where reliable delivery matters as much as model capability.



