Senior Systems Engineer · MSP delivery · 10 years in IT, started in the USAF · Colorado
Linux desktop systems · infrastructure visibility · human-gated AI development · local agent compute
I spend my days designing and delivering infrastructure for 80+ client environments. Off the clock I build the systems I want to run on: a desktop that understands the machine under it, a control plane that keeps AI agents on a leash, a live view of my lab, and an agent stack that runs on my own GPU.
These aren't four side projects. They're four layers of one environment.
┌──────────────────────────┐
│ T-CRYPT // SYSTEMS LAB │
└────────────┬─────────────┘
┌─────────────────────────────────┼─────────────────────────────────┐
│ BUILD │ OPERATE │ BUILD
┌─────────┴──────────┐ ┌──────────┴──────────┐ ┌──────────┴─────────┐
│ APHOTIC-HYPR │ │ AETHER │ │ GATE │
│ desktop platform │ │ live lab surface │ │ human-gated agent │
│ Quickshell · QML │ │ 4× Proxmox · ~50 │ │ control plane │
│ Resource Engine │ │ services · events │ │ (early build) │
└─────────┬──────────┘ └──────────┬──────────┘ └──────────┬─────────┘
│ claims VRAM from │ streams model loads, │ drives
│ loaded llama-swap models │ tickets, commits, evals │ Claude Code
└───────────────────────┐ │ ┌───────────────────────┘
▼ ▼ ▼
┌──────────────────────────────────────┐
│ COMPUTE // CLAUDE-LOCAL │
├──────────────────────────────────────┤
│ Claude Code harness · subagents │
│ claude-router local/cloud split │
│ llama-swap roles · lifecycle │
│ llama.cpp inference runtime │
│ RTX 4090 24GB VRAM │
└──────────────────────────────────────┘
| Layer | System | Status | What it is |
|---|---|---|---|
| BUILD | APHOTIC-HYPR | public · beta |
Arch / Hyprland / Quickshell desktop platform with resource arbitration |
| BUILD | GATE | public · early |
Localhost control plane: agents plan and execute, humans approve |
| OPERATE | AETHER | private · live lab |
Visual, agent-readable surface for the homelab |
| COMPUTE | CLAUDE-LOCAL | local architecture |
Claude Code subagents running on local models on my RTX 4090 |
A modular Hyprland environment where one Quickshell shell (300+ QML files) owns the whole desktop UI. It's the environment I run daily on the Arch side of my workstation, and it's built around one rule: features you aren't using shouldn't consume resources.
APHOTIC
├── Quickshell shell ─── bar (5 styles) · notch · launcher · notifications · OSD
│ lock · Command Center · settings · theme creator
├── Services ─────────── system / process / GPU VRAM usage · Hyprland IPC
│ audio · network · weather · plugin registry
├── Resource Engine ──── claim → contention → negotiation → apply → restore
├── Profiles ─────────── base (minimal|full) + Developer · Gaming · AI · Security
└── Plugins ──────────── everything outside the base shell, AI included
- Resource Engine. Workloads declare claims on finite resources (VRAM first). When a gaming claim contends with a loaded local model, the shell offers the negotiation (
Suspend/Keep Running/Ignore) instead of deciding for you. It never kills processes, never touches the kernel orsysctl, and stays dormant until something claims a resource. Ollama, llama-swap, and GameMode claimants ship today. - GPU-aware AI. The llama-swap claimant maps each loaded model to the
llama-serverprocess serving it, reads measured VRAM fromnvidia-smi, and unloads through llama-swap's API on Suspend. A separate service samples live tokens/sec and context use. - Agent visibility. Tracks running Claude Code / Codex / OpenCode / Gemini CLI sessions, separates harnesses (run sessions, execute tools) from providers (serve inference only), and renders tool calls live in the Agent Graph plugin.
- Theming pipeline. Eight shipped themes plus wallpaper-driven color generation, applied live across the shell, GTK3, and GTK4/libadwaita without a relog.
- Plugins. Manifest-based, installable and removable on a running desktop; a plugin can claim the full-screen Workspace plane. Separate repos for plugins, themes, and pets.
- System tooling.
aphoticCLI,--dry-run-first installer with layered package resolution, doctor checks, backups on uninstall. Installer changes run through a Proxmox VM before they ship.
→ Architecture · Resource Engine · Plugin System · Theming
AI can plan and execute. The human keeps authority over every transition that matters. GATE is a localhost control plane built on that rule. The name is literal: the core object is the gate, a checkpoint a step has to clear before the timeline lets it proceed. Agents can attach evidence to a gate; only a human can decide it.
goal ─► agent drafts plan ─► [HUMAN accepts] ─► step runs in its own worktree
│
agent submits evidence ◄─┘
│
[HUMAN decides gate] ─► next step unlocks
The shape is in place (proposed timelines, per-run worktrees, human-only approval, protected branches that are never touched, persistent project state) and the rest is still being built. It's early, and it's the direction I care most about for agentic development: autonomy inside explicit, reviewable boundaries.
Not a repository. AETHER is the live interface to my homelab.
┌─ AETHER // NODE ──────────────────────────────────────────────── LIVE ─┐
│ │
│ LAB (private) EDGE (public) │
│ ├── 4× Proxmox VE nodes ├── glass-cockpit view ◄── people │
│ ├── ~50 services └── /api/v1 · llms.txt ◄── agents │
│ ├── RTX 4090 · local models ▲ │
│ └── agents · tickets · evals │ │
│ │ │ │
│ └──► redaction gate ──► outbound-only publisher ──► relay │
│ │
│ model loads · ticket moves · commits · eval rounds │
└────────────────────────────────────────────────────────────────────────┘
Most homelab dashboards answer "what's running" for the owner on the LAN. AETHER shows the lab working, as it happens: models loading, agents moving tickets, commits landing, eval rounds deciding which local model earns which role. It's built for two audiences. People get a visual control surface; other people's agents get a machine-readable API and llms.txt so they can pull the same state and recipes.
The security boundary is the design constraint, not an afterthought. Nothing connects inbound to the lab, events are redaction-gated at the source before they leave, and the public edge only ever sees what an outbound-only publisher sends it.
Underneath it is an agent-operated lab: four standalone Proxmox VE nodes, Prometheus / Grafana / Uptime Kuma monitoring, CrowdSec, Security Onion, Proxmox Backup Server, and an operations repo that local and cloud agents work from directly.
Source is private and not published.
Running a local LLM isn't the interesting part. The interesting part is that my local models run inside Claude Code's native Agent workflow. Claude Code launches them as ordinary subagents, each in its own agent session, and the model doing the work is on my GPU.
LOCAL AGENT FABRIC
Claude Code (main session, cloud model)
│ ANTHROPIC_BASE_URL
▼
claude-router (loopback) ─── non-local model ──────────────► Anthropic API (passthrough)
│ local role
├── agent 01 ──┐
├── agent 02 ──┼──► llama-swap ──► llama-server (llama.cpp) ──► RTX 4090 · 24GB
└── agent 03 ──┘ role → model --parallel 1 one gen slot
execution |
local |
orchestrator |
Claude Code |
split |
claude-router: local roles → llama-swap, everything else → Anthropic unchanged |
router |
llama-swap |
runtime |
llama.cpp / llama-server |
interface |
OpenAI-compatible HTTP |
accelerator |
RTX 4090 · 24GB |
roles |
agent · agent-fast · coder · coder-fast · review · fast · thinker |
agents |
~3 concurrent streams, independently steerable |
Each role is a Claude Code subagent definition with its own tool set (Read, Grep, Edit, Bash, ...) and guardrails. A local subagent gets its own conversation loop, its own task and context lane, and real tool access: read files, run commands, run builds, do the development work. I can drop into any of those conversations and steer it directly, the same as a normal Claude Code subagent. The router strips auth from local calls, never logs cloud traffic, and has a one-variable kill switch: unset ANTHROPIC_BASE_URL and Claude Code goes straight to Anthropic again.
How three agents share one generation slot
llama-server runs with --parallel 1, so the loaded model has one generation slot. Three agents run concurrently at the workflow layer, not as three simultaneous GPU generations:
AGENT A ── reasoning ── tool call ───────── inference ── tool call
AGENT B ───── inference ── filesystem ───────────── inference
AGENT C ───────── tools ───── waiting ── inference ─────────────
▲ inference requests queue for the one slot; tool work overlaps
Agent work is mostly not token generation. While one agent holds the GPU, the others are reading files, running builds, or parsing output. The pipeline stays busy without claiming parallel generation it doesn't have.
Model routing and swap cost
llama-swap owns the model lifecycle. Agents ask for a role, not a model id, so a quant can change underneath without touching any harness config.
group "chat" swap · exclusive one chat model resident at a time
group "embeddings" persistent small embed model stays loaded alongside
- Agents on the same role share the loaded runtime. No swap.
- Agents on different chat models force an unload/load of a ~20GB model, and that latency dominates.
- So parallel agents get assigned to the same role on purpose when the work allows it. Role assignment is a scheduling decision.
- The embed model is pinned so a code-search call never evicts the chat model mid-session.
The config is managed as code: validated in the repo, checked for drift against the live file and the running server, promoted only when the GPU is idle, and pushed out to the other harnesses (OpenCode, Pi, DSH) that consume the same roles. One config serves both boots of the workstation.
Next: batched generation (experimental, not deployed)
llama-server --parallel 2 gives the model two slots and true batched concurrent generation. It's queued as an A/B test on the MoE agent / fast roles, with a unified KV cache splitting the context budget. The tradeoffs being measured:
- aggregate throughput vs per-stream speed
- context window split across slots
- VRAM headroom on a 24GB card
- gains that depend heavily on the workload mix
It ships only if aggregate throughput clearly improves inside the VRAM budget. Current config is --parallel 1.
| Layer | Tool | Role |
|---|---|---|
| Orchestration | Claude Code | agent harness, subagents, tool execution |
| Local/cloud split | claude-router | local roles to llama-swap, cloud traffic passed through |
| Routing | llama-swap | role aliases, swap groups, load/unload lifecycle |
| Runtime | llama.cpp / llama-server | GGUF inference, pinned builds, slot / context / reasoning-budget config |
| Serving (alt) | Ollama · LM Studio | quick model management, desktop serving |
| Training | Unsloth · PyTorch | fine-tuning experiments |
| Compute | RTX 4090 · CUDA | 24GB VRAM |
Working knowledge: GGUF and quantized inference (IQ / Q4 quants, MoE vs dense tradeoffs) · local model serving · GPU inference and VRAM budgeting · model routing · inference runtime configuration (slots, context, parallelism) · context management · agent orchestration · multi-agent workflows · local coding agents · local/cloud model interoperability over OpenAI-compatible APIs · task-level evals that pick which model earns which role
WORKSTATION ─────────────────────────────────────────────────────
CPU Intel Core i9-14900K
GPU NVIDIA RTX 4090 · 24GB
RAM 64GB
COOLING custom water loop
BOOT
├── Arch Linux
│ └── Aphotic-Hypr daily driver · Aphotic dev · Claude-Local host
└── Windows 11 when the job needs it · same llama-swap stack
ROLES development · local LLM inference · GPU compute · security lab
Hack The Box as 5H3LLKiller. Kali and BlackArch tooling, plus Aphotic's Security profile for offensive-research sublayers. Hack → Scream → Repeat.
Ask me about Linux, Proxmox, MSP infrastructure, HTB, or local inference. Help building out Aphotic-Hypr is always welcome.





