ENGINEERING / SRE / AUTOMATION / PLATFORMS

Reliability is
in the details.

More than two decades building the systems behind the experience.

I’ve built enterprise secrets platforms, operated observability at scale, and automated deployment. Today I lead a professional services team at HashiCorp (an IBM Company), and keep building: reliable infrastructure, AI systems, and tools that make complex work usable.

Utah, USA Open to interesting conversations

FIELD NOTES / 001INDEPENDENT LAB

A small datacenter.
A wide-angle mindset.

Explore the layers of my engineering practice.

Explore the system layers

01 / PLATFORM ENGINEERING

Own the whole lifecycle.

Scheduling, workload identity, service discovery, and storage. Build it, run it, learn from it.

Inside the datacenter

02 / AI INFRASTRUCTURE

A model is only the beginning.

GPU serving, consistent API access, and correlated metrics and logs. Make inference operable.

Inside the inference stack

03 / DEVELOPER TOOLS

Make the complexity useful.

Go MCP gateways and auditable agent coordination. Give people and tools a better way in.

Inside the tooling

01 / ENTERPRISE EXPERIENCE

Real scale.
Real responsibility.

My technical foundation was built inside enterprise systems. Reliability, security, and the people operating them have always been part of the work.

ADOBE / 2016–2022

Secrets infrastructure that could grow.

Technical lead for Adobe’s multi-tenant Vault service. I re-architected the deployment for resiliency, planned and executed its move from data centers to AWS, and built dashboards for leadership, consumers, and administrators.

Vault · AWS · Resiliency · Worldwide training

ADOBE / 2016–2022

Observability at enterprise scale.

I architected and operated Splunk Enterprise clusters ingesting multiple terabytes of logs daily across data center and cloud, managing 6,000+ forwarders. Python automation replaced recurring manual administration.

Splunk · Python · Multi-TB/day · 6,000+ forwarders

HASHICORP / STAFF ARCHITECT / 2022–2023

Platform adoption through self-service.

At a top-5 US bank, I designed a Vault cluster and namespace vending machine. At a top-10 bank, I built identity and authentication strategy with Sentinel policies enforcing secret schemas and validation at write time.

Vault · Identity · Sentinel · Enterprise architecture

XACTWARE / 2012–2016

Deployment without the interruption.

I built WAR and static-content deployment tooling integrated with Jenkins and NetScaler server draining for zero-downtime deployments. Application clustering, Tomcat session replication, and consolidated monitoring supported the operating model.

Java · Jenkins · NetScaler · Tomcat · On-call operations

Earlier chapters: Java development and QA at Xactware, PHP/Drupal at Astonish Design, and freelance web development. Download my IC resume (PDF) for the full career.

BUILT & OPERATED

Curiosity, backed
by real systems.

7

Nomad nodesSelf-hosted cluster

72+

TiB of storageTrueNAS + NVMe tier

40+

Self-hosted reposOn Forgejo

02 / CURRENT INDEPENDENT WORK

Built with intent.
Run with care.

Still building. Still accountable for the result.
Explore a focus, then open the engineering notes to see the decisions behind the build.

Showing three platform and reliability projects. Showing four AI and agent systems projects. Showing three software and tooling projects.
02 / AGENT ORCHESTRATIONV1 / testing

auto-agents

Give autonomous work an operating model.

A work-request bus and service directory for coding agents. Agents declare ownership, exchange requests, and carry conversations across stateless wake-ups; a human operator can monitor, intervene, approve, or stop the work.

  • Go
  • Nomad
  • Vault
  • PostgreSQL
Engineering notes

The coordination problem

Agents need explicit ownership and durable conversations when work spans services and separate sessions.

The design choices

A registry, inboxes, and threads coordinate work. Disposable Nomad workers use scoped identities. Concurrency, rate, and hop limits provide backstops; an operator interface and audited actions make oversight part of the architecture.

What this demonstrates

Designing the control and supervision model alongside the automation. V1 is feature-complete and end-to-end tested, with hands-on deployment testing as the next stage.

ORIGINAL INDEPENDENT PROJECT / OCT 2026

03 / PLATFORM ENGINEERINGAlways on

Pondside Datacenter

A place to build. A platform to depend on.

A seven-node Nomad cluster across mini PCs and Dell PowerEdge servers. I operate the stack from workload identity and service discovery through traffic routing, storage, and delivery.

  • Nomad
  • Vault
  • Consul
  • TrueNAS
  • Forgejo
Engineering notes

The operating challenge

A heterogeneous cluster still needs a consistent way to schedule workloads, deliver secrets, and expose services.

The design choices

Vault provides JWT workload identity; Consul handles discovery; Traefik routes services. HAProxy and Keepalived handle balancing and failover. TrueNAS provides roughly 72 TiB of usable storage plus a fast NVMe tier.

What this demonstrates

Hands-on ownership across compute, networking, identity, storage, and the lifecycle of the services running on them.

INDEPENDENT INFRASTRUCTURE / UPDATED OCT 2026

04 / AI INFRASTRUCTUREActive build

LLM inference on prosumer GPUs

Local models. Operational discipline.

A 27B-parameter model on a three-card AMD Radeon AI PRO R9700 system. Twin vLLM engines sit behind one LiteLLM endpoint, allowing either backend to be taken out for maintenance while the other keeps serving.

  • vLLM
  • LiteLLM
  • AMD RDNA4
  • VictoriaMetrics
Engineering notes

The serving challenge

Local inference needs more than a loaded model: it needs consistent API access, a maintenance path, and useful visibility.

The design choices

Qwen3.8-27B uses AWQ W4A16 quantization. Two cards run the same model as separate engines in one LiteLLM group. Metrics go to VictoriaMetrics, logs to VictoriaLogs, and Grafana connects the two with click-through log correlation and three alerts.

What this demonstrates

Applying reliability practices to emerging AI workloads, from hardware constraints to serving and observability.

INDEPENDENT PROJECT / UPDATED OCT 2026

05 / HUMAN-CENTERED TOOLINGAlways on

Homelab AI gateway platform

“Can you start a Minecraft server?”

Ten Go gateway services turn the homelab into model-accessible tools. A chat interface lets family members find and deploy Minecraft modpacks by asking, with orchestration, DNS, and persistent worlds handled behind the scenes.

  • Go
  • MCP
  • Nomad
  • Cloudflare
Engineering notes

The user problem

Non-technical users should not need to learn deployment tooling to play together.

The design choices

Gateways expose systems including Nomad, Vault, Postgres, and Cloudflare. The Minecraft flow finds packs across CurseForge, Modrinth, and FTB, launches Nomad job templates, manages Cloudflare DNS, and persists worlds on TrueNAS NFS.

What this demonstrates

Connecting infrastructure automation to a concrete user need through an approachable interface.

INDEPENDENT PROJECT / UPDATED OCT 2026

06 / OBSERVABILITYActive build

Homelab observability stack

Better questions. Better signals.

A self-hosted metrics and logs pipeline with a one-year retention target. Prometheus provides a short scrape buffer, VictoriaMetrics stores long-term metrics, and Grafana Alloy unifies collection.

  • Prometheus
  • VictoriaMetrics
  • Loki
  • Grafana Alloy
Engineering notes

The storage decision

Metrics and logs have different storage needs; a single storage choice need not fit both.

The design choices

Prometheus uses remote_write to VictoriaMetrics on NFS. Loki uses MinIO-backed object storage for logs. Shelly Gen 4 smart plugs bring power telemetry into the same pipeline. A one-year retention target guides storage planning.

What this demonstrates

Selecting storage and collection patterns around actual workloads, with hardware signals included in the operating picture.

INDEPENDENT PROJECT / UPDATED OCT 2026

07 / PHYSICAL RELIABILITYShipped

Pondside Power Backup

Reliability starts before the server.

A 5000W inverter/charger with 48V LiFePO4 rack batteries supports the lab through grid interruptions. Battery-management telemetry flows into Prometheus and Grafana, bringing power into the same operational view.

  • LiFePO4
  • RS485
  • Prometheus
  • Grafana
Engineering notes

The dependency

Service redundancy still depends on a working power supply.

The design choices

The battery-management system provides RS485 data to the monitoring pipeline. The setup is solar-ready; panel installation is future work.

What this demonstrates

Looking beyond software to the physical dependencies of availability, then instrumenting those dependencies.

INDEPENDENT PROJECT / UPDATED SEP 2026

Most of this work lives on my own infrastructure. My public contribution history is on GitHub (opens in a new tab)

02 / PUBLISHED PERSPECTIVE

Make the system
explain itself.

Good observability turns signals into useful answers.

HashiCorp Blog08 AUG 2023 / CO-AUTHOR

HashiCorp Vault observability: Monitoring Vault at scale (opens in a new tab)

Co-authored with JD Goins. A practical strategy that brings log analysis, telemetry, and API and synthetic monitoring together to help teams understand Vault health, usage, and the experience of its consumers.

QUESTIONS THE STRATEGY HELPS ANSWER

  • Is the cluster healthy under its current workload?
  • What are consumers actually experiencing?
  • Which usage patterns need a closer look?
Read the article (opens in a new tab)

04 / HOW I THINK

Curious by nature.
Practical by design.

I enjoy the deep technical work. The point is to make the resulting system easier to live with.

01

Start with the operator.

Think about deployment, visibility, maintenance, and recovery while designing the system. The person on call inherits every architectural decision.

02

Use AI with engineering discipline.

I use AI to increase what I can deliver. I own the problem framing, architecture, review, and verification: decision records, test-first development, measured evaluations, and staged rollouts turn generated code into systems I can stand behind.

03

Follow the user’s problem.

A chat request that becomes a running game server is the kind of result I care about. The infrastructure earns its place by making something useful possible.

AWAY FROM THE TERMINAL

D&D campaigns, live-fire cooking, Utah trails, electronic music, and two dogs with very little respect for uptime guarantees.

05 / NEXT CONVERSATION

Let’s build something
worth depending on.

Have a platform challenge, an AI infrastructure problem, or a team that cares about how systems run? I’d be glad to talk.