Bridging dev and ops with code
I am the sole DevOps engineer at my company and the single escalation point across engineering
for infrastructure, deployment and production incidents. Eight years across software and
infrastructure, the last three dedicated to DevOps and SRE.
Today the live estate runs on Vercel, Fly.io, Supabase and Cloudflare with self-hosted
Grafana, Loki and Prometheus, plus three AWS accounts. Before that I operated three
Flux CD-managed K3s clusters, held to a 99% uptime target, with SLOs, SLIs and error
budgets I defined myself. I own
on-call, write the postmortems and maintain the runbooks — and when an LLM or agent
misbehaves in production, the debugging is mine too, across
Anthropic Claude, OpenRouter, DeepInfra and Perplexity MCP.
Before infrastructure I spent five years building backend and frontend systems in Laravel,
Angular and Vue, which is why I can debug the application as well as the machine under it.
Where I've worked
- Sole DevOps engineer for the company and the single escalation point across engineering for infrastructure, deployment and production incidents.
- Built and operated two K3s clusters (Singapore and EU) end-to-end solo, taking over platform ownership after the engineer who originally built it left.
- Operated three Flux CD-managed K3s clusters — 33 workloads, 391 manifests — with Longhorn storage, Traefik ingress, cert-manager, Loki/Vector logging and SOPS-encrypted secrets. Live incident response included etcd disaster recovery and certificate rotation. (Oct 2025 – May 2026)
- Currently operate the live estate on Vercel, Fly.io, Supabase and Cloudflare with self-hosted Grafana/Loki/Prometheus, running distributed multi-LLM agent workloads on Anthropic Claude, OpenRouter, DeepInfra and Perplexity MCP.
- Built a Prometheus/Loki observability platform monitoring 45 scrape endpoints, and a daily automated SRE reporting system covering 10 production applications.
- Define SLOs, SLIs and error budgets against a 99% uptime target; own on-call incident response, write postmortems and maintain runbooks.
- Authored ~9,800 lines of Terraform across 22 reusable modules and ran GitOps delivery to K3s with Flux CD; maintained 197 CI/CD workflows across 65 repositories, including one service deploying four environments across two AWS regions.
- Instrument observability with Prometheus, Grafana, CloudWatch and Loki/OpenTelemetry; centralise secrets in HashiCorp Vault.
- Monitor cloud spend across AWS and Azure, and help select and size AWS Reserved Instance commitments.
- Managed high-availability blockchain RPC endpoints and node infrastructure for the Saakuru Wallet and SDK.
- Diagnosed distributed system failures — chain reorgs, transaction inconsistencies, indexing delays and explorer mismatches.
- Supported third-party node applications on Azure VMs alongside the primary AWS estate, including monitoring and operational maintenance.
- Helped decommission legacy blockchain stacks during the company-wide pivot to AI-focused infrastructure.
- Automated multi-environment deployments using Jenkins and GitHub Actions, reducing manual overhead and deployment errors for the MetaOne Wallet.
- Stabilized production environments by performing deep-dive troubleshooting into Node.js services and PostgreSQL database performance.
- Gained cross-cloud proficiency by managing AWS EC2/S3 assets and supporting initial Azure-based infrastructure operations.
- Standardized dev-to-prod parity by containerizing Laravel and Vue.js applications using Docker, improving developer onboarding speed.
- Developed internal REST APIs and admin dashboards, providing a strong software engineering foundation for future infrastructure automation roles.
Technical toolkit
A pragmatic toolkit built across cloud, container orchestration, IaC, and observability.
What I've built and run
Four things that say more about how I work than a tools list does.
Education & certifications
Certifications
Currently working toward cloud and Kubernetes credentials. This list grows.
Open for side gigs
Available for short-term DevOps consulting, infrastructure audits, and freelance contracts — on top of my full-time role. Outside of business hours & weekends.