DevOps Engineer · Infrastructure & Reliability

Budi Setiawan
Tanjung

Sole DevOps engineer running production infrastructure for multi-LLM AI agent workloads. Eight years across software and infrastructure, the last three dedicated to cloud operations, reliability and AI platform ops on AWS & Kubernetes.

Available for remote roles 📍 Surabaya, Indonesia 🌏 Open to global opportunities
Get in touch GitHub LinkedIn

Bridging dev and ops with code

I am the sole DevOps engineer at my company and the single escalation point across engineering for infrastructure, deployment and production incidents. Eight years across software and infrastructure, the last three dedicated to DevOps and SRE.

Today the live estate runs on Vercel, Fly.io, Supabase and Cloudflare with self-hosted Grafana, Loki and Prometheus, plus three AWS accounts. Before that I operated three Flux CD-managed K3s clusters, held to a 99% uptime target, with SLOs, SLIs and error budgets I defined myself. I own on-call, write the postmortems and maintain the runbooks — and when an LLM or agent misbehaves in production, the debugging is mine too, across Anthropic Claude, OpenRouter, DeepInfra and Perplexity MCP.

Before infrastructure I spent five years building backend and frontend systems in Laravel, Angular and Vue, which is why I can debug the application as well as the machine under it.

Where I've worked

DevOps Engineer (AI Infrastructure Focus)
SiloTech.xyz
Mar 2025 – Present
Lithuania · Remote
  • Sole DevOps engineer for the company and the single escalation point across engineering for infrastructure, deployment and production incidents.
  • Built and operated two K3s clusters (Singapore and EU) end-to-end solo, taking over platform ownership after the engineer who originally built it left.
  • Operated three Flux CD-managed K3s clusters — 33 workloads, 391 manifests — with Longhorn storage, Traefik ingress, cert-manager, Loki/Vector logging and SOPS-encrypted secrets. Live incident response included etcd disaster recovery and certificate rotation. (Oct 2025 – May 2026)
  • Currently operate the live estate on Vercel, Fly.io, Supabase and Cloudflare with self-hosted Grafana/Loki/Prometheus, running distributed multi-LLM agent workloads on Anthropic Claude, OpenRouter, DeepInfra and Perplexity MCP.
  • Built a Prometheus/Loki observability platform monitoring 45 scrape endpoints, and a daily automated SRE reporting system covering 10 production applications.
  • Define SLOs, SLIs and error budgets against a 99% uptime target; own on-call incident response, write postmortems and maintain runbooks.
  • Authored ~9,800 lines of Terraform across 22 reusable modules and ran GitOps delivery to K3s with Flux CD; maintained 197 CI/CD workflows across 65 repositories, including one service deploying four environments across two AWS regions.
  • Instrument observability with Prometheus, Grafana, CloudWatch and Loki/OpenTelemetry; centralise secrets in HashiCorp Vault.
  • Monitor cloud spend across AWS and Azure, and help select and size AWS Reserved Instance commitments.
DevOps Engineer
Saakuru Blockchain
Jan 2024 – Mar 2025
Singapore · Remote
  • Managed high-availability blockchain RPC endpoints and node infrastructure for the Saakuru Wallet and SDK.
  • Diagnosed distributed system failures — chain reorgs, transaction inconsistencies, indexing delays and explorer mismatches.
  • Supported third-party node applications on Azure VMs alongside the primary AWS estate, including monitoring and operational maintenance.
  • Helped decommission legacy blockchain stacks during the company-wide pivot to AI-focused infrastructure.
DevOps Engineer
AAG Ventures
Mar 2023 – Jan 2024
Singapore · Remote
  • Automated multi-environment deployments using Jenkins and GitHub Actions, reducing manual overhead and deployment errors for the MetaOne Wallet.
  • Stabilized production environments by performing deep-dive troubleshooting into Node.js services and PostgreSQL database performance.
  • Gained cross-cloud proficiency by managing AWS EC2/S3 assets and supporting initial Azure-based infrastructure operations.
Frontend & Laravel Developer
PT Saranan Integrasi Informatika Nusantara
Sep 2018 – Mar 2023
Indonesia
  • Standardized dev-to-prod parity by containerizing Laravel and Vue.js applications using Docker, improving developer onboarding speed.
  • Developed internal REST APIs and admin dashboards, providing a strong software engineering foundation for future infrastructure automation roles.

Technical toolkit

A pragmatic toolkit built across cloud, container orchestration, IaC, and observability.

Cloud & Orchestration
AWS EC2 ECS S3 Global Accelerator Kubernetes (K3s) Docker Azure VMs
IaC & CI/CD
Terraform GitHub Actions Jenkins Bash Node.js
📊
Observability & Security
Prometheus Grafana HashiCorp Vault CloudWatch
🤖
Niche Infrastructure
LLMOps Multi-LLM Backends Blockchain Nodes RPC / Indexing Supabase (Self-hosted)
💾
Data & Middleware
PostgreSQL Redis MinIO Nginx

What I've built and run

Four things that say more about how I work than a tools list does.

🤖
Multi-LLM agent infrastructure
Production infrastructure for distributed AI agent workloads integrated directly against Anthropic Claude, OpenRouter, DeepInfra and Perplexity MCP — multi-provider routing, per-provider failure modes, rate limits and cost across four vendors. Includes prompt-level behaviour debugging in production, not just uptime.
🛠
Internal developer-management platform
Principal contributor. Node/Express/TypeScript with Prisma and PostgreSQL, React 18/Vite frontend, deployed to Fly.io and Vercel. Magic-link authentication with server-side verification, a 25-entity domain model covering contracts through invoicing, and an MCP server so the whole workflow can be driven from an AI coding CLI.
Blockchain RPC & node operations
High-availability RPC endpoints and node infrastructure across two blockchain products. Chain reorgs, transaction inconsistencies, indexing delays, explorer mismatches, node performance tuning and chain upgrades — the failure modes most infrastructure engineers never meet.
🛡
Disaster recovery under pressure
Rebuilt compromised infrastructure inside-out to restore security and reliability, then carried two platforms through company pivots and the decommissioning that came with them — without losing the services still running on them.

Education & certifications

Bachelor of Information Technology
Universitas Widya Kartika
2014 – 2018
Blockchain Developer Certification
Pelita Bangsa Academy
2025
Fullstack Web Developer
Purwadhika Digital Technology School
2023

Certifications

Currently working toward cloud and Kubernetes credentials. This list grows.

Blockchain Developer
Pelita Bangsa Academy
Aug – Dec 2025
AWS Certified Cloud Practitioner
Amazon Web Services
In progress
AWS Certified Solutions Architect – Associate
Amazon Web Services
Planned
Certified Kubernetes Administrator (CKA)
CNCF / Linux Foundation
Planned

Open for side gigs

Accepting consultation & freelance work · Q2 2026

Available for short-term DevOps consulting, infrastructure audits, and freelance contracts — on top of my full-time role. Outside of business hours & weekends.

AWS Infrastructure Audit
Architecture review, IAM and security check, cost and Reserved Instance analysis.
💰
Cost Optimization (FinOps)
Identify waste, right-size resources, evaluate RIs / Savings Plans, S3 lifecycle.
🚀
CI/CD Pipeline Setup
GitHub Actions, Jenkins, multi-env deploys, secrets management.
📊
Observability Stack
Prometheus, Grafana, Loki, CloudWatch — dashboards, alerts, SLOs.
🐳
Kubernetes / K3s
Cluster setup, workload migration, troubleshooting, GitOps onboarding.
🤖
LLMOps Consulting
Multi-LLM backend orchestration, scaling AI workloads on AWS.
Or email directly
Replies typically within 24h · Free 30-min discovery call · NDA available on request