Staff Cloud Engineer · hybrid GCP and on-premises infrastructure, Kubernetes, Cloudflare, Terraform, hybrid cloud networking. Selected projects in more detail than a CV allows.
A Terraform module that describes a whole Cloudflare DNS zone as one map instead of a pile of per-record resources. Built as a product, not a snippet:
Records keyed in state by content, zone-file style, so adding, removing or reordering records never recreates their neighbours; an explicit key lets rotating values (DKIM, IPs) update in place.
Validation at plan: unknown attributes, IP addresses, TTLs, CNAME conflicts and duplicate records fail before the Cloudflare API, with the exact location of the mistake.
One schema for Cloudflare provider v4 and v5, with a state migration between them; adopting an existing zone takes one import block.
Records in HCL or YAML, with a JSON Schema for completion in the editor.
CI across Terraform and provider versions, plus end-to-end tests against a real Cloudflare zone: every record type, updates, import and the v4 to v5 migration.
17 releases; runs 1,750 production DNS records, moved onto it with zero attribute changes; fixes from that migration released upstream in 2.6.4
Cloudflare · Terraform · provider v4 to v5 · state migration
Context
Production Cloudflare managed with Terraform: 22 zones and ~18k lines of code on the v4 provider, which no longer gets new features.
Solution
Converted the code to provider v5 and fixed by hand every place where the plan showed a change to a live object.
Built a verification step: snapshots of the live configuration through the API before and after apply (DNS with record IDs, rules, Workers, pools, members), diffed automatically.
Moved all DNS onto my open-source module; long-standing CNAME conflicts went into an explicit allowlist instead of disabling the check, which immediately exposed one more hidden conflict.
Fixed long-running drift along the way and removed every temporary moved / import / removed block, so the migration is closed rather than left half done.
zero unplanned changes in production; 1,750 DNS records moved without a single attribute change
Firewall and network-policy changes are reviewed by people who see only the diff. The risky part is what the diff does not show: what an object name resolves to, which rules already allow the traffic, which deny will shadow a new rule.
Solution
A reviewer bot that reads the change together with the existing policy and comments on the exact lines: undefined objects, objects in the wrong zone, the same IP under another name, exposure to the internet, removal of the only matching rule, a rule that could extend an existing policy instead.
Never blocks the merge, reports clearly when it could not run, has a bypass for urgent fixes, and collects reactions to tune its checks.
Hardened as an untrusted-input system: path handling, quick-action and HTML injection through comments, and forged bot markers were closed before rollout.
Also reports unused objects and duplicate prefixes in the live configuration on every merge request.
running on 3 repositories; quality measured with 98 eval cases (mutations plus cases mined from past fixes)
AI agents with least privilege: homelab MCP server + RAG
I use AI agents (Claude Code) daily for infrastructure work. To let an agent investigate my Kubernetes homelab on its own, it needs real access to the cluster, GitOps repos and hosts, without being able to change or leak anything.
Solution
A Model Context Protocol server exposing read-only tools: kubectl get / describe / logs / events, ArgoCD application status, GitOps repository browsing, and host diagnostics over SSH.
Runs locally over stdio for Claude Code, and in-cluster over HTTP behind Traefik with cert-manager TLS. Same code, same permissions.
Retrieval-augmented search over the GitOps repositories with local embeddings (all-MiniLM-L6-v2, int8 on onnxruntime): no data leaves the lab, no external API.
Prometheus metrics for tool calls.
The hard part: guardrails that hold even if the agent misbehaves
Kubernetes access through a dedicated ServiceAccount with only get / list / watch, secrets explicitly excluded.
SSH key locked with a forced command on every host: it can run exactly the allowlisted diagnostic commands and nothing else, even if the pod is fully compromised.
Credentials live in Vault, fetched through the Kubernetes auth method, so there are no long-lived secrets on disk.
an agent that can debug the lab end to end, with a blast radius of zero writes
A global SaaS product runs across two on-premises data centers and GCP. Traffic between them has to be encrypted, redundant and fast to recover from link failures.
Solution
Partner Interconnect with HA VPN over Cloud Interconnect (IPsec on top of the interconnect).
ECMP across links, BGP routing between on-prem and Cloud Router.
Private access to Google APIs from on-prem over the private path.
The hard part
BFD was available only for the underlay BGP sessions. The overlay (IPsec) BGP sessions relied on default timers, so after a link failure traffic kept being sent into a dead path for 60-90 seconds.
Fix: Junos event scripts that react to BFD detecting an underlay failure and immediately tear down the corresponding overlay sessions, so routing converges without waiting for hold timers.
failover convergence 60-90 s to 1-3 s; several Gbit/s of production traffic with consistently stable availability
Cloudflare edge caching: ~160 ms to ~10 ms for European users
The product and the marketing website shared one domain, with routing patterns that changed with marketing needs. Many pages had no proper Cache-Control headers.
Solution
Proxied the traffic through Cloudflare and built a set of fine-tuned cache rules plus Cloudflare Workers to decide, per path, what can be served from the edge without breaking the product.
round-trip time for European users cut from ~160 ms to ~10 ms; origin load down ~95%; marketing pages served straight from the edge cache
Data center migration: from L2/STP to EVPN
EVPN · STP · data center fabric · firewalls
Context
A data center move with full equipment replacement, from a legacy L2 network built on STP.
Solution
Owned the network side: integrated the legacy L2 network with a new EVPN-based fabric so both could run side by side during the move, then transitioned traffic between firewalls.
delivered on time with zero service disruption
Access Review: centralized access auditing
Go · agent-server · Active Directory · PostgreSQL · OpenVPN · Linux
Context
Local accounts live on thousands of Linux VMs and bare-metal servers, in PostgreSQL databases and on OpenVPN instances. Inactive and orphaned accounts are a security risk that manual reviews do not catch.
Solution
Designed and built an agent-server tool, from the initial proposal to a full second-generation rewrite. Agents collect local accounts from every source; the server reconciles them against Active Directory and surfaces inactive and orphaned accounts. Written in Go with AI-assisted development; I owned the architecture, the agent-server protocol and the reconciliation logic.
The hard part
Consistent, low-overhead collection across very different sources at the scale of several thousand instances.
100% of infrastructure resources covered by automated access review
Protected web assets against large-scale scraping, credential stuffing and other automated attacks with Cloudflare Bot Management. Cut false positives by tuning custom firewall rules and Managed Challenges on bot-score analysis, and moved bot protection into the Terraform pipeline so every production environment gets the same policy.
Infrastructure & network as code
Terraform · GCP · Cloudflare · Junos · NETCONF
Refactored GCP and Cloudflare infrastructure from monolithic configurations into modules: VPCs, Interconnects, firewall rules, Cloudflare WAF / DNS / Workers.
Open-sourced terraform-cloudflare-easy-dns: a whole DNS zone as one nested map, keyed by content so edits never recreate neighbouring records.
Extended the same approach to on-prem Juniper devices with terraform-provider-junos, contributing edge cases and feature requests for enterprise networking.
Fleet memory reclamation
Linux · Proxmox · ZFS · zram · open source
Context
Hypervisors looked close to running out of memory in monitoring, which pointed towards buying more RAM. Much of that "used" memory was ZFS read cache: Proxmox's documented ARC default (10% of RAM, at most 16 GiB) is only written at install time when the root file system is ZFS, so pools added later fall back to OpenZFS's much larger default, and many tools count ARC as used memory.
Solution
Capped the ZFS ARC on hypervisors, sized to RAM and pool capacity, applied at runtime and persistently.
Disabled Transparent Huge Pages across the fleet, rolled out in stages.
Introduced zram on staging for higher workload density.
Shared
proxmox-zfs-arc-tuner: an interactive script that analyses host RAM and pools, calculates the ARC limit and applies it at runtime and persistently.