5xx.eu/projects

Nikita Puglachenko

Staff Cloud Engineer · hybrid GCP and on-premises infrastructure, Kubernetes, Cloudflare, Terraform, hybrid cloud networking. Selected projects in more detail than a CV allows.

Open source & writing

Terraform · Cloudflare

terraform-cloudflare-easy-dns

A Terraform module that describes a whole Cloudflare DNS zone as one map instead of a pile of per-record resources. Built as a product, not a snippet:

17 releases; runs 1,750 production DNS records, moved onto it with zero attribute changes; fixes from that migration released upstream in 2.6.4

Also

proxmox-zfs-arc-tuner and the article on Proxmox ZFS ARC defaults: see fleet memory reclamation.

Production Cloudflare migration to provider v5

Cloudflare · Terraform · provider v4 to v5 · state migration

Context

Production Cloudflare managed with Terraform: 22 zones and ~18k lines of code on the v4 provider, which no longer gets new features.

Solution

zero unplanned changes in production; 1,750 DNS records moved without a single attribute change

LLM review for network-policy merge requests

LLM · evals · Juniper SRX · GCP firewall · Cloudflare · GitLab CI

Context

Firewall and network-policy changes are reviewed by people who see only the diff. The risky part is what the diff does not show: what an object name resolves to, which rules already allow the traffic, which deny will shadow a new rule.

Solution

running on 3 repositories; quality measured with 98 eval cases (mutations plus cases mined from past fixes)

AI agents with least privilege: homelab MCP server + RAG

MCP · RAG · Claude Code · Kubernetes · ArgoCD · Vault · Node.js

Context

I use AI agents (Claude Code) daily for infrastructure work. To let an agent investigate my Kubernetes homelab on its own, it needs real access to the cluster, GitOps repos and hosts, without being able to change or leak anything.

Solution

The hard part: guardrails that hold even if the agent misbehaves

an agent that can debug the lab end to end, with a blast radius of zero writes

Hybrid connectivity: on-prem data centers and GCP

GCP · Partner Interconnect · HA VPN · IPsec · BGP · ECMP · BFD · Junos

Context

A global SaaS product runs across two on-premises data centers and GCP. Traffic between them has to be encrypted, redundant and fast to recover from link failures.

Solution

The hard part

BFD was available only for the underlay BGP sessions. The overlay (IPsec) BGP sessions relied on default timers, so after a link failure traffic kept being sent into a dead path for 60-90 seconds.

Fix: Junos event scripts that react to BFD detecting an underlay failure and immediately tear down the corresponding overlay sessions, so routing converges without waiting for hold timers.

failover convergence 60-90 s to 1-3 s; several Gbit/s of production traffic with consistently stable availability

Cloudflare edge caching: ~160 ms to ~10 ms for European users

Cloudflare · Workers · cache rules · Cache-Control

Context

The product and the marketing website shared one domain, with routing patterns that changed with marketing needs. Many pages had no proper Cache-Control headers.

Solution

Proxied the traffic through Cloudflare and built a set of fine-tuned cache rules plus Cloudflare Workers to decide, per path, what can be served from the edge without breaking the product.

round-trip time for European users cut from ~160 ms to ~10 ms; origin load down ~95%; marketing pages served straight from the edge cache

Data center migration: from L2/STP to EVPN

EVPN · STP · data center fabric · firewalls

Context

A data center move with full equipment replacement, from a legacy L2 network built on STP.

Solution

Owned the network side: integrated the legacy L2 network with a new EVPN-based fabric so both could run side by side during the move, then transitioned traffic between firewalls.

delivered on time with zero service disruption

Access Review: centralized access auditing

Go · agent-server · Active Directory · PostgreSQL · OpenVPN · Linux

Context

Local accounts live on thousands of Linux VMs and bare-metal servers, in PostgreSQL databases and on OpenVPN instances. Inactive and orphaned accounts are a security risk that manual reviews do not catch.

Solution

Designed and built an agent-server tool, from the initial proposal to a full second-generation rewrite. Agents collect local accounts from every source; the server reconciles them against Active Directory and surfaces inactive and orphaned accounts. Written in Go with AI-assisted development; I owned the architecture, the agent-server protocol and the reconciliation logic.

The hard part

Consistent, low-overhead collection across very different sources at the scale of several thousand instances.

100% of infrastructure resources covered by automated access review

Bot management

Cloudflare Bot Management · WAF · Managed Challenges · Terraform

Protected web assets against large-scale scraping, credential stuffing and other automated attacks with Cloudflare Bot Management. Cut false positives by tuning custom firewall rules and Managed Challenges on bot-score analysis, and moved bot protection into the Terraform pipeline so every production environment gets the same policy.

Infrastructure & network as code

Terraform · GCP · Cloudflare · Junos · NETCONF

Fleet memory reclamation

Linux · Proxmox · ZFS · zram · open source

Context

Hypervisors looked close to running out of memory in monitoring, which pointed towards buying more RAM. Much of that "used" memory was ZFS read cache: Proxmox's documented ARC default (10% of RAM, at most 16 GiB) is only written at install time when the root file system is ZFS, so pools added later fall back to OpenZFS's much larger default, and many tools count ARC as used memory.

Solution

Shared

5-25% of RAM freed per hypervisor by the ARC cap; RAM utilization down ~5 percentage points from THP