Modernizing Infrastructure at UNEP-WCMC with Ansible and AWX
How UNEP-WCMC leveraged Ansible and AWX to bring their infrastructure up to date.
Team leadership · Production infrastructure · AI operations
UK Indefinite Leave to Remain · No sponsorship required
I lead DevOps at UNEP-WCMC, looking after the infrastructure behind UN Biodiversity Lab, Protected Planet and Species+. Before that, I was Staff Engineer and Platform Team Manager at Agile Analog.
13 years running production systems across an ISP, e-commerce and UN public data platforms. I combine hands-on engineering in Terraform, Ansible, Linux, Go and Python with team leadership, incident response and a growing focus on AI infrastructure.
I build tools that make operations easier: self-service VMs, fleet automation and AI agents with human approval. I have worked with LLMs since 2023, use coding agents daily and test their output before it ships.
Lead a DevOps team of two at the UN Environment Programme's biodiversity centre, owning almost 60 cloud hosts across AWS, Azure and Linode, a 5-node Proxmox cluster, and the platforms behind UN Biodiversity Lab (UNBL), Protected Planet and Species+.
Joined as Platform Engineer in 2021, promoted to Senior DevSecOps Engineer in 2022 and to Staff Engineer and Platform Team Manager in 2023. Left in June 2023 when most of the Platform team was made redundant.
Cloud Engineer, then DevOps Engineer from 2018, in the Core Network team of an ISP serving hundreds of thousands of customers.
Production experience across infrastructure, delivery, security and AI.
Terraform, Ansible, AWX, Bicep, cloud-init, Puppet, Chef
Docker, Kubernetes (Linode LKE, Kustomize, cert-manager, kubectl), Proxmox VE, VMware
GitHub Actions (reusable workflows, self-hosted runners), GitLab CI, Kamal, Jenkins
AWS (EC2, S3, VPC, IAM, RDS, Route 53, CloudTrail, Security Hub), Azure, Linode
LLM agents, human-in-the-loop guardrails, MCP, multi-agent workflows, Claude Code, Anthropic and OpenAI APIs, RAG, self-hosted LLM inference (llama.cpp on GPUs), prompt caching
Zabbix, Prometheus, Grafana, Graylog
HashiCorp Vault, Teleport, NetBird, Active Directory / LDAP, SAST (gosec), UK-GDPR
Cloudflare (DNS, Tunnels, Workers, WAF, API), NGINX, Traefik, WireGuard, FreeRADIUS, PowerDNS
Go, Python, Bash, SQL
PostgreSQL, PostGIS, MySQL, MongoDB
Building with LLMs since 2023. My focus is useful operational tools, private inference and human control over infrastructure actions.
Set up an on-premises LLM platform on pooled GPUs and a RAG prototype over UNBL’s dataset catalogue, keeping the data on UNEP-WCMC hardware.
AI agents investigate incidents and tickets through SSH, Ansible and OpenTofu. Deployed at UNEP-WCMC with the organisation’s permission.
A compute marketplace for idle machines, with gVisor and Kata sandboxes. Its LLM deployment pipeline reads a repository, writes its Dockerfile and repairs failed builds.
Explore Hostirr →For my personal projects, I set the design and test every change; AI coding agents write most of the code.
CCNA · 2005 · Lapsed
CTI, Cape Town · 2003
English & Afrikaans · Driving licence
Engineering in practice
From infrastructure to incident response.
Built, operated and improved in production.
Monitoring is useful when it helps someone understand what broke and what to do next.
Selected production work
A mixed cloud estate needs a consistent operational picture, especially during an incident.
Rebuilt monitoring at UNEP-WCMC, replacing Nagios with Zabbix alongside Prometheus, Grafana and Graylog. Led responses to production and security incidents, including a botnet attack.
Introduced written incident reports so the team could learn from each response.
Infrastructure defined in code, secrets managed centrally, and environments the team can reproduce.
Selected production work
An inherited estate had inconsistent configuration and servers missing from the inventory.
Introduced fleet-wide Ansible with AWX on Linode LKE and Vault-backed secrets. Rebuilt UN Biodiversity Lab on Azure with Terraform and Bicep, including staging, production, backup and restore.
The first complete inventory uncovered untracked servers. A separate Species+ stack was provisioned alongside live production without downtime.
Reusable delivery workflows that reduce repeated setup and make production deployments easier to maintain.
Selected production work
Multiple production projects needed a common deployment approach and resilient runner infrastructure.
Wrote the organisation-wide GitHub Actions library for Kamal 2 deployments, hardened against CI injection. Built custom Ruby and Kamal images for two redundant sets of four self-hosted runners on Kubernetes.
A shared deployment library used by 8+ production projects, alongside a one-command developer setup.
Identity, secrets and isolation built into the way infrastructure is operated.
Selected production work
Engineers need practical access to systems while sensitive data and credentials remain controlled.
Introduced auditable zero-trust access with Teleport and NetBird. Built isolated AWS environments at Agile Analog and led a UK-GDPR disposal audit of a legacy account holding approximately 444 TB.
Read-only audit tooling across five AWS accounts, plus self-service Proxmox VMs using Active Directory sign-in.
Production operations experience applied to AI: private inference, useful tools and explicit approval.
OpenUniverse · Personal project
An agent needs enough access to investigate a real incident, with clear limits on the actions it can take.
Built OpenUniverse with provider-agnostic agents, SSH, Ansible and OpenTofu tools, logged calls and secret redaction. Deployed it at UNEP-WCMC with permission, using self-hosted models.
Human-in-the-loop operational workflows, backed by 1,000+ automated tests. Also built an on-premises inference platform and a RAG prototype over UNBL’s catalogue.
On your team
The work extends beyond infrastructure: helping colleagues get started, taking responsibility for change and giving teams tools they can use independently.
Mentored a junior engineer, created the team’s hiring test and onboarded a new colleague at UNEP-WCMC. Previously line-managed the Platform team at Agile Analog.
Mentoring · Hiring · Team leadershipRan the estate between hires, kept public URLs working through a major GIS re-platform, and upgraded Proxmox and GitLab across major versions without data loss.
Continuity · Migrations · Operational ownershipBuilt a one-command developer setup, shared deployment workflows used by 8+ production projects, and self-service virtual machines with Active Directory sign-in.
Developer experience · Reusable tools · Self-serviceA more reliable platform, a better developer experience, or a practical route into AI operations—these are the problems I work on.
Ely, Cambridgeshire · UK Indefinite Leave to Remain · No sponsorship required
An AI-powered DevOps operations platform for investigating tickets and acting through secure operational tools.
Explore →A community cloud marketplace where hosts offer spare compute and renters launch metered, isolated spaces.
Explore →A real-time multiplayer drawing and guessing game with public matchmaking, private rooms, and live canvases.
Explore →How UNEP-WCMC leveraged Ansible and AWX to bring their infrastructure up to date.
For DevOps, platform and AI infrastructure roles, tell me about the team, the challenge and the location or working arrangement.