Easter Egg : Alley Cat

Loading CAT.EXE...
Searching for game file...

Ruan du Toit

Senior DevOps & Platform Engineer

Team leadership · Production infrastructure · AI operations

Ruan du Toit

Ruan du Toit

Senior DevOps & Platform Engineer

UK Indefinite Leave to Remain · No sponsorship required

Production infrastructure. Practical AI.

I lead DevOps at UNEP-WCMC, looking after the infrastructure behind UN Biodiversity Lab, Protected Planet and Species+. Before that, I was Staff Engineer and Platform Team Manager at Agile Analog.

13 years running production systems across an ISP, e-commerce and UN public data platforms. I combine hands-on engineering in Terraform, Ansible, Linux, Go and Python with team leadership, incident response and a growing focus on AI infrastructure.

Almost 60cloud hosts across AWS, Azure & Linode
8+ projectsusing my reusable deployment workflows
15+ incidentsproduction & security responses led

I build tools that make operations easier: self-service VMs, fleet automation and AI agents with human approval. I have worked with LLMs since 2023, use coding agents daily and test their output before it ships.

Work experience

Senior DevOps Engineer and Team Lead
UNEP-WCMC, Cambridge
Jan 2024 – present

Lead a DevOps team of two at the UN Environment Programme's biodiversity centre, owning almost 60 cloud hosts across AWS, Azure and Linode, a 5-node Proxmox cluster, and the platforms behind UN Biodiversity Lab (UNBL), Protected Planet and Species+.

  • Re-platformed the public GIS service from Windows Server to Ubuntu on AWS, taking ArcGIS Enterprise and PostgreSQL up several major versions while keeping every public URL working.
  • Rebuilt UNBL, the UN's spatial-data platform for national biodiversity planning, on Azure with Terraform and Bicep: staging and production, auto-scaling, MongoDB backup and restore, and cost controls.
  • Mentored a junior engineer, ran the estate alone between hires and onboarded a new engineer. Created the team's hiring test, one-command developer setup and auditable zero-trust access (Teleport, NetBird).
  • Introduced fleet-wide Ansible automation with AWX on managed Kubernetes (Linode LKE) and Vault-backed secrets; the first complete server inventory uncovered servers nobody was tracking.
More delivery, security & AI work
  • Stood up a VPC-isolated Species+ stack alongside live production with no downtime, and a national environmental information system on Azure, both in Terraform.
  • Wrote the org-wide GitHub Actions library for Kamal 2 deployments, used by 8+ production projects and hardened against CI injection, running on two redundant sets of four self-hosted runners on Kubernetes with custom Ruby and Kamal images.
  • Upgraded the Proxmox cluster and self-hosted GitLab (with PostgreSQL) across major versions with no data loss.
  • Developed an in-house operations console (Python and Django, later rewritten in Go and Vue) covering fleet logs, web analytics, scheduled Ansible runs, backups and cloud billing, with Slack integration.
  • Delivered self-service VMs on Proxmox: staff sign in with Active Directory and get virtual machines on demand, with Slack integration.
  • Led the response to 15+ production and security incidents (including a botnet attack), introduced written incident reports, and rebuilt monitoring with Zabbix (replacing Nagios), Prometheus, Grafana and Graylog.
  • Cut idle AWS spend, led a UK-GDPR data-disposal audit of a legacy account holding ~444 TB, and wrote a read-only audit tool covering five AWS accounts.
  • Set up an on-premises LLM platform on pooled GPUs and a RAG prototype over UNBL's dataset catalogue, keeping all data on UNEP-WCMC hardware.
  • Deployed OpenUniverse, my own AI operations platform, with the organisation's permission: its agents run on GLM-4.7 and investigate ops tickets and incidents.
Staff Engineer (DevSecOps) and Platform Team Manager
Agile Analog, Cambridge
Sep 2021 – Jun 2023

Joined as Platform Engineer in 2021, promoted to Senior DevSecOps Engineer in 2022 and to Staff Engineer and Platform Team Manager in 2023. Left in June 2023 when most of the Platform team was made redundant.

  • Line-managed the Platform team and ran its planning, Jira projects and internal helpdesk.
  • Built isolated AWS project environments with Terraform, VPC peering and FortiGate access, monitored through Security Hub and CloudTrail alerts.
  • Ran configuration and secrets management with Puppet, Foreman and HashiCorp Vault, with Zabbix monitoring.
  • Cut idle EC2 spend with a custom AMI that supported hibernation, and introduced Kasm Workspaces for high-security work.
DevOps Engineer
RSAWEB, Cape Town
Sep 2017 – Sep 2021

Cloud Engineer, then DevOps Engineer from 2018, in the Core Network team of an ISP serving hundreds of thousands of customers.

  • Ran core network services (FreeRADIUS, PowerDNS with DNSSEC, Sandvine PacketLogic) and penetration-tested internal core services.
  • Automated Linux server builds with Chef and Terraform across AWS and VMware, built CI/CD in Jenkins and GitHub Actions, and ran a shared-hosting Docker cluster.
  • Operated MySQL clusters behind ProxySQL and Percona, plus MongoDB, with Zabbix and TICK-stack monitoring.
DevOps Engineer
InventCommerce, Cape Town
Nov 2013 – Sep 2017
  • Ran hosting infrastructure and architecture planning for e-commerce clients including Pearson, Danone and Edcon.
  • Ran Jenkins CI, advised clients on Docker, and load-tested client platforms with JMeter.

Tools I use to deliver

Production experience across infrastructure, delivery, security and AI.

DevOps
DevOps100%
CI/CD80%
Cloud
Amazon Web Services80%
Infrastructure
Terraform80%
Systems
Linux100%
Programming
Python60%
Scripting
Bash Scripting60%
Database
MySQL60%
Containers
Docker60%
Orchestration
Kubernetes80%
Configuration Management
Ansible80%
Web Servers
NGINX80%
Monitoring
Zabbix80%
Version Control
GitHub80%
Gitlab80%
100%
Systems
90%
DevOps
80%
Cloud
80%
Infrastructure
80%
Orchestration
80%
Configuration Management

Full toolkit · September 2026

Infrastructure as code

Terraform, Ansible, AWX, Bicep, cloud-init, Puppet, Chef

Containers and virtualisation

Docker, Kubernetes (Linode LKE, Kustomize, cert-manager, kubectl), Proxmox VE, VMware

CI/CD

GitHub Actions (reusable workflows, self-hosted runners), GitLab CI, Kamal, Jenkins

Cloud

AWS (EC2, S3, VPC, IAM, RDS, Route 53, CloudTrail, Security Hub), Azure, Linode

AI engineering

LLM agents, human-in-the-loop guardrails, MCP, multi-agent workflows, Claude Code, Anthropic and OpenAI APIs, RAG, self-hosted LLM inference (llama.cpp on GPUs), prompt caching

Observability

Zabbix, Prometheus, Grafana, Graylog

Security and identity

HashiCorp Vault, Teleport, NetBird, Active Directory / LDAP, SAST (gosec), UK-GDPR

Networking and edge

Cloudflare (DNS, Tunnels, Workers, WAF, API), NGINX, Traefik, WireGuard, FreeRADIUS, PowerDNS

Languages

Go, Python, Bash, SQL

Data

PostgreSQL, PostGIS, MySQL, MongoDB

AI infrastructure & agent engineering

Building with LLMs since 2023. My focus is useful operational tools, private inference and human control over infrastructure actions.

At UNEP-WCMC
Self-hosted inference

Set up an on-premises LLM platform on pooled GPUs and a RAG prototype over UNBL’s dataset catalogue, keeping the data on UNEP-WCMC hardware.

Proxmox → GPUs → llama.cpp → Internal inference
Personal project · Private alpha
OpenUniverse

AI agents investigate incidents and tickets through SSH, Ansible and OpenTofu. Deployed at UNEP-WCMC with the organisation’s permission.

  • 82 MCP infrastructure tools with logged calls and redacted secrets.
  • Allow/ask/deny guardrails and human approval in Slack.
  • Go and Vue 3, with 1,000+ automated tests.
  • Prompt caching reduced input-token cost by approximately 90%.
Explore OpenUniverse →
Personal project · Invite-only alpha
Hostirr

A compute marketplace for idle machines, with gVisor and Kata sandboxes. Its LLM deployment pipeline reads a repository, writes its Dockerfile and repairs failed builds.

Explore Hostirr →

For my personal projects, I set the design and test every change; AI coding agents write most of the code.

Education & training

Cisco Certified Network Associate

CCNA · 2005 · Lapsed

IT Engineering

CTI, Cape Town · 2003

English & Afrikaans · Driving licence

Engineering in practice

DevOps

From infrastructure to incident response.
Built, operated and improved in production.

Make the problem visible.

Monitoring is useful when it helps someone understand what broke and what to do next.

15+production & security incidents led
  1. 01CollectZabbix · Prometheus
  2. 02UnderstandGrafana · Graylog
  3. 03RespondInvestigation · recovery
  4. 04ImproveWritten incident reports

Selected production work

The challenge

A mixed cloud estate needs a consistent operational picture, especially during an incident.

ZabbixPrometheusGrafanaGraylog

What I delivered

Rebuilt monitoring at UNEP-WCMC, replacing Nagios with Zabbix alongside Prometheus, Grafana and Graylog. Led responses to production and security incidents, including a botnet attack.

Introduced written incident reports so the team could learn from each response.

Build it once. Make it repeatable.

Infrastructure defined in code, secrets managed centrally, and environments the team can reproduce.

Almost 60cloud hosts across AWS, Azure & Linode
  1. 01DefineTerraform · Bicep
  2. 02ConfigureAnsible · AWX
  3. 03ProtectVault-backed secrets
  4. 04OperateInventory · scheduled runs

Selected production work

The challenge

An inherited estate had inconsistent configuration and servers missing from the inventory.

TerraformBicepAnsibleAWXVaultProxmox

What I delivered

Introduced fleet-wide Ansible with AWX on Linode LKE and Vault-backed secrets. Rebuilt UN Biodiversity Lab on Azure with Terraform and Bicep, including staging, production, backup and restore.

The first complete inventory uncovered untracked servers. A separate Species+ stack was provisioned alongside live production without downtime.

Give every team a reliable release path.

Reusable delivery workflows that reduce repeated setup and make production deployments easier to maintain.

8+production projects using shared workflows
  1. 01CommitGitHub
  2. 02RunKubernetes runners
  3. 03DeployKamal 2
  4. 04ReuseShared workflow library

Selected production work

The challenge

Multiple production projects needed a common deployment approach and resilient runner infrastructure.

GitHub ActionsKubernetesKamal 2DockerGitLab CI

What I delivered

Wrote the organisation-wide GitHub Actions library for Kamal 2 deployments, hardened against CI injection. Built custom Ruby and Kamal images for two redundant sets of four self-hosted runners on Kubernetes.

A shared deployment library used by 8+ production projects, alongside a one-command developer setup.

Make access deliberate and auditable.

Identity, secrets and isolation built into the way infrastructure is operated.

5 accountscovered by my read-only AWS audit tool
  1. 01IdentifyActive Directory
  2. 02ConnectTeleport · NetBird
  3. 03IsolateVPC boundaries
  4. 04AuditAccess · data disposal

Selected production work

The challenge

Engineers need practical access to systems while sensitive data and credentials remain controlled.

TeleportNetBirdVaultAWS IAMActive Directory

What I delivered

Introduced auditable zero-trust access with Teleport and NetBird. Built isolated AWS environments at Agile Analog and led a UK-GDPR disposal audit of a legacy account holding approximately 444 TB.

Read-only audit tooling across five AWS accounts, plus self-service Proxmox VMs using Active Directory sign-in.

Give agents tools. Keep people in control.

Production operations experience applied to AI: private inference, useful tools and explicit approval.

82 toolsexposed through OpenUniverse’s MCP server
  1. 01InvestigateIncident · ticket
  2. 02DecideAllow · ask · deny
  3. 03ApproveHuman review in Slack
  4. 04Act & recordInfrastructure tools · audit log

OpenUniverse · Personal project

The challenge

An agent needs enough access to investigate a real incident, with clear limits on the actions it can take.

GoMCPSelf-hosted LLMsllama.cppRAG

What I delivered

Built OpenUniverse with provider-agnostic agents, SSH, Ansible and OpenTofu tools, logged calls and secret redaction. Deployed it at UNEP-WCMC with permission, using self-hosted models.

Human-in-the-loop operational workflows, backed by 1,000+ automated tests. Also built an on-premises inference platform and a RAG prototype over UNBL’s catalogue.

On your team

Build the platform. Support the people.

The work extends beyond infrastructure: helping colleagues get started, taking responsibility for change and giving teams tools they can use independently.

Help engineers grow.

Mentored a junior engineer, created the team’s hiring test and onboarded a new colleague at UNEP-WCMC. Previously line-managed the Platform team at Agile Analog.

Mentoring · Hiring · Team leadership

Own the difficult transitions.

Ran the estate between hires, kept public URLs working through a major GIS re-platform, and upgraded Proxmox and GitLab across major versions without data loss.

Continuity · Migrations · Operational ownership

Make other teams more effective.

Built a one-command developer setup, shared deployment workflows used by 8+ production projects, and self-service virtual machines with Active Directory sign-in.

Developer experience · Reusable tools · Self-service

What does your team need next?

A more reliable platform, a better developer experience, or a practical route into AI operations—these are the problems I work on.

Tell me about your team

Ely, Cambridgeshire · UK Indefinite Leave to Remain · No sponsorship required

Projects

Infrastructure experience, working products

Private alpha

OpenUniverse

An AI-powered DevOps operations platform for investigating tickets and acting through secure operational tools.

Explore →
Invite-only alpha

Hostirr

A community cloud marketplace where hosts offer spare compute and renters launch metered, isolated spaces.

Explore →
Live

Paint Off

A real-time multiplayer drawing and guessing game with public matchmaking, private rooms, and live canvases.

Explore →

Blog

Latest articles

View all articles →

Contact

Let’s talk about your team

For DevOps, platform and AI infrastructure roles, tell me about the team, the challenge and the location or working arrangement.

hello@ruandutoit.online · LinkedIn ↗