● available for services and projects
DevSecOps & SRE for reliable cloud platforms
I build and operate production infrastructure for startups and engineering teams — Kubernetes, multi-cloud, CI/CD, and observability.
status taking platform audits, automation work, and reliability engagements
services
Cloud Infrastructure Design
AWS, Azure, and GCP infrastructure defined as code — networking, compute, DNS, and secrets management.
Kubernetes & Platform Engineering
EKS, AKS, and GKE platforms with GitOps delivery, golden paths, and self-service environments.
CI/CD & Automation
Deployment pipelines on GitHub Actions, GitLab CI, or Jenkins — container builds, image scanning, release gates.
Observability & Reliability
SLOs, dashboards, incident runbooks, and alerting that pages for real problems — Prometheus, Grafana, New Relic.
Cloud Cost & Performance Optimization
Find where the spend actually goes — idle capacity, autoscaling gaps, logging and observability bills — then cut it.
Secure Your CI/CD Pipeline — DevSecOps
Security that runs on every merge instead of a quarterly audit — each control wired in as a pipeline gate with an owner and a documented bypass path. What gets wired in:
→ SAST on every pull request
→ SCA / dependency CVE gates
→ Container image scanning before push
→ IaC scanning: Terraform, Pulumi, Helm
→ Secrets detection: pre-commit + CI enforcement
→ SBOM generation and image signing
→ Keyless OIDC auth — no long-lived cloud keys in CI
→ Least-privilege runners and branch protection
→ Policy-as-code merge gates with documented exceptions
Migration Services
Getting off legacy platforms without a bad weekend — cutover plans sized to your downtime tolerance, with a rollback at every step. Work I take:
→ On-premises to Cloud migration
→ EC2 to Amazon EKS migration
→ EC2 to Amazon ECS Fargate migration
→ GitLab to GitHub migration
→ Legacy CI/CD to GitHub Actions migration
→ Database migration: CrunchyDB to CNPG
→ Containerization & Kubernetes workloads
→ Infrastructure as Code with Terraform
→ Observability: Prometheus, Grafana, CloudWatch, ELK
how i work
1
Assess
Map the current infrastructure, dependencies, risks, and downtime tolerance.
2
Design
Agree the target architecture and the sequence to get there.
3
Implement
Build it — workloads, pipelines, data, and the observability stack.
4
Handover
Verify performance and rollback readiness, write the runbooks, hand it to your team.
Know what you need, or want a second opinion first?
Start a conversation
client problems i solve
the problem
what you get
Deployments are too manual or risky
Blue-green and canary rollouts under ArgoCD, rollback at every step — change failure rate under 5%
Kubernetes is hard to operate
Running Kubernetes in production since 2017 — EKS, AKS and GKE upgrades, autoscaling, capacity, on-call
Cloud costs are increasing
Spend traced to idle capacity, autoscaling gaps and logging bills, then cut
Alerting is noisy and unreliable
SLOs and error budgets so a page means something — MTTR down 60–75%, Sev-2 incidents down 50%+
Infrastructure is not fully managed by code
Terraform and Pulumi modules with GitOps adoption — production apps under ArgoCD from day one
Security checks happen too late
SAST, SCA, IaC and container scanning wired in as merge gates
Developers wait too long for environments
Golden paths and self-service platform tooling — onboarding cut from days to hours
No clear SLOs, runbooks, or ownership model
SLOs, on-call runbooks and a named ownership model, handed to your team
case studies
3-5×
deploy frequency
EKS Multi-Region Platform Modernisation
problem
Elasticsearch/OpenSearch platforms were manually managed across EC2 and ECS with no consistent CI/CD or IaC.
solution
Migrated to EKS. Standardised CI/CD with GitHub Actions, Terraform, ArgoCD; Helm blue-green & canary, APM, feature flags.
outcome
Deployment frequency up 3–5×. Change failure rate down ~50%. Lead time cut 75%.
multi-region failover
click a region to simulate an outage
IaC
end-to-end deploys
ECS Fargate Platform Pipeline
problem
Container deployments lacked a standardised, auditable pipeline — manual ECS updates and inconsistent Terraform application across accounts.
solution
Jenkins pipeline building containers to ECR, applying Terraform modules (state in S3 with DynamoDB locking) for VPC/ECS/ALB/WAF, and deploying Fargate services behind ALB target groups with health checks.
outcome
Fully automated build-to-deploy flow with CloudTrail/Config/GuardDuty guardrails and New Relic APM/SLO coverage across the stack.
terraform plan
plan not yet run
24×
faster onboarding
Internal Platform & Golden Paths
problem
No standardised platform for service onboarding; manual provisioning led to drift and slow releases.
solution
Built Pulumi components (VPC, EKS/LKE/AKS, Grafana, CNPG) with ArgoCD app-of-apps and self-service environments.
outcome
Onboarding days → hours. Change failure rate under 5%. MTTR down 60–75%. Sev-2 down 50%+.
service onboarding
manual—
golden path—
Pulumi component + ArgoCD app-of-apps
100%
iac managed
Observability Stack — LKE + New Relic + OpenTofu
problem
Needed a reproducible, code-driven observability stack to demonstrate SLO-based alerting.
solution
Built a live K8s observability stack on Linode LKE with New Relic, OpenTofu, and GitHub Actions — dashboards as code.
outcome
Fully automated stack deployable end-to-end from a single pipeline run.
SLO error budget · 99.9%0.05%
budget holds — 30d window
Auto
agent eval gate
GCP Agentic CI/CD Pipeline
problem
CI pipelines verify code correctness but not AI agent behaviour — no automated gate for agent quality/regressions before deploy.
solution
Cloud Build pipeline runs an ADK evaluator as a sandboxed Cloud Run Job against golden datasets via Vertex AI (Gemini), streams trajectory metrics to BigQuery, and gates promotion through a pass/fail decision.
outcome
Passing builds auto-deploy the agent to Cloud Run and report status back to GitHub; failing evals block deployment before reaching production.
agent eval gatePASS
score above threshold — deploy to Cloud Run
Zero
downtime cutovers
SD-WAN & Network Modernisation
problem
Aging WAN infrastructure across multiple sites with manual cutovers and no runbook coverage.
solution
Led WAN/SD-WAN rollouts and data-centre moves. Documented runbooks and DR plans; ITIL-aligned incident/problem/change.
outcome
All site cutovers completed with zero downtime. Runbooks reduced recovery time and knowledge silos.
SD-WAN path selection
branchMPLSedgeDIAdc
click a link to cut it — traffic reconverges
experience
Platform Engineer / SRE — Hydrolix
- Defined platform roadmap around golden paths, SLOs, and SLAs
- Reduced new-service onboarding from days to hours via Pulumi + ArgoCD
- Cut MTTR 60–75%, Sev-2 incidents 50%+, change failure rate under 5%
DevSecOps / SRE — Omilia
- Operated conversational AI and NLU platforms across K8s, AWS, ELK
- Supported IVR, chat, and mobile virtual-assistant reliability
Senior DevOps — Asurion
- Built Elasticsearch platforms across EC2, ECS, EKS, OpenSearch Service
- Reduced change failure rate ~50% via blue-green, APM, feature flags
Network Engineer — Asurion
- Led 4-engineer team through ITIL incident, problem, change management
- Managed WAN/SD-WAN rollouts, DC moves, BGP/OSPF, on-call response
freelance packages
Infrastructure Audit
For teams that suspect reliability, cost, or security gaps and want evidence before committing to a build.
- Current-state review
- Risk and bottleneck report
- Prioritized recommendations
- Quick-win roadmap
CI/CD & DevSecOps Buildout
Best for teams that need secure and repeatable deployments.
- Pipeline design
- Build/test/deploy automation
- Security scan integration
- Deployment documentation
Kubernetes / Cloud Platform Build
For platform builds — golden paths, self-service environments, and an ownership model handed to your team.
- IaC modules
- Kubernetes or cloud architecture
- GitOps setup
- Monitoring and runbooks
- Handover documentation
contract works
Kubernetes Platform Audit
Placeholder: reviewed cluster topology, workload security posture, and upgrade strategy across three EKS clusters; delivered a prioritised remediation roadmap.
- Placeholder outcome — e.g. control-plane upgrade path unblocked
- Placeholder outcome — e.g. 14 findings closed within the engagement
EKS
audit
security
CI/CD Pipeline Buildout
Placeholder: replaced manual release process with GitHub Actions pipelines, Terraform-provisioned runners, and progressive delivery gates.
- Placeholder outcome — e.g. release cadence weekly to daily
- Placeholder outcome — e.g. rollback time under five minutes
GitHub Actions
Terraform
progressive delivery
Observability Stack Rollout
Placeholder: stood up Prometheus, Grafana, and Loki as code with SLO-based alerting tuned for signal over noise.
- Placeholder outcome — e.g. alert volume cut 80%, zero missed incidents
- Placeholder outcome — e.g. dashboards-as-code adopted by three teams
Prometheus
Grafana
SLOs
Cloud Cost Optimisation
Placeholder: right-sized compute, moved batch workloads to spot capacity, and added cost guardrails to the IaC pipeline.
- Placeholder outcome — e.g. monthly cloud spend down 35%
- Placeholder outcome — e.g. cost regression checks in every PR
FinOps
spot capacity
IaC guardrails
contact
Let’s improve your infrastructure
Tell me what’s breaking or slowing you down. I’ll assess the current setup and come back with a plan and a cost.