SRE / DevSecOps engineer

Infrastructure that teams
can operate.

I’m Alaric. I build cloud platforms, automate releases, and help engineering teams understand what happens in production.

Available for infrastructure and reliability projects.

EKS Multi-Region Platform Modernisation

problem
Elasticsearch/OpenSearch platforms were manually managed across EC2 and ECS with no consistent CI/CD or IaC.
solution
Migrated to EKS. Standardised CI/CD with GitHub Actions, Terraform, ArgoCD; Helm blue-green & canary, APM, feature flags.
outcome
Deployment frequency up 3–5×. Change failure rate down ~50%. Lead time cut 75%.
→ Try the simulation
stack
EKS
Terraform
ArgoCD
Helm
GitHub Actions
Prometheus
Grafana

ECS Fargate Platform Pipeline

problem
Container deployments lacked a standardised, auditable pipeline — manual ECS updates and inconsistent Terraform application across accounts.
solution
Jenkins pipeline building containers to ECR, applying Terraform modules (state in S3 with DynamoDB locking) for VPC/ECS/ALB/WAF, and deploying Fargate services behind ALB target groups with health checks.
outcome
Fully automated build-to-deploy flow with CloudTrail/Config/GuardDuty guardrails and New Relic APM/SLO coverage across the stack.
→ Try the simulation
stack ECS Fargate Jenkins
Terraform
ECR ALB WAF
New Relic
CloudWatch

Internal Platform & Golden Paths

problem
No standardised platform for service onboarding; manual provisioning led to drift and slow releases.
solution
Built Pulumi components (VPC, EKS/LKE/AKS, Grafana, CNPG) with ArgoCD app-of-apps and self-service environments.
outcome
Onboarding days → hours. Change failure rate under 5%. MTTR down 60–75%. Sev-2 down 50%+.
→ Try the simulation
stack
Pulumi
ArgoCD
EKS
LKE
AKS
Grafana
CNPG
GitHub Actions

Observability Stack — LKE + New Relic + OpenTofu

problem
Needed a reproducible, code-driven observability stack to demonstrate SLO-based alerting.
solution
Built a live K8s observability stack on Linode LKE with New Relic, OpenTofu, and GitHub Actions — dashboards as code.
outcome
Fully automated stack deployable end-to-end from a single pipeline run.
→ Try the simulation
stack
LKE
New Relic
OpenTofu
GitHub Actions
Kubernetes

GCP Agentic CI/CD Pipeline

problem
CI pipelines verify code correctness but not AI agent behaviour — no automated gate for agent quality/regressions before deploy.
solution
Cloud Build pipeline runs an ADK evaluator as a sandboxed Cloud Run Job against golden datasets via Vertex AI (Gemini), streams trajectory metrics to BigQuery, and gates promotion through a pass/fail decision.
outcome
Passing builds auto-deploy the agent to Cloud Run and report status back to GitHub; failing evals block deployment before reaching production.
→ Try the simulation
stack Cloud Build Cloud Run Vertex AI Gemini BigQuery ADK GitHub

SD-WAN & Network Modernisation

problem
Aging WAN infrastructure across multiple sites with manual cutovers and no runbook coverage.
solution
Led WAN/SD-WAN rollouts and data-centre moves. Documented runbooks and DR plans; ITIL-aligned incident/problem/change.
outcome
All site cutovers completed with zero downtime. Runbooks reduced recovery time and knowledge silos.
→ Try the simulation
stack SD-WAN BGP OSPF ITIL Nagios
Splunk
Linux

Azure Observability Platform

problem
Demonstrate how staging and production workloads connect to a shared monitoring platform, from telemetry collection to alert routing.
solution
Model resource diagnostics and application telemetry flowing into Log Analytics and Application Insights, with Azure Monitor alerts, Grafana, Azure Workbooks, and action groups. Monitoring profiles and Terraform gates describe the configuration workflow.
outcome
An interactive reference architecture for exploring signal paths, monitoring ownership, and critical, warning, and soak alert routes. This is a portfolio simulation; no live Azure deployment is claimed.
→ Explore the simulation
stack Azure Monitor Log Analytics Application Insights
Grafana
Azure Workbooks
Terraform

Scanner Platform

problem
Explain how multiple security scanners can feed a consistent verdict while adapting scan scope to repository capabilities and recent changes.
solution
A DevSecOps workflow with pull-request and deep-scan lanes for secrets, SAST, dependencies, and infrastructure. Optional container and DAST jobs feed shared reporting, where a central policy applies thresholds, exceptions, and coverage.
outcome
An interactive workflow walkthrough showing scanner selection, artifact collection, and block, warn, or report decisions. The diagram illustrates the workflow; it does not execute scans or demonstrate enforced merge protection.
→ Explore the simulation
stack
GitHub Actions
Gitleaks Semgrep Trivy Checkov TFLint ZAP

Cloud platforms & migrations

Build and maintain Kubernetes and cloud infrastructure with Terraform or Pulumi. Plan workload and database migrations with tested cutover and rollback procedures.

Delivery & security

Automate build, test, and release workflows. Add dependency, container, and infrastructure scanning; scoped access; and deployment gates your team can operate.

Reliability & operations

Define service objectives, investigate noisy alerts, and build useful dashboards and runbooks. Review capacity and cloud spend against actual workload needs.

I start with the current system and its constraints, agree the scope with your team, then implement and document the changes.

SRE — Hydrolix
June 2024 – May 2026 · Portland, Oregon
  • Defined platform roadmap around golden paths, SLOs, and SLAs
  • Reduced new-service onboarding from days to hours via Pulumi + ArgoCD
  • Cut MTTR 60–75%, Sev-2 incidents 50%+, change failure rate under 5%
DevSecOps / SRE — Omilia
Oct 2022 – Mar 2024 · Athens, Greece
  • Operated conversational AI and NLU platforms across K8s, AWS, ELK
  • Supported IVR, chat, and mobile virtual-assistant reliability
DevSecOps — Asurion
Oct 2017 – Oct 2022 · Manila
  • Built Elasticsearch platforms across EC2, ECS, EKS, OpenSearch Service
  • Reduced change failure rate ~50% via blue-green, APM, feature flags
Network Engineer — Asurion
Jun 2014 – Oct 2017 · Manila
  • Led 4-engineer team through ITIL incident, problem, change management
  • Managed WAN/SD-WAN rollouts, DC moves, BGP/OSPF, on-call response

Explore a fleet simulation, try interactive infrastructure examples, and follow architecture walkthroughs. These are demonstrations using synthetic data and simplified models.

FleetEngine — fleet intelligence

Explore a synthetic Metro Manila fleet, geospatial queries, delivery updates, and geofence events. Open the dashboard or inspect how its components fit together.

Live FleetEngine demonstrations are available on request. Explore the architecture below.

Local deployment architecture
GCP deployment architecture
MongoDB data flow
→ View repo

Azure Observability Platform

Trace telemetry from Azure workloads to operational views and alert routing. This interactive reference architecture uses a simulated platform.

Explore the observability platform

Scanner Platform

Follow pull-request scans, scheduled deep scans, optional container and DAST jobs, and the shared policy evaluation that produces a scan verdict.

Explore the scanner platform

CI/CD — architecture walkthroughs

Explore delivery pipelines, failover, service onboarding, and monitoring. Open a simulation or inspect how each architecture fits together.

AWS serverless pipeline
Multi-cloud DevSecOps pipeline
Agentic security harness
Multi-region failover

Select a region to simulate an outage and watch traffic move to the remaining regions. Reset to restore all regions.

Multi-region failover
click a region to simulate an outage
Terraform plan preview

Run a sample Terraform plan to inspect the resources that would be added, changed, or removed.

Terraform plan
plan not yet run
Service onboarding comparison

Compare manual service setup with an automated golden path. Start the comparison to see both workflows progress.

Service onboarding
manual—
golden path—
Pulumi component + ArgoCD app-of-apps
SLO error budget

Adjust the observed error rate to see how it affects a 99.9% SLO and its error budget.

SLO error budget · 99.9%0.05%
budget holds — 30d window
Agent evaluation gate

Adjust the agent score to see when the evaluation gate allows or blocks deployment.

Agent evaluation gatePASS
score above threshold — deploy to Cloud Run
Network path failover

Select a network link to simulate a failure and watch the preferred path change. Restore the links to try again.

SD-WAN path selection
branchedgedc
click a link to cut it — traffic reconverges
Have a platform problem to work through?
Tell me about your current setup, what needs to change, and your timeline.
© 2026 Alaric Acaac