# Arash Haghighat, CV

Tech Lead and Reliability Engineer, The Netherlands

Tech Lead and Reliability Engineer with 15+ years of experience delivering scalable, secure infrastructure across startups and enterprise environments. I lead by example, drive platform and DevEx improvements, and help teams adopt cloud-native, reliable systems through automation, observability, and a strong DevOps mindset.

- Email: hey@a12t.co
- LinkedIn: https://www.linkedin.com/in/a12t
- GitHub: https://github.com/irLinja
- PDF: https://a12t.co/files/CV-Arash-Haghighat.pdf

## Experience

### Senior Platform Engineer (Feb 2026 to present)

Just Eat Takeaway, via Xebia

I am building an ML and Data platform on top of Kubernetes, extending its APIs with custom resources and operators that abstract away cloud and third-party services such as VertexAI, SageMaker, MLflow, Amazon Glue, and Flink. The platform includes a control-plane API, an interactive frontend, an SDK for programmatic integration, and a kubectl-style CLI, providing data and ML teams with a self-service experience across AWS and GCP.

### Staff Reliability Engineer (Dec 2025 to Jan 2026)

DSM-Firmenich, via Xebia

I architected and implemented a continuous compliance monitoring platform for ~100 AWS accounts against well-known security benchmarks and the AWS Well-Architected framework. I built the deployment architecture on ECS Fargate and Lambda, enabling historical trend tracking so teams can measure remediation impact over time. Tuned controls to reduce false-positive noise and focus engineering attention on actual risk.

### Technical Lead (Apr 2025 to Dec 2025)

Dutch Railways (NS), via Xebia

I enabled the team building NS's mission-critical, multi-region Azure landing zone. I was involved in architectural and strategic decisions while ensuring platform scalability and resilience. I worked with platform stakeholders and internal customers to operate and monitor their services, ran incident response exercises, conducted chaos engineering experiments, and helped define onboarding and support processes. I actively advocated for platform engineering, DevOps practices, and improvements to developer experience across the organization.

### Senior Platform Engineer (Jul 2023 to Apr 2025)

Dutch Railways (NS), via Xebia

I helped build a self-service container platform using AKS and Azure. One of the major successes was implementing and migrating a cloud-agnostic, mission-critical Harbor registry, one of the largest multi-region, multi-cloud instances, into the new platform architecture and landing zones in 9 months. This platform provides customers with reusable components and golden paths to build and run secure, reliable, and scalable services. Observability shipped, with an OpenTelemetry Collector and auto-instrumentation so teams emitted consistent traces and metrics by default.

### Consultant and Trainer (Jul 2023 to present)

Xebia

Hands-on consulting across SRE and platform engineering initiatives. Engagements include multi-cloud Kubernetes (AWS, GCP, Azure), IaC building blocks, large-scale migrations, compliance automation on ECS/Lambda, and observability implementations. I focus on delivery first, and on improving how teams work along the way.

### Staff Site Reliability Engineer (Aug 2021 to Jul 2023)

Shypple

I designed and deployed new infrastructure, migrated legacy workloads to GKE, and led the rollout of CI/CD pipelines, Helm packages, and platform-wide Datadog monitoring. I improved platform resilience and agility by consolidating CI/CD on CircleCI and implementing automated unit/e2e testing, as well as supply chain security. These efforts raised platform availability from 99.8% to 99.99%. I also implemented regular pen tests, disaster recovery procedures, and infrastructure-as-code to ensure robust, secure, and scalable operations.

### Head of Infrastructure & Operations (Apr 2019 to Aug 2021)

AloPeyk

I led two cross-functional teams while coordinating with others to manage infrastructure across on-prem and cloud (GCP). I successfully reduced infrastructure costs by 40% by planning and executing a migration of ~300 monolithic VMs to Kubernetes microservices in a hybrid architecture. Before transitioning to the cloud, I redesigned our virtualization layer (Proxmox) to ensure GCP compatibility. My work covered everything from CI/CD and DevOps toolchains to observability, high availability, and performance tuning.

### Senior Site Reliability Engineer (Sep 2018 to Apr 2019)

AloPeyk

I implemented HAProxy as a load balancer and API gateway, scaling traffic up to 500K RPM. I also built fully automated GitLab CI pipelines that cut deployment time from a full day to under an hour, including QA gates and approval steps. I managed infrastructure across two data centers: backup, patching, monitoring, and scaling. I also operated clustered services such as MongoDB, Redis, and RabbitMQ, ensuring high availability and reliability.

### Senior DevOps Engineer (Jul 2018 to Mar 2020)

Pasargad Electronic Payment

I reduced the ops team's workload by 30% by automating server provisioning and configuration with Ansible. I also cut critical infrastructure deployment time from nearly a month to just a few hours by implementing GitLab CI pipelines. This helped accelerate agile delivery, improve consistency, and reduce human error in the production process.

### DevOps Engineer (Apr 2017 to Jul 2018)

Pasargad Electronic Payment

I implemented Zabbix for comprehensive infrastructure monitoring and maintained internal tools like JFrog Artifactory, SonarQube, GitLab, Jira, and the ELK stack. I led the migration from legacy version control systems (SVN, TFSVC) to Git, establishing scalable branching and tagging strategies and training developers on modern workflows.

### Earlier roles (Jun 2011 to Mar 2017)

Various companies

Supported telecom and enterprise infrastructure, with a focus on Linux systems, 2G/3G VAS platforms, and automation using tools like Zabbix, Ansible, LDAP, DNS, MySQL, and CI/CD pipelines (GitLab, Redmine).

Designed and deployed Virtual Desktop Infrastructure (VDI) solutions and private datacenter automation for in-house delivery.

Developed smart home automation apps and early mobile software using Windows CE, C#, and socket programming.

Contributed to cross-functional projects in telecom (MTN Irancell, Rahnema) and embedded systems, gaining strong foundations in infrastructure, deployment pipelines, and production operations.

## Skills

- **Leadership:** Team mentoring and onboarding, cross-team collaboration, roadmap planning, stakeholder management
- **Platform:** Platform engineering, self-service tooling, DevEx advocacy, container lifecycle
- **Cloud and infrastructure:** AWS, Azure, GCP, virtualization, Kubernetes, Helm, Harbor, GitOps, Terraform
- **Delivery:** GitHub, GitLab, Flux, Ansible, build pipelines, e2e testing, supply chain security
- **Observability:** OpenTelemetry, Prometheus, Grafana, Datadog, ELK Stack
- **Reliability:** Incident response, SLOs, chaos engineering, disaster recovery

## Certificates

- **Kubestronaut**: Cloud Native Computing Foundation, verified on Credly: https://www.credly.com/badges/ccf45199-8ebb-49e8-9357-20a7642fcb69
- **Cilium Cluster Mesh**: Isovalent, verified on Credly: https://www.credly.com/badges/538c4791-edce-4be4-9bb9-c8f741307684

## Languages

English, Farsi/Persian

## Volunteer work

Organizer, Site Reliability Engineering NL meetup (Apr 2024 to present)

Together with friends in the SRE world, I manage the Site Reliability Engineering NL Meetup group: planning, curating, and running our community events.

## Education

Azad University, BSc Software Engineering (Sep 2016)

Last updated June 2026.
