Arash Haghighat

Tech Lead and Reliability Engineer The Netherlands

Tech Lead and Reliability Engineer with 15+ years of experience delivering scalable, secure infrastructure across startups and enterprise environments. I lead by example, drive platform and DevEx improvements, and help teams adopt cloud-native, reliable systems through automation, observability, and a strong DevOps mindset.

Experience

to

Senior Platform Engineer

Just Eat Takeaway, via Xebia

I am building an ML and Data platform on top of Kubernetes, extending its APIs with custom resources and operators that abstract away cloud and third-party services such as VertexAI, SageMaker, MLflow, Amazon Glue, and Flink. The platform includes a control-plane API, an interactive frontend, an SDK for programmatic integration, and a kubectl-style CLI, providing data and ML teams with a self-service experience across AWS and GCP.

to

Staff Reliability Engineer

DSM-Firmenich, via Xebia

I architected and implemented a continuous compliance monitoring platform for ~100 AWS accounts against well-known security benchmarks and the AWS Well-Architected framework. I built the deployment architecture on ECS Fargate and Lambda, enabling historical trend tracking so teams can measure remediation impact over time. Tuned controls to reduce false-positive noise and focus engineering attention on actual risk.

to

Technical Lead

Dutch Railways (NS), via Xebia

I enabled the team building NS's mission-critical, multi-region Azure landing zone. I was involved in architectural and strategic decisions while ensuring platform scalability and resilience. I worked with platform stakeholders and internal customers to operate and monitor their services, ran incident response exercises, conducted chaos engineering experiments, and helped define onboarding and support processes. I actively advocated for platform engineering, DevOps practices, and improvements to developer experience across the organization.

to

Senior Platform Engineer

Dutch Railways (NS), via Xebia

I helped build a self-service container platform using AKS and Azure. One of the major successes was implementing and migrating a cloud-agnostic, mission-critical Harbor registry, one of the largest multi-region, multi-cloud instances, into the new platform architecture and landing zones in 9 months. This platform provides customers with reusable components and golden paths to build and run secure, reliable, and scalable services. Observability shipped, with an OpenTelemetry Collector and auto-instrumentation so teams emitted consistent traces and metrics by default.

to

Consultant and Trainer

Xebia

Hands-on consulting across SRE and platform engineering initiatives. Engagements include multi-cloud Kubernetes (AWS, GCP, Azure), IaC building blocks, large-scale migrations, compliance automation on ECS/Lambda, and observability implementations. I focus on delivery first, and on improving how teams work along the way.

to

Staff Site Reliability Engineer

Shypple

I designed and deployed new infrastructure, migrated legacy workloads to GKE, and led the rollout of CI/CD pipelines, Helm packages, and platform-wide Datadog monitoring. I improved platform resilience and agility by consolidating CI/CD on CircleCI and implementing automated unit/e2e testing, as well as supply chain security. These efforts raised platform availability from 99.8% to 99.99%. I also implemented regular pen tests, disaster recovery procedures, and infrastructure-as-code to ensure robust, secure, and scalable operations.

to

Head of Infrastructure & Operations

AloPeyk

I led two cross-functional teams while coordinating with others to manage infrastructure across on-prem and cloud (GCP). I successfully reduced infrastructure costs by 40% by planning and executing a migration of ~300 monolithic VMs to Kubernetes microservices in a hybrid architecture. Before transitioning to the cloud, I redesigned our virtualization layer (Proxmox) to ensure GCP compatibility. My work covered everything from CI/CD and DevOps toolchains to observability, high availability, and performance tuning.

to

Senior Site Reliability Engineer

AloPeyk

I implemented HAProxy as a load balancer and API gateway, scaling traffic up to 500K RPM. I also built fully automated GitLab CI pipelines that cut deployment time from a full day to under an hour, including QA gates and approval steps. I managed infrastructure across two data centers: backup, patching, monitoring, and scaling. I also operated clustered services such as MongoDB, Redis, and RabbitMQ, ensuring high availability and reliability.

to

Senior DevOps Engineer

Pasargad Electronic Payment

I reduced the ops team's workload by 30% by automating server provisioning and configuration with Ansible. I also cut critical infrastructure deployment time from nearly a month to just a few hours by implementing GitLab CI pipelines. This helped accelerate agile delivery, improve consistency, and reduce human error in the production process.

to

DevOps Engineer

Pasargad Electronic Payment

I implemented Zabbix for comprehensive infrastructure monitoring and maintained internal tools like JFrog Artifactory, SonarQube, GitLab, Jira, and the ELK stack. I led the migration from legacy version control systems (SVN, TFSVC) to Git, establishing scalable branching and tagging strategies and training developers on modern workflows.

to

Earlier roles

Various companies

Supported telecom and enterprise infrastructure, with a focus on Linux systems, 2G/3G VAS platforms, and automation using tools like Zabbix, Ansible, LDAP, DNS, MySQL, and CI/CD pipelines (GitLab, Redmine).

Designed and deployed Virtual Desktop Infrastructure (VDI) solutions and private datacenter automation for in-house delivery.

Developed smart home automation apps and early mobile software using Windows CE, C#, and socket programming.

Contributed to cross-functional projects in telecom (MTN Irancell, Rahnema) and embedded systems, gaining strong foundations in infrastructure, deployment pipelines, and production operations.

Last updated June 2026Download PDF