Senior DevOps Engineer
Job Description:
About the role
We’re looking for a Senior DevOps Engineer who can see the big picture and turn it into a practical platform strategy. You’ll own the infrastructure, reliability, security, with a strong focus on making everything faster, more efficient, and more cost-effective. This is a role for someone who looks beyond “keeping things running.”. You’ll need strong technical judgment, a strategic mindset, and the ability to work across teams to drive meaningful improvements. Consulting experience is important as we’re looking for someone who can quickly understand a complex environment, spot the biggest opportunities, and confidently drive change.
Responsibilities:
- Manager IaC for our GCP environment using Terraform/Terragrunt, including GKE, Cloud SQL, networking, CDN/NAT/DNS, IAM, Artifact Registry, and Secret Manager
- Manage Kubernetes setup with Helm and ArgoCD, including deployments, ingress, autoscaling, node-pools, workload isolation, and emergency changes.
- Build and maintain observability with Datadog (or similar): monitors, SLOs, dashboards, and logging.
- Lead production incidents, from live troubleshooting and mitigation through root-cause analysis, postmortems, and long-term fixes.
- Own cloud and SaaS cost optimization, including infrastructure right-sizing, vendor spend, and tooling to make costs visible across teams.
- Own identity and security infrastructure, including SSO, service authentication, vulnerability remediation, and keeping critical systems patched and secure.
- Support SaaS infrastructure needs, including access management, new product launches, partner integrations, and data/SQL infrastructure for cost and quality reporting.
Requirements:
- 7+ years in DevOps, SRE, Platform Engineering or a similar role, with significant production ownership
- Strong GCP and Terraform/Terragrunt experience.
- Strong Kubernetes experience, including Helm, ArgoCD/GitOps, ingress, and troubleshooting live systems.
- Hands-on experience building observability systems: monitoring, SLOs, dashboards, and logging.
- Experience leading production incidents and driving postmortems and remediation.
- Working knowledge of SSO/identity, service authentication, and vulnerability management.
- Comfortable with cloud/vendor cost optimization and strong SQL skills (BigQuery or similar).
- Highly autonomous, able to prioritize in a fast-moving environment, and comfortable explaining technical risks and trade-offs to teams.
Nice to have
- Experience with Pulumi
- Experience in building pipelines for training models
- Familiarity with GPU/ML infrastructure like Vertex AI for supporting AI/ML workloads
- Experience running API gateways, service meshes, or CDN/edge (e.g., Kong, Envoy, Linkerd)
- Experience with distributed databases or logging pipelines tooling.
- Experience supporting B2B/SDK launches and capacity planning for unpredictable traffic.
- Scripting experience for internal platform tools
- Comfortable owning infrastructure and security in small engineering teams.
Additional information
- Work with some of the most dynamic US tech companies, building and iterating on new features and platforms.
- Long-term projects with real technical challenges.
- Fully remote work with flexible hours.
- Collaboration flexibility: We work with PFA/SRL contracts.
- 30 paid days off per year.
- We provide equipment as needed (laptop, desktop, etc.).
- Continuous learning: We sponsor career-improving courses, seminars, and certifications.
- Opportunity for annual business visits to the US, depending on project needs.
- Get picky and choose a career that matches your mindset and lifestyle. Team up with a company that encourages you to do more and gives you the flexibility you need!