Job Openings Staff DevOps Engineer (GCP)

About the job Staff DevOps Engineer (GCP)

About the role

We're looking for a Staff Platform Engineer to own the infrastructure, reliability, and security foundation that the rest of the engineering organization builds on. This is a high-autonomy, high-leverage role: you'll be the primary architect and operator of our cloud infrastructure, our GitOps deployment pipeline, our observability stack, and a meaningful share of our production incident response — with direct influence over how the company balances reliability, cost, and security as it scales.

This role suits someone who thinks in systems, not tickets. You'll move fluidly between infrastructure-as-code, live production debugging, cost and vendor governance, and security hardening — often in the same week — and you'll be trusted to make judgment calls on ambiguous, cross-cutting problems without a large team behind you. You should be energized by ownership and comfortable being the final word on how the platform works, while also building the documentation, automation, and access model that lets others operate it safely alongside you.

Responsibilities:

  • Own infrastructure-as-code for our GCP environment (dev/staging/production) using Terraform/Terragrunt — networking, Cloud SQL, GKE, Cloud CDN/NAT/DNS, Artifact Registry, Secret Manager, and IAM — and continuously reduce the blast radius and bottleneck risk of infrastructure changes by extending safe self-service access to other engineers.
  • Run GitOps for the full service fleet: Kubernetes/Helm charts and ArgoCD-managed deployments across every production service, including ingress/gateway configuration, autoscaling, node-pool and workload isolation strategy, and emergency-change procedures when automation needs a human override.
  • Build and maintain observability as code: design and own monitors, SLOs, dashboards, and logging pipelines (e.g., Datadog or equivalent) so that degradations are caught before customers notice, and drive down alert noise and false positives over time.
  • Lead incident response: act as incident commander for production incidents, run live triage and mitigation, drive root-cause analysis grounded in real telemetry (not assumptions), and turn findings into durable fixes, monitors, and postmortems — while improving the incident process itself (escalation paths, on-call coverage, postmortem discipline).
  • Drive cost governance across cloud and vendor spend: monitor and optimize GCP, observability, and third-party SaaS costs; right-size infrastructure; negotiate vendor contracts; and build the tooling/dashboards needed to make spend visible and actionable across teams.
  • Own security and identity infrastructure: maintain SSO/identity provider infrastructure and service-to-service authentication, respond to and remediate vulnerabilities (including critical-severity findings) with urgency, and keep authentication/authorization systems patched and hardened.
  • Support cross-functional platform needs: administer access and provisioning across the company's core SaaS and cloud tooling, support new product launches (including partner/embedded integrations) with capacity planning and load-bearing infrastructure changes, and partner with data/analytics teams on the SQL and pipeline infrastructure that powers cost and quality reporting.

Requirements:

    • 6+ years in platform engineering, SRE, or DevOps roles, including significant ownership of production infrastructure at a company operating at meaningful scale.
    • Deep hands-on experience with a major cloud provider (GCP strongly preferred) and infrastructure-as-code tooling (Terraform/Terragrunt or equivalent).
    • Strong Kubernetes experience: Helm, GitOps tooling (ArgoCD or equivalent), ingress/gateway configuration, and troubleshooting live cluster issues under pressure.
    • Proven experience building and operating observability systems — monitors, SLOs, dashboards, and log pipelines — not just consuming a vendor's defaults.
    • Direct experience as an incident commander: leading live troubleshooting, writing rigorous postmortems backed by metrics/logs (not guesses), and following through on remediation.
    • Working knowledge of authentication/identity infrastructure (SSO/IdP systems, service-to-service auth) and a security-first mindset, including vulnerability triage and remediation.
    • Comfort with cost analysis and vendor/infrastructure spend optimization as a routine part of the job, not a one-off project.
    • Strong SQL skills (BigQuery or equivalent) for infrastructure cost analysis, audit logging, and data-driven decision-making.
    • Demonstrated ability to operate with a high degree of autonomy, prioritize across a wide and often reactive workload, and communicate infrastructure risk/trade-offs clearly to non-infrastructure stakeholders.

    Additional information

    - Work with some of the most dynamic US tech companies, building and iterating on new features and platforms.

    - Long-term projects with real technical challenges.

    - Fully remote work with flexible hours.

    - Collaboration flexibility: We work with B2B (PFA/SRL) contracts.

    - 30 paid days off per year.

    - We provide equipment as needed (laptop, desktop, etc.).

    - Continuous learning: We sponsor career-improving courses, seminars, and certifications.

    - Opportunity for annual business visits to the US, depending on project needs.

    Get picky and choose a career that matches your mindset and lifestyle. Team up with a company that encourages you to do more and gives you the flexibility you need!