Job Description:
Director of Infrastructure Engineering
https://jobs.micro1.ai/post/4ec0692c-6714-4c97-aa9d-406b0abd684e?referralCode=1761f2a
1-0970-4c4f-a273-a1b557d5ad9f&utm_source=referral&utm_medium=share&utm_campaign=job_referral
Job Type: Full-time
Location: Remote
The Role
We're looking for a Director of Infrastructure Engineering to build and scale the platform that powers production AI systems. You'll own the strategy and execution behind our cloud infrastructure, developer platform, and reliability practices, ensuring our systems remain secure, observable, and resilient as we grow.
This is a hands-on leadership role for someone who enjoys operating at every level—from defining long-term infrastructure strategy to solving production incidents and mentoring high-performing engineers.
What You'll Do
Own the architecture and evolution of our multi-cloud infrastructure across AWS and GCP, optimizing for scalability, reliability, security, and cost.
Lead and grow a high-performing Infrastructure/Platform Engineering team while establishing a strong engineering culture centered on ownership, operational excellence, and continuous improvement.
Build infrastructure as code using Terraform (or equivalent), ensuring reproducible, version-controlled, and auditable environments.
Design and improve CI/CD systems that enable fast, reliable, and secure software delivery.
Build world-class observability across metrics, logs, traces, and alerting to proactively detect and resolve production issues.
Define and drive reliability practices, including incident response, on-call operations, SLOs, error budgets, disaster recovery, and blameless postmortems.
Partner closely with Security and Engineering leadership to embed security by default and maintain compliance with frameworks such as ISO 27001, SOC 2, and CMMC.
Continuously improve the developer experience by reducing operational friction through automation and platform tooling.
What We're Looking For
8+ years building and operating production infrastructure, platform engineering, DevOps, or SRE systems.
3+ years leading engineering teams in high-growth environments.
Deep expertise with AWS, GCP, or multi-cloud production environments.
Strong experience with Terraform, Kubernetes, containers, and modern CI/CD platforms.
Proven experience building highly available, observable, and resilient production systems.
Strong understanding of infrastructure security, compliance, and operational risk management.
Experience scaling engineering organizations and infrastructure in fast-moving startup environments.
Excellent communication skills with the ability to influence technical strategy across engineering and executive stakeholders.
Preferred
Experience supporting AI/ML infrastructure or large-scale data platforms.
Familiarity with model training, inference, evaluation, or data pipeline infrastructure.
Experience operating regulated cloud environments such as FedRAMP, GovCloud, or CMMC Level 2.
Contributions to platform engineering, open-source infrastructure, or developer productivity initiatives.