Job Openings XTN-8429200 | DEVOPS ENGINEER

About the job XTN-8429200 | DEVOPS ENGINEER

We are looking for a DevOps Engineer to own and evolve the infrastructure that powers Anervea's AI platform and supporting services. You will be responsible for designing, deploying, and maintaining a reliable, secure, and cost-efficient cloud environment on AWS, with Docker-based workloads at its core.

This is a hands-on role for someone who enjoys the full lifecycle of infrastructure work — from provisioning and automation to monitoring, incident response, and continuous improvement. You will partner closely with engineering, AI/ML, and product teams to ship reliably and scale gracefully.

Tech Stack You’ll Support

  • You will be deploying, scaling, and maintaining infrastructure for the following stack:
  • Mobile applications built with Flutter.
  • Web frontend built with React.
  • Backend services built with Python (FastAPI).
  • AI/ML workloads including LLM inference, RAG pipelines, and supporting services.
  • Cloud platform: AWS (primary).
  • Containerization with Docker; orchestration via ECS or EKS.
  • Databases and storage: PostgreSQL/RDS, S3, vector stores, and caching layers (Redis).

Key Responsibilities

AWS Infrastructure

  • Design, provision, and maintain AWS infrastructure across services including EC2, ECS/EKS, S3, RDS, VPC, IAM, CloudFront, Route 53, ELB/ALB, Lambda, and CloudWatch.
  • Implement Infrastructure as Code (IaC) using Terraform, CloudFormation, or AWS CDK to ensure repeatable and version-controlled environments.
  • Manage networking (VPCs, subnets, security groups, NAT gateways, VPN/peering) with a strong focus on security and least-privilege access.
  • Optimize AWS spend through right-sizing, reserved instances/savings plans, autoscaling, and continuous cost monitoring

Docker & Containerization

  • Build, optimize, and maintain Docker images for backend services, AI/ML workloads, and supporting tooling.
  • Manage container orchestration on Amazon ECS (Fargate/EC2) or EKS (Kubernetes), including service definitions, task scaling, and rolling deployments.
  • Maintain private container registries (ECR) with tagging, scanning, and lifecycle policies.
  • Troubleshoot container-level issues across networking, storage, and runtime performance.
    CI/CD & Automation
  • Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline / CodeBuild.
  • Automate testing, container builds, image scanning, and zero-downtime deployments to staging and production environments.
  • Standardize deployment workflows across services so engineers can ship safely and quickly.

Monitoring, Reliability & Incident Response

  • Implement and maintain observability across the stack using CloudWatch, Prometheus, Grafana, ELK/OpenSearch, Datadog, or similar.
  • Define and track SLOs/SLAs, set up actionable alerting, and reduce noise in on-call workflows.
  • Lead incident response, conduct root-cause analyses, and drive postmortems with clear follow-up actions.
  • Plan and test backup, disaster recovery, and high-availability strategies.


Security & Compliance

  • Apply security best practices across IAM, secrets management (AWS Secrets Manager / Parameter Store), encryption at rest and in transit, and network security.
  • Manage SSL/TLS certificates, WAF rules, and DDoS protection.
  • Support compliance, audit, and data-protection requirements relevant to healthcare/AI workloads (e.g., data residency, access logging)

Generative AI Infrastructure

  • Support and maintain infrastructure for LLM and generative AI workloads on AWS, including Amazon Bedrock, SageMaker, and self-hosted model endpoints.
  • Help manage GPU-based compute (EC2 P/G instances, Inferentia) for model inference and training workloads, with a focus on cost and throughput optimization.
  • Set up and operate vector databases (OpenSearch, pgvector, Pinecone) and supporting RAG infrastructure.
  • Build secure, observable pipelines for model deployment, versioning, and rollback.
  • Monitor token usage, inference latency, and model-serving costs; implement guardrails and rate limiting where needed.
  • Manage API gateways and authentication for AI endpoints exposed to internal and external consumers.
    Collaboration
  • Work closely with backend, frontend, and AI/ML engineers to understand workload requirements and provide reliable infrastructure.
  • Document infrastructure, runbooks, and operational procedures so the team can operate confidently.
  • Mentor engineers on DevOps practices, deployment hygiene, and cloud-cost awareness.