Job Openings Site Reliability Engineer (Onsite, Lahore, PKR Salary)

About the job Site Reliability Engineer (Onsite, Lahore, PKR Salary)

Requirements:

  • 5+ years of experience in systems, infrastructure, or SRE engineering, operating production systems at scale.
  • Deep Linux troubleshooting skills across the OS, networking, storage, and performance, with hands-on experience working as root on production systems. Experience with Ubuntu is highly relevant, as it is used almost exclusively.
  • Hands-on experience operating GPU servers in production, including troubleshooting driver, device, and hardware-level issues, rather than only the workloads running on top of them.
  • Practical network troubleshooting experience, including diagnosing physical-layer faults.
  • Strong automation mindset with programming skills in Python or a comparable language.
  • Experience with configuration management, node provisioning, and infrastructure-as-code (IaC) using Ansible, Terraform, or similar tools.
  • Experience building observability and alerting solutions using Grafana and Prometheus.
  • Bachelor's degree in Computer Science or equivalent experience.
  • Experience operating GPU clusters or AI infrastructure at production scale.
  • Production experience with Kubernetes or Slurm; experience with both is a bonus.
  • Background in HPC or research computing.
  • Familiarity with the NVIDIA GPU stack, InfiniBand/RDMA, and NCCL.
  • Experience with CLI-based AI coding agents such as Claude Code, rather than browser-based assistants alone.
  • Contributions to open-source projects within the cloud-native, HPC, or AI infrastructure ecosystem.

Responsibilities:

  • Own the reliability, availability, and performance of production Linux GPU clusters, covering the operating system, drivers, GPUs, high-speed networking, and storage.
  • Lead deep, end-to-end troubleshooting of complex distributed systems, GPU nodes, networking, and storage issues.
  • Troubleshoot and resolve GPU rail and NCCL performance issues across multi-GPU and multi-node collective communication paths.
  • Diagnose network faults end to end, including configuration, routing, and physical-layer issues such as cabling, transceivers, and link errors.
  • Configure and maintain workload managers that schedule customer jobs, including Kubernetes, Slurm, or both, along with the identity, storage, and networking services they depend on.
  • Build automation and tooling to eliminate operational toil, using a modern language such as Python or Go to design solutions, review implementations, and redirect approaches when needed.
  • Use AI-assisted engineering tools such as Claude to accelerate automation, runbook development, and incident analysis.
  • Automate provisioning, image deployment, configuration, and remediation using Ansible and infrastructure-as-code.
  • Design and operate observability using Grafana, Prometheus, and Loki, while tuning alerts for meaningful signals and building self-healing capabilities that reduce the need for human intervention.
  • Lead incident response, on-call activities, blameless postmortems, and reliability improvements that maintain customer SLAs.
  • Partner with Platform and Systems Engineering teams on capacity planning, rollouts, and continuous improvement.

Working Hours:

8 PM - 4 AM