Job Openings Tech Lead, Machine Learning

About the job Tech Lead, Machine Learning


Our Client is a AI start-up company. As the Technical Lead, Machine Learning, you'll own the execution layer of the company's AI systems, translating research into reliable, scalable, production-grade machine learning solutions. This role sits at the intersection of research, infrastructure, and product engineering, with responsibility for making machine learning systems trainable, deployable, observable, and performant in production.

Key Responsibilities

  • Own the end-to-end machine learning lifecycle, including data pipelines, model training, evaluation, inference, deployment, and continuous improvement.
  • Fine-tune and adapt foundation models using modern techniques such as LoRA, QLoRA, SFT, DPO, and model distillation.
  • Design and operate scalable GPU-based inference systems while balancing latency, cost, and reliability.
  • Build and maintain data pipelines supporting both synthetic and real-world training datasets.
  • Develop evaluation frameworks covering model performance, robustness, safety, and bias.
  • Optimize production deployments, including GPU utilization, memory efficiency, latency reduction, and scaling strategies.
  • Collaborate closely with application engineers to integrate machine learning systems into backend, mobile, and desktop products.
  • Drive rapid iteration by making pragmatic engineering decisions based on production feedback.

Technology Stack: Python, PyTorch/JAX, GPU-based training and inference systems

Success Measures

You'll be successful in this role by:

  • Delivering production-ready ML systems with measurable performance improvements.
  • Building reliable, maintainable training and inference pipelines.
  • Quickly diagnosing and resolving production issues.
  • Providing technical leadership and helping other engineers deliver high-quality ML systems.
  • Continuously improving models and infrastructure based on production usage and user feedback.

Ideal Background

The team is looking for someone who:

  • Has built and deployed real-world machine learning systems used in production.
  • Has experience working with large language models and understands their limitations and failure modes.
  • Writes production-quality, scalable code.
  • Takes ownership and can work independently in an ambiguous environment.
  • Enjoys collaborating within a small, high-trust engineering team.