Job Openings Machine Learning Engineer

About the job Machine Learning Engineer

Our Client is a AI start-up company. As a Member of Technical Staff, Machine Learning, you'll help build the core machine learning systems that power AI products in production. This role is ideal for engineers who want to deepen their machine learning expertise while developing strong engineering judgment through building, debugging, and continuously improving production ML systems.

Responsibilities

  • Build and enhance machine learning components across data pipelines, model training, evaluation, and inference.
  • Fine-tune, adapt, and deploy models as part of production AI systems.
  • Develop evaluation frameworks and testing strategies to measure model quality and performance.
  • Build and maintain data pipelines for both real-world and synthetic datasets.
  • Investigate and resolve model performance issues, training failures, and production incidents.
  • Continuously improve models and systems based on production metrics and user feedback.
  • Collaborate closely with senior ML engineers, product managers, and cross-functional teams to deliver reliable AI features.
  • Design and optimize ML systems while balancing production requirements such as latency, cost, reliability, and safety.

Requirements

  • Strong foundation in machine learning, deep learning, and modern neural network architectures.
  • Hands-on experience training, fine-tuning, evaluating, or deploying machine learning models.
  • Ability to write clean, maintainable, production-quality Python code.
  • Eagerness to learn new tools, frameworks, and production ML best practices.
  • Comfortable working through ambiguity while steadily taking on greater ownership.
  • Curious, collaborative, and coachable, with a growth mindset.
  • A bias toward shipping, learning from production, and continuously improving systems.
  • Technologies and tools: Python, PyTorch / JAX, GPU-accelerated production ML systems

Success Looks Like

  • Machine learning models consistently achieve production targets for accuracy, latency, and reliability.
  • Model issues and production incidents are identified quickly, investigated thoroughly, and resolved effectively.
  • Data pipelines, training workflows, and inference systems are scalable, reproducible, and maintainable.
  • Strong collaboration with engineering, product, and research teams results in reliable, production-ready ML features.
  • Model and system improvements are driven by measurable production outcomes, user feedback, and continuous experimentation.