Job Openings
Backend Engineer, AI (Agent Systems)
About the job Backend Engineer, AI (Agent Systems)
Our Client is a AI start-up company. As a Backend Engineer, AI, you'll build and operate the backend infrastructure that powers every AI interaction within the product. You'll design and maintain production-grade systems that transform model capabilities into robust APIs and services consumed by mobile and desktop applications, with a strong focus on performance, observability, and operational excellence.
Responsibilities
- Build, deploy, and maintain backend services that power AI features in production.
- Design inference pipelines, orchesation layers, and service architectures around AI models.
- Develop scalable APIs that enable seamless integration between AI models, frontend applications, and internal services.
- Own production operations, including monitoring, logging, alerting, and incident response.
- Optimize inference performance through caching, batching, streaming, and other latency-reduction techniques.
- Improve system scalability, reliability, and cost efficiency based on production metrics and usage patterns.
- Collaborate closely with machine learning, product, and frontend teams to deliver high-quality AI-powered experiences.
Requirements
- Strong backend engineering experience building and operating production systems.
- Experience developing high-throughput, low-latency distributed services.
- Familiarity with AI inference architectures, including LLMs, embeddings, and multimodal models.
- Strong debugging and troubleshooting skills in distributed systems and production environments.
- Experience designing scalable APIs and backend services.
- A pragmatic, ownership-driven mindset with a focus on shipping reliable systems and continuously improving them through production insights.
- Technologies and tools: Python, Node.js, PyTorch, OpenAI, Anthropic, and open-source LLMs, SQL & NoSQL databases, Kubernetes, Docker
Success Looks Like
- AI backend services consistently deliver low-latency, high-throughput performance at production scale.
- APIs are reliable, well-designed, and enable seamless integration across frontend and machine learning systems.
- Production issues are detected early, resolved quickly, and followed by meaningful reliability improvements.
- Inference infrastructure is scalable, observable, and cost-efficient as usage grows.
- Continuous improvements driven by production data result in measurable gains in system performance, stability, and developer experience.