About the job AI Inference Engineer
About the Role
Premier Global Links LLC is seeking an experienced Member of Technical Staff, Inference Systems, to build and optimize a high-performance AI inference platform from the ground up.
This role is focused on LLM inference, model serving, distributed systems, and inference runtime performance. The ideal candidate has hands-on experience with production inference systems and strong systems engineering skills, with Rust experience highly valued.
Key Responsibilities
- Build and optimize production LLM inference and model-serving systems.
- Develop inference runtime components using Rust and other systems-level technologies.
- Design and implement batching, scheduling, request routing, and serving infrastructure.
- Build and optimize KV cache and prefix caching systems.
- Scale inference workloads across multi-GPU and multi-node environments.
- Profile, benchmark, and optimize latency, throughput, reliability, and cost.
- Work with inference engines such as vLLM, SGLang, or TensorRT-LLM.
- Investigate performance bottlenecks across the inference stack.
- Contribute to core architecture and technical decisions for the platform.
- Collaborate with a small, hands-on engineering team in a fast-paced environment.
Required Qualifications
- 2–10 years of experience in backend, distributed systems, or systems engineering.
- Hands-on experience building, operating, or optimizing LLM inference or serving systems.
- Deep understanding of transformer inference internals, including attention, KV cache, batching, and scheduling.
- Experience with a production inference engine such as vLLM, SGLang, or TensorRT-LLM.
- Strong programming experience with Rust, C++, Go, or systems-level Python/PyTorch.
- Experience building performance-critical systems where latency, throughput, and cost are important.
- Strong distributed systems and production software engineering fundamentals.
- Ability to work on-site 5 days per week in Palo Alto, CA.
Preferred Qualifications
- Production Rust experience.
- CUDA or Triton kernel development experience.
- Multi-GPU or multi-node serving experience.
- Experience with NCCL, NVLink, or RDMA.
- Experience with prefix caching, speculative decoding, or prefill/decode disaggregation.
- Contributions to open-source inference projects such as vLLM, SGLang, or Dynamo.
- Experience working on inference systems at an AI provider, accelerator company, research lab, or similar organization.
- Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related technical field.
Technology Environment
Rust | C++ | Go | Python | PyTorch | vLLM | SGLang | TensorRT-LLM | CUDA | Triton | NCCL | NVLink | RDMA | LLM Inference | Distributed Systems
Compensation & Benefits
- $230,000–$350,000 annual salary, based on experience and qualifications.
- Equity opportunity starting at approximately 0.5%, with flexibility based on experience.
- Professional growth and development opportunities.
- High-impact work within a fast-paced AI technology environment.
Work Arrangement
On-Site – Palo Alto, CA
Employees are expected to work from the Palo Alto office 5 days per week.
Equal Opportunity Employer
Premier Global Links LLC is an equal opportunity employer. Qualified applicants are considered based on their skills, experience, education, and qualifications.