Senior Engineer – AI Platform
Job Description:
Role Overview We are looking for a hands-on Senior Engineer to design, build, and operate an enterprise AI platform for a digital bank. Our goal is to provide a fully governed, safe, and efficient AI environment that staff actually want to use. You will take ownership across three core areas: an enterprise-grade inference gateway, identity and distribution systems, and self-hosted GPU inference clusters.
Core Responsibilities
Inference Serving & On-Premise GPU Operations
- Size hardware capacity and collaborate with infrastructure teams to handle hypervisors, rack/power setups, and GPU passthrough across bare-metal or data center environments.
- Deploy, operate, and fine-tune model serving frameworks using techniques like continuous batching, quantization, tensor parallelism, and KV cache optimization.
- Implement tenancy admission controls, quota management, queuing, and failover mechanisms to cloud inference when local capacity is saturated.
- Track and publish key performance indicators—such as p95 time-to-first-token, tokens per second, and cost per million tokens—to evaluate fixed OpEx break-even points against cloud alternatives.
Governed Gateway Architecture & Identity Integration
- Construct and maintain an inference gateway that handles rate limiting, per-user key management, model allowlists, and strict zero data retention policies.
- Connect governance systems directly to corporate single sign-on (OIDC / OAuth2) and automate onboarding, offboarding, and role claims via SCIM.
- Ensure total financial transparency by attributing real-time API spend, token usage, and operational costs down to individual users and departments.
Tooling Distribution & User Enablement
- Package, distribute, and manage agent configurations and execution harnesses across a diverse enterprise machine fleet.
- Maintain a secure, CI-driven registry for versioned skills, knowledge-base bundles, and agent modules featuring seamless rollback capabilities.
- Onboard, train, and support non-technical teams (such as risk, finance, and operations) to maximize productivity using approved AI tools.
What We Are Looking For
Essential Qualifications & Skills
- Production Platform Engineering: 5+ years of experience managing critical production infrastructure, with direct involvement in incident response and on-call rotations.
- Core Technical Stack: Proficiency in Python (primary) and TypeScript, alongside hands-on experience with Kubernetes, Linux, containers, Infrastructure as Code, and CI/CD pipelines.
- Enterprise Identity Systems: Strong grasp of JWT claims, OIDC, OAuth2, SCIM protocols, and enterprise directory integrations.
- Gateway & Multi-Tenancy Architecture: Background in building or maintaining API/LLM gateways, rate limits, quota systems, usage metering, and multi-tenant isolation.
- On-Premise Infrastructure Comfort: Experience navigating capacity planning, hypervisors, hardware passthrough, and enterprise change management processes.
- FinOps & Attribution Mindset: Ability to evaluate platforms through unit economics, capacity utilization, and cost attribution.
- Technical Enablement: Skill in simplifying complex AI tools and training non-technical colleagues to use them effectively.
Preferred / Bonus Qualifications
- Experience with Model Context Protocol (MCP), sovereign hosting constraints, model evaluation, or highly regulated financial services environments.
- Hands-on expertise tuning high-throughput GPU model serving stacks such as vLLM, SGLang, Triton, or TGI.
Platform Guiding Principles
- Usability First: Approved pathways must offer a better user experience than unapproved workarounds to prevent control bypasses.
- Attribution Before Scale: Total visibility into spend and usage per user is required before expanding system scale.
- Allocated Capacity: GPU resources are scarce and access is governed through explicit capacity allocation policies.
- Pragmatic Ownership: Open-source or licensed solutions are used for standard components, while internal focus is dedicated to security and governance boundaries.
Required Skills:
JWT Environment Hardware Data Center Incident Response Performance CI/CD pipelines Data Cloud API Support Access AI Usability Transparency Financial Services Pipelines Operations Ownership User Experience Onboarding CI/CD Components Architecture Change Management Optimization Infrastructure Economics Integration Kubernetes TypeScript Security Linux Finance Planning Design Engineering Training Python Management