About the Role
We're looking for an MLOps Engineer to build and operate the inference platforms our AI systems run on — the layer between "the model works in a notebook" and "the model serves production traffic reliably."
What You'll Do
- Design and operate GPU-backed inference infrastructure on Kubernetes across cloud providers
- Build CI/CD pipelines for model deployment, versioning, and rollback
- Set up monitoring, alerting, and autoscaling so inference systems stay within latency and cost budgets
- Work with ML engineers to optimize serving frameworks (Triton, vLLM, Ray Serve) for specific workloads
- Own infrastructure-as-code for the platforms you build, using Terraform
What We're Looking For
- 3+ years of experience running production infrastructure, with at least 1 year on ML/AI workloads specifically
- Strong Kubernetes fundamentals — you've debugged a cluster issue at 2am and lived to tell about it
- Experience with infrastructure-as-code (Terraform preferred) and CI/CD pipelines
- Comfortable working across AWS, Azure, or GCP
Nice to Have
- Direct experience with GPU cluster management and scheduling
- Familiarity with LLM inference frameworks (vLLM, Triton Inference Server, TGI)
- FinOps experience — you think about cost per request, not just uptime