About the Role
We're looking for a Senior Cloud Ops Engineer to own the reliability, performance, and cost-efficiency of production cloud environments for our enterprise clients. You'll be the person who keeps multi-cloud infrastructure running smoothly and gets called in when something needs to scale, migrate, or get fixed fast.
What You'll Do
- Own day-to-day operations of production cloud environments across AWS, Azure, and/or GCP — uptime, performance, and cost
- Design and implement infrastructure-as-code (Terraform, CloudFormation, or similar) for repeatable, auditable deployments
- Build and maintain CI/CD pipelines, container orchestration (Kubernetes/EKS/AKS/GKE), and automated deployment workflows
- Lead incident response for production issues — triage, root cause, and drive the fix, not just the workaround
- Implement monitoring, alerting, and observability so problems are caught before they become outages
- Drive FinOps practices to keep cloud spend aligned with actual usage and business value
- Mentor junior cloud engineers and set operational standards for the team
What We're Looking For
- 5–12+ years of experience in cloud infrastructure, DevOps, or site reliability engineering roles
- Deep hands-on experience with at least one major cloud provider (AWS, Azure, or GCP); multi-cloud experience is a strong plus
- Strong Kubernetes and containerization experience in production environments
- Proven track record owning infrastructure-as-code at scale
- Experience with incident management and on-call rotations for production systems
- Comfortable being the escalation point when things break at 2am — and comfortable making sure that happens less often over time
Nice to Have
- Cloud certifications (AWS Solutions Architect, Azure Administrator, GCP Professional Cloud Architect)
- Experience supporting GPU/AI workloads in the cloud
- FinOps certification or hands-on cost-optimization experience at scale
- Prior experience in a managed services or MSP environment