Sr. Cloud Infrastructure Engineer
About the role
Core Responsibilities
Design, implement, and maintain scalable, reliable infrastructure in AWS
Operate and improve Kubernetes-based environments, including production workloads
Build and maintain infrastructure as code using Terraform
Improve CI/CD pipelines, deployment workflows, and release automation in partnership with engineering teams
Build and maintain the packaging and reference architectures customers use to install our software in their own environments
Strengthen observability across the platform, including monitoring, logging, alerting, and actionable dashboards
Improve developer experience through tooling, environment automation, and self-service infrastructure
Operate within and preserve the established security and compliance posture of our environments
Monitor, troubleshoot, and resolve complex infrastructure issues with clear and timely communication
Participate in incident response and post-incident analysis
Develop and maintain documentation, runbooks, and technical standards
Identify opportunities to improve cost efficiency, performance, and resilience across environments
Required Qualifications
Minimum of 5 years of experience in DevOps, Infrastructure Engineering, Platform Engineering, or Site Reliability Engineering
Strong hands-on experience with AWS in production environments
Proven experience operating Kubernetes in production
Strong experience with Terraform and infrastructure-as-code practices
Proven experience building or improving CI/CD pipelines and deployment automation
Solid understanding of cloud networking, IAM, secrets management, and operational controls
Experience with monitoring, logging, and observability tooling
Scripting proficiency in Python, Go, Bash, or similar
Excellent troubleshooting and problem-solving skills in complex production environments
Strong communication skills with the ability to explain technical concepts to both technical and non-technical stakeholders
Must live/work in the U.S.
Experience operating stateful workloads on Kubernetes, such as databases or message queues, including persistent storage and backup/recovery
Experience with GitOps-based deployment workflows
Experience packaging software for customer-managed or self-hosted deployment (e.g., Helm charts)
Familiarity with compliance or security frameworks such as FedRAMP, NIST, SOC 2, or similar
Experience with PostgreSQL, cloud storage platforms, and production networking patterns
Experience with configuration management tools such as Ansible
Experience with additional cloud platforms such as Azure or GCP
Experience with service mesh or advanced Kubernetes networking
Experience supporting customer-facing or mission-critical production infrastructure
Top Secret Security Clearance