Senior DevOps Engineer
Remote (United States)
Job Details
Location: United States
Workplace: Remote
Experience: 4+ years in DevOps, Infrastructure, or Site Reliability Engineering
Core Areas: Kubernetes, Infrastructure as Code, Pulumi, CI/CD, Observability, Distributed Systems, Python, Infrastructure Security
Compensation: $170,000 – $185,000 annual base salary, plus potential bonuses and equity
Travel: Semi-frequent local and national travel, including some overnight travel
About the Role
This opportunity is for a Senior DevOps Engineer to own and evolve the infrastructure supporting an AI-powered cybersecurity platform. The role focuses on production monitoring and observability, containerized environments, Infrastructure as Code, CI/CD automation, internal tooling, system reliability, performance, scalability, and infrastructure security.
The position requires hands-on experience with Kubernetes, Docker, Pulumi, Python, distributed systems, and production infrastructure. Responsibilities include supporting high availability through a 24x7 on-call rotation, preparing infrastructure for continued customer growth, and contributing to future multi-cloud expansion across GCP and Azure.
What You'll Do
- Enhance production monitoring and observability across containerized environments.
- Improve Infrastructure as Code deployments using Pulumi to prepare infrastructure for continued customer growth and scale.
- Develop and refine internal tooling to improve engineering efficiency and infrastructure reliability.
- Strengthen the reliability, performance, and scalability of the core security platform that automates investigation techniques used to analyze security alerts.
- Participate actively in a 24x7 on-call rotation to maintain high availability and provide rapid response to production issues.
- Contribute to future multi-cloud infrastructure expansion efforts involving Google Cloud Platform (GCP) and Microsoft Azure.
- Contribute to new product features when interested in doing so.
Qualifications
Required Experience
- 4+ years of practical, hands-on experience as a DevOps Engineer, Infrastructure Engineer, or Site Reliability Engineer.
- Hands-on experience maintaining Infrastructure as Code and fully automated CI/CD deployment pipelines, with Pulumi preferred.
- Experience administering systems and managing production infrastructure.
Required Skills
- Strong skills in production observability and monitoring, including tools such as Prometheus and Grafana, as well as distributed tracing.
- Strong experience with containerization technologies, particularly Docker and Kubernetes (K8s).
- Production-level Python software development skills.
- Solid understanding of distributed systems, including troubleshooting, bottleneck identification, and performance optimization.
- Ability to maintain and improve fully automated CI/CD deployment pipelines.
- Strong system administration and infrastructure management capabilities.
- Familiarity with modern security best practices, common vulnerabilities, and methods for securing distributed environments.
- Ability to determine when a production-grade service is necessary versus when a one-time or temporary solution is appropriate.
- Ability to operate effectively in an early-stage startup environment with significant ambiguity and rapid execution.
- Ability to use data to guide technical and operational decisions.
- A strong opinion on Emacs versus Vim.
Benefits
- Company-paid health insurance.
- 401(k) plan with employer match.
- Self-managed paid time off.
- Parental leave.
- Company-provided equipment for working from home.
- Additional benefits as provided under the applicable benefits package.
Looking for more opportunities?
View All Jobs