SRE - Platform Engineer
Remote (United States)
Job Details
Location: United States
Workplace: Remote
Experience: Bachelor's degree in Computer Science, Computer Engineering, or a related field, or 8+ years of software engineering experience
Core Areas: Platform Engineering, Site Reliability Engineering (SRE), GCP, Kubernetes/GKE, Terraform, Observability, CI/CD, Infrastructure Automation
Compensation: $125,000–$150,000 per year
About the Role
This opportunity is for an SRE - Platform Engineer focused on the reliability, scalability, and performance of internal and client-facing IT infrastructure and the internal developer platform. The role combines platform engineering with Site Reliability Engineering (SRE), including cloud architecture, GKE cluster operations, infrastructure automation, observability, incident response, SLO/SLI management, capacity planning, and performance optimization.
The position works extensively with GCP, Kubernetes, Terraform, Linux, GitHub Actions, OpenTelemetry, Prometheus, Grafana, and Honeycomb. The engineering approach emphasizes self-service, security by default, least privilege, Test Driven Development (TDD), resilient systems, and scalable software delivery while providing technical direction and platform capabilities to software engineering teams.
What You'll Do
- Serve as a broad-domain architect for the internal developer platform and cloud engineering environment.
- Drive architecture decisions for engineering tooling and internally developed software.
- Mentor platform engineers and promote strong engineering practices across the team.
- Enable platform engineering capabilities for internal software engineering teams.
- Work as a peer with senior architects and engineers across software engineering.
- Design and engineer solutions within the GCP environment.
- Architect and oversee Google Kubernetes Engine (GKE) cluster operations and workload management.
- Provide technical feedback and participate in peer reviews and pair programming.
- Drive adoption of Test Driven Development by designing, developing, and debugging unit and integration tests for new and existing infrastructure and code.
- Continuously evaluate existing implementations and emerging technologies and share relevant knowledge with the engineering team.
- Apply continuous improvement practices across technical responsibilities and professional development.
- Communicate clearly with platform engineering teams and other stakeholders while providing technical direction.
- Stay current with platform changes and third-party libraries and proactively evaluate improved solutions for existing implementations.
- Apply OpenTelemetry and true observability principles, including an understanding of how observability differs from monitoring and logging.
- Help strengthen engineering practices and support the development of a high-performing engineering culture.
- Apply self-service, least privilege, and security-by-default principles across platform solutions.
- Define and maintain Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets.
- Lead incident response activities, including on-call rotations, root cause analysis, and post-mortem reviews.
- Implement and optimize monitoring, alerting, and observability systems to improve system reliability.
- Collaborate on capacity planning and performance optimization to support high availability.
- Perform other duties as assigned.
Qualifications
Required Experience
- Bachelor's degree in Computer Science, Computer Engineering, or a related field, or 8+ years of experience as a software engineer.
- Advanced experience creating and managing public cloud resources with Terraform or other Infrastructure as Code (IaC) tools.
- Experience participating independently in a 24/7 on-call schedule and successfully resolving issues without escalation.
- Proven experience with Site Reliability Engineering practices, including incident management and reliability engineering.
- Experience containerizing applications with Docker and orchestrating workloads with Kubernetes.
- Experience using Git in trunk-based development models.
- Experience using feature flagging for infrastructure and Kubernetes runtime environments.
- Experience using OpenTelemetry for observability and monitoring tools such as Datadog, New Relic, or similar technologies.
- Experience working with security compliance frameworks including FedRAMP, NIST, and SOC 2.
Required Skills
- Strong proficiency with Kubernetes; CKA or CKAD certification is optional.
- Extensive Unix/Linux knowledge and experience working predominantly through command-line interfaces.
- Polyglot development capability with proficiency in multiple languages, ideally including Golang, Node.js, Python, and HCL.
- Knowledge of multi-cloud environments, including familiarity with at least two of GCP, AWS, and Azure.
- Advanced ability to provision and manage public cloud infrastructure through Terraform or comparable Infrastructure as Code tools.
- Strong understanding of networking and routing principles.
- Familiarity with security configuration for web and API services, including SSL and access control.
- Ability to use JIRA or similar work-tracking systems, resolve tickets according to priority, and collaborate with a Technical Product Manager to adjust priorities.
- Strong technical documentation skills using Confluence or similar tools, including support notes, runbooks, and Architecture Decision Records (ADRs).
- Familiarity with creating end-to-end CI/CD pipelines using multiple tools and artifact storage.
- Familiarity with macOS as a desktop environment and predominantly CLI-based workflows.
- Ability to apply a product mindset by understanding stakeholder needs, priorities, and business value.
- Familiarity with Prometheus, Grafana, Honeycomb, or similar monitoring and observability technologies.
- Experience with chaos engineering, load testing, or reliability testing frameworks.
Additional Experience
- Experience with backend database technologies, including database support and performance improvements, is a plus.
Tooling Environment
- GitHub and GitHub Actions.
- Google Cloud Platform (GCP).
- Kubernetes through GKE, Helm, and Docker.
- Google Secret Manager (GSM).
- Terraform.
- Honeycomb.
- Grafana stack.
- Prometheus.
Security Responsibilities
- Maintain a high level of security for personal or private information accessed as part of the role, whether working remotely or at a facility.
- Participate in required security training, respect individual privacy rights, and comply with applicable security policies.
- Comply with additional regulatory or contractual requirements when accessing protected sensitive data, including information governed by HIPAA or contractual requirements related to credit card data.
Looking for more opportunities?
View All Jobs