Notifications

Loading notifications...
Garner Health Cover

Senior Site Reliability Engineer

Garner Health
Worldwide Full Time USD 191,000 - 226,000 7 days ago

About the Job

Garner is on a mission to transform the U.S. healthcare system — and we’re the only proven player doing exactly that. We partner with employers to redesign how healthcare works: applying 550+ proprietary clinical metrics across 80+ specialties to a dataset of 320M+ patients to identify the best-performing doctors, t...

The result is a rare “win win” — better care and lower costs for both members and employers. In just five years, our work has helped over 2.5 million people access higher-quality care and saved $1B in healthcare costs. We recently raised our Series E and have doubled five years running. If you've ever wanted your wo...

We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of the cloud infrastructure powering Garner’s products and AI/ML workloads. This role sits on our Platform Engineering team. You will run the machine: defining and upholding SLOs, leading incident response, and driv...

Key Responsibilities

Garner is headquartered in NYC, but this position is available for individuals who are comfortable with remote work and occasional travel to HQ. 
Run the Machine: Own the end-to-end reliability, performance, and resilience of Garner’s cloud environments (AWS, Kubernetes), including those powering AI/ML workloads; define, measure, and uphold SLOs across our critical services
Lead Incident Response: Serve in the on-call rotation, lead incident response, and drive deep-dive root cause analysis, seeing corrective actions through to resolution and rigorously reviewing infrastructure changes
Own Observability: Build and maintain the monitoring, alerting, and observability systems that let us detect and resolve issues before users feel them
Scale & Optimize: Translate ambiguous, high-performance scaling requirements into well-defined, automated, and composable infrastructure-as-code deliverables (Terraform); proactively identify and implement cost-efficiency and performance gains across the stack to maximize cloud ROI
Automate Away Toil: Pay down impactful tech debt and reduce operational toil, using AI tools and automation to convert repetitive operational work into hands-free, monitored processes, and holding our internal platform to the same rigorous standards as our customer-facing products
Enable Engineering: Build and maintain the deployment and observability standards that empower the broader engineering team to ship AI features faster and more reliably; communicate complex cloud and reliability concepts clearly to technical and non-technical stakeholders
Uphold Security & Compliance: Ensure our infrastructure and operations meet Garner’s security and HIPAA compliance obligations

Required Skills & Abilities

4+ years of hands-on experience operating production cloud infrastructure at scale in an SRE, DevOps, or platform engineering role
Deep expertise with Kubernetes and Terraform in a cloud-first environment (AWS preferred)
A strong track record with production observability: defining SLOs, building monitoring and alerting, and leading incident response and blameless post-incident reviews
Strong software engineering fundamentals in Python or Go, applied to infrastructure automation (experience with Kubernetes APIs a plus)
Experience driving cloud cost-efficiency and performance optimization across compute, storage, and networking
Experience supporting AI/ML or data-intensive workloads in production is a plus
Experience operating in a security-conscious or regulated environment (HIPAA, SOC 2) is a plus
Fluency with AI tools (e.g., Claude) applied to real engineering and operations workflows, or strong motivation to build it fast
A desire to be a part of a high-performing, mission-driven team that operates with intense urgency, a strong sense of individual accountability, and a commitment to authentic feedback
AWS, Kubernetes, Terraform, Istio, Python, Go, TypeScript, Postgres, NATS, Datadog, GitLab

This is a unique opportunity to join a fast-growing company in a transformative role, helping shape the future of healthcare.
We are unable to sponsor or take over sponsorship of an employment visa at this time.

Qualifications

Experience: 4 years experience

Apply now

Please let Garner Health know you found this job on Job Vista. This helps us grow!