Notifications

Loading notifications...
Horizon3.ai Cover

Senior Engineering Manager, Site Reliability

Horizon3.ai
Remote Full Time USD 260,000 - 280,000 9 days ago

About the Job

We are a fusion of former U.S. Special Operations cyber operators, startup engineers, and formerly frustrated cybersecurity practitioners. We're committed to helping solve our common security problems: ineffective security tools, false positives resulting in alert fatigue, blind spots, "checkbox” security culture, c...

We are looking for an engineering leader to build and lead the Site Reliability function at Horizon3. This leader is expected to hire a team and establish a Site Reliability function to define and drive investments in operational excellence to ensure Horizon3’s product and service offerings meet customer and busines...

Build a SRE team, from scratch. Hire experienced site reliability staff and build a team of 4-6 in year one.

Key Responsibilities

Professionalize incident management. Define and document incident processes and practices for your SRE team and for the application feature teams. Make tool and vendor decisions to support processes.
Drive incident professionalism and reliability culture across the engineering organization through training and process adoption.
Use design reviews, code reviews, and blameless retrospectives to drive a culture of quality and excellence in engineering.
Balance incident response while also executing on a roadmap of observability and reliability engineering initiatives.
Hire and directly manage site reliability engineers.
As a Manager, you will be responsible for:
Recruiting and onboarding talented individuals to support our organizational goals
Mentoring, coaching, equipping, and developing your team
Recognizing and retaining high performers
Leading horizontally with peer Management & Senior Leaders
Please note this job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee. Duties, responsibilities, and activities may change at any time with or without notice. 
In any materials you submit, you may redact or remove age-identifying information such as age, date of birth, or dates of school attendance or graduation. You will not be penalized for redacting or removing this information.

Required Skills & Abilities

Demonstrated experience leading hiring and growing SRE or Infrastructure teams. Experience leading or building SRE functions, including incident management processes, on-call programs, SLO/SLA definition, and operational runbooks.
Previous career experience as a Site Reliability Engineer. Comfortable in being hands-on while you grow and hire your team.
Deep hands on experience with observability: application performance management, logs and traces, and golden signals and service-specific metrics.
Experience in selecting and deploying incident management tooling (e.g., PagerDuty, FireHydrant,etc.) and creating decision making frameworks for vendor selection.
Strong working knowledge of at least one major cloud provider (AWS, GCP, or Azure) — infrastructure, networking, managed services, IAM, cost management. AWS preferred.

Apply now

Please let Horizon3.ai know you found this job on Job Vista. This helps us grow!