Senior Site Reliability Engineer

Posted

HavocAIRemote full_timesenior

Tech Stack

Solid badges = required, outlined = preferred

Responsibilities

  • Design and evolve reliability architecture for distributed and cloud-hosted systems.
  • Define and implement SRE best practices, including SLIs, SLOs, error budgets, and capacity planning.
  • Lead incident response processes including on-call rotations, escalation, and post-incident reviews.
  • Design and maintain observability systems for metrics, logging, tracing, and alerting.
  • Build automation to improve system reliability, deployment safety, and recovery processes.

Soft Skills

Incident ResponseCross-Functional CollaborationChaos Engineering

Benefits

  • 401k
  • Dental
  • Equity
  • Gym/Wellness
  • Health Insurance
  • Life Insurance
  • Parental Leave
  • Remote Stipend
  • Unlimited PTO
  • Vision

Culture

Mission-DrivenInnovationIntegrityTransparent LeadershipWork-Life Balance

Requirements

Regions: Us

Get jobs like this in your inbox

Weekly Distributed Systems, Linux, Networking hiring trends and salary data — free.

Join 8 engineers getting weekly insights

Get market intelligence in your inbox

Free weekly insights on tech hiring trends, salaries, and in-demand stacks.

Already a subscriber? Sign in

About HavocAI
Industry: defense technology
Size: startup

HavocAI is a leader in all-domain collaborative autonomy, building software-defined hardware for military and commercial-grade autonomous systems across sea, air, and land to sense, decide, and act together. The company was founded in 2024 and focuses on optimizing mission performance and minimizing human risk.

View company profile →
Compensation
Equity: Equity Package