Tech Stack
Solid badges = required, outlined = preferred
Responsibilities
- Design and evolve reliability architecture for distributed and cloud-hosted systems.
- Define and implement SRE best practices, including SLIs, SLOs, error budgets, and capacity planning.
- Lead incident response processes including on-call rotations, escalation, and post-incident reviews.
- Design and maintain observability systems for metrics, logging, tracing, and alerting.
- Build automation to improve system reliability, deployment safety, and recovery processes.
Soft Skills
Incident ResponseCross-Functional CollaborationChaos Engineering
Benefits
- 401k
- Dental
- Equity
- Gym/Wellness
- Health Insurance
- Life Insurance
- Parental Leave
- Remote Stipend
- Unlimited PTO
- Vision
Culture
Mission-DrivenInnovationIntegrityTransparent LeadershipWork-Life Balance
Requirements
Regions: Us
Get jobs like this in your inbox
Weekly Distributed Systems, Linux, Networking hiring trends and salary data — free.
Join 8 engineers getting weekly insights
Get market intelligence in your inbox
Free weekly insights on tech hiring trends, salaries, and in-demand stacks.
Already a subscriber? Sign in
About HavocAI
Industry: defense technology
Size: startup
HavocAI is a leader in all-domain collaborative autonomy, building software-defined hardware for military and commercial-grade autonomous systems across sea, air, and land to sense, decide, and act together. The company was founded in 2024 and focuses on optimizing mission performance and minimizing human risk.
View company profile →Compensation
Equity: Equity Package
Similar Jobs
Cloud Platform Tech Lead
HavocAI · Remote
Remote
DevOps Engineer
HavocAI · Remote
Remote
Senior Site Reliability Engineer, Data Infrastructure
CoreWeave · New York, NY / Bellevue, WA
$165k
Staff Site Reliability Engineer
Blink Health · Remote
Remote
Software Engineer, Site Reliability
Hebbia · New York City; San Francisco, CA
$160k