Skip to main content

Site Reliability Engineer | Hybrid

Job DescriptionJob Description

Hybrid — Berkeley, CA

Assignment: 10/26/2026 – 10/27/2027

$80/hr

 

Role Summary

 

As a Site Reliability Engineer on the Operations Technology team, you'll be part of a round-the-clock crew keeping a -scale HPC facility accessible, reliable, and secure. Working from advanced monitoring and data collection systems, you'll proactively catch issues before they escalate, triage and resolve alerts across compute, storage, and network systems, and build the automation that makes the whole environment more resilient over time. You'll also collaborate closely with cross-functional teams to coordinate maintenance, improve tooling, and ensure the infrastructure scales smoothly as demand grows, keeping the computational power behind critical scientific research running without interruption.

 

What You'll Own

  • Monitor and triage alerts across computer, storage, network, and facility systems in real time
  • Build automation that prevents issues before they become outages
  • Develop new tools and integrations across the monitoring pipeline (APIs → alerts → action)
  • Walk the data center floor to keep power, cooling, and environmental systems humming
  • Coordinate maintenance activities across teams and keep incidents accurately tracked
  • Dig into complex, ambiguous problems and drive them to resolution

 

What You Bring

  • Comfort working Owl shift (12am–8am), 5 days/week, hybrid onsite in Berkeley, CA
  • Solid Linux/command-line (SSH) chops
  • Programming/scripting experience — Python, C, C++, Perl, or Java
  • A self-starter mindset — eager to pick up Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, and building management/cooling systems
  • Network security fundamentals (ACLs, firewalls)
  • Strong cross-team communication and collaboration skills

 

Nice to Have

  • Experience building or deploying Agentic AI / autonomous automation for technical workflows
  • ServiceNow implementation experience
  • ITSM best-practice know-how

\nCompany Description

Great Organization!

Company DescriptionGreat Organization!

Site Reliability Engineer | Hybrid

Berkeley, CA
Full time

Published on 09/04/2026

Share this job now