LTD Global

Site Reliability Engineer

Berkeley, CA - Contracted

📍 Hybrid — Berkeley, CA
📅 1 Year Contract Assignment with possibility of extension based on performance and organizational needs.
💰 $80/hr

Ever wondered what powers breakthrough research in energy, physics, materials science, and chemistry? You're looking at it. This national HPC facility supports 11,000+ scientists pushing the boundaries of what's possible, and we need a sharp, self-motivated SRE to help keep that engine running without interruption.

If you love solving real problems on live infrastructure, thrive on ownership, and want your work to directly enable world-class science, this is your seat.


What You'll Own
  • Monitor and triage alerts across compute, storage, network, and facility systems in real time
  • Build automation that prevents issues before they become outages
  • Develop new tools and integrations across the monitoring pipeline (APIs → alerts → action)
  • Walk the data center floor to keep power, cooling, and environmental systems humming
  • Coordinate maintenance activities across teams and keep incidents accurately tracked
  • Dig into complex, ambiguous problems and drive them to resolution
What You Bring
  • Comfort working Owl shift (12am–8am), 5 days/week, hybrid onsite in Berkeley, CA
  • Solid Linux/command-line (SSH) chops
  • Programming/scripting experience: Python, C, C++, Perl, or Java
  • A self-starter mindset, eager to pick up Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, and building management/cooling systems
  • Network security fundamentals (ACLs, firewalls)
  • Strong cross-team communication and collaboration skills
Nice to Have
  • Experience building or deploying Agentic AI / autonomous automation for technical workflows
  • ServiceNow implementation experience
  • ITSM best-practice know-how
Apply: Site Reliability Engineer
* Required fields
First name*
Last name*
Email address*
Location
Phone number*
Resume*

Attach resume as .pdf, .doc, .docx, .odt, .txt, or .rtf (limit 5MB) or paste resume

Paste your resume here or attach resume file

This role requires working the Owl shift (12am–8am), 5 days/week, hybrid onsite in Berkeley, CA, for. Are you available and willing to commit to this schedule, location?*
Describe your hands-on experience working in a Linux/command-line (SSH) environment, including how long you've worked in this capacity.*
Which languages have you used to build tools or automation (e.g., C, C++, Perl, Java, Python)?*
Describe your experience supporting a 24/7 operations or data center environment, including how you've triaged alerts or incidents.*
What experience do you have with monitoring/observability tools such as Prometheus, VictoriaMetrics, Alertmanager, or Kubernetes?*
Describe your experience briefly with network security — specifically configuring or maintaining ACLs and working with firewalls.*
This position is unable to sponsor a visa of any kind, now or at any point in the future. Are you legally authorized to work in this role without the need for visa sponsorship, both currently and going forward?*
The following questions are entirely optional.
To comply with government Equal Employment Opportunity and/or Affirmative Action reporting regulations, we are requesting (but NOT requiring) that you enter this personal data. This information will not be used in connection with any employment decisions, and will be used solely as permitted by state and federal law. Your voluntary cooperation would be appreciated. Learn more.
Gender
Race/Ethnicity
Human Check*