Site Reliability Engineer (High Performance Computing) Verified today

SpaceX · Hawthorne, CA · Production · first seen 2026-09-24

About the role

SpaceX HPC is a shared compute platform used across the company for vehicle and structures simulation, machine learning, AI inference, and more. This Site Reliability Engineer role will establish a real SRE operating model on these capabilities: toil reduction, automation, observability, and sustainable incident processes.

You will own everything from Linux machines and Infrastructure as Code, storage, and user-facing applications - the whole ecosystem as a product, not as a ticket queue. You do not need prior HPC experience, but you do need production instincts: you have operated real infrastructure, you write code to delete toil, and you care whether users can actually get work done. You'll work alongside HPC systems engineers who design and commission clusters to make them more reliable and provide world class services.

Aerospace experience is not required. We value engineers who treat teammates with fairness and respect, who are self-critical, and who take ownership of hard production problems.

What you will do

  • Participate in on-call rotation; practice sustainable incident response and blameless postmortems
  • Manage node lifecycle with infrastructure as code: OS images, firmware, configuration management, kernel and driver stack
  • Build observability for both HPC administrators and end users - cluster, node, and storage health for operators, and job/workflow-level signal for people running work on the platform
  • Reduce toil with automation; split time between operating production systems and writing software that makes that work smaller
  • Sustainably manage resources, including compute and storage
  • Lead capacity planning with users across the company: understand what they will need next and turn that into a concrete picture of tomorrow's compute and storage
  • Collaborate with HPC systems engineers and engineers across all disciplines on operable, maintainable infrastructure

Basic qualifications

  • Bachelor's degree in computer science, engineering, math, or a scientific discipline; OR 2+ years of professional experience operating production infrastructure in lieu of a degree
  • 2+ years of experience with Linux operating systems in production
  • 2+ years of experience operating production infrastructure (servers, services, or networks), including monitoring, debugging, and repairing what you own

Preferred skills and experience

  • 2+ years of professional experience in SRE, DevOps, or production infrastructure engineering
  • Experience with monitoring and alerting (Prometheus, Grafana, Nagios, or similar)
  • Experience deploying and maintaining configuration management or infrastructure as code (Ansible, Puppet, Terraform, or similar)
  • Experience writing scripts/code (eg. Python or similar languages) to automate common tasks
  • Experience with containers (Docker, Podman, Singularity/Apptainer)
  • Experience with Kubernetes administration for on-premise deployment
  • Experience with distributed or high-performance storage (VAST or similar), including capacity, performance, and lifecycle management
  • Familiarity with HPC clusters, schedulers (Slurm, PBS, LSF), or GPU compute - not required; we will teach this
  • Familiarity with scientific computing (CFD, FEA) and/or ML training workloads (PyTorch, TensorFlow, CUDA) and/or AI inference workloads
  • Good understanding of version control, testing, continuous integration, build, deployment and monitoring
  • Ability to communicate clearly with users, peers, and vendors in both incident and design settings
  • Comfortable working with mission-critical and sensitive systems, with a sense of urgency appropriate to the responsibilities

Additional requirements

  • Position is based in Hawthorne, CA and is primarily on-site
  • Must be able to participate in an on-call rotation
  • Must be willing to work extended hours and weekends as needed for incidents, cluster bring-up, and time-critical failures
  • Eligibility for access to classified material up to TS/SCI with polygraph
  • Must be a U.S. citizen or national, U.S. lawful permanent resident (green card holder), Refugee under 8 U.S.C. 1157, or Asylee under 8 U.S.C. 1158, or be eligible to obtain required authorizations from the U.S. Department of State

Compensation

Level 1: $125,000 to $160,000 per year
Level 2: $145,000 to $195,000 per year

Your actual level and base salary will be determined on a case-by-case basis and may vary based on job-related knowledge and skills, education, and experience.

Base salary is one part of your total rewards package. You may be eligible for long-term incentives in the form of company stock or long-term cash awards, discretionary bonuses, and employee stock purchase plan discounts. Benefits include comprehensive medical, vision, and dental coverage; 401(k) retirement plan; short and long-term disability insurance; life insurance; paid parental leave; 3 weeks of paid vacation; 10 or more paid holidays per year; and paid sick leave in accordance with company policy and applicable law.