Sr. Site Reliability Engineer, Platform Infrastructure Verified today
Sr. Site Reliability Engineer, Platform Infrastructure
The application software team is the central nervous system of SpaceX. Manufacturing is how SpaceX turns designs into hardware. The compute, storage, and networking that run our factories must be as reliable as the products we build. This team owns infrastructure supporting Starship, Starlink, Starshield, and Terafab. This position will have a direct impact on factory uptime, throughput, and production scale across programs.
What you will do
- Deploy, upgrade, operate, maintain, and scale compute, storage, and networking for manufacturing systems across Starship, Starlink, Starshield, and Terafab
- Manage infrastructure as code and use observability to provide a complete picture of platform health
- Design for reliability, stability, and scale; find and remove bottlenecks with measurement and engineering
- Practice proactive maintenance: capacity planning, lifecycle management, and reducing toil before it becomes an incident
- Partner with software engineers, manufacturing stakeholders, and site teams to build operable, maintainable systems
- Improve the full lifecycle from design through deployment, operation, and continuous refinement
- Practice sustainable incident response and blameless postmortems
- Provide high-quality support to manufacturing and engineering users
- Communicate clearly with stakeholders and teammates
- Participate in on-call and travel to sites as needed for deployments, incidents, and cross-site reliability
Basic qualifications
- Bachelor's degree in computer science, information systems, or an engineering discipline; OR 7+ years of professional experience in SRE or DevOps in lieu of a degree
- 3+ years of experience with Python and Python-based development frameworks
- Experience with Linux operating systems
Preferred skills and experience
- Experience with compute, storage, and/or networking infrastructure in production
- Infrastructure as code (Terraform, Ansible, Puppet, or similar)
- Containers and virtualization (Docker, Kubernetes, vSphere, QEMU, KVM, etc.)
- Databases and data modeling (Postgres, Clickhouse, etc.)
- Ability to translate high-level requirements into implementations from first principles
- Comfort with mission-critical systems and appropriate urgency and care
- Skillful communication with customers, peers, and management
- Comfort operating across multiple sites and manufacturing programs
Requirements
- Must be able to work extended hours and weekends as needed
- Must be able to travel to different sites (Hawthorne, CA; Redmond, WA; Cape Canaveral, FL; Starbase, TX)
- Ability to pass Air Force background check for Cape Canaveral
- This role requires you to be on-site. Remote and/or hybrid work will not be considered
- Must be a U.S. citizen or national, U.S. lawful permanent resident, Refugee under 8 U.S.C. Section 1157, or Asylee under 8 U.S.C. Section 1158, or eligible to obtain required authorizations from the U.S. Department of State
