Site Reliability Engineer, AI Infrastructure (Starshield) Verified today
About the role
SpaceX is deploying Starshield, the world's largest US government satellite constellation, providing immediate access to critical intelligence and national security data for the US government globally. As a Site Reliability Engineer focused on Starshield's software and GPU infrastructure, you will design, operate and scale the infrastructure that supports critical national security missions.
What you will do
- Manage GPU/CPU infrastructure deployments to Top Secret data centers
- Manage and provide support for GPU as a service for external customers on bare metal hardware and virtualized platforms
- Design, validate, and productize solutions for AI clusters (100k+ GPU scale)
- Develop automation to deploy and manage on-premise Kubernetes AI clusters and operating systems
- Deploy and manage core infrastructure such as databases, monitoring and distributed storage
- Closely collaborate with AI engineers to create highly scalable, operable, and maintainable products
- Engage in the whole lifecycle of services from inception and design, through deployment, operation and refinement
- Implement monitoring and alerting to maintain high availability
- Identify areas for improvement and create innovative solutions enabling high system availability
Basic qualifications
- Bachelor's degree in computer science, information systems/IT, or engineering discipline plus 1+ years professional experience in site reliability engineering or DevOps; OR 3+ years professional experience in site reliability engineering or DevOps in lieu of a degree
- 1+ years professional experience with Linux operating systems
- Experience with Terraform, Ansible, or other infrastructure tools
- Experience with containerization technologies (OCI containers, Kubernetes)
- Experience scripting in Bash, Python, or similar languages
- Development experience in Python, C++, or Go
Preferred skills and experience
- 1+ years experience with Python and Python-based development frameworks
- Experience managing Kubernetes clusters, not just using them
- Knowledge of Linux boot process and systems configuration
- Deep understanding of testing, continuous integration, build, deployment and continuous monitoring
- Understanding of relevant build technologies such as Bazel and Makefiles
- Focus on performance bottlenecks and performance improvement techniques
- Understanding of distributed databases and data modeling
- Experience automatically managing thousands of servers (Terraform or Ansible)
- Strong networking knowledge of TCP/IP
- Cloud virtualization experience (as the cloud provider)
- Experience using NVIDIA GPU deployment stacks (Blackwell/Rubin)
- Excellent communications skills across customers, peers, and management
- Active Top Secret, Top Secret SCI, or DOE Level Q clearance
Additional requirements
- Must be willing to work extended hours and weekends as needed
- Must be willing to travel domestically and globally as needed
- This position requires successfully obtaining and maintaining a Top Secret Security Clearance as a condition of employment
Compensation
Level 1: $125,000 to $160,000 annually
Level 2: $145,000 to $195,000 annually
Actual level and base salary determined case-by-case based on job-related knowledge and skills, education, and experience.
Additional benefits include long-term incentives (company stock or cash awards), discretionary bonuses, Employee Stock Purchase Plan, comprehensive medical/vision/dental coverage, 401(k) retirement plan, disability and life insurance, paid parental leave, 3 weeks paid vacation, 10+ paid holidays annually, and paid sick leave. Those with active clearance receive a 10% differential, up to an additional $20,000 annually, once officially briefed into a classified program.
ITAR requirements
Applicant must be a US citizen or national, US lawful permanent resident (green card holder), Refugee under 8 U.S.C. § 1157, or Asylee under 8 U.S.C. § 1158, or be eligible to obtain required authorizations from the US Department of State.
