Sr. Site Reliability Engineer, AI Infrastructure (Starshield) Verified today
About the role
SpaceX is deploying Starshield, the world's largest US government satellite constellation providing critical intelligence and national security data access globally. As a Senior Site Reliability Engineer focused on Starshield's software and GPU infrastructure, you will design, operate and scale the infrastructure supporting these missions.
What you will do
- Manage GPU/CPU infrastructure deployments to Top Secret data centers
- Manage and provide support for GPU as a service for external customers on bare metal hardware and virtualized platforms
- Design, validate, and productize solutions for AI clusters (100k+ GPU scale)
- Develop automation to deploy and manage on-premise Kubernetes and AI clusters, and operating systems
- Deploy and manage core infrastructure such as databases, monitoring and distributed storage
- Closely collaborate with AI engineers to create highly scalable, operable, and maintainable products
- Engage in the whole lifecycle of services from inception through design, deployment, operation and refinement
- Implement monitoring and alerting to maintain high availability
- Identify areas for improvement and create innovative solutions for system availability
- Mentor and train junior engineers
- Lead the team to technical excellence through your technical decisions and guidance
Basic qualifications
- Bachelor's degree in computer science, information systems/IT, or engineering discipline plus 5+ years professional experience with Linux operating systems; OR 7+ years professional experience in software, DevOps, or site reliability engineering
- 5+ years experience with Kubernetes
- 5+ years managing Linux operating systems
- Experience with Terraform, Ansible, or other infrastructure tools
- Experience with containerization technologies (OCI containers, Kubernetes)
- Scripting experience in Bash, Python, or similar languages
- Development experience in Python, C++, or Go
Preferred qualifications
- 5+ years experience with Python and Python-based development frameworks
- Experience managing Kubernetes clusters, not just using them
- Knowledge of Linux boot process and systems configuration
- Deep understanding of testing, continuous integration, build, deployment and continuous monitoring
- Understanding of relevant build technologies such as Bazel and Makefiles
- Focus on performance bottlenecks and optimization techniques
- Understanding of distributed databases and data modeling
- Experience automatically managing thousands of servers with Terraform or Ansible
- Strong TCP/IP networking knowledge
- Cloud virtualization experience as the cloud provider
- Experience with NVIDIA GPU deployment stacks (Blackwell/Rubin)
- Active Top Secret, Top Secret SCI, or DOE Level Q clearance
- Excellent communication skills across customers, peers, and management
Additional requirements
- Willing to work extended hours and weekends as needed
- Willing to travel domestically and globally when needed
- This position requires successfully obtaining and maintaining a Top Secret Security Clearance as a condition of employment. While clearance may not be immediately necessary upon hire, you are encouraged to initiate the application process upon accepting the offer. Your ability to secure the necessary clearance is essential for fulfilling key responsibilities. Should you be unable to obtain it, SpaceX reserves the right to modify or terminate employment to align with operational needs.
Compensation
Level 3: $165,000 to $265,000 base salary. Actual level and base salary determined case-by-case based on job-related knowledge and skills, education, and experience.
Additional compensation may include long-term incentives in company stock or long-term cash awards, discretionary bonuses, and Employee Stock Purchase Plan discounts. Benefits include comprehensive medical, vision, and dental coverage, 401(k) retirement plan, short and long-term disability insurance, life insurance, paid parental leave, 3 weeks paid vacation, and 10+ paid holidays annually. Employees with an active clearance receive a 10% differential, up to an additional $20,000 annually, once briefed into a classified program.
US person requirement
To conform to U.S. Government export regulations, applicant must be a U.S. citizen or national, U.S. lawful permanent resident (green card holder), Refugee under 8 U.S.C. section 1157, or Asylee under 8 U.S.C. section 1158, or be eligible to obtain required authorizations from the U.S. Department of State.
