Sr. Site Reliability Engineer, AI Infrastructure (Starshield) Verified today

SpaceX · Redmond, WA · Production · first seen 2026-09-15

Sr. Site Reliability Engineer (Starshield)

SpaceX is deploying Starshield, the world's largest US government satellite constellation tasked with providing immediate access to critical intelligence and national security data globally. We design, build, test, and operate all parts of the system including receivers and software infrastructure.

As a senior engineer focused on Starshield's software and GPU infrastructure, you will design, operate and scale the infrastructure supporting critical national security missions. You will develop automation to deploy and manage on-premise compute resources, create highly scalable software products, and collaborate across engineering teams.

Responsibilities

  • Manage GPU/CPU infrastructure deployments to Top Secret data centers
  • Manage and provide support for GPU as a service for external customers on bare metal hardware and virtualized platforms
  • Design, validate, and productize solutions for AI clusters at 100k+ GPU scale
  • Develop automation to deploy and manage on-premise Kubernetes AI clusters and operating systems
  • Deploy and manage core infrastructure such as databases, monitoring and distributed storage
  • Closely collaborate with AI engineers to create highly scalable, operable, and maintainable products
  • Engage in the complete lifecycle of services from inception and design through deployment, operation and refinement
  • Implement monitoring and alerting to maintain high availability
  • Identify areas for improvement and create innovative solutions enabling high system availability
  • Mentor and train junior engineers
  • Lead the team to technical excellence; your decisions guide the team

Basic Qualifications

  • Bachelor's degree in computer science, information systems/IT, or engineering discipline with 5+ years of professional Linux operating systems experience; OR 7+ years of professional software, DevOps, or site reliability engineering experience without degree
  • 5+ years of experience with Kubernetes
  • 5+ years of experience managing Linux operating systems
  • Experience with Terraform, Ansible, or other infrastructure tools
  • Experience with containerization technologies such as OCI containers or Kubernetes
  • Experience scripting in Bash, Python, or similar languages
  • Development experience in Python, C++, or Go

Preferred Skills and Experience

  • 5+ years of Python and Python-based development frameworks
  • Experience managing Kubernetes clusters, not just using them
  • Knowledge of Linux boot process and systems configuration
  • Deep understanding of testing, continuous integration, build, deployment and continuous monitoring
  • Understanding of build technologies such as Bazel and Makefiles
  • Focus on performance bottlenecks and performance improvement techniques
  • Understanding of distributed databases and data modeling
  • Experience automatically managing thousands of servers using Terraform or Ansible
  • Strong TCP/IP networking knowledge
  • Cloud virtualization experience as the cloud provider
  • Experience with NVIDIA GPU deployment stacks (Blackwell/Rubin)
  • Active Top Secret, Top Secret SCI, or DOE Level Q clearance
  • Excellent communication skills with customers, peers, and management in formal and informal situations

Additional Requirements

  • Willingness to work extended hours and weekends as needed
  • Willingness to travel domestically and globally when needed
  • Must successfully obtain and maintain a Top Secret Security Clearance as a condition of employment. While clearance may not be immediately necessary upon hire, you are encouraged to initiate the application process promptly upon accepting the offer. Your ability to secure the necessary clearance is essential for key responsibilities. SpaceX reserves the right to modify or terminate employment if you are unable to obtain it.

Compensation and Benefits

Pay Range: Level 3: $165,000 - $270,000

Actual level and base salary determined on a case-by-case basis based on job-related knowledge and skills, education, and experience.

Base salary is one part of total rewards. You may be eligible for long-term incentives in the form of company stock or long-term cash awards, discretionary bonuses, and the ability to purchase stock at a discount through an Employee Stock Purchase Plan. Comprehensive medical, vision, and dental coverage; 401(k) retirement plan; short and long-term disability insurance; life insurance; paid parental leave; and various discounts and perks are included. You will accrue 3 weeks of paid vacation and be eligible for 10 or more paid holidays per year. Employees in Washington State accrue paid sick time in compliance with state and federal law. Company shuttles are offered for roundtrip travel from select Seattle locations to the Redmond office Monday to Friday.

Those with an active clearance will receive a 10% differential, up to an additional $20,000 annually, once officially briefed into a classified program.