Site Reliability Engineer Verified today
About Aerospacelab
Aerospacelab designs and manufactures small satellites combined with earth observation services. Founded in 2018, the company has grown to over 350 full-time employees across offices in Louvain-la-Neuve (Belgium), Toulouse (France), Lausanne (Switzerland), and US West Coast. The team spans 25+ engineering disciplines from hardware design to software development and data science.
Responsibilities
- Provide SRE expertise across relevant projects, contributing to reliability, performance, and operational excellence.
- Ensure clear and efficient communication between all relevant stakeholders including engineering, security, and business teams.
- Offer SRE guidance and best practices within internal R&D processes when required.
- Collaborate closely with Security, Development, and Operations teams to integrate reliability and observability best practices into the entire service lifecycle.
- Implement SRE principles such as SLIs/SLOs, error budgets, incident management, tooling automation, and proactive reliability improvements.
- Contribute to incident response, on-call rotation, root-cause analysis, and long-term remediation to strengthen production environments.
- Continuously expand knowledge through internal resources and learning opportunities.
Qualifications
- Master's, PhD, or Bachelor's degree in Computer Science, Engineering, or a related field.
- Strong knowledge of modern reliability, DevOps, and infrastructure technologies.
- Ability to drive improvements, implement process changes, and introduce new tools and automation.
- Strong organization, communication, and documentation skills with ability to convey complex concepts to non-technical audiences.
- Experience collaborating with a variety of stakeholders across organizations.
- Understanding of secure coding practices and modern security frameworks such as NIST and OWASP.
- Familiarity with SRE methodologies including observability, automation, performance engineering, system design, and reliability metrics.
Technical Skills
- Experience with cloud and on-premise infrastructure environments.
- Experience with container platforms such as Kubernetes; Tanzu experience is a plus.
- Experience with Docker and containerization.
- Experience with packaging and config tools such as Helm and Kustomize.
- Experience with Backup/DR for clusters and workloads such as Velero and Kasten.
- Experience with policy as code such as OPA/GateKeeper and Kyverno.
- Understanding and application of GitOps principles and tooling such as Flux and ArgoCD.
- Experience with CI/CD tools such as Jenkins, GitLab, and GitHub Actions.
- Experience with Git, branching strategies, and version control best practices.
- Experience with Infrastructure as Code tooling such as Terraform and Ansible.
- Experience with Linux system administration, troubleshooting, and performance tuning.
- Knowledge of distributed systems, load balancing, networking, and storage.
- Experience with monitoring and observability tools such as Prometheus, Grafana, Alloy, Mimir, and Loki.
- Understanding of container security, IAM, and secrets management such as Vault and AWS Secrets Manager.
- Performance and reliability testing experience such as k6 is a plus.
- Experience with cloud providers; certifications are an asset.
- Knowledge of databases including NoSQL, SQL, Postgres, and MongoDB is a plus.
- Experience with programming languages such as Python, Java, or C/C++/C# is a plus.
- Experience with big data technologies such as Azure Data Factory, AWS Data Pipeline, and GCP Dataflow is a plus.
- Knowledge of French and English.
