Site Reliability Engineer (DataCosmos) Verified today

Open Cosmos · Didcot, United Kingdom · Production · first seen 2026-07-16

About the role

You will own the reliability, performance and scalability of our data platform and processing pipelines as we grow. You will monitor systems end-to-end, ensuring full visibility across infrastructure and data flows. You will respond to incidents, troubleshoot issues and drive long-term fixes.

Key responsibilities

  • Own reliability, performance and scalability of the data platform and processing pipelines
  • Monitor systems end-to-end with full visibility across infrastructure and data flows
  • Respond to incidents, troubleshoot issues and drive long-term fixes
  • Improve deployments and contribute to CI/CD pipelines for safe, repeatable releases
  • Work closely with engineering teams to design resilient, scalable systems
  • Automate processes and reduce operational overhead
  • Support customer-impacting issues alongside Customer Success teams

What you bring

  • Strong demonstrable ability to work with Linux systems and cloud platforms (AWS, GCP or Azure)
  • Solid Kubernetes knowledge and ability to run production systems
  • Clear understanding of observability (monitoring, logging, tracing)
  • Capable of designing or operating high-availability, distributed systems
  • Mindset focused on automation, scalability and continuous improvement
  • Confidence working in fast-moving environments where reliability matters

Location and application

This role can be based in any of our European locations. You must have the legal right to work in your chosen location. Please submit your CV in English.