Senior Site Reliability Engineer
Our client is an exceptional technology-driven digital assets business that has established itself as a respected and growing player within the market. Backed by a highly successful investment platform, the business combines the agility of a lean engineering environment with the stability and resources of a larger organisation. As they continue to invest in their technology function, they are looking to hire a Senior Site Reliability Engineer to take ownership of the reliability, scalability, and performance of a critical front-office platform. This is a high-impact SRE position within a lean team, offering significant autonomy, visibility, and influence over the platform's future direction.
They are looking for a proactive, highly accountable engineer who thrives in fast-paced environments and enjoys solving complex problems before they become issues. The ideal candidate is a hands-on SRE with a strong engineering mindset, excellent stakeholder management skills, and a genuine passion for building resilient systems. This role would suit someone who enjoys balancing traditional SRE responsibilities such as platform reliability, incident management, and observability with infrastructure engineering and automation projects. You'll be comfortable working directly with business users, taking ownership of critical systems, and driving SRE best practices across the environment while seeing the direct impact of your work.
Key Responsibilities
- Design, build, and maintain highly scalable cloud infrastructure and Kubernetes environments.
- Drive automation across infrastructure, deployment, and operational workflows to improve efficiency and reduce manual effort.
- Own and continuously improve Infrastructure-as-Code practices using Terraform and modern cloud-native tooling.
- Develop and maintain monitoring, logging, and alerting capabilities to improve platform visibility and resilience.
- Support critical production systems, ensuring high availability and strong performance for business users.
- Define and track reliability metrics, including SLOs and SLIs, to proactively manage platform health.
- Lead incident investigation, root cause analysis, and post-incident improvements.
- Partner closely with developers and business stakeholders to support platform growth and ensure operational excellence.
- Take ownership of infrastructure and platform projects from design through to production delivery.
- Contribute to the long-term technical direction of a rapidly evolving platform.
Key Skills & Experience
- 5+ years' experience in Site Reliability Engineering, Platform Engineering, DevOps, or a closely related discipline.
- Strong hands-on experience managing production Kubernetes environments.
- Proven expertise with Infrastructure-as-Code, particularly Terraform.
- Solid Python development or scripting experience for automation and tooling.
- Experience working within cloud environments such as GCP, AWS, or Azure.
- Demonstrated ability to take ownership of complex technical challenges from investigation through to resolution.
- Excellent communication skills with the ability to engage confidently with both technical and non-technical stakeholders.
This is an outstanding opportunity for an engineer looking to move beyond a traditional infrastructure role and join a highly visible, business-facing environment where they will have genuine ownership, direct impact, and the chance to work on a modern cloud-native platform at the heart of a growing digital assets business.
