Senior DevOps Engineer
About us
At Xelix, we work with some of the world’s largest companies to automate and strengthen their financial controls. Our AI solutions redefine how Accounts Payable teams operate – moving from manual processes to automated, intelligent workflows.
Xelix is a fast-paced scale-up – things move fast and expectations are high. We raised our Series B with Insight Partners in June 2025 and are expanding aggressively. We have a team of over 150 talented people pulling together to achieve our goals. Everyone is trusted to take ownership, move fast and have a meaningful impact. We prioritise personal and professional growth, keep things fun, and love to celebrate a milestone together.
In this role you’ll grow, be challenged and help shape the future of Xelix. If you’re excited about building something special with us, we’d love to hear from you.
About the role
We are looking for a Senior DevOps Engineer to take ownership of the reliability, scalability, and security of our infrastructure. This role is for someone who has spent a decade or more building and operating production systems end-to-end — from cloud infrastructure and CI/CD pipelines through to databases, messaging, and the operational needs of machine learning workloads. You will be a technical anchor for the team, setting standards for infrastructure-as-code, security, and operational excellence, while remaining deeply hands-on with the platform day to day.
You’ll need to be based in or around London and able to work from our office several days a week, working closely with engineering, data, and security teams.
What you'll be doing
Design, build, and operate our cloud infrastructure on AWS, with a strong focus on container orchestration (ECS and EKS) and reliable, scalable production services.
Own our Infrastructure-as-Code estate in Terraform, driving consistent, version-controlled, repeatable provisioning across environments.
Build and maintain CI/CD pipelines using GitHub Actions, enabling teams to ship safely and frequently.
Operate and tune production RDS/Aurora databases, including performance, backups, migrations, partitioning, and scaling strategies.
Support ML Ops needs: help productionise machine learning workloads and ensure the ML infrastructure is reliable, observable, and cost-effective.
Write Python for automation, tooling, and infrastructure glue code across the platform.
Champion DevSecOps practices — embedding security into the pipeline, managing secrets, enforcing least-privilege access, and working with engineering teams to remediate vulnerabilities early.
Build and improve monitoring, alerting, and incident response so issues are caught and resolved before they become outages.
Mentor other engineers, contribute to architecture decisions, and help shape the team’s technical standards and roadmap.
Participate in on-call rotation and take ownership of incidents affecting production systems.
What you’ll bring
10+ years of experience in a DevOps, SRE, platform, or infrastructure engineering role, ideally including production ownership of business-critical systems.
Strong proficiency in Python, used for automation, tooling, and operational scripting.
Deep, hands-on experience with AWS, including container platforms such as ECS and EKS.
Proven expert-level skill with Terraform — this is a must-have, not a nice-to-have.
Solid experience with database operations, ideally PostgreSQL in AWS RDS/Aurora, including performance tuning, migrations, backups, and scaling large or high-throughput databases.
Practical experience with CI/CD pipelines, specifically GitHub Actions.
A strong DevSecOps mindset, with experience embedding security practices, secrets management, and compliance controls into infrastructure and pipelines.
Strong troubleshooting skills across Linux systems, networking, and distributed infrastructure.
Excellent communication skills and the ability to work cross-functionally with engineering, data, and security stakeholders.
Must be based in or commutable to London and available to work from the office several days per week.
Preferred / Nice to Have
Background in high-availability, low-latency, or highly regulated environments (e.g. fintech, trading, healthcare).
MLOps experience — deploying, monitoring, and scaling machine learning workflows in production.
Experience with ClickHouse, Amazon Redshift, or other data warehousing/analytics platforms.
Strong identity and access management platform experience and knowledge.
Experience working in environments with security compliance requirements and building tooling to meet these needs.
What we offer in return
💰 Competitive salary of £65,000 - £90,000 depending on experience
🏝️ 27 days of annual leave (including 3 days Christmas closing), with the option to roll over 3 days
🏡 Hybrid working from our dog-friendly Hoxton office and on-site gym
🏥 Comprehensive private medical & dental cover with Vitality
🍼 Enhanced parental leave pay
📚 Learning & development culture – £1,000 personal annual budget
🌍 We’re carbon-neutral and are working towards ambitious carbon reduction goals
🎯 Lots of team socials & activities
☀️ Annual team retreat
Want to learn more?
We believe that people from diverse backgrounds, with different identities and experiences make our company and product better. No matter your background, we'd love to hear from you! And if you have a disability, please let us know if there's any way we can make the interview process better for you - we're happy to accommodate!
If you're a recruiting agency - we have an existing list of agencies we work with and we are not currently planning on expanding the list. Neither the Talent team nor hiring managers or the Support team will respond to cold outreach.
- Department
- Engineering
- Location
- London
- Remote status
- Hybrid
- Salary
- £65,000 - £90,000/year