Description
We fuse together exceptional talent who deliver outstanding software solutions. Our approach has helped us grow 60% in 2021, 94% in 2022, while in 2023 we joined forces with Insight, a Fortune 500 company and a leading solutions and systems integrator. With exciting growth plans and cutting-edge projects, there has never been a better time to join our incredible team.
About the role
We are looking for a highly experienced Principal Devops Engineer to help lead the infrastructure and reliability architecture for the technology landscape as we accelerate our modernisation journey.
Our broader architectural direction is already taking shape, including DDD, microfrontends and Backend-for-Frontend (BFF), while our infrastructure direction is standardising around AWS, Terraform, GitHub Actions and ECS/Fargate. However, important decisions remain around areas such as gateways and routing, scalability, observability, resilience, deployment architecture, and how existing systems should progressively move toward the target state.
Your primary responsibility will be to consolidate this direction into a coherent infrastructure and reliability architecture and define a pragmatic path for its adoption.
You will operate between Enterprise Architecture and SRE/engineering teams, translating broader architectural direction into practical patterns, standards, reference implementations, and modernisation strategies.
This is a hands on Principal level individual contributor role with significant technical influence across infrastructure. You will be expected to challenge existing decisions where appropriate, validate important architectural choices through proofs of concept and reference implementations, and provide the technical direction that enables engineering teams to implement and adopt the target architecture successfully.
What you'll work on
Infrastructure modernisation & target architecture
- Define and evolve the target infrastructure and reliability architecture.
- Consolidate architectural decisions already underway into a coherent, scalable, and maintainable target state.
- Define pragmatic modernisation strategies for existing systems, balancing business value, technical risk, cost, and migration effort.
- Assess systems and recommend whether they should be incrementally modernised, aligned with the target architecture, temporarily retained, or eventually replaced.
- Define transition patterns that allow teams to modernise without unnecessary large scale rewrites.
- Establish the target architecture and infrastructure patterns as the default for new modules.
- Identify architectural gaps, risks, and cross-system dependencies across landscape.
AWS cloud & platform architecture
- Define scalable, resilient, secure, and cost-conscious architectures using AWS.
- Define approaches to service-to-service communication, external API exposure, routing, and gateway strategy.
- Guide architectural decisions around scalability, workload and capacity management, fault isolation, availability, and graceful degradation.
- Define appropriate environment, networking, and deployment strategies for services.
- Work closely with security and enterprise platform teams to ensure alignment with Pearson-wide standards.
Infrastructure as Code & CI/CD
- Establish Infrastructure as Code standards using Terraform, including reusable patterns, modules, and conventions that engineering teams can adopt consistently.
- Shape CI/CD architecture using GitHub and GitHub Actions as the standard delivery platform.
- Define reusable deployment patterns and approaches to environment promotion, rollback, and safe releases.
- Reduce infrastructure and CI/CD divergence between engineering teams through reusable standards and automation.
Reliability, observability & production readiness
- Shape and evolve wide approaches to reliability, resilience, observability, and production readiness.
- Shape standards for metrics, logs, traces, dashboards, and alerting across distributed systems.
- Help establish meaningful SLIs, SLOs, and reliability targets where appropriate.
- Guide architectural approaches to disaster recovery, failure handling, backups, and recovery strategies.
- Help teams design systems that remain operable and cost-effective as usage and complexity grow.
Architecture standards & AI-enabled engineering
- Translate infrastructure and SRE architecture decisions into reusable standards, reference implementations, Terraform patterns, and CI/CD practices.
- Collaborate with teams evolving AI enabled engineering framework so agreed infrastructure patterns and guardrails can be incorporated into engineering workflows.
- Create clear architecture decision records, reference architectures, and implementation guidance that teams can apply consistently.
Technical leadership
- Act as a senior technical authority for SRE and infrastructure architecture.
- Operate as a bridge between Enterprise Architects and the SRE and engineering teams responsible for implementation.
- Validate important architectural decisions through proofs of concept and reference implementations where appropriate.
- Review major infrastructure designs and provide technical direction across teams.
- Mentor senior engineers, Tech Leads, and SREs and help drive alignment on cross-team technical decisions.
Required skills & experience
Architecture & technical leadership
- Extensive professional experience designing and operating large-scale distributed systems in production.
- Strong track record defining cloud and infrastructure architecture across multiple teams or services.
- Experience leading or shaping modernisation across technology estates containing both legacy and modern systems.
- Ability to define a target architecture while creating realistic incremental migration paths toward it.
- Strong architectural judgment and the ability to balance technical quality, delivery speed, risk, cost, and organisational constraints.
- Strong technical depth and willingness to create prototypes or reference implementations when needed to validate architectural decisions.
AWS, infrastructure & delivery
- Deep expertise in AWS and cloud-native architecture.
- Strong experience designing and operating containerised workloads on AWS, particularly ECS/Fargate.
- Strong understanding of AWS networking, load balancing, routing, and secure connectivity.
- Experience designing scalable API, gateway, and ingress architectures.
- Strong experience with relational data platforms such as Amazon RDS/Aurora.
- Good understanding of event-driven and asynchronous architecture patterns.
- Strong experience with Terraform in production environments.
- Strong experience with GitHub and GitHub Actions, including reusable CI/CD workflows and deployment automation.
- Strong understanding of cloud native autoscaling, workload capacity management, and scaling strategies for containerised and event-driven systems.
SRE & reliability engineering
- Strong understanding of Site Reliability Engineering principles.
- Experience designing observability approaches for distributed production systems.
- Strong understanding of metrics, logging, tracing, alerting, and production monitoring.
- Experience working with SLIs, SLOs, reliability targets, and production-readiness practices.
- Strong knowledge of resilience patterns, failure modes, scalability, and disaster recovery.
- Ability to reason about trade-offs between reliability, performance, complexity, and cost.
Security
- Strong understanding of AWS security fundamentals, including IAM, secrets management, encryption, and secure infrastructure patterns.
- Ability to work effectively with specialist security teams and translate organisational requirements into practical engineering approaches.
General & soft skills
- Strong technical leadership skills without relying on formal line-management authority.
- Ability to influence engineering direction across multiple teams and organisational boundaries.
- Comfortable challenging existing approaches constructively and explaining the trade-offs behind alternative solutions.
- Pragmatic approach to modernisation, with the judgment to distinguish valuable change from unnecessary technical churn.
- Excellent communication skills with engineers, Tech Leads, Engineering Managers, Enterprise Architects, and other stakeholders.
- Ability to turn ambiguous technical problems into clear decisions and actionable next steps.
Nice to have
- Experience with Backend-for-Frontend and microfrontend architectures.
- Experience with event-driven architectures and Kafka/MSK.
- Experience with OpenTelemetry and modern observability platforms.
- Experience with progressive delivery approaches such as canary or blue/green deployments.
- Experience with cloud cost optimisation or FinOps practices.
- Experience defining architecture standards, paved roads, or engineering golden paths across multiple autonomous teams.
- Experience working in EdTech, digital learning, or another large global technology organisation.
BENEFITS:
- Opportunities for certification and training
- Dual monitor setup and high-spec workstations
- English courses
- Gym allowance
- Medical Reimbursement
- Full salary covered up to 20 days of sickness
- Flexible working hours
- Loyalty scheme
- Team-building activities, special events and conferences
- UK/EU Travel opportunities
- Snacks and drinks in the office
To see more roles, click here.
Similar jobs
We fuse together exceptional talent who deliver outstanding software solutions. Our approach has helped us grow 60% in 2021, 94% in 2022, while in 2023 we joined forces with Insight, a Fortune 500 company and a leading s…
We are looking for a Principal Engineer to help scale our engineering organization, reduce system complexity, and define a sustainable technical direction, building on top of a mature and evolving platform. You will oper…
Company Background Our client is a global technology company transforming the way brands and agencies make marketing decisions. Its platform enables marketers to plan and activate cross-channel, programmatic media campai…
Xebia is a global AI-first, digital transformation, and engineering partner. With over 25 years of experience and a team of 5,000 professionals across 16 countries, we help organizations design and build scalable product…
Est. 80,000 EUR
About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…
Est. 120,000 EUR
About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…
About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…
About impact.com impact.com is the world’s leading commerce partnership marketing platform, transforming the way businesses grow by enabling them to discover, manage, and scale partnerships across the entire customer jou…
Est. 90,000 EUR
About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…
Est. 80,000 EUR
About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…
Est. 120,000 EUR
About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…
About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…
Est. 120,000 EUR
Responsibilities Manage team capacity, resource allocation, and establish effective 24/7 on-call rotation processes for incident response. Take ownership of the infrastructure budget (FinOps), optimizing cloud costs acro…
Company Background Coherent Solutions is a global digital product engineering company founded in 1995 and headquartered in Minneapolis, Minnesota. With a presence in nine countries and a team of more than 1,700 engineers…
Est. 124,000 USD
Our Systems Engineering team owns the infrastructure that keeps a ~2,500-person global firm running around the clock: cloud (primarily AWS), Microsoft 365, Windows Server, identity platforms (Entra ID, Active Directory,…
Est. 120,000 EUR
This is a Staff Engineer role for a highly experienced software engineer who combines deep hands-on engineering expertise with strong system design skills. The ideal candidate contributes to architecture reviews and tech…
This is a Staff Engineer role for a highly experienced software engineer who combines deep hands-on engineering expertise with strong system design skills. The ideal candidate contributes to architecture reviews and tech…
Est. 55,000 EUR
About the Team We build full-stack products that combine highly interactive React interfaces with scalable Java backend systems. Because our software runs in demanding operational environments where connectivity can be h…
Est. 85,000 USD
We are looking for an experienced and proactive Infrastructure Support Engineer to provide technical support and administration across our cloud and on-prem IT environments. This role combines infrastructure operations,…
Company Background Our client is a global technology company transforming the way brands and agencies make marketing decisions. Its platform enables marketers to plan and activate cross-channel, programmatic media campai…
Est. 120,000 EUR
We are looking for a DevSecOps and Application Security Lead to join our team and build our application security from scratch. In this role, you will lead the security direction within our department, focusing on integra…
Est. 80,000 GBP
Company Overview Metrea delivers effects-as-a-service to national security partners across five domains and more than a dozen mission areas. These include airborne ISR, electronic warfare, secure communications, aerial r…
Est. 80,000 EUR
What you’ll be doing Work with a team of DevOps and DBA professionals Improve existing infrastructure and processes across the countries we’re deployed in, as well as streamlining processes to deploy to new countries in…
Est. 45,000 EUR
About Ubiquiti At Ubiquiti Inc., we create technology platforms for Businesses, Smart Homes, and Internet Service Providers, driven by our goal to connect everyone, everywhere. To date, Ubiquiti has shipped over 100 mill…
About Us The Financial Times is one of the world’s leading news organisations, globally recognised for its authority, integrity and accuracy, with a mission to deliver quality information and services worldwide. At the F…
Emergent builds autonomous coding agents that replace traditional software development by generating, testing, and deploying production applications directly from plain-language intent. Our systems run in production at g…
Est. 120,000 USD
N-iX is looking for Lead DevOps Engineer (AWS Migration) - to join the team About the client Our client is a leading European online car marketplace, serving over 30 million monthly users across 18 countries. As a Lead D…
The Ministry of Finance is responsible for the Government Procurement (GP) policies, which govern how government agencies conduct their procurement. With evolving needs of our public officers and changing procurement lan…
About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…
Est. 250,000 USD
Role: Principal Architect Reports to: Global Head of Technology (SVP Engineering)Function: Platform & Product Architecture (Cloud-Native Data/SaaS) About Reltio At Reltio®, an SAP Company, we believe data should fuel…