Description
About The Role
The Site Reliability Engineering (SRE) team architects, builds, and maintains the rock-solid infrastructure that applications rely on. At the Senior Level, you own reliability, performance, and cost outcomes for the systems under your area end-to-end, not just executing well-defined tasks, but deciding between tradeoffs, scoping ambiguous problems, and driving process and system improvements that span teams. You'll work closely with development, security, and product teams, and mentor other engineers as a technical point of reference for the team.
What You Will Do
- Own the availability, performance, scalability, and security of production systems end-to-end, across cloud (AWS/GCP) and on-premises environments.
- Design and evolve Kubernetes deployment strategy for production workloads.
- Own CI/CD and GitOps pipelines in production (ArgoCD or equivalent) and the Terraform that provisions the infrastructure behind them.
- Diagnose and resolve database performance issues.
- Build and maintain observability that surfaces problems before they become incidents.
- Seek out and implement process and system improvements affecting performance and security, coordinating with multiple stakeholders.
- Scope and lead medium-to-large infrastructure initiatives: gather requirements, prioritize by business impact, and communicate impact to stakeholders.
- Negotiate technical tradeoffs with stakeholders to meet business SLAs.
- Provide technical guidance and mentorship to peers and junior engineers; promote best practices and standards across the team.
- Maintain documentation and process discipline for the systems and incidents you own.
- Lead structured incident investigation, isolating server, database, and application layers with a metrics-first approach, including on-call during high-traffic events.
What Are We Looking For
- At least 4 years of experience in either SRE, DevOps, MLOps, or platform engineering, including senior-level scope at a high-traffic company.
- Deep expertise in one major cloud provider (preferably AWS), with a proven ability to ramp up on the other quickly.Production experience & expertise with Kubernetes & Linux fundamentals
- CI/CD & GitOps (ArgoCD or other equivalent stacks)
- Database performance analysis & monitoring (MySQL, Postgres)
- Observability tooling & standards (Datadog, OpenTelemetry)
- Infrastructure as Code (Terraform)
- Strong working English, verbal & written communication. Strong documentation and process discipline.
- On-call & incident handling experience during high-traffic events.
- Comfortable negotiating with stakeholders to propose technical compromises that meet business SLAs.
- Demonstrated growth mindset and proven ability to own ambiguous scope.
Similar jobs
About the Role The Site Reliability Engineering (SRE) team architects, builds, and maintains the rock-solid infrastructure that applications rely on. We work closely with development teams to ensure scalability, reliabil…
Job Title: Site Reliability Engineer (SRE)Key Skills: Kubernetes, AWS/Azure/GCP, Terraform, Python, Observability, CI/CDExperience: +6 YOE.Location: Costa Rica, Peru, Colombia, and Bolivia.Mode: Remote. We at Coforge are…
Sr. Director, Site Reliability Engineering Coupang operates one of the largest and most complex technology platforms in the world. We are seeking a Senior Director, Site Reliability Engineering (Head of SRE) to define an…
Sr. Director, Site Reliability Engineering Coupang operates one of the largest and most complex technology platforms in the world. We are seeking a Senior Director, Site Reliability Engineering (Head of SRE) to define an…
Est. 90,000 GBP
Who are we? Ensono is a global technology services provider dedicated to helping organizations navigate the complexity of digital transformation. Through Ensono Product, Consulting & Technology, our dedicated consult…
Reolink, a leader in intelligent visual technology for homes and businesses, was founded in 2009 by a group of engineers with a strong commitment to and passion for smarter security solutions.Our products are now trusted…
Job Title: DevOps EngineerKey Skills: DevOps, Cloud Operations, AWS, TerraformLocation: BrazilMode: RemoteWe at Coforge are hiring DevOps Engineer with the following skill set.Key Responsibilities: Participate in a bi-we…
We are representing a leading force in the decentralized exchange (DEX), and seeking a high-caliber technical leader to architect the backbone of a global financial ecosystem. In this role, you will bridge the gap betwee…
Reolink, a leader in intelligent visual technology for homes and businesses, was founded in 2009 by a group of engineers with a strong commitment to and passion for smarter security solutions.Our products are now trusted…
Est. 120,000 USD
We are representing a leading force in the decentralized exchange (DEX), and seeking a high-caliber technical leader to architect the backbone of a global financial ecosystem. In this role, you will bridge the gap betwee…
At OneSpan, we specialize in digital identity and anti-fraud solutions that create exceptional and secure experiences.We are looking for a Site Reliability Engineer to join our growing platform team in Delhi NCR. You wil…
Est. 124,000 USD
Application Support Engineer (Site Reliability Engineer) Location: USAJob Type: Full-Time, no visa sponsorship available Coforge is seeking a Senior Application Support Engineer (SRE) to join our dynamic team of consulta…
Est. 82,000 USD
SummaryWe are seeking an experienced Site Reliability Engineering (SRE) / DevOps Manager to lead our teams and drive reliability roadmaps. As a key player in our leading crypto tax and portfolio tracking platform, you wi…
Jump Trading Group is committed to world class research. We empower exceptional talents in Mathematics, Physics, and Computer Science to seek scientific boundaries, push through them, and apply cutting edge research to g…
Est. 180,000 USD
Summary: We are seeking an experienced Site Reliability Engineering (SRE) / DevOps Manager to lead our teams and drive reliability roadmaps. As a key player in our leading crypto tax and portfolio tracking platform, you…
Summary: We are seeking an experienced Site Reliability Engineering (SRE) / DevOps Manager to lead our teams and drive reliability roadmaps. As a key player in our leading crypto tax and portfolio tracking platform, you…
Senior Site Reliability Engineer I Location San Jose, Costa Rica - Remote Summary of role Own availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s plane…
Est. 95,000 USD
SummaryWe are seeking an experienced Site Reliability Engineering (SRE) / DevOps Manager to lead our teams and drive reliability roadmaps. As a key player in our leading crypto tax and portfolio tracking platform, you wi…
Est. 155,000 USD
Recognized as the No. 1 site trusted by real estate professionals, Realtor.com® has been at the forefront of online real estate for over 25 years, connecting buyers, sellers, and renters with trusted insights and expert…
Title: Senior Site Reliability Engineer - I, Product Area FocusLocation: Noida (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excellence o…
Title: Staff Site Reliability Engineer, Product Area FocusLocation: Noida / Bangalore (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excel…
Est. 120,000 USD
SummaryWe are seeking an experienced Site Reliability Engineering (SRE) / DevOps Manager to lead our teams and drive reliability roadmaps. As a key player in our leading crypto tax and portfolio tracking platform, you wi…
We are seeking a skilled and passionate Engineer to join our team to build and operate a Whole-of-Government (WoG) runtime platform. As a Site Reliability Engineer, you will be responsible for designing and operating Git…
Title: Staff Site Reliability Engineer, Product Area FocusLocation: Noida/ Bangalore (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excell…
Est. 70,000 USD
SummaryWe are seeking an experienced Site Reliability Engineering (SRE) / DevOps Manager to lead our teams and drive reliability roadmaps. As a key player in our leading crypto tax and portfolio tracking platform, you wi…
We are representing a leading force in the decentralized exchange (DEX), and seeking a high-caliber technical leader to architect the backbone of a global financial ecosystem. In this role, you will bridge the gap betwee…
Est. 110,000 USD
Summary: We are seeking an experienced Site Reliability Engineering (SRE) / DevOps Manager to lead our teams and drive reliability roadmaps. As a key player in our leading crypto tax and portfolio tracking platform, you…
About impact.com impact.com is the world’s leading commerce partnership marketing platform, transforming the way businesses grow by enabling them to discover, manage, and scale partnerships across the entire customer jou…
At Reltio®, an SAP Company, we believe data should fuel your success in the enterprise AI era. Our Context Intelligence Platform turns fragmented data into a trusted, connected context so AI agents and systems can act wi…
Est. 140,000 USD
BeyondTrust is a place where you can bring your purpose to life through the work that you do, creating a safer world through our cybersecurity SaaS portfolio. Our culture of flexibility, trust, and continual learning mea…