inhousefyi
← Back to listings

Site Reliability Engineer Senior

CobreColombia · Posted 1 month ago
Full-time
Apply now

Description

What is Cobre, and what do we do?
Cobre is Latin America’s leading instant b2b payments platform. We solve the region’s most complex money movement challenges by building advanced financial
infrastructure that enables companies to move money faster, safer, and more efficiently.
We enable instant business payments—local or international, direct or via API—all from a single platform.
Built for fintechs, PSPs, banks, and finance teams that demand speed, control, and efficiency. From real-time payments to automated treasury, we turn complex financial processes into simple experiences.
Cobre is the first platform in Colombia to enable companies to pay both banked and unbanked beneficiaries within the same payment cycle and through a single interface.
We are building the enterprise payments infrastructure of Latin America!
What we are looking for:
The Cobre Infrastructure team and their SRE engineers are professionals who face the daily challenges that allow us to improve the technological level of our products. We enjoy
each project or task, giving 100% and learning from each other. Our main goal is to maintain the reliability of our systems. To achieve this, we collaborate with other teams to find the most effective solutions, maintain high-reliability processes, adopt the necessary safety measures and optimize time and cost in every decision.
What would you be doing:
● Support teams in defining the infrastructure that will support the solution architecture.
● Support all the infrastructure (aws services and k8s clusters) and company products Culture zero-downtime deployments.
● Assisting with troubleshooting application issues and incidents related with infrastructure services.
● Review code instrumentation with development teams and ensure necessary dashboards are created to monitor.
● Document and maintain runbooks and procedures, automate as much as possible (AI-driven auto-remediation for incidents) to reduce MTTR.
● Perform periodic load and scalability testing to establish baselines, drift, and capacity planning.
● Design and implement peak readiness reviews for anticipated high-volume times.
● Lead weekly operational state reviews covering performance trends, anomalies, errors and other availability events with SREs, product owners, and development teams.
● Contribute to incident management, postmortems, and reliability reviews.
● Socialize SRE culture across teams within the organization to publicize the value of SRE, mentor and train other engineers around proactive reliability decision making and planning.

What do you need:
● Proven experience of at least 5 years as SRE or DevOps, with a strong focus on highly available and scalable environments, cloud infrastructure, observability, and incident management.
● In-depth technical knowledge of microservices architecture and cloud platforms (e.g., AWS, Kubernetes), along with proficiency in Infrastructure as Code (IaC) tools (e.g., Terraform, Pulumi).
● Strong mindset for automation and continuous improvement with a huge interest in AIOps / AI-driven auto-remediation (n8n, aws bedrock, python scripting …)
● Understanding of secure-by-design infrastructure principles.
● Exposure to GitOps and declarative configuration patterns.
● Basic knowledge in Port.io or any Internal Developer Platform (Backstage, Cortex).
● Strong understanding of monitoring, logging, and alerting tools, with a track record of improving system reliability and performance. (e.g., NewRelic, Datadog, Cloudwatch…)
● Proven experience troubleshooting, mitigating, and resolving issues in a distributed system.
● Ability to define and execute the SRE strategy, aligning it with company goals and driving the adoption of SRE practices across multiple teams.
● Resilience in facing challenges and promoting a fail-fast, learn-fast culture that embraces innovation and experimentation.
● Exceptional communication skills to effectively convey complex technical concepts to both technical and non-technical stakeholders.
● Ability to actively listen and understand diverse team and stakeholder needs,demonstrating empathy in decision-making and conflict resolution

Similar jobs

CobreColombia

What is Cobre, and what do we do? Cobre is Latin America’s leading instant b2b payments platform. We solve the region’s most complex money movement challenges by building advanced financial infrastructure that enables co…

Full-time
CobreRemote

What is Cobre, and what do we do? Cobre is Latin America’s leading instant b2b payments platform. We solve the region’s most complex money movement challenges by building advanced financial infrastructure that enables co…

Full-timeRemote
CobreRemote

Est. 40,000 USD

What is Cobre and what do we do? Cobre is Latin America’s leading instant b2b payments platform. We solve the region’s most complex money movement challenges by building advanced financial infrastructure that enables com…

Full-timeRemote
EncoraRemote

Job Title: Site Reliability Engineer (SRE)Key Skills: Kubernetes, AWS/Azure/GCP, Terraform, Python, Observability, CI/CDExperience: +6 YOE.Location: Costa Rica, Peru, Colombia, and Bolivia.Mode: Remote. We at Coforge are…

Full-timeRemote
EncoraRemote

Job Title: DevOps EngineerKey Skills: DevOps, Cloud Operations, AWS, TerraformLocation: BrazilMode: RemoteWe at Coforge are hiring DevOps Engineer with the following skill set.Key Responsibilities: Participate in a bi-we…

Full-timeRemote
CobreBogotá, Cundinamarca, Colombia

What is Cobre and what do we do? Cobre is Latin America’s leading instant b2b payments platform. We solve the region’s most complex money movement challenges by building advanced financial infrastructure that enables com…

Full-time
NEORISColombia

NEORIS is a Digital accelerator that helps companies enter the future, having 20 years of experience as Digital Partners of some of the largest companies in the world. We have more than 4,000 professionals in 11 countrie…

Full-time
Frontend Engineer9 months ago
CobreRemote

What is Cobre, and what do we do? Cobre is a corporate treasury and payments platform designed to elevate the way Latin American companies manage their finances. At Cobre, we build CFO-tech on top of owned payment rails…

Full-timeRemote
CobreCiudad de México, México

What is Cobre and what do we do? Cobre is Latin America’s leading instant b2b payments platform. We solve the region’s most complex money movement challenges by building advanced financial infrastructure that enables com…

Full-time
EncoraUnited States

Est. 124,000 USD

Application Support Engineer (Site Reliability Engineer) Location: USAJob Type: Full-Time, no visa sponsorship available Coforge is seeking a Senior Application Support Engineer (SRE) to join our dynamic team of consulta…

Full-time
StraitsXJakarta, Jakarta, Indonesia

About The Role The Site Reliability Engineering (SRE) team architects, builds, and maintains the rock-solid infrastructure that applications rely on. At the Senior Level, you own reliability, performance, and cost outcom…

Full-time
Coupang InternalSeoul, South Korea

Sr. Director, Site Reliability Engineering Coupang operates one of the largest and most complex technology platforms in the world. We are seeking a Senior Director, Site Reliability Engineering (Head of SRE) to define an…

Full-time
CoupangSeoul, South Korea

Sr. Director, Site Reliability Engineering Coupang operates one of the largest and most complex technology platforms in the world. We are seeking a Senior Director, Site Reliability Engineering (Head of SRE) to define an…

Full-time
CobreColombia

What is Cobre, and what do we do? Cobre is Latin America’s leading instant b2b payments platform. We solve the region’s most complex money movement challenges by building advanced financial infrastructure that enables co…

Full-time
CobreRemote

Est. 120,000 USD

What is Cobre, and what do we do? Cobre is Latin America’s leading instant b2b payments platform. We solve the region’s most complex money movement challenges by building advanced financial infrastructure that enables co…

Full-timeRemote
NICEMedellin, Colombia

At NiCE, we don’t limit our challenges. We challenge our limits. Always. We’re ambitious. We’re game changers. And we play to win. We set the highest standards and execute beyond them. And if you’re like us, we can offer…

Full-time
Spin CareersRemote

Objective of the RoleThe Senior SRE Engineer is a highly experienced role responsible for leading the enhancement and maintenance of the reliability, availability, and performance of the company's IT infrastructure and a…

Full-timeRemote
CobreRemote

What is Cobre, and what do we do? Cobre is Latin America’s leading instant b2b payments platform. We solve the region’s most complex money movement challenges by building advanced financial infrastructure that enables co…

Full-timeRemote
IC5 - Staff SRE3 months ago
Spin CareersRemote

Objective of the RoleThe Staff Infrastructure Engineer – SRE is a senior technical leader responsible for architecting and scaling the reliability, availability, and performance of critical infrastructure platforms. This…

Full-timeRemote

Est. 120,000 USD

We are representing a leading force in the decentralized exchange (DEX), and seeking a high-caliber technical leader to architect the backbone of a global financial ecosystem. In this role, you will bridge the gap betwee…

Full-timeRemote
CobreBogotá, Bogotá, Colombia

About Cobre Cobre is Latin America's leading instant B2B payments platform. We're building the financial infrastructure that enables companies across the region to move money faster, safer, and more efficiently—whether l…

Full-time
EnsonoRemote

Est. 90,000 GBP

Who are we? Ensono is a global technology services provider dedicated to helping organizations navigate the complexity of digital transformation. Through Ensono Product, Consulting & Technology, our dedicated consult…

Full-timeRemote
EncoraBogotá, Colombia

Job Title: Technical Program Manager – Cloud, Infrastructure & SREKey Skills: Cloud, Infrastructure, DevOps, SRE, Program Governance, Risk Management, Stakeholder ManagementExperience: 7+ years (minimum 5+ years in T…

Full-time
CobreRemote

Est. 120,000 USD

What is Cobre, and what do we do? Cobre is Latin America’s leading instant b2b payments platform. We solve the region’s most complex money movement challenges by building advanced financial infrastructure that enables co…

Full-timeRemote
CobreRemote

What is Cobre and what do we do Cobre is Latin America’s leading instant b2b payments platform. We solve the region’s most complex money movement challenges by building advanced financial infrastructure that enables comp…

Full-timeRemote
Hyphen Connect LimitedSingapore, Singapore

We are representing a leading force in the decentralized exchange (DEX), and seeking a high-caliber technical leader to architect the backbone of a global financial ecosystem. In this role, you will bridge the gap betwee…

Full-time
StraitsXJakarta, Jakarta, Indonesia

About the Role The Site Reliability Engineering (SRE) team architects, builds, and maintains the rock-solid infrastructure that applications rely on. We work closely with development teams to ensure scalability, reliabil…

Full-time
EncoraBogotá, Colombia

Job Title: Technical Program Manager Key Skills: Technical Program Management, Cloud, Infrastructure, DevOps, Site Reliability Engineering (SRE), Program Governance, Risk Management, Stakeholder Management, Cloud Transfo…

Full-time
Impact.comCape Town, South Africa

About impact.com impact.com is the world’s leading commerce partnership marketing platform, transforming the way businesses grow by enabling them to discover, manage, and scale partnerships across the entire customer jou…

Full-time
OneSpanNoida, Uttar Pradesh, India

At OneSpan, we specialize in digital identity and anti-fraud solutions that create exceptional and secure experiences.We are looking for a Site Reliability Engineer to join our growing platform team in Delhi NCR. You wil…

Full-time