inhousefyi
← Back to listings

Director - Backend Engineering - AI Infra

CoupangBengaluru, Karnataka, India · Posted 2 months ago
Full-time
Apply now

Description

Job Description: Director of Backend Engineering (AI Infrastructure)

Company Introduction

We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did I ever live without Coupang?” Born out of an obsession to make shopping, eating, and living easier than ever, we are collectively disrupting the multi-billion-dollar commerce industry from the ground up and establishing an unparalleled reputation for being leading and reliable force in South Korean commerce.

We are proud to have the best of both worlds — a startup culture with the resources of a large global public company. This fuels us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurs surrounded by opportunities to drive new initiatives and innovations. At our core, we are bold and ambitious people that like to get our hands dirty and make a hands-on impact. At Coupang, you will see yourself, your colleagues, your team, and the company grow every day.

Role Overview

We are seeking a visionary Director of Backend Engineering to lead the teams responsible for the software "brain" that manages our global AI Physical Infrastructure. You will oversee the development of the SDN orchestrators, automated fleet management systems, and the high-performance storage backends that power our AI training and inference clusters.

Your mission is to abstract the complexity of specialized hardware (NVIDIA/HPC) into a seamless, automated, and hyper-reliable cloud platform.

Key Responsibilities

1. Strategic Leadership & Fleet Orchestration

  • Software-Defined Infrastructure: Lead the design and delivery of an SDN Orchestrator to automate complex GPU networking (InfiniBand/RoCE/NVLink) and core DC routing.
  • Fleet Health Automation: Oversee the development of backend services for GPU Health & Fault Detection, automating the lifecycle from burn-in and diagnostics to global RMA workflows.
  • Capacity & Traffic Engineering: Drive the backend logic for global traffic routing, load balancing (NGINX/Kong), and IPAM to ensure zero-bottleneck training environments.

2. Data & Storage Systems

  • HPC Data Pipelines: Collaborate with storage engineers to build backend interfaces for Parallel File Systems (Lustre, Weka, VAST etc.), ensuring high-throughput data delivery to compute nodes.
  • Storage Durability: Direct the backend strategy for AI Object Storage, focusing on high durability and low-latency retrieval for massive datasets.

3. Engineering Excellence

  • Scalable Architecture: Act as the final technical authority for AI Infra Architecture, ensuring systems are resilient, multi-region, and capable of sub-millisecond coordination.
  • DevOps & IaC Culture: Champion a "Hardware-as-Code" mindset, utilizing Python, Ansible, and Terraform to eliminate manual intervention in DC operations.

4. Team Development

  • Lead a multi-disciplinary org including Backend Developers, SDN Engineers, and Infra Ops teams, AI Infra Engineering
  • Establish 24/7 L1/L2/L3 operational standards to maintain > 99.99% availability of the AI fleet.

Required Qualifications

  • Experience: 15+ years in Backend Engineering, with at least 5 years in a leadership role managing complex infrastructure (Cloud, FinTech, or HPC).
  • Deep Infrastructure Knowledge: Proven experience with Linux internals, hardware-software interfaces (drivers/firmware), and distributed systems.
  • Networking Mastery: Solid understanding of L2/L3 networking, and ideally, specialized fabrics like InfiniBand or RoCE.
  • The Stack: Professional proficiency in Python, Go, or C++, and deep experience with Terraform, Kubernetes, and Ansible.
  • Large-Scale Data: Experience managing high-performance storage backends (GPFS, Lustre, or equivalent parallel systems).
  • Hardware Savvy: You don't just write code; you understand power envelopes, liquid cooling constraints, and GPU architecture (NVIDIA/HPE/Dell).

Preferred Skills

  • Experience building custom SDN controllers or orchestration layers from scratch.
  • Direct experience with NVIDIA or GPUDirect technologies.
  • Previous success in a "Hyper-scale" environment (AWS, Azure, GCP, Meta, AI Cloouds etc.).

Recruitment Process

  • Application Review - Phone Interview - Onsite (or Virtual Onsite) Interview – Offer

The exact nature of the recruitment process may vary according to the specific job and may be changed due to scheduling or other circumstances.

  • Interview schedules and the results will be informed to the applicant via the e-mail address submitted at the application stage.

Details to Consider

  • This job posting may be closed prior to the stated end date for application if all openings are filled.
  • Coupang has the right to rescind an offer of employment if a candidate is found to have submitted false information as part of the application process.
  • Those eligible for employment protection (recipients of veteran’s benefits, the disabled, etc.) may receive preferential treatment for employment in accordance with applicable laws.

Privacy Notice

Your personal information will be collected and managed by Coupang as stated in the Application Privacy Notice located below.

https://privacy.coupang.com/en/land/jobs/

Similar jobs

Scale AILondon, United Kingdom

Est. 120,000 GBP

At Scale AI, our mission is to accelerate the development of AI applications. For 10 years, Scale has been the leading AI data foundry, helping fuel the most exciting advancements in AI, including: generative AI, defense…

Full-time
NscaleHouston, Texas, United States

Est. 200,000 USD

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time
NscaleLondon, United Kingdom

Est. 80,000 GBP

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time
NscaleHouston, Texas, United States

Est. 182,500 USD

. About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superio…

Full-time
Scale AIBudapest, Hungary

Est. 84,000 EUR

At Scale AI, we’re not just building AI tools—we’re pioneering the next era of enterprise AI. As businesses race to harness the power of Generative AI, Scale is at the forefront, delivering cutting-edge solutions that tr…

Full-time
Scale AISan Francisco, California, United States

Est. 340,200 USD

Director of Engineering, Physical AI Role Overview The Director of Engineering will report to the General Manager of Physical AI, and will be responsible for leading a multi-disciplinary engineering organization. In this…

Full-time
Scale AILondon, United Kingdom

Est. 150,000 GBP

Director, Forward Deployed Engineering At Scale AI, we're not just building AI tools—we're pioneering the next era of enterprise AI. The world's largest enterprises are racing to deploy Generative AI into the core of the…

Full-time
NscaleBatam, Riau Islands, Indonesia

Note to candidates: this role is based out of Batam, Riau Islands About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise c…

Full-time
Scale AILondon, United Kingdom

Est. 95,000 GBP

At Scale AI, we’re not just building AI tools—we’re pioneering the next era of enterprise AI. As businesses race to harness the power of Generative AI, Scale is at the forefront, delivering cutting-edge solutions that tr…

Full-time
NscaleNarvik, Norway

Est. 840,000 NOK

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time
Scale AILondon, United Kingdom

Est. 120,000 GBP

As a Software Engineer on the AI Infrastructure team, you'll help build and evolve our agent sandboxing platform — the secure, high-performance code execution layer powering our agentic workflows, deployed across both in…

Full-time
NscaleIceland

Please note: This role requires you to be working on-site in the Datacenter About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large en…

Full-time
NscaleHouston, Texas, United States

Est. 140,000 USD

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time

Est. 150,000 USD

Staff Platform Engineer Location: This position is a remote role based in the US Company Overview Robots & Pencils is an applied AI engineering firm building the next frontier of business architecture. We design and…

Full-timeRemote
NscaleHouston, Texas, United States

Est. 165,000 USD

Role Overview As a Technical Program Manager (TPM) for AI Infrastructure Operations, you will be the operational backbone of our high-scale, high-performance AI and High-Performance Computing (HPC) environment. You will…

Full-time
Scale AISan Francisco, California, United States

Est. 202,500 USD

AI Infrastructure Engineer - Agent Sandbox Platform As a Software Engineer on the AI Infrastructure team, you'll help build and evolve our agent sandboxing platform — the secure, high-performance code execution layer pow…

Full-time
NscaleRemote

Est. 215,000 USD

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-timeRemote
NscaleHouston, Texas, United States

Est. 175,000 USD

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time
DataikuRemote

Est. 80,000 EUR

Dataiku is the Platform for AI Success, the enterprise orchestration layer for building, deploying, and governing AI. In a single environment, teams design and operate analytics, machine learning, and AI agents with the…

Full-timeRemote
Scale AINew York, New York, United States

Est. 243,225 USD

Scale GP is building the infrastructure that makes enterprise AI seamless. We are looking for a Senior Infrastructure Engineer to act as a primary technical lead, engineering the 'paved road' for our knowledge retrieval…

Full-time
NscaleLondon, United Kingdom

Est. 120,000 GBP

About Nscale Nscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform. We own the data centres, software, and applications that power today's AI stack using sustainable technology sol…

Full-time
NscaleHouston, Texas, United States

Est. 205,000 USD

About the role: The AI infrastructure market moves weekly. As the Director of Competitive Insights, you will not just track competitors; you will translate complex market movements—such as new Nvidia allocations, legacy…

Full-time

Robots & Pencils is an applied AI engineering firm building the next frontier of business architecture. We design and ship AI co-workers that integrate into enterprise operations and deliver measurable results for ou…

Full-timeRemote
Scale AISan Francisco, California, United States

Est. 307,800 USD

Role Overview The Director of Product Management, Physical AI will report to the General Manager of Physical AI, and will be responsible for defining the long-term strategy and business positioning of our frontier data p…

Full-time
NscaleUnited States

Est. 125,000 USD

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time
NscaleUnited States

Est. 230,000 USD

About Nscale NScale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. NScale enables AI-focused companies to achieve superior…

Full-time
Scale AISan Francisco, California, United States

Est. 201,825 USD

Scale GP (Scale Generative AI Platform) is an enterprise-grade AI platform that provides APIs for knowledge retrieval, inference, evaluation, and more. We are looking for a strong engineer to join our team and help us bu…

Full-time
NscaleSingapore, Singapore

About NscaleNscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior r…

Full-time

Est. 120,000 PLN

Xebia is a global AI-first, digital transformation, and engineering partner. With over 25 years of experience and a team of 5,000 professionals across 16 countries, we help organizations design and build scalable product…

Full-timeRemote

Est. 84,000 PLN

Xebia is a global AI-first, digital transformation, and engineering partner. With over 25 years of experience and a team of 5,000 professionals across 16 countries, we help organizations design and build scalable product…

Full-timeRemote