inhousefyi
← Back to listings

Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)

SUBMERRemote · Posted 3 months ago
Full-timeRemoteEst. 80,000 EUR
Apply now

Description

About Radian Arc.
Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments.

What impact you will have

Mission: Install, validate, operate, and maintain regional and core GPU cloud deployments across datacenter environments.

This role owns the practical deployment and operational readiness of the platform stack in the field. It bridges infrastructure engineering, datacenter operations, networking, host systems, storage, and platform operations. The engineer is responsible for taking a validated architecture and BOM and turning it into a working production environment, including physical deployment coordination, rack and cabling validation, switch and host bring-up, firmware and BIOS validation, operating-system installation, GPU and DPU validation, storage integration, platform stack installation, acceptance testing, and operational handover.

The role is intentionally hands-on and cross-domain. It is not a pure datacenter technician role, and it is not a pure platform engineering role. It is the person who can work in a datacenter, understand cabling and optics, debug host and network issues, validate GPU servers, support platform installation, and coordinate with engineering when problems arise.

The role is especially important as the platform evolves from regional deployments toward HGX-based GPU systems, east-west fabrics, and core AI infrastructure. This profile is focused on actual installation, commissioning, maintenance, and operational readiness.

What you’ll do

Datacenter deployment and commissioning

  • Coordinate physical deployment activities with datacenter providers, integrators, logistics teams, and internal engineering.
  • Validate rack layouts, elevations, power feeds, airflow assumptions, cable paths, and labeling before installation.
  • Support rack-and-stack activities for GPU nodes, CPU nodes, storage nodes, switches, firewalls, routers, serial/OOB equipment, PDUs, and supporting infrastructure.
  • Validate fibre and copper cabling against the deployment design, including OOB, north-south, east-west, storage, and management networks.
  • Validate optics, transceivers, link speeds, breakout cables, port mappings, and redundancy assumptions.
  • Maintain accurate as-built documentation, including rack elevations, cable maps, port maps, serial numbers, asset records, IP allocations, and change records.

Host and hardware bring-up

  • Bring up GPU servers, platform servers, storage nodes, and supporting infrastructure.
  • Validate BIOS, BMC, firmware, NIC, DPU, GPU, NVMe, RAID/HBA, and platform firmware versions.
  • Configure and validate BMC access using Redfish/IPMI and OOB management networks.
  • Validate GPU visibility, PCIe topology, NUMA layout, thermals, power behavior, and hardware health.
  • Run hardware acceptance tests, burn-in tests, GPU stress tests, network tests, and storage validation before handover.
  • Troubleshoot hardware issues across servers, GPUs, DPUs, NICs, optics, cables, disks, memory, firmware, and BIOS.

Network deployment support

  • Support deployment and validation of OOB, north-south, storage, and east-west networking.
  • Work with networking engineering to apply and validate switch configurations.
  • Validate BGP, ECMP, VLAN/VRF segmentation, EVPN/VXLAN where applicable, OVS/OVN integration, and routing reachability.
  • Validate VyOS routers, OOB firewalls, transit routers, Citrix NetScaler/WAF, and customer connectivity.
  • Support RoCE/RDMA fabric validation for distributed AI workloads where applicable.
  • Troubleshoot practical network issues such as link flaps, optics issues, incorrect polarity, MTU mismatches, route leaks, VLAN errors, packet loss, PFC/ECN issues, and fabric congestion.
  • Support integration with NVIDIA Cumulus / Spectrum-X environments, and assist with Cisco or SONiC-based alternatives if those become part of the roadmap.

Platform stack installation and validation

  • Support installation and validation of the platform stack across regional and core deployments.
  • Install and validate host operating systems, kernel versions, NVIDIA drivers, Mellanox/NVIDIA OFED or inbox drivers, CUDA compatibility, Docker/containerd, KVM/QEMU, and platform agents.
  • Support CloudStack-based deployments and Kubernetes/KubeVirt-based deployments.
  • Validate GPU passthrough, SR-IOV, BlueField NIC/DPU behavior, VM networking, and container networking.
  • Support Kubernetes node registration, GPU Operator validation, CSI validation, CNI validation, and node lifecycle workflows.
  • Support storage integration with StorPool, Weka, local NVMe, or other supported storage platforms.
  • Execute acceptance tests and produce deployment readiness reports.

Operational maintenance and Day-2 support

  • Perform controlled maintenance activities such as firmware upgrades, switch upgrades, host OS updates, GPU driver updates, BIOS changes, and hardware replacements.
  • Support incident response for infrastructure issues affecting GPU nodes, hosts, networking, storage, or platform components.
  • Perform root-cause analysis for deployment and operational failures.
  • Maintain runbooks for installation, validation, upgrade, rollback, troubleshooting, and handover.
  • Work with engineering to turn repeated operational issues into automation, better validation, or platform improvements.
  • Participate in on-call or escalation rotations for regional and core environments where appropriate.

Platform observability and validation

  • Ensure telemetry is correctly configured for hosts, GPUs, DPUs, switches, storage, OOB devices, and platform components.
  • Validate Zabbix, Prometheus, Grafana, Loki, DCGM/NVML, NVIDIA NetQ or equivalent telemetry sources.
  • Confirm that deployment health checks, hardware alerts, performance dashboards, and operational alarms work before production handover.
  • Support performance baseline testing for GPU, network, storage, and host layers.
  • Assist with NCP-related validation and benchmarking.

Cross-team coordination

  • Work closely with the Senior Director of Infrastructure Operations.
  • Work closely with Staff Network, Staff Storage, Sr Hardware/Infrastructure, Sr Platform, Sr Fleet Automation, Observability, Product Engineering, Sales Engineering, and Service Delivery roles.
  • Provide field feedback into reference architectures, BOMs, rack layouts, cabling standards, deployment playbooks, and validation procedures.
  • Coordinate with external vendors including datacenter providers, systems integrators, server vendors, storage vendors, NVIDIA, and networking vendors.
  • Act as the practical field escalation point when architecture, BOM, datacenter conditions, and platform implementation do not align.

Technical Stack

Hardware and datacenter

  • GPU servers: L40S, RTX 6000 Pro, H200, B200/B300-class systems, HGX systems, and future NVL72-style rack-scale systems.
  • CPU/platform servers.
  • Storage nodes and JBODs.
  • PDUs, BMCs, serial/OOB, firewalls, routers, switches.
  • Rack layouts, power feeds, airflow, cold/hot aisle containment, cabling, optics.
  • DTC/DLC cooling.

Host and systems

  • Ubuntu Linux.
  • Linux networking.
  • BIOS/BMC/firmware lifecycle.
  • Redfish, IPMI.
  • NVIDIA drivers, CUDA, DCGM/NVML.
  • Mellanox/NVIDIA NICs, BlueField DPUs.
  • KVM/QEMU, VFIO, PCI passthrough.
  • Docker/containerd.

Networking

  • NVIDIA Spectrum/Cumulus.
  • VyOS.
  • OVS/OVN.
  • BGP, ECMP, VLAN, VRF, EVPN/VXLAN.
  • RoCE/RDMA.
  • SR-IOV.
  • Citrix NetScaler / WAF.
  • OOB and break-glass access.

Platform

Similar jobs

SUBMERRemote

Est. 120,000 EUR

Location & work modality: Europe/ Remote Start: Aug 2026 Type of Contract: Full time or Contract About Radian Arc Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificia…

Full-timeRemote
Firmus TechnologiesLaunceston, Tasmania, Australia

Firmus Technologies Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific. Founded in Australia in 2019, our mission is to create the most effi…

Full-time
SUBMERRemote

Est. 120,000 EUR

About Radian Arc.Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. Our teams…

Full-timeRemote
NebiusRemote

Est. 175,000 USD

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to…

Full-timeRemote
NscaleSingapore, Singapore

About NscaleNscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior r…

Full-time
Astera LabsSan Jose, California, United States

Est. 141,000 USD

Astera Labs (NASDAQ: ALAB) provides rack-scale AI infrastructure through purpose-built connectivity solutions. By collaborating with hyperscalers and ecosystem partners, Astera Labs enables organizations to unlock the fu…

Full-time
AgodaBangkok, Thailand

About Agoda At Agoda, we bridge the world through travel. Our story began in 2005, when two lifelong friends and entrepreneurs, driven by their passion for travel, launched Agoda to make it easier for everyone to explore…

Full-time
Five9Bengaluru, India

Join us in bringing joy to customer experience. Five9 is a leading provider of cloud contact center software, bringing the power of cloud innovation to customers worldwide. Living our values everyday results in our team-…

Full-time
Firmus TechnologiesSingapore, Singapore

Firmus Technologies Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure from model to grid. Founded in Australia in 2019, our mission is to create the most effic…

Full-time
PlayStation GlobalBerlin, Germany

Est. 95,000 EUR

Why Sony Interactive Entertainment? Sony Interactive Entertainment isn’t just the Best Place to Play — it’s also the Best Place to Work. Sony Interactive Entertainment (SIE) is the company behind the PlayStation brand. A…

Full-time
Ursa MajorBerthoud, Colorado, United States

Est. 155,000 USD

The future of aerospace and defense starts here. Ursa Major was founded to revolutionize how America and its allies access and apply high-performance propulsion, from hypersonics to solid rocket motors, satellite maneuve…

Full-time
Five9Bengaluru, India

Join us in bringing joy to customer experience. Five9 is a leading provider of cloud contact center software, bringing the power of cloud innovation to customers worldwide. Living our values everyday results in our team-…

Full-time
NebiusAlabama, United States

Est. 115,000 USD

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to…

Full-time
NscaleSingapore, Singapore

About NscaleNscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior r…

Full-time
Firmus TechnologiesSingapore, Singapore

Firmus Technologies Firmus Technologies is a global leader pioneering the solution to AI’s energy challenge, founded in Australia in 2019 by a visionary team of entrepreneurs. Our mission is to create the most energy-eff…

Full-time
Coupang InternalSeattle, Washington, United States

Est. 165,000 USD

Please complete the attached the Internal Transfer Request Form and submit it.Please make sure you are applying with your Coupang e-mail address. Job Overview As the Director of Product Management for High Performance Co…

Full-time
VectraAustin, Texas, United States

Est. 175,000 USD

Vectra® is the leader in AI-driven threat detection and response for hybrid and multi-cloud enterprises. The Vectra AI Platform delivers integrated signal across public cloud, SaaS, identity, and data center networks in…

Full-time
CoupangSeattle, Washington, United States

Est. 165,000 USD

Company Introduction We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did we ever live without Coupang?” Born out of an obsession to make shopping, eating, and livin…

Full-time
CoupangMountain View, California, United States

Est. 165,000 USD

Company Introduction We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did we ever live without Coupang?” Born out of an obsession to make shopping, eating, and livin…

Full-time
Astera LabsSan Jose, California, United States

Est. 165,000 USD

Astera Labs (NASDAQ: ALAB) provides rack-scale AI infrastructure through purpose-built connectivity solutions. By collaborating with hyperscalers and ecosystem partners, Astera Labs enables organizations to unlock the fu…

Full-time
Myriad360Remote

Est. 165,000 USD

Who You Are We are seeking a Senior Director of Enterprise Infrastructure to lead strategy, architecture, and delivery across enterprise storage, virtualization, hyperconverged infrastructure (HCI), disaster recovery (DR…

Full-timeRemote
NavVisMunich, Germany

Est. 120,000 EUR

OUR VISION When people use our technology to "bridge the gap" between the physical and digital worlds, they don’t just capture reality - they create a new one. In this new reality, they are smarter, more productive, more…

Full-time
SB EnergySan Diego, California, United States

Est. 125,000 USD

Do you want to work with high-caliber professionals in a dynamic and growing company? Are you entrepreneurial, hard-working, and collegial? Join us at SB Energy, a leading infrastructure company backed by SoftBank Group…

Full-time
QuEra Computing, Inc.Tsukuba, Japan

Role Overview We are looking for a Sr. Control System Engineer/Site Reliability Engineer (SRE) to integrate and

Full-time
QuEra Computing, Inc.Boston, Massachusetts, United States

Est. 124,000 USD

Role Overview We are looking for a Control System Engineer/Site Reliability Engineer (SRE) to integrate and maintain

Full-time
GlanceBangalore, India

Glance AI is an AI commerce platform shaping the next wave of e-commerce with inspiration-led shopping, less about searching for what you want and more about discovering who you could be. Operating in 140 countries, Glan…

Full-time
GlanceBangalore, India

Glance AI is an AI commerce platform shaping the next wave of e-commerce with inspiration-led shopping, less about searching for what you want and more about discovering who you could be. Operating in 140 countries, Glan…

Full-time
NebiusIsrael

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to…

Full-time
NebiusRemote

Est. 140,000 USD

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to…

Full-timeRemote
NavVisMunich, Germany

Est. 90,000 EUR

OUR VISION When people use our technology to "bridge the gap" between the physical and digital worlds, they don’t just capture reality - they create a new one. In this new reality, they are smarter, more productive, more…

Full-time