inhousefyi
← Back to listings

Staff Network Engineer (AI Fabric, Datacenter and Edge Networking) - Radian Arc (EMEA)

SUBMERRemote · Posted 1 month ago
Full-timeRemoteEst. 120,000 EUR
Apply now

Description

Location & work modality: Europe/ Remote

Start: Aug 2026

Type of Contract: Full time or Contract

About Radian Arc

Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments.

What impact you will have

Design, implement, and operate the network infrastructure powering the GPU cloud platform, including high-performance AI fabrics as well as classical datacenter networking components such as routing, security, and external connectivity. This role spans both high-performance east-west networking for distributed AI workloads and north-south connectivity, security, and inter-datacenter transport.

As the first dedicated networking role in the organization, the Staff Network Engineer combines Staff-level architectural ownership, technical direction, and cross-functional influence with hands-on execution across design, deployment, troubleshooting, automation, and operational improvement.

The Staff Network Engineer owns the long-term technical direction and operational strategy for Radian Arc’s AI interconnect networks, designing scalable GPU fabrics and ensuring predictable low-latency performance across distributed training and inference workloads. The role includes designing large-scale RoCE and Ethernet fabrics, guiding architecture decisions, and ensuring operational excellence across global deployments, from hyperscale datacenters to smaller edge locations.

You will collaborate closely with platform, compute, storage, observability, and operations teams to ensure networking is deeply integrated into the overall infrastructure architecture. This role also acts as the senior escalation point for complex networking incidents, driving deep technical investigations and systemic improvements that increase reliability, latency consistency, and operational maturity across the platform.

Because this is currently the primary networking role in the company, the position is intentionally hybrid: you are expected to operate at L6 / Staff in terms of technical direction, standards, cross-team influence, and long-term design, while also directly executing critical networking work that, in a larger organization, would be distributed across multiple engineers.

What you’ll do

AI Fabric & HPC Networking

  • Design and operate high-performance GPU networking fabrics supporting distributed AI workloads.
  • Architect large-scale RoCE fabrics optimized for distributed training and inference.
  • Optimize network performance for GPU communication patterns and east-west traffic.
  • Design fabric topologies such as:

○ Leaf-spine

○ Fat-Tree

○ Rail architectures

○ Multi-plane

    • Implement high-performance networking technologies including:

○ RDMA

○ RoCE

○ High-bandwidth east-west fabrics

○ Spectrum-X

  • Collaborate with compute teams to support distributed training frameworks and GPU communication libraries.
  • Define reference architectures and design principles for AI fabrics so future deployments follow reusable standards rather than one-off implementations.
  • Evaluate architectural trade-offs across performance, resilience, cost, operability, and deployment speed, and make clear recommendations to stakeholders.

Datacenter Networking

  • Design and operate Layer-2 and Layer-3 datacenter networks.
  • Implement scalable routing architectures based on BGP and ECMP.
  • Design tenant network isolation mechanisms across multi-tenant environments.
  • Implement and maintain:

○ Network bridges

○ Routing stacks

○ Overlay networking systems

  • Maintain north-south ingress/egress routing and traffic management.
  • Define standards and reusable patterns for segmentation, routing, and overlay integration across platform deployments.

Technologies include:

  • VyOS routers
  • Linux networking stacks
  • OVS / OVN
  • BGP / ECMP
  • VLAN / VRF segmentation

Security & Edge Connectivity

  • Deploy and maintain north-south security infrastructure
  • Implement WAF and application-layer protections
  • Integrate security controls with platform services

Technologies include:

  • Citrix NetScaler / Citrix WAF
  • TLS termination
  • DDoS mitigation
  • API and proxy gateway protection

Inter-Datacenter Networking

  • Design and operate private interconnects between datacenters
  • Implement and maintain dark fiber ring architectures
  • Operate high-capacity WAN connectivity between regions
  • Integrate datacenter fabrics into a global backbone network
  • Define scalable design principles for backbone evolution, inter-site routing, redundancy, and failure-domain isolation

Technologies include:

  • DWDM / dark fiber transport
  • BGP inter-site routing
  • Redundant fiber ring architectures
  • 100–400G optical transport
  • Spectrum-XGS

Engineering Execution & Delivery

  • Lead end-to-end engineering delivery of networking infrastructure, from design and labvalidation to production deployment
  • Validate network BOMs together with procurement and deployment teams
  • Provide detailed input into datacenter layouts and rack elevations
  • Drive capacity planning, performance modeling, and scaling strategies
  • Ensure network changes are executed safely with minimal customer impact
  • Act as both the architectural owner and the practical execution lead for critical network initiatives during the build-out phase of the networking function
  • Establish deployment standards, validation criteria, rollback approaches, and acceptance patterns that future engineers and teams can reuse

Operational Excellence & Reliability

  • Own operational performance and reliability of networking infrastructure
  • Drive automation for:

○ Provisioning

○ Configuration management

○ Monitoring

○ Lifecycle management

  • Improve day-2 operations through automation and operational tooling
  • Lead incident response and root-cause analysis for major network events
  • Define and track SLAs, SLOs, and reliability metrics
  • Translate major incidents and operational pain points into durable standards, design changes, and long-term architectural improvements
  • Establish measurable benchmarks for reliability, latency consistency, operability, and recovery behavior across network deployments.

Cross-Functional Collaboration

  • Work closely with infrastructure, platform, SRE, compute, storage, observability, and datacenter operations teams.
  • Provide technical leadership across infrastructure initiatives.
  • Communicate architectural decisions, trade-offs, and risks clearly to stakeholders.
  • Influence the long-term platform networking roadmap and architecture.
  • Act as the primary networking design authority across the organization, guiding adjacent teams on how networking constraints and capabilities should shape platform decisions.
  • Raise the technical bar by mentoring engineers in adjacent domains and helping build the future networking function.

Technical Stack

Datacenter Networking

  • BGP
  • EVPN / VXLAN
  • ECMP
  • VLAN / VRF
  • OVS / OVN
  • Linux networking
  • BlueField DPU

Routing & Control Plane

  • VyOS
  • BGP-based routing architectures
  • ECMP fabrics

Security

  • Citrix NetScaler / WAF
  • DDoS protection

Transport & Backbone

  • Dark fiber
  • Metro fiber rings
  • DWDM transport
  • 100–1600G optical networking

AI Networking

  • RDMA
  • RoCE
  • GPU fabrics
  • Large-scale east-west compute networking
  • Congestion control

What you'll need

Core Experience

  • Strong hands-on experience designing and operating large-scale datacenter networks
  • Expert knowledge of modern networking protocols including:

○ BGP

○ OSPF

○ ECMP

○ EVPN / VXLAN

  • Proven experience operating high-speed Ethernet networks in production environments
  • Experience operating NVIDIA / Mellanox networking platforms
  • Experience owning both architecture and direct implementation in lean or fast-scaling environments is strongly preferred

Advanced AI Fabric Networking Expertise

The candidate should have deep expertise in designing and operating networking fabrics optimized for large-scale GPU clusters and distributed AI workloads.

This includes a strong understanding of GPU communication patterns and the networking requirements of distributed training and inference systems.

Relevant expertise includes:

  • Deep understanding of NCCL communication patterns and their impact on network topology and performance.
  • Experience tuning RoCE fabrics for large-scale GPU clusters.
  • Strong knowledge of RDMA transport behavior and failure modes.
  • Practical experience implementing and tuning PFC and ECN for congestion management.
  • Understanding of GPU collective communication patterns such as all-reduce, all-gather, broadcast, reduce-scatter, and their impact on east-west network traffic.
  • Experience designing rail-optimized GPU networking fabrics for distributed training and inference clusters.
  • Familiarity with diagnosing performance issues related to:

○ NCCL stalls

○ RDMA congestion

○ Fabric hotspots

○ Packet loss impacting distributed training

  • Understanding of how networking performance affects distributed AI frameworks such as PyTorch and TensorFlow.

The candidate should also be able to collaborate closely with compute platform teams to ensure that networking infrastructure is optimized for distributed training, distributed inference, andhigh-throughput AI workloads.

Systems & Troubleshooting

  • Ability to debug complex cross-layer issues spanning:

○ Hardware

Similar jobs

SUBMERRemote

Est. 80,000 EUR

About Radian Arc.Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. Our teams…

Full-timeRemote
Gore Mutual InsuranceCambridge, Ontario, Canada

At Gore Mutual, we’ve always set ourselves apart as a modern mutual that does good. Now, we’re proudly building on that legacy to transform our company—and our industry—for the better. Effective January 1, 2026, Gore has…

Full-time
Astera LabsSan Jose, California, United States

Est. 141,000 USD

Astera Labs (NASDAQ: ALAB) provides rack-scale AI infrastructure through purpose-built connectivity solutions. By collaborating with hyperscalers and ecosystem partners, Astera Labs enables organizations to unlock the fu…

Full-time
Lightning AINew York, New York, United States

Est. 190,000 USD

Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…

Full-time
AviatrixBengaluru, Karnataka, India

Job Title: Senior MTS – Cloud Test Location: Bangalore, India Who We Are: Aviatrix® is pioneering the Cloud Native Security Fabric — the architecture the Containment Era requires. The Cloud Native Security Fabric governs…

Full-time
NscaleUnited States

Est. 125,000 USD

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time
Ultragenyx PharmaceuticalBedford, Massachusetts, United States

Est. 140,000 USD

Why Join Us? Be a hero for our rare disease patients At Ultragenyx, we fundamentally believe that taking real impactful action to care for the needs of patients and our people is always the right thing to do. To achieve…

Full-time
Squarepoint CapitalNew York, New York, United States

Est. 120,000 USD

Position Overview: Squarepoint is looking for a highly skilled and detail-oriented Network Core Specialist to architect, develop, optimize and secure scalable networks for Datacenter, Campus and Cloud infrastructures. Th…

Full-time
Myriad360Remote

Est. 140,000 USD

Location: Remote (U.S.) Who You Are You are a Senior Network and Security architect who will build and operate the Networking & Security managed service in collaboration with the Networking & Security Practice Ma…

Full-timeRemote
Network Engineer1 month ago
Firmus TechnologiesSingapore, Singapore

Firmus Technologies Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific. Founded in Australia in 2019, our mission is to create the most effi…

Full-time
TelnyxRemote

About Telnyx Telnyx is an industry leader that's not just imagining the future of global connectivity—we're building it. From architecting and amplifying the reach of a private, global, multi-cloud IP network, to bringin…

Full-timeRemote
NscaleHouston, Texas, United States

Est. 175,000 USD

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time
SUBMERRemote

Est. 120,000 EUR

About Radian Arc.Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. Our teams…

Full-timeRemote
NebiusRemote

Est. 175,000 USD

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to…

Full-timeRemote
Network Specialist3 months ago
Aylo CareersNicosia, Cyprus

Est. 55,000 EUR

Established in 2004, we are a tech pioneer offering world-class adult entertainment and games on some of the internet’s safest and most popular platforms. With the support of an international team of dynamic and collabor…

Full-time
Lightning AIRemote

Est. 170,000 USD

Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…

Full-timeRemote
NscaleHouston, Texas, United States

Est. 200,000 USD

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time
NscaleSeattle, Washington, United States

Est. 300,000 USD

Principal Network Engineer, AI Infrastructure & High-Performance Networking About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and larg…

Full-time
NscaleUnited States

Est. 180,000 USD

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time
NscaleUnited States

Est. 180,000 USD

Senior Network Engineer About the Role We are seeking a Senior Network Engineer to lead the design, implementation, and operation of high-performance network infrastructure across large-scale data center and cloud enviro…

Full-time
Firmus TechnologiesSingapore, Singapore

Firmus Technologies Firmus Technologies is a global leader pioneering the solution to AI’s energy challenge, founded in Australia in 2019 by a visionary team of entrepreneurs. Our mission is to create the most energy-eff…

Full-time
Betsson GroupStockholm, Sweden

Est. 840,000 SEK

About the role As a Senior Network Engineer within the Connectivity area at Betsson Group, you will work hands-on with production network infrastructure that underpins our customer-facing platforms. You will operate and…

Full-time
Betsson GroupValletta, Malta

Est. 55,000 EUR

About the role As a Senior Network Engineer within the Connectivity area at Betsson Group, you will work hands-on with production network infrastructure that underpins our customer-facing platforms. You will operate and…

Full-time
Network Engineer1 month ago
TrexonRemote

Est. 120,000 USD

Trexon's Network Engineer is a critical part of our IT infrastructure team. This position's primary responsibilities are the design, implementation, and maintenance of our global network infrastructure. This will leverag…

Full-timeRemote
Network Specialist3 months ago
Aylo CareersMontréal, Quebec, Canada

Established in 2004, we are a tech pioneer offering world-class adult entertainment and games on some of the internet’s safest and most popular platforms. With the support of an international team of dynamic and collabor…

Full-time
Astera LabsSan Jose, California, United States

Est. 165,000 USD

Astera Labs (NASDAQ: ALAB) provides rack-scale AI infrastructure through purpose-built connectivity solutions. By collaborating with hyperscalers and ecosystem partners, Astera Labs enables organizations to unlock the fu…

Full-time
LendingTreeCharlotte, North Carolina, United States

Est. 125,000 USD

**PLEASE NOTE: This is a hybrid role requiring in-office presence 3 days per week (Monday-Wednesday or Tuesday-Thursday) with 2 days remote. The candidate must be in or near Charlotte, NC or have plans to relocate. This…

Full-time
SB EnergySan Diego, California, United States

Est. 125,000 USD

Do you want to work with high-caliber professionals in a dynamic and growing company? Are you entrepreneurial, hard-working, and collegial? Join us at SB Energy, a leading infrastructure company backed by SoftBank Group…

Full-time
Coupang InternalBangalore, India

Developer Platform – Gateway Services – Job Description Staff Engineer – Gateway Services Level: L7-1 Location: Bangalore p

Full-time
CoupangIndia

Senior Staff Engineer – Gateway Services The Gateway Services team is responsible for building Coupang’s traffic infrastructure layer using Service Mesh. Key aspects of the infrastructure layer include service discovery,…

Full-time