Member of Technical Staff — Training
RadixArkPalo Alto, California, United States · Posted 6 months agoDescription
About the Role
In This Role, You Will
- Design, build, and operate distributed training, rollout, and orchestration systems for large-scale LLM and multimodal post-training across multi-GPU, multi-node environments.
- Profile and optimize performance across the full-stack — model implementation, parallelism strategies, communication libraries, and GPU kernels — to improve throughput, latency, memory efficiency, hardware utilization, and cost.
- Investigate numerical correctness and low-precision issues in distributed training and inference, including train–inference consistency for reinforcement learning.
- Improve the reliability of long-running workloads through checkpointing, fault recovery, observability, and operational tooling.
- Build supporting infrastructure for reinforcement learning and agentic post-training, including asynchronous rollout, trajectory collection, sandboxed execution, evaluation harnesses, and data pipelines.
- Contribute to open-source training and inference systems, including Miles and SGLang, and partner with researchers to turn experimental requirements into production systems.
Minimum Qualifications
- 3+ years of experience building or operating distributed machine learning systems, large-scale training infrastructure, or high-performance inference systems.
- Hands-on experience with post-training systems, training backends, or inference systems for large language models (e.g., Megatron-LM, FSDP, SGLang, TensorRT-LLM, vLLM).
- Experience in at least two of the following areas:
- Performance, efficiency, and scalability of multi-GPU, multi-node workloads
- Numerical correctness or low precision
- Stability, reliability, or fault tolerance
- Post-training algorithm recipes and orchestration infrastructure for large training runs
- Multimodal training or inference, including vision-language models and multimodal generation
- Agent infrastructure, including sandboxes, harnesses, and eval systems
- Building and maintaining open-source projects widely adopted in industry and academia
Preferred Qualifications
- Familiarity with RL algorithms such as PPO, GRPO, and their variants, and experience applying them in large-scale post-training.
- Experience with modern post-training frameworks (e.g., Miles, slime, AReaL, verl, Prime-RL).
- Key open-source contributions to training or inference frameworks (e.g., SGLang, vLLM, Megatron-LM).
- GPU kernel development (e.g., CUDA, Triton, CUTLASS) or communication-layer optimization (e.g., NCCL, RDMA, NVLink/NVSwitch).
- Experience training or serving models at very large scale (e.g., Mixture-of-Experts models on clusters of thousands of GPUs).
- Top-tier publications in ML systems or other systems fields.
Even if you don't meet every qualification above, we encourage you to apply — we care most about demonstrated ability to build and reason about large-scale systems.
About RadixArk
Compensation
Equal Opportunity
RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
Similar jobs
Est. 300,000 USD
About the Role RadixArk is seeking a Member of Technical Staff, Developer Technology (DevTech) to make LLM inference and training dramatically faster, cheaper, and more accessible on modern GPU hardware. Our systems sit…
Est. 300,000 USD
About the Role RadixArk is seeking a Member of Technical Staff — Inference-Multimodal & Diffusion to advance the frontier of generative modeling. You will work on cutting-edge diffusion and flow-based models for imag…
Est. 300,000 USD
About the Role RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference. You will work on the core systems that serve frontier models at scale, optimizing performance, la…
Est. 300,000 USD
About the Role RadixArk is looking for a Member of Technical Staff Cluster Infrastructure to architect and scale the core compute platform that powers frontier-level AI training and inference. You will design and operate…
Est. 300,000 USD
About the Role RadixArk is looking for a Member of Technical Staff — Backend/API Platform Engineer to build the API layer, control plane, and platform services that power SGLang and Miles in production. You'll design and…
Est. 300,000 USD
About the Role RadixArk is seeking a Member of Technical Staff — Kernel / Compiler / Communication to push the limits of performance for frontier AI systems. You will work at the lowest layers of the stack — kernels, run…
Est. 300,000 USD
About the Role RadixArk is hiring a Member of Technical Staff — Performance in Palo Alto, CA — someone who can push LLM inference and training systems to the limit across real production workloads. You’ll work on the per…
Est. 165,000 USD
About the Role As a Technical Program Manager at RadixArk, you'll drive the execution of complex, cross-functional programs across our inference and training infrastructure. You'll partner closely with Product Management…
Est. 300,000 USD
About the Role RadixArk is seeking a Developer Advocate to build and engage our technical community around SGLang, Miles, and our open source infrastructure. SGLang already has 30K+ GitHub stars and serves billions of to…
Est. 300,000 USD
About the Role RadixArk is seeking a Member of Technical Staff - Inference-Multi-Hardware to push the limits of performance for frontier AI systems. Most performance engineering assumes a single vendor's stack. This role…
Est. 300,000 USD
About the Role RadixArk is seeking experienced product-focused engineers to join our team in building the developer-facing surfaces of our inference and training infrastructure. As a Member of Technical Staff — Product,…
Est. 140,000 USD
About The Role RadixArk is launching a full-time, paid, 1-year residency program for aspiring AI infrastructure engineers. You'll rotate across inference, training, kernels, compilers, and cluster infrastructure, working…
Est. 144,000 USD
Key Responsibilities Product Strategy & Roadmap Define, prioritize, and drive the product roadmap for inference and training infrastructure. Stay ahead of AI trends, including new model architectures, hardware optimi…
Est. 300,000 USD
About the Role RadixArk is hiring a Member of Technical Staff — CI / Infrastructure to own the infrastructure that keeps SGLang moving. Our CI system runs 300+ GPU tests across NVIDIA, AMD, Intel, and Ascend hardware poo…
Est. 300,000 USD
About the Role RadixArk is looking for a Member of Technical Staff — TPU Systems to build high-performance inference and training systems using JAX, XLA, and Pallas. You'll push model workloads to their limits on TPU har…
Est. 155,000 USD
About the Role RadixArk is seeking a Product Marketing Manager to own how SGLang, Miles, and our open source infrastructure are positioned and perceived across the market. SGLang already has 20K+ GitHub stars and serves…
Est. 150,000 USD
About the Role We're looking for a Head of Business Development to build the BD function at RadixArk from the ground up. The BD team is the institutional memory of this company — maintaining active relationships across e…
Est. 159,000 USD
About the Role We're looking for a hands-on Talent Operations Specialist to build and run the machinery behind talent and people ops as we scale. This isn't a traditional HR generalist role - it's for someone who treats…
Est. 115,000 USD
About the Role RadixArk builds the open-source AI infrastructure behind SGLang and Miles, used by developers and enterprises around the world. We're looking for a visual designer to join our design team and to give our b…
Est. 315,000 USD
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
Est. 190,000 USD
About NewsBreak Founded in 2015, NewsBreak is the Content Intelligence platform shaping the future content economy. With over 40 million monthly active users, our flagship platform delivers highly personalized local news…
Est. 326,700 USD
Scale's LLM post-training platform team builds our internal distributed framework for large language model training. The platform powers MLEs, researchers, data scientists, and operators for fast and automatic training a…
Est. 315,000 USD
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
Est. 237,500 USD
Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…
Est. 289,800 USD
XPENG is a leading smart technology company at the forefront of innovation, integrating advanced AI and autonomous driving technologies into its vehicles, including electric vehicles (EVs), electric vertical take-off and…
Est. 297,900 USD
AI is becoming vitally important in every function of our society. At Scale, our mission is to accelerate the development of AI applications. For 9 years, Scale has been the leading AI data foundry, helping fuel the most…
Est. 235,200 USD
XPENG is a leading smart technology company at the forefront of innovation, integrating advanced AI and autonomous driving technologies into its vehicles, including electric vehicles (EVs), electric vertical take-off and…
Est. 310,000 USD
SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organi…
Est. 239,000 USD
Your Impact at LILA The AI Research team is tackling one of the most exciting, open problems in AI: training LLMs to run long-horizon scientific discovery tasks. Our approach spans the full post-training stack - from SFT…
Est. 328,650 USD
XPENG is a leading smart technology company at the forefront of innovation, integrating advanced AI and autonomous driving technologies into its vehicles, including electric vehicles (EVs), electric vertical take-off and…