inhousefyi
← Back to listings

Infrastructure, Large-scale Training

HarkSan Jose, California, United States · Posted 5 months ago
Full-timeEst. 315,000 USD
Apply now

Description

About Hark

Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.

We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines. While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world.

To get there, we're developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems.

About the Role

We are looking for a Member of Technical Staff, Infrastructure Compute to lead and manage large-scale GPU computing clusters powering our AI training and deployment workloads. You'll work at the intersection of systems engineering and machine learning infrastructure, owning the reliability, scalability, and efficiency of the compute platform that our research and engineering teams depend on. This is a high-impact, highly technical role suited for someone who thrives in complex distributed systems environments and cares deeply about infrastructure as a product.

Responsibilities

  • Design, implement, and maintain Infrastructure as Code (IaC) best practices to enable repeatable, auditable, and scalable cluster provisioning.
  • Enhance and harden CI/CD deployment pipelines to ensure robust, secure, and low-latency model service delivery across production environments.
  • Own and evolve stable training infrastructure operating at the scale of 10,000+ GPUs, including job scheduling, fault tolerance, and network fabric optimization.
  • Partner closely with ML researchers and engineers to understand compute bottlenecks and translate them into infrastructure improvements.
  • Monitor system health, define SLOs, and lead incident response for critical training and inference workloads.
  • Drive capacity planning, cost efficiency initiatives, and hardware lifecycle management across the GPU fleet.
  • Contribute to internal tooling and platform abstractions that improve developer experience for teams consuming compute resources.

Requirements

  • 5+ years of experience in infrastructure, systems, or platform engineering, with at least 2 years working in ML or HPC environments.
  • Demonstrated experience managing GPU clusters or large-scale distributed compute infrastructure.
  • Strong proficiency in at least one systems or infrastructure programming language.
  • Deep understanding of networking fundamentals (RDMA, InfiniBand, or RoCE a plus) relevant to high-throughput training workloads.
  • Experience with container orchestration, job scheduling, and multi-tenant resource management.
  • Proven track record owning production systems with high reliability requirements.
  • Strong debugging and observability skills across the full infrastructure stack.

Bonus Qualifications

  • Kubernetes (K8s) — particularly experience operating large, GPU-aware clusters.
  • Pulumi or similar modern IaC tooling.
  • Rust and/or Go for systems-level tooling and performance-critical services.
  • Familiarity with PyTorch and Ray for understanding workload patterns and integration requirements.

Compensation

The US base salary range for this full-time position is between $180,000 - $450,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components and benefits depending on the specific role. This information will be shared if an employment offer is extended.

Similar jobs

NscaleHouston, Texas, United States

Est. 205,000 USD

About the role: The AI infrastructure market moves weekly. As the Director of Competitive Insights, you will not just track competitors; you will translate complex market movements—such as new Nvidia allocations, legacy…

Full-time
NscaleHouston, Texas, United States

Est. 242,500 USD

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time
NscaleHouston, Texas, United States

Est. 182,500 USD

. About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superio…

Full-time
NscaleHouston, Texas, United States

Est. 165,000 USD

Role Overview As a Technical Program Manager (TPM) for AI Infrastructure Operations, you will be the operational backbone of our high-scale, high-performance AI and High-Performance Computing (HPC) environment. You will…

Full-time
NscaleIceland

Please note: This role requires you to be working on-site in the Datacenter About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large en…

Full-time
NscaleLondon, United Kingdom

Est. 80,000 GBP

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time
NscaleUnited States

Est. 225,000 USD

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale simplifies AI development while enabling superior…

Full-time
NscaleSeattle, Washington, United States

Est. 150,000 USD

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time
NscaleNarvik, Norway

Est. 840,000 NOK

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time
NscaleHouston, Texas, United States

Est. 235,000 USD

About the role: You are the commercial architect of our physical supply chain. Balancing long-lead and massive GPU deployments with explosive customer demand requires flawless yield management. You are a player-coach who…

Full-time
NscaleRemote

Est. 120,000 USD

About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack — energy, data centres, GPU superclusters, orchestration, and AI services — delivering high-performance infra…

Full-timeRemote
NscaleHouston, Texas, United States

Est. 250,000 USD

About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack — energy, data centres, GPU superclusters, orchestration, and AI services — delivering high-performance infra…

Full-time
NscaleRemote

Est. 145,000 USD

About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack — energy, data centres, GPU superclusters, orchestration, and AI services — delivering high-performance infra…

Full-timeRemote
CoreWeaveBellevue, Washington, United States

Est. 254,500 USD

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI…

Full-time
NuroMountain View, California, United States

Est. 242,540 USD

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the wo…

Full-time
NscaleNew York, New York, United States

Est. 250,000 USD

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time
ATOMS Careers pageSan Francisco, California, United States

Est. 254,000 USD

Who we are Atoms is building the machines that power the next era of progress. Over the last decade, software has transformed the digital world. But the physical world, where food is made, minerals are mined, goods are m…

Full-time
CoreWeaveSan Francisco, California, United States

Est. 231,000 USD

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI…

Full-time
NscaleHouston, Texas, United States

Est. 250,000 USD

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time
CoreWeaveSunnyvale, California, United States

Est. 241,000 USD

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI…

Full-time
NscaleHouston, Texas, United States

Est. 140,000 USD

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time
NuroMountain View, California, United States

Est. 200,450 USD

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the wo…

Full-time
NscaleSeattle, Washington, United States

Est. 300,000 USD

Principal Network Engineer, AI Infrastructure & High-Performance Networking About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and larg…

Full-time
NscaleUnited States

Est. 180,000 USD

About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…

Full-time
NscaleHouston, Texas, United States

Est. 256,667 USD

About Nscale Nscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform. We own the data centers, software, and applications that power today’s AI stack using sustainable technology sol…

Full-time
NscaleLondon, United Kingdom

Est. 120,000 GBP

About Nscale Nscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform. We own the data centres, software, and applications that power today's AI stack using sustainable technology sol…

Full-time
CoreWeaveSunnyvale, California, United States

Est. 269,500 USD

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI…

Full-time
NscaleUnited States

Est. 230,000 USD

About Nscale NScale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. NScale enables AI-focused companies to achieve superior…

Full-time
NscaleSingapore, Singapore

About NscaleNscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior r…

Full-time
NscaleAustin, Texas, United States

Est. 215,000 USD

About the role: As the Director of Revenue Enablement, your mandate is to ensure our GTM team is the most technically proficient, strategically lethal, and commercially astute sales force in the AI infrastructure market.…

Full-time