Description
About 10a Labs: 10a Labs is the safety and threat-intelligence layer trusted by frontier AI labs, AI unicorns, Fortune 10 companies, and leading global technology platforms. Our adversarial red teaming, model evaluations, and intelligence collection enable engineering, safety, and security teams to stay ahead of evolving threats and deploy AI systems safely.
Software Engineer, Infrastructure & Platform
About the Role
We are seeking a Software Engineer, Infrastructure & Platform to build the systems and infrastructure that power advanced AI evaluations, including evaluations focused on autonomous model behavior, agentic systems, and loss-of-control risks.
This is a hands-on engineering role at the intersection of backend systems, infrastructure, and AI. You will build secure and reproducible environments where frontier models can interact with tools, execute code, complete complex tasks, and operate across realistic multi-step workflows, including machine learning research and engineering.
The ideal candidate has strong backend and infrastructure fundamentals with attention to security, enjoys debugging complex distributed systems, and is excited to apply those skills to difficult problems in AI safety and evaluation.
What You'll Do
- Design and build sandboxed evaluation environments where AI models can safely execute code, use tools, interact with services, and complete complex tasks.
- Build backend services and infrastructure supporting large-scale, repeatable AI and agentic evaluations.
- Develop agent scaffolding and evaluation harnesses, including tool-use loops, context management, retries, state management, token budgets, and multi-agent or subagent workflows.
- Build systems for provisioning and orchestrating isolated environments using technologies such as Docker, Kubernetes, VMs, and cloud infrastructure.
- Design secure approaches to networking, permissions, secrets, credentials, and resource isolation for model-driven environments.
- Develop APIs, internal tools, and automation that allow analysts, engineers, and subject-matter experts to create and run evaluations efficiently.
- Improve the reliability and reproducibility of evaluations through logging, observability, snapshotting, debugging tools, and automated testing.
- Build systems capable of running thousands of evaluation tasks reliably and capturing the artifacts and telemetry needed to understand model behavior.
- Partner with analysts, red teamers, and domain experts to translate complex evaluation ideas into robust technical systems.
- Investigate failures across the evaluation stack and distinguish between model limitations and infrastructure, harness, or environment failures.
What We're Looking For
- 3–5+ years of professional software engineering experience, particularly in backend, infrastructure, platform, SRE, or distributed systems engineering.
- Strong programming skills in Python and experience building production-quality software.
- Experience designing and operating backend services or distributed systems.
- Hands-on experience with Docker, Kubernetes, virtual machines, or other container/orchestration technologies.
- Experience working with GCP, AWS, or similar cloud infrastructure.
- Strong understanding of Linux systems, networking, authentication, permissions, and infrastructure security.
- Experience with infrastructure-as-code or automation tools such as Terraform.
- Strong debugging skills and comfort diagnosing failures across application, infrastructure, and networking layers, especially in agentic loops.
- Ability to build systems that are reproducible, observable, scalable, and secure.
- Comfort working on ambiguous technical problems where the architecture and requirements may evolve quickly.
- Interest in AI systems, agentic workflows, AI security, or model evaluations. Prior professional AI experience is helpful but not required.
Nice to Have
- Experience building developer platforms, CI/CD systems, test infrastructure, sandboxes, or ephemeral compute environments.
- Experience with agent frameworks, LLM APIs, tool-calling systems, or AI evaluation infrastructure.
- Experience designing secure execution environments for untrusted or semi-trusted code.
- Background in SRE, platform engineering, cloud infrastructure, cybersecurity, or developer tooling.
- Experience with distributed task execution, queues, workflow orchestration, or large-scale automated testing.
- Familiarity with AI safety, adversarial testing, model evaluations, or autonomous-agent systems.
- Familiarity with agentic AI fundamentals, including common harnesses, Model Context Protocol, agent benchmarks, and security risks to AI agents.
Compensation & Benefits
- Salary Range: $110K–$160K, depending on experience and location
- Bonus: Performance-based annual bonus
- Professional Development: Support for conferences, continuing education, or leadership training
- Work Environment: Fully remote, U.S.-based
- Health Benefits: Comprehensive health, dental, and vision coverage
- Time Off: Generous PTO and paid holiday schedule
- Retirement: 401(k) plan
Similar jobs
Est. 147,500 USD
About 10a Labs: 10a Labs is the safety and threat-intelligence layer trusted by frontier AI labs, AI unicorns, Fortune 10 companies, and leading global technology platforms. Our adversarial red teaming, model evaluations…
Est. 165,000 USD
About 10a Labs: 10a Labs is the safety and threat-intelligence layer trusted by frontier AI labs, AI unicorns, Fortune 10 companies, and leading global technology platforms. Our adversarial red teaming, model evaluations…
Est. 110,000 USD
About 10a Labs: 10a Labs is the safety and threat-intelligence layer trusted by frontier AI labs, AI unicorns, Fortune 10 companies, and leading global technology platforms. Our adversarial red teaming, model evaluations…
Est. 60,000 USD
About 10a Labs: 10a Labs is the safety and threat-intelligence layer trusted by frontier AI labs, AI unicorns, Fortune 10 companies, and leading global technology platforms. Our adversarial red teaming, model evaluations…
Est. 127,000 USD
About 10a Labs: 10a Labs is the safety and threat-intelligence layer trusted by frontier AI labs, AI unicorns, Fortune 10 companies, and leading global technology platforms. Our adversarial red teaming, model evaluations…
Est. 120,000 USD
About CodeRoad CodeRoad provides end-to-end software development services, helping businesses scale with ideal infrastructure solutions. From staff augmentation to dedicated IT teams and general software engineering, our…
Est. 80,000 EUR
Our client is a fast-growing European fintech company in the business spend management space. We're looking for a Senior Applied AI Engineer to design, build and ship customer-facing AI features end-to-end. Responsibilit…
About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…
Est. 144,000 USD
Location: United States – RemoteClearance: Ability to obtain and maintain a Public Trust LTS is seeking a highly skilled Agentic AI Security Engineer to ensure our AI systems are secure, trustworthy, resilient, and gover…
Est. 115,000 USD
About 10a Labs: 10a Labs is the safety and threat-intelligence layer trusted by frontier AI labs, AI unicorns, Fortune 10 companies, and leading global technology platforms. Our adversarial red teaming, model evaluations…
Est. 120,000 USD
The Team At Coderoad, we're more than just a software development company—we're your gateway to the global tech world. Whether you're looking to skill up or level up your career, we offer the challenges you’ve been searc…
Est. 205,500 USD
Principal AI Engineering Architect We're looking for a Principal AI Engineering Architect to lead the design and delivery of complex, multi-domain systems spanning cloud, data, and AI — with deep, hands-on mastery of mul…
Est. 120,000 EUR
N-iX is a global software development company founded in 2002, connecting over 2,400+ tech professionals across 40+ countries. We deliver innovative technology solutions in cloud computing, data analytics, AI, embedded s…
Est. 140,000 USD
Location: United States – RemoteClearance: Ability to obtain and maintain a Public Trust LTS is seeking a highly skilled Senior Applied AI Engineer to focus on continuously improving the intelligence behind the platform.…
About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…
Est. 52,000 USD
About 10a Labs: 10a Labs is the safety and threat-intelligence layer trusted by frontier AI labs, AI unicorns, Fortune 10 companies, and leading global technology platforms. Our adversarial red teaming, model evaluations…
We are looking for an LLM / Agentic Evaluation Rig Engineer to build the system that decides whether our AI output is good enough to ship. Because our commentary sits next to externally reported financials, we cannot rel…
Est. 170,000 USD
Location: United States – RemoteClearance: Ability to obtain and maintain a Public Trust LTS is seeking a Senior Agentic AI Software Engineer to build the intelligence behind the platform—the autonomous agents, orchestra…
Est. 120,000 USD
About CodeRoad CodeRoad provides end-to-end software development services, helping businesses scale with ideal infrastructure solutions. From staff augmentation to dedicated IT teams and general software engineering, our…
We're looking for an engineer with hands-on experience building and evaluating GenAI services - from RAG and agentic reasoning systems to production-grade LLM deployments. You'll work closely with Frontend and Backend te…
Est. 120,000 GBP
As a Software Engineer on the AI Infrastructure team, you'll help build and evolve our agent sandboxing platform — the secure, high-performance code execution layer powering our agentic workflows, deployed across both in…
Est. 193,500 USD
Staff Forward Deployed AI Engineer We're looking for a Staff Forward Deployed AI Engineer to lead the design, delivery, and adoption of AI/ML systems directly within client environments. This role is ideal for an experie…
Est. 402,500 USD
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of co…
Est. 202,500 USD
AI Infrastructure Engineer - Agent Sandbox Platform As a Software Engineer on the AI Infrastructure team, you'll help build and evolve our agent sandboxing platform — the secure, high-performance code execution layer pow…
Est. 140,000 USD
Location: Remote (U.S. or Canada) Type: US Applicants – Full-Time; Canadian Applicants – Independent Contractor About Human Agency We're scaling rapidly and have a growing pipeline of opportunities that demand exceptiona…
Est. 675,000 USD
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of co…
Est. 155,000 USD
About You You've built tools that engineers rely on. You think agentic AI and developers working together is one of the most interesting opportunities of this decade, and you want your fingerprints on the tools, standard…
Est. 150,000 USD
Staff Platform Engineer Location: This position is a remote role based in the US Company Overview Robots & Pencils is an applied AI engineering firm building the next frontier of business architecture. We design and…
Est. 140,000 USD
LTS is seeking an AI Platform and Harness Engineer to develop and maintain the infrastructure, tooling, and evaluation frameworks that power enterprise AI solutions. This role is responsible for building the AI platform…
Est. 175,000 USD
Gradial is the marketing operations system of work that helps marketers and creatives move from idea to execution faster. Our platform orchestrates across martech stacks, workflows, and people to automate marketing execu…