Description
Firmus Technologies
Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure from model to grid. Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.
At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. This co-designed approach from model to grid allows us to make every watt count and deliver low-cost AI tokens globally.
Firmus AI Cloud
Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale. It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.
Why you’ll love working here
At Firmus, you’ll work at the intersection of sustainability and artificial intelligence in a fast-paced environment powered by next-generation technology. You’ll be helping to transform an entire industry — and you’ll feel it every day.
Our team is made up of true innovators and leaders in their fields, and as an emerging company, you won’t be lost in a crowd. You’ll work closely with the founders, build a strong network, and see the impact of your work first-hand as we democratise AI tools for everyone — more sustainably and more affordably.
We believe great things happen when people from diverse backgrounds come together to do their best work and be their authentic selves. We are proud to be an equal opportunity employer.
ROLE
Firmus Technologies is seeking a Senior Platform Engineer to join our Engineering and Technology team. You will drive the design and implementation of our MLOps capability. You will also collaborate with other engineers and make technical decisions on scaling Firmus AI factory platform engineering capabilities to planet scale, from IaC, container orchestration, observability, self-service portal to platform security. This role is ideal for a self-starter with passion for building things from first principles. You naturally break down complex problems into their fundamental truths to uncover novel and elegant solutions - rather than relying on conventional patterns.
KEY RESPONSIBILITIES
- Build MLOps capabilities from the ground up, enabling reproducible, scalable, and secure ML workflows across internal and customer-facing environments.
- Continuously improve our DevOps platform to ensure reliability, scalability, security, and seamless integration with CI/CD pipelines and infrastructure services.
- Design, implement, operate and secure Kubernetes-based production infrastructure for high reliability, performance and security, including clusters supporting NVIDIA GB300 NVL72 systems with NVIDIA Quantum-X800 InfiniBand or Spectrum-X Ethernet.
- Develop world-class observability platforms for internal and external customers
- Integrate Firmus central services with NVIDIA’s software stack, including Mission Control, NETQ, UFM, and NMX.
- Lead the enhancement and evangelism of internal platform products that provide cohesive, composable, secure-by-default, and low-friction self-service experiences that accelerates time to market and reduce engineers' cognitive load.
- Drive incident response efforts, participate actively in the on-call rotation, and lead detailed Root Cause Analysis (RCA) to continuously improve system reliability, operational maturity, and incident handling processes.
SKILLS AND EXPERIENCE
- Bachelor's degree in computer science or a related technical field.
- 7+ years of experience as Platform Engineer, Site Reliability Engineer, DevOps engineer, MLOps Engineer or Observability Engineer.
- Demonstrated strong proficiency on the following areas:
- Infrastructure-as-Code, configuration management and CI/CD (e.g., Terraform, Ansible, GitHub Actions, Jenkins, ArgoCD).
- Containerization technologies (e.g., Docker), Kubernetes networking and cluster management, including upgrades and troubleshooting.
- Observability stack design and scaling (e.g., Loki, Grafana, Tempo, Prometheus, Thanos, ClickHouse).
- Telemetry solutions using various technology (e.g., Redfish, gNMI, SNMP, eBPF, streaming analytics).
- Unified telemetry collection with OpenTelemetry.
- Compliance automation (e.g., OPA, Kyverno).
- Competent in scripting and programming skills (e.g., Bash, Python, Go).
- Systems knowledge on Linux internals, networking stacks, and distributed storage.
- Clear and effective English communication, written and spoken.
- Bonus: Experience in high-growth startups or regulated industries with robust security and data privacy requirements, including SOC 2 Type 2 and ISO 27001.
Similar jobs
Firmus Technologies Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific. Founded in Australia in 2019, our mission is to create the most effi…
Firmus Technologies Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific. Founded in Australia in 2019, our mission is to create the most effi…
Est. 165,000 USD
Firmus Technologies Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure from model to grid. Founded in Australia in 2019, our mission is to create the most effic…
Est. 165,000 USD
Firmus Technologies Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific. Founded in Australia in 2019, our mission is to create the most effi…
Firmus Technology Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific. Founded in Australia in 2019, our mission is to create the most effici…
Firmus Technologies Firmus Technologies is a global leader pioneering the solution to AI’s energy challenge, founded in Australia in 2019 by a visionary team of entrepreneurs. Our mission is to create the most energy-eff…
Est. 156,000 USD
Sustainable Talent is partnering with Nvidia a global leader who's been transforming computer graphics, PC gaming, and accelerated computing for over 25 years. We are looking for a Site Reliability Engineer to support ou…
Role Summary Firmus is seeking a highly skilled and driven Senior Engineer to play a key role in designing, building, and operating software-defined infrastructure, including high-performance AI storage platforms. You wi…
Est. 250,000 USD
Firmus Technologies Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure from model to grid. Founded in Australia in 2019, our mission is to create the most effic…
Est. 650,000 DKK
Who We Are VML is a leading creative company that combines brand experience, customer experience, and commerce, creating connected brands to drive growth. VML is celebrated for its innovative and award-winning human-firs…
Role Summary Every AI feature we ship touches thousands of GPUs. The Senior DevOps Engineer will build the release engineering backbone—CI/CD pipelines, automated testing gates, one-click deployments with instant rollbac…
Est. 175,000 USD
About SHEIN SHEIN is a global online fashion and lifestyle retailer, offering SHEIN branded apparel and products from a global network of vendors, all at affordable prices. Headquartered in Singapore, with more than 15,0…
Principal Software Engineer Looking for engineering problems that actually matter? Join Sand's Decision Intelligence Platform team and help build the semantic engines, data platforms, and AI-powered systems that enable g…
About Nintex: At Nintex, we are transforming the way people work, everywhere. As the global standard for process intelligence and automation, we're trusted by over 10,000 public and private sector organizations across 90…
Est. 144,000 USD
About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…
Firmus Technologies Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific. Founded in Australia in 2019, our mission is to create the most effi…
About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…
Join us in bringing joy to customer experience. Five9 is a leading provider of cloud contact center software, bringing the power of cloud innovation to customers worldwide. Living our values everyday results in our team-…
Firmus Technologies Firmus Technologies is a global leader pioneering the solution to AI’s energy challenge, founded in Australia in 2019 by a visionary team of entrepreneurs. Our mission is to create the most energy-eff…
Est. 90,000 EUR
About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…
Est. 120,000 EUR
About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…
About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…
About NscaleNscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior r…
Role Summary The AI Engineer will establish Firmus AI Factory as the foundation for efficient, production-grade distributed training by delivering pre-built training recipes (TorchTitan, Megatron etc.), evaluation benchm…
Est. 175,000 USD
Vectra® is the leader in AI-driven threat detection and response for hybrid and multi-cloud enterprises. The Vectra AI Platform delivers integrated signal across public cloud, SaaS, identity, and data center networks in…
Emergent builds autonomous coding agents that replace traditional software development by generating, testing, and deploying production applications directly from plain-language intent. Our systems run in production at g…
Est. 140,000 USD
Company Overview Allen Control Systems (ACS) is a cutting-edge defense startup founded by two former Navy electrical engineers with a proven track record in robotics and software. We are developing a small, autonomous gu…
Job Overview We are seeking a self-driven Principal Site Reliability Engineer with a strong technical background and excellent communication skills. This individual will lead the development, construction, and management…
Est. 144,000 USD
Emergent builds autonomous coding agents that replace traditional software development by generating, testing, and deploying production applications directly from plain-language intent. Our systems run in production at g…
Est. 80,000 EUR
About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…