Senior Observability Platform Engineer
United States · Posted 6 days agoDescription
Senior Observability Platform Engineer – Nscale
About Nscale
Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale simplifies AI development while enabling superior results, supporting strategic business outcomes such as cost management, rapid innovation, and environmental responsibility.
We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency while contributing to the technology that powers the future.
About the Role
As a Senior Observability Platform Engineer, you’ll play a key role in designing, building, and scaling Nscale’s observability platform. You’ll focus on delivering reliable, high-quality visibility into GPU clusters, AI workloads, and the infrastructure that powers them.
You approach observability as a product—balancing usability, scalability, and operational efficiency. You build systems that reduce cognitive load for engineers, surface meaningful signals, and enable fast, confident debugging when things go wrong.
You’ll contribute to platform direction, implement critical systems, and collaborate closely with SRE, infrastructure, and AI/ML teams to ensure observability is embedded into everything we run.
This is a hands-on engineering role with meaningful influence over platform design and evolution.
What You’ll Do
- Design, build, and operate scalable observability systems across metrics, logs, traces, and alerting
- Contribute to architectural decisions around tooling, data pipelines, storage, and retention strategies
- Improve signal quality by reducing noise, managing cardinality, and refining alerting practices
- Help identify and address observability gaps before they impact reliability
- Partner with SRE, infrastructure, and AI/ML teams to integrate observability into services and platforms
- Develop reusable patterns, libraries, and best practices that improve consistency across teams
- Participate in incident response and postmortems, driving actionable improvements
- Evaluate and adopt tools that improve developer experience, scalability, and operational efficiency
- Support and mentor engineers within the team through code reviews and knowledge sharing
About You
- 5+ years in SRE, infrastructure engineering, platform engineering, or observability-focused roles
- Experience operating and scaling observability systems in production environments
- Strong understanding of monitoring concepts: metrics, logs, traces, alerting, and SLOs
- Hands-on experience with several of: Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic
- Solid programming skills (Python, Go, or similar) with the ability to build and maintain production systems
- Experience working with Kubernetes-based infrastructure
- Familiarity with Infrastructure-as-Code (Terraform, Ansible, or similar)
- Pragmatic mindset with a focus on simplicity, reliability, and maintainability
- Strong collaboration skills and ability to work across teams
Preferred
- Experience with observability data pipelines (Kafka, Vector, Fluent Bit, etc.)
- Exposure to AI/ML infrastructure or GPU-based systems
- Familiarity with performance monitoring for distributed systems
- Experience improving developer experience through observability tooling
Equal Opportunities Statement
We strongly encourage applications from people of color, the LGBTQ+ community, people with disabilities, neurodivergent individuals, parents, carers, and people from lower socio-economic backgrounds.
If there’s anything we can do to accommodate your specific situation, please let us know.
Note: Responsibilities outlined are not exhaustive and may evolve as business needs change
The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.
Similar jobs
Est. 225,000 USD
About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale simplifies AI development while enabling superior…
Est. 245,000 USD
Principal Observability Platform Engineer – Nscale About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscal…
Est. 240,000 USD
About Nscale Nscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform. We own the data centres, software, and applications that power today's AI stack using sustainable technology sol…
Est. 120,000 USD
About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack — energy, data centres, GPU superclusters, orchestration, and AI services — delivering high-performance infra…
Est. 145,000 USD
About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack — energy, data centres, GPU superclusters, orchestration, and AI services — delivering high-performance infra…
Est. 145,000 USD
About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…
Est. 250,000 USD
About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…
Est. 75,000 USD
About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…
Est. 182,500 USD
. About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superio…
Est. 242,500 USD
About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…
Est. 60,000 USD
About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…
Est. 250,000 USD
About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…
Est. 217,500 USD
About NscaleNscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-nativestartups and global enterprises, from bare metal up through the platform services teams actually build…
Est. 250,000 USD
About Nscale Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack — energy, data centres, GPU superclusters, orchestration, and AI services — delivering high-performance infra…
Est. 140,000 USD
About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…
Est. 120,000 EUR
About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…
About NscaleNscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior r…
Est. 225,000 USD
About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…
About NscaleNscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior r…
Est. 140,000 USD
. About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superio…
Est. 150,000 USD
About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…
Est. 140,000 USD
About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…
Est. 180,000 USD
About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…
Est. 160,000 USD
ROLE OVERVIEW Nscale is seeking a Development Program Manager to be the organizational backbone of our Americas data center development team — building and maintaining the processes, folder structures, dashboards, and re…
Please note: This role requires you to be working on-site in the Datacenter About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large en…
Est. 145,000 USD
About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…
Est. 215,000 USD
About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…
Est. 650,000 NOK
About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…
Est. 300,000 USD
Principal Network Engineer, AI Infrastructure & High-Performance Networking About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and larg…
Est. 195,000 USD
About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior…