Description
About the Role
Abnormal AI is looking for an experienced and driven Platform & Infra software engineer to join the PI team. Join us and help build the platforms that power Abnormal's growth
- Observability Platform - Own and evolve the monitoring, metrics, and alerting infrastructure that every engineering team at Abnormal depends on. You'll work across the Prometheus, Chronosphere, and Grafana stack to ensure engineers can see what their systems are doing in real time — building dashboards, managing metric pipelines at scale, operating the PagerDuty alerting pipeline, and driving cost-efficient observability across all production environments (US, EU, and GovCloud).
Your Impact
- Own the observability stack (Prometheus, Chronosphere, Grafana, PagerDuty) that every team relies on to detect, diagnose, and resolve production issues — when you make it better, every engineer at Abnormal gets faster.
- Design platforms and developer tooling that remove friction — reducing deployment times, simplifying pipeline authoring, and letting product teams focus on building rather than firefighting.
- Drive SLAs and SLOs for critical shared infrastructure ensuring the systems behind our products are resilient and cost-efficient.
- Your architectural decisions on alerting pipelines and cross-environment deployments will define what products we can build and how quickly we deliver them to customers.
What you will do
- Work with the Tech Lead, Engineering Manager, and Product Manager to design, develop, and deliver key platform features — from technical design docs through production rollout
- Own features end-to-end: scoping, implementation, testing, deployment, and post-launch monitoring across multiple environments (US, EU, GovCloud)
- Take ownership of 1-3 key services within Observability (Prometheus, Chronosphere, Grafana, PagerDuty pipeline) or Data Infra (Airflow, Spark) and be accountable for their reliability, performance, and evolution
- Participate in on-call rotations — triage, diagnose, and resolve production issues independently, building deep operational knowledge of the systems you own
- Improve system resilience by converting runbooks into automated solutions, refining SLAs/SLOs, and proactively identifying performance bottlenecks and failure modes
- Assume ownership of the reliability of everything you build, including comprehensive unit tests, integration testing, and observability instrumentation
- Build platforms, tooling, and APIs that make it easier for other engineering teams to ship — whether that's faster pipeline deployments, better dashboards, or simpler alerting configuration
- Partner with internal customers (product and engineering teams) to understand their needs and translate them into scalable platform capabilities
- Communicate effectively in an async-first, distributed environment — proactively providing updates, discussing challenges, and proposing solutions without prompting
- Mentor junior engineers on the team, helping them ramp up on service operations and development practices
- Raise the bar of engineering excellence through code reviews, knowledge sharing, design discussions, and contributing to team best practices
Must Haves
- Backend Engineering & Distributed Systems (4+ years)
- 4+ years of hands-on backend engineering experience designing, building, and operating production-grade distributed systems
- Strong proficiency in Python — the primary language for Airflow DAGs, platform services, and automation tooling
- Working proficiency in Golang — used for high-performance infrastructure components, metric pipelines, and platform services
- Experience building systems that process data at scale — whether metric ingestion pipelines, stream/batch processing, or high-throughput API services
- Demonstrated experience owning a service or platform end-to-end — from technical design through production deployment, monitoring, and iteration
- Comfortable balancing feature development with operational responsibilities: you've shipped features and kept them running reliably at scale
- Experience writing technical design documents that articulate trade-offs, propose solutions, and get buy-in from peers and tech leads
- Track record of breaking down ambiguous problems into concrete, deliverable milestones
- Experience with fault tolerance patterns — retries, circuit breakers, graceful degradation, backpressure — and knowing when to apply each
- Proven incident response capability: you've been on-call, diagnosed production issues under pressure, and driven them to resolution
- Strong testing discipline — unit tests, integration tests, and an understanding of what to test and how to keep test suites maintainable
- Ability to design systems with a forward-looking perspective — thinking about how your architecture handles 10x growth, multi-region deployment, and evolving requirements
- Ability to contribute to and influence cross-team technical direction — you're not just implementing specs, you're shaping the solution
- Async-first communication excellence — strong written communication skills for design docs, Slack discussions, PR reviews, and status updates across time zones
- Proactive communicator — you surface blockers early, share con
- Solid understanding of monitoring, alerting, and observability principles — you've instrumented services, set up dashboards, defined SLIs/SLOs, or triaged production incidents using metrics and logs
Nice to Have
- Hands-on experience with Prometheus — PromQL queries, recording rules, alerting rules, relabeling configs, and understanding metric cardinality challenges at scale
- Experience with Grafana — building dashboards, templating, managing datasources, and creating meaningful visualizations for operational and business metrics
- Familiarity with commercial observability platforms like Chronosphere, Datadog, New Relic, or Honeycomb — understanding trade-offs between self-hosted and managed solutions
- Experience designing or operating an alerting pipeline — PagerDuty, OpsGenie, or similar — including alert routing, escalation policies, and reducing noise/alert fatigue
Cloud Infrastructure & Kubernetes
- Familiarity with AWS services — EC2, ECS, EKS, S3, RDS, IAM, CloudWatch, Lambda, SQS/SNS — and understanding how to architect cost-effective, secure cloud infrastructure
- Experience with Kubernetes (K8s) — deploying and operating workloads, understanding pods/services/deployments, Helm charts, and debugging cluster-level issues
- Exposure to Infrastructure-as-Code tools — Terraform, Pulumi, or CloudFormation — and understanding the value of declarative infrastructure management
- Experience with CI/CD pipelines — GitHub Actions, Jenkins, or similar — and optimizing build/deploy times for platform services
Programming & Framework
- Experience with Django or similar Python web frameworks — building APIs, managing migrations, and understanding ORM performance characteristics
- Familiarity with gRPC or protobuf for inter-service communication in a microservices architecture
Technical Leadership & Platform Thinking
- Experience leading a small team (2-4 engineers) to build a feature or component from scratch — scoping, task breakdown, code reviews, and delivery management
- Experience building internal developer platforms or tooling — CLIs, SDKs, self-service portals, or automation that improved developer productivity
- Track record of reducing operational toil — automating runbooks
#LI-FS1
A note on AI in our process:
Abnormal AI uses AI-assisted tools to help our recruiting team prepare for candidate interviews. These tools analyze resume content and role requirements to suggest interview questions and areas for the interviewer to explore.They do not make hiring decisions or screen candidates automatically. Every decision about a candidacy is made by a person. Further, if your application is successful and Abnormal AI makes a conditional offer of employment, we will carry out pre-employment checks which must be successfully completed to progress to a final offer. All processes and pre-employment checks are in line with prevailing legislation and Abnormal AI's policies relevant to our security and privacy standards.
Abnormal AI is an equal opportunity employer. Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability, protected veteran status or other characteristics protected by law. For our EEO policy statement please click here. If you would like more information on your EEO rights under the law, please click here.
Similar jobs
The Opportunity We are looking for a Software Engineer II, Backend to join our Federated Intelligence Platform team in India. The team serves as the federated intelligence layer within the Abnormal Data Platform, managin…
Est. 186,500 USD
About the Role Abnormal AI is looking for a Senior Cloud Security Engineer to help build the next generation of security capabilities at scale. This is a senior IC-level role that blends deep security expertise with the…
Est. 144,550 USD
About the Role The Enterprise Platform Engineer helps build and run secure, reliable, and increasingly AI-enabled workforce technology across identity, endpoints, SaaS, collaboration, network, and office IT. You will tur…
Est. 181,850 USD
About the Team Abnormal AI is looking for a Software Engineer II to join the Dev Accelerator Team. The Dev Accelerator team builds and operates our internal developer platform: the tools, libraries, and infrastructure th…
About the Role At Abnormal AI, we are on a thrilling mission to safeguard the world's largest enterprises against a vast range of relentless email and collaboration application cyber security attacks. Our relentless purs…
About the Role At Abnormal AI, we are on a thrilling mission to safeguard the world's largest enterprises against a vast range of relentless email and collaboration application cyber security attacks. Our relentless purs…
Est. 135,200 USD
About the Role Abnormal AI is hiring a Technical Program Manager to lead delivery of high-impact programs across our AI Platform and Data Systems teams. You'll own strategic initiatives from CIO leadership — spanning AI…
About the Team Abnormal AI is looking for a Software Engineer II to join the Dev Accelerator Team. The Dev Accelerator team builds and operates our internal developer platform: the tools, libraries, and infrastructure th…
About the Role Abnormal AI is looking for a Staff Software Engineer to serve as the technical lead for the APAC leg of our Multi-Product Platform division. Our team owns the full customer journey — from the moment a pros…
Est. 181,850 USD
About the Role The AI Transformation Office at Abnormal owns how the company redesigns work around AI: mapping every business function for AI opportunity, running structured workstreams across Sales, CS, Support, Marketi…
Est. 181,850 USD
About the Role Abnormal AI is looking for a Software Engineer II to join the Detection Team. The Detection Division is focused on building the world’s most advanced technology to identify and stop email and cloud-based a…
Est. 217,400 USD
About the Role Abnormal AI is an AI-native cybersecurity company. We use behavioral AI to stop the attacks that get past everything else, and our products depend on processing data responsibly, at scale, across a growing…
Est. 186,500 USD
About the Role Abnormal AI is hiring a Senior Cloud Security Engineer to build and operate the security systems that protect our FedRAMP authorized environment. You will own secure delivery pipelines, infrastructure as c…
Job Description Abnormal AI is revolutionising cybersecurity by leveraging AI-native technologies to combat modern cyber threats. Behind every fast-moving go-to-market motion is a platform that makes it possible — and th…
Est. 120,000 GBP
At Abnormal AI, we protect our customers against nefarious adversaries who are constantly evolving their techniques and tactics to outwit and undermine the traditional approaches to Security. Abnormal is recognized as a…
Software Engineer II, Full Stack Location: India — Hybrid About Abnormal AI Abnormal AI is redefining cybersecurity with an AI-native platform that protects modern enterprises from sophisticated email, collaboration, and…
About Abnormal AI Abnormal AI is the leading AI-native security platform, protecting enterprise organizations from the full spectrum of cyber attacks. Powered by behavioral AI that understands human and machine identity,…
The Mission We are building the high-throughput, low-latency infrastructure that powers our real-time threat detection capabilities. This isn't just about maintaining stable backend systems; it’s about architecting a cut…
Est. 181,850 USD
About the Role At Abnormal AI, our mission is to protect the world’s largest enterprises from advanced email and collaboration attacks. The Threat Narrative team transforms complex signals from our detection systems into…
Est. 256,450 USD
About the Role Abnormal AI is seeking a Staff Software Engineer to join the Multi-Product Platform division, owning Security & Privacy as a platform capability. We're looking for a distributed systems expert who has…
About the Role Abnormal AI is looking for a Senior Software Engineer to join the Multi-Product Platform (MPP) organization, which owns the product platform enabling scalable, consistent product development across our ent…
Est. 158,550 USD
About the Role Abnormal AI is looking for an Application Security Engineer II to secure the AI-powered systems at the core of our AWS-based platform (LLM-integrated features, agentic workflows, MCP connectors, and the mo…
Est. 110,075 USD
About the Role Abnormal AI is seeking an Mid-Market Sales Engineer to join our growing Sales Engineering team. As an Mid-Market Sales Engineer, you will serve as our customer’s primary technical contact, acting as a trus…
Est. 130,550 USD
The Critical Situation Manager owns the customer-facing response when a customer's trust in the platform has been shaken by a significant efficacy or product concern. This role engages when a customer's feedback is subst…
Est. 219,150 USD
Senior Software Engineer - Identity Security Enterprises are undergoing two simultaneous transformations: moving critical infrastructure to the cloud and adopting AI agents at scale. Both create an identity explosion—hum…
Est. 110,075 USD
About the Role Abnormal AI is seeking an Mid-Market Sales Engineer to join our growing Sales Engineering team. As an Mid-Market Sales Engineer, you will serve as our customer’s primary technical contact, acting as a trus…
Est. 60,000 EUR
About the Role Abnormal AI is looking for a Professional Services Consultant to join the Self-Serve Professional Services team. The Self-Serve team owns the post-sales technical deployment for our smaller, lower-touch ac…
Est. 55,000 GBP
About the Role Abnormal AI is looking for a Professional Services Consultant to join the Self-Serve Professional Services team. The Self-Serve team owns the post-sales technical deployment for our smaller, lower-touch ac…
Senior Software Engineer - Identity Security Posture Management (ISPM) With identity now serving as the main attack vector for modern enterprises, having complete visibility to understand every identity in your environme…
Est. 145,705 USD
About the Role Abnormal AI is seeking an Enterprise Sales Engineer to join our growing Sales Engineering team. As an Enterprise Sales Engineer, you will serve as our customer’s primary technical contact, acting as a trus…