Director, AI Platform Reliability

LogicMonitorSan Francisco, CA
$247,500 - $275,000Hybrid

About The Position

LogicMonitor is seeking an accomplished and hands-on Director of AI Platform Reliability to lead the architecture, development, and operation of highly scalable, distributed software platforms. This leader will be responsible for systems that process hundreds of millions/billions of transactions and events, manage terabytes to petabytes of data, and deliver reliable, low-latency services to enterprise customers. The ideal candidate combines strong engineering depth in Java, Kafka, distributed systems, and cloud-native microservices with a demonstrated ability to build and lead high-performing engineering organizations. This is a strategic leadership role, but it requires a leader who can remain close to the technology, participate in architecture reviews, challenge design decisions, guide teams through complex production problems, and establish the engineering practices required to operate mission-critical platforms at scale.

Requirements

  • 10+ years of professional software-engineering experience, including significant experience building large-scale distributed systems.
  • Experience leading engineering teams, architects, and staff engineers.
  • Demonstrated success delivering and operating platforms that process hundreds of millions of transactions, requests, or events.
  • Deep technical expertise in Java, JVM performance, concurrency, multithreading, memory management, and application profiling.
  • Strong experience designing and operating microservice-based and event-driven architectures.
  • Extensive production experience with Apache Kafka or a comparable distributed streaming platform.
  • Strong understanding of Kafka partitioning, replication, consumer groups, offset management, ordering, delivery semantics, schema evolution, and reprocessing.
  • Experience designing low-latency, highly available APIs and backend services.
  • Experience managing terabyte- or petabyte-scale datasets across relational, NoSQL, streaming, and object-storage technologies.
  • Strong understanding of distributed-systems concepts, including consensus, replication, partitioning, consistency models, idempotency, backpressure, and fault tolerance.
  • Experience operating cloud-native applications using Kubernetes, containers, infrastructure as code, and automated CI/CD pipelines.
  • Experience with at least one major cloud platform, such as AWS, Google Cloud, or Microsoft Azure.
  • Strong knowledge of observability practices involving metrics, logs, traces, profiling, alerting, dashboards, and service-level objectives.
  • Demonstrated experience improving system reliability, scalability, latency, cost efficiency, and engineering productivity.
  • Strong written and verbal communication skills, including the ability to explain complex technical decisions to engineering teams, executives, and business stakeholders.
  • Proven ability to build inclusive, accountable, and high-performing engineering organizations.

Responsibilities

  • Lead and scale multiple engineering teams responsible for high-volume, business-critical distributed systems and data platforms.
  • Define the technical strategy and architecture for platforms processing hundreds of millions of transactions and terabytes of data.
  • Guide the development of Java-based microservices, APIs, Kafka streaming pipelines, batch-processing workflows, and cloud-native services.
  • Build and evolve scalable data lake and Data Lakehouse platforms supporting real-time, near-real-time, and batch analytics workloads.
  • Establish reliable data ingestion, transformation, storage, governance, lineage, retention, and data-quality practices across streaming and batch pipelines.
  • Build low-latency, highly available, fault-tolerant systems with strong scalability, resiliency, and disaster-recovery capabilities.
  • Define and own operational SLAs, SLOs, availability targets, recovery objectives, and performance metrics for critical services and data pipelines.
  • Drive capacity planning, load testing, throughput optimization, and improvements to p95 and p99 latency.
  • Ensure effective Kafka design, including partitioning, consumer groups, ordering, schema evolution, replay, and lag management.
  • Establish engineering standards for architecture, coding, testing, security, observability, and production readiness.
  • Partner with Product, Architecture, SRE, Security, Data, and Infrastructure teams to deliver strategic platform initiatives.
  • Strengthen operational excellence through monitoring, incident management, on-call practices, root-cause analysis, and continuous reliability improvements.
  • Recruit, mentor, and develop engineering managers, architects, and senior technical leaders.
  • Improve developer productivity, CI/CD automation, deployment safety, and release predictability.
  • Manage technical debt, platform modernization, cloud costs, and long-term scalability investments.

Benefits

  • Comprehensive health, dental and vision coverage
  • Generous parental leave policies
  • Access to our Employee Assistance Program
  • Various Wellness programs
  • 401K with company matching
  • Lifestyle Spending Account
  • Unlimited vacation policy
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service