Principal Engineer - Observability

TargetBrooklyn Park, MN
$168,000 - $303,000Hybrid

About The Position

The Observability team builds and operates the enterprise platform that enables engineering teams to understand, measure, and improve the reliability and performance of their products. We provide standardized telemetry, logging, tracing, metrics, alerting, system maps, and operational insights across more than 20,000 services. Our platform enables: Enterprise-scale metrics, logs, and traces built on OpenTelemetry standards; Real-time system maps and dependency graphs across complex service ecosystems; SLO management, error budget tracking, and reliability showback; Integrated alerting, on-call alignment, and actionable operational insights; AI-driven root cause analysis and automated remediation capabilities. We operate at massive scale in 24x7 production environments. Technologies commonly used across the team include Golang, Kubernetes, Kafka, ClickHouse, InfluxDB, Grafana, React, and distributed systems patterns that support high-volume telemetry pipelines. We are moving beyond traditional monitoring: our ambition is an intelligent, agent-enabled observability platform that proactively detects degradation, explains system behavior, and recommends or executes recovery actions before guests are impacted. As a Principal Engineer, you are the senior-most individual-contributor technical leader on the Observability team. You partner closely with three Senior Engineering Managers and product leadership to set the technical vision, architecture, and standards for a unified, intelligent reliability platform serving all of Target Tech. This is a hands-on technical leadership role: you lead through depth, design, and influence rather than through direct reports. Principal Engineers operate at the overall Infrastructure level. While you are assigned to a domain — in this case Observability — you are expected to contribute meaningfully to broader Infrastructure priorities, including initiatives and problems that lie outside your assigned team. Your expertise is applied wherever it moves the wider organization forward.

Requirements

  • Deep, current expertise in distributed systems and high-volume telemetry pipelines — OpenTelemetry internals, columnar and time-series storage and query engines (e.g., ClickHouse, InfluxDB), streaming with Kafka, and performance, cost, and reliability engineering at scale.
  • Proven ability to influence across an organization without direct authority — setting standards, driving architectural consensus, and mentoring senior engineers.
  • A product mindset toward internal platforms. You measure success by adoption and developer experience, not just capability — designing self-service, reducing friction, and making reliability and instrumentation the path of least resistance for thousands of engineers.
  • In-depth knowledge of system design, build, test, and operational debugging practices
  • Demonstrated experience architecting and delivering production distributed systems at scale
  • A track record of sustained technical influence and standard-setting across multiple engineering teams
  • Commitment to continuous learning and staying current with evolving technologies
  • Building and operating large-scale telemetry or data-intensive distributed systems
  • 4 year degree or equivalent experience. Continuing education to maintain thorough knowledge of technical domains along with staying current in latest technologies
  • 12+ years of experience in technology development or services
  • 4+ years of experience in strategic planning and setting technical direction
  • Supporting and operating 24x7 business-critical capabilities
  • Instrumenting products with metrics, logs, and traces to make them observable by default
  • Designing and deploying scalable APIs and microservices
  • Working with streaming and columnar/time-series data stores (e.g., Kafka, ClickHouse, InfluxDB)
  • Depth in one or more platform languages such as Golang, and orchestration with Kubernetes
  • Using version control (Git) and working in Linux/Unix-based environments
  • Communicating complex technical solutions clearly across diverse technical and non-technical audiences
  • Working effectively across multiple teams in a product-oriented environment
  • Experience in Java/ J2EE, sql/ no-sql (postgre, mongoDB, Cassandra, graph structure, etc.), Python, Ruby, Chef, Drone, Kubernetes containers, Cloud tech, etc.

Responsibilities

  • Owning the end-to-end architecture of the observability platform — metrics, logs, traces, alerting, SLOs, system maps, and AI-driven root cause analysis — across teams and roadmaps.
  • Solving the distributed-systems and high-volume telemetry-pipeline challenges where vendor defaults break down at Target scale and no standard playbook exists.
  • Driving observability as an internal product with first-class developer experience, self-service adoption, and “observable by default” instrumentation embedded directly in developer workflows.
  • Setting technical standards, leading design and architecture reviews, and multiplying the effectiveness of engineers across all three teams.
  • Contributing to Infrastructure-wide technical priorities beyond Observability, applying your judgment and expertise to problems that span teams and domains.
  • Representing Observability in cross-domain and enterprise architecture forums, and aligning the platform’s direction with the broader Target Tech strategy.

Benefits

  • medical
  • vision
  • dental
  • life insurance
  • 401(k)
  • employee discount
  • short term disability
  • long term disability
  • paid sick leave
  • paid national holidays
  • paid vacation
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service