About The Position

We are looking for a Senior Platform Engineer with at least 7 years of experience to join our Platform Engineering team, with a strong focus on observability. In this role, you will build, operate, optimize, and enhance the platforms that provide visibility into our applications and infrastructure. You will work closely with Staff and Principal Engineers who define architecture and technical direction, and you will be responsible for implementing, operating, and continuously improving observability solutions across the organization.

Requirements

  • 7+ years of hands-on experience in Platform Engineering, DevOps, SRE, Infrastructure Engineering, or Observability Engineering.
  • Strong hands-on experience with at least one major observability stack (e.g., Elastic, Grafana, Datadog, Splunk, or equivalent).
  • Experience operating and troubleshooting production observability platforms at scale.
  • Strong understanding of logs, metrics, distributed tracing, telemetry pipelines, and core observability concepts.
  • Experience onboarding telemetry/data sources and supporting internal customers through integration and troubleshooting.
  • Hands-on experience with Kubernetes, containers, Linux, networking, and cloud infrastructure.
  • Familiarity with OpenTelemetry or similar instrumentation and collection frameworks.
  • Advanced troubleshooting skills across applications, infrastructure, and distributed systems.
  • Hands-on experience with Infrastructure as Code tools such as Terraform.
  • Experience building CI/CD pipelines using GitHub Actions, Azure DevOps, Jenkins, or similar tools.
  • Hands-on automation and scripting experience using Python, Go, Bash, or similar languages.
  • Experience with performance tuning, capacity planning, data lifecycle management, and platform optimization.
  • Working knowledge of telemetry collection, signal processing, cross-signal analysis, alerting, and production troubleshooting.
  • Consistently demonstrated ability to build, operate, troubleshoot, and improve production-grade platform solutions.
  • Experience implementing SLIs, SLOs, alerting standards, and reliability monitoring.

Nice To Haves

  • Experience with modern observability tooling such as Grafana, Prometheus, ClickHouse, or pipeline/stream processing platforms.
  • Experience with log/metric/trace collectors and streaming technologies (e.g., Fluent Bit, Kafka).
  • Experience with high-volume telemetry environments and cost optimization.
  • Experience with GitOps workflows and observability-as-code practices.
  • Experience creating onboarding documentation, runbooks, and self-service guidance for platform users.

Responsibilities

  • Build, operate, maintain, and enhance enterprise observability platforms for applications and infrastructure.
  • Operate and optimize observability backends, ingestion pipelines, agents/collectors, and data lifecycle management.
  • Implement observability solutions and standards defined by Staff and Principal Engineers.
  • Build and maintain telemetry pipelines for logs, metrics, traces, and events.
  • Lead telemetry and data onboarding for applications, infrastructure, Kubernetes, cloud, and platform teams — including integration guidance, pipeline configuration, and validation of data quality.
  • Provide product support to internal teams — troubleshoot telemetry ingestion issues, query performance problems, missing or incorrect data, and platform usage questions.
  • Monitor and improve platform availability, performance, capacity, scalability, and reliability.
  • Troubleshoot complex production issues related to telemetry ingestion, processing, storage, indexing, and query performance.
  • Optimize data stores, retention policies, indexing strategies, and storage utilization.
  • Build and maintain dashboards, alerts, integrations, and operational tooling.
  • Automate platform provisioning, configuration, and deployments using Infrastructure as Code and CI/CD.
  • Partner with Application, SRE, Infrastructure, and Security teams to improve telemetry quality, coverage, and incident response.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service