Senior Platform Engineer - Observability

Capital GroupCharlotte, NC
Hybrid

About The Position

As a Senior Platform Engineer on our observability platform team, you’ll design, build, and operate the observability capabilities that thousands of engineers rely on to understand the health, performance, and behavior of their applications and infrastructure. You’ll build on open telemetry standards to give teams a unified, vendor-flexible experience across metrics, events, logs, and traces—turning raw signals into actionable insight that accelerates troubleshooting, strengthens reliability, and improves the developer experience. This is a hands-on senior engineering role. You’ll engineer telemetry collection and pipelines, automate onboarding so teams get observability out-of-the-box, and deliver Observability-as-Code—monitors, dashboards, and alerts managed as versioned, reusable modules. A central goal is a true single pane of glass: you’ll correlate metrics, events, logs, and traces into one unified experience so engineers can move seamlessly from signal to root cause without switching tools. You’ll integrate with leading observability backends (for example, a SaaS APM/metrics platform such as Datadog alongside log platforms and open-source stacks like Prometheus and Grafana) while keeping the architecture standards-based and portable, so we’re never locked to a single vendor. You’ll define golden-signal and SLI/SLO standards, tune cost and cardinality, advance AIOps and anomaly detection, and coach engineering teams to raise observability maturity across the organization. You’ll also develop with AI—using AI-assisted coding tools and agentic workflows to build and refactor platform tooling faster, while keeping quality and security high.

Requirements

  • Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent technical experience.
  • At least 8 years in Software Engineering, Platform Engineering, SRE, or Cloud Engineering, with a strong software background designing, building, and operating production-grade systems at scale.
  • Strong software engineering fundamentals.
  • Solid foundation in networking, security, and cloud-native architectures.
  • Deep, hands-on observability platform experience—instrumenting applications and building telemetry pipelines and backends across metrics, events, logs, and traces (MELT), grounded in OpenTelemetry standards with a vendor-neutral, portable design mindset—using modern tooling such as a leading SaaS APM/metrics platform (e.g., Datadog), open-source stacks (Prometheus, Grafana, OpenTelemetry Collector), and log platforms (e.g., Splunk, CloudWatch).
  • Proven experience delivering unified, single-pane-of-glass observability—correlating metrics, events, logs, and traces (including trace-to-log correlation and consistent tagging/context propagation) so engineers move from signal to root cause in one experience without switching tools.
  • Expert in at least one of Python or Go.
  • Comfortable reading and instrumenting services in additional languages such as Java, .NET (C#), and Node.js/TypeScript.
  • Strong telemetry query-language skills (e.g., PromQL, SQL, or equivalent).
  • Develop with AI as part of your craft—using AI-assisted coding tools (e.g., GitHub Copilot or equivalent) and agentic workflows to accelerate development, testing, and refactoring of platform tooling—while applying sound engineering judgment to review, validate, and secure AI-generated code in line with approved enterprise AI platforms and controls.
  • Strong Infrastructure-as-Code skills (Terraform/OpenTofu).
  • Deliver Observability-as-Code—monitors, dashboards, and alerts as versioned, reusable modules deployed through CI/CD.
  • Telemetry collection and pipeline engineering (collector/agent fleet management, data hygiene and tagging standards, and cost/cardinality/sampling optimization) on cloud-native AWS (EKS/Kubernetes, containers).

Responsibilities

  • Design, build, and operate observability capabilities for thousands of engineers.
  • Build on open telemetry standards for a unified, vendor-flexible experience across metrics, events, logs, and traces.
  • Engineer telemetry collection and pipelines.
  • Automate onboarding for out-of-the-box observability.
  • Deliver Observability-as-Code (monitors, dashboards, alerts as versioned, reusable modules).
  • Correlate metrics, events, logs, and traces into a unified single pane of glass experience.
  • Integrate with leading observability backends (e.g., Datadog, Prometheus, Grafana, Splunk, CloudWatch).
  • Define golden-signal and SLI/SLO standards.
  • Tune cost and cardinality.
  • Advance AIOps and anomaly detection.
  • Coach engineering teams to raise observability maturity.
  • Develop with AI using AI-assisted coding tools and agentic workflows.
  • Define reliability and observability standards (golden signals, distributed tracing/context propagation, SLIs/SLOs, structured logging, RUM/synthetics).
  • Apply AIOps/anomaly detection to accelerate detection and root-cause analysis.
  • Lead technical workstreams independently.
  • Mentor engineers.
  • Communicate clearly with technical teams and stakeholders (Agile/SCRUM).

Benefits

  • Competitive salary
  • Bonuses
  • Generous time-away
  • Health benefits from day one
  • Opportunity for flexible work options
  • 2-for-1 matching gifts for charitable contributions
  • Opportunity to secure annual grants for charitable organizations
  • On-demand professional development resources
  • Company-funded retirement contribution (15% of eligible earnings)
  • Annual performance bonus
  • Capital's annual profitability bonus
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service