Principal Engineer Software, Dev & Production Infrastructure (Chronosphere)

Palo Alto NetworksBoston, MA
$147,000 - $237,500Remote

About The Position

Chronosphere, a Palo Alto Networks company, is the observability platform built for control in the modern, containerized world. Chronosphere empowers customers to focus on the data and insights that matter by reducing complexity, optimizing costs, and remediating issues faster. Chronosphere reduces data volumes and associated costs by 84% on average while saving developers thousands of hours. Recognized as a leader by major analyst firms, Chronosphere is trusted by the world's most innovative brands, including DoorDash, Affirm, and Zillow. Our Infrastructure organization is scaling rapidly, and we are looking for world-class Principal Engineers to drive the future of our platform. We are actively hiring for two distinct pillars within the same org and will consider all applicants for both tracks. During our unified interview process, we will work closely with you to determine which team—or combination of responsibilities—aligns best with your technical strengths and career aspirations. The Two Areas You Will Be Considered For: Production Engineering (Reliability & Scale) The Core Focus: Ensuring all live components of our our global platform are optimized, highly reliable, and battle-tested. This role focuses on pushing our massive distributed cloud resources to their absolute limits. Dev Infrastructure (Developer Velocity & Platforms) The Core Focus: Building the platform and tooling that enables developer velocity and software reliability for the entire Chronosphere engineering organization. This team owns the full end-to-end SDLC. Combined Principal-Level Leadership Regardless of which pillar you align with most, as a Principal Engineer you will: Own massive technical initiatives from inception to delivery, balancing feature velocity with long-term technical debt. Define platform standards and reference architectures that span a 1–3 year horizon. Act as the "glue" across the organization, consulting on infrastructure best practices and up-leveling the team through dedicated mentorship.

Requirements

  • 8–10+ years of relevant experience in high-scale infrastructure, production engineering, systems architecture, or developer platform roles.
  • 7 + years of experience in at least one backend language (e.g., Go, Python, Java, C++, C#, or Rust).
  • 5 + years of deep expertise building and debugging systems that deal with CAP theorem trade-offs, eventual consistency, and distributed tracing.
  • Robust experience working with AWS or GCP and navigating Kubernetes clusters.
  • Deep knowledge of Linux internals, process management, resource isolation, and networking protocols (OSI model, load balancing, service meshes).
  • A proven track record of estimating complex work effectively, debugging distributed architectures via logs/traces, and proactively identifying design edge cases.

Nice To Haves

  • Active engagement or experience with open-source communities.
  • Experience or active interest in using AI coding assistants (like Cursor or Claude) to accelerate engineering velocity and automate boilerplate tasks.

Responsibilities

  • Solve complex distributed systems problems at the largest scales in the world.
  • Scale our production infrastructure and architecture globally using Kubernetes, GCP, and AWS.
  • Use Go to build and operate highly scalable, resilient backend services.
  • Develop the internal tools necessary to monitor, alert, and optimize our live production footprint.
  • Architect & Build: Design high-scale developer tooling, dynamic testing environments, and CI/CD pipelines in a 100% modern, containerized microservices ecosystem.
  • Infrastructure as Code (IaC): Treat infrastructure as a first-class citizen by defining and managing entire ephemeral environments using declarative IaC (Terraform).
  • Drive Systemic Quality: Identify and eliminate systemic bottlenecks across the development lifecycle through architectural changes, advanced tooling, and near-real-time telemetry processing.
  • Own massive technical initiatives from inception to delivery, balancing feature velocity with long-term technical debt.
  • Define platform standards and reference architectures that span a 1–3 year horizon.
  • Act as the "glue" across the organization, consulting on infrastructure best practices and up-leveling the team through dedicated mentorship.

Benefits

  • Restricted stock units
  • Bonus
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service