Principal Core Engineer — Infra / SRE

Edgescale AIDenver, CO

About The Position

We’re looking for a Core Engineer at the Principal Infra / SRE level to own the reliability, scalability, upgradeability, and operational excellence of our edge platform at fleet scale. In this role, you’ll be the technical authority for designing and operating compound capabilities that span software, infrastructure, networking, security, data, and hardware—ensuring we can reliably deploy, upgrade, and manage fleets of thousands of devices with the highest technical rigor. You will set and enforce production standards, and you have the authority to stop changes that would put fleet safety or reliability at risk. During high-severity incidents, you are the technical owner—leading root-cause analysis and driving fixes across teams. This is a hands-on role for someone who thrives in a high-ownership setting and wants to build the infrastructure that makes real-world AI possible. You’ll operate in an AI-native way, using AI to assist diagnostics and operations while ensuring all production changes remain governed, reviewed, and auditable.

Requirements

  • 10+ years building and operating production infrastructure and distributed systems, including reliability engineering at scale across complex, multi-tenant or fleet environments.
  • Deep experience with SRE practices: SLOs/SLIs, error budgets, observability, incident response, postmortems, and operational automation (e.g., Kubernetes-based platforms, Linux systems, and automation through infrastructure-as-code).
  • Strong systems thinking across software, infrastructure, networking, and security, with the ability to drive outcomes across multiple domains and enforce production standards.
  • Proven ability to lead ambiguous, high-impact initiatives end-to-end with strong technical judgment, crisp execution, and disciplined change management.
  • Clear communicator and trusted technical partner to engineering leadership, with the ability to lead high-severity incident response and drive cross-team alignment.
  • Ownership mindset: outcomes over tasks.

Nice To Haves

  • Designing and operating fleet management and upgrade systems at scale, including safe rollout/rollback, configuration management, and health monitoring (e.g., canary deployments, staged rollouts, and verifiable rollback mechanisms).
  • Building observability platforms that make complex systems diagnosable and measurable across large distributed deployments (e.g., metrics/logs/tracing pipelines, alerting, and dashboards that drive action).
  • Security-first operations experience (secure boot, signed updates, audit logging, default-deny posture) and working in compliance-sensitive environments with governed production changes.
  • Experience operating systems under real-world edge constraints (limited connectivity, bandwidth limits, variable environments, high reliability requirements) and building automation that reduces operational variance.
  • Applying AI to operations and engineering workflows (automated diagnostics, agentic triage, runbook generation, anomaly detection) to increase rigor and speed while keeping production pathways reviewed and auditable.

Responsibilities

  • Own platform-wide reliability and scalability architecture across the fleet, including upgradeability, rollback safety, resilience, observability, and incident response.
  • Lead the design and delivery of compound capabilities that span multiple specialist domains (hardware, networking, security, data, infrastructure, and AI runtime).
  • Set and enforce production-grade standards for operational excellence, including SLOs/SLIs, error budgets, on-call readiness, change management, incident management, and postmortem practices, with the authority to stop changes that introduce unacceptable risk.
  • Serve as the technical owner during high-severity incidents, leading diagnosis, root-cause analysis, and coordinated remediation across teams.
  • Design and operate secure, automated fleet lifecycle systems for deployment, updates, configuration management, and health management at scale.
  • Drive the evolution of observability and telemetry systems (metrics, logs, traces, audit, fleet state) so issues are detectable, diagnosable, and preventable.
  • Partner with engineering and commercial teams to translate real-world constraints into platform-level requirements and prioritization decisions.
  • Operate in an AI-native way: develop and use AI systems to accelerate diagnostics, automate operational workflows, and increase engineering velocity, while ensuring all production changes remain governed, reviewed, and auditable.
  • Mentor senior engineers across domains, review technical designs, and raise the quality bar for architecture and reliability across the organization.

Benefits

  • Meaningful equity through stock options in an early-stage, high-growth company.
  • Health, dental, and vision coverage.
  • 401(k) with company match.
  • Flexible PTO.
  • Paid parental leave.
  • Commuter benefits.
  • Relocation and visa support for eligible roles.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service