Site Reliability Engineer Jobs

1,132 jobs found — updated daily

Senior Site Reliability Engineer (SRE)

Instrumental, Inc.Palo Alto, CA
$175,000 - $229,000

About The Position

Instrumental builds the manufacturing acceleration platform behind the world’s complex electronics. We capture digital exhaust and engineering context from assembly lines—images, test logs, BOM data, performance, repair cycles—and our AI engines identify insights that are difficult or impossible for human engineers to find. We accelerate the companies building the AI era by improving manufacturing yield, throughput, and ramp. NVIDIA, Meta, L3Harris, and their manufacturing partners rely on Instrumental to accelerate new product introduction and production. The Instrumental platform collects, intelligently transforms, and contextually presents manufacturing data to technical end-users, enabling them to optimize their manufacturing process in real-time. Our core technology is proprietary ML algorithms, packaged in an accessible, user-centric user interface – we believe we must have both the best technology and the best access to that technology to win.

Requirements

  • 5 or more years of DevOps or SRE experience
  • Expert knowledge with Linux, shell, containerization, Kubernetes, IaC (terraform preferred), monitoring, logging, and APM tools.
  • Proven ability to take initiative and drive impactful projects to completion efficiently and independently.
  • Comfort with ambiguity, pace, and frequent pivots inherent in a startup environment, with a track record of creating clarity for teams.
  • Experience introducing and integrating AI tools/processes into development and operation workflows.
  • Demonstrated skill in setting, iterating on, and measuring KPIs to ensure ongoing performance, reliability and efficiency.
  • U.S. citizen only (due to government contract requirements).

Nice To Haves

  • Network/application security and compliance experience is a plus.

Responsibilities

  • Deploying and operating commercial SaaS platforms on public cloud infrastructure, AWS preferred.
  • Utilizing expert knowledge with Linux, shell, containerization, Kubernetes, IaC (terraform preferred), monitoring, logging, and APM tools.
  • Taking initiative and driving impactful projects to completion efficiently and independently.
  • Creating clarity for teams in a startup environment with ambiguity, pace, and frequent pivots.
  • Introducing and integrating AI tools/processes into development and operation workflows.
  • Setting, iterating on, and measuring KPIs to ensure ongoing performance, reliability and efficiency.
  • Ensuring compliance with company engineering, security, access control, and privacy policies.
  • Promptly reporting suspected security incidents or policy violations.

Benefits

  • health
  • vision
  • dental
  • commuter plans
  • parental leave

Career Resources

Build a Resume for Site Reliability Engineer

The resume builder that gets results.

  • Get clear feedback so you look as qualified as you are
  • Align your resume with the job to get further in the process, faster
  • Take the guesswork out of resume writing

Explore Related Job Searches

© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service