Staff Site Reliability Engineer

Veeam Software
•$172,400 - $441,500•Remote

About The Position

Veeam is launching a global Site Reliability Engineering (SRE) function to support the rollout and operation of our new SaaS offering: the Veeam Data Cloud. SRE at Veeam is a software engineering discipline that uses code, data, and systems thinking to make product delivery safe and fast. As a Staff Site Reliability Engineer, you’ll lead by building reliable-by-default platforms, tooling, and patterns that product teams adopt at scale. You will serve as a hands-on technical leader within the SRE team, guiding senior engineers, influencing product development teams, and ensuring the systems we operate are built to be reliable, scalable, and observable from the ground up. You will drive strategic initiatives, mentor others in the practice of SRE, and help define architectural best practices across our platform. This role is pivotal in aligning teams, enforcing high standards, and scaling SRE principles globally within Veeam.

Requirements

  • 8+ years in software engineering for cloud-based products; significant time designing/operating distributed systems at scale.
  • Strong proficiency in at least one backend language (C#, Java, Go, TypeScript/Node.js) and in writing production‑grade services/libraries.
  • Deep hands‑on with Kubernetes, IaC (Terraform or Pulumi), and CI/CD (e.g., GitHub Actions, GitLab, ArgoCD).
  • Practical observability expertise (metrics, tracing, logging) and experience turning SLOs/error budgets into engineering workflows.
  • Ability to lead cross‑team initiatives, influence architecture, and deliver measurable reliability outcomes.
  • Comfortable with a follow‑the‑sun on‑call model (8 hour daytime rotations) and coverage.

Nice To Haves

  • Built reliability platforms (SLO/SLO policy engines, progressive delivery, chaos/validation) used by multiple teams.
  • Multi‑cloud or advanced Azure networking/traffic management (cross‑region failover, DNS, gateway, service mesh).
  • Performance engineering at scale (workload modeling, cost/perf tradeoffs, regression detection).
  • Security/compliance aware delivery (SOC 2/ISO/SOx/FedRAMP patterns) as code.

Responsibilities

  • Reliability features as productized code: libraries, services, and controllers (e.g., deployment safety guards, rate‑limiters, circuit breakers, load‑shedding adapters, back‑pressure controls) that product teams import and extend.
  • Observability platform: define the data model and implement telemetry pipelines (metrics + logs + traces), SLI/SLOs, and error‑budget policies; ship SDKs/CLI/plugins that let teams declare SLOs in code and gate releases.
  • Change safety toolchain: progressive delivery primitives (canary, blue/green, feature flags), automated rollback, and release validation - delivered as reusable services/operators and CI/CD integrations.
  • Resilience automation: fault‑injection APIs, chaos experiments, traffic shadowing, and load/perf harnesses baked into pre‑prod and prod pipelines.
  • Golden‑path platform components: Terraform/Pulumi modules, Kubernetes operators, Helm charts, and reference microservice templates (authn/z, config, tenancy, telemetry) with paved‑road docs.
  • Incident learning systems: post‑incident automation (context capture, timeline, action tracking), and code changes that remove classes of failure (not just runbooks).
  • Write high‑quality code in one or more of: Go, TypeScript/Node.js, C#, or Java. Design APIs, write tests, and ship iteratively.
  • Lead designs for distributed, multi‑region services (initially on Azure) with a focus on failure modes, graceful degradation, and operability.
  • Partner with Staff/Principal peers across product and platform to align on reliability standards and drive cross‑team adoption.
  • Instrument systems deeply (metrics, logs, traces) and automate detection/response; keep alerting actionable.
  • Lead complex incidents during your daytime; drive blameless learning and land systemic fixes in code.
  • Mentor senior engineers; raise the bar via design reviews, ADRs, and pair programming.

Benefits

  • Unlimited paid time off
  • 12 paid holidays including 4 global VeeaMe Days for self-care
  • 24 paid volunteer hours annually through Veeam Cares
  • Paid parental leave: 8 weeks for all parents, 16 weeks for birthing parents
  • Medical, dental, and vision coverage starting on your first day
  • Mental health support, therapy sessions, and digital wellness tools via our Employee Assistance Program
  • 401(k) retirement plan with company matching contributions
  • Fertility, adoption, and surrogacy support through Maven
  • AirVet: 24/7 virtual veterinary care at no cost
  • Legal services, identity protection, and supplemental health insurance options
  • Tax-advantaged spending accounts for healthcare, dependent care, and commuting
  • Opportunities to learn and grow through on-demand libraries (LinkedIn Learning, O’Reilly), mentoring, workshops, and learning events like our annual Global Day of Learning
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service