SRO Lead

VersantEnglewood Cliffs, NJ
$170,000 - $190,000

About The Position

The System Reliability Engineering (SRE) Lead is a hands-on technical leader responsible for improving the reliability, performance, and scalability of VERSANT’s software, production, and platform systems. Reporting to the VP of Infrastructure, this role works closely with Software Engineering, Production Engineering, Platform Engineering, and Infrastructure teams to implement reliability best practices, drive end-to-end testing, and ensure systems perform under real-world conditions. This is a player-coach role focused on execution—building testing frameworks, improving observability, and helping teams proactively identify and resolve system weaknesses before they impact production.

Requirements

  • 5+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, or Infrastructure roles.
  • Strong hands-on experience operating and troubleshooting production systems.
  • Experience implementing integration testing, E2E testing, or performance/load testing frameworks.
  • Familiarity with observability tools (metrics, logging, tracing) and monitoring systems.
  • Experience with cloud platforms (AWS, GCP, or Azure) and distributed systems.
  • Experience working with CI/CD pipelines and automation.
  • Strong debugging and problem-solving skills in complex systems.

Nice To Haves

  • Experience supporting media, broadcast, or real-time production systems.
  • Familiarity with high-throughput or low-latency systems.
  • Exposure to SRE concepts such as SLIs/SLOs and incident management practices.
  • Experience with containerized environments (Kubernetes, Docker) is a plus.
  • Familiarity with infrastructure as code (Terraform, CloudFormation).
  • Strong collaboration skills and ability to work across multiple engineering teams.

Responsibilities

  • Partner with engineering teams to improve system reliability, availability, and performance.
  • Help define and implement SLIs, SLOs, and basic reliability standards across services.
  • Identify reliability gaps and work with teams to address risks in system design and operations.
  • Contribute directly to code, tooling, and automation that improves system resilience.
  • Design and implement end-to-end (E2E) testing workflows across distributed systems.
  • Build and maintain integration testing frameworks validating cross-service dependencies.
  • Execute and scale load and performance testing to validate systems under peak conditions.
  • Partner with teams to integrate automated testing into CI/CD pipelines.
  • Help establish practical testing standards and ensure adoption across teams.
  • Support performance benchmarking and system capacity planning efforts.
  • Analyze system performance and identify bottlenecks across application and infrastructure layers.
  • Partner with infrastructure and platform teams to optimize system throughput and latency.
  • Implement and improve monitoring, logging, and alerting across services.
  • Help ensure systems are observable, debuggable, and well-instrumented.
  • Participate in incident response and support root cause analysis efforts.
  • Contribute to post-incident reviews and track follow-up actions to improve reliability.
  • Work closely with software, platform, enterprise and production engineering teams to embed reliability practices into day-to-day development.
  • Provide guidance and hands-on support for testing, observability, and performance improvements.
  • Help standardize tools, frameworks, and processes used across teams.
  • Mentor engineers on reliability engineering fundamentals and testing best practices.

Benefits

  • health insurance
  • retirement plans
  • paid time off
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service