SRO Engineer

VersantEnglewood Cliffs, NJ
$130,000 - $160,000Onsite

About The Position

The System Reliability Operations (SRO) Engineer is a hands-on operational engineering role responsible for helping establish and scale reliability practices across VERSANT’s technology organization. As a member of the newly formed Platform SRO team, this role will develop standards, documentation, operational procedures, and reusable practices that improve the reliability, resilience, and supportability of enterprise platforms and services. Reporting to the Director, Platform SRO and working closely with the SRE Lead, the SRO Engineer will partner with Software Engineering, Platform Engineering, Infrastructure, Cloud Engineering, Security, Enterprise Technology, and Production Operations teams. The engineer will help translate reliability principles into clear, practical guidance that teams can consistently apply across cloud, on-premises, and hybrid environments. This role is well suited for an engineer who combines technical troubleshooting experience with strong documentation, organization, and collaboration skills. Success will require building repeatable operational practices, improving service documentation, supporting incident and problem management, and helping engineering teams prepare their systems for reliable production operation.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience.
  • 3+ years of experience in System Reliability Operations, Site Reliability Engineering, IT Operations, Production Support, Platform Engineering, DevOps, Infrastructure Engineering, or a related technical role.
  • Experience supporting production applications, infrastructure, cloud platforms, or enterprise technology services.
  • Demonstrated experience creating technical documentation, operational runbooks, troubleshooting guides, or standard operating procedures.
  • Working knowledge of incident management, problem management, root cause analysis, and operational escalation practices.
  • Familiarity with reliability concepts such as SLIs, SLOs, availability, resiliency, service ownership, and operational readiness.
  • Experience with monitoring, logging, alerting, or observability tools.
  • Familiarity with cloud platforms such as AWS, Azure, or GCP and modern distributed systems.
  • Experience using enterprise collaboration and workflow tools such as ServiceNow, Jira, Confluence, or similar platforms.
  • Strong analytical and troubleshooting skills, with the ability to organize complex technical information into clear, actionable guidance.
  • Strong written and verbal communication skills with the ability to collaborate across engineering, infrastructure, security, and operations teams.
  • Ability to manage multiple documentation and operational improvement initiatives in a developing organization.

Nice To Haves

  • Experience supporting media, broadcast, streaming, digital publishing, or other highly available and time-sensitive environments.
  • Familiarity with ITIL Incident, Problem, Change, and Knowledge Management practices.
  • Experience supporting operational readiness reviews, disaster recovery exercises, or production service onboarding.
  • Familiarity with observability tools such as Datadog, Splunk, Grafana, Prometheus, New Relic, Elastic, or similar platforms.
  • Experience with scripting or automation using Python, PowerShell, Bash, or comparable technologies.
  • Familiarity with CI/CD pipelines, Infrastructure as Code, Kubernetes, or containerized platforms.
  • Experience developing documentation standards, knowledge-management structures, or enterprise operating procedures.
  • Relevant cloud, ITIL, ServiceNow, or SRE-related certifications are a plus.

Responsibilities

  • Help establish, document, and maintain reliability engineering standards for applications, platforms, infrastructure, and shared services.
  • Translate SRO and SRE principles into practical guidelines, checklists, templates, and operating procedures that can be adopted across engineering teams.
  • Support the definition and implementation of service-level indicators, service-level objectives, availability targets, and operational health measures.
  • Develop standard criteria for production readiness, operational acceptance, service ownership, monitoring, escalation, and support.
  • Identify opportunities to standardize reliability practices across teams while accounting for different platforms, technologies, and business requirements.
  • Research operational trends and lessons learned from incidents and recommend updates to standards and documentation.
  • Create and maintain operational runbooks, troubleshooting guides, service support documentation, escalation procedures, and standard operating procedures.
  • Develop reusable templates for production readiness reviews, incident response, root cause analysis, service inventories, dependency mapping, and operational handoffs.
  • Partner with engineering teams to improve the quality, consistency, and accessibility of technical and operational documentation.
  • Help establish documentation ownership, review cycles, version control, and governance practices.
  • Organize operational knowledge in platforms such as Confluence, ServiceNow, or similar enterprise knowledge-management tools.
  • Ensure critical services have current documentation covering architecture, dependencies, monitoring, recovery procedures, support contacts, and known risks.
  • Promote documentation as an ongoing engineering responsibility rather than a one-time delivery activity.
  • Support operational readiness reviews for new services, major releases, migrations, and significant platform changes.
  • Validate that services have appropriate monitoring, alerting, dashboards, support procedures, dependency documentation, and recovery plans before production launch.
  • Work with engineering teams to identify and close readiness gaps.
  • Help define consistent service onboarding and operational acceptance processes for the Platform SRO organization.
  • Maintain service inventories and ownership information for critical enterprise platforms.
  • Support disaster recovery, failover, capacity, and resilience testing activities.
  • Participate in incident response for production and platform issues, supporting technical investigation, coordination, communications, and documentation.
  • Assist incident commanders and technical teams during high-severity events by maintaining timelines, tracking actions, and organizing relevant service information.
  • Support post-incident reviews and root cause analysis, ensuring findings, contributing factors, and corrective actions are clearly documented.
  • Track corrective and preventive actions through completion and escalate overdue or recurring reliability risks.
  • Analyze incident patterns and recurring issues to identify opportunities for improved documentation, automation, monitoring, or engineering standards.
  • Help maintain incident-response playbooks, severity definitions, escalation paths, and communication templates.
  • Partner with engineering teams to ensure applications and platforms have appropriate monitoring, logging, alerting, and dashboard coverage.
  • Help document observability requirements and recommended practices for enterprise services.
  • Review alerts for clarity, actionability, ownership, and alignment with documented response procedures.
  • Assist with the development of service-health dashboards and operational reporting.
  • Identify monitoring gaps and help teams create supporting runbooks and troubleshooting instructions.
  • Contribute to efforts that reduce alert noise and improve the quality of operational signals.
  • Identify repetitive operational tasks that can be standardized or automated.
  • Develop scripts, workflow automations, templates, or lightweight tools that improve documentation, readiness assessments, incident follow-up, and operational reporting.
  • Support integration of reliability checks and operational requirements into CI/CD and change-management workflows.
  • Help measure adoption of SRO standards and identify areas requiring additional guidance or enablement.
  • Contribute to the continuous improvement of Platform SRO processes, tools, and operating practices.
  • Partner with technical teams to understand their services, operational challenges, and support requirements.
  • Facilitate working sessions to create runbooks, document dependencies, define service objectives, and improve operational readiness.
  • Provide guidance to engineers on documentation, incident response, observability, and reliability fundamentals.
  • Develop internal reference materials and contribute to training or enablement sessions for engineering and operations teams.
  • Promote consistent reliability practices while building strong relationships across a matrixed enterprise organization.
  • Communicate operational risks, documentation gaps, and improvement recommendations clearly to technical stakeholders and Platform SRO leadership.

Benefits

  • health insurance
  • retirement plans
  • paid time off
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service