SRE/ Site Reliability Engineer - Jersey City, NJ

Innova SolutionsJersey City, NJ
Onsite

About The Position

As a(n) SRE/ Site Reliability Engineer you will: Define and govern SLIs, SLOs, and SLAs for supported services. Centralized ownership of reliability practices across multiple applications and platforms. Lead incident management, major incident response, postmortems, and root cause analysis. Implement and maintain monitoring, observability, and alerting solutions using enterprise-standard tools. Drive toil reduction and automation through scripting and engineering solutions. Manage and improve AWS/cloud infrastructure reliability. Build and maintain CI/CD pipelines and infrastructure automation using tools such as Terraform. Perform capacity planning, performance tuning, and resiliency testing. Partner with application teams to improve reliability, availability, scalability, and operational excellence. Establish SRE best practices based on principles from the Google SRE Handbook. Act as an internal consulting team, helping development teams adopt reliability standards rather than owning application feature development.

Requirements

  • Strong Site Reliability Engineering experience
  • AWS cloud expertise
  • Python/scripting skills
  • Monitoring and observability experience
  • Incident management and production support
  • Terraform and infrastructure automation
  • CI/CD exposure
  • Knowledge of SLI/SLO/SLA concepts
  • Ability to work across multiple application teams

Responsibilities

  • Define and govern SLIs, SLOs, and SLAs for supported services.
  • Centralized ownership of reliability practices across multiple applications and platforms.
  • Lead incident management, major incident response, postmortems, and root cause analysis.
  • Implement and maintain monitoring, observability, and alerting solutions using enterprise-standard tools.
  • Drive toil reduction and automation through scripting and engineering solutions.
  • Manage and improve AWS/cloud infrastructure reliability.
  • Build and maintain CI/CD pipelines and infrastructure automation using tools such as Terraform.
  • Perform capacity planning, performance tuning, and resiliency testing.
  • Partner with application teams to improve reliability, availability, scalability, and operational excellence.
  • Establish SRE best practices based on principles from the Google SRE Handbook.
  • Act as an internal consulting team, helping development teams adopt reliability standards rather than owning application feature development.

Benefits

  • Medical & pharmacy coverage
  • Dental/vision insurance
  • 401(k)
  • Health saving account (HSA)
  • Flexible spending account (FSA)
  • Life Insurance
  • Pet Insurance
  • Short term and Long term Disability
  • Accident & Critical illness coverage
  • Pre-paid legal & ID theft protection
  • Sick time
  • other types of paid leaves (as required by law)
  • Employee Assistance Program (EAP)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service