Sr Software Engineer - Reliability Engineering

Cox EnterprisesHoltsville, NY
$121,800 - $203,000Hybrid

About The Position

We're hiring a Sr. Software Engineer - Reliability Engineer who can code across the stack and cares deeply about reliability. You'll design and build infrastructure, observability tooling, and operational systems—treating resilience and debuggability as first-class concerns. You'll own projects end-to-end: from architecture → code → deployment → production. You'll split time between infrastructure-as-code, incident response, system improvements, and mentoring. You'll work on a team that ships quality systems while maintaining operational excellence across a platform serving millions of dealership transactions daily.

Requirements

  • 5+ years software engineering, platform engineering, or infrastructure engineering experience.
  • Strong coding in Python, Go, Java, or equivalent; writes clean, testable code.
  • AWS hands-on: EC2, RDS, DynamoDB, S3, Aurora, Lambda, VPCs, Athena.
  • Terraform or equivalent infrastructure-as-code experience.
  • Docker and container orchestration (Kubernetes or similar).
  • Debugging on Linux and Windows platforms; able to troubleshoot complex systems using logs, metrics, and architectural knowledge.
  • System design thinking: can architect scalable systems and reason about trade-offs.
  • Bachelor’s degree in a related discipline and 4 years’ experience in a related field. The right candidate could also have a different combination, such as a master’s degree and 2 years’ experience; a Ph.D. and up to 1 year of experience; or 16 years’ experience in a related field.

Nice To Haves

  • Experience with observability tools (New Relic, Splunk, Prometheus).
  • Incident response experience; familiar with postmortem practices.
  • Interest in or hands-on experience with SRE concepts (SLOs, resilience, failure modes).
  • Cost optimization mindset; has identified and eliminated cloud waste.
  • Windows and Linux system troubleshooting and performance analysis.
  • Experience with CI/CD pipelines and deployment automation.

Responsibilities

  • Design and implement resilience initiatives: redundancy, failover, disaster recovery, data protection.
  • Write infrastructure-as-code (Terraform); manage 50+ AWS accounts with infrastructure patterns.
  • Own system health: proactively maintain application performance, minimize downtime, ensure consistent user experience.
  • Evolve team's SRE standards and practices.
  • Build observability into systems: logging, metrics, distributed tracing, alert design.
  • Improve monitoring frameworks; enable faster incident detection and resolution.
  • Design dashboards and alerts that help teams understand system behavior.
  • Partner with application teams on service instrumentation.
  • Drive significant reductions in cloud spend through architectural improvements and resource utilization.
  • Review infrastructure for efficiency; identify and eliminate waste.
  • Balance cost, performance, and reliability in design decisions.
  • Build production systems, APIs, internal tools, and automation with clean, well-tested code.
  • Design for maintainability, operational simplicity, and reliability.
  • Participate in code review and technical design discussions.
  • Mentor junior engineers on code quality and architectural thinking.
  • Participate in on-call rotations; debug and resolve production incidents.
  • Conduct postmortem analysis; drive systemic improvements.
  • Develop operational procedures and runbooks.

Benefits

  • Flexible vacation with pay
  • Seven paid holidays
  • Up to 160 hours of paid wellness annually
  • Bereavement leave
  • Time off to vote
  • Jury duty leave
  • Volunteer time off
  • Military leave
  • Parental leave
  • Health care insurance (medical, dental, vision)
  • Retirement planning (401(k))
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service