About The Position

Our Platform Engineering team is looking for a Senior Site Reliability Engineer to help run, maintain, and improve the performance of our Ruby on Rails Stack and Infrastructure. You will help scale our application using best in class architecture and software design. This includes training, software engineering, system design, and operational practices that support the needs of our engineers and customers while accounting for future growth. You will be entrusted with proactively identifying and owning initiatives that will help improve the performance, reliability, and scalability of our application stack and databases. Our team treats AI as a core part of how we engineer. We build and use AI agents, skills, and automated workflows to handle toil, speed up investigation and remediation, and give our engineers more time for high-leverage work.

Requirements

  • 5+ years of Ruby/Rails Experience
  • 3+ years of AWS Experience
  • Kubernetes experience
  • Experience with profiling and benchmarking source code
  • Effective at code review and identifying potential performance problems before they reach production
  • Experience with Datadog or other APM tools
  • Excellent written and verbal communication skills

Nice To Haves

  • Experience building AI agents, LLM-powered automations, or integrations (e.g., using tool/function calling or MCP) for engineering or operations workflows
  • Infrastructure as Code tools (Terraform)
  • Deep understanding of cloud network fundamentals (routing, firewalls, load balancers, CDNs, VPCs, etc.)
  • Experience with distributed event and data stores, such as Kafka, Redis, Elasticsearch, Memcached, and TimescaleDB
  • You know a thing or two about the fleet management industry

Responsibilities

  • Proactively identify, triage, and resolve performance issues
  • Enhance system observability by monitoring performance metrics across Ruby, Rails, and database systems, including SLOs and SLIs
  • Build and maintain AI agents, skills, and automations that reduce operational toil across incident response, triage, and routine maintenance
  • Use AI-assisted tooling to accelerate performance analysis, root-cause investigation, and code review
  • Collaborate with other SREs to proactively identify and address performance bottlenecks
  • Help product engineers adopt AI-driven workflows for performance and reliability best practices
  • Lead database capacity planning and upgrade initiatives
  • Manage the database-specific components of disaster recovery planning and execution
  • Oversee backup systems and pre-production databases
  • Create and maintain infrastructure and operations documentation, including runbooks and context that both engineers and AI agents can act on
  • Participate in the on-call rotation

Benefits

  • Multiple health/dental coverage options (100% coverage for employee, 50% for family)
  • Vision insurance
  • Incentive stock options
  • 401(k) match of 4%
  • PTO - 4 weeks (increases at year two!)
  • 12 company holidays + 2 floating holidays
  • Parental leave - birthing parent (16 weeks paid) non-birthing (4 weeks paid)
  • FSA & HSA options
  • Short and long term disability (short term 100% paid)
  • Community service funds
  • Professional development funds
  • Wellbeing fund - $150 quarterly
  • Business expense stipend - $125 quarterly
  • Mac laptop + new hire equipment stipend
  • Fully stocked kitchen with tons of drinks & snacks (BHM only)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service