About The Position

As Site Reliability Engineer (SRE), you'll be part of the action-working closely with application teams and highly skilled engineers from diverse backgrounds to automate operations, optimize infrastructure, and troubleshoot issues in an exciting, fast-paced environment. You'll play a vital role in ensuring that our systems are reliable, scalable, and high-performing. This role is designed for driven individuals who: - Love learning new technologies and thrive in solving complex challenges. - Are independent, motivated, and excited to take on ambitious projects. - Excel at collaborating with engineering teams and can stay calm under pressure. - Have a passion for delivering quality, reliable solutions in a dynamic, high-energy workplace. - Excel in critical thinking, writing solid and clean code, planning, documenting, and communicating effectively both within and outside the team.

Requirements

  • BS degree or higher in Computer Science or a related field
  • Minimum 1-2 years of experience in Site Reliability Engineering, or an Infrastructure-focused role.
  • Proficiency in one or more programming languages (eg. Java, Python).
  • Understanding of data structure and algorithms, software development life cycle (SDLC).
  • Knowledge on fundamentals of network, databases, system administration.
  • Understanding on version control systems like Github.
  • Familiarity with UNIX/Linux Operating Systems.
  • Knowledgeable with container based technologies such as Docker, Kubernetes, or EKS.
  • Knowledgeable with modern web services architectures and cloud platforms such as AWS, GCP.
  • Exceptional analytical and troubleshooting skills in complex Unix/Linux systems environment and applications implementations.
  • Experience with telemetry and monitoring/observability tools such as Splunk, Grafana or Prometheus.
  • Ability to build tools from scratch.
  • Ability to work in a collaborative environment.

Nice To Haves

  • Course work on Web technologies, Machine Learning will be a plus.
  • Strong problem-solving, communication skills

Responsibilities

  • Build and maintain robust and highly available Continuous Integration pipelines.
  • Identify and implement automation solutions to eliminate repetitive manual processes.
  • Automate build & deployment processes.
  • Design and implement new software to streamline manual operations and enhance developer productivity for application management.
  • Troubleshoot and handle platform issues for both cloud and on premises infrastructures.
  • Improve operational readiness by performing root cause analysis of critical issues and creating long-term solutions.
  • Maintain scalable logging and monitoring infrastructure.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service