SRE Technical Lead

Bright Vision TechnologiesTown of Guilderland, NY
Hybrid

About The Position

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential. The Senior IT Technical Lead for Site Reliability Engineering (SRE) is a hands-on, expert-level role. This individual is responsible for ensuring the operational excellence of critical production systems. This role is less about managing project timelines and resources, and more about creating innovative solutions to improve a mission critical service. The technical lead will drive automation, optimize performance, lead incident response from a technical standpoint.

Requirements

  • At least 8 years of experience in IT, with significant time spent in a senior-level SRE or similar role.
  • Deep SRE and Systems Expertise: Comprehensive knowledge and hands-on experience applying SRE principles to manage the reliability and scalability of enterprise-level systems. This includes cloud platforms (e.g., AWS, GCP, Azure), microservices, containers (Kubernetes).
  • Automation and Tooling: Proven ability to write and implement code (e.g., Python) and leverage automation tools to eliminate manual toil and create scalable, self-healing systems.
  • Observability and Analysis: Expertise in implementing robust monitoring, alerting, logging, and tracing systems (e.g., Dynatrace, Splunk, ELK Stack).

Nice To Haves

  • Exceptional leadership, communication, and interpersonal skills, with the ability to influence and collaborate effectively with both technical and non-technical stakeholders, including senior leadership.

Responsibilities

  • Ensuring the operational excellence of critical production systems.
  • Creating innovative solutions to improve a mission critical service.
  • Driving automation.
  • Optimizing performance.
  • Leading incident response from a technical standpoint.
  • Implementing robust monitoring, alerting, logging, and tracing systems.
  • Analyzing system metrics to proactively identify issues and drive data-backed technical decisions for performance tuning and capacity planning.
  • Building and maintaining internal tools that enhance infrastructure and operations.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service