Senior Site-Reliability Engineer

Entarian•Remote, VA
•$105,000 - $140,000

About The Position

We are seeking an experienced Senior Site-Reliability Engineer to join our infrastructure team. The ideal candidate will be responsible for ensuring the reliability, availability, and performance of our Windows-based production environments. You will bridge development and operations to deliver highly available services while maintaining operational excellence.

Requirements

  • 5+ years of experience in Systems Administration, DevOps, or Site-Reliability Engineering roles
  • Strong expertise in Windows Server environments (2016+), including Active Directory, IIS, and MS SQL.
  • Strong cloud skills, AWS experience preferred.
  • Advanced proficiency in scripting, including module development and integration with REST APIs
  • Hands-on experience with configuration management tools: Terraform (infrastructure provisioning), Puppet or Chef (configuration management)
  • Experience with monitoring and observability platforms (e.g., Prometheus, Grafana, Datadog, New Relic)
  • Solid understanding of networking concepts (DNS, TCP/IP, Load Balancing, VPN)
  • Strong problem-solving skills with the ability to troubleshoot complex issues across multiple technology layers

Nice To Haves

  • Bachelor's degree in Computer Science, Information Technology, or related field (or equivalent professional experience)
  • Certifications such as AWS Solutions Architect, Microsoft Certifications, or HashiCorp Certified: Terraform Associate
  • Experience with containerization technologies (Docker, Kubernetes)
  • Familiarity with CI/CD tools (Gitlab Pipelines, Jenkins, GitHub Actions)
  • Knowledge of security best practices and compliance frameworks
  • Experience with log aggregation and analysis tools (ELK Stack, Splunk)

Responsibilities

  • Design, implement, and maintain scalable infrastructure using Infrastructure as Code (IaC) practices
  • Develop and maintain automation scripts using various scripting languages (PowerShell, Python, Ruby, etc) for OS provisioning, configuration management, and operational tasks
  • Implement and manage configuration management solutions (Terraform, Puppet, and/or Chef) across hybrid environments
  • Monitor system health, performance, and availability using industry-standard tools and practices
  • Establish and enforce SLAs, SLOs, and error budgets for production services
  • Participate in on-call rotation and respond to incidents with a focus on rapid restoration and root cause analysis
  • Collaborate with development teams to improve deployment pipelines and release processes
  • Document operational procedures, runbooks, and architectural decisions
  • Conduct post-mortem reviews and implement corrective actions to prevent recurrence

Benefits

  • medical, dental, and vision insurance
  • life, AD&D, and disability insurance
  • paid time off and 11 company holidays
  • 401(k) retirement plan with company matching
  • additional employee benefits and wellness resources
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service