Sr. Site Reliability Engineer

SS&C TechnologiesWaltham, MA
$130,000 - $140,000Hybrid

About The Position

SS&C is a leading financial services and healthcare technology company headquartered in Windsor, Connecticut, with over 27,000 employees in 35 countries. Approximately 20,000 financial services and healthcare organizations rely on SS&C for expertise, scale, and technology. The Sr. Site Reliability Engineer will act as the guardian of the products, ensuring systems are reliable, scalable, and efficient enough to meet business goals. This role involves investigating customer-reported issues, providing technical support for information systems, and collaborating with the Operations Team for production deployments. The engineer will participate in production issue bridges, work with R&D and Operations teams to resolve customer issues, and respond to escalated incidents and monitoring alerts. A key responsibility is building diagnostic and analytical tools to improve Mean Time to Acknowledge (MTTA), Mean Time to Detect (MTTD), and Mean Time to Resolve (MTTR). Additionally, the role requires configuring and integrating monitoring tools to enhance system availability, scalability, and latency, and participating in an on-call rotation for critical issues.

Requirements

  • BS or MS in Computer Science or similar discipline
  • Strong work experience in Unix/Linux
  • Strong knowledge of Java Web-based enterprise applications, Python, or Bash to automate tasks and build tooling.
  • Strong work experience and troubleshooting skills in Kubernetes and Docker.
  • Working experience of Azure, AWS with (CloudWatch, EKS, EFS, S3, RedShift and other AWS services) and Infrastructure as Code (IaC) tools like Terraform or Ansible.
  • Working Knowledge in basic networking and various application and transport protocols. HTTP(s), JMS etc. TCP, UDP etc
  • Experience working with one or more of the following: Splunk, Datadog, Dynatrace, Zabbix, Prometheus, etc.
  • Working Experience on Akamai/Cloudflare DNS, CDN , DataStream & WAF.
  • Experience working with one or more of the following: Oracle, PostgreSQL, MongoDB
  • Experience working with messaging subsystems: RabbitMQ, Interconnect, AMQ.
  • Familiar with AI tools.

Nice To Haves

  • Minimum 7 years of experience in developing Software projects and/or DevOps/SRE.

Responsibilities

  • Investigate issues raised by customers, database administrators, and application support and suggest short term and long term solutions.
  • Provide day-to-day technical support in maintaining the information system, including responsibility for ensuring processes and outputs are complete and error-free.
  • Work with Operation Team for deployment and validation of changes to production during release/deployment/Change Request Implementations.
  • Participate in production issue bridge and work with different R&D and Operation teams to resolve customer issues.
  • Responding to and resolving escalated incidents for customer issues or monitoring alerts.
  • Building diagnostic and analytical tools that improve the MTTA, MTTD, and MTTR.
  • Configure and integrate commercially available monitoring tools into the production systems to improve availability, scalability, and latency.
  • Participate in on-call rotation and handling time critical issues during on-call schedule.
  • In-depth analysis of incident RCA.
  • Working with R&D and architecture teams on defects and runtime inefficiencies identified in the production environment.
  • Building systems/site monitoring tools for system health and APIs to ensure smooth operations of production systems.
  • Validate and Verify software deliverables for production readiness.
  • Risk assessment and mitigation of changes to the production systems.
  • Conducts project planning, cost analysis and vendor comparisons (POC/POV) and works on project implementation.
  • Works with development teams to enhance and improve system operability.
  • Conducts tests of network redundancy, resilience and failover of network elements to ensure up-time standards are fully achieved.
  • May be required to provide on-call service coverage with other department employees.

Benefits

  • medical, dental, and vision coverage
  • a 401(k) plan with company match
  • paid time off, holidays, and parental leave
  • professional development reimbursement opportunity
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service