Sr. Business Analyst - SRBIZA

TekWissen•Hartford, CT
•$93•Hybrid

About The Position

This role serves as the highest level technical escalation point (L3 SME) for Linux/Unix infrastructure. The position involves managing and supporting large-scale Red Hat Enterprise Linux (RHEL), Oracle Linux, CentOS, SUSE, and Unix environments. Key responsibilities include advanced troubleshooting of OS, kernel, filesystem, storage, networking, and performance issues, as well as leading OS patching, upgrades, vulnerability remediation, and lifecycle management. The role ensures system availability, stability, scalability, and operational compliance, while also leading the resolution of critical incidents and major outages. Additionally, the position involves performing root cause analysis (RCA), implementing preventive measures, reviewing and approving changes related to compute infrastructure, and driving continuous service improvement initiatives. Capacity planning, resource optimization, analyzing system utilization, bottlenecks, and trends are also crucial. The role will implement proactive monitoring and self-healing mechanisms, and drive toil reduction through automation and operational innovation. Developing and maintaining shell scripting and automation solutions, automating provisioning, patching, compliance checks, and operational tasks are key. Contribution to Infrastructure as Code practices using Terraform and supporting CI/CD integration for infrastructure deployment activities are also expected.

Requirements

  • Compute - Linux
  • AWS
  • Terraform
  • Shell scripting and automation solutions
  • Infrastructure as Code practices

Responsibilities

  • Serve as the highest level technical escalation point (L3 SME) for Linux/Unix infrastructure.
  • Manage and support large-scale Red Hat Enterprise Linux (RHEL), Oracle Linux, CentOS, SUSE, and Unix environments.
  • Perform advanced troubleshooting of OS, kernel, filesystem, storage, networking, and performance issues.
  • Lead OS patching, upgrades, vulnerability remediation, and lifecycle management.
  • Ensure system availability, stability, scalability, and operational compliance.
  • Lead resolution of critical incidents and major outages.
  • Perform root cause analysis (RCA) and implement preventive measures.
  • Review and approve changes related to compute infrastructure.
  • Drive continuous service improvement initiatives.
  • Conduct capacity planning and resource optimization.
  • Analyze system utilization, bottlenecks, and trends.
  • Implement proactive monitoring and self-healing mechanisms.
  • Drive toil reduction through automation and operational innovation.
  • Develop and maintain shell scripting and automation solutions.
  • Automate provisioning, patching, compliance checks, and operational tasks.
  • Contribute to Infrastructure as Code practices using Terraform.
  • Support CI/CD integration for infrastructure deployment activities.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service