Site Reliability Engineer

InfosysChandler, AZ
Onsite

About The Position

In the assigned Job Role of Infrastructure Consultant 2, your Area Of Responsibility will be as below: ⦁ Collaborate with internal and client teams to resolve complex incidents, conduct root cause analyses, and document findings with preventive recommendations ⦁ Participate in evaluation of client IT infrastructure, prepare actionable assessment reports, and support due diligence to document infrastructure maturity and improvement opportunities ⦁ Contribute to the design of scalable, cost-effective IT infrastructure solutions, review reusable components, and develop technical documentation for deployed systems ⦁ Align release schedules and environment readiness, execute deployments as per protocols, perform post-deployment testing, and manage version control to track changes ⦁ Co-ordinate maintenance schedules, emergency fixes, and technology upgrades while ensuring uninterrupted integration into existing systems and processes ⦁ Facilitate performance data analysis across systems, coordinate insights on system behavior, and support capacity planning to optimize performance ⦁ Conduct security checks, recovery drills, and compliance audits, implement security measures, and coordinate continuity plans to maintain adherence to standards ⦁ Gather feedback to identify automation opportunities, analyze existing infrastructure processes, and propose enhancements for efficiency gains ⦁ Act as liaison with onsite, offshore, and vendor teams to document project requirements, ensuring effective collaboration ⦁ Develop a centralized repository of technical and procedural knowledge, leveraging insights from other projects to drive efficiency and retain organizational expertise Your contribution to the team: ⦁ A collaborative spirit and excellent communication skills. ⦁ Ability to handle complex incidents and implement resolutions ⦁ A knack for conducting IT infrastructure assessment and identifying key optimization opportunities ⦁ Focused approach towards deployment management, system optimization, and process automation initiatives including sector specific focus ⦁ The ability to work with cross-functional teams

Requirements

  • Strong experience in Site Reliability Engineering / Production Engineering.
  • Hands-on expertise with: IBM MQ (queue managers, clustering, channels, DLQ management).
  • Kafka / Confluent platform (topics, brokers, partitions, consumer groups).
  • Large-scale distributed messaging systems and runtime management.
  • Deep understanding of: System reliability, scalability, and high availability design.
  • Messaging reliability patterns (guaranteed delivery, retry handling, replay, ordering).
  • Incident management, root cause analysis, and problem management.
  • Experience with: Observability tools (Dynatrace, Splunk, Prometheus, Grafana) for messaging platforms.
  • Event and anomaly detection in high-volume systems.
  • Strong scripting/automation skills: Shell, Python, PowerShell.
  • Experience managing Linux/Unix and Windows production environments.
  • Knowledge of: Event-driven architecture and messaging-based integration patterns.
  • Understanding of: Messaging platform security (TLS, certificates, channel auth, encryption).
  • Vulnerability remediation and risk mitigation in production systems.
  • Excellent troubleshooting skills in high-pressure, real-time environments (e.g., message backlog, latency spikes, connection failures.
  • Bachelor’s degree or foreign equivalent required from an accredited institution. Will also consider three years of progressive experience in the specialty in lieu of every year of education.

Nice To Haves

  • Clear, concise communication with technical and non technical stakeholders.
  • Ability to work effectively across engineering, infrastructure, security, and application teams

Responsibilities

  • Collaborate with internal and client teams to resolve complex incidents, conduct root cause analyses, and document findings with preventive recommendations
  • Participate in evaluation of client IT infrastructure, prepare actionable assessment reports, and support due diligence to document infrastructure maturity and improvement opportunities
  • Contribute to the design of scalable, cost-effective IT infrastructure solutions, review reusable components, and develop technical documentation for deployed systems
  • Align release schedules and environment readiness, execute deployments as per protocols, perform post-deployment testing, and manage version control to track changes
  • Co-ordinate maintenance schedules, emergency fixes, and technology upgrades while ensuring uninterrupted integration into existing systems and processes
  • Facilitate performance data analysis across systems, coordinate insights on system behavior, and support capacity planning to optimize performance
  • Conduct security checks, recovery drills, and compliance audits, implement security measures, and coordinate continuity plans to maintain adherence to standards
  • Gather feedback to identify automation opportunities, analyze existing infrastructure processes, and propose enhancements for efficiency gains
  • Act as liaison with onsite, offshore, and vendor teams to document project requirements, ensuring effective collaboration
  • Develop a centralized repository of technical and procedural knowledge, leveraging insights from other projects to drive efficiency and retain organizational expertise

Benefits

  • Medical/Dental/Vision/Life Insurance
  • Long-term/Short-term Disability
  • Health and Dependent Care Reimbursement Accounts
  • Insurance (Accident, Critical Illness , Hospital Indemnity, Legal)
  • 401(k) plan and contributions dependent on salary level
  • Paid holidays plus Paid Time Off
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service