Senior Infrastructure Support Engineer

ThoughtworksChicago, IL
Remote

About The Position

As a senior Infrastructure Support Engineer, you play a vital role in maintaining technical excellence and operational efficiency, with a primary focus on cloud environments. You'll help clients through the transition to agile, value-focused practices, emphasizing shared responsibility and continuous improvement. You will monitor infrastructure performance, respond to incidents promptly and maintain resources in line with modern standards, incorporating sustainable practices.

Requirements

  • Hands-on experience in using CI/CD tools such as Jenkins, CircleCI or Gitlab for executing deployments.
  • Knowledge of Infrastructure as Code (IAC) tech stacks such as Terraform, Ansible, ARM or Cloudformation to provision and manage infrastructure.
  • Working experience in using observability tools for logging, monitoring, tracing and alerting, e.g.: Datadog/PrometheusGrafana, ELK/EFK/Splunk.
  • Experience in supporting at least one public cloud, e.g.: AWS, Azure or GCP.
  • Hands-on experience executing most common operations in managing workloads on any container ecosystem tech stacks. e.g.: Docker, Kubernetes, Openshift, etc.
  • Understand system performance tuning and scaling to handle common heavy load scenarios along with concepts of highly available systems and basics of disaster recovery solutions, and are familiar with failover, backup and recovery concepts.
  • Experience operating a Linux OS such as RHEL or a Debian-Based OS and are familiar with most common Linux OS operations and commands, reading and tweaking Bash scripts and managing runtime environment configurations such as Env Vars, Logs, etc.
  • Experience supporting backend storage solutions such as SQL and NoSQL databases, e.g.: Postgres and MongoDB, and caching solutions such as Redis and Memcached.
  • Experience in networking configuration and security, and are familiar with common networking setup and security practices, e.g.: loading, balancing, proxies, transport layer security (TLS) and certificate management, and an understanding of standard network protocols and configurations.
  • Good understanding of fundamental concepts of APIs such as request, response, headers, authentication, JSON payloads, etc.
  • Strong communication and articulation skills, proficient in English and able to confidently hold a Q&A discussion with senior stakeholders.
  • People skills with an emphasis on close collaboration with multiple, cross-functional teams from the client side or Thoughtworks.
  • Ability to work under pressure and with composure during production incidents.
  • Strong analysis, deduction and reasoning skills, with the ability to identify patterns in data and draw conclusions.
  • Strong drive and ownership to sign up and deliver work when called upon without being too concerned with role boundaries.
  • Willing to be part of a rotation- and need-based 24x7 available team.

Responsibilities

  • Keep a vigilant eye on the operations of shipped products and services following the agreed upon “Eyes on glass/Follow the sun” engagement models.
  • Monitor product/service operations against key performance indicators defined by the business and take necessary actions in response to detected deviations.
  • Define and document the appropriate responses to various kinds of incident scenarios in collaboration with the Service Reliability Engineering (SRE) team and client stakeholders, and prepare runbooks.
  • Reduce the human effort in day-to-day operations by automating operations, using the latest tech stacks befitting the task and improving the overall efficiency of the entire team as time progresses.
  • Be the first responder to incidents in production/other high-value environments and execute the appropriate response as established by runbooks or based on your judgment of incidents.
  • Initiate or establish communication to the support teams across all service functions, setting up war rooms for incident response, coordinating with tech leads, SRE leads and development teams to resolve incidents, as necessary.
  • Prepare incident root cause analysis (RCA) and postmortem reports, explaining analyses and outlining preventive measures to clients; Collaborating with SRE, development teams or independently, ensure clear communication and proactive steps for future incident prevention.
  • Implement service/product reliability improvement in collaboration with service reliability engineers by writing infrastructure/observability configuration code.

Benefits

  • Learning & Development programs
  • Interactive tools
  • Teammates who want to help you grow
  • Empowering employees in their career journeys
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service