Senior DevOps/SRE Engineer

SEIChicago, IL
$140,000 - $170,000Onsite

About The Position

We are looking for a Senior Site Reliability Engineer to work as part of a lean, product‑focused engineering organization. This role is about building and operating reliable cloud‑based systems by writing code, automating infrastructure and delivery workflows, and reducing friction for developers and users. You will work closely with product and application engineers to design, deploy, and operate systems with clear ownership and practical engineering judgment. We expect you to use modern tooling, including AI‑assisted tools where appropriate, to speed up automation, troubleshooting, and operations while remaining accountable for correctness, security, and reliability. This role favors simple, effective solutions, hands‑on ownership, and continuous improvement within small Agile teams.

Requirements

  • BA/BS, in a related technical field; or the equivalent in education and work experience
  • 8+ years of experience in DevOps, SRE, platform engineering, or similar roles supporting application teams running production services
  • Strong CI/CD experience (Jenkins and Git-based workflows preferred), including building secure, reliable pipelines and enabling teams to ship safely
  • Experience implementing and operating observability platforms (logging/metrics/alerting); Elastic Stack/OpenSearch experience is a plus
  • Hands-on, demonstrable experience designing and operating AWS environments, and enabling application teams to adopt AWS correctly (networking, IAM, security, reliability, and cost awareness)
  • Infrastructure as Code experience (Terraform preferred; CloudFormation acceptable), including building reusable modules/patterns and managing changes through review and automation
  • Experience supporting CI/CD builds and deployment patterns for common application stacks (for example Java, NodeJS, and .NET)
  • Experience scripting in Bash, Python, or PowerShell
  • Experience working on large scale cloud-based web applications

Nice To Haves

  • Ability to clearly communicate both verbally and in writing with client and team members, including experience documenting and presenting findings
  • Excellent analytical skills, organizational abilities, and problem-solving skills
  • Familiarity with AI agentic development tools (e.g., Claude Code, GitHub Copilot, Windsurf) and practical experience applying them to infrastructure and operations workflows
  • Self-starter who works efficiently in a fast-paced environment with changing priorities and a geographically distributed team
  • Ability to think creatively and seek optimum solutions
  • Ability to grasp loosely defined concepts and transform them into tangible results and key deliverables
  • Diagnostic skills with the ability to analyze technical, business and financial issues and options
  • Ability to infer from previous examples, willingness to understand how an application is put together
  • Action-oriented, with the ability to quickly deal with change
  • Someone who will embody our SEI Values of courage, integrity, collaboration, inclusion, connection and fun.

Responsibilities

  • Design, build, and operate cloud infrastructure for critical production and non‑production applications with reliability and simplicity as primary goals
  • Architect and evolve multi-account AWS foundations (organizations, accounts, IAM boundaries, guardrails, and environment separation) to enable secure, scalable delivery
  • Design and operate cloud networking architecture (VPCs, routing, segmentation, ingress/egress, connectivity patterns) to support reliability, security, and compliance requirements
  • Treat reliability, security, and compliance as first‑class design concerns throughout the system lifecycle
  • Build tooling and automation that reduces errors, shortens recovery time, and improves day‑to‑day operations
  • Implement monitoring, logging, and alerting that make system behavior observable and actionable
  • Use AI‑assisted tools to accelerate infrastructure delivery, automation, troubleshooting, and root‑cause analysis, applying engineering judgment to validate outcomes
  • Implement reliability guardrails for releases (progressive delivery, safe rollbacks, change risk controls) and provide production support during deployments.
  • Participate in incident response, perform root cause analysis, and drive durable improvements that prevent recurrence
  • Work closely with application engineers to co‑own system design, operation, and continuous improvement
  • Maintain clear, lightweight documentation that supports shared ownership and effective on‑call operations

Benefits

  • comprehensive care for your physical and mental well-being
  • a strong retirement plan
  • tuition reimbursement
  • support for working parents
  • flexible Paid Time Off (PTO)
  • healthcare (medical, dental, vision, prescription, wellness, EAP, FSA)
  • life and disability insurance (premiums paid for base coverage)
  • 401(k) match
  • education assistance
  • commuter benefits
  • up to 11 paid holidays/year
  • 21 days PTO/year pro-rated for new hires which increases over time
  • paid parental leave
  • back-up childcare arrangements
  • paid volunteer days
  • a discounted stock purchase plan
  • investment options
  • access to thriving employee networks
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service