Production Engineer (IC4)

Ontrac SolutionsChicago, IL
$75 - $85Hybrid

About The Position

Ontrac Solutions is seeking a high-aptitude Production Engineer to support a large-scale enterprise OS modernization and infrastructure hardening program for one of our enterprise clients. This role is built for an engineer with real software engineering foundations — specifically Python — who has since moved into infrastructure and is comfortable working across OS modernizations (RHEL7 → EL8/EL9), packaging migrations (Chef → CINC), and CI/CD hardening, while simultaneously executing hands-on runbooks, automation, and service onboarding. This is a genuine hybrid role: roughly half build — Python tooling, RPM packaging, pipeline and rollback work — and half operate — runbooks, service onboarding, and Tier-2 break/fix. You will work directly with the client's SRE organization, internal engineering teams, and customer stakeholders from initial definition through final delivery. Candidates who are pure application developers with no Linux fleet exposure, and candidates who are pure operations with no real software engineering behind them, will not clear screening.

Requirements

  • Software engineering foundation: 3+ years of professional software engineering experience, with strong hands-on Python. You have written and maintained code other engineers depended on — modules, packaging, tests, code review, and version control, not just single-file scripts.
  • Enterprise OS modernization: Hands-on experience with enterprise Linux at fleet scale and with large-scale OS upgrade programs — RHEL7 → EL8/EL9 or an equivalent major-version migration you executed rather than observed.
  • Packaging migrations: Experience building RPM packages to replace legacy configuration, and with large-scale packaging or configuration-management migrations such as Chef → CINC.
  • Independent bug ownership: Able to triage, own, and resolve bugs end to end without hand-holding — reproduce, isolate, fix, test, and ship.
  • CI/CD and release safety: Ability to harden CI/CD pipelines, observability frameworks, and rollout/rollback mechanisms specifically tailored for legacy-to-modern infrastructure transitions.
  • Tier-2 operational support: Willingness and experience to partner closely with an SRE team providing "follow-the-sun" tier-2 support, including hands-on incident response and break/fix operations on existing platforms.
  • Service onboarding: Experience onboarding services to newly established monitoring and logging stacks.
  • Automation and documentation: A demonstrated habit of automating repetitive operations and documenting technical procedures for others to run.
  • Must be located in the United States and authorized to work in the US.

Nice To Haves

  • Proven experience planning and executing logging and monitoring tool rollouts end to end — not only operating a stack someone else stood up.
  • Experience supporting a team through cloud cutovers and component migrations to cloud environments.
  • Perl scripting experience — legacy tooling in this environment is Perl, and the ability to read and safely modify it is a real advantage.
  • Hands-on exposure to modern observability tooling — Chronosphere, Prometheus, or Grafana — and to Splunk integrations.
  • Provisioning and image work: Kickstart/PXE, golden images, or repository and mirror management.

Responsibilities

  • Drive the technical transition of legacy systems to modern enterprise Linux environments, including RHEL7 → EL8/EL9 upgrade paths.
  • Build and maintain RPM packages to replace legacy configuration, and carry the packaging migration from Chef to CINC.
  • Develop and execute automated runbooks that make the migration repeatable rather than manual.
  • Triage, own, and resolve migration bugs independently, from first report through verified fix.
  • Transition monitoring infrastructure to a modern stack — Chronosphere, Prometheus, and Grafana — and manage Splunk integrations.
  • Onboard services to the newly established monitoring and logging stacks, including metrics, dashboards, and alert rules.
  • Harden observability frameworks alongside the pipelines they instrument, so regressions surface before users find them.
  • Harden CI/CD pipelines for legacy-to-modern infrastructure transitions, including build, test, and package promotion stages.
  • Design and maintain rollout and rollback mechanisms that make large-fleet changes reversible.
  • Automate repetitive operational work and replace manual runbook steps with tested, reviewed code.
  • Provide Tier-2 operational support and incident response under a follow-the-sun model, in close partnership with the client's SRE team.
  • Perform hands-on break/fix operations on existing platforms while the modernization proceeds in parallel.
  • Assist application developers with architectural support and troubleshooting during cloud migration phases.
  • Author and maintain technical documentation — runbooks, migration procedures, and package and pipeline ownership notes.

Benefits

  • Ontrac reimburses the full cost of your verification if you are hired.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service