Infrastructure Services Engineer (Hybrid Eligible)

Oak Ridge National LaboratoryOak Ridge, TN
Hybrid

About The Position

We are seeking an Infrastructure Services Engineer who will focus on specializing in monitoring and observability. This position resides in the Infrastructure Operations Center (IOC) in the Digital Services Infrastructure & Operations division of the Information Technology Services Directorate, at Oak Ridge National Laboratory (ORNL). As part of our team, you will design, operate, and continuously improve monitoring solutions across on-premises, cloud, and containerized environments. The IOC provides 24/7/365 monitoring and operational support for ORNL’s enterprise infrastructure and business-essential systems and services.

Requirements

  • BS degree in information technology or a related technical field and 2 years of relevant experience.
  • Experience operating, administering, or engineering enterprise monitoring platforms for infrastructure, applications, networks, or cloud environments.
  • Experience supporting enterprise Windows and Linux server environments, including performance analysis and troubleshooting.
  • Experience developing automated solutions using PowerShell, Python, or similar scripting tools.
  • Working knowledge of cloud infrastructure, container platforms, orchestration technologies, virtualization, and virtual-machine lifecycle operations.
  • Understanding of networking fundamentals, system performance indicators, telemetry, and diagnostic methodologies.

Nice To Haves

  • Strong analytical and problem-solving skills, including the ability to use operational data to identify issues and recommend improvements.
  • Strong written and verbal communication, customer service, collaboration, and technical documentation skills.
  • Ability to prioritize responsibilities and balance project work, operational support, and incident response in a fast-paced environment.
  • Demonstrated experience automating repetitive work or improving technical and operational processes.
  • Experience engineering and administering one or more enterprise-scale monitoring platforms, such as Prometheus, Grafana, Elastic, SolarWinds, or Dynatrace.
  • Experience with observability concepts and technologies, including metrics, logs, traces, baselining, synthetic monitoring, and service-level objectives.
  • Experience with anomaly detection, predictive analytics, AIOps, or automated remediation.
  • Knowledge of automation and infrastructure-as-code frameworks, such as Ansible, Terraform, or Azure Automation.
  • Experience using version-control systems to maintain scripts, configurations, dashboards, or infrastructure code.
  • Experience with virtualized or clustered compute environments, including performance tuning and lifecycle automation.
  • Familiarity with enterprise storage technologies, including direct-attached, SAN, and object storage, and their monitoring requirements.
  • Knowledge of enterprise server, storage, network hardware, and platform-level instrumentation.
  • Experience with enterprise backup, patching, configuration, or lifecycle management practices.
  • Understanding of change management and controlled operational workflows.
  • Experience working in regulated, scientific, government, or similarly complex technical environments.
  • Motivated self-starter with the ability to work independently and participate creatively in collaborative teams across the laboratory.

Responsibilities

  • Design, implement, administer, and maintain enterprise monitoring and observability solutions across on-premises, cloud, and containerized environments.
  • Develop and optimize alerts, dashboards, reports, synthetic monitors, and telemetry pipelines to identify degradation early and accelerate incident triage and root-cause analysis.
  • Evaluate monitoring coverage, gaps, overlaps, and underused capabilities, and implement tools, integrations, and data sources that improve operational visibility.
  • Automate monitoring deployment, configuration, data collection, and remediation using PowerShell, Python, or other scripting and automation tools.
  • Evaluate and apply AI-assisted capabilities for anomaly detection, predictive analytics, and operational efficiency.
  • Collaborate with infrastructure, network, application, security, and other technical teams to improve system health and observability.
  • Support incident and problem management by providing relevant metrics, logs, performance trends, and historical analysis.
  • Work with vendors and internal subject matter experts to troubleshoot monitoring agents, collectors, integrations, and platform components.
  • Establish and maintain monitoring standards, topology diagrams, technical documentation, runbooks, and team procedures.
  • Support patching, backup, upgrade, and lifecycle activities for monitoring platforms and related infrastructure components.
  • Continuously improve alert thresholds, dashboards, data quality, automated remediations, and monitoring workflows to reduce noise and repetitive operational work.
  • Provide escalated support for monitoring-related issues and participate in an on-call or planned maintenance rotation as required.
  • Deliver ORNL’s mission by aligning behaviors, priorities, and interactions with our core values of Impact, Integrity, Teamwork, Safety, and Service. Promote equal opportunity by fostering a respectful workplace – in how we treat one another, work together, and measure success.

Benefits

  • Prescription Drug Plan
  • Dental Plan
  • Vision Plan
  • 401(k) Retirement Plan
  • Contributory Pension Plan
  • Life Insurance
  • Disability Benefits
  • Generous Vacation and Holidays
  • Parental Leave
  • Legal Insurance with Identity Theft Protection
  • Employee Assistance Plan
  • Flexible Spending Accounts
  • Health Savings Accounts
  • Wellness Programs
  • Educational Assistance
  • Relocation Assistance
  • Employee Discounts
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service