Cyberinfrastructure Engineer

University of Chicago•Hyde Park, IL
•Onsite

About The Position

The MANIAC Lab, located within the Enrico Fermi Institute of the Physical Sciences Division at the University of Chicago, builds and operates advanced cyberinfrastructure for scientific instruments investigating the fundamental mysteries of nature. The Lab operates data-intensive high-throughput computing facilities that are part of a global computing grid used to reconstruct and analyze particle collisions recorded by the ATLAS detector at the Large Hadron Collider (LHC) at CERN in Geneva, Switzerland. The Lab also serves as a gateway for U.S. researchers to national cyberinfrastructure resources through the Open Science Grid Consortium (OSG) and the Institute for Research and Innovation in Software for High Energy Physics (IRIS-HEP). The Lab operates the NSF Shared Tier 3 Analysis Facility. The facility supports ATLAS physicists analyzing the complete LHC Run 2 and Run 3 datasets as the collaboration prepares for the High-Luminosity LHC (HL-LHC), which begins with Run 4 around 2030. It provides computing infrastructure and services for the South Pole Telescope (SPT-3G) and operates a worldwide distributed data network for the XENON dark matter search experiment at the Gran Sasso National Laboratory in Italy. The Lab is increasingly focused on AI-enabled research computing. Its GPU platforms support machine learning training and inference, including models served through NVIDIA Triton and large language models hosted on facility hardware. The Lab has also built an agentic AI platform based on the Model Context Protocol (MCP). It gives researchers' AI assistants secure, identity-brokered access to facility services such as HTCondor, Kubernetes, Rucio, and ServiceX. The same platform supports AI-assisted facility operations, where sandboxed agents analyze operational metrics and propose corrective actions for human review. The Lab provides the Scalable Systems Laboratory (SSL), a software testing and integration platform for IRIS-HEP. IRIS-HEP develops software and computing solutions for the HL-LHC era. Through the SSL Deployment Factory, the Lab is packaging the Kubernetes infrastructure it already operates into tested, versioned bundles that other facilities can deploy. These bundles are bootstrapped with Kubespray and managed through GitOps with Flux, and cover services such as JupyterHub, BinderHub, ServiceX, and the MCP platform. The goal is to shorten a process that currently takes months of manual work at each site.

Requirements

  • Minimum requirements include a college or university degree in related field.
  • Minimum requirements include knowledge and skills developed through 2-5 years of work experience in a related job discipline.
  • Strong oral and written communication skills.
  • Initiative and capacity for teamwork and creativity.
  • Ability to effectively communicate and collaborate with team members , supervisors, and researchers.
  • Ability to manage complex technical details and switch between projects.
  • Ability to work independently with minimal supervision, take ownership of issues through resolution, and keep the team informed.
  • Comfortable collaborating in a distributed team through weekly Zoom meetings and daily Slack communication.

Nice To Haves

  • Bachelor’s degree in Computer Science, Computer Engineering, Physics or related field.
  • Strong experience managing Linux operating systems.
  • Configuration management and build systems for large numbers of computers using tools such as Puppet, Chef and Ansible.
  • Experience operating distributed storage at scale, particularly Ceph (including Rook on Kubernetes).
  • Hands-on data center hardware experience, including server diagnostics, component replacement, and working with vendor support (e.g., Dell iDRAC/warranty processes).
  • Experience with monitoring and alerting tools such as Nagios, Prometheus, Grafana, and Alertmanager.
  • Unix/Linux operating systems administration tools and shell scripts.
  • Distributed storage systems such as Ceph; local file systems such as ZFS.
  • Git version control, scripting (Bash, Python) and automation.
  • Knowledge and expertise in technologies such as TCP/IP and related protocols; networked file systems, including NFS.
  • Knowledge or experience with batch scheduling systems such as Slurm or HTCondor.
  • Knowledge of container technologies such as Docker, Kubernetes, Helm, OpenShift/OKD, OpenStack.
  • Familiarity with identity and access management (e.g., Keycloak, OAuth/OIDC).
  • Knowledge or experience with Spark, Dask or Ray.
  • Knowledge of GitOps and cluster lifecycle tooling such as Flux, Argo CD, or Kubespray.
  • Experience with GPU servers, including NVIDIA drivers, CUDA, and GPU scheduling in Kubernetes or HTCondor.
  • Familiarity with AI/ML infrastructure, such as model serving, LLM-based tooling, or MCP, is a plus.

Responsibilities

  • Provides systems administrative services for Linux compute clusters (CPU and GPU), storage systems, and related support servers. These systems support ATLAS production and analysis at the Midwest Tier 2 Center and the Analysis Facility.
  • Operates, upgrades, and scales Ceph distributed storage (including Rook-managed Ceph on Kubernetes). The goal is to meet growing HL-LHC capacity and throughput requirements.
  • Performs on-site hardware work in campus data centers, including installing, retrofitting, and replacing servers, storage, and GPUs.
  • Manages vendor support cases and spare-parts inventory.
  • Maintains monitoring, dashboards, and alerting (e.g., Prometheus, Grafana) for storage, compute, network, and hardware health.
  • Applies operating system and service security patches and vulnerability mitigations in coordination with University security requirements.
  • Performs network diagnostics, throughput measurement, and analysis for both LAN and WAN.
  • Supports the facility's participation in WLCG Data Challenges and HL-LHC readiness scale tests.
  • Participates in the team's operations support rotation and maintains documentation and operational runbooks.
  • Deploys and operates the infrastructure for agent-driven physics analysis. This includes the Lab's MCP gateway, the MCP servers it brokers access to, and the credential services behind them. Together, these let researchers' AI assistants securely use facility resources such as HTCondor, Kubernetes, Rucio, and ServiceX.
  • Supports IRIS-HEP integration challenges and demonstrations of end-to-end agentic analysis workflows. The work spans dataset discovery, batch processing, analysis, and inference run through AI agents.
  • Operates GPU platforms for machine learning and AI workloads, including locally hosted large language models and sandboxed agent environments.
  • Helps develop and operate AI-assisted operations agents that monitor HTCondor, Kubernetes, Ceph, and related services. These agents propose corrective actions under human review.
  • Deploys and operates services using container-based approaches (Docker, Kubernetes, Helm) and GitOps workflows (e.g., Flux).
  • Contributes reusable deployment bundles to the SSL Deployment Factory so other facilities can adopt these services.
  • Learns new distributed computing, AI/ML, and infrastructure-as-a-service technologies.
  • Maintains complex system and network administration functions.
  • Works with moderate guidance to administer simple systems and assists in the administration of larger systems.
  • Installs, configures, and maintains operating system workstations and servers.
  • Performs software installations and upgrades to operating systems and layered software packages.
  • Monitors and tunes the system to achieve optimum performance levels, acquiring higher-level skills in the process.
  • Performs other related work as needed.

Benefits

  • health
  • retirement
  • paid time off
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service