Storage & Virtualization Engineer

Advanced Micro Devices, Inc•Austin, TX
•Remote

About The Position

We are seeking a Storage & Virtualization Engineer to provide L2 technical and incident-management support for storage, virtualization, and related compute dependencies used by AMD Fleet Services. This offshore role is an escalation point for incidents that exceed L1 runbooks and is responsible for advanced diagnosis, coordinated restoration, authorized remediation, recovery validation, and complete escalation to the accountable platform or engineering owner. The engineer will improve the reliability, availability, and operational readiness of storage and virtual infrastructure while helping the IOQ Modern NOC convert recurring issues into monitoring, automation, SOP, runbook, capacity, and resiliency improvements.

Requirements

  • Hands-on infrastructure engineer with strong storage, virtualization, Linux, and systems troubleshooting fundamentals.
  • Ability to use evidence to isolate failures across hosts, hypervisors, virtual machines, networks, and storage paths.
  • Ability to balance restoration speed with change discipline.
  • Ability to communicate ownership, impact, and risk clearly during high-priority incidents.
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field preferred; equivalent relevant experience considered.
  • Storage, virtualization, Linux, cloud, automation, or IT service-management certification is beneficial.

Nice To Haves

  • Experience supporting enterprise storage and virtualization in a production data center, GPU, HPC, private-cloud, or large-scale infrastructure environment.
  • Hands-on experience with storage platforms such as Weka, NetApp, Pure, Ceph, Lustre, NFS, block storage, or comparable technologies.
  • Hands-on experience with virtualization platforms such as Xen, VMware, KVM, or comparable technologies.
  • Strong understanding of storage protocols, multipathing, volumes, filesystems, snapshots, replication, performance, capacity, and data-protection concepts.
  • Strong Linux administration and troubleshooting across services, filesystems, devices, networking, permissions, performance, and logs.
  • Experience with server hardware, hypervisors, clustering, virtual-machine lifecycle operations, and infrastructure management tooling.
  • Automation experience with Python, shell, Ansible, REST APIs, Git, CI/CD, or configuration-management systems.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Zabbix, or comparable platforms.
  • Experience with Jira or comparable ticketing systems, major-incident response, change control, postmortems, and problem management.
  • Ability to work independently during assigned offshore coverage and provide complete handoffs to United States-based IOQ and engineering teams.

Responsibilities

  • Provide L2 incident management and technical support for in-scope storage, virtualization, hypervisor, virtual-machine, and related compute services used by AMD Fleet Services.
  • Own technical investigation of assigned incidents from L1 escalation through diagnosis, mitigation, recovery validation, documentation, and handoff or closure.
  • Assess incident impact, priority, affected assets, dependencies, recovery options, risks, and required owners; provide concise updates during incident bridges and follow-the-sun handoffs.
  • Validate L1 evidence, isolate storage, virtualization, host, guest, network, operating-system, capacity, or performance failure domains, and execute approved recovery actions.
  • Investigate storage availability, latency, throughput, capacity, pathing, mount, protocol, volume, snapshot, replication, and data-access symptoms using platform telemetry and system evidence.
  • Support operational troubleshooting for storage platforms such as Weka, NetApp, Pure, Ceph, Lustre, NFS, block storage, or comparable technologies within assigned scope.
  • Troubleshoot virtualization platforms such as Xen, VMware, KVM, or comparable environments, including host health, guest state, resource allocation, datastore access, migration, and recovery validation.
  • Diagnose Linux host and virtual-machine issues involving services, filesystems, multipathing, device state, permissions, resource contention, performance, networking, and logs.
  • Correlate storage and virtualization telemetry with compute, network, identity, cluster, and application evidence to identify cross-domain dependencies and the accountable owner.
  • Coordinate restoration and escalation with storage engineering, platform engineering, compute, network, cloud, identity, security, facilities, and vendor support teams while maintaining clear IOQ scope and ownership boundaries.
  • Create complete escalation packages with impact, timeline, affected systems, evidence, actions attempted, change references, residual risk, and recommended next steps.
  • Participate in postmortems and problem-management reviews; identify recurring failure modes and track corrective actions to the accountable owner.
  • Develop and maintain SOPs, runbooks, validation checks, troubleshooting guides, and knowledge articles that enable safe and consistent L1 execution.
  • Automate repeatable diagnostics, evidence collection, capacity checks, configuration validation, health verification, and approved remediation using Python, shell, Ansible, APIs, or comparable tools.
  • Improve observability by defining actionable storage and virtualization telemetry, alerts, dashboards, service-health indicators, capacity thresholds, and diagnostic evidence requirements.
  • Maintain accurate Jira records, operational documentation, shift handoffs, and service-status updates throughout the incident lifecycle.

Benefits

  • AMD benefits at a glance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service