AI Infra Staff Researcher

Lenovo•Morrisville, NC
•Hybrid

About The Position

The Staff Researcher in AI Compute and Data Infrastructure will conduct applied research and hands-on development for intelligent, efficient, and resilient Hybrid AI systems. This position works across AI algorithms, computer systems, distributed computing, and data infrastructure to address performance, scalability, reliability, and energy-efficiency challenges. The successful candidate will independently own substantial research and development workstreams, build production-quality software, characterize AI workloads, diagnose infrastructure issues, and develop cross-layer optimization technologies spanning GPUs and other accelerators, CPUs, memory, storage, networking, system software, data pipelines, and AI frameworks.

Requirements

  • Bachelor's degree in computer science, computer engineering, artificial intelligence, electrical engineering, applied mathematics, or a related field, or equivalent practical experience.
  • Three or more years of relevant experience in AI compute and data infrastructure, machine learning systems, distributed systems, data platforms, performance engineering, reliability engineering, or advanced software development.
  • Strong programming skills in Python, C++, Java, Go, Rust, Scala, or a comparable language.
  • Demonstrated ability to design, implement, test, debug, profile, and optimize reliable software systems.
  • Experience with system profiling, telemetry analytics, observability, performance diagnosis, or failure analysis.
  • Technical expertise in at least two of the following areas: Machine learning or deep learning, GPU or accelerator optimization, Distributed training or inference systems, Large-scale data processing, Hardware/software co-optimization, Time-series modeling or signal processing, Infrastructure reliability and fault tolerance, Causal inference, Knowledge graphs or graph machine learning.
  • Strong analytical, experimental, communication, and cross-functional collaboration skills.
  • Must be a US citizen or US national; US permanent residents or candidates requiring sponsorship cannot be considered.

Nice To Haves

  • Experience with PyTorch, TensorFlow, JAX, CUDA, ROCm, Spark, Flink, Ray, Kafka, Kubernetes, or related technologies.
  • Experience with cloud, edge, on-premises, or hybrid AI infrastructure.
  • Experience delivering research prototypes or advanced software into production environments.
  • Publications, patents, open-source contributions, or demonstrated product impact.

Responsibilities

  • Research and develop technologies for AI compute and data infrastructure, distributed AI systems, and intelligent infrastructure management.
  • Design and implement production-quality software, system components, services, APIs, diagnostic tools, and scalable data-processing pipelines.
  • Characterize AI training, inference, and data-processing workloads using profiling, tracing, benchmarking, telemetry, logs, and hardware performance counters.
  • Diagnose performance bottlenecks and reliability issues across GPUs, accelerators, CPUs, memory hierarchy, storage, networking, operating systems, runtimes, and AI frameworks.
  • Develop hardware/software co-optimization solutions for GPU utilization, workload scheduling, resource allocation, memory and cache management, communication, data movement, storage access, and model execution.
  • Optimize large-scale data ingestion, preprocessing, transformation, storage, retrieval, and delivery for AI training, inference, and analytics workloads.
  • Build intelligent infrastructure diagnostics for anomaly detection, root-cause analysis, performance regression detection, system health assessment, capacity forecasting, and predictive maintenance.
  • Develop fault-tolerance and resilience mechanisms, including fault detection and isolation, checkpointing, recovery, retry, failover, graceful degradation, and automated remediation.
  • Apply machine learning and deep learning to workload modeling, performance prediction, resource optimization, failure prediction, and operational decision-making.
  • Apply time-series analysis and signal processing to infrastructure telemetry, event detection, change-point detection, workload forecasting, and system health monitoring.
  • Apply causal inference to performance attribution, root-cause analysis, intervention evaluation, and infrastructure optimization.
  • Develop knowledge graphs to model infrastructure topology, hardware/software dependencies, workloads, operational events, and failure relationships.
  • Optimize systems for throughput, latency, scalability, availability, resource utilization, energy consumption, and total cost of ownership.
  • Collaborate with hardware, systems, software, architecture, and product teams to transition research technologies into Enterprise AI and Personal AI products.
  • Contribute to patents, invention disclosures, technical publications, internal reports, and reusable software assets.
  • Provide technical guidance and mentorship to junior researchers and engineers.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service