About The Position

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference. In AI infrastructure organization, simplifying large hardware deployments with push button, single pane of glass for observability/monitoring and software capabilities for build-in resiliency are some of the key focus areas. As senior software development engineer in Test, we are looking for a candidate who can make a big impact on how we test and validate thousands of nodes in large deployments to ensure the cluster is 99.999% reliable.

Requirements

  • Bachelor's or master's degree in engineering in computer science, electrical, AI, data science or related field.
  • 5+ years of experience in testing one of areas like enterprise software, distributed systems, datacenter hardware and software.
  • Strong coding skills in one of the programming languages like python, golang and C/C++.
  • Strong debugging skills to debug issues in large distributed systems, hardware, and software.
  • Experience with debugging tools like pdb, gdb, strace and network monitors.
  • Strong understanding of operating systems internals like memory management, file system working, security and performance.
  • Strong understanding of datacenter layout, device performance characteristics like Servers, Memory, BIOS, PCIe, networking and storage.
  • Experience with cloud technologies like AWS, kubernetes and dockers.

Nice To Haves

  • Monitoring tools like grafana, prometheus is huge plus.
  • Understanding and experience of ML model training and inference is a huge plus.
  • Understand of ML hardware accelerators like GPU, custom accelerator ASIC is a huge plus.

Responsibilities

  • Innovate and execute tests on cutting edge AI infrastructure.
  • Define optimized test strategies and methodologies.
  • Adapt to new technologies and bring expertise.
  • Deep understanding of how large-scale distributed ML training and inference works.
  • Break down large distributed systems challenges into smaller components that can be unit tested.
  • Automate all cluster features in areas of high availability, failure scenarios, performance, stress and security.
  • Champion cluster security, reliability for uptime of 99.9999% and ease of use with observability.
  • Test all components of AI cluster including but not limited to cluster software involving kubernetes, prometheus and grafana.
  • Test cluster hardware components like ML wafer scale accelerators, CPU runtime nodes, High speed swarmx interconnect, High speed data transfer of weights through memoryx interconnect.

Benefits

  • Job stability with startup vitality.
  • Simple, non-corporate work culture that respects individual beliefs.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service