Staff Software Engineer, Ray Data

AnyscaleSan Francisco, CA
$240,000 - $270,000

About The Position

Anyscale is commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Anyscale is building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. The Ray Data team develops and maintains Ray Data, a Python-native data processing engine and a one-stop shop for all AI data processing needs. Ray Data provides performant, first-class integration with cutting-edge AI frameworks using both multimodal and structured data. The team is looking for exceptional engineers to build, optimize, and scale Ray Data for increasingly complex AI workloads, including multimodal data processing and large-scale batch inference.

Requirements

  • 6+ years of experience building production-grade software, infrastructure, or developer-facing systems, with strong Python engineering experience.
  • 6+ years of experience personally owning core architectural decisions within a distributed data or compute engine, rather than primarily operating or using a platform someone else designed.
  • Deep experience with distributed systems internals, such as scheduling, fault tolerance, data partitioning, distributed execution, performance optimization, or database and query engine internals.
  • A track record of reasoning through system-level tradeoffs and defending architectural decisions, such as batch vs. streaming, static vs. dynamic resource allocation, or consistency vs. availability.
  • Passion for solving the unsolved problems in large-scale AI infrastructure and building systems that enable the next generation of AI applications.

Responsibilities

  • Design, build, and improve the core systems that power Ray Data, with a focus on performance, scalability, and reliability.
  • Design and optimize distributed execution across different stages of data pipelines in heterogeneous environments.
  • Build data loading and processing solutions for production training and inference workloads.
  • Solve challenging problems in distributed execution, scheduling, resource management, data partitioning, fault tolerance, and performance optimization.
  • Make system-level architectural decisions and reason through tradeoffs in areas such as resource allocation, execution models, batch vs. streaming workloads, and consistency and availability.
  • Work with customers and new-age AI-native companies to understand and solve challenges in scaling their AI workloads.

Benefits

  • Competitive salary and equity
  • Health/dental/vision coverage (many plans up to 99% employer-covered)
  • Flexible time off
  • Paid parental leave
  • Mental health support
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service