Google-posted about 4 hours ago
Full-time • Mid Level
Sunnyvale, CA

The Core ML team contributes to frameworks and compilers that support the Google Cloud Platform (GCP) Cloud Tensor Processing Unit (TPU) service and related Machine Learning (ML) models and frameworks. It provides ML infrastructure customers with large-scale, cloud-based access to Google’s first-party ML supercomputers (TPUs and TPU Pods) to run training and inference workloads using PyTorch and JAX. In this role, you will be responsible for the PyTorch ML framework, processes, ecosystem, and model performance, as well as engagements with customers who take advantage of Google’s TPUs to achieve massive scale and speed in their ML workloads.The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide. We're the driving force behind Google's groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.

  • Work on AI framework development to enable PyTorch models to run on Google Cloud's TPUs and GPUs and tune for peak performance.
  • Provide comprehensive support for ML frameworks and compilers on Cloud TPUs and Graphics Processing Units (GPUs), enabling the training and deployment of the most advanced machine learning models, managing innovation and breakthroughs.
  • Enable PyTorch models for generative models, computer vision (image recognition, object detection, image generation), machine translation, language modeling, rankings and recommendations, speech recognition, etc.
  • Collaborate with other Google teams and leading researchers across the industry to continuously bring ML capabilities to our PyTorch in Cloud offering.
  • Design, develop, test, deploy, maintain, and improve software while contributing to open-source software development.
  • Bachelor’s degree or equivalent practical experience.
  • 5 years of experience with ML design and ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).
  • 5 years of experience in software development.
  • 5 years of experience testing, and launching software products, and 3 years of experience with software design and architecture.
  • Master’s degree or PhD in Engineering, Computer Science, or a related technical field.
  • 8 years of experience with data structures/algorithms.
  • 3 years of experience in a technical leadership role leading project teams and setting technical direction.
  • 3 years of experience working in an organization involving cross-functional, or cross-business projects.
  • Experience with compilers or ML frameworks.
© 2024 Teal Labs, Inc
Privacy PolicyTerms of Service