Senior Machine Learning Engineer

CloudflareSan Francisco, CA
Onsite

About The Position

You’ll help define how machine learning models run across Cloudflare’s global network, from frontier open LLMs and real-time voice models to customer-deployed models served on heterogeneous GPUs and next-generation accelerators. You’ll work with systems engineers, product teams, hardware partners, and AI/ML engineers to bring models into production with low latency, strong reliability, and efficient resource use. This role combines applied ML, inference optimization, evaluation, and production engineering, with a focus on benchmarking models, improving serving performance, validating quality, and building tooling that helps Cloudflare and its customers ship AI applications at Internet scale.

Requirements

  • Experience building, optimizing, and operating machine learning models in production environments.
  • Strong proficiency with Python and modern ML frameworks such as PyTorch, TensorFlow, JAX, or equivalent.
  • Hands-on experience with inference optimization techniques for large-scale models, including quantization, batching, caching, compilation, and serving runtime tuning.
  • Experience with large-scale inference serving frameworks or runtimes such as SGLang, vLLM, TensorRT-LLM, ONNX Runtime, Triton, llama.cpp, or similar.
  • Familiarity with LLMs, speech models, vision models, embeddings, multimodal models, retrieval-augmented generation, or other modern deep learning architectures.
  • Experience optimizing models for GPUs or specialized accelerators.
  • Strong understanding of production ML concerns, including evaluation, monitoring, model regressions, rollout safety, and reliability.
  • Ability to work across ML and systems boundaries, including familiarity with distributed systems, networking, or serverless platforms.
  • Track record of leading complex technical projects and mentoring other engineers.

Nice To Haves

  • Experience contributing to open source ML tooling, model serving frameworks, or inference runtimes.

Responsibilities

  • Develop, optimize, and productionize machine learning models for Cloudflare’s serverless inference platform, with a focus on performance, reliability, and model quality.
  • Build benchmarking and evaluation frameworks to measure latency, throughput, cost efficiency, and model behavior across LLMs, speech, vision, and other model families.
  • Improve inference performance through quantization, batching, caching, model compilation, runtime tuning, and accelerator-aware optimization.
  • Partner with systems engineers to integrate models into Cloudflare’s distributed inference infrastructure across a heterogeneous fleet of GPUs and next-generation accelerators.
  • Drive improvements to model deployment workflows, including validation, rollout safety, observability, regression testing, and operational readiness.
  • Collaborate with product and engineering teams to translate customer requirements into scalable ML capabilities for Workers AI.
  • Mentor engineers, contribute to technical direction, and raise the quality bar for production ML engineering practices across the team.

Benefits

  • We’re not just a highly ambitious, large-scale technology company. We’re a highly ambitious, large-scale technology company with a soul.
  • Fundamental to our mission to help build a better Internet is protecting the free and open Internet.
  • Project Galileo
  • Athenian Project
  • 1.1.1.1
  • We don’t store client IP addresses never, ever.
  • We will continue to abide by our privacy commitment and ensure that no user data is sold to advertisers or used to target consumers.
  • Cloudflare is proud to be an equal opportunity employer.
  • We are committed to providing equal employment opportunity for all people and place great value in both diversity and inclusiveness.
  • All qualified applicants will be considered for employment without regard to their, or any other person's, perceived or actual race, color, religion, sex, gender, gender identity, gender expression, sexual orientation, national origin, ancestry, citizenship, age, physical or mental disability, medical condition, family care status, or any other basis protected by law.
  • We are an AA/Veterans/Disabled Employer.
  • Cloudflare provides reasonable accommodations to qualified individuals with disabilities.
  • Please tell us if you require a reasonable accommodation to apply for a job.
  • Examples of reasonable accommodations include, but are not limited to, changing the application process, providing documents in an alternate format, using a sign language interpreter, or using specialized equipment.
  • If you require a reasonable accommodation to apply for a job, please contact us via e-mail at [email protected] or via mail at 101 Townsend St. San Francisco, CA 94107.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service