About The Position

Quora is a remote-first company with a mission to grow the world's collective intelligence through its platforms: Quora, a global knowledge sharing platform, and Poe, a cloud workspace for AI agents. This role focuses on the Quora product and involves working on the Machine Learning platform and ranking infrastructure. The team is responsible for serving reliability, ML engineer enablement, developer velocity, business impact, and cost efficiency. As a New Grad Software Engineer, you will work at the intersection of machine learning, distributed systems, and GPU serving performance, with the opportunity to ship to production within weeks. No prior ML infrastructure experience is required, as you will receive mentorship and technical guidance from senior engineers. The tech stack includes Python, Go, C++, PyTorch, Kubernetes/EKS, NVIDIA Triton, Ray, and AWS.

Requirements

  • Availability for meetings and impromptu communication during Quora's "coordination hours" (Mon-Fri: 9am-3pm Pacific Time).
  • A 2025 or 2026 graduate with or pursuing a B.S., M.S., or Ph.D. in Computer Science, Engineering, or a related technical field.
  • Genuine interest in large-scale distributed systems, infrastructure, and machine learning.
  • Knowledge of Python, Go, or C++, or the ability to learn them quickly.
  • A passion for learning and always improving yourself and the team around you.

Nice To Haves

  • Previous software engineering experience via an internship, work experience, open-source contribution, or coding competition.
  • Coursework or hands-on experience with ML frameworks such as PyTorch or TensorFlow.
  • Exposure to Kubernetes, Docker, or cloud technologies like AWS.
  • Experience with low-level performance work of any kind: profiling, benchmarking, optimization.
  • Passion for Quora's mission and goals.

Responsibilities

  • Help build and maintain the core infrastructure that powers Quora's ML platform, ensuring high availability, scalability, and performance.
  • Build and improve the distributed systems that serve our ML models in production, from Large Recommendation Models (LRMs) to Large Language Models (LLMs).
  • Work on GPU model serving, optimizing latency, throughput, and cost to support larger and more capable models.
  • Contribute to platform initiatives such as PyTorch-first standardization and ML ecosystem modernization.
  • Improve ML developer velocity by building tooling that helps ML engineers develop, test, and deploy models more efficiently.
  • Modernize our feature store so ML engineers can get new features into production faster.
  • Participate in the team's on-call rotation, helping resolve production issues as you grow your knowledge and ownership of the platform.

Benefits

  • Medical/dental/vision coverage
  • Equity refreshers
  • Remote work reimbursement
  • Paid time off
  • Employee assistance programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service