AI Engineer, RL & Evals

Recruiting From Scratch•San Francisco, CA
•Onsite

About The Position

This is a fast-growing, seed-stage AI company building the data and infrastructure layer that helps improve AI model performance in subjective and difficult-to-evaluate domains. The company works directly with leading AI research organizations and application-layer companies on post-training, reinforcement learning environments, evaluation systems, and high-quality data infrastructure. The engineering team operates at the intersection of backend engineering, applied machine learning, reinforcement learning, and product development. Engineers have significant ownership over the systems and environments that help evaluate and improve modern AI models. This is not a pure research role. The ideal candidate is a product-oriented AI engineer who enjoys building production systems, shipping end-to-end software, and working hands-on with RL environments, evaluation frameworks, agent systems, and ML infrastructure. You will work on technically challenging problems where traditional automated evaluation is difficult, including creative and subjective domains where quality cannot always be measured through simple deterministic metrics. The ideal candidate is someone who can move comfortably between backend engineering and applied ML, take ownership of ambiguous problems, and turn research concepts into reliable production systems.

Requirements

  • 2+ years of professional experience as an AI Engineer, ML Engineer, RL Engineer, or strong software engineer working on AI systems.
  • Strong production engineering experience.
  • Experience building reinforcement learning environments, evaluation systems, or ML post-training infrastructure.
  • Experience shipping production software rather than working exclusively on research.
  • Strong backend or full-stack engineering experience.
  • Experience owning technical projects from design through production.
  • Experience working with ambiguous technical problems.
  • Experience working in fast-paced, engineering-driven environments.
  • Strong ability to bridge software engineering and applied ML.
  • Experience collaborating with research, product, or engineering teams.
  • Strong Python experience.
  • Strong PyTorch experience.
  • Hands-on reinforcement learning experience.
  • Experience building ML evaluation frameworks or evaluation systems.
  • Experience working with LLMs.
  • Strong backend engineering fundamentals.
  • Experience with distributed systems.
  • Experience building data pipelines.
  • Experience with APIs and production services.
  • Experience with data processing, indexing, embeddings, or retrieval systems.
  • Strong testing and debugging practices.
  • Experience building production ML or AI infrastructure.
  • Experience building reinforcement learning environments.
  • Experience designing or implementing evaluation frameworks.
  • Understanding of RL concepts and practical application.
  • Experience evaluating LLM or agent behavior.
  • Experience building systems for model post-training.
  • Experience designing tasks, benchmarks, or grading systems.
  • Comfortable working on domains where objective evaluation is difficult.
  • Strong understanding of the interaction between models, data, environments, and evaluation.
  • Experience turning ML research concepts into production systems.
  • Ability to work across model-facing and software infrastructure layers.
  • High ownership and accountability.
  • Strong technical judgment.
  • Product-oriented mindset.
  • Comfortable operating in ambiguity.
  • Strong problem-solving and systems-thinking skills.
  • Able to move quickly from prototype to production.
  • Strong communication skills.
  • Comfortable collaborating with research and engineering teams.
  • Able to explain technical concepts clearly.
  • Strong English communication skills.
  • Curious and motivated by difficult AI problems.
  • Willing to work on-site 5 days per week in San Francisco.

Nice To Haves

  • Experience working with agentic systems or agent frameworks is a strong plus.
  • Experience building agent harnesses or context layers is a strong plus.

Responsibilities

  • Build & Scale RL Environments and Evaluation Systems: Build and scale reinforcement learning environments for complex and subjective domains. Design evaluation frameworks that measure model performance beyond traditional benchmark metrics. Create unique tasks, grading systems, and evaluation methodologies for difficult-to-verify domains. Build systems that allow AI models and agents to be tested systematically. Develop infrastructure for evaluating model behavior, quality, reliability, and performance. Design automated evaluation workflows that reduce reliance on manual human assessment. Iterate on environments and evaluation systems based on model performance and research findings. Build reliable production infrastructure supporting RL and post-training workflows.
  • Build Backend & AI Infrastructure: Design and build production backend systems supporting AI and ML workflows. Build agent harnesses, context layers, APIs, and supporting infrastructure. Develop data pipelines for scraping, indexing, embedding, processing, and evaluating large datasets. Build distributed systems capable of supporting large-scale ML and evaluation workloads. Develop infrastructure connecting models, environments, datasets, agents, and evaluation systems. Work across backend engineering and ML systems to turn ideas into production-ready products. Improve system scalability, reliability, performance, and maintainability. Own backend and infrastructure components from architecture through production.
  • Ship Product-Oriented AI Systems: Ship production code end-to-end rather than working exclusively on research prototypes. Build AI-powered products and infrastructure used by internal teams and external partners. Develop systems that improve the quality and usefulness of AI models in real-world applications. Build agentic systems and evaluation infrastructure for environments where traditional automated metrics are insufficient. Translate ambiguous product and research requirements into practical engineering solutions. Work across product, engineering, and research requirements to deliver production systems. Rapidly prototype, validate, and productionize new ideas. Balance technical experimentation with reliability and production quality.
  • Collaborate With Research & Frontier AI Teams: Collaborate with internal research teams on RL, post-training, evaluation, and model improvement. Work directly with leading AI labs and technical partners to develop environments and evaluation frameworks. Translate research concepts into production engineering systems. Help define tasks, environments, evaluation methodologies, and technical requirements. Communicate technical decisions and tradeoffs clearly across engineering and research teams. Contribute to technical strategy around AI evaluation and post-training infrastructure. Work across backend, ML, data, and product teams to solve complex AI problems. Take significant ownership over systems that directly influence model performance.

Benefits

  • $175,000 – $275,000 base salary.
  • Flexibility to go higher for exceptional candidates.
  • Competitive equity.
  • Full-time position.
  • On-site work model — 5 days per week in San Francisco.
  • Open to visa transfers, including OPT and H-1B transfers.
  • Opportunity to work directly on RL, AI evaluation, post-training, and AI infrastructure.
  • Significant ownership over production AI systems.
  • Opportunity to work with leading AI research organizations and technical partners.
  • Opportunity to build infrastructure for difficult and emerging AI evaluation problems.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service