Senior Research Engineer, Agentic Data and Tooling, DeepMind

GoogleNew York, NY
$174,000 - $252,000

About The Position

At Google, research-focused Software Engineers are embedded throughout the company, allowing them to setup large-scale tests and deploy promising ideas quickly and broadly. Ideas may come from internal projects as well as from collaborations with research programs at partner universities and technical institutes all over the world. From creating experiments and prototyping implementations to designing new architectures, engineers work on real-world problems including artificial intelligence, data mining, natural language processing, hardware and software performance analysis, improving compilers for mobile platforms, as well as core search and much more. But you stay connected to your research roots as an active contributor to the wider research community by partnering with universities and publishing papers. The Google Deepmind (GDM) Agent Data and Tooling team within the Human Data Platform organization builds the environments, pipelines, and tooling that power Gemini's frontier agentic capabilities. We operate under an active data ownership model — moving beyond commodity data collection to build high-fidelity interactive worlds, capture complex multi-turn trajectories, and land data into model training and evaluation pipelines to drive measurable hillclimbing on coding, computer control, tool-use benchmarks, and more. Build the critical infrastructure and interactive environments that directly drive Gemini's agentic and reasoning capabilities. In this role, you will sit at the intersection of software engineering and model training: writing high-velocity production code to create rich interactive worlds. We are looking for an engineer who loves to deliver code and build 0-1 systems at lightning pace and under high pressure, making heavy use of AI tools to boost velocity/output. Artificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority. We are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort. Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

Requirements

  • Bachelor's degree in Computer Science, Information Technology, a related technical field, or equivalent practical experience.
  • 5 years of experience working with large language models (LLMs).
  • 2 years of experience developing and training machine learning models.
  • Experience with Agentic integrations and Model Context Protocol.

Nice To Haves

  • Master's degree or PhD in Electrical Engineering, Computer Science, or equivalent practical experience.
  • 2 years of experience with full-stack development.
  • Excellent analytical, problem-solving and communication skills with demonstrated attention to detail.
  • A deep passion for AI technology and all of its possibilities .

Responsibilities

  • Build agentic data infrastructure at high velocity: Own and deliver key components across the agentic data stack. Rapidly prototype, iterate, and ship robust production code to generate, capture, and curate complex multi-turn agent trajectories at scale.
  • Bridge research and engineering to drive model hillclimbing: Collaborate closely with Gemini research teams to close the loop between data creation and model quality. Understand how multi-turn trajectory design, environment complexity, and reward signals impact SFT, RL training, and capability hillclimbing. Directly integrate curated data into training pipelines and evaluate downstream model performance on frontier benchmarks.
  • Build quality tooling and support gold-standard benchmarks: Create human-in-the-loop annotation tooling and interactive trajectory review surfaces, working in tandem with automated validation checkers leveraging adversarial LLM judges and programmatic verifiers. Support the creation and curation of gold-standard evaluation sets for flagship benchmarks.

Benefits

  • 15% bonus target
  • equity
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service