Senior AI Data Engineer

Cooley AI•Washington, DC
•$220,000 - $250,000

About The Position

Cooley AI is seeking a Senior AI Data Engineer to join the Innovation team. Cooley AI is the wholly-owned subsidiary of Cooley LLP. The frontier development arm of the leading technology law firm in the world. Our mission is to bring the best-in-class practices from the technology field into the legal domain to help lead our firm and the legal industry into the AI era. All employees of Cooley AI are employees of Cooley and are seconded to Cooley AI. As a leading technology law firm, Cooley is determined to become a leader in the digital practice of law. The Senior AI Data Engineer is a senior technical individual contributor within the Data Products function, responsible for building and maintaining the AI infrastructure layer that powers the firm's AI-powered data applications on the Northstar Lakehouse. Working as a peer to the Senior Data Application Engineer, this role is responsible for the intelligence layer beneath the application: the RAG pipelines, embedding infrastructure, vector search implementation, LLM integration patterns, and AI feature observability that make AI-powered legal intelligence products reliable, performant, and trustworthy in production. The Senior AI Data Engineer brings production AI engineering depth to a greenfield AI application build, operating at the intersection of data engineering, applied machine learning, and enterprise application development. This is a high-impact individual contributor role and will build production RAG and LLM-powered systems with the understanding of what it takes to make AI features work reliably at scale.

Requirements

  • After orientation at Cooley LLP, exhibit proficiency in the Microsoft Office suite, iManage, and other firm applications
  • Ability to work extended and/or weekend hours, as required
  • Ability to travel, as required
  • 5+ years of data engineering or AI engineering experience with hands-on production responsibility for RAG pipelines, LLM integration, or AI-powered application data infrastructure in a cloud data platform environment
  • Demonstrated production RAG pipeline experience including document ingestion, chunking strategies, embedding generation, vector index design, retrieval quality monitoring, and ongoing pipeline maintenance at scale
  • Strong LLM API integration experience in production applications including prompt engineering, output validation, fallback handling, latency management, and cost-aware usage patterns using OpenAI, Anthropic, Google, or comparable LLM APIs
  • Hands-on vector search implementation experience including Pinecone, Databricks Vector Search, pgvector, Chroma, or comparable vector database platforms in a production application context
  • Strong Python proficiency including modular package design, test-driven development, and production-grade code standards for AI pipeline and infrastructure development
  • Databricks experience including Delta Lake, Unity Catalog, Databricks SQL, and medallion architecture conventions.
  • Experience with AI feature observability including LLM response quality monitoring, retrieval relevance metrics, token consumption tracking, and latency alerting for production AI systems
  • Production data engineering experience including pipeline development, data quality validation, CI/CD integration, and version control discipline in a cloud data platform environment
  • Comfort using GitHub for version control, pull requests, and CI/CD pipeline integration in a collaborative engineering context
  • Demonstrated Agile delivery fluency including sprint ceremonies, ticket writing, estimation, and definition of done discipline

Nice To Haves

  • Experience with Databricks Vector Search, FMAPI, Mosaic AI, or Databricks Model Serving in a production AI application context
  • Familiarity with LangChain, LangGraph, LlamaIndex, or comparable AI orchestration frameworks for RAG pipeline development and LLM workflow integration
  • Experience with embedding model evaluation, selection, and fine-tuning including open-source embedding models and the trade-offs between model quality, latency, and cost in production retrieval systems
  • Legal sector, professional services, or enterprise SaaS experience with exposure to complex document processing, knowledge retrieval, or AI-powered workflow automation in a regulated or high-stakes data environment
  • Familiarity with React application architecture and how AI features are surfaced to users through front-end application components, sufficient to collaborate effectively with the Senior Data Application Engineer on the application and intelligence layer interface
  • Experience with Databricks Appkit or Databricks Apps for AI-powered application hosting
  • Exposure to agentic AI patterns including tool-calling architectures, agent memory and state management, and multi-step LLM workflow orchestration using LangGraph, OpenAI Assistants, Claude with tools, or comparable frameworks
  • Experience with PySpark or distributed data processing for large-scale document ingestion and embedding pipeline development at volumes that exceed single-node processing capacity
  • Comfort using AI-assisted development tools such as Claude Code, GitHub Copilot, Cursor, or comparable to accelerate AI pipeline development, documentation, and code review workflows

Responsibilities

  • Design, build, and maintain production RAG pipelines that ingest, process, chunk, embed, and index firm data assets for retrieval by AI-powered legal intelligence applications. Own the full RAG pipeline lifecycle from document ingestion through retrieval quality monitoring
  • Own the vector search implementation including index design, embedding model selection and maintenance, retrieval strategy configuration, and ongoing retrieval quality assessment. Ensure the vector search layer returns accurate, contextually relevant results
  • Build and maintain embedding pipelines that transform structured and unstructured firm data assets into vector representations consumable by retrieval systems and LLM-powered features. Apply chunking strategies, metadata enrichment, and embedding quality validation appropriate to legal sector document complexity
  • Monitor and maintain retrieval quality in production, establishing metrics for retrieval relevance, embedding freshness, index coverage, and pipeline execution health that surface degradation before it affects application users
  • Collaborate with the Data Engineering & Platform function to ensure RAG pipeline data sources are correctly integrated with the Northstar Lakehouse Silver and Gold layer data assets, consuming governed, certified data rather than bypassing platform governance controls
  • Build and maintain the LLM integration layer for Data Applications squad products, implementing prompt engineering patterns, output validation frameworks, fallback handling, latency management, and cost-aware API usage patterns that make LLM-powered features reliable in production
  • Implement AI output handling patterns that ensure application features degrade gracefully when model responses are unexpected, latency is high, or retrieval quality issues affect the context provided to the LLM. Design for failure from the outset rather than retrofitting resilience
  • Build and maintain AI feature observability tooling that provides the Data Applications squad with visibility into LLM response quality, retrieval relevance, token consumption trends, latency patterns, and user interaction signals that inform ongoing AI feature improvement
  • Evaluate and integrate Databricks-native AI capabilities including FMAPI, Mosaic AI, Databricks Model Serving, and Databricks Vector Search as they mature, making architectural recommendations on adoption timing and integration patterns alongside the Platform Architect
  • Apply responsible AI practices to all AI feature implementations including LLM output validation, bias awareness in retrieval systems, appropriate user-facing communication of AI output reliability, and compliance with the firm's AI governance framework
  • Stay current on the rapidly evolving AI engineering landscape including advances in RAG architecture, embedding models, LLM APIs, and AI application patterns, bringing relevant innovations to the squad's technical roadmap
  • Build and maintain data pipelines that prepare, transform, and deliver data assets to AI consumption layers, ensuring data flowing into RAG pipelines, embedding infrastructure, and LLM context windows is accurate, current, and appropriately governed
  • Work within Unity Catalog access controls and data classification policies, understanding the governance framework and ensuring AI pipeline implementations respect data classification, access control, and audit logging requirements rather than working around them
  • Collaborate with the Data Engineering & Platform function on the data contracts between the Northstar Lakehouse Silver and Gold layers and the AI infrastructure layer, ensuring AI pipelines consume data through governed interfaces and surface data freshness and quality signals to the application layer
  • Implement pipeline-level data quality checks within AI ingestion pipelines, validating that data entering RAG pipelines and embedding infrastructure meets the completeness, accuracy, and freshness standards required for reliable AI feature performance
  • Apply production engineering discipline to all AI pipeline implementations including version control, CI/CD integration, testing frameworks, documentation standards, and deployment through the established platform CI/CD framework
  • Work as a close technical partner to the Senior Data Application Engineer, owning the AI and data infrastructure layer that the React application on Databricks Appkit depends on. Coordinate proactively on the interface between the application layer and the intelligence layer so neither engineering track becomes a blocker for the other
  • Participate fully in Data Applications squad Agile ceremonies including sprint planning, daily standups, sprint reviews, and retrospectives. Write clear, well-estimated tickets that accurately reflect the complexity of AI infrastructure work and distinguish spike work from build work
  • Communicate proactively about AI infrastructure dependencies, retrieval quality issues, LLM API changes, and embedding pipeline health that may affect application behavior or sprint commitments
  • Contribute to the Data Applications squad engineering culture by modeling production-grade AI engineering practices, sharing knowledge of RAG architecture and LLM integration patterns, and peer reviewing code with teaching intent
  • Engage with the Director, Data Products on the AI feature roadmap, providing honest technical assessments of AI feature feasibility, implementation complexity, and the infrastructure investment required to make AI-powered features production-grade
  • All other duties as assigned or required

Benefits

  • medical
  • health savings account (with applicable medical plan)
  • dental
  • vision
  • health and/or dependent care flexible spending accounts
  • pre-tax commuter benefits
  • life insurance
  • AD&D
  • long-term care coverage
  • backup care for children and/or adults
  • other parental support benefits
  • firm-paid life insurance
  • firm-paid AD&D
  • firm-paid LTD
  • short term medical benefits
  • 21 days of Paid Time Off (“PTO”)
  • 10 paid holidays
  • generous parental leave
  • fertility benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service