Reflection's Data team builds the training corpora our frontier models learn from. Before a model can learn anything, the data has to be found, fetched, extracted, and delivered reliably, responsibly, and at enormous scale. The ingestion layer is the machinery that turns the open web, licensed corpora, and other large-scale sources into well-structured, versioned, auditable datasets for pre-training. As Data Ingestion Lead, you'll provide front-line leadership of the team that builds this layer, spanning all three of its pillars web crawl, data ingestion pipelines, and data lakes. You'll build, mentor, and grow a team of data ingestion engineers, guide the technical and architectural decisions across crawling, extraction, and corpus storage/delivery, and work closely with the pre-training research, data quality, and data partnerships teams that depend on what you ship. You'll stay close enough to the stack to make targeted contributions as an individual contributor and to maintain a deep understanding of the team's technical work. Subtle decisions at the ingestion layer what we crawl, how we extract, what we keep ripple through training and directly affect where our models are strong, safe, and where they fail.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed