We are looking for a Software Engineer to build the common infrastructure for data curation at MEDC. Today, data curation work across MEDC and our modeling partners (e.g., semantic and content QA datasets, generative retrieval evals) happens largely in ad-hoc notebooks, with no shared capabilities or standardization. This makes every new curation project slow to launch, hard to discover, and labor-intensive to productionize. You will turn that into a coherent, reusable platform. This is not a pure data engineering role. The core of the work is using LLMs to transform data, for example turning the Netflix catalog and member signals into question-answer pairs, and then deciding what to keep. That means designing sampling strategies and filtering methods, often based on evaluation models, that maximize data quality, and proving that those choices actually improve downstream models. You will work hand in hand with researchers, so modeling intuition matters as much as engineering skill.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed