As a Senior Member of Technical Staff specializing in web data for pre-training, you will play a pivotal role in developing the large scale web data pipeline that underpins Cohere’s advanced language models. In this role, you will work extensively with Common Crawl and other large-scale web corpora, transforming raw, noisy internet data into high-quality training data for pretraining. You will own key components of the data pipeline, including extraction, parsing, deduplication, and filtering. You will also analyze the composition and quality of web data, study its impact on downstream model performance, and collaborate closely with the broader data and evaluation teams to iterate on the training corpus. Your work will be essential to Cohere’s mission of delivering efficient and reliable language understanding and generation capabilities, driving innovation in natural language processing. If you are passionate about transforming data into the foundation of AI systems, this role offers a unique opportunity to make a meaningful impact.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed