Datalab is building the core infrastructure for how enterprises process and understand documents at scale. We’re at an 8-figure run rate with a team of 7. Anthropic is a customer. And we have hundreds more across FAANG, frontier AI labs, healthcare, finance, government, and legal. Our models - Chandra, Surya, Marker, and Lift - have significant adoption, with 60,000+ GitHub stars and broad developer mindshare. We’re backed by founding members of OpenAI, FAIR, and Hugging Face. We move fast, ship often, and we're hiring builders who do the same. We're looking for a Research Engineer to own problems end to end across our models, inference service, and product. You won't just train a model and hand it off. You'll take it from training through benchmarking, into our inference stack, and work with the team to integrate it into our products. We're a small team that has shipped the current state of the art OCR model, Chandra. Our models collectively have 50k+ Github stars. Our tools are used internally at frontier AI labs like Anthropic, and Fortune 500 enterprises like Siemens. Our team focuses on training small, efficient models that outperform much larger LLMs on domain-specific tasks (like OCR, structured extraction, tables). We move fast, prioritize practical results, and build tools that are open, reproducible, and built to last. You'll test hypotheses quickly, iterate on results, and balance experimental rigor with shipping to customers.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed