We are looking for exceptional engineers to join our team. You, alongside the team, will own the platform that runs our benchmarks. This spans everything needed to evaluate LLMs at scale: Python libraries, a web platform, distributed systems, cloud infrastructure, and tooling. You'll work across the stack—whatever needs to be built to run benchmarks reliably and efficiently. At Vals, we believe in autonomy. You will be given a high degree of independence to make decisions on tech stacks, system architecture, and code structure. You will also provide guidance to others on the team, both through informal feedback and formal processes like architecture reviews and code reviews. Our platform serves startups, enterprises, and research labs measuring model performance. We work with all the major foundation model labs, some of the largest financial institutions, and hospital systems in the world. Our work has been featured by the Wall Street Journal, Washington Post, and Bloomberg. We are building the standard for evaluating the ability of LLMs to perform real-world tasks. You will contribute directly to the infrastructure that makes this possible.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed