Aaru builds simulations of human behavior using AI agents to model real-world populations and decision-making. Companies use these simulations to test choices before committing to product launches, pricing, communications, or policy changes. Aaru emphasizes the need for simulations to represent real people, be calibrated, remain coherent, and present evidence legibly. The company is a small, in-person team in New York, valuing urgency, high ownership, and intellectual honesty, expecting team members to surface inconvenient evidence, adapt quickly, and see work through to completion. Evaluation Research at Aaru focuses on ensuring that the company's populations, predictions, and simulations accurately reflect the real world to support critical decisions. This team defines measurement standards, develops methodologies, and generates evidence for system improvement and capability description. The function operates as both a creator of reusable evaluation infrastructure ('rails') such as datasets, harnesses, and reporting systems, and a performer of specific evaluations ('carts') like backtests, forecast studies, and coherence tests. It is distinct from conventional QA or internal approval, functioning as an independent research unit that collaborates with development teams while maintaining the ability to communicate potentially unfavorable findings.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Manager
Education Level
No Education Listed