Arcade is building the world’s first AI physical product creation platform, where imagination becomes reality. Our platform lets anyone design, purchase, and sell custom, manufacturable products using natural language and generative AI. We believe everyone should have the power to create physical goods as easily as they post online, and we’re building the infrastructure to make that real for both consumers and businesses. We’ve raised $42M from a world-class group of investors, including Reid Hoffman, Forerunner Ventures (Kirsten Green), Canaan Partners (Laura Chau), Adverb Ventures (April Underwood), Factorial Funds (Sol Bier), Offline Ventures (Brit Morin), Sound Ventures (Ashton Kutcher), Inspired Capital (Alexa von Tobel), and Torch Capital (Jonathan Keidan). Our angel investors include Elad Gil, Ev Williams, Marissa Mayer, Sara Beykpour, Kayvon Beykpour, Anna Veronika Dorogush, Eugenia Kuyda, David Luan, Sharon Zhou, Kelly Wearstler, Karlie Kloss, Colin Kaepernick, Christy Turlington Burns, and Jeff Wilke. Arcade is headquartered in San Francisco’s Presidio and led by serial entrepreneur Mariam Naficy (Minted, Eve), and a mission-driven team from Google, Apple, Stability AI, Glean, NVIDIA, Databricks, LinkedIn, Stanford, MIT, Berkeley, and more. Arcade’s Chief AI Officer is Varun Jampani, a leading researcher who co-authored Dreambooth and created Stable Diffusion 3.5, among other things. Raghudeep Gadde, Head of Research at Arcade, was formerly a Principal Scientist at Amazon. Together, we’re pioneering a new category at the intersection of AI, personal expression, and on-demand manufacturing, and we’re building fast. The Role We are seeking a high-caliber, deeply innovative Vision-Language Model (VLM) expert to lead our efforts in teaching foundation models how to evaluate consumer products like an expert appraiser. This is not a standard implementation role. You will be expected to invent new technologies, design novel architectures, and author proprietary training paradigms when existing open-source or commercial models fall short. You will go beyond simple object detection, engineering systems that can estimate highly abstract and valuable aspects of product images—such as generating rich, context-aware captions, predicting precise market price points, and evaluating subjective aesthetic quality or "beauty" scores. If you are a pioneer who thrives on solving unsolved multimodal problems and wants your inventions to power a groundbreaking production platform, we want to hear from you.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
Ph.D. or professional degree