This role is on-site M-F in the Flatiron District - New York City, NY. Our client is a Series A point-of-sale software company in NYC. They move fast and ship daily. AI-assisted engineering isn't a talking point here; agents like Devin, Claude, and Codex are part of how the team works every day, and engineers who master that leverage ship multiples of what they could alone. About the Team Applied AI owns the intelligence layer, including aspects of Merchandising, Custom Reporting, Ecommerce, and Agentic workflows. This is the team behind our AI Manager — a suite of production agents that handle end-to-end workflows like retail pricing strategies, creating marketing campaigns, maintaining SEO (eCommerce) and developing custom reports that give store owners the insight they need to run a stronger business. This team also owns invoice OCR/CV (structured extraction from messy distributor invoices, with selectable LLM engines) and embedding-based product matching. Nothing here is a demo: every model output lands in front of a real store owner making a real pricing decision. About the Role You'll design, build, and iterate on the agents powering our core merchant experience. Working on this mission-critical team, you'll develop production agent and context-layer applications — turning each store's sales history, catalog, and market context into decisions that make independent retailers money. You'll own your models from research to production, and your work ships into stores the same week. In this role, you will Apply state-of-the-art ML and LLM techniques to problems spanning: Merchandising intelligence (slow-mover detection, price and promotion recommendation, competition and seasonality signals); Document understanding (invoice OCR and structured extraction across LLM engines); Retrieval and ranking (embedding-based product matching on pgvector, catalog dedup, contextual recommendations) Build agent capabilities on top of our Manager Agent platform — task generation, review workflows, and chat over each store's own data Build the evaluation harness for both offline and online techniques, designing experiments and metrics (evals, QA playbooks, Langfuse tracing) that provide deep insight into recommendation quality and merchant impact Own the entire model lifecycle from research to production: data analysis, modeling, evaluation, offline/online testing, and iterative improvement — and build autonomous harnesses that let agent squads explore new problem spaces in parallel Collaborate cross-functionally with engineers, PMs, and store owners to ensure our AI drives measurable improvements in merchant revenue and hours saved Stay at the forefront of ML/AI innovation by evaluating and incorporating emerging research, models, and techniques into the product lifecycle
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
High school or GED