Nearly every company in the world runs on custom software for critical operations like tracking performance metrics, handling support workflows, building admin dashboards, and countless processes you might never have thought of. But most companies don't have the resources to properly invest in these tools, leading to a lot of old, clunky internal software, or worse, teams still stuck in manual and spreadsheet workflows. AI has changed who gets to build software. The definition of "developer" now includes analysts, operators, and domain experts creating solutions directly—and the tools they reach for are multiplying by the week. That's both an opportunity and a challenge: as more people build with more AI tools, the risk of shipping ungoverned software into production grows just as fast. At Retool, we're building the platform that makes all of it safe to ship. Build with any AI tool you want, then deploy into one place that connects to your real business data, enforces enterprise policies automatically, and lets teams create once and reuse everywhere with shared, trusted components. The cost of building software has collapsed. The cost of governing it hasn't—and that's the problem we solve. Developers and domain experts have already automated over 100 million hours of work on our platform, freeing them to focus on creative problem-solving and strategic work that drives real business value. The people closest to the problem can now build the software to solve it, safely, and within enterprise guardrails. Let's build the future together. Why we're looking for you We build products where agents write and run code, and where the environment that code runs in is ours to own. As agents get more capable, users go from prompt to working app in minutes instead of hours, and increasingly they expect agents that don't just generate the app but run the work inside it. We're looking for engineers who have shipped agentic products into production and kept them running reliably, affordably, and at scale, and who want to bring that experience to developer-facing surfaces used every day by real engineering teams. What you'll do As an AI Engineer, you'll build the agent platform behind Retool's AI products and own model-driven behavior in production. You will work across the product, infrastructure, and evaluation layers. Your work shapes what users experience and how confidently the team can ship. You might: * Own the behavior of agentic features across multiple product surfaces, including quality, safety, variance, and failure modes, and shape the tool and harness surface agents operate against, including MCP servers, sub-agents, and skills * Work across the product and infrastructure boundary, shaping agent behavior while understanding what it costs at runtime, and serve as the infrastructure team's technical counterpart on agent workloads * Design and evolve prompting, context construction, retrieval, routing, and tool-use strategies for long-horizon workflows, and build the evaluation systems that measure them through statistical signals, distributions, and trends rather than pass/fail tests * Detect, diagnose, and resolve non-deterministic failures such as hallucinations, partial correctness, instruction drift, or context sensitivity, working from transcripts and traces rather than logs alone * Partner closely with product and infrastructure teams on how agent workloads are provisioned, isolated, and rolled out, including for self-hosted customers, and set the pattern for how we ship agentic products safely You'll work across the stack (TypeScript, Node.js, React), but your leverage won't come from code volume alone. It will come from shaping runtime behavior with precision, measurement, and intent. What this role is, and is not It is * Accountable for agent behavior, not just system correctness * Designing, Building, and Deploying agentic products in both cloud and self-hosted environments * Grounded in evaluation, iteration, and regression prevention under non-determinism * Comfortable designing systems where outputs vary, confidence is probabilistic, and correctness is contextual It is not * Adding LLM calls to existing features and moving on * Shipping AI features without owning their long-term reliability, drift, or user trust * An SRE role, though you'll own the agent-side bugs that surface as infrastructure incidents * Model training or research, though you'll shape model behavior, selection, and tool design
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed