We are building a recursive self-improvement system — a machine learning system that iteratively improves itself through feedback, evaluation, and automated learning loops. You will help build the engineering pipeline that keeps these loops fast, reliable, and trustworthy: the training and evaluation pipelines, the reward and feedback signals, and the safeguards that prevent a self-improving system from silently degrading or gaming its objectives. This is an engineering-first role with deep reinforcement learning requirements. You should be equally comfortable writing robust production ML code and reasoning about reward design, credit assignment, and why feedback-driven systems become unstable.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior