As Sweden's national center for applied AI, AI Sweden is seeking a Master's thesis student to work on optimizing Large Language Models (LLMs) for effective use in public administration and enterprise workflows. The project focuses on Reinforcement Learning with Verifiable Rewards (RLVR) to improve multi-step agentic trajectories and error recovery in function calling, specifically comparing RLVR training against Supervised Fine-Tuning (SFT) alone. The student will adapt the oellm-rlvr framework to execute tool-use rollouts and evaluate performance in both Swedish and English across various administrative workflows.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Career Level
Entry Level