As Sweden's national center for applied AI, AI Sweden is on a mission to accelerate the use of AI to benefit society, competitiveness, and everyone living in Sweden. They drive impactful initiatives in areas such as healthcare, energy, and public services while pushing the boundaries of AI research in fields such as natural language processing, machine learning, and AI security. This master's thesis project focuses on post-training approaches using Reinforcement Learning (RL), specifically Group Relative Policy Optimization (GRPO), to enhance the reasoning capabilities of Large Language Models (LLMs) in multilingual contexts. The project aims to validate whether GRPO can effectively utilize translated and native math/logic rollouts to improve Prelude 9B's reasoning capabilities in Swedish and other EU languages, addressing the gap in understanding GRPO's efficacy on non-English reasoning tasks.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Career Level
Entry Level