About The Position

As Sweden's national center for applied AI, AI Sweden is on a mission to accelerate the use of AI to benefit society, competitiveness, and everyone living in Sweden. They drive impactful initiatives in areas such as healthcare, energy, and public services while pushing the boundaries of AI research in fields such as natural language processing, machine learning, and AI security. This master's thesis project focuses on post-training approaches using Reinforcement Learning (RL), specifically Group Relative Policy Optimization (GRPO), to enhance the reasoning capabilities of Large Language Models (LLMs) in multilingual contexts. The project aims to validate whether GRPO can effectively utilize translated and native math/logic rollouts to improve Prelude 9B's reasoning capabilities in Swedish and other EU languages, addressing the gap in understanding GRPO's efficacy on non-English reasoning tasks.

Requirements

  • Ongoing Master’s studies in Computer Science, Data Science, Machine Learning, Engineering Physics, or a related quantitative field.
  • Proficiency in Python and hands-on experience with modern deep learning frameworks (PyTorch, Hugging Face ecosystem).
  • Familiarity with LLM post-training alignment (e.g., SFT, DPO, RLHF/RLVR) or context-extension, alongside comfort running distributed GPU training in Linux/HPC environments.

Nice To Haves

  • Curious, self-driven MSc students eager to work at the frontier of open-weight European AI research (LLMs).
  • Students who thrive on empirical discovery, design rigorous experiments, and let data challenge their assumptions.

Responsibilities

  • Analyze RLHF, PPO, GRPO methodologies, and reasoning benchmarks (e.g., GSM8k, MATH).
  • Adapt the oellm-rlvr framework to perform GRPO on Prelude 9B using Swedish reasoning datasets (e.g., translated Dolci-Think).
  • Benchmark the resulting model on multilingual reasoning tasks, explicitly measuring the trade-off between training TFLOPs/s, memory constraints, and final accuracy.

Benefits

  • Work alongside leading AI scientists and change leaders.
  • Engage with research questions that have both a long shelf-life and are widely applicable to Swedish industry and the public sector.
  • Opportunity for publications at competitive venues.
  • Experience a culture of research excellence.
  • Work in an organization with governmental influence and startup agility.
  • Close contact with government, academia, and private and public sectors.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service