About The Position

As Sweden's national center for applied AI, AI Sweden is seeking a Master's thesis student to work on optimizing Large Language Models (LLMs) for effective use in public administration and enterprise workflows. The project focuses on Reinforcement Learning with Verifiable Rewards (RLVR) to improve multi-step agentic trajectories and error recovery in function calling, specifically comparing RLVR training against Supervised Fine-Tuning (SFT) alone. The student will adapt the oellm-rlvr framework to execute tool-use rollouts and evaluate performance in both Swedish and English across various administrative workflows.

Requirements

  • Ongoing Master’s studies in Computer Science, Data Science, Machine Learning, Engineering Physics, or a related quantitative field.
  • Proficiency in Python and hands-on experience with modern deep learning frameworks (PyTorch, Hugging Face ecosystem).
  • Familiarity with LLM post-training alignment (e.g., SFT, DPO, RLHF/RLVR) or context-extension, alongside comfort running distributed GPU training in Linux/HPC environments.

Nice To Haves

  • Curious, self-driven MSc students eager to work at the frontier of open-weight European AI research (LLMs).
  • Thrive on empirical discovery, design rigorous experiments, and let data challenge your assumptions.

Responsibilities

  • Review agentic environment design, function-calling benchmarks (BFCL, Terminal-Bench), and RLVR mechanics.
  • Adapt oellm-rlvr to execute tool-use rollouts with deterministic execution verifiers on Prelude 9B.
  • Benchmark tool-calling precision, schema validity, and Model FLOPs Utilization (MFU) across single-turn and multi-turn administrative workflows.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service