About The Position

As Sweden's national center for applied AI, AI Sweden is seeking a master's thesis student to accelerate the use of AI for societal benefit, competitiveness, and innovation. This thesis focuses on improving the understanding of Large Language Model (LLM) agents, which increasingly rely on retrieved documents, tool outputs, and past memory. When these agents make incorrect or harmful decisions, it is often difficult to pinpoint the exact input that caused the issue. While observability platforms offer automated root-cause suggestions, their reliability is questionable due to the lack of known causes in real-world incidents. Existing benchmarks often identify the responsible agent or step rather than the specific input content, and their accuracy remains low. This thesis aims to build a benchmark of executed agent traces with planted, verified causes to evaluate existing attribution methods. The project involves a literature study, building a test set with known causes (e.g., hidden instructions, false tool results, outdated memory), and testing existing attribution methods on this benchmark to measure their accuracy, computational needs, and sensitivity to factors like trace length and model randomness.

Requirements

  • Ongoing Master’s studies in Computer Science, Data Science, Engineering Physics, Complex Adaptive Systems, Machine Learning, or a related field.
  • Curiosity, independence, and self-driven attitude.
  • Comfortable with Python.
  • Comfortable with deep learning.
  • Comfortable with the reality that experiments might yield unexpected results.

Responsibilities

  • Conduct a literature study on current methods for identifying causal steps or inputs in agent failures.
  • Build a simple agent that searches documents, calls tools, and remembers past sessions.
  • Plant inputs (hidden instructions, false tool results, outdated memory) into the agent's document store to cause data leaks.
  • Verify planted inputs by running the agent multiple times with and without them.
  • Create test cases with single causes, multiple interacting causes, and no planted cause.
  • Test existing attribution methods by providing them with recorded agent runs and evaluating their accuracy in identifying the causal input.
  • Measure the computational resources required by the tested methods.
  • Analyze how results change with longer runs and increased model randomness.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service