As Sweden's national center for applied AI, AI Sweden is seeking a master's thesis student to accelerate the use of AI for societal benefit, competitiveness, and innovation. This thesis focuses on improving the understanding of Large Language Model (LLM) agents, which increasingly rely on retrieved documents, tool outputs, and past memory. When these agents make incorrect or harmful decisions, it is often difficult to pinpoint the exact input that caused the issue. While observability platforms offer automated root-cause suggestions, their reliability is questionable due to the lack of known causes in real-world incidents. Existing benchmarks often identify the responsible agent or step rather than the specific input content, and their accuracy remains low. This thesis aims to build a benchmark of executed agent traces with planted, verified causes to evaluate existing attribution methods. The project involves a literature study, building a test set with known causes (e.g., hidden instructions, false tool results, outdated memory), and testing existing attribution methods on this benchmark to measure their accuracy, computational needs, and sensitivity to factors like trace length and model randomness.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Entry Level