Topic: Improving the Reliability of LLM-as-a-Judge

Topic: Improving the Reliability of LLM-as-a-Judge

Personal details

Title Improving the Reliability of LLM-as-a-Judge
Description

Large language models are increasingly used as automated evaluators to assess the quality, correctness, safety, or relevance of other AI-generated responses. However, an LLM judge may exhibit systematic biases, such as preferring longer answers, being influenced by the order of responses, favoring outputs generated by similar models, or being sensitive to irrelevant formatting and wording. These biases can reduce the reliability and fairness of LLM-based evaluation.

The goal of this thesis is to investigate methods for identifying and reducing biases in LLM-as-a-judge systems. The student will:

  • review existing LLM-as-a-judge methods and known sources of evaluation bias;
  • design experiments to measure biases such as position, verbosity, style, and self-preference bias;
  • implement and compare mitigation techniques, such as response-order randomization, score calibration, multi-judge aggregation, structured evaluation criteria, or consistency checking;
  • evaluate the proposed methods on established datasets and different language models; and
  • analyze the trade-off between evaluation reliability, computational cost, and agreement with human judgments.
  • develop new techniques to push state-of-the-art.

The expected outcome is an improved evaluation framework that makes LLM-based judging more robust, transparent, and less sensitive to irrelevant characteristics of the evaluated responses.

Home institution Department of Computing Science
Associated institutions
Type of work practical / application-focused
Type of thesis Bachelor's or Master's degree
Author Prof. Dr. Chih-Hong Cheng
Status available
Problem statement
Requirement

The daily research communication will be done in English. 

Created 22/07/26