Personal details
| Title | Improving the Reliability of LLM-as-a-Judge |
| Description | Large language models are increasingly used as automated evaluators to assess the quality, correctness, safety, or relevance of other AI-generated responses. However, an LLM judge may exhibit systematic biases, such as preferring longer answers, being influenced by the order of responses, favoring outputs generated by similar models, or being sensitive to irrelevant formatting and wording. These biases can reduce the reliability and fairness of LLM-based evaluation. The goal of this thesis is to investigate methods for identifying and reducing biases in LLM-as-a-judge systems. The student will:
The expected outcome is an improved evaluation framework that makes LLM-based judging more robust, transparent, and less sensitive to irrelevant characteristics of the evaluated responses. |
| Home institution | Department of Computing Science |
| Associated institutions |
|
| Type of work | practical / application-focused |
| Type of thesis | Bachelor's or Master's degree |
| Author | Prof. Dr. Chih-Hong Cheng |
| Status | available |
| Problem statement | |
| Requirement | The daily research communication will be done in English. |
| Created | 22/07/26 |