Off-campus Eastern Washington University users: To download EWU Only theses, please use your EWU SSO (Single Sign-on) credentials. Clicking the blue “Download” button below will prompt you to log in in order to access the thesis document.

Non-EWU users: Please talk to your local librarian about requesting this thesis through Interlibrary loan.

Date of Award

Spring 2026

Rights

Access restricted for 5 years to EWU users with an active EWU NetID

Date Available to Non-EWU Users

2031-06-26

Document Type

Thesis: EWU Only

Degree Name

Master of Science (MS) in Computer Science

Department

Computer Science

First Advisor

Dr. Sanmeet Kaur

Second Advisor

Dr. Abinash Borah

Third Advisor

Mrs. Lynnae Daniels

Abstract

Large language models (LLMs) have emerged as powerful tools for recommendation systems as they leverage deep semantic understanding to model user preferences. However, these systems are susceptible to popularity bias. This is where they over-recommend mainstream items at the expense of relevant niche content, degrading both the user experience and provider fairness. While existing research has explored bias mitigation through prompt engineering as seen in Hamad (2025) or through algorithmic debiasing (e.g., PBiLoss, 2025; DACRec, 2025) in isolation, the integration of both strategies in a single training pipeline remains largely unexamined (see Section 2.6). This study addresses the primary research question of how popularity bias in LLM-based recommendation systems can be more effectively mitigated through a combined approach of structured prompt engineering and customized loss functions without significantly degrading recommendation relevance. This research proposes and evaluates a dual-mitigation framework built upon BLAIR (Hou et al., 2024), a RoBERTa-based sentence encoder pre-trained on the Amazon Reviews 2023 dataset. First, we design Prompt A, an offline LLM augmentation strategy that generates inclusive, stereotype-free item descriptions for training data enrichment, and Prompt B, a gentle inference-time re-ranking instruction. We then introduce a multi-objective loss function combining Inverse-Propensity-Weighted Scaled Cross-Entropy (IPW-SCE), a popularity calibration penalty, a diversity regularizer, and an augmentation consistency term, trained under a two-stage continued pre-training and supervised fine-tuning protocol. Lastly, we develop a comprehensive evaluation framework measuring accuracy, popularity bias via Log Popularity Difference (LPD), diversity, and catalog coverage.

Share

COinS