Off-campus Eastern Washington University users: To download EWU Only theses, please use your EWU SSO (Single Sign-on) credentials. Clicking the blue “Download” button below will prompt you to log in in order to access the thesis document.
Non-EWU users: Please talk to your local librarian about requesting this thesis through Interlibrary loan.
Date of Award
Spring 2026
Rights
Access restricted for 1 year to EWU users with an active EWU NetID
Date Available to Non-EWU Users
2027-06-22
Document Type
Thesis: EWU Only
Degree Name
Master of Science (MS) in Computer Science
Department
Computer Science and Electrical Engineering
First Advisor
Sanmeet Kaur
Second Advisor
Antonio Espinoza
Third Advisor
Lynnae Daniels
Abstract
Prompt refinement is widely used to improve large language model (LLM) performance, yet the conditions under which such improvements remain stable, transferable, and dependent on task family remain poorly understood. This thesis tests the hypothesis that autonomous prompt refinement produces reliable held-out performance improvements across rule-governed task families. Under this framing, stable held-out improvement would support the view that refinement can generalize beyond the failures used to revise a prompt; when such improvement does not persist, the study examines whether the result reflects saturation, overfitting, or task-family-specific failure boundaries.
The study empirically evaluates three autonomous refinement strategies: a deterministic heuristic refiner that applies fixed policy edits in response to observed failures, a structured DSPy-based refiner that searches over candidate policy revisions, and an evolutionary refiner that mutates and selects prompt policies across generations. These refiners are tested across representative classes of rule-governed tasks spanning Syntax-Constrained Generation (SCG), Deductive Logic, and Constraint Satisfaction Problems (CSPs). The evaluation uses strict machine-checkable validation, fixed development and held-out splits, difficulty-tiered analysis where applicable, and task-family-specific validators. To isolate the effects of prompt revision, refinement is evaluated under unassisted prompt-only generation, without external tools or task-specific solvers, and assessed in terms of development-set improvement, held-out transfer, and refinement overfitting.
The results show that autonomous prompt refinement does not function as a uniform improvement method. Instead, its effects vary systematically with task family, baseline error rate, and the source of the observed failures. Across SCG tasks, including JSON schema generation and a custom domain-specific language (DSL), baseline performance was high, suggesting that these tasks were largely within the models’ existing capabilities; the comparatively few remaining errors were often correctable through prompt refinement. All three refiners successfully repaired generated-JSON failures, while in the DSL task family, refined prompts converged on semantically similar path-reversal instructions without fully correcting the underlying model’s execution of the required operation.
By contrast, compact CSP tasks remained difficult: models produced plausible but constraint-violating outputs despite prompt changes, indicating a failure boundary for sustained constraint maintenance. Deductive Logic tasks produced the richest refinement signal. Deterministic heuristic refinement produced the smallest but most stable held-out gains, DSPy produced mixed but occasionally transferable policy interventions, and evolutionary refinement produced the largest development-set gains, but the selected policies did not preserve those gains under held-out evaluation. Natural-language deduction tasks exposed model-specific prompt sensitivity, failure-conditioned recovery, uncertainty-threshold effects, and tier-dependent overfitting; symbolic entailment and equisatisfiability tasks showed that apparent small-sample gains can collapse under broader evaluation because of label-threshold shifts, reasoning-budget effects, and development-set overfitting. Overall, the findings characterize autonomous prompt refinement as a bounded diagnostic method for identifying saturation, errors correctable through prompt revision, failure boundaries, overfitting, and threshold shifts across rule-governed task families.
Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-No Derivative Works 4.0 International License.
Recommended Citation
Locke, Robert B., "Autonomous Prompt Refinement for Discrete, Rule-Governed Tasks" (2026). EWU Masters Thesis Collection. 1018.
https://dc.ewu.edu/theses/1018