This work explores the role of prompt design and judge selection in LLM-as-a-Judge evaluations of free text legal question answering. We examine whether automatic task prompt optimization improves over human-centered design, whether optimization effectiveness varies by judge feedback style, and whether optimized prompts transfer across judges. We systematically address these questions on the LEXam benchmark by optimizing task prompts using the ProTeGi method with feedback from two judges (Qwen3-32B, DeepSeek-V3) across four task models, and then testing cross-judge transfer. Automatic optimization consistently outperforms the baseline, with lenient judge feedback yielding higher and more consistent gains than strict judge feedback. Prompts optimized with lenient feedback transfer better to strict judges than the reverse direction. Analysis reveals that lenient judges provide permissive feedback, yielding prompts with broader applicability, whereas strict judges produce restrictive feedback, leading to judge-specific overfitting. Our findings demonstrate algorithmically optimizing prompts on training data can outperform human-centered prompt design and that judges' dispositions during optimization shape prompt generalizability. Code and optimized prompts are available at https://github.com/TUMLegalTech/icail2026-llm-judge-gaming.
翻译:本研究探讨了在自由文本法律问答的LLM-as-a-Judge评估中,提示设计与法官选择的作用。我们考察了自动任务提示优化是否优于以人为中心的设计,优化效果是否因法官反馈风格而异,以及优化后的提示是否能在不同法官间迁移。我们通过在LEXam基准上系统性地解决这些问题,采用ProTeGi方法,使用两位法官(Qwen3-32B、DeepSeek-V3)的反馈对四个任务模型进行任务提示优化,然后测试跨法官的迁移效果。自动优化始终优于基线,宽松的法官反馈相比严格的法官反馈能带来更高且更一致的增益。由宽松反馈优化的提示向严格法官的迁移效果优于反向迁移。分析表明,宽松法官提供宽容反馈,生成了适用性更广的提示,而严格法官产生限制性反馈,导致针对特定法官的过拟合。我们的研究结果证明,在训练数据上通过算法优化提示可以超越以人为中心的提示设计,并且优化过程中法官的倾向性会影响提示的泛化能力。代码和优化后的提示可从https://github.com/TUMLegalTech/icail2026-llm-judge-gaming获取。