When Preferences Fail to Become Incentives: A Utility–Behavior Gap in Large Language Models | ResearchPod