ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning | ResearchPod