Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions | ResearchPod