Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos | ResearchPod