Ashok Prasad Neupane, Dipan Bartaula, Ankit Belbase, Saugat Adhikari, Samip Ghimire, Saroj Poudel, Binod Bhattarai, Danda Pani Paudel
5 min
Abstract
Existing streaming vision-language models (VLMs) continuously perceive and reason over visual streams, but their computational pathways remain fixed throughout inference. Consequently, they cannot adapt computation to evolving scene dynamics, where different future events demand different levels and forms of perception. We show that streaming VLMs inherently possess the ability to anticipate the immediate future, and leverage this capability to dynamically configure future computation in a training-free manner. Realizing such anticipatory computation, however, is very challenging: future anticipation must be sufficiently reliable to guide computation, planning must run concurrently with streaming inference, and online reconfiguration must incur negligible overhead. To address these challenges, we introduce FORESIGHT, a dual-stream architecture comprising two Siamese LLMs with shared weights, input encoders, and KV cache. The first LLM continuously processes incoming tokens, while the second runs ahead of the stream to anticipate future context, plan future computation, and generate task responses without interrupting streaming inference. Each plan decides when to reason next, what to check then, and how densely to sample, keeping transient evidence separate from persistent control. The resulting computation plan is executed online through an efficient reconfiguration protocol with schemaguided decoding and lightweight diff-based updates, enabling dynamic adaptation with low overhead. With a frozen Qwen3-VL-8B backbone, FORESIGHT achieves 23.0 mean joint F1 on OmniPro Online evaluation beating strongest trained baseline by 9.5%, while improving the backbone by 6.7 on StreamingBench and 15.4 on OVO-Bench, with the largest gain of 18.7 when evidence arrives later in the video stream. Our source code will be made publicly available.
Sam: That's hysteresis, then. It forces the model to be sure before acting. But doesn't that risk missing the onset of a fast event?
Alex: That's the central tension. The authors fit the thresholds on a calibration split to balance latency against precision. It isn't perfect, but it cuts down on reporting events that aren't there.
Sam: The benchmark pattern interests me. Why does it do so much better on event narration than on deduplicated counting?
Alex: I'd attribute it to what each task demands. Narration is driven by high-saliency events the system is well placed to catch. Counting requires a consistent state over time, and that sits badly with aggressive token pruning.
Sam: Pruning is destructive. Drop the wrong tokens and you lose the history you need to count accurately. So it trades long-term state for fast reaction.
Alex: Yes. It's optimized for proactive response, not exhaustive analysis. For a streaming assistant that's often the right priority, but it marks a real limit for approaches built on aggressive pruning.
Sam: What about the cost side? Is the ingest pass a fixed expense regardless of how much video the planner actually needs?
Alex: Largely, yes. The appendix cost model shows compute dominated by the initial ingest, with the planning checks adding only marginal overhead. The shared cache means visual tokens are prefilled once, and the read-only planning loop avoids the redundant work of models that re-read the whole history.
Sam: And the check rate? Does the model learn to adjust it?
Alex: It's reactive, not learned. The planner lengthens the check interval when visual evidence is low-saliency. There's no policy being trained in the reinforcement learning sense. It's using the backbone's latent foresight to time the next check.
Sam: So the contribution isn't a new VLM. It's the control logic around a frozen one.
Alex: That's how I read it. It reframes anticipation from a prediction target into an inference-time controller. The bottleneck in streaming may be less what the model knows than how it allocates computation given expected future states. The caveats are the dependence on backbone grounding, static thresholds, and the cost to long-horizon state.
Sam: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Alex: Thanks for listening.