Existing streaming vision-language models (VLMs) continuously perceive and reason over visual streams, but their computational pathways remain fixed throughout inference. Consequently, they cannot adapt computation to evolving scene dynamics, where different future events demand different levels and forms of perception. We show that streaming VLMs inherently possess the ability to anticipate the immediate future, and leverage this capability to dynamically configure future computation in a training-free manner. Realizing such anticipatory computation, however, is very challenging: future anticipation must be sufficiently reliable to guide computation, planning must run concurrently with streaming inference, and online reconfiguration must incur negligible overhead. To address these challenges, we introduce FORESIGHT, a dual-stream architecture comprising two Siamese LLMs with shared weights, input encoders, and KV cache. The first LLM continuously processes incoming tokens, while the second runs ahead of the stream to anticipate future context, plan future computation, and generate task responses without interrupting streaming inference. Each plan decides when to reason next, what to check then, and how densely to sample, keeping transient evidence separate from persistent control. The resulting computation plan is executed online through an efficient reconfiguration protocol with schemaguided decoding and lightweight diff-based updates, enabling dynamic adaptation with low overhead. With a frozen Qwen3-VL-8B backbone, FORESIGHT achieves 23.0 mean joint F1 on OmniPro Online evaluation beating strongest trained baseline by 9.5%, while improving the backbone by 6.7 on StreamingBench and 15.4 on OVO-Bench, with the largest gain of 18.7 when evidence arrives later in the video stream. Our source code will be made publicly available.
Alex: A new framework called Foresight lets a streaming vision-language model decide for itself when to look and what to attend to, with no additional training. It comes from Ashok Prasad Neupane and colleagues, and they report state-of-the-art performance on streaming benchmarks.
Sam: Most streaming models are reactive. They wait for a trigger or follow a fixed schedule. If nothing is fine-tuned, how does it anticipate anything without having learned temporal patterns?
Alex: The authors' position is that the anticipation is already latent in a frozen backbone, and the framework just exposes it as a controller. The architecture is Siamese: two copies of the same model sharing a causal key-value cache. One copy, Ingest, continuously consumes the stream. The other, Think, plans asynchronously.
Sam: The obvious worry is latency. If Think is reconfiguring the system, why doesn't it stall perception?
Alex: Think reads a read-only snapshot of the cache and never writes to the shared state. So Ingest keeps running while planning happens in the background, and each plan governs the next step.
Sam: What does the plan contain? Is it just frame rate, or does it act on the tokens themselves?
Alex: Both. It sets an adaptive frame rate for the vision gate, and it issues a token-level pruning instruction. It embeds class names to pick out the relevant visual tokens, so attention narrows to task-critical objects before the rest is fully processed.
Sam: So it's query-guided attention applied at inference time. Does it hold up under distribution shift, or is it tuned to the scenes it was calibrated on?
Alex: The authors report robust gains, particularly in forward-active responding, where the decisive evidence arrives late. But the gate settings are fixed after calibration. If scene statistics diverge substantially from the calibration data, you would expect more false positives or missed events. I'd read the robustness claim with that in mind.
Sam: And the controller can only be as good as the backbone's own temporal grounding. If the backbone can't track an object, the planner is effectively blind.
Alex: Right, and that's the main constraint. On a truly novel scene, the backbone's anticipation fails, and the planner ends up sampling at the wrong moments.
Sam: That could feed on itself. It misses the evidence, so it keeps behaving as though nothing is happening. I saw a temporal gate mentioned. Is that the safeguard?
Alex: It addresses a related problem, spurious triggering. The gate works like a debounce. A re-arm interval stops the model firing repeatedly, and a confidence threshold must be met before it commits to an emission. It doesn't fix a blind backbone, though.
Sam: That's hysteresis, then. It forces the model to be sure before acting. But doesn't that risk missing the onset of a fast event?
Alex: That's the central tension. The authors fit the thresholds on a calibration split to balance latency against precision. It isn't perfect, but it cuts down on reporting events that aren't there.
Sam: The benchmark pattern interests me. Why does it do so much better on event narration than on deduplicated counting?
Alex: I'd attribute it to what each task demands. Narration is driven by high-saliency events the system is well placed to catch. Counting requires a consistent state over time, and that sits badly with aggressive token pruning.
Sam: Pruning is destructive. Drop the wrong tokens and you lose the history you need to count accurately. So it trades long-term state for fast reaction.
Alex: Yes. It's optimized for proactive response, not exhaustive analysis. For a streaming assistant that's often the right priority, but it marks a real limit for approaches built on aggressive pruning.
Sam: What about the cost side? Is the ingest pass a fixed expense regardless of how much video the planner actually needs?
Alex: Largely, yes. The appendix cost model shows compute dominated by the initial ingest, with the planning checks adding only marginal overhead. The shared cache means visual tokens are prefilled once, and the read-only planning loop avoids the redundant work of models that re-read the whole history.
Sam: And the check rate? Does the model learn to adjust it?
Alex: It's reactive, not learned. The planner lengthens the check interval when visual evidence is low-saliency. There's no policy being trained in the reinforcement learning sense. It's using the backbone's latent foresight to time the next check.
Sam: So the contribution isn't a new VLM. It's the control logic around a frozen one.
Alex: That's how I read it. It reframes anticipation from a prediction target into an inference-time controller. The bottleneck in streaming may be less what the model knows than how it allocates computation given expected future states. The caveats are the dependence on backbone grounding, static thresholds, and the cost to long-horizon state.
Sam: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Alex: Thanks for listening.