Pengcheng Pan, Yonekura Shogo, Kuniyoshi Yasuo
6 min
Abstract
Optimizing vision models purely for classification accuracy can impose an alignment tax, degrading human-like scanpaths and limiting interpretability. We introduce EVA, a neuroscience-inspired hard-attention mechanistic testbed that makes the performance-human-likeness trade-off explicit and adjustable. EVA samples a small number of sequential glimpses using a minimal fovea-periphery representation with CNN-based feature extractor and integrates variance control and adaptive gating to stabilize and regulate attention dynamics. EVA is trained with the standard classification objective without gaze supervision. On CIFAR-10 with dense human gaze annotations, EVA improves scanpath alignment under established metrics such as DTW, NSS, while maintaining competitive accuracy. Ablations show that CNN-based feature extraction drives accuracy but suppresses human-likeness, whereas variance control and gating restore human-aligned trajectories with minimal performance loss. We further validate EVA's scalability on ImageNet-100 and evaluate scanpath alignment on COCO-Search18 without COCO-Search18 gaze supervision or finetuning, where EVA yields human-like scanpaths on natural scenes without additional training. Overall, EVA provides a principled framework for trustworthy, human-interpretable active vision.
Alex: So rewards shape the looking path, losses sharpen guesses. What happens if you pull out pieces, like that gate you mentioned?
Sam: Ablations show the detailed crop network drives most accuracy gains—without it, performance drops sharply. The gate and variance tweaks mainly boost human-like paths with little accuracy hit. On CIFAR-10, full EVA tops hard-attention rivals in accuracy while staying competitive on path similarity.
Alex: So the trade-off is tunable—stronger features help spotting but need controls to keep paths natural. That explains the visuals too, fixations hitting faces or wheels like humans.
Sam: Yes, and it scales: on bigger images, EVA uses few glimpses efficiently versus dense scanning rivals. Cross-task, paths transfer to search without retraining. The paper suggests this makes AI vision more auditable, revealing evidence steps we can trace.
Alex: You mentioned cross-task transfer—paths working for search without retraining. How well does that hold up on COCO-Search18, where humans hunt for specific objects?
Sam: COCO-Search18 has human eye tracks from people hunting for objects like cars in photos. EVA, trained just on basic labels from COCO, matches those human search paths about as well as models built specially for gaze prediction—using measures that line up paths like stretching rubber bands.
Alex: So no special eye-training data, yet it transfers to search tasks decently. What do the internal states reveal—are the brain-like layers doing distinct jobs?
Sam: Researchers plotted the model's hidden activity using a tool that spots main patterns, like summarizing a video by its key motions. The upper layer clusters by object type, gathering proof; the lower handles where to look next, keeping local details for movement.
Alex: That separation makes the process traceable. But the paper flags some limits too, right?
Sam: Yes, training stays finicky to tweaks in optimization, so stability could improve. EVA prioritizes useful brain-like shortcuts over exact biology, and gaze data has biases like center-looking that need fixes. Classification training doesn't perfectly match search behaviors.
Alex: Fair points—these controls make the trade-off adjustable, but real-world tweaks are next. So the scanpaths aren't just pretty trails—they carry info about how the model builds its decision step by step. What does looking inside one reveal?
Sam: At each step, the model only sees evidence from glimpses so far, so its guess starts rough and shifts with new info. Early on, it might lean one way; later glimpses can flip that. This shows the scanpath as a timeline of belief changes, unlike a static heat map without order.
Alex: Like watching a detective narrow suspects as clues pile up, instead of just a list of places checked.
Sam: Yes, and order matters. The paper tests this by shuffling the model's fixation spots—keeping the same places but scrambling the sequence. Performance drops sharply for EVA, proving timing shapes the final call. Other models barely budge.
Alex: So shuffling hurts EVA more because its strategy relies on sequence. Those tests—like forcing center-only looks—back up the paths doing real work?
Sam: They do. Center-fixed or corner-only gazing tanks accuracy, as expected. Shuffling a good path hurts EVA about twice as much as rivals. This confirms scanpaths aren't decorative; they're a process we can probe.
Alex: That strengthens the case for these models in clinics or safety checks—trace not just where, but the buildup. A solid way to build trust.
Sam: Precisely. By making trade-offs explicit through tweaks like those controls, it offers a clear path to auditable active vision.
Alex: That's a grounded take on bridging AI performance with human-style reasoning. Thanks, Sam—this has clarified why traceable paths matter. Thanks for listening to ResearchPod.