Optimizing vision models purely for classification accuracy can impose an alignment tax, degrading human-like scanpaths and limiting interpretability. We introduce EVA, a neuroscience-inspired hard-attention mechanistic testbed that makes the performance-human-likeness trade-off explicit and adjustable. EVA samples a small number of sequential glimpses using a minimal fovea-periphery representation with CNN-based feature extractor and integrates variance control and adaptive gating to stabilize and regulate attention dynamics. EVA is trained with the standard classification objective without gaze supervision. On CIFAR-10 with dense human gaze annotations, EVA improves scanpath alignment under established metrics such as DTW, NSS, while maintaining competitive accuracy. Ablations show that CNN-based feature extraction drives accuracy but suppresses human-likeness, whereas variance control and gating restore human-aligned trajectories with minimal performance loss. We further validate EVA's scalability on ImageNet-100 and evaluate scanpath alignment on COCO-Search18 without COCO-Search18 gaze supervision or finetuning, where EVA yields human-like scanpaths on natural scenes without additional training. Overall, EVA provides a principled framework for trustworthy, human-interpretable active vision.
Alex: Welcome to another episode of ResearchPod. Sam, I've noticed that AI systems for looking at images and spotting objects often get things right, but their "viewing patterns" don't match how people actually scan a picture with their eyes—what's going on there?
Sam: This paper from researchers at the University of Tokyo introduces a system called EVA to tackle that issue. When you train AI vision models to classify images accurately—like spotting a dog—they often stop jumping around like human eyes do to gather clues. The authors call this an "alignment tax," where better accuracy makes the AI's looking patterns less human-like and harder to inspect.
Alex: So the paper asks why accurate AI vision doesn't mimic human eye movements, and proposes a fix? And that matters in real situations, like checking medical scans, where we need to see what evidence the AI used?
Sam: Yes. Humans can't process a whole image at full detail at once, so our eyes make quick jumps called saccades to focus on key spots one by one. AI models that copy this—hard-attention models—leave a trail of focus points, or scanpaths, that people can follow. But pushing for top accuracy usually breaks that trail, turning the AI into a black box.
Alex: Right, so the challenge is designing an AI that stays accurate and keeps human-style scanpaths, without extra training data on where humans look.
Sam: Precisely. EVA uses brain-inspired tweaks to make that tradeoff adjustable, trained only on basic labels like "dog" or "cat," no gaze data needed. On datasets with human eye-tracking, it produces scanpaths much closer to people's while holding strong accuracy.
Alex: Okay, so EVA holds accuracy while matching human scanpaths better. But inside, how does it process each glimpse—what makes one spot high-detail versus background?
Sam: Humans see sharp detail only straight ahead, with fuzzy edges around the sides, so the eye grabs a crisp center patch and rough surround for context. EVA copies that: a detailed crop from the focus point goes through a strong network, while the wider blurry area uses a simple encoder. It's like zooming your phone camera on one part of a photo while seeing the rest dimly.
Alex: And the training—how does it learn where to look next without human eye data?
Sam: It uses a reward setup where getting the final label right sends a thumbs-up signal back through time. The system picks spots based on past glimpses, adjusts to favor good paths—like a game where winning teaches better moves over many tries. They call this REINFORCE. A separate loss fine-tunes class predictions, and extras encourage smart info flow.
Alex: So rewards shape the looking path, losses sharpen guesses. What happens if you pull out pieces, like that gate you mentioned?
Sam: Ablations show the detailed crop network drives most accuracy gains—without it, performance drops sharply. The gate and variance tweaks mainly boost human-like paths with little accuracy hit. On CIFAR-10, full EVA tops hard-attention rivals in accuracy while staying competitive on path similarity.
Alex: So the trade-off is tunable—stronger features help spotting but need controls to keep paths natural. That explains the visuals too, fixations hitting faces or wheels like humans.
Sam: Yes, and it scales: on bigger images, EVA uses few glimpses efficiently versus dense scanning rivals. Cross-task, paths transfer to search without retraining. The paper suggests this makes AI vision more auditable, revealing evidence steps we can trace.
Alex: You mentioned cross-task transfer—paths working for search without retraining. How well does that hold up on COCO-Search18, where humans hunt for specific objects?
Sam: COCO-Search18 has human eye tracks from people hunting for objects like cars in photos. EVA, trained just on basic labels from COCO, matches those human search paths about as well as models built specially for gaze prediction—using measures that line up paths like stretching rubber bands.
Alex: So no special eye-training data, yet it transfers to search tasks decently. What do the internal states reveal—are the brain-like layers doing distinct jobs?
Sam: Researchers plotted the model's hidden activity using a tool that spots main patterns, like summarizing a video by its key motions. The upper layer clusters by object type, gathering proof; the lower handles where to look next, keeping local details for movement.
Alex: That separation makes the process traceable. But the paper flags some limits too, right?
Sam: Yes, training stays finicky to tweaks in optimization, so stability could improve. EVA prioritizes useful brain-like shortcuts over exact biology, and gaze data has biases like center-looking that need fixes. Classification training doesn't perfectly match search behaviors.
Alex: Fair points—these controls make the trade-off adjustable, but real-world tweaks are next. So the scanpaths aren't just pretty trails—they carry info about how the model builds its decision step by step. What does looking inside one reveal?
Sam: At each step, the model only sees evidence from glimpses so far, so its guess starts rough and shifts with new info. Early on, it might lean one way; later glimpses can flip that. This shows the scanpath as a timeline of belief changes, unlike a static heat map without order.
Alex: Like watching a detective narrow suspects as clues pile up, instead of just a list of places checked.
Sam: Yes, and order matters. The paper tests this by shuffling the model's fixation spots—keeping the same places but scrambling the sequence. Performance drops sharply for EVA, proving timing shapes the final call. Other models barely budge.
Alex: So shuffling hurts EVA more because its strategy relies on sequence. Those tests—like forcing center-only looks—back up the paths doing real work?
Sam: They do. Center-fixed or corner-only gazing tanks accuracy, as expected. Shuffling a good path hurts EVA about twice as much as rivals. This confirms scanpaths aren't decorative; they're a process we can probe.
Alex: That strengthens the case for these models in clinics or safety checks—trace not just where, but the buildup. A solid way to build trust.
Sam: Precisely. By making trade-offs explicit through tweaks like those controls, it offers a clear path to auditable active vision.
Alex: That's a grounded take on bridging AI performance with human-style reasoning. Thanks, Sam—this has clarified why traceable paths matter. Thanks for listening to ResearchPod.