Taorui Huang, Namita Gaidhani, Ritvik Bansal, S M Jubaer, Regina Lim, Rhett Zhao, Gavin Rafael Selin, Sunnie Deng Gao, Hasnain N Syed
5 min
Abstract
In our prior work, Pedestrian Archetypes, we defined pedestrian archetypes as collections of behaviors that uniquely identify a specific type of pedestrian. The first paper proposed 12 pedestrian archetypes, including the Wanderer, Drunk, Distracted, Flash, Indecisive, Blind, Flock, Jaywalker, Elderly, Kid, Eventful, and Parked Pedestrian. These archetypes were introduced to move beyond single behavior labels and provide a more natural way to describe how dangerous pedestrians actually behave progressively in real-world traffic scenarios. However, upon further annotation of YouTube dash-cam videos, we identified 7 additional pedestrian archetypes with observable and significant behavioral differences from the previously proposed ones. These new archetypes capture pedestrian behavior patterns that could not be fully explained by the original taxonomy. In this pre-print, we introduce each new archetype, define its essential and optional behaviors, explain how it differs from previously proposed archetypes, and provide video-frame evidence showing the archetype in action.
Alex: Are there archetypes that are more deliberately disruptive?
Sam: Several. The "Influencer" stops in traffic to film content. The "Protester" occupies a lane on purpose. Both are intentionally ignoring traffic flow—but their goal is different from someone who's simply distracted. And that goal matters, because it changes what the car should expect them to do next.
Alex: So the car isn't just asking "will they cross?"—it's asking "what are they trying to accomplish?"
Sam: That's the core shift. The "Street Vendor," for instance, is actually trying to get closer to the car—to make a sale. Most pedestrians move away from approaching vehicles. Recognizing that difference in intent changes how the car should respond. Moving away slowly might not be a threat. Moving closer deliberately is a different situation entirely.
Alex: Where does this data come from? How did researchers actually build these archetypes?
Sam: Largely from sources like YouTube dash-cam footage—which is where a significant limitation comes in. That kind of footage tends to capture dramatic, unusual moments. It probably doesn't represent the subtle, everyday risks of normal traffic. The researchers acknowledge this. The archetypes may skew toward rare, extreme cases rather than the quiet ambiguity of an ordinary street crossing.
Alex: So this isn't meant to replace how a car sees the world—it's more like a specialized layer for the most unpredictable situations?
Sam: That's a fair reading. One of the practical applications the paper highlights is testing. Engineers can simulate a "Con Artist" or a "Pseudo Pedestrian"—someone who approaches a car but isn't actually crossing—thousands of times in a virtual environment, without putting anyone at risk. That kind of structured testing is hard to do when you're just working with generic "pedestrian" models.
Alex: It's a meaningful idea—teaching machines to interpret the intent behind movement, not just the movement itself.
Sam: And it points to something broader. A lot of human safety on the road depends on unspoken social cues—eye contact, hesitation, body language. We read those instinctively. Encoding them into a structured system is genuinely difficult, but this paper is an attempt to at least name them clearly enough to work with.
Alex: So the goal is to move from a car that sees "moving objects" to one that can ask: what does this person want, and what are they likely to do next?
Sam: That's it. By structuring the unpredictable—giving it names, patterns, and testable definitions—the hope is that we make these systems more robust, and ultimately, the road safer for everyone.
Alex: Thanks for walking me through it, Sam. And thank you to our listeners for joining us on ResearchPod.