In our prior work, Pedestrian Archetypes, we defined pedestrian archetypes as collections of behaviors that uniquely identify a specific type of pedestrian. The first paper proposed 12 pedestrian archetypes, including the Wanderer, Drunk, Distracted, Flash, Indecisive, Blind, Flock, Jaywalker, Elderly, Kid, Eventful, and Parked Pedestrian. These archetypes were introduced to move beyond single behavior labels and provide a more natural way to describe how dangerous pedestrians actually behave progressively in real-world traffic scenarios. However, upon further annotation of YouTube dash-cam videos, we identified 7 additional pedestrian archetypes with observable and significant behavioral differences from the previously proposed ones. These new archetypes capture pedestrian behavior patterns that could not be fully explained by the original taxonomy. In this pre-print, we introduce each new archetype, define its essential and optional behaviors, explain how it differs from previously proposed archetypes, and provide video-frame evidence showing the archetype in action.
Alex: Welcome to another episode of ResearchPod. Today, we're looking at how self-driving cars can better handle one of their trickiest challenges: unpredictable human behavior on the road.
Sam: The core problem is this: self-driving cars are very good at calculating physics—speed, distance, stopping time. What they struggle with is human irrationality. Someone might step into traffic on purpose, or freeze up unexpectedly, or try to trick the car entirely. This paper argues that to handle those situations, we need to give cars a way to recognize types of people and types of behavior, not just track moving objects.
Alex: So rather than just seeing "a person is in the road," the car tries to figure out why they're there?
Sam: Exactly. The researchers argue that isolated actions aren't enough. A person standing in a lane could be confused, protesting, or staging a fake accident. The same physical position means something completely different depending on context. So they propose that behaviors form patterns—and those patterns define what they call an "archetype," a kind of character class for pedestrians.
Alex: Like character classes in a video game? A rogue plays differently than a tank, even if they're standing in the same spot?
Sam: That's a solid way to put it. Once the car identifies which class a person belongs to, it can anticipate their likely next move—and adjust how cautiously it responds.
Alex: So how does the car actually recognize these classes? What's it looking at?
Sam: The researchers built what they call an "ontology"—think of it as a standardized dictionary of behavior tags. Things like "flinching," "retreating," "approaching," or "standing still in a lane." On their own, those tags don't mean much. But the system looks for which tags tend to show up together. That combination—what they call a statistical co-occurrence—is what defines an archetype.
Alex: So the ontology is the dictionary, and the archetypes are the characters built from it?
Sam: Exactly. And they've defined nineteen of them. Take the "Con Artist"—someone who stages a fake collision to claim insurance money. Because they want to avoid real injury, they typically wait until a car is nearly stopped before jumping in front of it. The system recognizes that pattern: standing in a lane, then jumping, then falling—movements that are statistically unlikely to happen together by accident.
Alex: How does it tell that apart from someone who just tripped?
Sam: That's where the co-occurrence matters. Someone who trips doesn't tend to wait first. The timing and sequence of movements is different. It's not any single action—it's the combination and order that flags it as intentional.
Alex: Are there archetypes that are more deliberately disruptive?
Sam: Several. The "Influencer" stops in traffic to film content. The "Protester" occupies a lane on purpose. Both are intentionally ignoring traffic flow—but their goal is different from someone who's simply distracted. And that goal matters, because it changes what the car should expect them to do next.
Alex: So the car isn't just asking "will they cross?"—it's asking "what are they trying to accomplish?"
Sam: That's the core shift. The "Street Vendor," for instance, is actually trying to get closer to the car—to make a sale. Most pedestrians move away from approaching vehicles. Recognizing that difference in intent changes how the car should respond. Moving away slowly might not be a threat. Moving closer deliberately is a different situation entirely.
Alex: Where does this data come from? How did researchers actually build these archetypes?
Sam: Largely from sources like YouTube dash-cam footage—which is where a significant limitation comes in. That kind of footage tends to capture dramatic, unusual moments. It probably doesn't represent the subtle, everyday risks of normal traffic. The researchers acknowledge this. The archetypes may skew toward rare, extreme cases rather than the quiet ambiguity of an ordinary street crossing.
Alex: So this isn't meant to replace how a car sees the world—it's more like a specialized layer for the most unpredictable situations?
Sam: That's a fair reading. One of the practical applications the paper highlights is testing. Engineers can simulate a "Con Artist" or a "Pseudo Pedestrian"—someone who approaches a car but isn't actually crossing—thousands of times in a virtual environment, without putting anyone at risk. That kind of structured testing is hard to do when you're just working with generic "pedestrian" models.
Alex: It's a meaningful idea—teaching machines to interpret the intent behind movement, not just the movement itself.
Sam: And it points to something broader. A lot of human safety on the road depends on unspoken social cues—eye contact, hesitation, body language. We read those instinctively. Encoding them into a structured system is genuinely difficult, but this paper is an attempt to at least name them clearly enough to work with.
Alex: So the goal is to move from a car that sees "moving objects" to one that can ask: what does this person want, and what are they likely to do next?
Sam: That's it. By structuring the unpredictable—giving it names, patterns, and testable definitions—the hope is that we make these systems more robust, and ultimately, the road safer for everyone.
Alex: Thanks for walking me through it, Sam. And thank you to our listeners for joining us on ResearchPod.