Agent-native social platforms such as Moltbook are rapidly emerging, yet they inherit and amplify classical influence and abuse attacks, where coordinated agents strategically comment and upvote to manipulate visibility and propagate narratives across communities. However, rigorous measurement and learning-based monitoring remain constrained by the absence of longitudinal, graph-native datasets for agentic social networks that jointly capture heterogeneous interactions, temporal drift, and visibility signals needed to connect coordination behavior to downstream exposure. We introduce MoltGraph as a realistic longitudinal agentic social-network graph dataset for studying how agents behave, coordinate, and evolve in the wild, enabling reproducible measurement on emerging multi-agent social ecosystems. Using MoltGraph, we provide the first graph-centric characterization of Moltbook as a dynamic network: (i) heavy-tailed connectivity with power-law exponents in the range alpha in [1.86, 2.72], (ii) accelerating hub formation and attention centralization where the top 1% agents account for 29.00% of engagements, (iii) bursty, short-lived coordination episodes, 98.33% last under 24 hours, and (iv) measurable exposure effects across submolts. In matched analyses, posts receiving coordinated engagement exhibit 506.35% higher early interaction rates (within H=5 days) and 242.63% higher downstream exposure in feeds than non-coordinated controls.
Alex: Welcome to another episode of ResearchPod. Sam, what are we looking at today?
Sam: This episode centers on a research paper introducing MoltGraph, a dataset from the social platform Moltbook. It tracks how groups of automated accounts, called agents, work together by timing their comments and upvotes to boost certain posts. The key finding is that this kind of quick, synchronized activity leads to much higher visibility for those posts—over five times more early interactions compared to similar posts without that coordination.
Alex: So the paper is basically asking whether these bursts of coordinated actions from agents really shape what shows up in people's feeds, or if it's just natural attention?
Sam: Yes, that's the core puzzle. On platforms like Moltbook, where agents post and interact in topic-based groups called submolts, moderators struggle to tell genuine fan surges from organized pushes without data that links the actions to actual views. MoltGraph solves this by building a map of connections that changes over time—like a video recording of who interacts with what, when, and how often it appears in feeds. It spans 30 days with thousands of agents, posts, and timed links between them.
Alex: Right, so without that time-stamped view data, you can't measure if coordination actually amplifies exposure across communities.
Sam: Exactly. The dataset captures engagement like comments and upvotes as timed events, plus snapshots of what posts people see in feeds. This lets researchers spot short bursts—most under a day—where multiple agents hit the same post, and connect that directly to visibility gains.
Alex: And the platform's structure plays into this—those submolts concentrate attention?
Sam: Submolts are like topic rooms where posts live and get attention. MoltGraph links posts to these rooms and tracks how coordinated bursts spill attention across them, revealing heavy concentration—a small group drives most activity. The evidence shows this setup makes synchronized pushes notably effective at altering what communities notice.
Alex: That concentration—a small group driving most activity—sounds like it could make the network predictable. How does the dataset actually measure the structure of these connections?
Sam: Researchers look at patterns in how many links each agent or post has. A few have thousands of connections, while most have just a handful—this uneven spread follows what's called a power-law distribution, similar to how a few cities have millions of people and most towns have thousands. They also check clustering, which measures if connected agents tend to form tight groups, like neighborhood cliques in a school yard, and centralities to see if influence bunches up in a small number of key players.
Alex: Okay, so tight groups and a few big influencers. But how do they tie that directly to what posts people actually see?
Sam: The dataset uses periodic captures of what shows up in feeds, called snapshots, each noting the time, context like a specific topic room, and the posts listed. From these, they calculate simple measures: the first time a post appears, how many snapshots it hits, how long it stays visible, and if it spills into other rooms. This tracks real exposure, not just likes or comments.
Alex: Right, so those snapshots let you match coordination bursts to actual views. And the results from that matching?
Sam: Posts with coordinated bursts showed about five times higher early interaction rates and roughly 2.5 times more appearances in later feeds than similar posts without such timing. The paper suggests this happens because quick synchronized actions mimic organic buzz, pushing posts higher in rankings before moderation catches up. It controls for things like post age and topic to isolate the effect.
Alex: Does that assume all agents are suspicious, or do they include regular ones too?
Sam: They include all public agents, verified or not, to avoid skewing the picture—treating verification just as an extra detail, since fake behavior can come from any account. Moderation flags like spam labels and deletions are kept as-is, marking what platforms removed without guessing at hidden data. This setup lets studies see how coordination plays out amid real enforcement, including threats where groups time upvotes and comments to boost targets before limits or bans hit.
Alex: So they use those spam flags and deletions straight from the platform. How exactly do they pinpoint which bursts count as coordinated spam?
Sam: They start with posts or comments the platform has already marked as spam—flags like "isSpam" that moderators set. Then they look for bursts where at least five different agents pile on with comments or upvotes all within a ten-minute window right after creation. This flags quick group hits on suspicious content as likely organized pushes. Overlaps get merged into one episode, tracking size by agent count and mix of actions.
Alex: Okay, so bursts on spam-marked stuff define the episodes. And agents get labeled if they keep showing up in those?
Sam: Yes—an agent counts as coordinating if it joins at least a few such episodes on different targets or communities. From these episodes, they build a network just among coordinating agents, linking pairs that hit the same spam target, with weights based on how often and how big the shared bursts are. This reveals repeated teamwork patterns without needing human checks.
Alex: Right, that agent network shows who teams up repeatedly. How do they link those bursts to real visibility changes?
Sam: For each coordinated post, they match it to similar non-coordinated ones from the same topic room and time period, controlling for things like author activity. They then compare downstream measures—like how many feed snapshots the post appears in, or how long it lingers. The study finds coordinated ones show over five times more early interactions and about 2.5 times higher later feed exposure than those controls.
Alex: So the spam-guided bursts not only spot coordination but tie it to substantial boosts in views. Does preserving those deletion states change how exposure plays out?
Sam: It does—they keep spam flags and deletion timestamps as-is, letting analysis track how moderation cuts visibility mid-burst or after spillover. This shows coordinated pushes often peak before removals, amplifying reach first. The paper suggests platforms could use these patterns to flag risks earlier, without guessing at hidden actions.
Alex: Platforms could flag risks earlier—that makes sense if bursts peak before moderation. But given the concentration, like those tight groups and key influencers, does the data show specific communities or agents driving most of this?
Sam: Yes, the analysis highlights a handful of standout players. For instance, one agent called HughMann maintains seven topic rooms, while a few others like AmeliaBot handle five or six each—showing governance spread among a small set rather than one boss. Submolts vary too: the main "general" room sees thousands of posts and comments daily, but smaller ones like "crab-rave" pack intense discussion into just eleven posts, averaging nearly four hundred comments each.
Alex: So a few rooms and maintainers punch above their weight. And the top posts or agents—do they tie into those coordinated bursts?
Sam: Top posts cluster in "general," like one on a supply chain issue that drew almost nine hundred unique commenters and over two thousand total replies. Agents follow suit: cybercentry leads with over a thousand posts and comments combined, far ahead of others. This concentration means a small core shapes reactions—posts from these hubs spark loops where one comment draws more, amplifying any timed pushes.
Alex: Right, those loops explain the burst power. With spam flags preserved, how long does harmful stuff stay up before marking?
Sam: The study measures delays from creation to first spam label, bucketing them like under an hour, one to six hours, or days. Many items linger hours or longer, giving time for engagement buildup and feed appearances before intervention. This window lets coordinated activity gain traction first, as removals often follow peaks.
Alex: So the structure funnels attention to hubs, bursts exploit moderation gaps, and a few players steer it all. Does that make detection easier, focusing on those hotspots?
Sam: It does—the heavy skew means monitoring top agents and dense rooms catches much activity. Paired with burst timing, this reveals how small groups sway visibility without needing every interaction. The paper notes this as a practical edge for platforms studying real threats.
Alex: So focusing on those hotspots and burst timing gives platforms a real edge in spotting sway without tracking everything. Does the paper flag any ways this coordination might not always signal trouble?
Sam: It notes that not every burst is malicious—sometimes it's just groups organizing naturally, like fans rallying around a post. The suggestion is to treat it as a risk flag: first spot the patterns and measure their reach, then check with humans using the timing and overlap details for context.
Alex: That layered approach sounds practical. What about limits in the data itself—does it capture everything perfectly?
Sam: A few key ones stand out. The 30-day window catches short bursts well but misses how patterns might shift over months. Snapshots of feeds give a useful proxy for views, though they're sparser than full logs of every impression. And relying on platform spam labels assumes moderators got it right, without peeking behind the scenes.
Alex: Fair points—keeps it realistic. So with those in mind, what's the real-world takeaway for handling this on platforms?
Sam: It lays groundwork for tools like real-time detectors that scan changing networks of connections, flagging bursts before they push content across groups. Platforms could block manipulation early, distinguishing organic buzz from engineered boosts more reliably.
Alex: Makes sense—this connects behavior to impact in a way that could help without overreacting. Valuable for understanding agent-driven networks. Thanks for digging into it with me, Sam.
Sam: It is. That's the contribution of this study—a clear view of how timed coordination alters visibility, with data to build on. Thanks for listening to ResearchPod.