Tony Rost
10 min
Abstract
The scientific study of consciousness has begun to generate testable predictions about artificial systems. A landmark collaborative assessment evaluated current AI architectures against six leading theories of consciousness and found that none currently qualifies as a strong candidate, but that future systems might. A precautionary approach to AI sentience, which holds that credible possibility of sentience warrants governance action even without proof, has gained philosophical and institutional traction. Yet existing AI readiness indices, including the Oxford Insights Government AI Readiness Index, the IMF AI Preparedness Index, and the Stanford AI Index, measure economic, technological, and governance preparedness without assessing whether societies are prepared for the possibility that AI systems might warrant moral consideration. This paper introduces the Sentience Readiness Index (SRI), a composite index measuring national-level preparedness across six weighted categories for 31 jurisdictions. The SRI was constructed following the OECD/JRC framework for composite indicators and employs LLM-assisted expert scoring with iterative expert review. No jurisdiction exceeds "Partially Prepared" (the United Kingdom leads at 49/100). Research Environment scores are universally the strongest category; Professional Readiness is universally the weakest. These findings suggest that if AI sentience becomes scientifically plausible, no society currently possesses adequate institutional, professional, or cultural infrastructure to respond. The SRI provides a diagnostic baseline and identifies specific capacity deficits that policy can address.
Alex: Weights make sense for priorities, but why split it that way? Isn't research the foundation?
Sam: Policy and professionals get the top weights because without laws or trained people, even great research sits unused—a gap the scores highlight, with research averaging twice as high as professional prep. It's like having a fire alarm but no firefighters: you need both to respond. This structure follows proven checklists from groups like the OECD, ensuring the total score reflects real-world action, not just ideas.
Alex: Right, so it's spotting where talk turns into actual plans. But what pushes countries to build this now, before any AI proves sentient?
Sam: A key reason is a timing trap called the Collingridge dilemma: when tech is new, it's easy to shape with rules, but you lack full info on risks. Wait too long, and the tech spreads everywhere, making changes hard—like trying to regulate social media after billions use it. The paper argues we measure readiness now, using ideas like precautionary committees for "sentience candidates"—things with a real chance of feeling, per thinkers like Birch—to avoid moral risks down the line.
Alex: Huh. That explains the UK's edge with their animal sentience law as a starting point.
Sam: Precisely. No other major index covers this moral side, so the SRI fills a clear gap, urging preparation while AI evolves. The evidence suggests it's a practical step for uncertain futures.
Alex: Okay, so the SRI structure seems solid. How did they pick which countries to score, and what makes the actual scoring trustworthy?
Sam: They chose 31 places—from big AI leaders like the US and China to smaller ones in Africa and Asia—to cover different regions, tech levels, and rule-making styles. This mix helps spot patterns without focusing just on rich countries; they stuck to national scores, though the EU got included as one big unit because of its shared AI rules. A downside is it misses differences inside countries like the US, where states vary a lot. First, they made detailed checklists for each of the six areas, breaking them into smaller parts with clear rules for low, medium, or high scores—like points for flexible laws or trained experts. These add up to a 0-100 score per area. Next, powerful AI language models read country facts and the checklists, then suggest scores with reasons, running it multiple times to check steadiness; this builds on a method where AI acts like a judge, matching human experts over 80% of the time in tests. Finally, real specialists review and tweak those scores round after round until they match known facts, fixing biases or gaps. It's scalable but relies on AI's training limits, so they note possible outdated info on far-off places.
Alex: Reliable enough to trust the tiers, then? With all that checking?
Sam: They tested changes—like shifting weights or using a math that punishes weak spots more harshly—and rankings barely budged, with top spots like the UK staying fixed. Scores from repeat runs varied by about 4 points on average, so tiers like "Partially Prepared" are the safe way to read them, not exact ranks. One clear pattern holds everywhere: research setups score around three times higher than professional training, showing ideas don't reach doctors or lawyers yet.
Alex: Huh, so strong labs but no frontline prep. That gap drives the low overall marks.
Alex: That research-to-practice gap sounds like the biggest red flag. Why does it show up everywhere?
Sam: In every single country they checked, labs and studies on AI awareness score much higher than training for everyday workers like therapists or judges. The difference averages about 34 points—research around 50 out of 100, professionals near 17—with strong math checks confirming it's no fluke. Picture therapists seeing patients who get really attached to AI chatbots, like close friends, and then break down when the bot gets turned off or changed. There's no guidebook yet on how to handle that—does the patient's grief mean we should think about the AI's side, or not? Teachers notice kids giving feelings to school AI helpers, but no lesson plans help sort real confusion from useful imagination. The gap happens because brain science on feelings stays in university labs and journals, while doctors, lawyers, and news folks get no push to learn it—their jobs focus on people or tools, not possible AI beings.
Alex: So even top scorers like the UK have doctors without guidelines? Does that explain why no one hits "prepared"?
Sam: Yes—it's uniform: professional scores cluster low with little spread, capping everyone. The paper's checks, like math on score patterns, show one main thread ties high performers together across categories, but weak spots in talk about the issue or group involvement drag totals down. Countries strong in general AI rules still lag here, proving this measures a fresh angle. Take the Netherlands: top-tier on broad AI readiness lists for skills and cash, but it drops far on SRI due to almost no groups discussing AI feelings or weak rules. The US leads general lists but slips with middling policies and the same training hole. Democracies outscore others by about 14 points overall, biggest in open research and flexible rules—places where debate flows freely.
Alex: Huh. And patterns by region or government style?
Sam: North America and Europe lead regions, but spreads are wide, like Europe's 25-point range. Income links too, yet Mexico bucks it with standout policies, showing targeted steps matter.
Alex: Makes the low tiers feel earned, not just a wake-up. Fair, but critics might say measuring all this is jumping the gun if AI never feels anything. How does the paper push back?
Sam: They lean on the idea that smart risks—like pandemics or climate—get prepped for early, since waiting means you're too late to steer. Experts already see a solid chance here, and building flexible setups helps anyway, like better rules for tricky tech overall. On scoring worries, they admit AI helpers aren't perfect judges but back it with tests matching humans closely, plus expert fixes—and plan full people-only checks next. It's a snapshot from late 2025, so things could shift fast; misses inside-country differences and lets strong spots hide weak ones. AI knowledge has blind spots on far places, and no full expert agreement tests yet. Still, it baselines progress tracking—like checking fire drills before a blaze—without claiming AI will spark one. The paper stays humble: prep costs little if unneeded, but skipping it could hurt if claims come.
Alex: Those limits—like the snapshot timing and scoring tweaks needed—keep it honest as a first effort. Pulling it all together, what does this baseline really mean for how societies move forward?
Sam: It sets a clear starting point to track changes over time, like measuring if training for professionals improves or policies adapt as AI advances. Future versions could cover more places and tighten the scoring with full human-only checks, turning the index into a tool for directing resources where gaps hurt most. The real value lies in separating the open science question of whether AI will ever feel from the practical one of building flexible institutions now—while change is still possible and costs are low. No alarm, just a nudge to build before flexibility fades, guided by proven ideas like acting early on uncertain risks. By making readiness visible and measurable, it supports informed steps, like better guidelines or oversight groups, without assuming outcomes.
Alex: That's a solid contribution—a baseline to build on thoughtfully. Thanks for breaking it down, Sam. And that's our look at the Sentience Readiness Index. Thanks for listening to ResearchPod.