ResearchPod Summary
Understanding the prevalence of vulnerability—such as mental ill health, substance misuse, alcohol dependence, and homelessness—in police interactions is critical for resource allocation and policy development. However, structured administrative data are often incomplete or inconsistently recorded. This study investigates whether Large Language Models (LLMs) can extract these indicators from unstructured, free-text police incident logs. The researchers developed a multi-stage pipeline using a locally-hosted 8B parameter model to ensure data security. The process involved a two-step classification method: an initial screening to filter out irrelevant logs, followed by a nuanced assessment of vulnerability indicators. To ensure robustness, the team used knowledge distillation to train the model, performed repeated inferences for each log to measure stability, and applied statistical corrections to mitigate systematic biases identified through human review.
The study demonstrates that LLMs can generate meaningful prevalence estimates at scale, identifying mental ill health in approximately one in five incidents. However, the researchers emphasize that naive deployment is unreliable. Single-pass classifications were found to be unstable, and the model exhibited a systematic tendency to over-assign vulnerability indicators compared to human judgment. Achieving defensible results required substantial human oversight and statistical adjustment, highlighting that while LLMs are powerful tools for information extraction, they cannot currently replace human analysis for operational decision-making or individual-level assessments.
This research provides a realistic assessment of the potential and limitations of using AI to analyze sensitive public sector data. It underscores that while LLMs can help bridge the gap between unstructured narrative data and policy-relevant insights, they are not "plug-and-play" solutions. The findings suggest that for police forces and other public agencies, the path to using AI for evidence-based policy requires significant investment in methodological rigor, human-in-the-loop validation, and careful handling of model uncertainty.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.