Using Fine-Tuned LLMs to Identify Indicators of Vulnerability in UK Police Incident Logs

summary

Video file (mp4)

The gist

This research explores whether fine-tuned Large Language Models (LLMs) can be adapted to estimate the prevalence of four key vulnerability indicators—mental ill health, substance misuse, alcohol

In short

Researchers tested if fine-tuned Large Language Models (LLMs) could estimate the prevalence of mental ill health, substance misuse, alcohol dependence, and homelessness from unstructured UK police incident narratives. The study used nearly 3,000 logs to create estimates like 8% for alcohol dependence. While LLMs can provide scalable estimates for policy planning, their individual predictions are imperfect and require human oversight.

Key concepts

LLM Classification Pipeline
A two-step process where an LLM first highlights text snippets potentially related to a vulnerability and then assesses the evidence found using a three-level scale: STRONG (direct discussion), MODERATE (most likely explanation), or WEAK (alternative explanations). This structured approach helps turn unstructured text into quantifiable data.
Knowledge Distillation
A training technique used to adapt smaller, local LLMs. A larger, powerful 'teacher' model (like GPT-4o-mini) is used to train and guide smaller models (8B, 3B, 1B parameters) by teaching them how to follow specific classification instructions on a dataset.
Monte Carlo Simulation
A statistical method used to correct initial raw prevalence estimates. Because the LLM outputs were initially overestimated, this simulation uses human review data and transition probabilities between model labels and true labels to generate more defensible, adjusted prevalence measurements.

Terminology used across episodes

This episode discusses

The paper

Using Fine-Tuned LLMs to Identify Indicators of Vulnerability in UK Police Incident Logs · Read on arXiv

ESRC Vulnerability and Policing Futures Research Centre · School of Law, University of Leeds

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Using Fine-Tuned LLMs to Identify Indicators of Vulnerability in UK Police Incident Logs".

Jane: This research explores whether fine-tuned Large Language Models (LLMs) can be adapted to estimate the prevalence of four key vulnerability indicators—mental ill health, substance misuse, alcohol dependence,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, we’re starting with the title "Using Fine-Tuned LLMs to Identify Indicators of Vulnerability in UK Police Incident Logs," which tells us exactly what this research is aiming at—using fine-tuned models to spot vulnerability indicators inside police logs. The authors are Sam Relinsa and Daniel Birksa from the University of Leeds, and they’re looking at how these narrative analyses might change our understanding of frontline policing demands.

Jane: That title really captures the core idea: taking those long, descriptive incident reports and using an AI to pinpoint specific issues like substance misuse or alcohol dependence that might be hidden in plain sight. It’s about making sense of the volume and variety of text police officers deal with every day.

Lu: The authors are specifically testing if an LLM pipeline can estimate four key indicators—mental ill health, substance misuse, alcohol dependence, and homelessness—within unstructured UK police incident narratives using nearly three thousand de-identified logs. That scope is quite broad for a single study; it’s trying to cover a wide spectrum of social needs impacting officers.

Meng: I'm interested in the methodology mentioned on page one, which involves a multi-stage pipeline combining repeated model inference, label aggregation, structured human review, and statistical correction to assess output stability. How did they manage the computational demands of running these multiple steps on those large texts?

Lalam: That multi-stage approach is really smart because it doesn't just give you one answer; it checks for consistency across different model runs and then verifies those findings with human input, which builds a much more reliable output.

The paper's summary: Tom: The main summary of the paper explains that they are exploring whether an LLM-based classification pipeline can estimate the prevalence of these four vulnerability indicators within UK police incident narratives, and crucially, when those outputs can be considered defensible measurements. They aren't just trying to guess; they’re trying to establish if the AI’s output is actually useful for informing resourcing decisions.

Jane: Exactly, Tom; the paper summarizes that while administrative data has limitations, this LLM approach uses narrative text to explore if we can get better insights into the actual needs of vulnerable people encountered by officers. It moves us from just looking at numbers to understanding the context behind those numbers.

Lu: The summary highlights a prior study they referenced where an instruction-tuned LLM was found to be highly effective at identifying narratives where vulnerabilities were absent, acting as a reliable negative filter, which sets a good baseline for evaluating how well this current pipeline filters out noise.

Meng: That mention of the prior study is important because it suggests they are building on existing knowledge about LLM capabilities in qualitative coding before applying it to this specific, sensitive domain. I wonder how much bias from that previous work might carry over into their current system design.

Lalam: The summary really emphasizes the goal: informing resourcing and policy debates regarding frontline policing demands, which is where the real-world impact lies—if we can quantify these needs better, we can advocate for better support structures.

The paper's improvements: Tom: Now that we know what they did, the paper points out specific improvements they made to their pipeline. They addressed issues like systematic overestimation in certain areas, especially for substance misuse and alcohol dependence, by employing a Monte Carlo simulation to adjust for known classification errors using transition probabilities derived from human review data.

Jane: That statistical correction step sounds really important because it acknowledges that the initial model outputs aren't perfect, so they have a mathematical way to correct the errors based on what humans actually found during review. It’s like applying a quality control check after the AI has made its initial pass.

Lu: They also focused heavily on adapting models for this task using knowledge distillation, training smaller local models like 8B and 3B parameters with LoRA adapters, using GPT-4o-mini as a teacher model to teach them how to follow those structured classification instructions. That’s a clever way to make the model efficient enough for secure compute environments.

Meng: From an engineering perspective, distilling knowledge from a larger model into smaller ones makes sense for deployment; it reduces the computational footprint while hopefully retaining much of the performance needed for this specific classification task. I'm curious if that distillation process introduced any new types of error or instability compared to just fine-tuning directly.

Lalam: The focus on distilling knowledge means they are creating a system that is lightweight enough to run securely where it needs to be, while still being trained specifically on the required logic for these four vulnerability indicators. It makes the tool practical for deployment rather than just an academic exercise.

Conclusion: Tom: So, wrapping up on "Using Fine-Tuned LLMs to Identify Indicators of Vulnerability in UK Police Incident Logs," the authors conclude that while LLMs can produce meaningful, if imperfect, prevalence estimates at scale, they aren't valid measurements without careful methodological support. They stress that errors remain frequent and unpredictable at the individual level.

Jane: It sounds like the main implication is a clear boundary: we can get large-scale population-level estimates for planning purposes, but we have to be extremely cautious when using those numbers to make specific decisions about an individual case because of those inherent uncertainties.

Lu: The future work mentioned suggests that the real value lies in applying this at a greater scale to inform demand modelling and strategic planning, which is where the paper places its primary contribution rather than trying to achieve perfect individual accuracy.

Meng: I think for practical implementation, it means we should treat these prevalence estimates as high-level indicators for resource allocation strategy, not as definitive diagnostic tools for any single officer’s interaction with someone. It sets realistic expectations for what the system can reliably deliver in a high-stakes environment.

Lalam: I feel this work is incredibly valuable because it proves that AI can be a tool to supplement, and not replace, human judgment when dealing with such complex social issues on the ground; it gives us a quantitative voice for those difficult conversations.

Tom: That’s a powerful way to put it; we've seen how the paper "Using Fine-Tuned LLMs to Identify Indicators of Vulnerability in UK Police Incident Logs" uses narrative text and statistical correction to generate estimates for mental ill health, alcohol dependence, substance misuse, and homelessness. What an interesting piece of research.

More episodes

← Home