AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement

summary

Video file (mp4)

The gist

AnchorSIPS is a synthetic dataset and evaluation resource for evidence-supported psychosis-risk symptom measurement, consisting of 10,000 synthetic structured-interview bundles.

In short

The episode discusses AnchorSIPS, a synthetic dataset of ten thousand psychosis-risk interviews designed to test AI systems. Hosts analyze findings showing that while current large language models can predict diagnoses, they fail at citing specific evidence from the transcript to support their claims. The resource aims to measure true evidence-supported reasoning.

Key concepts

AnchorSIPS
A synthetic dataset containing ten thousand structured interviews for assessing psychosis-risk symptoms (SIPS). It was created because real psychiatric data is too sensitive to share, allowing researchers to train AI systems without privacy concerns.
Structured Interview for Prodromal Syndromes (SIPS)
A method used by clinicians to assess whether an individual may be at risk for developing psychosis. The dataset uses this structure to simulate the detailed clinical assessments required in real-world settings.
Synthetic Dataset
Fake, yet highly realistic, data created using a 'plan-then-realize' pipeline. This method ensures the ground truth (labels and diagnoses) is fixed by logic before an LLM generates dialogue, making the results auditable and reproducible.
Evidence Linking
The ability of an AI model to cite the exact lines of dialogue or transcript excerpts that support a specific clinical decision or diagnosis. The episode notes this is the critical bottleneck for current models.

Terminology used across episodes

This episode discusses

The paper

AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement · Read on arXiv

Guilherme C. Oliveira, Stephanie Fong, Zimu Wang, Clarice Lee, Xiangyu Zhao, Duy Khoa Pham, Duong Nhu, Yiwen Jiang, Jiahe Liu, Zhongxing Xu, Dwarikanath Mahapatra, Dominic Dwyer, Zongyuan Ge

Monash University · Orygen and The University of Melbourne · Swinburne University of Technology · University of Liverpool · Khalifa University

Progress on AI for psychosis-risk assessment is limited by a data-access bottleneck. Real clinical interviews are difficult to share because of privacy, governance, and consent constraints. We present AnchorSIPS, a synthetic dataset of 10K structured psychosis-risk interviews with transcript-grounded measurement targets. Each interview is modeled on Mini-SIPS, a clinician-administered psychosis-risk interview. It captures history, 24 symptom questions, follow-up evidence for items the patient affirms, decisions about delusion-like symptoms (unusual beliefs), hallucination-like symptoms (unusual perceptions), and disorganized communication, exclusion of clear psychotic-level symptoms ("frank psychosis"), and a final attenuated psychosis syndrome (APS) diagnosis, a high-risk state of milder or early psychotic symptoms. The APS diagnosis is not a standalone label. It depends on earlier endorsements, supporting follow-up details, symptom-class decisions, and the frank-psychosis check. Every intermediate decision is anchored to its supporting transcript turns. AnchorSIPS is generated by a plan-then-realize pipeline. A hidden case sheet specifies the patient's clinical state, a deterministic planner fixes the interview structure, and an LLM realizes only the patient utterances under validation and bounded repair. Fixing labels and structure before generation avoids the inter-turn inconsistencies typical of multi-turn LLM dialogue. Across seven LLM baselines, models recover coarse decisions but fail to extract follow-up details or cite supporting transcript turns, so final-label performance overstates interview competence. AnchorSIPS is intended for research on evidence extraction, transcript-grounded measurement, and uncertainty under partial disclosure.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement".

Jane: The paper was written by Guilherme C. Oliveira, Stephanie Fong, Zimu Wang, Clarice Lee, Xiangyu Zhao et al. from Monash University and Orygen and The University of Melbourne and Swinburne University of Technology and University of Liverpool and Khalifa University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's got a real mouthful of a title — "AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement." Jane, I need you to unpack that for me, because I'm already lost at "SIPS."

Jane: Happy to, Tom. So SIPS stands for the Structured Interview for Prodromal Syndromes — it's a way clinicians assess whether someone might be at risk for psychosis. And this paper built a huge synthetic dataset of those interviews, ten thousand of them, so researchers can train AI systems to understand how these assessments actually work.

Tom: Ten thousand fake clinical interviews? That's wild. But why synthetic instead of just using real ones?

Jane: Because real psychiatric interviews are incredibly sensitive. You can't just share those recordings or transcripts — privacy laws, consent issues, governance. So the authors created realistic fake ones that preserve the structure of the real thing without exposing any actual patient.

Lu: And that's the clever part, Tom. They didn't just ask a language model to make up conversations. They built a "plan-then-realize" pipeline. First, a deterministic planner decides everything about the case — what symptoms the patient has, what the diagnosis should be, what questions get asked. Then, and only then, does an LLM write the patient's actual words.

Tom: So the AI is basically an actor reading a script, not the director?

Lu: Exactly. The script is fixed before any dialogue exists. That means the labels — the diagnoses, the symptom endorsements — are all determined by logic, not by whatever the language model happens to say. It's auditable and reproducible.

Jane: And that matters because the final diagnosis isn't just one label. It depends on a whole chain of intermediate steps — which symptoms the patient endorses, what follow-up details they give, whether certain criteria are met, whether something else is ruled out. The dataset captures all of that structure.

Tom: So it's not just "here's a transcript, here's the answer." It's "here's every step along the way and the exact lines of dialogue that support each step." That sounds like a much harder test for an AI.

Jane: That's precisely the point. And it's why this paper could be a big deal — we'll get into how the models actually performed on that test next.

Summary: Tom: So we've established what AnchorSIPS is — ten thousand synthetic psychosis-risk interviews with a fixed structure. But Jane, what did the paper actually find when they tested AI systems on this dataset?

Jane: The headline finding is a gap, Tom. A big one. They tested seven different language models, and the models were pretty good at the coarse stuff — recognizing which symptoms were mentioned, predicting whether the patient had frank psychosis. But when it came to the detailed work, they fell apart.

Tom: How bad are we talking?

Jane: Look at the numbers. Query endorsement — just detecting which of the twenty-four symptom questions were answered yes — that was near perfect, up to one point zero F1. But evidence linking, which means citing the exact transcript turns that support a decision, topped out at zero point two nine. That's low.

Lu: And that's the critical insight, Tom. A model can say "this patient has attenuated psychosis syndrome" and be right. But if you ask it to show its work — to point to the specific lines where the patient described the symptom, the frequency, the distress — it can't do it reliably. The final answer overstates the model's actual competence.

Tom: So it's like a student who guesses the right answer on a multiple-choice test but can't explain why. They look smart until you ask them to show their reasoning.

Jane: Exactly. And the paper calls this out directly. The final-label performance — the diagnosis prediction — makes the models look better than they actually are. The real bottleneck is extracting the follow-up details and grounding them in the transcript.

Meng: From where I sit, that's a practical problem too. If you're building a system to help clinicians review interviews, you don't just want a diagnosis. You want the system to pull out the relevant evidence so the clinician can verify it. If the AI can't cite its sources, it's not actually saving anyone time.

Tom: And the models were all over the place, right? No single model dominated?

Jane: Right. The top models — GPT-five point five, DeepSeek V4 Flash, Claude Opus four point seven — were clustered together on the composite score. But one smaller model, Ministral 8B, actually got the highest score on the final APS diagnosis. Yet its evidence linking was the worst of the group. So you can have a model that nails the diagnosis but can't tell you why.

Tom: That's the kind of result that keeps engineers up at night.

Meng: It does. Because it means you can't just trust the final output. You need to build systems that force the model to ground every claim in the transcript. And this dataset gives you a way to measure whether you're actually succeeding at that.

Improvements: Tom: So AnchorSIPS exposes a real problem — models can guess diagnoses but can't cite evidence. What does the paper suggest we actually do about it? What are the improvements?

Jane: Well, the paper is honest that this is a benchmark, not a finished solution. But it points toward some clear directions. One is that the dataset itself is designed to stress-test models on partial disclosure — patients who are guarded, vague, inconsistent, or delay revealing important information.

Lu: And that's a genuinely hard problem, Tom. In real psychosis-risk interviews, young people often downplay their symptoms. They're scared of stigma. They don't want to admit they hear things or believe unusual things. So the information comes out gradually, in fragments, sometimes contradictorily.

Tom: So the dataset deliberately includes cases where the patient is being difficult?

Jane: Exactly. Thirty percent of the interviews are marked "guarded," thirty-five percent are "vague," twenty percent have strong delayed revelation. The idea is to see whether a model can stay uncertain when the evidence isn't there, rather than committing to an early answer and never recovering.

Meng: That's the part I care about. A model that overconfidently says "no psychosis risk" because the patient was vague in the first few turns is dangerous. The paper's finding that models are bad at evidence linking suggests they're also bad at knowing when they don't know.

Lu: And that's why the paper proposes measuring things like "grounded correctness" — a prediction only counts if the model cites the right evidence. And "unsupported claim rate" — how often the model makes a claim with no supporting transcript. These are the metrics that matter for real-world safety.

Tom: So the improvement isn't just "make the model better at diagnosing." It's "make the model better at knowing what it actually knows."

Jane: Right. And the paper also flags that their baseline results are just that — baselines. They used a single-pass prompting approach. They explicitly say that chain-of-thought reasoning, few-shot prompting, retrieval-augmented citation, and multi-pass extract-then-cite pipelines would likely do better. The dataset is a measuring stick, not a ceiling.

Lu: And that's the beauty of it. Because the labels are fixed deterministically before any text is generated, you can trust the benchmark. When a model improves on AnchorSIPS, you know it's genuinely getting better at evidence-supported reasoning, not just pattern-matching to plausible-sounding dialogue.

Tom: So the next step is for researchers to build better systems and measure them against this benchmark. But before we get there — what about the human side? The paper mentions an expert audit. How did that go?

Jane: That's actually our next segment. The authors brought in trained psychosis-risk interviewers to check whether the synthetic evidence actually supports the labels. And the results are... interesting.

First Page: Tom: So we're looking at the first page of AnchorSIPS now, and Jane, you mentioned the expert audit. What did the researchers actually do?

Jane: They recruited three reviewers with real experience conducting structured psychosis-risk interviews — including PSYCHS interviews with young people at ultra-high risk. These aren't random annotators. These are people who've done this work in the field. And they were asked to check whether the transcript evidence in the synthetic interviews actually supports the labels.

Tom: So it's a quality check. Are the fake interviews realistic enough that real experts agree with the labels?

Jane: Exactly. They looked at fifteen interview-by-symptom-class decision records, each rated by all three reviewers. That's forty-five item-level reviews and two hundred seventy criterion-level ratings. And the overall finding was that reviewers usually agreed the cited transcript snippets supported the intended labels.

Tom: Usually. So there were problem areas?

Jane: Two, specifically. The clearest evidence was for whether symptoms worsened in the past year and whether they caused distress. But ratings were less decisive for "functional influence" — whether the symptom affects daily life — and for "alternative explanation" — whether stress, sleep, or substances could explain it. Those are harder to infer from short transcript excerpts.

Lu: And that makes clinical sense, Tom. In real interviews, assessing functional impact takes careful probing. You need to ask about school, work, relationships, daily routines. And ruling out alternative explanations requires a differential conversation. A brief excerpt often just doesn't contain enough to make that call confidently.

Tom: So the audit found the dataset is mostly solid, but those two areas are where the synthetic interviews might be thinner?

Jane: Right. And the paper is upfront about it. They call functioning and rule-out judgments "important targets for larger future audits." The pairwise agreement between reviewers was fifty-seven percent exact, with a mean absolute difference of zero point five one on the No-Partly-Yes scale. So there's real ambiguity there.

Meng: For me, that's actually reassuring. If the audit had come back with perfect agreement, I'd be suspicious. Real clinical judgment has uncertainty. The fact that experts disagree on the harder criteria means the synthetic data is capturing some of that genuine difficulty.

Tom: That's a good point. A perfect dataset would be a red flag. This one has realistic rough edges.

Jane: And it also tells us where the benchmark is hardest — which is exactly what you want from a stress test. The paper isn't trying to hide its weaknesses. It's documenting them so researchers know where to focus.

Lu: And that connects back to the bigger picture. The first page frames this whole thing around evidence-supported measurement — the idea that an AI system should be judged on whether it can recover the intermediate clinical state, not just the final label. That's the philosophical core of the paper.

Conclusion: Tom: Alright, let's wrap this up. We've been talking about "AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement." Jane, give us the final summary.

Jane: The big picture is this: the paper built ten thousand synthetic psychosis-risk interviews with a fixed, auditable structure. The labels are determined before any AI writes a word. And the key finding is that current language models can guess diagnoses but can't reliably cite the transcript evidence that supports them.

Tom: And that's the gap that matters for real-world use.

Jane: Exactly. The dataset gives researchers a way to measure whether a system is actually doing evidence-supported reasoning or just pattern-matching. And the expert audit shows the synthetic data is mostly realistic, with known weak spots in functional impact and alternative explanations.

Lu: I'd add that the plan-then-realize approach is a template for other sensitive domains. If you can't share real clinical data, you can build synthetic data where the ground truth is fixed by logic, not by the generator. That's a reusable idea.

Meng: And from a practical standpoint, this benchmark will push engineers to build systems that force models to cite their sources. That's the only way these tools become trustworthy enough for clinicians to actually use.

Tom: So what's the takeaway for our listeners? This paper is a measuring stick. It shows us where AI is failing at psychosis-risk assessment — not at the final answer, but at the evidence. And it gives us a tool to track improvement.

Jane: And that's a genuinely important contribution. Because you can't fix what you can't measure. AnchorSIPS gives the field a way to measure the thing that matters most: whether an AI can back up its claims with evidence.

Tom: Well said. That's a wrap on AnchorSIPS. Thanks for joining us, everyone — we'll be back next time with another paper from the arXiv. Until then, keep questioning the answers.

More episodes

← Home