2608.07202-Authoring and Management of Transparent Research Integrity Assessments of Randomised Clinical Trial Publications Using LLM-assisted Tools and Provenance Knowledge Graphs

page_by_page

Video file (mp4)

In short

The episode discusses a paper on the RIPE Observatory, a system using LLM-assisted tools and provenance knowledge graphs to assess the integrity of randomized clinical trial publications. Hosts highlight the pipeline's components, the importance of human oversight, and the pilot's low cost and feasibility.

Key concepts

Provenance knowledge graph
A machine-readable record of how each integrity assessment was conducted, including evidence, automated suggestions, and human decisions. It allows tracing why a trial was flagged, ensuring transparency and enabling comparison between human and AI judgments.
INSPECT-SR checklist
A Cochrane-endorsed framework for assessing research integrity in systematic reviews. It guides reviewers through questions about trial registration, data trustworthiness, and other integrity concerns, and is used by the INSPECT-AI tool to structure assessments.
Human-in-the-loop
A design where an AI tool proposes assessments but a human must confirm or override each decision. This ensures final judgments are made by people, addressing disagreements between AI and humans and maintaining accountability in the integrity assessment process.

This episode discusses

Transcript

Introduction to the show: ident: Paper Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Authoring and Management of Transparent Research Integrity Assessments of Randomised Clinical Trial Publications Using LLM-assisted Tools and Provenance Knowledge Graphs".

Jane: The paper was written by Milan Markovic, Goutham Indukuri, Somayajulu Sripada, Colby J. Vorland, Jack Wilkinson et al. from University of Aberdeen and Indiana University and University of Manchester and University of Auckland.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Paper Summary: Tom: We've got a fascinating paper to open with today, all about keeping unreliable clinical trial results from quietly leaking into medical practice. The author list runs across computing science, epidemiology, and biostatistics, which tells you straight away this is a practical problem. And the whole thing hangs together as one pipeline.

Jane: It really is a systems story. There's an eye-assisted assessment tool, a formal ontology for recording how each assessment happened, and a public knowledge graph holding the results. Those three pieces make up the RIPE Observatory.

Lu: The problem they're tackling is serious. Systematic reviews of randomized trials feed directly into clinical guidelines, so a falsified trial doesn't just waste money, it can change what doctors recommend. That's the stakes for patients.

Tom: The paper cites earlier work showing 27 trials with integrity concerns had found their way into 88 systematic reviews or clinical guidelines, changing findings in over half of them. In 87 percent of those cases, the change was substantial enough to shift the direction of effect.

Meng: That's a scary number to open with.

Tom: It is. And it's exactly the kind of downstream harm that makes integrity assessment more than an academic exercise.

Meng: Which is why they built INSPECT-AI. It walks a human reviewer through the INSPECT-SR checklist, the Cochrane-endorsed integrity framework, while an LLM pulls evidence from the PDF and from external registries and databases. Every suggestion has to be confirmed or overridden by a person.

Jane: And the clever bit is the provenance side. Every piece of evidence, every automated suggestion, every override gets recorded. You can trace why a trial was flagged, step by step.

Lalam: That's the piece that's been missing everywhere. There are integrity checklists, and there are eye assistants, but nobody has been publishing the full audit trail of the assessment in a machine-readable, FAIR-aligned form.

Tom: And the released dataset is substantial — 140 assessments of 95 trial publications sitting in an open knowledge graph with a SPARQL endpoint.

Lu: But the results also show why the human in the loop matters. Automated suggestions and human reviewers disagreed on about 14 percent of individual check outcomes. And on trial registration questions, human reviewers disagreed with each other more than half the time.

Jane: That sounds like a weakness at first, but it's actually the paper's strongest argument. You need the provenance precisely because the humans disagree, and because the eye disagrees with the humans.

Meng: And the cost numbers make it feasible. The whole pilot ran on less than eighty dollars, with the language model portion just over a dollar and each paper coming in below ten cents.

Lalam: That combination — rigorous provenance plus trivial cost — makes me think this is a template for how any evidence-based field could audit its literature.

Tom: The abstract already tells you how carefully they've positioned this work. Let's look at that opening page.

Page 1 of the Paper: Tom: So we've sketched the whole arc, and now we can read the opening page the way the authors intended. The abstract's first sentence describes INSPECT-eye as an interactive tool, and that word carries real weight. The system proposes, but a person confirms or overrides every decision before it counts.

Jane: It also says the INSPECT-SR framework is community approved, which is a crucial detail. That checklist went through a formal consensus process with international integrity researchers before Cochrane endorsed it. You can't claim that for most instruments in this space.

Lu: The abstract ties the tool to the honesty principle in research publishing. There are several principles of research integrity, but this work deliberately focuses on one question: are the data and findings trustworthy enough for evidence synthesis?

Meng: And then the scale: 140 expert assessments of 95 publications. Those assessments were done by real reviewers using the tool, and each one is captured in the ontology's structure. The word "initial" matters too — this is a starting point, not a finished corpus.

Lalam: What I find telling is the permanent identifier at the bottom — w3id.org/ripe. That's infrastructure thinking. The authors are saying this resource is meant to persist and be referenced, not to live and die with a grant cycle.

Lu: And that's rare for a research prototype.

Lalam: Exactly. The choice of a stable web identifier signals intent from day one.

Jane: The author list tells the same story. Computing scientists from Aberdeen sit alongside epidemiologists and biostatisticians from Indiana, Manchester, and Auckland. That mix explains why the human workflow and the technical machinery stay in view together.

Tom: The keywords are equally revealing — research integrity, large language model, ontology, provenance, knowledge graph. That's a signal to two communities at once. The integrity world gets a tool, and the semantic web world gets a model it can build on.

Meng: One phrase in the abstract deserves attention too. It says the assessments were "described using RIPE-O," which implies the ontology came after the assessments, shaped by what actually happened. That's the opposite of a top-down modeling exercise.

Jane: From that framing, the paper moves to laying out the RIPE Observatory itself — three named components, each with a permanent identifier. That's where the architecture gets concrete.

Page 2 of the Paper (Discussing Page 3): Tom: The next page introduces the RIPE Observatory as a collection of semantic resources, and each component gets its own permanent identifier — the tool, the ontology, and the knowledge graph. That's the architecture made explicit.

Jane: There's a refreshingly honest division of labor here. They name the ontology and the knowledge graph as the core semantic contributions, while the tool is included because it shaped them and completes the pipeline. But they also say the ontology is meant to work with other integrity tools, not just theirs.

Meng: That openness is the difference between a project deliverable and actual infrastructure. There are other checklists in this space — TRACT, RIA, REAPPRAISED, and the tool used in Cochrane pregnancy research. If each one produces its own incompatible data format, you haven't solved the problem.

Lu: The related work section draws the line against the existing scholarly knowledge graphs. OpenAIRE, Scholia, Semantic Scholar, the Open Research Knowledge Graph, SemOpenAlex — they all integrate enormous amounts of bibliographic data.

Jane: But none of them carry data about trustworthiness. No retraction records, no expressions of concern, no integrity flags. That gap is what this paper claims.

Tom: And the authors point out that discovering problematic trials today is still a manual slog. Integrity sleuths alert journals, but publishers can be slow to act, and a high proportion of untrustworthy publications remain unflagged years later.

Jane: Which seems backwards for a world that tracks citations in real time.

Tom: Completely. The scholarly web knows more about who cites whom than about whether a study deserves to be cited at all.

Meng: What they're adding is the layer that records the checks being performed, in a form computers can reason over. That's a different kind of contribution from another checklist.

Lalam: And they're aware of adjacent attempts, including an early experiment semi-automating the TRACT checklist with GPT-4o. So they're positioning themselves against known work rather than ignoring it.

Lu: The paragraph about the checklists also explains why INSPECT-SR matters — it's the one endorsed by Cochrane, which gives the assessments institutional weight from the start.

Jane: Having situated themselves against the field, the next section explains how they actually built the system. And the methodology is as interesting as the results.

Page 3 of the Paper (Discussing Page 5): Tom: The methodology section reads like a case study in interdisciplinary tool building, and the biggest outcome is the decision to keep a human in the loop from day one. That decision emerged from the development process rather than being a default starting point.

Jane: They're open about how it happened. The non-technical team members understood the strengths and limits of eye only after seeing real, semi-functional prototypes. Paper mock-ups and technical explanations didn't land the same way.

Lu: And this is where having biomedical researchers in the room mattered. They supplied fifty real problematic publications with detailed guidance on how to assess them. That gave the computing team the domain grounding no textbook could provide.

Meng: The order of operations is also surprising. They built the tool first, and the ontology afterwards, derived from the logs the running application actually produced. Most ontology work starts with a model and then hunts for data to fit it.

Jane: That's the reverse of the usual pattern.

Meng: It is. And it's probably why the model feels grounded in real assessment practice rather than abstract theorizing.

Lalam: The provenance graphs were generated from those JSON logs using declarative YARRRML mappings. So the audit trail is a byproduct of the system in use, not a reconstruction attempted after the fact.

Jane: For the ontology itself they followed a formal engineering framework called LOT, with competency questions framed around the seven Ws of provenance — who, what, where, which, why, when, how.

Tom: And the paper makes a smart reuse decision by building on TIDO, an ontology from threat intelligence. That brings a forensic framing with it — an integrity assessment is treated like an investigation, with evidence, hypotheses, and evaluation activities.

Meng: TIDO already aligns with the W3C provenance standard, PROV-O, so all of this inherits standard semantics. Interoperability comes almost for free.

Lu: They also reuse the SPAR ontologies for bibliographic details rather than reinventing how to describe authors and publications. That's the kind of hygiene that makes an ontology adoptable beyond the team that built it.

Jane: All of that design work then got exercised in a real pilot with real reviewers. And the numbers from that pilot are surprisingly good.

Page 4 of the Paper (Discussing Page 7): Tom: The pilot deployment is a genuine field test. Thirteen volunteers from three universities spent October and November last year using the tool on real publications.

Jane: And these weren't random volunteers. 77 percent had experience conducting systematic reviews, and nearly 39 percent had ten or more years of that experience. When people like that say a tool is usable, it carries weight.

Lu: They assessed 69 publications and produced 104 assessment traces, with some papers assessed by multiple reviewers. That overlap is exactly what you need if you want to study disagreement later.

Meng: The usability results were strong. Nobody found the tool difficult to use, 69 percent called it user-friendly, and 69 percent rated the response time as quick or very quick.

Tom: And no negative ratings on any of those scales.

Meng: Right — zero difficulty complaints. That's a clean result for a prototype.

Lalam: One detail I like is that participants were given publications with known trustworthiness issues, but they were also encouraged to choose their own. The traces reflect genuine reviewer behavior, not a scripted exercise.

Tom: The paper is also honest that the tool didn't remove the complexity of integrity assessment. It made the process faster and better supported, which is a more realistic promise than total automation.

Jane: There's a subtle point about who these reviewers were. Only 61 point 5 percent had done integrity assessments before, although 77 percent had done risk of bias assessments. Related skills, but not the same thing.

Lu: That actually strengthens the pilot. If health researchers with mixed backgrounds can pick up the tool and produce assessments, that's decent evidence for wider adoption.

Meng: And the data from this pilot became the raw material for the knowledge graph. The volunteer traces, combined with assessments by the expert core team, grew into the 140 published assessments.

Jane: With that data in hand, the paper moves to the ontology. And the model turns out to be a lot more generic than you'd expect from a tool built around one checklist.

Page 5 of the Paper (Discussing Page 9): Tom: The ontology section opens with a design claim that's easy to miss but really important: RIPE-O doesn't prescribe which questions are asked. It models the structure of an assessment, not the content of any particular checklist.

Jane: So a question like "are there concerns about the timing of study registration" becomes an integrity assessment question. Evidence is gathered, an evaluation activity weighs it, and the answer is recorded as a hypothesis with an outcome.

Meng: The consequence is that the same structure could document integrity concerns beyond clinical trials. Raising a question, weighing evidence, recording a verdict — that pattern is generic.

Lalam: They're drawing on forensic modeling through TIDO, so every assessment is a case. Evidence supports or challenges hypotheses, evaluation is an activity, and all of it connects back to PROV-O semantics.

Lu: The ontology distinguishes different kinds of signals — retraction notices, corrections, expressions of concern, peer comments. Each is its own class, so the graph can record precisely what kind of post-publication attention a paper received.

Jane: And the agents are separated too. Human reviewers and automated agents are both recorded as the sources of the hypotheses they generate.

Tom: Without that separation, you couldn't compare machine decisions with human ones. Both sets of hypotheses live in the same structure, which makes the comparison direct.

Jane: Exactly. That parallel design is what powers the disagreement analysis later.

Tom: The model also gives study design evidence and registry evidence their own properties — registration identifiers, dates, recruitment timelines. That tells you the ontology was shaped by what the tool actually had to extract.

Meng: They validated the ontology with a standard pitfall scanner and turned their competency questions into SPARQL queries. Verification like that gives confidence the model behaves as claimed.

Lu: One more thing I'd flag: the authors argue several of their checks are universally applicable to all academic literature, not just randomized trials. Retractions and expressions of concern exist in every field.

Jane: Once you have a structure like that, the natural step is to fill it with real assessments and publish them as a knowledge graph. And that's exactly what RIPE-KG does.

Page 6 of the Paper (Discussing Page 11): Tom: The knowledge graph is the public face of the entire project. At the time of writing it contains 140 assessments of 95 distinct publications, and more than 1,200 individual integrity hypotheses.

Jane: Reviewer anonymity is handled with pseudonymised IDs and roles. The graph can distinguish different reviewers without exposing them, which is sensible given how contentious integrity assessments can be.

Meng: The author disambiguation problem is substantial, and the paper is transparent about the numbers. GROBID extracted over 1,200 author mentions from the PDFs, which collapsed into 875 distinct author instances.

Lu: Of those, 766 got linked to SemOpenAlex author identifiers using owl:sameAs links. That's what enables federated queries across their graph and the wider scholarly graph.

Tom: So roughly 12 percent stayed unlinked.

Lu: Yes, 109 authors, and they don't hide it. Missing OpenAlex metadata and mismatched author lists are the reasons given.

Tom: They also note that the OpenAlex API proved more reliable than the SemOpenAlex SPARQL endpoint, a practical detail that will save people headaches.

Jane: The transformation pipeline is declarative. JSON logs from the tool map into RDF through YARRRML rules, so the graph is regenerated from the process records rather than hand-assembled.

Lalam: For me, the web interface matters as much as the SPARQL endpoint. A knowledge graph only experts can query is a walled garden. The browser view lets editors and reviewers see the assessments and the evidence behind them.

Meng: And because the graph is aligned with SemOpenAlex from the start, you can combine integrity outcomes with fields of study, citation data, and author networks. That's where the analytical power comes from.

Jane: Which brings us to the analysis section. The numbers there are genuinely provocative, because they quantify just how much disagreement exists at every level of the process.

Page 7 of the Paper (Discussing Page 13): Tom: The analysis section is the payoff. Across 514 paired automated and human outcomes, they agreed 86 point 4 percent of the time and disagreed 13 point 6 percent of the time.

Jane: The distribution of those disagreements is the real story. Retraction checks were close, at 8 point 5 percent disagreement. Research-team-related checks ran around 15 percent. And study registration hit 26 point 2 percent disagreement — the weakest spot by far.

Lu: And that makes clinical sense. Judging whether registration was properly timed requires knowing how trials actually operate — whether a delay is acceptable, whether the reported timeline is plausible. That's expertise an LLM can't reliably pull from a PDF.

Meng: The inter-reviewer numbers are even more striking. Across the 22 publications with multiple assessments, reviewers agreed with each other 94 point 4 percent of the time on research-team concerns, but only 44 point 4 percent of the time on registration.

Jane: More than half of the time they disagreed.

Meng: More than half. On a question that can determine whether a trial gets included in an evidence synthesis.

Tom: It's also interesting that the paper suspects some reviewers accepted the automated retraction suggestions without doing the extra check of publisher websites, which the tool can't always reach. The human side can be too trusting too.

Jane: The overall verdicts tell a similar story. Of the 22 works reviewed more than once, over a quarter drew different overall conclusions from different reviewers.

Lalam: My favorite detail is the single publication that received eight assessments spanning all three verdict categories — no concerns, some concerns, serious concerns. That example is the entire argument for provenance in miniature.

Meng: Without an audit trail, that paper looks like noise. With the trail, you can study why experienced people read the same evidence so differently.

Lu: It also pushes back on any idea that the eye should replace the reviewers. The data says the opposite — you need the human, and you need the record of what the human did.

Jane: Then the paper turns to what all this costs to run. And that's the question everyone asks about eye systems in practice.

Page 8 of the Paper (Discussing Page 15): Tom: The cost section is wonderfully concrete. For the entire two-month pilot period, the total operational cost was about .53, and that includes everything.

Jane: The largest line item was infrastructure — a single virtual server running all the services in Docker containers, at . The LLM calls, using Gemini 2 point 0 Flash through Google eye Studio, came to just .34.

Meng: Doing the arithmetic, that's below ten cents per paper analyzed. When the authors say this makes integrity assessment scalable, they show the invoice to prove it.

Lalam: Then the limitations section brings the tone back to earth. Only four of the INSPECT-SR checks are automated so far, and the tool inherits whatever gaps exist in OpenAlex, SemOpenAlex, Retraction Watch, and PubPeer.

Lu: Those external sources are real dependencies. If Retraction Watch is incomplete, or publication data in OpenAlex is wrong, the assessments reflect it. The paper acknowledges this directly.

Jane: There's also the firewall problem. The tool can't always reach publisher websites to verify retraction status, which is why the guidance tells reviewers to do that extra check manually.

Tom: The honest framing is that this is decision support with a recorded audit trail, not an automated judge of scientific trustworthiness. Given the disagreement data we just discussed, that's the right line to draw.

Meng: And the federated query examples on this page show the payoff of linking everything together. They pull endocrinology publications and their integrity outcomes straight out of the combined graph.

Lu: So a funder or a journal could ask, across an entire field, which publications carry integrity concerns. That query is now routine.

Jane: The authors close by drawing conclusions that pull the threads together rather than overclaiming. And that's where we should land too.

Conclusion: Tom: We've walked through the whole pipeline, so let's pull the threads together. What this paper delivers is an end-to-end system for research integrity assessment: an eye-assisted tool, an ontology for provenance, and a public knowledge graph holding 140 assessments of 95 trial publications.

Jane: The core lesson is that integrity assessment is variable by nature, and the right response to that variability is to record it openly rather than smooth it over. Disagreement becomes data.

Lu: For clinical evidence, that means a realistic path toward scaling integrity checks that currently take experts hours per paper, at a cost that's remarkably low.

Meng: For the knowledge graph community, it's a demonstration that trustworthiness can be modeled using established standards — PROV-O, TIDO, the SPAR ontologies — rather than building in isolation.

Lalam: And it's a template for any field that depends on published evidence, from education research to environmental policy. The plumbing is generic even though the checks are clinical.

Tom: The human-in-the-loop result is the thing I'll carry with me. The system found real value, but the disagreements with humans, and among humans, make the case for keeping people central to the process.

Jane: The permanent identifiers, the documented ontology, the queryable graph — this is infrastructure built to be extended. The authors are explicit in inviting other integrity tools to plug into the same framework.

More episodes

← Home