MatchMiner-AI: Open-source, Privacy-preserving Cancer Clinical Trial Matching using Artificial Intelligence
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "MatchMiner-AI: Open-source, Privacy-preserving Cancer Clinical Trial Matching using Artificial Intelligence".
Jane: The paper was written by Jennifer Altreuter, Pavel Trukhanov, Morgan A. Paul, Michael J. Hassett, Irbaz B. Riaz et al. from Dana-Farber Cancer Institute and Mayo Clinic.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's got a pretty hefty title: "MatchMiner-AI: Open-source, Privacy-preserving Cancer Clinical Trial Matching using Artificial Intelligence." Jane, I gotta say, just reading that title gets me excited, because it's tackling two huge problems at once.
Jane: It really does, Tom. And let's break down what that title actually means, because there's a lot packed in there. So "clinical trial matching" is the process of figuring out which experimental cancer treatments a patient might be eligible for. And the paper's authors, this big team from Dana-Farber and Mayo Clinic, they're trying to automate that with AI.
Tom: Right, and the "open-source" and "privacy-preserving" parts are the real kickers. You see, most of these AI tools are built by companies that keep everything locked up, and they often need to send your private medical data to their servers. This team wanted to build something that any hospital could use without giving up patient privacy.
Jane: Exactly. And that's a huge deal, Tom, because right now, fewer than one in ten adults with cancer actually enroll in a clinical trial. One big reason is that oncologists just don't have the time to manually search through thousands of complex eligibility criteria for each patient. So this tool is trying to make that process faster and more accessible.
Tom: And the "privacy-preserving" bit is clever, because they trained the whole system on synthetic data. That means they generated fake patient records, not real ones, so the AI never saw a single real person's medical history during training. That way, they can share the models freely without worrying about leaking patient information.
Jane: That's the part that really blew my mind when I read it. They basically created a whole fake hospital full of fake patients to teach the AI how to do its job. And then they tested it on real patients at Dana-Farber, and it worked remarkably well. We're going to get into the actual numbers later, but the approach itself is so clever.
Tom: It's a classic "teach the model in a sandbox, then let it loose in the real world" strategy. And the fact that they're giving away all the code and the models for free means that any hospital in the world could potentially set this up. That's democratizing access to cutting-edge cancer research.
Jane: And that's the big picture here, Tom. This isn't just a tech demo. This is a tool that could actually help get more patients into trials, which means faster drug development and more treatment options for people who are running out of standard options. We'll talk about how it actually works in a second, but first, let's just appreciate the ambition of this title.
Tom: Absolutely. And speaking of how it works, the paper walks us through a pretty sophisticated pipeline. Stick around, because we're about to get into the details of how they built this thing.
Abstract: Tom: So Jane, we've set the stage with the title, but the abstract of "MatchMiner-AI" really lays out the whole story. Let's talk about what they actually did and what they found.
Jane: Yeah, so the abstract gives us the core numbers. They built this pipeline that uses a large language model to do two main things. First, it reads a patient's messy electronic health record, all those notes and reports, and summarizes it into a clean, structured format. Second, it reads the eligibility criteria for thousands of clinical trials and extracts the "target populations," which they call "trial spaces."
Tom: And that's the key innovation, right? Instead of trying to check every single eligibility criterion, which is super complex and error-prone, they focus on the core things that matter most: cancer type, age, sex, prior treatments, and key biomarkers. It's like finding the right neighborhood before you start looking at specific houses.
Jane: Exactly. And then they have this custom AI model, called TrialSpace, that acts like a super-fast search engine. It embeds both the patient summaries and the trial spaces into a mathematical space where similar things are close together. So when you have a patient, you can instantly find the trials that are most relevant.
Tom: And the results are pretty striking. In their retrospective evaluation, they compared their model to a baseline text-embedding model. The baseline scored a mean average precision at twenty of about zero point four four for trial-enrolled patients. Their full pipeline, with the TrialChecker re-ranker, jumped that up to zero point nine five. That's a massive improvement.
Jane: For our listeners who aren't statisticians, that means the old model was often pulling up irrelevant trials, but the new one is almost always pulling up trials that actually make sense for the patient. And they saw similar improvements for patients who received standard of care, not just those who enrolled in trials.
Tom: And here's a really important detail from the abstract. They compared their AI tool to an older, rules-based tool that relied on tumor genomic data. That old tool could only work for nineteen out of fifty patients, because it needed that specific genomic test. MatchMiner-AI worked for all fifty patients, because it can use the entire medical record, not just one test result.
Jane: And even when they looked at just those nineteen patients where both tools could work, MatchMiner-AI was better. An independent AI judge rated eighty percent of its trial suggestions as reasonable, compared to only fifty-three percent for the old rules-based approach. So it's not just covering more patients, it's also making better suggestions.
Tom: That's the kind of evidence that makes you sit up and pay attention. It's one thing to build a cool demo, but they've shown that this thing actually outperforms existing methods on real patient data. And they did it all with models that are small enough to run on a regular computer, which is huge for practical adoption.
Jane: And that's the beauty of the abstract. It's a promise that this open-source, privacy-preserving approach isn't just theoretically nice, it's actually better. We're going to dig into the methodology next, because the way they trained this thing is fascinating.
Improvements: Tom: Alright Jane, we've covered the title and the abstract. Now let's talk about what "MatchMiner-AI" actually improves upon. Because this paper isn't just saying "look, we built a thing." It's saying "here's how we made it better than everything else."
Jane: Right, and the biggest improvement is in how they handle the messy reality of patient data. A lot of previous AI trial-matching systems assumed you already had a clean, short summary of the patient's history. But in the real world, you have years of progress notes, imaging reports, and pathology reports. So they built a system that can ingest all of that raw text and create a summary on the fly.
Tom: And that's where the synthetic data training comes in. They didn't have a real patient dataset they could share, so they used a teacher LLM to invent thousands of fake patient histories and full-length clinical documents. That's over a million synthetic documents, by the way. And they trained their smaller, faster models to mimic the teacher's judgment on this fake data.
Jane: It's like an apprenticeship. The big, smart teacher model shows the small student models millions of examples of how to match patients to trials. And by the end, the student models are almost as good, but they're fast enough and small enough to run in a real hospital setting.
Tom: And they also improved on the "boilerplate" problem. Most trials have common exclusion criteria, like "no uncontrolled brain metastases" or "no recent heart failure." Instead of making the AI check every single one of these for every trial, they built a separate, dedicated checker. It's a simple binary classifier that just says "yes, this patient is excluded" or "no, they're not."
Jane: That's a smart division of labor. The main matching model focuses on the core clinical question, like "does this patient have the right cancer type and the right prior treatment?" And then the boilerplate checker handles all the generic safety stuff. It's more efficient and easier to debug.
Tom: And they didn't just build it in a vacuum. They actually piloted it with eight practicing oncologists at community-based sites. And those doctors gave feedback, like "hey, you're showing me pediatric trials for my adult patients" or "you're missing brand-name drugs in the summaries." And the team went back and fixed those issues.
Jane: That iterative process is so important. You can have the best algorithm in the world, but if it doesn't fit into the doctor's workflow, it's useless. And they listened. They even integrated a link directly into the electronic health record system, so the oncologist can just click a button and see the trial matches for the patient they're looking at.
Tom: And the final improvement is in the evaluation itself. They didn't just trust their own models. They used a frontier commercial LLM as a judge to check the quality of their patient summaries and trial matches. And they had human oncologists manually review a sample of the matches. So they've got multiple layers of validation.
Jane: And that's what gives me confidence in the results. It's not just one metric saying "this works." It's the AI agreeing with the doctors, and the doctors agreeing with each other, and the whole thing being consistent across different patient demographics. This is a thoroughly tested system.
Tom: So we've got a system that's better, faster, and more private than what came before. The implications here are massive, and we'll wrap up with those next.
Conclusion: Tom: Alright, we've spent this whole episode on "MatchMiner-AI: Open-source, Privacy-preserving Cancer Clinical Trial Matching using Artificial Intelligence," and I think it's time to step back and think about what this really means for the world.
Jane: Yeah, Tom, and I think the biggest implication is that it could finally start to move the needle on that dismal ten percent clinical trial enrollment rate. If an oncologist can get a list of reasonable trial options for a patient in seconds, instead of spending an hour digging through databases, they're going to have those conversations a lot more often.
Tom: And it's not just about the patient's benefit. It's about the whole research enterprise. Trials that fail to accrue enough patients are a massive waste of time and money. If this tool helps fill trials faster, we could see new cancer drugs getting approved sooner.
Jane: And the fact that it's open-source and privacy-preserving is the real game-changer. It means a small community hospital in rural America, or even a hospital in a developing country, could potentially deploy this without having to sign a huge contract with a tech company or worry about sending patient data across borders.
Tom: That's the democratization angle, and it's huge. The authors have put all the code, the model weights, and even the synthetic training data online for free. So any research group with some technical skill can pick it up and run with it.
Jane: And there's a broader lesson here about using synthetic data in medicine. This paper shows that you can train a highly effective clinical AI without ever touching a single real patient record. That's a blueprint that could be applied to lots of other problems, like predicting disease progression or identifying adverse drug reactions.
Tom: Now, we should be clear that this isn't a replacement for a doctor. It's a tool to help them. The paper itself says it's not a medical device or formal decision support. It's a way to surface options that the oncologist then has to evaluate with their clinical judgment.
Jane: And that's the right framing. It's an assistant, not an autopilot. But it's an assistant that could save doctors hours every week and make sure no patient is overlooked.
Tom: So as we say goodbye to "MatchMiner-AI," I think the takeaway is that this is a genuinely important step forward. It's a practical, tested, and shareable solution to a real problem that affects millions of people. And it's a great example of what's possible when researchers prioritize openness and privacy.
Jane: Couldn't agree more, Tom. It's a paper that gives you hope that AI in healthcare can be both powerful and responsible. We'll be keeping an eye on how this gets adopted in the real world.
Tom: And that's a wrap on this one. Thanks for listening, everyone. We'll see you next time with another fascinating paper.
Jane: Take care, everyone.
Jennifer Altreuter, Pavel Trukhanov, Morgan A. Paul, Michael J. Hassett, Irbaz B. Riaz, Muhammad Umar Afzal, Arshad A. Mohammed, Ayub Umair, Huan He, Chueh Husan Hsu, Sarah Sammons, James Lindsay, Emily Mallaber, Harry R. Klein, Gufran Gungor, Matthew Galvin, Michael D'Eletto, Sabrina Y. Camp, Stephen C. Van Nostrand, James Provencher, Joyce Yu, Naeem Tahir, Alexandra S. Bailey, Jonathan Wischhusen, Olga Kozyreva, Taylor Ortiz, Hande Tuncer, Jad El Masri, Alys Malcolm, Tali Mazor, Ethan Cerami, Kenneth L. Kehl
Dana-Farber Cancer Institute · Mayo Clinic
cs.AI, cs.LG
Submitted: 2026-08-12
Updated: 2026-08-14
Code: https://github.com/dfci/matchminer-ai-training
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 64/100
The gist: MatchMiner-AI is an open-source, privacy-preserving clinical trial matching pipeline for oncology, developed by researchers at Dana-Farber Cancer Institute and Mayo Clinic.
Key concepts
- Clinical Trial Matching
- This is the process of determining which experimental cancer treatments a patient might be eligible for. The AI tool automates this complex task, replacing manual searching through thousands of eligibility criteria to make the process faster and more accessible.
- Privacy-Preserving AI
- This approach ensures that medical data is handled securely. MatchMiner-AI was trained on synthetic (fake) patient records, meaning the AI never saw a single real person's medical history, allowing it to be shared freely without risking patient information.
- Synthetic Data Training
- The authors used a large language model to generate over a million fake patient histories and clinical documents. They trained their smaller, faster AI models on this synthetic data to teach them how to accurately match patients to trials without using real medical records.
- TrialSpace
- This is a custom AI model used in the pipeline. It embeds both patient summaries and trial eligibility criteria into a mathematical space. This allows the system to instantly find the most relevant trials for a patient by identifying similar characteristics.
Terminology
Summary
MatchMiner-AI is an open-source, privacy-preserving clinical trial matching pipeline for oncology, developed by researchers at Dana-Farber Cancer Institute and Mayo Clinic. The platform uses open-weight large language models (LLMs) to summarize patient histories from unstructured electronic health record (EHR) text and to extract target populations from trial eligibility documents. It was co-developed with practicing clinical oncologists and trained on fully synthetic EHR data, with no grounding in protected health information (PHI), to overcome restrictions on sharing solutions built on individual-level patient data.
The pipeline includes modules to: (1) extract core eligibility criteria (age, sex, cancer type/histology, cancer burden/treatment intent, biomarker requirements, and prior treatment requirements) for target patient populations, or spaces,
for each trial; (2) summarize longitudinal patient histories from unstructured EHR text using an LLM; (3) rank candidate patient-trial space
matches; (4) predict, on a 0–5 ordinal scale, whether and how specifically a given candidate patient meets the core eligibility criteria for a specific trial; and (5) predict whether a patient meets common exclusion criteria, such as uncontrolled brain metastases or a history of a recent second primary cancer, for a specific trial unrelated to the core eligibility criteria (a boilerplate exclusions
check).
For training, the clinicaltrials.gov API was queried in October 2025 to identify 13,160 active phase I-IV cancer clinical trials, yielding 34,519 extracted trial spaces. A subset of 3,000 trials (7,695 spaces) was randomly sampled for training. For each space, the Gemma-4-31B-IT teacher LLM was prompted to invent longitudinal sequences of clinical events for 5 hypothetical patients who met the space-defining eligibility criteria and 5 patients who did not quite
meet those criteria. This yielded 76,935 synthetic patient histories. For each clinical event corresponding to an imaging report, pathology report, oncologist assessment, or genomic sequencing report, Gemma-4-31B-IT was prompted to invent a full-length clinical document, yielding 1,295,607 synthetic documents. Text augmentation steps were applied, including randomly replacing generic anti-cancer drug names with brand names or common abbreviations, and incorporating a prior history of a second primary cancer for 10% of patients.
Patient histories were summarized using Gemma-4-31B-IT, processing records in chronological, overlapping chunks (approximately 50,000 tokens with a 500-token overlap) and maintaining a structured running summary. The summarization prompt instructed the LLM to generate semi-structured output capturing the same core clinical concepts used to define trial spaces, plus any history of conditions that might meet common boilerplate
exclusion criteria.
A TrialSpace
text embedding model was developed by fine-tuning Qwen3-0.6B-Embedding to embed patient summaries and trial spaces into the same mathematical vector space, using multiple negatives ranking loss and CoSENT loss. The model was fine-tuned in three iterations, with each iteration using the previous model to identify top candidate matches for further teacher-LLM scoring, improving discrimination among highly ranked candidates. A TrialChecker
re-ranking/regression model was distilled by fine-tuning ModernBERT-large to predict the Gemma-4-31B-IT 0–5 ordinal trial match quality score for candidate patient-trial space combinations. A BoilerplateChecker
classification model was also distilled onto ModernBERT-large to generate a predicted probability of exclusion based on patient and trial boilerplate text strings.
The pipeline was piloted with eight medical oncologists at DFCI-operated community-based sites over a six-month period, with feedback leading to modifications such as providing a direct link in the production EHR, incorporating age and sex into prompts, modifying synthetic note generation to include brand names and abbreviations, improving biomarker matching logic, redesigning patient summarization to process all clinical text, and placing inactive cancers in the boilerplate exclusions section.
Evaluation was performed using DFCI EHR data for 7,982 therapeutic trial enrollment events for 7,056 patients and 11,393 standard-of-care treatment starts for 5,758 patients. For trial-enrolled patients, the denominator included 814 trials (2,990 spaces) open at DFCI at the time of each patient's actual enrollment; for standard-of-care patients, a random sample of 500 trials (1,347 spaces) not open at DFCI was used. Evaluation trials were not included in training.
Across retrospective evaluations of distillation fidelity, the pipeline outperformed a baseline text-embedding model (Llama-Embed-Nemotron-8B), improving mean average precision (MAP) at 20 from 0.44 (95% CI 0.44-0.45) to 0.95 (95% CI 0.95-0.96) for trial-enrolled patients and from 0.38 (95% CI 0.37-0.38) to 0.94 (95% CI 0.93-0.94) for patients who received standard of care therapies. For the trial-enrolled cohort, the ModernBERT TrialChecker reproduced the teacher's binarized reasonable consideration
label with an AUROC of 0.95 (95% CI 0.95–0.95) in the patient-centric and 0.98 (95% CI 0.98–0.98) in the trial-centric use case; agreement with the teacher's full 0–5 ordinal score was strong (Spearman ρ 0.84, 95% CI 0.84–0.84; and 0.81, 95% CI 0.81–0.81 respectively). The BoilerplateChecker reproduced the teacher boilerplate exclusion
judgment with AUROC 0.95 (95% CI 0.95–0.96) and 0.95 (95% CI 0.95-0.95) for the patient-centric and trial-centric use cases respectively. Metrics stratified by patient age, sex, race, and ethnicity were broadly consistent across subgroups.
For the standard-of-care treatment cohort, the TrialChecker reproduced the teacher's binarized reasonable consideration
label with an AUROC of 0.92 (95% CI 0.92–0.92) in the patient-centric and 0.96 (95% CI 0.96–0.96) in the trial-centric use case; agreement with the teacher's full 0–5 ordinal score remained strong (Spearman ρ 0.82, 95% CI 0.82–0.82; and 0.85, 95% CI 0.84–0.85, respectively). The BoilerplateChecker reproduced the teacher's boilerplate exclusion
judgment with AUROC 0.97 (95% CI 0.97–0.98) in the patient-centric and 0.97 (95% CI 0.97–0.97) in the trial-centric use case.
In a 50-patient sample selected for comparison between MatchMiner-AI and a rules-based tumor genomic trial matching algorithm (MatchMiner-Genomics), MatchMiner-AI retrieved trials for all patients, as opposed to 19 patients (38%) who had tumor genomic data available. Among those 19 patients, 80% of 256 trial suggestions retrieved by MatchMiner-AI were deemed reasonable considerations by a frontier LLM (Claude Opus 4.8), vs 53% of 113 suggestions retrieved by the rules-based approach. MatchMiner-AI surfaced at least one reasonable trial for 100% of patients (19/19), while MatchMiner-Genomics did so for 74% (14/19).
LLM-as-judge evaluations using Claude Opus 4.8 and Gemini 3.5 Flash were performed. For patient summarization, after oncologist adjudication, 44 of 50 (88%, 95% CI 76%–94%) summaries were judged clinically correct by each judge, with defects concentrated in treatment-history and biomarker sections. For trial match quality scoring, Opus 4.8 agreed exactly with the pipeline on the 0–5 score for 65% of 100 patient–trial pairs (95% CI 55%–74%) and within one point for 91% (95% CI 84%–95%; quadratic-weighted κ 0.74, 95% CI 0.59–0.84; Spearman ρ 0.75, 95% CI 0.60–0.87), with 95% agreement (95% CI 89%–98%; κ 0.86, 95% CI 0.71–0.97) on the binary question of whether a trial was a reasonable consideration. Gemini 3.5 Flash agreed exactly with the pipeline in 72% of pairs (95% CI 63%–80%), within one point for 92% of pairs (95% CI 85%–96%; quadratic-weighted κ 0.76, 95% CI 0.62–0.88; Spearman ρ 0.76, 95% CI 0.61–0.88), and had 92% agreement (95% CI 85%–96%) and κ 0.79 (95% CI 0.62–0.92) for the binary reasonable consideration outcome. For boilerplate exclusion, Opus–pipeline agreement was 88% (95% CI 80%–93%; κ 0.69, 95% CI 0.51–0.84); Gemini-pipeline agreement was 82% (95% CI 73%–88%; κ 0.51, 95% CI 0.31–0.69).
A sample of 20 patient summaries was taken for independent dual annotation of trial match score components by two oncologists and three LLMs (Claude Opus 4.8, Gemma-4-31B-IT, and Gemma-4-E2B-IT). Across all score components, two oncologists agreed with one another numerically more often (91%, 95% CI 86%-96%) than any judge agreed with the index oncologist (84%, 95% CI 78%-90%), and the three judges did not differ detectably from one another, spanning under two percentage points across 125 concept pairs (McNemar's exact test, all p ≥ 0.68). All three judges were more self-consistent on test–retest than the oncologists were with one another. Scoring judges against the two-oncologist consensus increased every judge's agreement, indicating that a material share of apparent judge error was human disagreement.
The authors note limitations: MatchMiner-AI does not replace clinicians in weighing treatment options, is not a medical device or formal clinical decision support tool, the small TrialChecker and BoilerplateChecker text classifiers directly output predictions without generating reasoning traces, training and evaluation datasets were in English only, and generalizability to patient histories from other institutions requires further evaluation. The study demonstrates utility of synthetic EHR data for training AI pipelines for certain real-world applications. Evaluation and impact assessment among clinicians, research staff, and clinical trialists is ongoing.
Synthetic data, model weights, and demonstration apps are available at http://huggingface.co/ksg-dfci. Code for training and inference is available at https://github.com/dfci/matchminer-ai-training and https://github.com/dfci/matchminer-ai-inference respectively. The distilled task-specific models are released in the ksg-dfci 0526
model collection on Hugging Face (TrialSpace-0526, TrialChecker-0526, and BoilerplateChecker-0526). Funding was provided by Meta Corp Llama Impact Grant, Nancy Lurie Marks Family Foundation, United States NIH/NCI (R37CA295653), and United States Department of Defense (W81XWH2210086).
Improvements for AI systems
Based on the MatchMiner-AI paper, here are specific improvements I can implement in AI systems, along with what the improved system can do.
Improvement: Replace PHI-grounded training data with a synthetic data generation pipeline. Use a teacher LLM (e.g., Gemma-4-31B-IT) to generate: (a) longitudinal clinical event sequences for hypothetical patients matching trial criteria, and (b) full-length clinical documents (progress notes, imaging/pathology/NGS reports) from those sequences. Apply text augmentation (brand names, abbreviations, second-primary-cancer histories) to increase realism.
What the improved system can do: Train and share clinical AI models without exposing protected health information, eliminating memorization and membership-inference risks. Models can be openly distributed, enabling multi-institution collaboration and reproducibility without data-use agreements.
Abstract
Background: Clinical trials are essential to advancing cancer treatments, but fewer than 10% of adults with cancer enroll in therapeutic trials. Open-source AI trial matching tools could democratize access to trial options. Methods: We created MatchMiner-AI, co-developed with practicing clinical oncologists and trained on synthetic electronic health record (EHR) data. It uses open-weight LLMs to summarize patient histories from unstructured EHR text and extract target populations from trial eligibility documents. Embedding and re-ranking models were distilled to retrieve and rank trial and patient suggestions. Multifaceted evaluation was performed, including retrospective quantification of distillation fidelity; applying a closed-source LLM as judge of patient summarization and matching; and evaluation of candidate matches by oncologists. Results: Across retrospective evaluations of distillation fidelity, the pipeline outperformed a baseline text-embedding model, improving mean average precision (MAP) at 20 from 0.44 (95% CI 0.44-0.45) to 0.95 (95% CI 0.95-0.96) for trial-enrolled patients and from 0.38 (95% CI 0.37-0.38) to 0.94 (95% CI 0.93-0.94) for patients who received standard of care therapies. In a 50-patient sample selected for comparison between MatchMiner-AI and a rules-based tumor genomic trial matching algorithm, MatchMiner-AI retrieved trials for all patients, as opposed to 19 patients (38%) who had tumor genomic data available. Among those 19 patients, 80% of 256 trial suggestions retrieved by MatchMiner-AI were deemed reasonable considerations by a frontier LLM, vs 53% of 113 suggestions retrieved by the rules-based approach. Conclusion: MatchMiner-AI is an open-source, open-weights, clinical trial matching AI pipeline for oncology. Synthetic training data, model weights, inference tools, and demonstration frontends are publicly available.
Sources
- DeepEnroll: Patient-Trial Matching with Deep Embedding and Entailment Prediction
- COMPOSE: Cross-Modal Pseudo-Siamese Network for Patient Trial Matching
- Utilizing ChatGPT to Enhance Clinical Trial Enrollment
- Zero-Shot Clinical Trial Patient Matching with LLMs
- Efficient Natural Language Response Suggestion for Smart Reply
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
- In Defense of Cross-Encoders for Zero-Shot Retrieval
- Gemma: Open Models Based on Gemini Research and Technology
- Llama-Embed-Nemotron-8B: A Universal Text Embedding Model for Multilingual and Cross-Lingual Tasks
- The challenge of uncertainty quantification of large language models in medicine
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection