Extracting Information in a Low-resource Setting: Case Study on Bioinformatics Workflows

arXiv:2411.19295 · cs.CL · Submitted 2025-03-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Extracting Information in a Low-resource Setting: Case Study on Bioinformatics Workflows".

Jane: The paper was written by Clémence Sebe, Sarah Cohen-Boulakia, Olivier Ferret and Aurélie Névéol from University of Paris-Saclay and National Centre for Scientific Research (CNRS) and Commissariat à l'énergie et aux atomique (CEA).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: So, we’ve established the problem; now let’s look at what the paper actually summarized. They started by framing this as a low-resource extraction task and tested four main strategies to solve it.

Jane: It sounds like they were trying every angle—from using generative models to using standard NLP techniques—to see what would work best on this specific data.

Lu: The introduction of BioToFlow, their own annotated corpus, is a massive step toward providing that structured data needed for reliable AI performance.

Meng: I'm interested in how they created it; fifty-two articles covering Nextflow and Snakemake workflows, which totals over seventy-eight thousand tokens.

Lalam: It’s a very focused and high-quality dataset, ensuring the machine learning model isn't just guessing based on general medical text but is specifically looking at workflow components.

Tom: The results showed that while few-shot models like Llama-three were challenging, the overall performance was much better when they used an encoder-based approach.

Jane: The core finding here is that you simply can't rely on massive, general language models for this specialized domain without significant effort.

Lu: They had to build their own resource because standard industry datasets weren't designed to capture the specific entities needed, like 'ManagementSystem' or 'Container.'

Meng: This tells us that when we want specialized extraction, we need highly targeted data collection rather than just relying on what’s already available.

Lalam: We are learning that the quality of our training data dictates whether AI can handle the complexity of scientific documentation.

Improvements and Methods: Tom: Now, let’s dig into the methods, because there were several approaches they tested and see which really delivered results. They tried decoder-based approaches using autoregressive models first.

Jane: Those initial results were pretty low, hovering around forty percent F1 in relaxed mode. It seems like the complexity of the language was just too much for those models to handle consistently.

Lu: That low performance is a clear signal that these types of generative models need a very specific type of input or prompt structure that we haven't fully standardized yet.

Meng: The breakthrough, I think, was when they started using the biLSTM-CRF neural model with SciBERT and merging it with their BioToFlow data.

Lalam: This fusion of SoftCite and BioToFlow showed the potential of combining external knowledge with internal domain expertise to improve extraction.

Tom: And they weren't stopping there; they tried integrating external knowledge, like lists of bioinformatics tools, directly into the language models.

Jane: It was interesting to see how that helped them specifically—the addition of tool vocabulary improved recognition for key entities like 'Tool' and 'Data.'

Lu: That’s a classic case of domain-specific knowledge injection boosting performance, proving that specialized knowledge is often more valuable than general model power.

Meng: It’s a practical win; we know exactly which parts of the model need targeted updates to help us automate this extraction process successfully.

Lalam: By showing that specific improvements work, they are providing a clear roadmap for how AI can better serve scientific communication and documentation.

Practical Application: Tom: We’ve seen the methods, but let’s talk about the real-world application of this research. How does it actually help someone running a bioinformatics pipeline?

Jane: Imagine a researcher finds a great paper describing how they processed their DNA sequences using BWA, but they need to know exactly which specific version or configuration was used.

Lu: The ability to extract that precise information is what turns this abstract research into practical utility; it helps reconstruct the experimental environment.

Meng: That's the engineering value—we can now automate the creation of execution environments based on extracted parameters and tool versions, making reproducibility much easier.

Lalam: It allows us to connect the theoretical description in a paper to the physical reality of running code in a public repository like GitHub.

Tom: So, by identifying those core entities—the tools, the data, and the environment—they are creating a digital twin of that scientific workflow.

Jane: They are making sure that even if someone is using Nextflow or Snakemake, they can pull out the necessary details to understand how it was set up.

Lu: This moves us away from just "reading" a description to actively "parsing" and the specific components of a functional system.

Meng: We could use this to build automated documentation services for public repositories that currently lack human-readable explanations for their scripts.

Lalam: This greatly reduces the barrier to entry, allowing researchers to share their methods without having to write exhaustive, standardized documentation themselves.

Conclusion: Tom: As we wrap up this deep dive into "Extracting Information in a Low-resource Setting: Case Study on Bioinformatics Workflows," we see that while the field is highly specialized, effective solutions exist.

Jane: The conclusion seems to be that for complex tasks like this, you need both a small, high-quality domain corpus and powerful language models that are tailored to the specific knowledge of the bioinformatics community.

Lu: I think we should look forward to more advanced few-shot learning techniques being applied here.

Meng: And I agree; exploring ways to merge information from different sources is the next practical step for automation.

Lalam: We must ensure that these insights lead to a world where scientific knowledge is shared not just quickly, but completely and accurately documented.

Tom: It’s been a fascinating look at how we can apply AI to solve some of the most challenging data extraction problems in science.

Jane: Thank you for listening, everyone, and we hope this paper inspires more advanced work in the future are exploring these methodologies.

Clémence Sebe, Sarah Cohen-Boulakia, Olivier Ferret, Aurélie Névéol

University of Paris-Saclay · National Centre for Scientific Research (CNRS) · Commissariat à l'énergie et aux atomique (CEA)

cs.CL

Submitted: 2025-03-10

Updated: 2026-08-25

Journal ref: Advances in Intelligent Data Analysis XXIII. IDA 2025

DOI: 10.1007/978-3-031-91398-3_21

Code: https://github.com/marconaguib/autoregressive_ner

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 76/100

The gist: This paper investigates methods for extracting detailed information from scientific articles describing bioinformatics workflows, a task complicated by the "low-resource context" caused by a lack of

Key concepts

BioToFlow
This is a specific, annotated corpus created by the researchers. It consists of fifty-two articles detailing Nextflow and Snakemake workflows, totaling over seventy-eight thousand tokens. This focused dataset provides the structured data necessary for reliable AI performance in this specialized domain.
Low-resource extraction task
This refers to the challenge of extracting specific, complex entities (like 'ManagementSystem' or 'Container') from scientific documentation. Standard industry datasets are often inadequate for these tasks, requiring highly targeted data collection rather than relying on general knowledge bases.
Domain-specific knowledge injection
This method involves integrating specialized information, such as lists of bioinformatics tools, directly into the language models. This injection boosts performance significantly, improving the recognition of key entities like 'Tool' and 'Data' within scientific text.

Terminology

Summary

This paper investigates methods for extracting detailed information from scientific articles describing bioinformatics workflows, a task complicated by the low-resource context caused by a lack of annotated corpora. Because these workflows are essential for complex biological data analyses but are often described in non-executable forms or poorly documented code repositories, automating their documentation is critical for improving accessibility and reusability in data-intensive sciences.

The Research Problem

Bioinformatics workflows consist of numerous steps that orchestrate tool execution and require specific computational environments to ensure reproducibility. Currently, researchers face two primary challenges: (1) an increase in articles describing workflows in a descriptive form without executable code, and (2) a rise in programmatic implementations on repositories like GitHub that lack sufficient documentation. To address this, the authors frame the problem as a low-resource extraction task and test four specific strategies:

** Creating a tailored annotated corpus; 1) few-shot named-entity recognition (NER) with an autoregressive language model; 2) NER using masked language models with existing and new corpora; and 3) integrating workflow knowledge into NER models. 1**

The BioToFlow Corpus

To facilitate this research, the authors developed BioToFlow, a new corpus composed of 52 articles (balanced between Nextflow and Snakemake workflows) containing approximately 78,419 tokens. The researchers developed a Workflow Representation Schema to categorize entities into three main groups:

** Core entities: representing major components such as the workflow name, bioinformatics tools (classified as BioInfo, Lab, Context, or General), methods, and data. 1**

** Environment entities: capturing resources required for execution, such as management systems or programming languages. 1**

** Specific details: including versioning information, bibliographic references (Biblio), and parameters. 1**

Experimental Methodologies

The study evaluates several NLP approaches to perform Named-Entity Recognition (NER). First, they tested a decoder-based approach using Llama-3-8B-Instruct for few-shot NER, which yielded an overall performance below 40% in relaxed mode. Second, they explored encoder-based approaches using SciBERT and the NLStruct library. This included experiments on a large existing corpus (SoftCite) and the BioToFlow corpus itself. The researchers also investigated the injection of domain-specific knowledge directly into language models by adding bioinformatics tool names from knowledge bases like Biotools, Bioconda, Biocontainers, and Bioweb into the SciBERT vocabulary.

Key Findings and Results

The results demonstrate that while generative models struggle with the specificity of these entities, encoder-based models achieve much higher accuracy. Training a model on the BioToFlow corpus yielded an overall F1-measure of 70.4%, which is comparable to inter-annotator agreement. The integration of domain knowledge proved particularly effective for specific tasks; specifically, adding bioinformatics tool vocabulary improved the extraction of tool names, increasing the F1 score from 74.8% to 77.0% when fine-tuning was applied. However, this knowledge integration was less effective across the entire information schema, as some entities experienced slight declines in performance. The study concludes that high-performance information extraction is achievable but requires domain-specific annotated data and tailored language models.of the paper"

Improvements for AI systems

To improve existing Artificial Intelligence systems based on the methodologies presented in this paper, I propose the following specific architectural and procedural enhancements:

  1. Implement a Hybrid Corpus Training Strategy for specialized domains. Instead of relying solely on small, high-quality domain-specific datasets (which are expensive to produce) or large general datasets (which lack precision), AI systems should be trained using a multi-stage approach: first training on large, partially aligned silver corpora (like SoftCite) and then fine-tuning on a smaller, expert-annotated gold corpus (like BioToFlow).

  2. Integrate Vocabulary Injection via Domain Knowledge Bases. For Named Entity Recognition (NER) tasks in niche fields, the system should not rely solely on subword tokenization. Instead, it should explicitly inject specialized vocabularies—derived from external authoritative registries and knowledge bases (e.g., Bioconda, Biotools)—directly into the model's embedding layer to prevent information loss during tokenization.

  3. Develop Multi-Level Entity Recognition Architectures capable of handling nested entities. Standard NER models often struggle with hierarchical structures; systems should be specifically architected (using frameworks like NLStruct) to recognize that one entity can exist within another (e.g., a bibliographic reference nested within a workflow name).

  4. Adopt Encoder-based Masked Language Model Pre-training for low-resource extraction tasks. The paper demonstrates that while generative/decoder models (like Llama-3) are versatile, encoder-based models (like SciBERT) paired with biLSTM-CRF layers provide superior precision and recall for specific entity extraction in scientific contexts.


By implementing these improvements, the resulting AI system will be able to:

  1. Perform highly accurate, automated extraction of complex technical workflows from unstructured scientific literature.

  2. Identify and categorize highly specialized entities (tools, parameters, software versions, and hardware requirements) that are typically missed by general-purpose LLMs.

  3. Systematically transform descriptive scientific text into structured data formats suitable for automated registration in technical repositories (e.g., GitHub or workflow management systems).

  4. Operate effectively in low-resource settings where massive amounts of annotated training data are unavailable, by leveraging the strategic fusion of existing datasets and domain-specific knowledge injection.

Abstract

Bioinformatics workflows are essential for complex biological data analyses and are often described in scientific articles with source code in public repositories. Extracting detailed workflow information from articles can improve accessibility and reusability but is hindered by limited annotated corpora. To address this, we framed the problem as a low-resource extraction task and tested four strategies: 1) creating a tailored annotated corpus, 2) few-shot named-entity recognition (NER) with an autoregressive language model, 3) NER using masked language models with existing and new corpora, and 4) integrating workflow knowledge into NER models. Using BioToFlow, a new corpus of 52 articles annotated with 16 entities, a SciBERT-based NER model achieved a 70.4 F-measure, comparable to inter-annotator agreement. While knowledge integration improved performance for specific entities, it was less effective across the entire information schema. Our results demonstrate that high-performance information extraction for bioinformatics workflows is achievable.

Sources

Related papers