InforMask: Unsupervised Informative Masking for Language Model Pretraining
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "InforMask: Unsupervised Informative Masking for Language Model Pretraining".
Jane: The paper was written by Nafis Sadeq, Canwen Xu and Julian McAuley from University of California, San Diego.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Hey everyone, welcome back to the show. Today we’re digging into a paper that’s got a pretty catchy name — InforMask: Unsupervised Informative Masking for Language Model Pretraining.
Jane: And Tom, I gotta say, the title alone tells you a lot. It’s about how we train language models — you know, the big ones like BERT — by hiding words and asking the model to guess them.
Tom: Right, that’s the masked language model trick. But the twist here is the word “informative.” Instead of just hiding random words, these authors from UC San Diego — Nafis Sadeq, Canwen Xu, and Julian McAuley — figured out a way to hide the words that actually matter.
Jane: Exactly. Think of it like studying for a test. If you quiz yourself on the easy stuff every time, you’re not really learning. But if you focus on the tricky, important facts, you get better faster.
Tom: And that’s what InforMask does. It uses something called Pointwise Mutual Information — basically a measure of how strongly two words are connected — to decide which words are worth masking.
Jane: So instead of masking “the” or “and” all the time, it goes after names, places, and other high-value tokens. That’s a big deal because those are the words that carry actual knowledge.
Tom: And the best part? It’s fully unsupervised. No need for human labels or external knowledge bases. It just looks at the corpus and figures it out on its own.
Jane: That’s huge for scalability. You can apply this to any language or domain without extra annotation cost.
Tom: And the results are pretty wild — they show big gains on factual recall benchmarks. But we’ll get into that in a bit.
Jane: For now, let’s just say this paper is asking a simple question: what if we stopped masking randomly and started masking smart?
Tom: And the answer seems to be — you get a model that learns more, faster. Stick around, because we’re about to break down how they actually did it.
Summary: Tom: So Jane, we’ve got the title — InforMask — but let’s talk about what the paper actually claims in its summary. This is one of those papers where the abstract packs a punch.
Jane: It really does. The core claim is that random masking is suboptimal. When you mask tokens randomly, you waste training signal on words that are too easy to guess.
Tom: And that’s not just a hunch — they back it up with experiments. They compare random masking against their InforMask strategy, and InforMask wins on both the LAMA benchmark for factual recall and SQuAD for question answering.
Jane: LAMA is like a pop quiz for language models — it asks things like “Thomas Edison was born in MASK” and checks if the model knows the answer.
Tom: And on that benchmark, InforMask beats random masking, span masking, and even PMI-Masking, which is a previous attempt at smarter masking.
Jane: What’s really impressive is that their model — they call it InformBERT — outperforms BERT and even RoBERTa on LAMA, even though RoBERTa was trained on ten times more data.
Tom: Ten times. That’s not a small gap. RoBERTa-base was trained on one hundred sixty gigabytes of text, and InformBERT only saw sixteen gigabytes.
Jane: And yet it still comes out ahead on factual recall. That tells you the masking strategy itself is doing a lot of the heavy lifting.
Tom: The paper also introduces a clever efficiency trick — they precompute token-specific masking rates once, so that during training, masking is just as fast as random masking.
Jane: That’s important because if your masking strategy is too slow, it becomes a bottleneck for large-scale training.
Tom: So the summary is: smarter masking leads to better knowledge retention, and they made it efficient enough to actually use in practice.
Jane: And that’s the kind of result that makes you wonder why we’ve been masking randomly for so long.
Tom: Good question. And that’s exactly what we’re going to dig into next — the improvements they propose over existing methods.
Improvements: Tom: Alright Jane, so we’ve covered the big picture. Now let’s get into the improvements — what does InforMask actually do differently?
Jane: So the key idea is something they call “Informative Relevance.” It’s a score they calculate for each token based on how much information it shares with the rest of the sentence.
Tom: And they use Pointwise Mutual Information to compute that. PMI tells you how surprising it is to see two words together — like “Harry” and “Potter” have high PMI because they co-occur a lot.
Jane: Right. So if you’re masking one word, you want to pick the one that has the highest total PMI with all the unmasked words. That way, the model has enough hints to make a reasonable guess, but the task is still challenging.
Tom: And that’s the sweet spot — you don’t want it too easy, but you also don’t want it impossible.
Jane: Exactly. But here’s the tricky part — finding the best set of tokens to mask in a sentence is computationally expensive. You can’t try every combination.
Tom: So they came up with a clever workaround. Instead of searching all possibilities, they randomly sample a bunch of masking candidates — like thirty per sentence — and then score each one.
Jane: And they pick the one with the highest informative score. That gives you a good approximation without the heavy computation.
Tom: But there’s another problem — if you run this every epoch, it’s still slow. So they precompute token-specific masking rates from a small sample of the corpus.
Jane: That’s the part I love. They run the algorithm once, count how often each token gets masked, and then use those rates for all future epochs.
Tom: So during training, masking is just as fast as random masking — but the tokens being masked are much more informative.
Jane: And they show this approximation works — it actually beats repeating the same masked data every epoch, because it introduces more diversity.
Tom: So the improvements are threefold: a smarter scoring metric, a fast sampling strategy, and a one-time preprocessing step that makes it scalable.
Jane: And the results speak for themselves — better factual recall, better question answering, and better training efficiency.
Tom: Now, let’s take a closer look at the first page of the paper to see how they frame the problem.
First Page: Tom: So Jane, we’re flipping to page one of InforMask, and the authors start by laying out the problem with random masking.
Jane: And they make a really good point — random masking sometimes produces masks that are too easy. If you mask a stop word like “the,” the model can guess it without even looking at the context.
Tom: That’s wasted training signal. The model isn’t learning anything new.
Jane: But there’s another issue too — some tokens are just more important than others. Named entities like “Thomas Edison” carry a lot of factual weight, and they should be masked more often.
Tom: And the paper shows a really nice example with the sentence “Thomas Edison was an inventor and businessman.” They generate four different masking candidates and score them.
Jane: The best one masks “Thomas Edison” together, because that forces the model to use the context — “inventor and businessman” — to figure out who it is.
Tom: And the worst one masks “businessman” — that’s too easy because the word “inventor” is right there.
Jane: So the idea is to find masks that are interesting and challenging — not too easy, not impossible.
Tom: And they also point out that previous methods like span masking have a weakness — they mask whole spans, which can break up named entities like “Mount Fuji” or “Mona Lisa.”
Jane: That’s a great observation. If you mask the whole span, you lose the relationship between the words inside it. InforMask avoids that by focusing on individual tokens.
Tom: And that’s why it does so well on knowledge-intense tasks — because it’s specifically targeting the words that carry knowledge.
Jane: The first page really sets the stage — it’s clear, it’s motivated, and it makes you want to read on.
Tom: And it also sets up the key question: can we automatically identify the most informative tokens without any supervision?
Jane: And the answer, as we’ve seen, is yes — using PMI and a smart sampling strategy.
Tom: Alright, let’s wrap this up and talk about what it all means.
Conclusion: Tom: So Jane, we’ve spent the whole episode on InforMask: Unsupervised Informative Masking for Language Model Pretraining. Let’s pull it all together.
Jane: The big takeaway is that masking strategy matters a lot. It’s not just about how many tokens you mask — it’s about which ones.
Tom: And InforMask solves that by using Pointwise Mutual Information to find the most informative tokens, then sampling and scoring masking candidates to pick the best one.
Jane: And they made it efficient with token-specific masking rates, so you get the benefit without the slowdown.
Tom: The results are genuinely impressive — InformBERT beats BERT and RoBERTa on LAMA despite being smaller and trained on less data.
Jane: And on SQuAD, it holds its own too. That’s real-world question answering, not just a toy benchmark.
Tom: The implications are pretty big. If we can make pretraining more efficient just by changing how we mask, that saves compute, time, and money.
Jane: And it’s fully unsupervised, so it can be applied to any language or domain without extra annotation.
Tom: There are some limitations, of course — the authors mention they couldn’t scale to larger models or tune hyperparameters due to compute constraints.
Jane: But even so, the gains are clear, and the method is simple enough to be adopted widely.
Tom: I think this is one of those papers that could quietly change how everyone does pretraining.
Jane: Totally. It’s a small tweak with a big payoff.
Tom: Alright, that’s a wrap on InforMask. Thanks for listening, and we’ll see you next time with another paper.
Jane: Bye everyone!
Nafis Sadeq, Canwen Xu, Julian McAuley
University of California, San Diego
cs.CL
Submitted: 2022-10-21
Journal ref: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP 2022)
DOI: 10.18653/v1/2022.emnlp-main.395
Code: https://github.com/NafisSadeq/InforMask
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 59/100
The gist: InforMask is a new unsupervised masking strategy for training masked language models, proposed to address the suboptimality of random masking, which allocates an equal masking rate for all tokens.
Key concepts
- InforMask
- A method for language model pretraining that improves upon random word masking. It uses Pointwise Mutual Information (PMI) to identify the most informative tokens, ensuring the model learns from challenging, knowledge-carrying words.
- Pointwise Mutual Information (PMI)
- A measure used in InforMask to calculate how strongly two words are connected or co-occur. High PMI indicates that two words are related and therefore good candidates for masking.
- Masked Language Model (MLM)
- A common language model training technique where the model is trained by hiding certain words (masking them) and then predicting what those missing words should be.
Terminology
Summary
InforMask is a new unsupervised masking strategy for training masked language models, proposed to address the suboptimality of random masking, which allocates an equal masking rate for all tokens. The method exploits Pointwise Mutual Information (PMI) to select the most informative tokens to mask, aiming to improve performance on knowledge-intense tasks such as factual recall and question answering.
The paper introduces Informative Relevance,
a metric based on PMI that measures the quality of a masking choice by summing the PMI between each masked word and all unmasked words in a sentence. This metric ensures the informativeness of the masked token while maintaining a moderate difficulty for the model to predict it. The PMI matrix is calculated corpus-wise using skip-gram co-occurrence within a window, enabling sentence-level and local co-occurrence to be considered.
To address the computational challenge of maximizing total Informative Relevance for a text sample with multiple masks, the authors propose a sample-and-score algorithm. This algorithm randomly generates s masking candidates per document, calculates their informative scores, and selects the candidate with the highest score. This reduces time complexity to O(kn) and introduces diverse masking patterns. For training over multiple epochs, they further propose token-specific masking rates, computed once as a preprocessing step by counting the frequency of each token being masked according to the algorithm. This approximation allows masking to be as fast as random masking without further overhead.
Experiments were conducted on the LAMA factual recall benchmark and SQuAD v1 and v2 question answering benchmarks. The authors pretrained base-size BERT models with different masking strategies (random, SpanBERT, PMI-Masking, and InforMask) for 3 epochs using the same corpus (Wikipedia and BookCorpus, 3.3B tokens) and hyperparameters. InforMask outperformed all other masking strategies on all subsets of LAMA and on both SQuAD datasets. For example, on LAMA, InforMask achieved an overall MRR of 0.591, compared to 0.549 for random, 0.495 for span, and 0.522 for PMI-Masking. On SQuAD v1, InforMask achieved an F1 of 80.47 and EM of 71.41, outperforming random (79.08 F1, 69.44 EM), span (78.88 F1, 69.04 EM), and PMI-Masking (80.31 F1, 70.98 EM).
The authors also trained a 40-epoch model, InformBERT, and compared it to BERT-base (trained with the same corpus and epochs), BERT-large, RoBERTa-base, and RoBERTa-large. InformBERT outperformed BERT-base by 0.145 overall on LAMA (0.698 vs. 0.553 MRR) and outperformed both RoBERTa-base and RoBERTa-large despite having 1/3 the parameters and 1/10 the corpus size of RoBERTa. On SQuAD v1 and v2, InformBERT also outperformed BERT-base (e.g., SQuAD v1 F1: 81.22 vs. 81.07; SQuAD v2 F1: 72.71 vs. 72.35).
Training dynamics analysis showed InforMask outperforms other masking strategies from the beginning of training and maintains the lead. The model trained with InforMask outperforms BERT and RoBERTa with fewer than 15% of the training steps. The token-specific masking rate approximation was shown to be more effective than looping the same masked data each epoch, as it introduces more diverse patterns.
The paper also analyzes why PMI-Masking underperforms random masking on LAMA. PMI-Masking increases masking rates for tokens in correlated spans but decreases rates for unigram named entities not in such spans, whereas InforMask increases masking rates for tokens with high informative saliency regardless of span membership. Case studies show InformBERT correctly predicts answers like ESPN
and bishops
where RoBERTa fails.
The authors note limitations due to computational budget, including not scaling to larger models or full pretraining for all baselines, and not performing hyperparameter tuning. They also acknowledge potential social biases in the model, similar to other language models.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems:
Improvement: Replace random token masking with InforMask's PMI-based informative masking strategy during pretraining.
What the improved system can do:
-
Automatically identify and prioritize masking of high-information tokens (e.g., named entities like
Voldemort
,Microsoft
) while reducing masking of stop words (the
, "of", "in") -
Achieve superior factual recall performance: 0.698 MRR on LAMA benchmark vs 0.553 for BERT-base and 0.592 for RoBERTa-base, despite using only 1/10th of RoBERTa's training data
-
Reach comparable performance to RoBERTa-large (355M params, 160GB corpus) using only 125M params and 16GB corpus, after just 120k training steps
These improvements are directly implementable using the paper's Algorithm 1 and the released code/checkpoints, requiring no additional supervision or external resources beyond the training corpus itself.
Sources
- Common Sense or World Knowledge? Investigating Adapter-Based Knowledge Injection into Pretrained Transformers
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- ERNIE: Enhanced Representation through Knowledge Integration
- Should You Mask 15% in Masked Language Modeling?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering