Unsupervised Improvement of Factual Knowledge in Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Unsupervised Improvement of Factual Knowledge in Language Models".
Jane: The paper was written by Nafis Sadeq, Byungkyu Kang, Prarit Lamba and Julian McAuley from Intuit and University of California, San Diego.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Alright, welcome back to the show, everybody. Today we’re digging into a paper that’s got a wonderfully straightforward title: “Unsupervised Improvement of Factual Knowledge in Language Models.” I’m Tom, and as always, I’ve got Jane here with me.
Jane: Hi, folks! And Tom, I gotta say, just reading that title makes me happy. We spend so much time hearing about bigger models, more data, fancier architectures. This one is about tweaking how we train the models we already have, and it’s all done without any external knowledge bases. That’s a big deal.
Tom: Right, and the team behind it is from UC San Diego and Intuit—Nafis Sadeq, Byungkyu Kang, Prarit Lamba, and Julian McAuley. They’re essentially asking a simple question: when a language model is learning to predict masked words, can we make it focus on the words that actually carry factual weight, like names, dates, and places?
Jane: And their answer is yes, and they do it in a totally unsupervised way. No knowledge graphs, no entity embeddings, no expensive hand-crafted resources. Just a clever reweighting of the pretraining objective itself.
Tom: So for anyone tuning in late, the core idea is that traditional masked language modeling treats every word equally. But in a sentence like “Marie Curie was born in MASK,” the word “Warsaw” is way more informative than the word “was.” The model should be penalized more for getting “Warsaw” wrong than for missing a stopword.
Jane: Exactly. And the authors show that by making the model prioritize those informative tokens, it gets significantly better at recalling facts later on. We’re talking about a seventeen point five percent relative improvement in Mean Reciprocal Rank on the ConceptNet portion of the LAMA benchmark, which is a solid jump.
Tom: And that’s just the start. They also see gains on question answering, sentiment analysis, and natural language inference, all in a closed-book setting. So the model isn’t looking anything up; it’s just using what it stored during pretraining.
Jane: I love that this is a pretraining-only change. You don’t need to change your model architecture or your downstream fine-tuning. You just train it smarter from the get-go.
Tom: And that’s the hook for today’s episode. We’re going to break down exactly how they compute this “informative relevance” and why it works so well. Stick around.
Summary: Tom: So, Jane, we’ve set the stage. Let’s get into the meat of the paper. The authors propose two key changes to the standard masked language modeling recipe.
Jane: Right. First, instead of masking fifteen percent of tokens uniformly, they mask tokens with a variable rate. Tokens that are more “informative” get masked more often, up to fifty percent of the time, while the average stays around nineteen percent. The second change is that they use a weighted cross-entropy loss, so mistakes on informative tokens are penalized more heavily, with weights normalized between one and five.
Tom: And how do they figure out which tokens are informative? That’s the clever part.
Jane: They use Pointwise Mutual Information, or PMI. Basically, they look at word co-occurrence statistics across the whole Wikipedia corpus. If a word frequently appears near specific other words, it gets a high PMI score. For example, “Nintendo” probably has high PMI with “console” or “Mario,” whereas a word like “the” has low PMI with everything because it’s everywhere.
Tom: So they build a giant matrix of PMI values between all pairs of words in the vocabulary—one hundred thousand by one hundred thousand—and then for each document, they sum up the PMI scores for each token against all the other tokens in that document. That gives a per-token informative relevance score.
Jane: Exactly. And that one-pass computation takes about two hours and eleven gigabytes of memory. It’s not trivial, but it’s a one-time cost before pretraining, and it’s completely unsupervised. No human labels, no knowledge bases.
Tom: So they train four models to isolate the effects: a baseline with uniform masking and uniform penalty, one with just weighted penalty, one with just variable masking, and then their full model with both. The full model, which they call BERTvw, consistently wins.
Jane: And the ablation study shows that variable masking helps a bit more than the weighted penalty, but combining them gives the best results. It’s a nice, clean experimental design.
Tom: I also appreciate that they didn’t just test on LAMA. They fine-tuned on SQuAD and GLUE, and they used AutoPrompt for zero-shot sentiment and NLI. The gains on fine-tuning are smaller, but they’re still positive on most tasks.
Jane: Right, and that’s actually a really interesting finding. The authors point out that fine-tuning can wash out some of the factual knowledge gained during pretraining, which is something other researchers have observed too. But in prompt-based settings, the knowledge shines through.
Tom: So the summary is: a smarter masking strategy, computed for free from the corpus, leads to measurably better factual recall and downstream performance. That’s a strong result.
Jane: And it opens the door for a lot of follow-up work. Which is exactly what we’re going to talk about next.
Improvements: Tom: Welcome back. We’ve covered the basics, but now I want to bring in our senior researcher, Lu, and our engineer, Meng, because this paper has implications that go beyond just the numbers.
Lu: Thanks, Tom. I’ve been thinking about this all morning. The key improvement here isn’t just the performance boost—it’s the philosophy. The authors are saying that not all tokens are created equal, and the training objective should reflect that. This is a step toward more efficient learning, where the model spends its capacity on what matters.
Jane: And Lu, you’re right. It’s like studying for an exam. If you spend equal time on every page of the textbook, you’ll be prepared for the easy questions but might miss the key facts that show up on the final. This paper is essentially telling the model to highlight the important passages.
Meng: From an engineering standpoint, I’m curious about the practical side. The PMI computation is a one-time cost, but it’s still a 100k by 100k matrix. For a smaller team or a niche domain, is that feasible?
Tom: That’s a great question, Meng. The authors mention it takes about two hours and eleven GB of RAM. That’s not nothing, but it’s also not prohibitive. And you could imagine using a smaller vocabulary or a subsampled corpus to bring that cost down.
Meng: Right, and the other thing is that this is a drop-in replacement for the MLM objective. You don’t need to change the model architecture or the training pipeline. You just swap in the new masking rates and loss weights. That’s a huge win for adoption.
Lu: And the implications go further. Think about multilingual models. Knowledge bases are expensive to build for low-resource languages, but this method only needs text. So you could apply the same trick to a Swahili or Hindi corpus and get better factual recall without any external resources.
Jane: That’s a really exciting point, Lu. The authors explicitly mention that existing knowledge bases may not be available for all languages and domains. This method sidesteps that entirely.
Meng: I also want to highlight the case studies in the paper. They show that the baseline model often produces generic words like “computer” or “american” for a prompt about a gaming company, while their model produces “nintendo,” “walt,” and “atari.” That’s a qualitative improvement in specificity, not just a metric bump.
Tom: And that specificity is what makes the model useful in real-world applications. If you’re building a question-answering system, you want it to say “Nintendo,” not “computer.”
Lu: Exactly. And I think the next step is to apply this to generative models like T5. The authors mention that as future work, and I bet we’ll see results there within the year.
Jane: So the improvements are clear: better performance, no external resources, and a clear path to broader applications. But what does this mean for the field as a whole? That’s our final segment.
Conclusion: Tom: Alright, we’re in the home stretch. Let’s wrap up our discussion of “Unsupervised Improvement of Factual Knowledge in Language Models.”
Jane: So, to recap: the paper introduces a way to make masked language modeling smarter by focusing on informative tokens. They compute token importance using PMI, then use that to drive both masking rates and loss weights. The result is a model that stores more factual knowledge and performs better on a range of tasks.
Tom: And the beauty is that it’s fully unsupervised. No knowledge graphs, no entity embeddings, no human annotations. Just a clever reweighting of the training objective.
Lu: I’d add that this is a reminder that pretraining objectives are still an underexplored frontier. We’ve been using the same MLM recipe for years, and this paper shows that a thoughtful tweak can yield significant gains.
Meng: From a practical standpoint, it’s a low-risk, high-reward change. Any team training a BERT-style model can adopt this with minimal effort. And the code is public, so you can try it yourself.
Jane: And the implications for the world? More efficient language models that can answer factual questions without needing to look things up. That’s valuable for everything from search engines to virtual assistants to educational tools.
Tom: There are limitations, of course. The model underperforms on some syntax-heavy tasks like CoLA, and there’s a risk of amplifying biases if the informative tokens are biased. But these are solvable problems.
Lu: I’m excited to see where this goes. If the authors extend this to generative models, we could see even bigger gains in text-to-text tasks.
Tom: Well said, Lu. So that’s a wrap on “Unsupervised Improvement of Factual Knowledge in Language Models.” Thanks to everyone who tuned in. We’ll be back next time with another paper to dissect.
Jane: Until then, keep asking questions and keep learning. Goodbye, everyone!
Nafis Sadeq, Byungkyu Kang, Prarit Lamba, Julian McAuley
Intuit · University of California, San Diego
cs.CL
Submitted: 2023-04-04
Updated: 2026-08-18
Code: https://github.com/intuit/wMLM
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 69/100
Key concepts
- Masked Language Modeling (MLM)
- A training technique where the model learns by predicting masked or hidden words within a sentence. The paper improves this standard method by making the model focus more heavily on factually important tokens like names and dates.
- Pointwise Mutual Information (PMI)
- A statistical measure used to determine how frequently two words appear together in a corpus. High PMI scores indicate that a word is highly informative because it frequently co-occurs with specific other words, rather than being general stopwords.
- Unsupervised Improvement
- The method improves model performance without requiring human labeling, external knowledge graphs, or expensive hand-crafted resources. It relies only on the statistical properties of the existing text corpus (like Wikipedia).
- Weighted Cross-Entropy Loss
- A modification to the standard training loss function that penalizes mistakes differently. By assigning higher weights to informative tokens, the model is forced to pay more attention and improve its accuracy on crucial factual words.
Terminology
Summary
Summary
This paper proposes a novel, fully unsupervised pretraining strategy for masked language models (MLMs) that improves their performance on knowledge-intensive tasks without relying on external knowledge bases. The authors argue that the traditional Masked Language Modeling (MLM) objective is sub-optimal for learning factual knowledge because it is often dominated by high-frequency words. To address this, they introduce two key modifications to the MLM objective: (1) masking tokens with higher informative relevance
more frequently, and (2) penalizing mistakes on these informative tokens more severely during loss computation.
The informative relevance of a token is computed in a completely unsupervised manner using Pointwise Mutual Information (PMI). Specifically, the authors compute word co-occurrence statistics within a skip-gram window (size 10) on the pretraining corpus. They then build a PMI matrix for all word pairs in the vocabulary (100k × 100k). For each document, they construct a pairwise PMI matrix between all words in that document and compute the row-wise sum, which reflects the token-specific informative relevance within that document. These values are averaged across the corpus, then normalized and converted into token-specific masking rates (ranging from 15% to 50%, with an average of 19%) and token-specific penalty weights (normalized within the range [1, 5]).
The proposed loss function is a weighted cross-entropy loss: L MLM = -sum i=1 N w y i e x i,y i over sum v in V e x i,v, where w y i is a penalty weight specific to the output token y i. The authors train four models on the Wikipedia corpus (Hugging Face) with a wordpiece tokenizer (vocab size 100k) and a BERT-base architecture (12 layers, hidden dimension 768). The models are: (a) BERTuu (uniform masking, uniform penalty – baseline), (b) BERTuw (uniform masking, weighted penalty), (c) BERTvu (variable masking, uniform penalty), and (d) BERTvw (variable masking, weighted penalty – the proposed approach). Training uses a batch size of 128, learning rate 5e-5, AdamW optimizer, 10 epochs, and max document length of 128. The increased masking rate and penalty weight apply only to whole-word tokens; subword tokens use the minimum masking rate (15%) and penalty weight (1).
The proposed model (BERTvw) significantly outperforms the baseline (BERTuu) on factual recall tasks in the LAMA benchmark. The relative improvement in Mean Reciprocal Rank (MRR) over the baseline is 17.5% for ConceptNet, 6% for GoogleRE, and 8.1% for TREx. On the SQuAD portion of LAMA (zero-shot closed-book QA), the proposed model achieves a 19.9% relative improvement over the baseline. Case studies show that the proposed model is more likely to rank the ground truth label higher and produces more specific words given a particular context (e.g., predicting Nintendo
instead of generic words like computer
for a prompt about a gaming company).
For closed-book sentiment analysis and natural language inference (NLI) using AutoPrompt, the proposed system achieves 8.1%, 21.1%, and 14.7% relative improvement in accuracy over the baseline for sentiment analysis (SST2), 3-way NLI, and 2-way NLI, respectively. When fine-tuned on SQuAD v1 and v2, the proposed model outperforms the baseline on both tasks (F1 scores of 72.61 vs. 69.96 for v1, and 85.28 vs. 83.22 for v2). On the GLUE benchmark, it outperforms the baseline on seven out of nine tasks, though the relative improvement is less significant than in zero-shot or prompt-tuning scenarios. The authors attribute this to findings from prior work (Wallat et al., 2020) that factual knowledge learned during pretraining may be lost during fine-tuning.
An ablation study comparing BERTuw and BERTvu shows that a variable masking rate performs slightly better than a weighted penalty in most cases. The authors conclude that their proposed pretraining strategy is effective in storing factual knowledge within language models, leading to better performance on knowledge-intensive tasks. They also note a limitation: the proposed training objective reduces the importance of stopwords, which may negatively impact performance on syntax-sensitive tasks like CoLA. Future work aims to extend the approach to text-to-text models such as T5.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, and the resulting capabilities:
Improvements to the Pretraining Objective:
-
Implement a Variable Masking Rate (VMR): Instead of masking 15% of tokens uniformly, I will compute a token-specific masking probability based on Pointwise Mutual Information (PMI) with neighboring words. Tokens with high average PMI (e.g., named entities, rare technical terms) will be masked at a rate up to 50%, while common stopwords will be masked at the minimum 15%. This forces the model to learn stronger contextual representations for information-dense tokens.
-
Implement a Weighted Cross-Entropy Loss (WCE): I will modify the MLM loss function to apply a token-specific penalty weight (normalized between 1 and 5) to the loss contribution of each masked token. Mistakes on high-PMI tokens (e.g.,
Nintendo
in a gaming context) will be penalized 5x more severely than mistakes on low-PMI tokens (e.g.,the
). This directly biases gradient updates toward encoding factual associations. -
Combine VMR and WCE into a Unified Objective: I will integrate both mechanisms into a single training loop. The masking rate determines which tokens are hidden, and the penalty weight determines how much the model is punished for getting them wrong. This dual-pressure approach ensures the model allocates more capacity to factual recall.
-
Compute PMI Statistics in a One-Pass Preprocessing Step: I will build a 100k x 100k co-occurrence matrix using a skip-gram window of 10 on the pretraining corpus. From this, I will derive a per-token
informative relevance
score by summing the row-wise PMI values within each document and averaging across the corpus. This requires 11GB of RAM and 2 hours of compute, which is a one-time cost.
What the Improved AI System Can Do:
-
Achieve Higher Factual Recall Accuracy: The system will demonstrate a 17.5% relative improvement in Mean Reciprocal Rank (MRR) on the LAMA benchmark (ConceptNet) compared to a standard BERT baseline. For example, given the prompt
Photosynthesis releases [MASK] into the Earth's atmosphere,
the improved model will rankoxygen
first with a score of 0.21, whereas the baseline only ranks it third with a score of 0.09. -
Produce More Specific and Correct Answers in Zero-Shot QA: When probed with
During Super Bowl 50 the [MASK] gaming company debuted their ad for the first time,
the improved model will generate specific candidates likeNintendo,
Walt,
andAtari
(withNintendo
ranked first), whereas the baseline only produces generic words likecomputer
andelectronic.
-
Improve Closed-Book Prompt-Based Performance: The system will achieve 8.1% relative improvement in accuracy on SST2 sentiment analysis, 21.1% on 3-way NLI, and 14.7% on 2-way NLI using AutoPrompt templates, without any fine-tuning. This is because the model stores more relational knowledge during pretraining, which can be elicited via trigger tokens.
-
Enhance Fine-Tuned Performance on Extractive QA: When fine-tuned on SQuAD v1 and v2, the system will achieve 72.61 F1 (vs. 69.96 baseline) and 85.28 F1 (vs. 83.22 baseline), respectively. The improvement is modest but consistent, indicating better initialization for downstream reading comprehension.
-
Maintain or Improve Performance on Most GLUE Tasks: The system will outperform the baseline on 7 out of 9 GLUE tasks (e.g., 89.91 accuracy on SST2, 88.49 on QNLI, 56.34 on WNLI), while showing a slight regression on syntax-heavy tasks like CoLA (28.93 vs. 31.06), which is an acceptable trade-off for knowledge-intensive applications.
Specific Technical Implementation Details:
-
Vocabulary Size: 100k wordpiece tokens to ensure broad entity coverage.
-
Masking Range: 15% (minimum) to 50% (maximum) based on PMI score.
-
Penalty Weight Range: 1.0 (minimum) to 5.0 (maximum), normalized.
-
Architecture: BERT-base (12 layers, 768 hidden dim), trained for 10 epochs on Wikipedia with a batch size of 128, learning rate 5e-5, and AdamW optimizer.
-
Subword Handling: Whole-word tokens receive the variable masking/weighting; subword tokens default to 15% masking and weight 1.0 to avoid fragmentation issues.
Caveat to Monitor: The system may underperform on tasks requiring strong syntactic understanding (e.g., CoLA) due to reduced emphasis on stopwords. If this is critical, I will add a small regularization term to the loss to preserve some weight on low-PMI tokens, or use a softer penalty range (e.g., 1 to 3) to balance factual recall with grammaticality.
Sources
- GPT Understands, Too
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- ERNIE: Enhanced Representation through Knowledge Integration
- Modifying Memories in Transformer Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering