Unsupervised Improvement of Factual Knowledge in Language Models

summary

Video file (mp4)

In short

The episode discusses 'Unsupervised Improvement of Factual Knowledge in Language Models,' which proposes a smarter way to train models using masked language modeling (MLM). The method uses Pointwise Mutual Information (PMI) to identify factually informative tokens, adjusting both masking rates and loss weights to improve factual recall without external knowledge bases.

Key concepts

Masked Language Modeling (MLM)
A training technique where the model learns by predicting masked or hidden words within a sentence. The paper improves this standard method by making the model focus more heavily on factually important tokens like names and dates.
Pointwise Mutual Information (PMI)
A statistical measure used to determine how frequently two words appear together in a corpus. High PMI scores indicate that a word is highly informative because it frequently co-occurs with specific other words, rather than being general stopwords.
Unsupervised Improvement
The method improves model performance without requiring human labeling, external knowledge graphs, or expensive hand-crafted resources. It relies only on the statistical properties of the existing text corpus (like Wikipedia).
Weighted Cross-Entropy Loss
A modification to the standard training loss function that penalizes mistakes differently. By assigning higher weights to informative tokens, the model is forced to pay more attention and improve its accuracy on crucial factual words.

Terminology used across episodes

This episode discusses

The paper

Unsupervised Improvement of Factual Knowledge in Language Models · Read on arXiv

Nafis Sadeq, Byungkyu Kang, Prarit Lamba, Julian McAuley

Intuit · University of California, San Diego

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Unsupervised Improvement of Factual Knowledge in Language Models".

Jane: The paper was written by Nafis Sadeq, Byungkyu Kang, Prarit Lamba and Julian McAuley from Intuit and University of California, San Diego.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Alright, welcome back to the show, everybody. Today we’re digging into a paper that’s got a wonderfully straightforward title: “Unsupervised Improvement of Factual Knowledge in Language Models.” I’m Tom, and as always, I’ve got Jane here with me.

Jane: Hi, folks! And Tom, I gotta say, just reading that title makes me happy. We spend so much time hearing about bigger models, more data, fancier architectures. This one is about tweaking how we train the models we already have, and it’s all done without any external knowledge bases. That’s a big deal.

Tom: Right, and the team behind it is from UC San Diego and Intuit—Nafis Sadeq, Byungkyu Kang, Prarit Lamba, and Julian McAuley. They’re essentially asking a simple question: when a language model is learning to predict masked words, can we make it focus on the words that actually carry factual weight, like names, dates, and places?

Jane: And their answer is yes, and they do it in a totally unsupervised way. No knowledge graphs, no entity embeddings, no expensive hand-crafted resources. Just a clever reweighting of the pretraining objective itself.

Tom: So for anyone tuning in late, the core idea is that traditional masked language modeling treats every word equally. But in a sentence like “Marie Curie was born in MASK,” the word “Warsaw” is way more informative than the word “was.” The model should be penalized more for getting “Warsaw” wrong than for missing a stopword.

Jane: Exactly. And the authors show that by making the model prioritize those informative tokens, it gets significantly better at recalling facts later on. We’re talking about a seventeen point five percent relative improvement in Mean Reciprocal Rank on the ConceptNet portion of the LAMA benchmark, which is a solid jump.

Tom: And that’s just the start. They also see gains on question answering, sentiment analysis, and natural language inference, all in a closed-book setting. So the model isn’t looking anything up; it’s just using what it stored during pretraining.

Jane: I love that this is a pretraining-only change. You don’t need to change your model architecture or your downstream fine-tuning. You just train it smarter from the get-go.

Tom: And that’s the hook for today’s episode. We’re going to break down exactly how they compute this “informative relevance” and why it works so well. Stick around.

Summary: Tom: So, Jane, we’ve set the stage. Let’s get into the meat of the paper. The authors propose two key changes to the standard masked language modeling recipe.

Jane: Right. First, instead of masking fifteen percent of tokens uniformly, they mask tokens with a variable rate. Tokens that are more “informative” get masked more often, up to fifty percent of the time, while the average stays around nineteen percent. The second change is that they use a weighted cross-entropy loss, so mistakes on informative tokens are penalized more heavily, with weights normalized between one and five.

Tom: And how do they figure out which tokens are informative? That’s the clever part.

Jane: They use Pointwise Mutual Information, or PMI. Basically, they look at word co-occurrence statistics across the whole Wikipedia corpus. If a word frequently appears near specific other words, it gets a high PMI score. For example, “Nintendo” probably has high PMI with “console” or “Mario,” whereas a word like “the” has low PMI with everything because it’s everywhere.

Tom: So they build a giant matrix of PMI values between all pairs of words in the vocabulary—one hundred thousand by one hundred thousand—and then for each document, they sum up the PMI scores for each token against all the other tokens in that document. That gives a per-token informative relevance score.

Jane: Exactly. And that one-pass computation takes about two hours and eleven gigabytes of memory. It’s not trivial, but it’s a one-time cost before pretraining, and it’s completely unsupervised. No human labels, no knowledge bases.

Tom: So they train four models to isolate the effects: a baseline with uniform masking and uniform penalty, one with just weighted penalty, one with just variable masking, and then their full model with both. The full model, which they call BERTvw, consistently wins.

Jane: And the ablation study shows that variable masking helps a bit more than the weighted penalty, but combining them gives the best results. It’s a nice, clean experimental design.

Tom: I also appreciate that they didn’t just test on LAMA. They fine-tuned on SQuAD and GLUE, and they used AutoPrompt for zero-shot sentiment and NLI. The gains on fine-tuning are smaller, but they’re still positive on most tasks.

Jane: Right, and that’s actually a really interesting finding. The authors point out that fine-tuning can wash out some of the factual knowledge gained during pretraining, which is something other researchers have observed too. But in prompt-based settings, the knowledge shines through.

Tom: So the summary is: a smarter masking strategy, computed for free from the corpus, leads to measurably better factual recall and downstream performance. That’s a strong result.

Jane: And it opens the door for a lot of follow-up work. Which is exactly what we’re going to talk about next.

Improvements: Tom: Welcome back. We’ve covered the basics, but now I want to bring in our senior researcher, Lu, and our engineer, Meng, because this paper has implications that go beyond just the numbers.

Lu: Thanks, Tom. I’ve been thinking about this all morning. The key improvement here isn’t just the performance boost—it’s the philosophy. The authors are saying that not all tokens are created equal, and the training objective should reflect that. This is a step toward more efficient learning, where the model spends its capacity on what matters.

Jane: And Lu, you’re right. It’s like studying for an exam. If you spend equal time on every page of the textbook, you’ll be prepared for the easy questions but might miss the key facts that show up on the final. This paper is essentially telling the model to highlight the important passages.

Meng: From an engineering standpoint, I’m curious about the practical side. The PMI computation is a one-time cost, but it’s still a 100k by 100k matrix. For a smaller team or a niche domain, is that feasible?

Tom: That’s a great question, Meng. The authors mention it takes about two hours and eleven GB of RAM. That’s not nothing, but it’s also not prohibitive. And you could imagine using a smaller vocabulary or a subsampled corpus to bring that cost down.

Meng: Right, and the other thing is that this is a drop-in replacement for the MLM objective. You don’t need to change the model architecture or the training pipeline. You just swap in the new masking rates and loss weights. That’s a huge win for adoption.

Lu: And the implications go further. Think about multilingual models. Knowledge bases are expensive to build for low-resource languages, but this method only needs text. So you could apply the same trick to a Swahili or Hindi corpus and get better factual recall without any external resources.

Jane: That’s a really exciting point, Lu. The authors explicitly mention that existing knowledge bases may not be available for all languages and domains. This method sidesteps that entirely.

Meng: I also want to highlight the case studies in the paper. They show that the baseline model often produces generic words like “computer” or “american” for a prompt about a gaming company, while their model produces “nintendo,” “walt,” and “atari.” That’s a qualitative improvement in specificity, not just a metric bump.

Tom: And that specificity is what makes the model useful in real-world applications. If you’re building a question-answering system, you want it to say “Nintendo,” not “computer.”

Lu: Exactly. And I think the next step is to apply this to generative models like T5. The authors mention that as future work, and I bet we’ll see results there within the year.

Jane: So the improvements are clear: better performance, no external resources, and a clear path to broader applications. But what does this mean for the field as a whole? That’s our final segment.

Conclusion: Tom: Alright, we’re in the home stretch. Let’s wrap up our discussion of “Unsupervised Improvement of Factual Knowledge in Language Models.”

Jane: So, to recap: the paper introduces a way to make masked language modeling smarter by focusing on informative tokens. They compute token importance using PMI, then use that to drive both masking rates and loss weights. The result is a model that stores more factual knowledge and performs better on a range of tasks.

Tom: And the beauty is that it’s fully unsupervised. No knowledge graphs, no entity embeddings, no human annotations. Just a clever reweighting of the training objective.

Lu: I’d add that this is a reminder that pretraining objectives are still an underexplored frontier. We’ve been using the same MLM recipe for years, and this paper shows that a thoughtful tweak can yield significant gains.

Meng: From a practical standpoint, it’s a low-risk, high-reward change. Any team training a BERT-style model can adopt this with minimal effort. And the code is public, so you can try it yourself.

Jane: And the implications for the world? More efficient language models that can answer factual questions without needing to look things up. That’s valuable for everything from search engines to virtual assistants to educational tools.

Tom: There are limitations, of course. The model underperforms on some syntax-heavy tasks like CoLA, and there’s a risk of amplifying biases if the informative tokens are biased. But these are solvable problems.

Lu: I’m excited to see where this goes. If the authors extend this to generative models, we could see even bigger gains in text-to-text tasks.

Tom: Well said, Lu. So that’s a wrap on “Unsupervised Improvement of Factual Knowledge in Language Models.” Thanks to everyone who tuned in. We’ll be back next time with another paper to dissect.

Jane: Until then, keep asking questions and keep learning. Goodbye, everyone!

More episodes

← Home