Thinking Outside the (Gray) Box: A Context-Based Score for Assessing Value and Originality in Neural Text Generation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Thinking Outside the (Gray) Box: A Context-Based Score for Assessing Value and Originality in Neural Text Generation".
Jane: The paper was written by Giorgio Franceschelli and Mirco Musolesi from Alma Mater Studiorum Università di Bologna and University College London.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the show, everyone. Today we’re looking at a paper that’s got a wonderfully playful title: “Thinking Outside the (Gray) Box: A Context-Based Score for Assessing Value and Originality in Neural Text Generation.” Jane, that title alone tells me these authors are trying to shake things up.
Jane: Oh, absolutely, Tom. And I love the little pun in there. We usually talk about black-box models, right? Things we can’t see inside. But these authors, Giorgio Franceschelli and Mirco Musolesi from Bologna and UCL, they’re saying, "Hey, let’s not just look inside the box, let’s rethink what we’re even asking the box to do."
Tom: Right, and they’re not just asking it to be smart or accurate. They want it to be creative. They want it to be valuable *and* original. And that’s a tough combination, because usually, if you push a model to be more original, it starts making mistakes, and if you push it to be accurate, it gets boring and repetitive.
Jane: Exactly. And that’s the core tension they’re tackling. They’ve come up with a score, called CoVO, which stands for Context-based Value and Originality. And the clever part is how they define those two things. They’re using information theory, which is basically the math of how much we learn from a message.
Tom: So, for the listeners at home, can you break that down a little more? What does "value" mean in this math, and what does "originality" mean?
Jane: Sure. Think of it like this: you ask the model to write a poem about a rainy day. The "value" part is checking if the model’s poem actually sounds like a poem about a rainy day. They do this by flipping the question around. They ask the model, "Given this poem, can you guess what the original request was?" If the model can easily guess it was about a rainy day, then the poem is valuable.
Tom: That’s a neat trick. And the originality part?
Jane: The originality part is the opposite. They look at how *surprised* the model is by its own poem. If the poem is full of words and phrases the model would normally never use for a rainy day, then it’s original. It’s unexpected. So, they’re trying to maximize the chance that the poem is clearly about rain, while also maximizing how unlikely that specific poem was in the first place.
Tom: So it’s a balancing act. It’s like a tightrope walk between "that makes perfect sense" and "wow, I never would have thought of that." And they’re using this score to actually train the model, not just evaluate it.
Jane: Precisely. And that’s what we’re going to dig into next. They use this score as a reward in a reinforcement learning setup, which is a fancy way of saying they’re teaching the model to chase this specific goal. We’ll see how they did it and what happened when they let the model loose on poetry and math problems.
Tom: I’m already hooked. So they’re not just describing a problem, they’ve built a solution and tested it. Let’s get into the meat of the paper in the next segment.
Summary of the Paper: Tom: So, Jane, we’ve got the title and the basic idea. Now, for the summary. The paper “Thinking Outside the (Gray) Box” doesn’t just stop at defining this CoVO score. They actually use it to fine-tune large language models using a reinforcement learning algorithm called GRPO.
Jane: Right. And for our listeners, GRPO is a way to train a model by giving it a reward for good behavior. In this case, the reward is the CoVO score. They run the model on a prompt, it generates a few different outputs, and then they score each one. The model then learns to produce more of the outputs that score high.
Lu: And what I find really interesting, Tom, is that they’re not just using this for one task. They tested it on three very different things. They tried poetry generation, they tried solving math word problems, and they even used a benchmark called NoveltyBench, which is specifically designed to test how diverse and novel a model’s responses are.
Tom: Lu, that’s a great point. It’s not just a one-trick pony. And the results seem to back that up. In the poetry task, they found that their method led to poems that were more original, meaning they were less likely to accidentally copy famous works, while also adhering to the requested tone better.
Meng: But I’m curious about the practical side, Lu. You mentioned math problems. In math, you can’t just be original; the answer has to be right. Did they see a trade-off there? Did the model get more creative but start giving wrong answers?
Jane: That’s the million-dollar question, Meng. And the paper shows it’s a real balancing act. When they used only the CoVO score, the model’s accuracy on the math tests did drop a bit. But when they combined the CoVO score with a simple reward for getting the correct answer, they actually got the best of both worlds. The accuracy stayed high, and the solutions were more diverse than the baseline model.
Lu: And that diversity is key. It means the model is finding different ways to solve the same problem, which is a sign of deeper understanding, not just memorization. It’s not just regurgitating the most common solution from its training data.
Tom: So it’s like they’ve found a dial. You can turn up the "originality" knob, but you have to be careful not to break the "value" knob. And they’re showing you how to tune it for different tasks.
Meng: So the math results are promising, but I’m still wondering about the engineering side. How heavy is this to implement? You need a second model to compute the score, right? That sounds like it could double the compute cost.
Jane: That’s a really good practical question, Meng. And we’ll get into the details of how they actually implemented it in the next segment, because they do have a clever way of handling that. But for now, let’s just say the potential payoff in terms of creativity seems worth the extra cost.
Tom: And that’s the hook. Let’s take a closer look at the "how" in the next segment. We’ll talk about the improvements and the clever engineering that makes this work.
Improvements Suggested by the Paper: Tom: Alright, we’re back. So we’ve established that “Thinking Outside the (Gray) Box” has this score and it works on different tasks. Now, let’s talk about the clever improvements and the practical details. Jane, you mentioned the second model earlier. How do they actually compute this score?
Jane: So, they use a frozen reference model, which is just the original, pre-trained model that they started with. This model is used to compute both parts of the score. For the "value" part, they do something really clever. They take the generated poem, and they append a little question to it, like "What is the style of this text?" Then they ask the model to generate the original prompt.
Meng: So they’re using the model itself to check if the output is on-topic. That’s a neat trick to avoid needing a separate, specially-trained classifier. It’s like using the model as its own judge.
Jane: Exactly. And for the "originality" part, they just look at how surprised the model is by the generated text given the prompt. That’s the surprisal we talked about. The less likely the text is, the more original it’s considered.
Lu: And this is where the improvement gets interesting. They don’t just use this score in a vacuum. They plug it into GRPO, and they also have the option to add a KL divergence term. For the listeners, that’s a way to keep the new, fine-tuned model from drifting too far away from the original, sensible model.
Tom: So it’s a safety rail. It keeps the model from going completely off the rails in its quest to be original.
Lu: Precisely. And their experiments show this trade-off beautifully. Without the KL term, the model is more original but sometimes breaks the rules, like writing a "sonnet" that doesn’t have the right number of lines. With the KL term, the model is more likely to follow the rules, but it’s a bit less surprising.
Meng: So, for a practical engineer like me, this means you have two knobs to turn. You have the CoVO score itself, which balances value and originality, and then you have this KL knob, which controls how far you’re willing to let the model wander. That’s a really flexible framework.
Jane: And it’s not just for poetry. In the math experiments, they showed that adding the CoVO score on top of the standard correctness reward led to more diverse problem-solving strategies. The model was finding new ways to get to the right answer, which is a huge deal for educational tools or for generating synthetic training data.
Tom: So it’s not just about making a model that writes pretty poems. It’s about making models that can think more flexibly across the board. And they even tested it on NoveltyBench, which is a benchmark that directly measures if a model can produce ten different, high-quality answers to the same prompt. And their method improved both the novelty and the quality scores there.
Lu: And that’s the real promise here. It’s a general-purpose tool for injecting creativity into any text generation task, without sacrificing quality. It’s a way to push back against the tendency of these models to just give you the most average, boring answer.
Tom: I love that. So we’ve got the theory, the implementation, and the results. Let’s wrap this up in our final segment and talk about what this means for the world.
Conclusion: Tom: And we’re back for the final word on “Thinking Outside the (Gray) Box: A Context-Based Score for Assessing Value and Originality in Neural Text Generation.” Jane, give us the big picture.
Jane: The big picture, Tom, is that this paper gives us a concrete, mathematical way to define and optimize for creativity. It’s not just about making models that are "less boring." It’s about giving them a goal that explicitly rewards both being useful and being surprising. And they’ve shown it works across poetry, math, and general open-ended tasks.
Lu: And I think the impact goes beyond just making chatbots more fun. Think about scientific discovery. If you have a model that’s trying to generate a new hypothesis, you don’t want the most obvious one. You want one that is plausible but also unexpected. This CoVO score is a direct way to reward that kind of thinking.
Meng: From a practical standpoint, I’m excited about the flexibility. The fact that you can combine this score with other rewards, like a correctness check in math, means it can be dropped into existing training pipelines without a complete overhaul. It’s a modular component for creativity.
Lalam: And from a cultural perspective, this is about enriching our creative landscape. These models are becoming tools for writing, for composing, for brainstorming. If we can give them a nudge toward genuine originality, they can help us explore ideas we might never have considered ourselves. They become partners in creativity, not just autocomplete machines.
Tom: That’s a beautiful way to put it, Lalam. So, to sum it up: the paper introduces a score, CoVO, that balances value and originality. They use it to fine-tune models with reinforcement learning. And they show it makes models better at being creative without losing their usefulness.
Jane: And while there are still open questions, like how to perfectly balance those two forces or how well it aligns with human taste, the framework is solid. It’s a significant step forward in our quest to build genuinely creative AI.
Tom: Well said. We’ve had a great time with this one. A big thank you to our listeners for joining us. We’ll be back soon with another paper from the arXiv. Until then, keep thinking outside the box. Goodbye, everyone.
Giorgio Franceschelli, Mirco Musolesi
Alma Mater Studiorum Università di Bologna · University College London
cs.CL, cs.AI, cs.CY, cs.LG
Submitted: 2026-08-17
Updated: 2026-08-18
Code: https://github.com/aparrish/gutenberg-dammit
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 66/100
The gist: Despite the increasing use of large language models for creative tasks, their outputs often lack diversity.
Key concepts
- CoVO Score
- Context-based Value and Originality (CoVO) is a metric designed to reward AI outputs that are both useful (valuable) and surprising (original). It helps balance the tension between accuracy and creativity in generated text.
- Value
- In this context, value is measured by checking if the model's output aligns with the original request. This is done by asking the model to guess the initial prompt based on its generated text.
- Originality
- Originality is assessed by looking at how 'surprised' the model is by its own output. A less likely, more unexpected piece of text scores higher for originality.
- Reinforcement Learning (GRPO)
- The CoVO score is used as a reward in a reinforcement learning setup (specifically GRPO). This trains the model to consistently produce outputs that maximize both value and originality.
Terminology
Summary
Despite the increasing use of large language models for creative tasks, their outputs often lack diversity. Common solutions, such as sampling at higher temperatures, can compromise the quality of the results. Dealing with this trade-off is still an open challenge in designing AI systems for creativity. Drawing on information theory, we propose a context-based score to quantitatively evaluate value and originality. This score incentivizes accuracy and adherence to the request while fostering divergence from the learned distribution. We show that our score can be used as a reward in a reinforcement learning framework to fine-tune large language models for maximum performance. We validate our strategy through experiments considering a variety of creative tasks, such as poetry generation and math problem solving, demonstrating that it enhances the value and originality of the generated solutions.
The nature of the self-supervised learning algorithms used for the training of these models tends to make their sampling distribution as close as possible to the training data distribution. In addition, fine-tuning, such as that based on reinforcement learning from human feedback, is often necessary to generate appropriate and accurate responses. However, this process tends to reduce output diversity further, and linguistic creativity tends to be lower than that of humans. On the contrary, LLMs for creative tasks should produce more novel and surprising texts that maintain a high level of correctness and adherence to the request. One typical solution is to sample at a higher temperature to increase diversity. However, this might lead to generating less coherent text.
To address the issues described above, we propose a new training approach for creative tasks based on CoVO, a Context-based score for Value and Originality, with the goal of taking into consideration both value and originality of the neurally-generated text in the optimization of LLMs. The definition of CoVO is grounded in the analysis of mutual information between the model’s outputs and inputs, and vice versa. More specifically, we formulate a new optimization objective where, given a specific input, the desired output is derived by simultaneously maximizing the conditional probability of the input given the output and minimizing the conditional probability of the output given the input under the generative model. In this way, we optimize for solutions that are appropriate for the input request but also different from the outputs we would normally obtain from the model. In particular, we show that our information-theoretic score can be used as a reward in RL-based fine-tuning algorithms, guiding pre-trained models toward more diverse yet valuable solutions.
Our key contributions are the following: We present the theoretical foundations of our approach, deriving our context-based score for value and originality from the concept of mutual information. We discuss how the score can be practically computed in the case of autoregressive models, and how it can be used as a reward with Group Relative Policy Optimization, a state-of-the-art reinforcement learning algorithm for fine-tuning LLMs. We evaluate our GRPO-based method on mathematical problem solving, poetry generation, and the tasks included in NoveltyBench, demonstrating that our approach can enhance both the quality and diversity of generated outputs, positioning it as a strong candidate for creativity-focused applications of current foundation models.
The CoVO score for a target y given a source x on a reference probability distribution p is formally defined as: sCoVO(x, y, p) = λv log p(xy) − λo log p(yx), where the first term represents Value and the second term represents Originality. Solving this maximization problem involves finding the target y that maximizes the posterior probability of x while also being unlikely given x. In other words, the optimal y∗ must be unexpected and diverse from p(yx), but it must also be explainable by x. − log p(yx), commonly known as surprisal, is widely used to measure diversity and surprise, and adheres to the first requirement from the standard definition of creativity, i.e., originality. Conversely, log p(xy) can be used to measure value or effectiveness, the second requirement of the definition. If the request can be inferred from the outcome, the latter constitutes an appropriate instance of that task or a correct, useful solution to that problem.
For autoregressive models, the score is implemented as: sAR CoVO(x, y, pθ) = λv sv(x, y, pθ) + λo so(x, y, pθ), where the value component is the normalized average log probability of the source given the target, and the originality component is the normalized average log probability of the target given the source. Calculating pθ(xy) is not trivial, so we consider an approximation pθ(xy′), where y′ = y + q, with q representing an additional question designed solely to increase the likelihood of generating the source text x.
Once the CoVO score has been defined, its adoption in an RL framework is straightforward. We can directly utilize our CoVO score as the final reward for the generated sequence. Then, the model can be trained with any policy gradient method. Our experiments leverage GRPO, which is a state-of-the-art choice for training language models. GRPO adds a per-token KL divergence term to the loss rather than to the single rewards. In particular, GRPO aims to minimize the KL divergence, thus the second term − log rref = − log πref + log πθ can be seen as made of two optimization problems: the minimization of the originality component − log πref from Equation 9, and the maximization of surprisal, or self-information, − log πθ under the current model.
We evaluate the effectiveness of our RL strategy through three case studies: poetry generation, mathematical problem resolution, and the tasks included in NoveltyBench. In all experiments, we employ two settings, i.e., GRPO to maximize the score from Equation 11 without the KL divergence loss, i.e., with β = 0.0 (CoVO); and GRPO to maximize the score from Equation 11 with the KL divergence loss, i.e., with β = 0.05 (CoVO + KL). Both methods assume λv = λo = 1.0.
For poetry generation, we consider the Meta-Llama-3-8B model as our pre-trained agent and use Low-Rank Adaptation instead of fine-tuning the entire network. We perform a quantitative evaluation where we compute poetical metrics for quality (lexical correctness of poems, adherence to line- and syllable-level constraints, and tone adherence to the request through zero-shot classification) and for originality (accidental reproduction of existing poems). Overall, our CoVO-based fine-tuning leads to a higher tone adherence and lower reproduction rate, at the potential cost of metric adherence, especially without the KL loss. Indeed, its role seems to foster quality, trading off some originality. On the contrary, not using the KL loss arguably avoids any significant reproduction, as demonstrated by the very low maximum token-based longest common substring.
For math problem resolution, we focus on the Mistral-based MetaMath-Mistral-7B model, fine-tuned with self-supervised learning on the MetaMathQA dataset. The RL problem is formulated considering up to two rewards: our CoVO score computed on the procedure and a verifiable, extrinsic reward based on the correctness of the answer. The evaluation considers both GSM8K and MATH test sets. For the GSM8K test set, while all methods achieve similar results, using the CoVO score only leads to greater EAD diversity and to diverge more from the original model, while the presence of the math reward leads to greater accuracy, especially without KL. The results for the MATH test confirm that the most accurate method is the one trained to optimize the CoVO score and the extrinsic reward, with a negligible trade-off in terms of diversity, since the cross-input EAD and the against-pretrained scores are still substantially higher than those from the baselines. However, removing the extrinsic reward likely pushes the model too far from its pre-trained version, causing the accuracy to decrease.
For NoveltyBench, we fine-tune the Llama-3.2-3B-Instruct model for a single epoch on the ‘wildchat’ partition, then compute the novelty and quality scores on the ‘curated’ partition. Optimizing for the CoVO score results in substantial improvements in both novelty and quality metrics, with the greatest gains in novelty. Moreover, these results underscore the interplay between KL loss and our CoVO score: incorporating the KL penalty tends to improve quality, but at the cost of reduced novelty.
The CoVO score captures properties that are functionally aligned with creativity-relevant aspects of language generation, as demonstrated through both conceptual analysis and empirical results. When used as a reward function in Group Relative Policy Optimization, CoVO drives improvements in the novelty and quality of model outputs, even with a relatively small number of optimization steps. Although KL-divergence regularization is typically employed to constrain policy shifts and preserve alignment with the base model distribution, CoVO can contribute independently to several desirable behaviors: reducing the risk of inadvertent memorization of copyrighted material, promoting diversity in generated outputs, and mitigating undesirable inductive biases introduced during pretraining.
There are a few limitations worth noting. Firstly, our score represents only a quantifiable approximation of a particular theoretical perspective on creativity, grounded in the dimensions of value and originality. For example, value has been considered from the perspective of effectiveness, while other dimensions have been proposed as well; and other classic definitions of creativity add a third requirement by splitting originality into novelty and surprise. Moreover, our score reflects a specific view of the evaluation of creativity based on the generated outputs and does not account for potential alternative theories and perspectives. Finally, our experiments are currently limited to only three relatively short-form text generation tasks. While their generalizability is supported by the theoretical framework discussed above, the resulting performance was experimentally evaluated for a finite number of scenarios.
In conclusion, we presented CoVO, a novel score that quantifies the value and originality of neurally-generated text. The definition of CoVO is based on the analysis of mutual information between the model’s outputs and inputs, and vice versa. We also proposed an optimization problem where a generative model aims to maximize this score to generate more creative products, and detailed how to use it in language modeling. We conducted experiments on poetry generation, math problem solving, and tasks included in NoveltyBench, exploring trade-offs in accuracy vs diversity. Effectively balancing value and originality maximization remains an open question, but our score seems to relate to domain-specific measures appropriately. In addition, fine-tuning to maximize it improves quality- and diversity-related metrics. Our research agenda aims to extend our method to other models and tasks, to include inference-level strategies such as creativity-oriented sampling schemes, and to explore its use for evaluation rather than solely for optimization. We also plan to investigate the definition of additional scores for capturing other potentially relevant aspects of the creative process. Despite being costly and inherently constrained, assessing whether our creativity score aligns with human judgment is another key direction for future work.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems:
1. Implement a CoVO-based reward function for RL fine-tuning
-
Add a dual-term reward:
λv * log p(xy) - λo * log p(yx)wherexis the prompt andyis the generated output -
Use the normalized, length-adjusted version from Equation 11 to avoid sequence-length bias
-
Integrate this reward into GRPO (Group Relative Policy Optimization) instead of standard PPO, eliminating the need for a value network
2. Add a frozen reference model for score computation
-
Keep a frozen copy of the pre-trained model to compute both
p(yx)(originality) andp(xy)(value) -
For
p(xy), append a question likeHow would you describe this text?
to the output before computing the reverse probability, as this improves likelihood estimation -
Use the max-token normalization trick (
max v∈V log p) to balance the magnitude between value and originality terms
3. Enable dual-mode training with adjustable KL regularization
-
Offer two configurations:
CoVO(β=0.0, maximizing originality) andCoVO + KL(β=0.05, balancing value and originality) -
This lets users choose between higher diversity (no KL) or higher quality/coherence (with KL)
4. Apply to specific creative tasks
-
Poetry generation: Fine-tune a base LLM (e.g., Meta-Llama-3-8B) with LoRA to generate poems with higher tone adherence and lower accidental reproduction of existing works (reducing T-LCS scores)
-
Math problem solving: Fine-tune a math-specialized model (e.g., MetaMath-Mistral-7B) to produce more diverse solution procedures while maintaining accuracy, using CoVO as a secondary reward alongside answer correctness
-
Open-ended generation: Fine-tune an instruction-tuned model (e.g., Llama-3.2-3B-Instruct) on a diverse prompt set to improve both novelty (more distinct outputs per prompt) and quality (higher utility scores)
5. Add diversity metrics for evaluation
-
Use expectation-adjusted distinct N-grams (EAD) and sentence embedding cosine similarity (SBERT) to measure syntactic and semantic diversity
-
Compute both cross-input diversity (across outputs for the same prompt) and against-pretrained diversity (comparing to the original model's output)
What the improved AI system can do:
-
Generate poems that are more original (lower reproduction of existing works) while maintaining or improving tone adherence
-
Solve math problems with more diverse solution strategies, reducing repetitive reasoning patterns
-
Produce multiple distinct, high-quality responses to the same prompt (improving NoveltyBench scores)
-
Avoid memorization of copyrighted or existing text, reducing legal and ethical risks
-
Allow fine-grained control over the creativity-quality trade-off via the KL regularization parameter
Sources
- Emergent autonomous scientific research capabilities of large language models
- On the Opportunities and Risks of Foundation Models
- Training Verifiers to Solve Math Word Problems
- CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
- Gemini: A Family of Highly Capable Multimodal Models
- The Llama 3 Herd of Models
- Creative Preference Optimization
- Mistral 7B
- Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
- GPT-4 Technical Report
- Mind the Gap: Conformative Decoding to Improve Output Diversity of Instruction-Tuned Large Language Models
- Code Llama: Open Foundation Models for Code
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- LLaMA: Open and Efficient Foundation Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering