Hardening Soft Information: Evidence on Analyst Integration Costs

arXiv:2505.12269 · econ.GN, cs.AI, cs.CL, math.LO, q-fin.EC, q-fin.GN · Submitted 2025-05-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Hardening Soft Information".

Jane: Analysts communicate their subjective expectations for firms’ future using natural language rather than numerical forecasts, and this linguistic expression predicts errors in those numerical forecasts.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So we've talked about how analysts use language to communicate vague expectations, but let's get a clearer picture of what "Hardening Soft Information: Evidence on Analyst Integration Costs" is actually saying about this topic. Jane, can you give us the main thesis in simple terms?

Jane: The paper argues that people often have vague knowledge about future outcomes, which makes it hard to quantify those outcomes precisely, so they communicate these expectations using natural language instead of exact numbers. The central claim is that this linguistic expression plays an important role in how analysts form and revise their expectations.

Lu: Essentially, the authors are testing the idea that we should relax standard economic models by including vague knowledge, and they provide empirical evidence showing that analysts use language to convey these vague expectations instead of relying solely on numerical forecasts.

Meng: So, what's the significance of this finding for the financial world? Why should anyone care if analysts are using words instead of precise figures? Is this a small academic curiosity or something more substantial?

Lalam: It matters because it suggests that useful information isn't always available as a clean number; sometimes it's embedded in the way people speak about things, and we need to learn how to read that linguistic signal. This has implications for how we build systems to understand human behavior.

Tom: Right, so the paper is showing that when you look at analyst reports, they include valuable information in their textual descriptions but leave out the precise numerical forecasts, and this language itself has predictive power regarding those forecasts.

Jane: That means the way an analyst frames their view—whether it's optimistic or cautious—can actually influence whether their final numerical estimate ends up being too high or too low, depending on the tone they use.

Lu: The paper proposes that a rough set, which is vague knowledge, can be represented by text rather than precise numbers; this shows that linguistic representation is a valid way to handle uncertainty in information structures.

Meng: From an engineering perspective, if we can't rely on crisp sets for everything, we have to design systems that are robust enough to handle fuzzy inputs effectively. It shifts our focus from just finding the perfect number to interpreting the whole context of the communication.

Lalam: This suggests a direction where AI could focus less on pure mathematical optimization and more on capturing and processing these qualitative, textual signals effectively, which feels like a significant shift in how we approach complex data interpretation.

Conclusion: Tom: So, we’ve looked at how analysts use language to communicate vague expectations, why this matters for their forecasts, and what that suggests about how we model information. Now let’s wrap up with the conclusion of "Hardening Soft Information: Evidence on Analyst Integration Costs." Jane, what's your take on the authors and the overall message?

Jane: The paper by Kerry Xiao and Amy Zang presents a really important idea: that some useful information is known vaguely and is only communicated through language. The implication is that things like corporate culture or disruptive innovation aren't reflected in standard financial reports but are instead conveyed through narrative.

Lu: That speaks to the fundamental limitation of current accounting systems, which seem designed for precise, quantifiable data rather than capturing these kinds of subjective experiences. It challenges the assumption that all relevant information must be measurable in a numerical sense.

Meng: So, if this is true, it means any system we build to analyze business or culture needs to account for this linguistic layer; we can't just ignore the stories people tell us because they aren't in a spreadsheet. What does that mean for practical application?

Lalam: For AI development, it means moving beyond just pattern matching numbers and learning to genuinely understand the context and tone of communication, which is where cultural insights come from. This could really enhance the richness of our models.

Tom: Exactly! It suggests that analysts are making rational choices between using soft language for subjective knowledge and hard numbers for precise data, depending on how well they understand the underlying precision of what they're trying to convey.

Jane: So, the authors are suggesting that this phenomenon isn't just a quirk of finance; it’s a way real-world knowledge gets shared when quantification is too hard or undesirable, and we need to acknowledge that channel in our analysis.

Lu: It reinforces the idea that we have to respect the structure of how information is communicated, rather than assuming that the best information is always numerical. This paper opens up avenues for incorporating this vague knowledge directly into modeling subjective expectations.

Meng: It’s a lot to take in, but from an engineering viewpoint, it tells us that the most valuable signal might not be the one we can measure with a single metric, but the entire communication structure surrounding it. That's something to keep in mind when designing next-generation analytical tools.

Lalam: I think this is exciting because it validates the idea that high-quality, complex insights are often found in the 'soft' layer of data, and if we can decode that language better with AI, we can build systems that capture a much fuller picture of how things actually work.

Tom: Fantastic discussion on "Hardening Soft Information: Evidence on Analyst Integration Costs." It’s clear that the authors are pushing us to look beyond just the numbers and pay close attention to the language analysts use, because that language tells a story about what they actually know.

Hong Kong University of Science and Technology

econ.GN, cs.AI, cs.CL, math.LO, q-fin.EC, q-fin.GN

Submitted: 2025-05-18

Updated: 2026-10-01

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 74/100

The gist: Analysts communicate their subjective expectations for firms’ future using natural language rather than numerical forecasts, and this linguistic expression predicts errors in those numerical

Key concepts

Vague Knowledge
This refers to the way people understand the world, which is often imprecise. Standard economic models assume all observed signals are perfectly clear, but real-world knowledge frequently involves 'borderline cases' where things cannot be definitely classified as belonging to a set.
Linguistic Communication
People convey this vague knowledge through natural language rather than precise numbers. This is advantageous because it allows analysts to describe concepts like 'management ability' or 'innovation' where a specific mathematical measure is unavailable, avoiding false clarity from numerical estimates.
Forecast Error Prediction
The study found that the tone of the language used in an analyst report significantly predicts errors in their later numerical forecasts. Optimistic language leads to forecasts that are systematically too low, demonstrating how linguistic expression influences the final quantitative prediction.

Terminology

Summary

Analysts communicate their subjective expectations for firms’ future using natural language rather than numerical forecasts, and this linguistic expression predicts errors in those numerical forecasts. This finding suggests that useful information is often vaguely known and only communicated through language, which has significant implications for how financial analysts form and revise their expectations.

Theoretical Framework of Vague Knowledge

The study incorporates vague knowledge into models of subjective expectations by relaxing standard information structures, arguing that people often observe and understand the state of the world in vague ways. This is summarized in Proposition 1: A well-defined boundary is not a necessary condition for an information set to be informative. Standard economic theory assumes all observed signals are precise (crisp sets), but real-world knowledge often involves vagueness, where objects cannot be definitely classified as belonging to an information set or not, leading to borderline cases. This distinction is reflected in the physical world: natural language is vague, while numbers are precise.

Mechanism of Linguistic Communication

The paper proposes that people convey their vague knowledge through natural language. Proposition 2 states that A rough set cannot be faithfully represented by any crisp set (e.g., numbers), but it can be faithfully represented by another rough set (e.g., text). Linguistic expressions are advantageous because they:

  1. Allow analysts to describe information for which a relevant precise measure is unavailable, such as concepts like management ability, innovation, brand names, governance, and synergy.

  2. Convey imprecision truthfully when the underlying information base is too limited to warrant mathematical modeling and precise predictions, avoiding numerical estimates that may lend false clarity.

  3. Reduce production costs and enable prompt communication within analysts’ processing capacity because quantification incurs additional processing costs.

Empirical Predictions on Forecast Errors

The theory leads to several empirical hypotheses concerning the relationship between textual tone and forecast errors:

  1. H1 predicts a systematical negative relation between errors in numerical forecasts and textual opinions in the same report. This is tested by finding that the tone is significantly and negatively related to the forecast error of the same report, suggesting that analysts’ numerical forecasts are too low (high) when they include more optimistic statements in natural language.

  2. H2 predicts that vague expectation in earlier reports, though expressed in natural language, precedes numerical forecasts. This is supported by finding that the positive (negative) textual tone in the earlier analyst reports to be predictive of subsequent upward (downward) revisions in the analysts’ numerical research outputs.

  3. H3 predicts that linguistic expressions convey a vague expectation that is too costly to be incorporated into numerical forecasts, and this effect is stronger when analysts are constrained by time or processing capacity.

Moderating Effects of Uncertainty and Constraints

The predictive power of linguistic expressions is not constant; it varies with external factors:

  1. The relation between errors in numerical forecasts and textual opinions is stronger when the report exhibits a higher degree of vagueness. Specifically, the interaction term between textual tone and the indicator variable for vagueness shows a significant negative coefficient.

  2. The positive relation between current textual opinions and future revisions is stronger when the analyst faces a higher degree of uncertainty in forecasting tasks, such as longer forecast horizons, high macroeconomic/market volatility (VIX), or poor information environments (small firm size, low analyst coverage).

  3. The effect of linguistic expressions on subsequent revisions is also moderated by processing capacity: the positive relation between current textual opinions and future revisions is stronger when the analyst's processing capacity is more constrained.

Alternative Explanations and Conclusion

The study tests alternative explanations for the observed effects:

  1. Optimistic bias in numerical forecasts (using percentages of positive/negative sentences) was shown not to fully explain the results, as both good news and bad news conveyed linguistically can predict errors in respective directions.

  2. Offering clients an information advantage (delaying quantification of precise signals) was ruled out by testing subsamples where analysts were less likely to have such incentives (e.g., firms with low institutional ownership).

The overall conclusion is that some useful information is vaguely known and only communicated through language, and this finding has implications for accounting systems, suggesting that intangibles such as breakthrough and disruptive innovation, customer relations, and corporate culture are not reflected in financial reports but are instead communicated linguistically. Analysts make rational choices between soft (linguistic) and hard (numerical) communication formats based on the precision of underlying knowledge.

Key Empirical Findings Summary

**: The overall textual tone is significantly and negatively related to the forecast error, meaning optimistic language leads to forecasts that are too low relative to actual earnings. **

**: The relationship between current textual opinions and future revisions is positive, indicating that analysts wait for new evidence before quantifying vague expectations. **

**The predictive power of linguistic expressions increases when reports are vaguer, when uncertainty is higher, and when analysts are busier.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed the core theoretical contributions and empirical findings of this paper, Vague Knowledge: Evidence from Analyst Reports. The central insight is that analysts communicate useful information about future states not through precise numerical forecasts (hard information), but through vague natural language (soft information), which predicts errors in those numerical forecasts.

Here are specific improvements to AI systems based on this research, and what the improved AI system can achieve:


) 1. Implementation of Vague Expectation Modeling for Uncertainty Quantification:

The paper develops a formal decomposition of future states into precise expectations (based on numbers), vague expectations (based on language/rough sets), and unpredictable residual, leading to Equation (3):

Forecast Error = - [Vague Expectation] + Bias(Precise) + Noise.

Improvement: AI models used for financial forecasting should incorporate a mechanism to explicitly separate hard quantifiable signals from soft linguistic signals. The system must be able to estimate the Vague Expectation term—the portion of the forecast error attributable to unquantifiable, vague knowledge expressed in textual reports.

Improved System Capability: An AI financial model would not only output a numerical forecast but would also provide a confidence interval for that forecast, explicitly quantifying the uncertainty arising from qualitative analyst commentary. This allows downstream risk management systems to distinguish between risks driven by quantifiable market volatility and risks driven by ambiguous, unmeasurable sentiment expressed in reports (e.g., predicting earnings revisions based on the tone of language).

) 2. Linguistic Signal Extraction and Tone-Based Error Prediction:

The study uses FinBERT to quantify textual tone (positive, negative, neutral) and empirically finds a significant negative relationship between this tone and forecast errors (Prediction 1). Furthermore, it tests how the degree of vagueness (measured by the percentage of non-numerical sentences or hedging words) modulates this predictive power (Prediction 2).

Improvement: AI systems designed to process analyst reports must move beyond simple sentiment analysis. They need a layered linguistic processing pipeline that extracts not just what is said but how it is said. This involves fine-tuning large language models (like FinBERT) specifically for financial text and developing classifiers sensitive to linguistic vagueness (identifying hedging words, approximators, and qualitative descriptors).

Improved System Capability: An AI research assistant could ingest thousands of analyst reports and automatically flag reports where the vague expectation is likely to cause forecast errors. If a report uses high levels of vague language or hedging (high vagueness indicator), the system should assign a higher weight to its textual signal when predicting subsequent numerical revisions, effectively predicting which qualitative statements will lead to future quantitative forecast adjustments.

) 3. Uncertainty-Dependent Information Prioritization:

The research shows that analysts wait for new evidence when uncertainty is high (H2), and that the predictive power of linguistic expressions increases with the degree of forecasting uncertainty (Prediction 4).

Improvement: The AI's information-gathering module should dynamically adjust its reliance on different data types based on perceived future uncertainty. If a forecast horizon is long or macroeconomic volatility is high, the system must prioritize analyzing qualitative textual signals over refining precise numerical parameters until new, more precise evidence becomes available.

Improved System Capability: A predictive AI for long-term investment strategy would use linguistic cues (e.g., shifts in tone toward caution) as leading indicators that a period of vague knowledge is prevalent, prompting the system to reduce reliance on short-term numerical models and instead focus on scenario planning or qualitative risk assessment until the uncertainty resolves.

) 4. Processing Capacity and Cost-Benefit Analysis for Communication:

The paper suggests that linguistic expressions are economically feasible because they have lower production costs than quantification when dealing with vague knowledge (H3, Prediction 5).

Improvement: AI systems tasked with generating analyst reports or summarizing complex data should incorporate a cost-benefit layer. When presenting information, the system should weigh the precision gained from numerical output against the processing cost of producing that precision. If the underlying knowledge is inherently vague (as identified by linguistic features), the system should favor lower-cost, linguistically expressed summaries over high-precision numerical ones.

Improved System Capability: An automated reporting tool could generate a summary for a client, choosing between two outputs: (A) A highly precise numerical projection with associated high computational cost, or (B) A linguistically nuanced summary that conveys the expected vague expectation with lower processing overhead, thereby optimizing the communication of imperfect knowledge.

Abstract

We examine how the cost of transforming qualitative information into precise numerical estimates--a form of integration cost--creates a structural friction in expectations formation. To isolate this integration cost from the costs of information awareness and acquisition, we exploit sell-side analyst reports, in which the same forecaster simultaneously produces textual narratives and numerical forecasts. Because the information underlying the text has already been acquired, any systematic gap between the two outputs can be attributed to integration costs. We document systematic quantification inefficiency: an analyst's textual tone negatively predicts her contemporaneous forecast errors and positively predicts her subsequent numerical revisions, revealing that analysts leave part of their qualitative insights unquantified until further evidence arrives. Consistent with this integration-friction explanation, this inefficiency intensifies when reports are linguistically vaguer, environmental uncertainty is higher, or analysts' processing capacity is more constrained, and it persists where strategic and behavioral explanations are weaker. Our findings provide direct, large-sample evidence that integration costs constitute a distinct economic friction, explaining why soft information carries value-relevant content beyond contemporaneous hard numbers.

Related papers