Revolutionizing Finance with LLMs: An Overview of Applications and Insights
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Revolutionizing Finance with LLMs: An Overview of Applications and Insights".
Jane: The paper was written by Huaqin Zhao, Zhengliang Liu, Zihao Wu, Yiwei Li, Tianze Yang et al. from University of Georgia and Northwestern University and University of Texas at Arlington.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title and Authors: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's been making the rounds on arXiv, and it's called "Revolutionizing Finance with LLMs: An Overview of Applications and Insights." Jane, I have to say, the title alone gets me excited — we're talking about taking these massive language models and pointing them at the world of money.
Jane: Oh, absolutely, Tom. And before we get into the weeds, let's talk about who's behind this. The paper comes from a big team led by Huaqin Zhao and Zhengliang Liu at the University of Georgia, with collaborators from Northwestern, UT Arlington, and a bunch of other institutions. It's one of those sprawling collaborations that you see a lot in AI research these days.
Tom: Right, and what I love about this paper is that it's not just theory. They actually ran tests. They took GPT-four and threw it at six different financial datasets — sentiment analysis, named entity recognition, question answering, even stock movement prediction. And the results are genuinely impressive.
Jane: Yeah, let's talk about those numbers, because they're pretty striking. On the Financial Phrase Bank, GPT-four hit seventy-eight percent accuracy on sentiment analysis. On FiQA-SA, it got seventy-nine percent. And on the NER task — that's named entity recognition, finding companies and people in financial documents — it scored an eighty-one percent F1 score.
Tom: And those are zero-shot results, Jane. That means GPT-four wasn't fine-tuned on any of these financial datasets. It just read the instructions and went to work. For someone like me who remembers when we had to train a separate model for every single task, that's wild.
Jane: It really is. And the paper frames this as a survey plus an evaluation. So they're not just saying "look what GPT-four can do" — they're mapping out the whole landscape of how LLMs are being used in finance, from quantitative trading to fraud detection to robo-advisors.
Tom: And that's the part that gets me thinking about the real-world impact. If a general-purpose model can walk into financial tasks cold and perform this well, what does that mean for the industry? I mean, we're not talking about replacing human analysts tomorrow, but this is a serious signal.
Jane: Exactly. And the authors are careful to point out the limitations too. GPT-four isn't doing the heavy computational lifting in quantitative trading — it's more of a sentiment interpreter that feeds into existing models. But that's still a huge deal. It's like having a research assistant who can read every earnings report, every news article, every tweet, and tell you what the market mood is.
Tom: So when we look at the authors and the scope here, this feels like a foundational paper. It's laying out the map for where LLMs and finance intersect. And I think the next question is, what exactly did they find when they dug into the different financial tasks? That's coming up next.
Summary of the Paper: Jane: So we've talked about the authors and the big picture, but let's get into what this paper actually covers. "Revolutionizing Finance with LLMs" breaks the financial applications down into four main buckets: financial engineering, financial forecasting, financial risk management, and real-time question answering.
Tom: And each of those buckets has some really concrete use cases. Like in financial engineering, they're talking about quantitative trading and portfolio optimization. The idea is that LLMs can read through analyst reports and news articles to pick up on sentiment that traditional quantitative models just can't see.
Jane: Right, and that's the part I find fascinating. Traditional trading models work with numbers — prices, volumes, historical data. But so much of what moves markets is narrative. It's how people talk about a company, whether the tone in an earnings call is confident or nervous. And that's exactly what LLMs are good at.
Tom: Then you've got financial forecasting, which includes stuff like predicting mergers and acquisitions, forecasting insolvency, and market trend prediction. The paper actually shows GPT-four doing stock movement prediction on the BigData22 dataset, and it hit fifty-three percent accuracy. Now, that doesn't sound like much, but in a binary up-or-down prediction, anything above fifty percent is meaningful.
Jane: And let's not forget the risk management side. They looked at credit scoring, ESG scoring, fraud detection, and compliance checks. The fraud detection part is really cool — they used a simulated mobile money transactions dataset called PaySim, and GPT-four correctly identified all five transactions in their test case.
Tom: I loved that example in the paper. They gave GPT-four a table of transactions and asked it to flag suspicious ones. And it didn't just say "this one's fraud" — it explained why. It noticed that in one transaction, the entire balance was transferred out, which is unusual. In another, the destination account was zeroed out immediately after receiving funds, which is a classic money-laundering pattern.
Jane: That's the thing about LLMs that's different from traditional fraud detection systems. They don't just give you a flag — they give you a rationale. And in a compliance context, that's huge. Regulators want to understand why a transaction was flagged, not just that it was.
Tom: And then there's the question answering side, which covers financial education. The paper talks about using GPT-four to explain complex financial concepts in plain language, customize learning for different skill levels, and create interactive learning experiences. On the FinQA and ConvFinQA datasets, GPT-four got sixty-four percent and seventy-three percent exact match accuracy respectively.
Jane: So the paper is really painting a picture of an AI that can read, reason, and explain across the entire financial landscape. But here's the thing — they also ran into some limitations. And that's what I want to dig into next, because the paper is pretty honest about where LLMs fall short.
Improvements Suggested by the Paper: Tom: Alright, so we've covered what the paper found, but now let's talk about where it says we need to go from here. Because "Revolutionizing Finance with LLMs" isn't just a victory lap — it's pretty clear about the gaps.
Jane: Yeah, and the biggest gap is that LLMs can't do the actual computational work. They're great at reading text and understanding sentiment, but they can't directly optimize a portfolio or execute a trade. The paper calls their role "auxiliary" — they feed insights into existing quantitative models, but they're not standalone solutions.
Tom: And that's where I want to bring in Lu, because I know you've been thinking about this. Lu, what do you make of the paper's suggestion that we need hybrid systems — combining LLMs with traditional quantitative models?
Lu: I think that's exactly the right direction, Tom. The paper hints at it in the future work section — developing systems where the LLM handles the qualitative side and the quantitative models handle the math. But I'd push it further. What if the LLM isn't just feeding sentiment into a model, but actually helping to design the trading strategy itself? We're starting to see work on LLMs that can generate code for backtesting strategies, and that's a natural extension of what this paper is describing.
Meng: But hold on, Lu. From my side, the practical concern is reliability. If an LLM is generating trading strategies, how do you verify that it's not hallucinating? The paper mentions accuracy and reliability as major challenges, but I don't think it fully addresses how you'd build a system that can catch those errors before they cost real money.
Jane: That's a really good point, Meng. And the paper does acknowledge this — it talks about combining LLMs with expert systems and manual review mechanisms. But you're right that the verification problem is still open. How do you know when the LLM is confidently wrong?
Tom: And that's where the paper's suggestion about interpretability comes in. One of the things they highlight is that GPT-four can explain its reasoning. In the fraud detection example, it didn't just flag a transaction — it walked through why it was suspicious. That transparency is valuable, but it also means we need to check whether the reasoning is actually sound, not just whether it sounds plausible.
Lu: Exactly, Tom. And I think the next step is what the paper calls "enhancing the interpretability and reliability of LLM outputs in financial contexts." That means building tools that can audit the LLM's reasoning, cross-check its claims against verified data, and flag when it's operating outside its knowledge boundary.
Meng: And there's also the data problem. The paper uses datasets like FinQA and ConvFinQA, but those are relatively small and clean. Real financial data is messy, noisy, and constantly changing. If we're going to deploy these systems in production, we need to know how they perform when the data isn't curated.
Jane: So the improvements the paper is suggesting really come down to three things: integration with quantitative models, better interpretability, and robustness in real-world conditions. And I think that last point is going to be the hard one.
Conclusion: Tom: Alright, we're wrapping up our discussion of "Revolutionizing Finance with LLMs: An Overview of Applications and Insights." And I have to say, this paper left me feeling genuinely optimistic about where this technology is headed.
Jane: Me too, Tom. Let's recap what we learned. The paper surveyed four major areas of finance — engineering, forecasting, risk management, and question answering — and showed that GPT-four can handle tasks across all of them with zero-shot learning. We saw seventy-eight percent accuracy on financial sentiment, eighty-one percent F1 on named entity recognition, and even fifty-three percent on stock movement prediction.
Tom: And we talked about the limitations too — LLMs can't do the heavy computational lifting, they need to be paired with quantitative models, and we still have to solve the reliability problem. But the paper's vision of hybrid systems, where LLMs handle the language and traditional models handle the math, feels like the right path forward.
Jane: And let's not forget the human impact. The paper talks about using LLMs for financial education — helping people understand complex concepts, getting personalized learning experiences, and making financial knowledge more accessible. That's a real benefit for everyday people, not just Wall Street.
Tom: Absolutely. And I think the biggest takeaway for me is that this paper is a roadmap. It shows us where we are, where we need to go, and what the obstacles are. It's not promising magic — it's showing us the hard work that needs to happen.
Jane: Well said. So we're going to say goodbye to "Revolutionizing Finance with LLMs" and get ready for the next paper on our list. Thanks to Lu and Meng for joining us today, and thanks to all our listeners for tuning in.
Tom: And remember, if you're curious about this paper, it's on arXiv — go check it out. The authors did some really solid work, and there's a lot to dig into. Until next time, keep asking questions.
Jane: See you all on the next episode.
Huaqin Zhao, Zhengliang Liu, Zihao Wu, Yiwei Li, Tianze Yang, Peng Shu, Shaochen Xu, Haixing Dai, Lin Zhao, Hanqi Jiang, Yi Pan, Junhao Chen, Yifan Zhou, Wei Ruan, Zheyuan Zhang, Zeyu Zhang, Ruitong Sun, Gengchen Mai, Ninghao Liu, Tianming Liu
University of Georgia · Northwestern University · University of Texas at Arlington
cs.CL
Submitted: 2026-08-12
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 53/100
Key concepts
- Zero-Shot Results
- This refers to the impressive ability of GPT-four to perform financial tasks without being fine-tuned on specific datasets. It means the model can successfully complete a complex task simply by reading instructions and applying its general knowledge.
- Named Entity Recognition (NER)
- In finance, NER is the process of identifying and extracting key entities—such as company names, people, or financial instruments—from large documents. The paper noted GPT-four scored an 81% F1 score on this task.
- Hybrid Systems
- The paper suggests that LLMs should not be standalone solutions. Instead, they must be combined with traditional quantitative models (which handle the math) to create robust financial systems that leverage both language understanding and computation.
- Interpretability
- This refers to the ability of an AI model, like GPT-four, to explain its reasoning or rationale for a decision. In fraud detection, this transparency is crucial because regulators need to know *why* a transaction was flagged.
Terminology
Summary
Summary
In recent years, Large Language Models (LLMs) like ChatGPT have seen considerable advancements and have been applied in diverse fields. Built on the Transformer architecture, these models are trained on extensive datasets, enabling them to understand and generate human language effectively. In the financial domain, the deployment of LLMs is gaining momentum. These models are being utilized for automating financial report generation, forecasting market trends, analyzing investor sentiment, and offering personalized financial advice. Leveraging their natural language processing capabilities, LLMs can distill key insights from vast financial data, aiding institutions in making informed investment choices and enhancing both operational efficiency and customer satisfaction. In this study, we provide a comprehensive overview of the emerging integration of LLMs into various financial tasks. Additionally, we conducted holistic tests on multiple financial tasks through the combination of natural language instructions. Our findings show that GPT-4 effectively follow prompt instructions across various financial tasks. This survey and evaluation of LLMs in the financial domain aim to deepen the understanding of LLMs’ current role in finance for both financial practitioners and LLM researchers, identify new research and application prospects, and highlight how these technologies can be leveraged to solve practical challenges in the finance industry.
Finance is a highly specialized and complex field that involves a great deal of data analysis, prediction, and decision making. LLM’s ability to process large-scale text data makes it a promising application in the financial field. For example, by analyzing financial reports, market news, investor communications, etc., LLMs can provide insights into market trends, perform risk assessments, and even assist in investment decisions. In addition, LLMs can process natural language queries and provide instant financial advice and support, which is a big step forward for the financial services industry. However, applying LLMs to the financial sector also faces several challenges. First, data in the financial domain is highly specialized and complex. Financial terminology, regulations, and market dynamics require a high level of model comprehension. In addition, financial decision-making usually involves high risk, which requires a high degree of accuracy and reliability in prediction. Therefore, it is a major challenge to ensure that the output of LLMs is both accurate and reliable. To address these issues, researchers and developers are continuously refining the algorithms of LLMs to improve its understanding and processing of specialized domain knowledge. With a large amount of specialized training data, the model can better grasp specific knowledge in the financial domain. At the same time, the combination of expert systems and manual review mechanisms can further improve the accuracy and reliability of the model’s application in the financial domain.
The significant contributions of this article are distilled into four primary points, each focusing on the synergy between LLMs and financial applications. First, we meticulously survey and synthesize existing LLMs for finance literature, exploring the latest advancements in four independent task categories: financial engineering, financial forecasting, financial risk management, and financial real-time question answering. Second, we summarize the primary technical approaches that LLMs offer to the realm of finance, examine the potential in the investment field, and provide a foundational survey for researchers in this domain. Third, we assess the effectiveness of GPT-4 in various tasks. Fourth, we concisely overview of the most significant results from our research, discuss the major unresolved issues that should be addressed in subsequent efforts, and offer insights into future directions and possibilities in this field.
The scope of finance tasks covered includes Financial Engineering, which is a multidisciplinary field that combines finance, mathematics, and computer science to create and implement innovative financial strategies and products. LLMs assist in Financial Engineering by enhancing two key subtasks: Quantitative Trading and Portfolio Optimization. In quantitative trading, LLMs, with their advanced natural language processing capabilities, play a pivotal role in effectively extracting and utilizing implicit sentiment information in investment strategies. By analyzing vast amounts of textual data, LLMs can identify subtle, often nuanced sentiments embedded in analysts’ reports, market news, and financial statements. The integration of LLMs into quantitative trading strategies represents a significant advancement in the field, allowing for a more holistic approach to investment decisions. In portfolio optimization, LLMs excel in processing and analyzing vast amounts of unstructured data, including market reports, news articles, and financial statements, providing deeper insights and supplementary analysis crucial for risk assessment. By augmenting quantitative data with qualitative insights derived from LLMs, investors can achieve a more holistic approach to portfolio optimization. Additionally, leveraging the analytical power of LLMs and artificial intelligence, robo-advisors are making significant strides in reshaping the world of financial investing. The essence of robo-advisors’ appeal lies in their computational power, which allows them to tailor portfolios to the individual user’s circumstances, taking into account market dynamics and personal risk preferences. The LLMs is critical in this context, parsing extensive data sets to discern complex financial market patterns, allowing robo-advisors to provide informed investment guidance.
Financial Forecasting includes Merge and Acquisition Forecasting, Insolvency Forecasting, and Market Trend Forecast. In Mergers and Acquisitions forecasting, NLP offers pivotal tools for mining and interpreting vast arrays of textual data. LLMs can adeptly analyze financial reports, news articles, and press releases to unearth underlying trends or strategic shifts that may hint at forthcoming M&A activities. For insolvency forecasting, language models can analyze a myriad of textual sources to gauge a company’s financial health accurately. By evaluating financial disclosures, news articles, and statements from corporate leaders, these models can detect early signs of financial distress. Incorporating GPT-4’s capabilities into market trend analysis represents a significant leap forward in the application of artificial intelligence within the domain of financial forecasting. GPT-4 brings to the table its formidable prowess in processing vast datasets, extracting nuanced patterns, and synthesizing this information to generate predictions. Its capacity to parse through disparate data sources, including real-time financial news, historical price data, and burgeoning trends on social media platforms, allows it to construct a multi-faceted view of market conditions. Furthermore, GPT-4 transcends mere predictive output; it provides the underlying rationale for its forecasts, thereby granting investors and analysts a window into the ’thought process’ of the AI.
Financial Risk Management includes Credit Scoring, ESG Scoring, Fraud Detection, and Compliance Check. The advent of LLMs offers a promising avenue to transcend the limitations of traditional rule-based or machine learning algorithms in credit and risk assessment. LLMs, with their prowess in multitask learning and few-shot generalization, present an opportunity to redefine the landscape of financial assessments. The integration of GPT-4 into the process of ESG scoring remains an abundant blank area deserving to explore. GPT-4 assists with enhanced data processing and analysis, allowing it to process vast amounts of unstructured data rapidly including corporate sustainability reports, news articles, social media posts, and other relevant documents. Leveraging advanced reasoning and text mining capabilities, LLMs can significantly contribute to the identification of financial fraud in various domains including transactions, emails, profiles, contractors, and decentralized finance. These LLMs serve as an initial filter, learning from customer transaction histories and detailed transaction information to isolate highly suspicious transactions from the billions processed, thereby substantially alleviating the manual labor burden involved in investigating vast quantities of transaction data. LLMs with zero-shot learning capabilities are becoming indispensable in the dynamic world of financial compliance, where regulations are in a constant state of flux. Zero-shot LLMs can adapt to new standards without the need for fine-tuning, which traditionally demands regular updates and a wealth of annotated data.
Financial Real-Time Question Answering includes Financial Education. GPT-4 is an advanced artificial intelligence language model developed by OpenAI that is capable of understanding and generating human-like natural language. This feature makes it a powerful tool for financial education, especially when it comes to explaining complex financial concepts, providing customized learning experiences, and enhancing user interaction. GPT-4 can simplify complex financial concepts into easy-to-understand language, has unique advantages in providing customized learning experience, and plays an important role in enhancing user interaction. However, although GPT-4 has many advantages in financial education, it also has some limitations, including reliance on existing knowledge bases and data, and the need to consider ethical and compliance issues.
In the evaluation section, the authors meticulously chose six diverse datasets to showcase the extensive capabilities of GPT-4 in the financial sector. These datasets encompass a wide range of text types, including news articles, analytical reports, and social media posts like tweets. The datasets include FPB and FiQA-SA for sentiment analysis, NER for named entity recognition, FinQA and ConvFinQA for question answering, and BigData22 for stock movement prediction. The authors examined various prompting strategies, including vanilla zero-shot prompting, Chain-of Thought enhanced zero-shot prompting, and one-shot prompting to investigate their impact on GPT’s performance in the stated financial tasks. The formulation of prompts is essential in interacting with LLMs, and the prompt design includes three parts: System Role Explanation, Response Format for Different Tasks, and Example and Output.
In the tested financial tasks, LLMs demonstrated precise execution capabilities. Based on the responses gathered, the authors believe that LLMs exhibit exceptional zero-shot learning and mathematical reasoning abilities, along with their strongest suit, language sentiment analysis. The effectiveness of LLMs in financial tasks is quantitatively assessed by comparing their recommendations against real-world financial data and historical market performance. The results show that GPT-4 achieves an accuracy of 0.78 on FPB, 0.79 on FiQA-SA, an EntityF1 of 0.81 on NER, an EmAcc of 0.64 on FinQA, an EmAcc of 0.73 on ConvFinQA, and an accuracy of 0.53 on BigData22. For financial tasks lacking dedicated datasets, the authors curated case studies to showcase the capabilities of GPT-4, such as fraud detection on the PaySim dataset.
The limitations of LLMs are evident in areas such as optimization and quantitative trading. While they can assist in identifying market sentiments, LLMs cannot directly engage in computational tasks. Their role is more auxiliary, aiding in sentiment analysis which then feeds into existing models that handle quantitative variables. This indicates that LLMs, as of now, are not standalone solutions for computational finance tasks but rather powerful tools for augmenting existing models. For future work, there is immense potential in integrating LLMs with advanced quantitative models. One promising direction could be the development of hybrid systems that combine the text processing prowess of LLMs with sophisticated quantitative trading algorithms. Another area could be enhancing the interpretability and reliability of LLMs outputs in financial contexts, ensuring that the insights generated are not only accurate but also actionable. Moreover, exploring the application of LLMs in predictive analytics for market trends, based on historical data and current events, can open new avenues in financial forecasting. This integration of qualitative and quantitative analysis could revolutionize how financial markets are analyzed and traded.
In conclusion, the article delves into the multifaceted application of GPT-4 across a spectrum of 11 financial tasks, shedding light on the capabilities and constraints of LLMs in the financial domain. Central to the findings is the remarkable adeptness of LLMs in text processing, sentiment analysis, and their zero-shot learning abilities. The proficiency of LLMs in sifting through and interpreting extensive textual data is unmatched, thus playing a pivotal role in decoding market dynamics and investor sentiment. However, it is crucial to acknowledge the limitations of LLMs in direct computational tasks, particularly in optimization and quantitative trading, where their role remains largely supplementary. Despite these constraints, the potential of LLMs in enhancing financial models and decision-making processes is undeniable. As we advance, the integration of LLMs with quantitative models and the refinement of their application in finance will be areas of significant interest. The continual evolution of LLMs promises to not only bolster existing financial methodologies but also to pave the way for innovative approaches in financial analysis and strategy.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems, along with what the improved system can do:
Improvements:
-
Enhanced Zero-Shot Financial Instruction Following: Implement a prompt-engineering framework that combines system role explanation, task-specific response formatting, and one-shot examples. This improves the model's ability to execute financial tasks without fine-tuning, as demonstrated by GPT-4's performance on FPB (78% accuracy), FiQA-SA (79%), and NER (81% EntityF1).
-
Hybrid Sentiment Analysis Integration: Combine LLM-based sentiment extraction with traditional lexicon-based and machine learning approaches (e.g., Naive Bayes, SVM, KNN) to create a multi-layered sentiment scoring system. This addresses the limitations of unweighted word scoring and neutral sentiment detection, improving market trend prediction accuracy.
-
Structured Sequential Prompting for Numerical Reasoning: Implement chain-of-thought prompting that breaks down complex multi-step financial problems (e.g., from FinQA and ConvFinQA datasets) into sequential sub-tasks. This improves arithmetic, deductive reasoning, and common-sense understanding, as evidenced by GPT-4 achieving 64% and 73% exact match accuracy on these datasets respectively.
-
Zero-Shot Compliance Adaptation Module: Develop a system that leverages GPT-4's zero-shot learning to automatically adapt to regulatory changes without retraining. This is critical for compliance checks where fine-tuned models become obsolete when regulations update, reducing the risk of outdated assessments.
-
Fraud Detection Pattern Recognition Enhancement: Integrate LLM-based contextual analysis with existing transaction monitoring systems. The model learns from transaction histories and detailed transaction information to isolate highly suspicious transactions, reducing manual investigation burden. This is validated by GPT-4 correctly identifying 5 out of 5 fraudulent transactions in the PaySim dataset.
-
Multimodal Financial Data Fusion: Build a system that processes both structured tabular data (e.g., earnings reports) and unstructured text (e.g., news, tweets) simultaneously, using LLMs to bridge the gap between symbolic information and natural language. This enables more comprehensive analysis than traditional single-modality models.
What the Improved AI System Can Do:
-
Accurately classify financial news sentiment (positive/negative/neutral) with 78-79% accuracy, enabling automated market sentiment monitoring and early warning systems.
-
Extract key financial entities (organizations, persons, locations) from SEC filings and agreements with 81% F1-score, facilitating automated knowledge graph construction and compliance monitoring.
-
Answer complex numerical financial questions from earnings reports with 64-73% exact match accuracy, supporting automated financial analysis and investor decision support.
-
Predict stock price movements with 53% accuracy using a combination of tweets and historical price data, providing a baseline for algorithmic trading strategies.
-
Detect fraudulent transactions in real-time by analyzing transaction patterns and contextual cues, flagging suspicious activities for further investigation.
-
Adapt to new regulatory requirements instantly without retraining, ensuring compliance checks remain current and reducing the risk of penalties.
-
Provide personalized financial education by simplifying complex concepts, adjusting content difficulty based on user progress, and offering interactive Q&A.
-
Analyze ESG practices from unstructured data (sustainability reports, news, social media) to provide dynamic, real-time ESG scoring that reflects the most current information.
-
Assist in M&A forecasting by analyzing financial reports, news, and social media to identify strategic alignments and potential acquisition targets.
-
Support insolvency forecasting by detecting early signs of financial distress from corporate communications, regulatory filings, and sentiment shifts.
Abstract
In recent years, Large Language Models (LLMs) like ChatGPT have seen considerable advancements and have been applied in diverse fields. Built on the Transformer architecture, these models are trained on extensive datasets, enabling them to understand and generate human language effectively. In the financial domain, the deployment of LLMs is gaining momentum. These models are being utilized for automating financial report generation, forecasting market trends, analyzing investor sentiment, and offering personalized financial advice. Leveraging their natural language processing capabilities, LLMs can distill key insights from vast financial data, aiding institutions in making informed investment choices and enhancing both operational efficiency and customer satisfaction. In this study, we provide a comprehensive overview of the emerging integration of LLMs into various financial tasks. Additionally, we conducted holistic tests on multiple financial tasks through the combination of natural language instructions. Our findings show that GPT-4 effectively follow prompt instructions across various financial tasks. This survey and evaluation of LLMs in the financial domain aim to deepen the understanding of LLMs' current role in finance for both financial practitioners and LLM researchers, identify new research and application prospects, and highlight how these technologies can be leveraged to solve practical challenges in the finance industry.
Sources
- GPT-4 Technical Report
- FinBERT: Financial Sentiment Analysis with Pre-trained Language Models
- Coarse-to-fine Knowledge Graph Domain Adaptation based on Distantly-supervised Iterative Training
- LLM4TS: Aligning Pre-Trained LLMs as Data-Efficient Time-Series Forecasters
- A Survey on Evaluation of Large Language Models
- FinQA: A Dataset of Numerical Reasoning over Financial Data
- ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering
- Empowering Many, Biasing a Few: Generalist Credit Scoring through Large Language Models
- Beam Search Strategies for Neural Machine Translation
- OpenAGI: When LLM Meets Domain Experts
- GPT-InvestAR: Enhancing Stock Investment Strategies through Annual Report Analysis with Large Language Models
- Time-LLM: Time Series Forecasting by Reprogramming Large Language Models
- Multimodality of AI for Education: Towards Artificial General Intelligence
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
- Mask-guided BERT for Few Shot Text Classification
- Understanding LLMs: A Comprehensive Overview from Training to Inference
- Context Matters: A Strategy to Pre-train Language Model for Science Education
- Transformation vs Tradition: Artificial General Intelligence (AGI) for Arts and Humanities
- Radiology-GPT: A Large Language Model for Radiology
- Can ChatGPT Forecast Stock Price Movements? Return Predictability and Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering