LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers

arXiv:2601.03013 · cs.CR · Submitted 2026-01-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers".

Elias: The gist The paper discusses designing and evaluating Large Language Models (LLMs) for analyzing security operation center (SOC) reports by creating an Analyst-wise Checklist and a novel framework called MESSALA…

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So, we're looking at this paper called "LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers". The main idea is that LLMs are used to analyze security reports from SOCs, but there aren't really any established criteria on how good those reports actually are.

Elias: Exactly. They argue that because those evaluation criteria are missing, it’s unclear if the AI can properly judge these reports based on what actual SOC practitioners know.

Nadia: The paper addresses this by first creating an Analyst-wise Checklist, which is meant to capture the knowledge of real SOC folks through looking at literature and talking to analysts. That's their answer to the first research question.

Priya: So, they are basically trying to build a standard for what makes a good analysis report based on expert input before they even try to use an AI on it.

Elias: Right, and then they propose this whole framework called MESSALA, which uses that checklist to guide the LLM in giving feedback from those expert perspectives. It’s designed to imitate how a real SOC analyst thinks.

Nadia: The core claim is that by using this multi-perspective approach—the checklist and MESSALA—they can actually do two things: quantitatively score the reports and qualitatively give actionable feedback based on expert judgment.

Priya: That sounds like they are trying to bridge the gap between what an LLM spits out and what a human expert would actually care about in their daily work.

Elias: It seems important because it shows that you can ground the AI's evaluation not just in surface-level text, but in these structured, multi-faceted criteria derived from real practice.

Conclusion: Nadia: So, looking at this paper's title, "LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers," it really boils down to giving the AI a proper way to judge security reports.

Elias: The authors are Okada and his colleagues from Panasonic Holdings Corporation in Osaka. They developed this framework called MESSALA as their main contribution.

Nadia: What this means in simpler terms is that instead of just letting an LLM read a report and guess how good it is, you give the LLM a detailed set of rules based on what experts actually look for—things like decision support or technical understanding.

Priya: So, the implication here is that if we want to use AI tools to help us manage security incidents, we have to first invest time in defining those expert criteria so the AI isn't just guessing randomly.

Elias: That makes sense. The paper confirms that this method allows LLMs to produce both a measurable score and specific comments that analysts can actually use when they review their work.

Nadia: It suggests that for any tool meant to assist in security operations, the design phase needs to focus heavily on integrating real-world human judgment into the evaluation process from the start.

Panasonic Holdings Corporation

cs.CR

Submitted: 2026-01-06

Updated: 2026-04-08

Journal ref: European Symposium on Research in Computer Security (ESORICS 2026)

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 92/100

The gist: The gist The paper discusses designing and evaluating Large Language Models (LLMs) for analyzing security operation center (SOC) reports by creating an Analyst-wise Checklist and a novel framework

Key concepts

Analyst-wise Checklist
This is a set of evaluation criteria for security analysis reports created by reviewing public guidelines and interviewing SOC practitioners. It organizes criteria into three key areas: decision support, technical understanding, and accountability, reflecting real-world SOC knowledge.
MESSALA Framework
A novel framework that maximizes report evaluation by integrating the Analyst-wise Checklist with a Granularization Guideline. It uses two LLMs—one for high-level assessment and one for in-depth evaluation—to imitate how veteran SOC practitioners think.
Multi-perspective Evaluation
This involves using different lenses to evaluate a report, combining superficial information assessment with deep, detailed evaluation guided by the checklist. This process aims to reproduce the complex cognitive process of human experts when judging security analysis reports.

Terminology

Summary

The gist The paper discusses designing and evaluating Large Language Models (LLMs) for analyzing security operation center (SOC) reports by creating an Analyst-wise Checklist and a novel framework called MESSALA to provide expert-level feedback.

How it works

  1. The process begins with designing the Analyst-wise Checklist, which contains evaluation criteria for analysis reports in order to reflect real-world SOC practitioners’ knowledge through literature review and interviews with SOC practitioners, as the answer to RQ1 (Page 3). This checklist is constructed by collecting candidate items from a literature review of 13 public guidelines and handbooks and semi-structured interviews with 15 SOC practitioners, as described in Section 3.2 (Page 5). The criteria are organized into three standpoints: Decision support and action planning for stakeholders, Technical understanding and root-cause clarification of the event for analysts, and Accountability and quality assurance of the analysis process for each organization (Page 15).

  2. Next, the paper proposes MESSALA, a novel framework that maximizes evaluation by leveraging the Analyst-wise Checklist to imitate SOC practitioners’ cognition [77] (Page 3). MESSALA consists of three components: the Analyst-wise Checklist, the Granularization Guideline, and multi-perspective evaluation (Page 18). The Granularization Guideline steers LLM inference by translating checklist items into actionable evaluations, focusing on specific parts requiring detailed evaluations of veteran SOC practitioners [77] (Page 19).

  3. The Multi-perspective Evaluation LLM integrates two evaluation processes: a High-level Evaluation LLM based on superficial information and an In-depth Evaluation LLM using the Granularization Guideline, with their outputs integrated to return a final score and feedback [22] (Page 18). This design aims to reproduce human cognitive process of text recognition [74] (Page 18).

Quantitative Evaluation

The paper conducts quantitative evaluations with MESSALA against existing LLM-based methods to answer RQ2, which asks if LLMs can quantitatively evaluate reports from the perspective of SOC practitioners under defined criteria (Page 3). The evaluation uses two datasets: Real-world Analysis Reports and Pseudo-Reports generated by an LLM using a chain-of-thought (CoT) style prompt [23] (Page 24). The primary evaluation metrics include Spearman’s rank correlation ρ, Kendall’s tau τ, and Pearson’s correlation r, as well as the root mean square error (RMSE) to measure deviation in scores (Page 25). The results demonstrate that MESSALA consistently outperforms baseline methods in almost all settings, achieving a high correlation with human gold evaluations, such as correlations of 0.6 on several models [26] (Page 27).

Qualitative Evaluation

The paper conducts qualitative evaluations to answer RQ3, which asks if LLMs can qualitatively evaluate reports with feedback from the perspective of SOC practitioners under defined criteria (Page 3). This is done through two complementary evaluations: examining how LLM-generated feedback comments are useful for SOC practitioners and whether the LLM can correctly point out defects within analysis reports using a dataset consisting of defect-injected analysis reports [28] (Page 28). In the multi-metric rating, feedback comments generated by MESSALA were consistently rated higher by veteran SOC practitioners across all categories except for Accuracy and Validity, indicating that LLMs can provide actionable feedback for analysis reports compared with existing methods [30] (Page 31). Furthermore, the defect-injection evaluation showed that MESSALA identifies 86.4% of the injected defects, exceeding Method 1 across all defect categories [35].

Limitations and Ethical Considerations

The study notes several limitations, including the use of private datasets which are limited in number and length, necessitating extension into more diverse settings (Page 37). The Analyst-wise Checklist contains a large number of items, suggesting future work should focus on a mechanism that automatically selects these items (Page 37). Ethical considerations involved informed consent for interviews and obtaining consent from stakeholders responsible for each SOC before using analysis reports in experiments, adhering to the principle of data minimization (Page 38).

The paper concludes that MESSALA can quantitatively evaluate reports and qualitatively provide actionable feedback comments under the evaluation criteria as the answer to RQ2 and RQ3, respectively (Page 8). The findings show that integrating high-level and in-depth evaluations enables judgments to be grounded in concrete descriptions rather than superficial characteristics (Page 26). Finally, MESSALA can identify defects on analysis itself in addition to providing plausible explanations, demonstrating its ability to generate meaningful feedback [33] (Page 35). The paper plans further experiments across diverse settings and additional deployments in real-world SOC environments (Page 8).

--- Page 1 ---

LLMs, You Can Evaluate It! Design of Multi-perspective Evaluation for Reports in Security Operation Centers

Hiroyuki Okada, Tatsumi Oba, and Naoto Yanai Panasonic Holdings Corporation, Osaka, Japan okada.hiroyuki001@jp.panasonic.com Abstract. Security operation centers (SOCs) often produce analysis reports on security alerts, and large language models (LLMs) will likely be used for this task; While producing analysis reports with actionable insights requires constantly evaluating the reports of SOC practitioners, it is unclear whether LLMs appropriately evaluate the reports due to a lack of evaluation criteria; In this paper, we discuss evaluations using LLMs for analysis reports. To this end, we first design Analystwise Checklist that contains SOC practitioners’ evaluation criteria for analysis reports in the real world through literature review and interviews with SOC practitioners as a user study; Next, we design a novel LLM-based multi-perspective evaluation framework, named MESSALA, that maximizes report evaluation and provides feedback with analysis reports on veteran SOC practitioners’ perceptions by leveraging the checklist above; We also conduct quantitative evaluation and qualitative evaluation with real-world analysis reports. Keywords: security operation centers, analysis reports, large language models, report evaluation, user study; 1 Introduction Cyberattacks have increased across many organizations, and establishing a security operation center (SOC) to analyze security alerts and provide their responses has become urgent and crucial for each organization in recent years; An important mission for SOCs is, in addition to providing security alerts and their responses, to write analysis reports with clear and actionable insights into their underlying cyberattacks [66]; Interestingly, in proportion to the quality of analysis reports, various SOC practitioners, including not only analysts but also their managers, can understand the content of security alerts; However, writing analysis reports for security alerts forces a heavy workload on SOC practitioners because these reports must contain actionable insights [46]; To reduce this workload, tools such as large language models (LLMs), that introduce actionable insights into analysis reports are in high demand [13]; Although there are several tools arXiv:2601.03013v4 [cs.CR] 8 Apr 2026 to generate analysis reports for security alerts [53, 54], automatically generated reports often contain inaccurate information, which may degrade performance in SOCs [46]; Even when state-of-the-art LLMs are used, they often cause hallucinations, such as inaccurate information [36]; As another motivation, writing analysis reports to the exact requirements of report-writing guidelines is stressful [57]; The reason is that evaluation of analysis reports should identify any lack of descriptions as well as judgments from the perspectives of veteran SOC practitioners [62]. Based on the above background, in this paper, we discuss evaluations for analysis reports in SOCs using LLMs in order to improve the quality of the reports; Specifically, we discuss the following research questions in this paper: RQ1 What are the evaluation criteria that guarantee the quality of analysis reports from SOC practitioners’ knowledge?; RQ2 Can LLMs quantitatively evaluate the quality of analysis reports from the perspective of SOC practitioners under the defined evaluation criteria?; RQ3 Can LLMs qualitatively evaluate the quality of analysis reports with feedback from the perspective of SOC practitioners under the defined evaluation criteria? We first design the Analyst-wise Checklist, which contains the evaluation criteria in its specific items to reflect SOC practitioners’ knowledge for evaluating analysis reports in the real world, as an answer to RQ1; To design this, we conducted a literature review of 13 public guidelines and handbooks and semi-structured interviews with 15 SOC practitioners; Next, based on the above checklist, we propose a novel framework, called Multiperspective Evaluation System for Security Analysis using Llm Assistance (MESSALA), to maximize the evaluation of analysis reports in SOCs; In a nutshell, MESSALA guides LLMs in evaluating analysis reports by leveraging the Analyst-wise Checklist from multiple perspectives; As described in Section 4, MESSALA can imitate SOC practitioners’ cognition [77] to obtain expert knowledge from the checklist, and therefore, can return appropriate evaluation results with actionable feedback for each analysis report; When we conduct quantitative evaluations with MESSALA, we show that it can rate analysis reports for the evaluation criteria in comparison with the existing methods [26, 51]; Thus, to answer RQ2, we confirm if MESSALA enables LLMs to quantitatively evaluate analysis reports from the perspective of veteran SOC practitioners; Third, we conduct qualitative evaluations to identify whether the feedback generated by MESSALA aligns with judgments from the perspective of veteran SOC practitioners; As a result, we confirm if MESSALA can provide more actionable feedback to both novice and veteran SOC practitioners than the existing methods, as it is more specific and easier to understand; It means that MESSALA also enables LLMs to qualitatively evaluate analysis reports through its feedback [31].

--- Page 2 ---

H. Okada et al.

Improvements for AI systems

  1. No single-perspective evaluation is sufficient; MESSALA integrates High-level Evaluation LLM with an In-depth Evaluation LLM to avoid blind spots and biased judgments arising from a single viewpoint by virtue of crossreferencing outputs from different levels of evaluation [52]. This allows the system to evaluate reports by mimicking human cognitive processes, as described in the paper: "This design imitates the cognitive process of human text recognition [74]: humans first activate knowledge fragments through superficial features, and then integrate their contexts to understand the meaning of the given text."

  2. The Granularization Guideline enables deeper context-aware evaluation by translating checklist items into concrete confirmation points. For example, when evaluating If the incident is assumed to be an operational effect, is this explained?, the system generates specific points such as whether the maintenance activity is concretely described and checks if the timing of the anomalous communications is consistent with the maintenance notice.

  3. MESSALA provides actionable qualitative feedback by focusing on low-scoring items and using direct language. The system's feedback is restricted to checklist items with scores less than or equal to “3, focusing on points that received particularly low evaluation scores, as noted in Table 10, which helps users derive concrete improvement points and reproducible advice."

  4. The defect-injection evaluation demonstrates the system's ability to identify subtle quality issues beyond superficial detection. MESSALA identifies 86.4% of the injected defects, exceeding Method 1 across all defect categories, proving it can capture issues like Opaque Decision Rationale by explicitly pointing out when the report fails to explain why the communication occurred and lacks consideration of whether the communication or the involved endpoint is anomalous.

Sources

Related papers