LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers
summary
The gist
The gist The paper discusses designing and evaluating Large Language Models (LLMs) for analyzing security operation center (SOC) reports by creating an Analyst-wise Checklist and a novel framework
In short
The paper designs an Analyst-wise Checklist based on SOC practitioner knowledge and creates MESSALA, a novel framework for LLMs to evaluate security reports. MESSALA uses this checklist to provide expert-level, multi-perspective feedback, proving that LLMs can quantitatively score reports and qualitatively offer actionable insights superior to existing methods.
Key concepts
- Analyst-wise Checklist
- This is a set of evaluation criteria for security analysis reports created by reviewing public guidelines and interviewing SOC practitioners. It organizes criteria into three key areas: decision support, technical understanding, and accountability, reflecting real-world SOC knowledge.
- MESSALA Framework
- A novel framework that maximizes report evaluation by integrating the Analyst-wise Checklist with a Granularization Guideline. It uses two LLMs—one for high-level assessment and one for in-depth evaluation—to imitate how veteran SOC practitioners think.
- Multi-perspective Evaluation
- This involves using different lenses to evaluate a report, combining superficial information assessment with deep, detailed evaluation guided by the checklist. This process aims to reproduce the complex cognitive process of human experts when judging security analysis reports.
Terminology used across episodes
This episode discusses
- LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers · Paper Radio
- Towards AI-Driven Human-Machine Co-Teaming for Adaptive and Agile Cyber Security Operation Centers
- Can Large Language Models Be an Alternative to Human Evaluations?
- MARG: Multi-Agent Review Generation for Scientific Papers
- Human-like Summarization Evaluation with ChatGPT
- ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing
- HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition
- LOCALINTEL: Generating Organizational Threat Intelligence from Global and Local Cyber Knowledge
- Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
- Decoding BACnet Packets: A Large Language Model Approach for Packet Interpretation
- ContextBuddy: AI-Enhanced Contextual Insights for Security Alert Investigation (Applied to Intrusion Detection)
- Automated Alert Classification and Triage (AACT): An Intelligent System for the Prioritisation of Cybersecurity Alerts
The paper
LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers · Read on arXiv
Panasonic Holdings Corporation
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers".
Elias: The gist The paper discusses designing and evaluating Large Language Models (LLMs) for analyzing security operation center (SOC) reports by creating an Analyst-wise Checklist and a novel framework called MESSALA…
Nadia: First, who's behind it and why it matters.
Paper summary: Nadia: So, we're looking at this paper called "LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers". The main idea is that LLMs are used to analyze security reports from SOCs, but there aren't really any established criteria on how good those reports actually are.
Elias: Exactly. They argue that because those evaluation criteria are missing, it’s unclear if the AI can properly judge these reports based on what actual SOC practitioners know.
Nadia: The paper addresses this by first creating an Analyst-wise Checklist, which is meant to capture the knowledge of real SOC folks through looking at literature and talking to analysts. That's their answer to the first research question.
Priya: So, they are basically trying to build a standard for what makes a good analysis report based on expert input before they even try to use an AI on it.
Elias: Right, and then they propose this whole framework called MESSALA, which uses that checklist to guide the LLM in giving feedback from those expert perspectives. It’s designed to imitate how a real SOC analyst thinks.
Nadia: The core claim is that by using this multi-perspective approach—the checklist and MESSALA—they can actually do two things: quantitatively score the reports and qualitatively give actionable feedback based on expert judgment.
Priya: That sounds like they are trying to bridge the gap between what an LLM spits out and what a human expert would actually care about in their daily work.
Elias: It seems important because it shows that you can ground the AI's evaluation not just in surface-level text, but in these structured, multi-faceted criteria derived from real practice.
Conclusion: Nadia: So, looking at this paper's title, "LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers," it really boils down to giving the AI a proper way to judge security reports.
Elias: The authors are Okada and his colleagues from Panasonic Holdings Corporation in Osaka. They developed this framework called MESSALA as their main contribution.
Nadia: What this means in simpler terms is that instead of just letting an LLM read a report and guess how good it is, you give the LLM a detailed set of rules based on what experts actually look for—things like decision support or technical understanding.
Priya: So, the implication here is that if we want to use AI tools to help us manage security incidents, we have to first invest time in defining those expert criteria so the AI isn't just guessing randomly.
Elias: That makes sense. The paper confirms that this method allows LLMs to produce both a measurable score and specific comments that analysts can actually use when they review their work.
Nadia: It suggests that for any tool meant to assist in security operations, the design phase needs to focus heavily on integrating real-world human judgment into the evaluation process from the start.
More episodes
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits
- 2610.10742-BRANCH: Bypassing Multi-Scanner AI Guardrails
- 2610.10752-Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection
- 2610.10766-CPU-Auth: Device Fingerprinting for Authentication via DVFS Side-Channel