How to Ask the AI: A User Perspective Survey for Large Language Model Prompting
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "How to Ask the AI: A User Perspective Survey for Large Language Model Prompting".
Jane: The paper was written by Yiqun Zhang, Yunfan Zhang, Mingjie Zhao, Sen Feng and Yiu-ming Cheung from Guangdong University of Technology and Hong Kong Baptist University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everyone. We are looking at a paper that basically everyone who has ever typed something into ChatGPT needs to see. It’s called "How to Ask the AI: A User Perspective Survey for Large Language Model Prompting." Jane, that title alone is doing a lot of work.
Jane: It really is, Tom. And I love that they put "user perspective" right in there. So many papers are written by engineers for engineers, but this one is saying, hey, we’re looking at this from the side of the person who just wants to get a good answer. It’s about the art of asking the question, not the machinery behind it.
Tom: Exactly. And the authors—Yiqun Zhang, Yunfan Zhang, Mingjie Zhao, Sen Feng, and Yiu-ming Cheung—they’re not just theorists. They’ve built this as a living project, which I think is a fantastic idea. It means the advice isn’t frozen in time.
Jane: Right, because the models change so fast. What works for one version of a model might not work for the next. So having a guide that can be updated is almost as important as the guide itself. They’re basically saying, we’ll keep teaching you how to ask, even as the answers get better.
Tom: And the implication here is huge. If we can lower the barrier to effective prompting, we’re not just helping people write better emails. We’re helping them think better. The way you frame a question to a model forces you to clarify what you actually want.
Jane: That’s such a good point. It’s like the old saying about teaching—if you can’t explain it simply, you don’t understand it. Prompting is like that. You have to know what you’re looking for before you can ask for it. This paper is basically a manual for that process.
Tom: So we’ve got a manual, we’ve got a taxonomy, and we’ve got a promise to keep it current. I’m curious about what’s actually inside. What are the strategies they’re talking about?
Jane: We’re going to get into that. But first, I want to say, the fact that they’re focusing on the user experience, not just the model’s performance, is a breath of fresh air. It’s about making the technology work for us, not the other way around.
Tom: Couldn’t agree more. And that’s the hook for our next segment, where we actually break down what the paper’s summary tells us about how to approach these tools. Stick around.
Summary: Tom: So, Jane, we’ve set the stage. Now let’s talk about what this paper actually says. The core message is that prompting isn’t just typing a question. It’s a skill, and it can be learned. The authors break it down into organizational tasks and innovative tasks.
Jane: I love that distinction. Organizational tasks are things like asking for a summary, getting a translation, or just having a conversation. Innovative tasks are where you’re asking the model to reason, to write a story, or to generate code. And the paper says, you should prompt differently depending on which one you’re doing.
Tom: Right. For the simple stuff, you can just ask directly. Zero-shot prompting, they call it. But for the hard stuff, like solving a math problem or planning a trip, you need to give the model more to work with. You might need to break the task down into steps.
Jane: And that’s where the "chain-of-thought" idea comes in. You’re literally asking the model to show its work. And the paper has a great way of showing this. They ran experiments where they asked models to draw a Christmas tree, and the difference in results based on the prompt was wild.
Tom: That was a great example. When they just said "draw a Christmas tree," the model did okay. But when they gave it a complex set of instructions about layers and colors, it actually got worse. It got confused trying to follow all the rules.
Jane: That’s the counterintuitive part. More instructions aren’t always better. Sometimes, giving the model too many constraints makes it fail. It’s like telling someone to walk a tightrope while juggling and reciting poetry. They’re going to drop something.
Tom: So the paper isn’t just a list of tricks. It’s a guide to thinking about what the model needs. And they’ve got this evaluation system where they use LLMs to judge other LLMs. It’s a clever way to get a consistent score.
Jane: And they compared that to human judges too. The scores were pretty aligned, which gives us confidence that the evaluation is fair. So we’re not just taking one model’s word for it. We have a multi-model consensus.
Tom: That’s a solid approach. It’s like getting a second opinion, but from six different doctors at once. And the takeaway is that there’s no single best prompt. It depends on what you’re trying to do.
Jane: Exactly. And that’s what we’re going to dig into next. We’re going to look at the specific improvements they suggest. How do you actually get better at this? What are the templates they’re offering?
Tom: So stick around. We’re about to get into the practical advice, the stuff you can use tonight when you’re trying to get a model to write a cover letter or debug your code.
Improvements: Tom: Alright, Jane, we’ve talked about the big picture. Now let’s get into the meat. The paper suggests that the biggest improvement comes from matching the prompt to the task. And they’ve created a decision table to help you do that.
Jane: That table is gold. It basically says, if you’re doing a knowledge Q andA, you should use retrieval-augmented generation, or RAG. That means telling the model to look things up before it answers. If you’re doing reasoning, you should use chain-of-thought. If you’re doing creative writing, you might want to use self-refine, where the model writes a draft and then critiques itself.
Tom: And the key improvement is that they don’t just tell you the strategy. They give you actual prompt templates. Like, here’s exactly what you type to get a better result. That’s the "replicable template" that so many users are missing.
Jane: Right. And one of my favorite examples from the paper is the email drafting. They show how a zero-shot prompt gets you a basic, functional email. But if you add a persona, like "you are a supportive team leader," the email changes completely. It becomes warmer, more encouraging.
Tom: That’s the persona prompting. And it’s such an easy improvement. You just add one sentence at the beginning, and the whole tone shifts. It’s like the difference between asking a friend for advice and asking a stranger.
Jane: And they also show the power of instruction prompting, where you list out the specific requirements. That gives you the most structured output. But it also takes more effort to write. So there’s a trade-off between effort and control.
Tom: And that’s the real improvement the paper offers. It’s not just "here’s the magic prompt." It’s "here’s how to think about the trade-offs." Do you want speed? Do you want accuracy? Do you want creativity? You have to pick your priority.
Jane: And they’ve quantified this. They had LLMs rate the different strategies on things like task performance, factuality, and ease of use. Direct prompting scores high on ease of use but lower on complex tasks. Reasoning-based prompting scores high on complex tasks but takes more time.
Tom: So it’s a menu, not a prescription. You look at what you need, and you pick the strategy that fits. And that’s a huge improvement over the trial-and-error that most of us are doing right now.
Jane: Absolutely. And this leads us to the first page of the paper, where they lay out the problem they’re trying to solve. We’re going to look at that next, and it’s a real eye-opener.
First Page: Tom: So, Jane, we’ve been talking about the solutions. But let’s go back to the beginning. The first page of "How to Ask the AI" paints a pretty stark picture of the problem. They talk about how users feel the LLMs are powerful but hard to control.
Jane: And that’s such a common feeling. You know the model can do amazing things, but you can’t get it to do what you want. It’s like having a super-smart assistant who keeps misunderstanding your instructions. Frustrating doesn’t even cover it.
Tom: They also mention that users are confused about how much detail to include. Do you give a one-sentence prompt or a full paragraph? And every time you start a new task, it feels like you’re starting from scratch because you don’t have a template that works.
Jane: That’s the "replicable template" problem we mentioned earlier. And the paper says this is a major barrier. People are spending hours on trial and error, just trying to get a decent response. That’s a huge waste of time and energy.
Tom: And the first page also introduces the core idea that prompts are like a language. We’re not just talking to the model; we’re learning to speak its language. And like any language, it has rules and patterns.
Jane: That’s a beautiful way to put it. And it makes the whole thing less intimidating. You’re not a programmer. You’re just learning a new way to communicate. And the paper is like a phrasebook for that language.
Tom: They also mention that the interaction is fundamentally different from human-to-human communication. With a person, you can rely on shared context and unspoken assumptions. With an LLM, you have to be much more explicit.
Jane: Right. You can’t just say "you know what I mean." The model doesn’t know. You have to spell it out. And that’s a skill. It’s the skill of being clear and precise, which is actually a great skill to have in general.
Tom: So the first page sets up the problem perfectly. It validates the frustration that so many users feel, and it promises a solution. And the rest of the paper delivers on that promise.
Jane: It does. And it does it in a way that’s accessible to everyone, not just AI researchers. That’s the real achievement here. They’ve taken a complex topic and made it usable.
Tom: And with that, we’re ready to wrap up. But before we go, let’s bring in Lu, Meng, and Lalam to get their take on the bigger picture.
Lu: Thanks, Tom. From a research perspective, I think the most exciting implication is that this survey could actually influence how we design the next generation of models. If we know how users prompt, we can train models to be more robust to different prompting styles.
Meng: And from an engineering standpoint, I love the idea of the living GitHub project. It means we can keep the advice current as the models evolve. That’s practical. That’s something we can actually use.
Lalam: And from my perspective, the cultural impact is significant. By making prompting accessible, we’re democratizing the ability to use AI effectively. That means more people can leverage these tools for education, for creativity, for problem-solving. It’s a step toward making AI a true collaborator, not just a tool.
Tom: Great points, all of you. And that brings us to our conclusion.
Conclusion: Tom: Well, we’ve reached the end of our discussion on "How to Ask the AI: A User Perspective Survey for Large Language Model Prompting." And I have to say, this paper feels like a gift to everyone who’s ever felt frustrated by a chatbot.
Jane: It really does. We’ve covered a lot of ground today. We talked about the taxonomy of prompting strategies, from direct to reasoning-based to self-improvement. We saw how the right prompt can turn a generic response into something tailored and useful.
Tom: And we saw how the wrong prompt, or an overly complex one, can actually break the model. That Christmas tree example was perfect. It showed that more isn’t always better.
Jane: Exactly. The paper’s biggest contribution is giving us a framework to think about prompting. It’s not about memorizing magic phrases. It’s about understanding what the model needs and giving it that.
Tom: And they’ve backed it up with real experiments and an evaluation system that uses multiple LLMs to judge the results. That gives us confidence that the advice is solid, not just one person’s opinion.
Jane: And the fact that they’re keeping it updated as a living project means we can come back to it as the technology changes. That’s forward-thinking.
Tom: So, as we say goodbye to this paper, I want to leave our listeners with one thought. The next time you’re stuck with a bad response from an LLM, don’t blame the model. Take a step back and think about your prompt. Are you being clear? Are you giving enough context? Are you using the right strategy for the task?
Jane: That’s the takeaway. And it’s a powerful one. This paper gives us the tools to be better communicators, not just with AI, but with ourselves. Because to ask a good question, you have to know what you want.
Tom: Well said, Jane. And with that, we’re ready to move on to our next paper. Thanks for listening, and we’ll see you next time.
Jane: Bye, everyone!
Yiqun Zhang, Yunfan Zhang, Mingjie Zhao, Sen Feng, Yiu-ming Cheung
Guangdong University of Technology · Hong Kong Baptist University
cs.HC, cs.AI
Submitted: 2026-06-17
Updated: 2026-08-11
Comments: Accepted to TAI
Code: https://github.com/Torantulino/Auto-GPT
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 43/100
The gist: This survey explores the principles, taxonomy, and organization of prompts from a user-centered perspective.
Key concepts
- Organizational Tasks
- These are simple tasks for an AI, such as asking for a summary, getting a translation, or having a conversation. For these types of requests, direct prompting is usually sufficient.
- Innovative Tasks
- These involve more complex requests where the model needs to reason, write stories, or generate code. For these tasks, users should provide more detail and break down the task into steps.
- Chain-of-Thought
- This is a prompting technique where you ask the model to show its work by explaining its reasoning process. This is useful for complex problems like math or planning a trip.
- Replicable Template
- The paper stresses the need for templates so users don't start from scratch every time. These templates provide structured ways to ask questions, improving consistency and reducing trial-and-error.
Terminology
Summary
This survey explores the principles, taxonomy, and organization of prompts from a user-centered perspective. Differing from the existing surveys that primarily focus on technical principles and application scenarios of LLMs, this paper provides actionable guidelines for formulating effective LLM prompts across diverse real-world tasks and specifically contributes by: 1) developing an intuitive evaluation strategy for prompt efficacy, 2) providing prompting workflow demonstrations on representative applications, and 3) maintaining a dynamically updated open-source project to ensure the core takeaways remain up-to-date. These measures lower the threshold for users to correctly understand and craft prompts that align with evolving application scenarios. This work will be maintained as a living GitHub project here.
The paper states that "AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses simply by typing requests, such as 'plan a three-day Vienna trip', 'solve the attached mathematical problem', 'draft an email to inquire review progress', etc., which are also known as LLM prompts. Crafting clear and well-structured prompts leads to more appropriate LLM feedback, which effectively bridges human-LLM interaction. Although prompting appears accessible to non-expert users, precisely organizing effective prompts is a highly systematic and skillful process, presenting potential challenges even for experienced users."
The paper identifies three typical dilemma scenarios for users: "1) Felt that LLMs are powerful but hard to control, always struggling to get LLMs to produce exactly what they need; 2) Confused by the required length and detail for prompts, as well as what kind of information must be included for LLMs to properly comprehend the intention; 3) Every time prompting with LLMs feels like starting from scratch due to the lack of a 'replicable template', with trial and error dominating the process."
The main contributions are summarized as four-fold: "Users centered narrative. This survey represents the first review of prompting techniques from a non-AI researcher's perspective. By systematically examining the role of prompts across diverse routine tasks, it lowers the entry barrier, enabling readers to quickly grasp prompting and leverage the power of LLMs." "Task-oriented taxonomy. Prompting strategies are categorized by task types: 1) organizational tasks: conversation, Q&A, language translation, and 2) innovative tasks: content generation, reasoning, and coding. This helps readers easily map their specific LLM usage needs to appropriate prompting strategies." "Intuitive demonstrations. An intuitive quantitative evaluation system and a set of practical task demonstrations are designed to showcase how different prompting styles perform in specific scenarios. The original prompts in the tasks are also demonstrated, closing the final replication gap in real-world applications." "Sustainability. The core content of this survey is designed to be continuously updated: Taxonomy, prompting templates, key examples, and case studies will be maintained as an open GitHub project, ensuring that the audience of this paper keeps up with timely and reliable prompting knowledge and guidance."
The paper introduces a taxonomy of prompting strategies, stating: Prompting strategies can be broadly categorized into several types according to their mechanisms and principles, which are summarized in Table I.
The categories include Direct Prompting (Zero-/Few-shot, Persona, Instruction Prompting), Reasoning-based (Prompt chaining, Chain-of-thought, Tree-of-thought, Least-to-most), Retrieval/Tool-Augmented (Retrieval-Augmented Generation (RAG), Reasoning and Acting (ReAct)), Ensemble-based (Context Calibration, Multi-template Voting), Self-improve (Self-consistency, Self-refine), Agent-based (AutoGPT, BabyAGI, LangChain Agents), Learnable Prompts (Prompt Tuning, Prefix Tuning), and Automated Construction (AutoPrompt, Generated Knowledge Prompting).
The paper further categorizes tasks into organizational and innovative tasks: "Organizational tasks can be grouped into three categories: general conversation, knowledge Q&A, and natural language assistance." Innovative tasks can also be grouped into three categories: reasoning and logical analysis, content generation, and code assistance.
A decision table (Table II) guides users to determine suitable prompting strategies for corresponding tasks.
The paper provides detailed prompting manuals for each task category. For example, for General Conversation, it recommends Zero-shot, Few-shot, Persona, Prompt chaining, and Multi-template Vote prompting. For Knowledge Q&A, it recommends Zero-shot, Few-shot, Retrieval-Augmented Generation, Reasoning and Acting, and Context Calibration. For Natural Language Assistance, it recommends Zero-shot, Few-shot, Chain-of-thought, and Self-refine. For Reasoning and Logical Analysis, it recommends Chain-of-thought, Tree-of-thought, Least-to-most, and Self-consistency. For Content Generation, it recommends Zero-shot, Few-shot, Instruction, and Least-to-most. For Code Assistance, it recommends Zero-shot, Few-shot, Chain-of-thought, and Self-refine.
The paper conducts a Cross-Task Prompts Rating Evaluation using an LLMs-as-Judges framework: this part lets the LLMs rate each other... the LLMs-as-Judges evaluation framework is utilized, i.e., using multiple large models to score each prompt.
Six LLMs (ChatGPT-5, Claude Sonnet 4, Gemini 2.5 Flash, Qwen3-Max, Grok 4 Fast, and DeepSeek-V3.1) assess four prompting categories (direct, reasoning-based, ensemble-based, and self-improve) from 1 to 10 across five user-oriented aspects: task performance, factuality, compliance, easy-using, and time cost. Key observations include: "In Fig. 4, the ensemble-based and self-improve prompting categories obtain a high score on organizational tasks in terms of task performance. The reasoning-based prompting strategy achieves the highest scores on innovative tasks in terms of task performance." As shown in Fig. 5, direct prompting obtains the overall highest scores, especially in terms of easy-using and time cost.
Prompting strategies also follow the 'no one-size-fits-all' principle. Each prompting strategy has its unique advantages, with no single dominating across all dimensions.
The paper presents five real-world demonstrations: Image Generation Evaluation, Code-Generated Christmas Tree Challenge, Daily Planning Evaluation, Story-Telling Challenge, and Email Drafting Demonstration. For image generation, the paper finds that "the evaluation scores exhibit a successive increase across the zero-shot, few-shot, and least-to-most promptings in both LLMs and humans' judgments, suggesting that employing reasoning-based prompting is more suitable for enhancing content generation performance. For the coding challenge, the paper observes that
Zero-shot prompting generally achieves the best performance across all results for the Christmas tree task, while for the LeetCode algorithmic challenge,
CoT demonstrates the best coding performance in terms of Lines and Acc, especially in hard problems. SR performs well in most cases in terms of accuracy, but suffers from the slowest feedback time and longest output. For daily planning, persona prompting
introduces empathetic and motivational content... making responses feel more encouraging and
contextualizes advice within realistic scenarios. For story-telling,
multi-template voting achieves a better compliance, as the multi-template prompts give LLMs more freedom in generation, enabling LLMs to select the best-suited feedback. For email drafting,
persona prompting offers the best trade-off between efficiency and effectiveness in this simple task."
The paper concludes: "This survey systematically examined prompting techniques for LLMs from users' perspectives. It categorizes prompting strategies by the organizational and creative task types, evaluates their effectiveness through a novel mutual evaluation mechanism to ensure quantification and fairness, and visually demonstrates the prompting process and feedback through a variety of case studies. Key findings reveal that: 1) direct prompting (e.g., zero-shot, persona-based) excels in simplicity and speed for routine tasks, while reasoning-based ones (e.g., chain-of-thought) enhance complex problem-solving; 2) context and constraints in prompts significantly improve output quality, as seen in the case study of storytelling and image generation applications; and 3) no one-size-fits-all approach exists, i.e., optimal prompting depends on the demands of tasks and the goals of users."
The paper also discusses future prospects: "1) Automatic prompts optimization based on user behavior can reduce trial-and-error. However, excessive automation could diminish user control and, in the long term, potentially impair the development of human cognitive abilities; 2) Multimodal prompts integration that combines text, voice, and visual prompts could unlock richer interactions. Nevertheless, biases and misinterpretations dominated by specific modalities will be difficult to eliminate due to the natural gaps between modalities; 3) Research into prompt engineering reveals the mechanisms of LLM-prompt interactions, supporting the crafting of customized and functionally diverse prompts to enhance LLM efficacy. However, this also lowers the barrier to 'jailbreaking' of LLMs, potentially leading to harmful outputs. Promising future research orientations include
1) trustworthy adaptive prompting models, 2) multimodal prompt fusion, and 3) prompt-based attack and defense techniques."
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:
1. Implement a Task-Aware Prompt Strategy Selector
-
Improvement: Build a routing layer that classifies incoming user requests into organizational (conversation, Q&A, language assistance) or innovative (reasoning, content generation, coding) tasks, then automatically applies the optimal prompting strategy (e.g., zero-shot for simple Q&A, chain-of-thought for complex reasoning, least-to-most for multi-step generation).
-
What it can do: The system will no longer rely on users to know which prompting technique to use. It will automatically choose the best approach, improving output quality and reducing user trial-and-error.
2. Add a Persona and Context Injection Module
-
Improvement: Before generating a response, the system will detect the user’s implicit needs (e.g., tone, expertise level, urgency) and inject relevant persona or context cues into the prompt (e.g., “You are a supportive career coach,” “The user is a beginner in Python”).
-
What it can do: The system will produce more tailored, empathetic, and instruction-following responses, especially in organizational tasks like email drafting and daily planning, where the paper shows persona prompting significantly improves compliance and user engagement.
3. Integrate a Self-Refine Loop for High-Stakes Outputs
-
Improvement: For tasks involving code generation, formal writing, or factual claims, the system will automatically run a two-pass process: first generate a draft, then review and refine it (as demonstrated in the paper’s self-refine prompting examples).
-
What it can do: The system will catch edge cases, improve code efficiency, and polish language, reducing errors in critical applications like debugging, academic writing, or medical advice.
4. Deploy a Multi-Template Voting Mechanism for Creative Tasks
-
Improvement: For content generation (e.g., stories, marketing copy, image prompts), the system will generate multiple variations of the same prompt, produce several candidate outputs, and select the most consistent or highest-quality response based on internal scoring.
-
What it can do: The system will deliver more creative, coherent, and detailed outputs, as shown in the paper’s story-telling challenge where multi-template voting outperformed zero-shot in narrative depth and compliance.
5. Add a Confidence and Calibration Indicator
-
Improvement: The system will append a confidence score (high/medium/low) to knowledge Q&A responses, based on internal consistency checks and retrieval grounding (as suggested by context calibration prompting).
-
What it can do: Users will be able to assess the reliability of answers, especially in domains like finance or health, reducing the risk of acting on hallucinated information.
6. Implement a Dynamic Difficulty-Aware Coding Assistant
-
Improvement: The system will classify coding problems by difficulty (easy/medium/hard) and automatically switch between zero-shot (for simple tasks), chain-of-thought (for medium/hard), and self-refine (for edge-case-heavy problems), based on the paper’s LeetCode evaluation results.
-
What it can do: The system will provide faster, more accurate code for routine tasks while maintaining high accuracy on complex algorithmic challenges, matching or exceeding human performance on hard problems.
7. Add a Privacy-Aware Prompting Mode
-
Improvement: The system will detect when a user’s input contains sensitive data (e.g., personal health, financial, or proprietary information) and automatically switch to a local or anonymized processing mode, or warn the user about the risks of retrieval-augmented prompting.
-
What it can do: The system will protect user privacy and prevent data leakage, addressing the critical limitation of cloud-based RAG and tool-augmented prompting highlighted in the paper.
-
Automatic adaptation to task type, user expertise, and data sensitivity.
-
Higher output quality in reasoning, coding, and creative tasks through dynamic strategy selection.
-
Reduced user effort by eliminating the need for manual prompt engineering.
-
Improved trust through confidence indicators and privacy safeguards.
-
Faster, more reliable performance across routine and complex tasks, validated by the paper’s quantitative evaluations (e.g., 100% accuracy on medium LeetCode problems with chain-of-thought, and higher human-judged scores for least-to-most image generation).
Sources
- GPT-4 Technical Report
- A Survey of Large Language Models
- Large Language Models: A Survey
- Advancing AI Research Assistants with Expert-Involved Learning
- A Survey of Prompt Engineering Methods in Large Language Models for Different NLP Tasks
- An Empirical Categorization of Prompting Techniques for Large Language Models: A Practitioner's Guide
- A Brief History of Prompt: Leveraging Language Models. (Through Advanced Prompting)
- A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications
- What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering
- Evaluating Large Language Models Trained on Code
- A Survey of Uncertainty Estimation in LLMs: Theory Meets Practice
- Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities
- Efficient Prompting Methods for Large Language Models: A Survey
- A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
- A Survey on Large Language Model Benchmarks
- Prompt Design and Engineering: Introduction and Advanced Methods
- LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
- Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
- Reframing Instructional Prompts to GPTk's Language
- On the Opportunities and Risks of Foundation Models
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support