LLARS: Enabling Domain Expert & Developer Collaboration for LLM Prompting, Generation and Evaluation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "LLARS: Enabling Domain Expert & Developer Collaboration for LLM Prompting, Generation and Evaluation".
Tom: LLARS (LLM Assisted Research System) is an open-source platform designed to bridge the gap between domain experts and developers by unifying prompt engineering, batch generation,
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, moving on from the overview, let’s talk about what LLARS actually claims this system is capable of doing based on the paper. The central thesis is that LLARS provides an open-source platform that bridges the gap between domain experts who have deep knowledge and developers who build the technical systems needed to work with Large Language Models.
Jane: That’s right, Tom, and it claims this system achieves this by unifying three tightly connected modules into one end-to-end pipeline: Collaborative Prompt Engineering, Batch Generation, and Hybrid Evaluation. The authors argue that this integration is what allows for rigorous iteration in building LLM-based systems where both deep domain expertise and technical skill are required.
Lu: They are proposing a system where the prompt creation process flows directly into batch generation, and those generated batches then become evaluation scenarios with just one click. This structure ensures that there's a clear path of provenance maintained throughout the entire process, which is what they’re aiming for.
Meng: Essentially, they claim that this unified pipeline allows users to move from ideating a prompt to scaling it across many outputs and finally assessing those outputs using both human and AI raters in a structured way. This sounds like a streamlined approach compared to juggling separate tools for each stage of the process.
Lalam: The paper emphasizes that this unification is what makes the platform matter because it tackles the practical need for collaborative iteration in sensitive domains where practitioners aren't always sharing language or tools between their respective disciplines.
Tom: So, if we boil it down to what they’re claiming, LLARS is a system that combines real-time co-authoring for prompts, configurable output production across prompts and models and data, and structured evaluation with live metrics to help you find the best model-prompt combination.
Jane: That's the essence of it. It focuses on giving users a single environment where they can test their ideas systematically, generate many results efficiently, and then rigorously evaluate those results using a hybrid approach involving both human and AI participants.
Lu: The paper suggests that this end-to-end pipeline is designed to provide the necessary structure for building trustworthy LLM-based systems by ensuring that every single output has a complete record of its history.
Meng: From an engineering viewpoint, the claim is that this system manages complexity by handling the different aspects—prompt design, scaling, and testing—as one cohesive unit rather than three separate tools operating in silos.
Lalam: It's about creating a verifiable workflow where the output isn't just something that exists, but something whose entire creation history is documented for review. That documentation aspect is crucial for building trust in AI systems that people actually rely on.
Tom: So, to summarize this section of the paper, LLARS claims it offers a complete system for domain experts and developers to collaborate efficiently across the prompt design, generation scaling, and evaluation stages using integrated features like version control and hybrid testing scenarios.
Jane: And they emphasize that this structure is what matters because it directly addresses the need for rigorous iteration in complex domains where expertise needs to be combined with technical implementation skills.
Conclusion: Tom: Looking at the title of "LLARS: Enabling Domain Expert and Developer Collaboration for LLM Prompting, Generation and Evaluation," it really captures the essence of what this platform is trying to achieve—it’s about making sure these two groups can work together effectively on building AI systems.
Jane: It does a good job because it points directly to the core function: enabling collaboration between domain experts and developers in the context of LLM prompting, generation, and evaluation. It highlights the key areas where this synergy is most needed.
Lu: I think it’s important that they are emphasizing collaboration rather than just focusing on automation; they want to show that this isn't just about replacing one person with a machine, but about augmenting the work of both sides.
Meng: That makes sense because when you look at real-world application, the paper suggests that the value comes from having those two different perspectives contributing their unique strengths in a way that neither could achieve alone.
Lalam: I think it speaks to a bigger cultural shift we need to make—one where the expertise of practitioners is actively integrated into the technical build process, not just consulted after everything is already done.
Tom: So, what does this mean practically? In simple terms, LLARS suggests that instead of developers writing prompts in isolation and then handing them off for testing, they should be working with the people who know the domain from the very beginning.
Jane: That’s right; it means moving away from a process where translation happens between disciplines and instead having everyone operating within a single workspace where they can iterate together on what is being built.
Lu: It implies that the platform is helping to create a shared language for building these AI tools, which simplifies the communication barrier significantly because you don't have to constantly re-explain concepts to each other.
Meng: And from an engineering perspective, that simplified communication translates into faster development cycles because less time is spent on clarifying assumptions and more time is spent actually building the working system.
Lalam: It suggests a future where AI systems are built with more inherent domain knowledge baked in, rather than being bolted on later as an afterthought.
Tom: So, the implication here is that this platform supports a way of working where building complex AI tools becomes more iterative and less fragmented because the different expert roles can operate in close partnership within one integrated system.
Jane: And ultimately, it’s about making sure that the people who understand *what* needs to be built are deeply involved in *how* it gets built through every step of the process.
Technische Hochschule Nurnberg Georg Simon Ohm
cs.AI, cs.CL, cs.HC, cs.SE
Submitted: 2026-05-11
Updated: 2026-05-11
Code: https://github.com/th-nuernberg/llars
Importance score: 77/100
The gist: LLARS (LLM Assisted Research System) is an open-source platform designed to bridge the gap between domain experts and developers by unifying prompt engineering, batch generation, and hybrid
Key concepts
- Collaborative Prompt Engineering
- This module lets domain experts and developers co-author prompts in a shared editor. It supports real-time editing, version history tracking, and using template variables to substitute data from a shared palette into prompts. This ensures that the prompt creation process is a joint effort between technical skills and deep domain knowledge.
- Batch Generation
- This feature combines any set of prompts, models, and data items into their full combinations (Cartesian product). It estimates costs for these combinations and generates hundreds of outputs from a single prompt. Every output is tagged with full provenance, tracking the exact source item and parameters used.
- Hybrid Evaluation
- Evaluation is structured as a scenario where human evaluators and LLM evaluators assess randomized results. The system supports various rating methods, allowing users to inspect aggregated results in real time, including automatically computed reliability metrics between the two groups.
Terminology
Summary
LLARS (LLM Assisted Research System) is an open-source platform designed to bridge the gap between domain experts and developers by unifying prompt engineering, batch generation, and hybrid evaluation into a single system. This platform matters because it addresses the need for rigorous, collaborative iteration in building LLM-based systems for sensitive domains where practitioners must balance deep domain expertise with technical skill.
The gist
LLARS unifies collaborative prompt engineering, batch generation and hybrid human–LLM evaluation with agreement analytics in a single open-source platform.
Platform Overview and Integration
LLARS runs as a containerised web application for textual data, integrating three tightly connected modules: Collaborative Prompt Engineering, Batch Generation, and Hybrid Evaluation. The system is co-designed with developers, social science researchers, and practicing counsellors to ensure seamless interdisciplinary collaboration. The modules are tightly integrated so that prompts created in the editor flow directly into batch generation and completed batches become evaluation scenarios with a single click.
This structure ensures that provenance is maintained throughout the entire pipeline.
Collaborative Prompt Engineering
This module allows domain experts and developers to co-author prompts in a shared real-time editor. Key features include:
-
Every keystroke appearing instantly for all connected users, enabling
diff comparison and rollback without affecting other blocks
within version history that tracks insertions and deletions. -
The ability to use template variables (e.g.,
a single click sends the assembled prompt with sample values substituted to a selected model
) as placeholders for external data, which are collected in a shared Variable Palette annotated with sample values. -
Prompts can be exported as JSON and are directly available in the Batch Generation module, where template variables are filled from uploaded data to
scale a single prompt to hundreds of outputs.
Batch Generation
Batch generation extends testing by combining any set of prompts, models, and data items into their Cartesian product (prompts × models × data items).
This process includes cost estimation and budget control through a generation matrix that previews all combinations. For instance, 50 email threads with two prompts and two models produce 200 outputs. Each result is tagged with full provenance (source item, prompt version, model, parameters, tokens and cost)
and is exportable as CSV or JSON. Crucially, Completed batches become evaluation scenarios in one click,
preserving attribution end-to-end.
Hybrid Evaluation
Evaluation is organized as a scenario—a self-contained campaign defining the evaluation type and assigning evaluators. Items are presented in randomized order without provenance information so evaluators cannot identify the source. LLARS supports multi-dimensional rating with configurable Likert scales, ranking into ordinal buckets or traditional ranking, categorical labelling, pairwise comparison, mail assessment and authenticity detection.
-
LLM evaluators are treated as
full participants alongside humans,
receiving the same items and using the same setup. -
Scenario owners can inspect aggregated results in real time, including
automatically computed inter-rater reliability (e.g., Krippendorff’s α),
filtered by human evaluators only, LLM evaluators only, or both combined. -
Provenance analysis leverages each item’s generating model and prompt to display the
top-bucket hit rate and full bucket distribution per model–prompt pair,
surfacing thebest performer.
Conclusion and Impact
The platform has been actively used in online counselling research, such as the Virtual Client project, where it helped determine how small a model could be while still meeting quality requirements.
Interviews confirmed that centralizing all stages in one workspace saves considerable time and makes interdisciplinary collaboration seamless,
eliminating the translation work
between disciplines. Future work includes extending LLARS to multi-turn conversational evaluation and automated calibration of LLM evaluators against human ratings to surface systematic biases in real time. Ultimately, LLARS bridges the gap between domain expertise and technical implementation by enabling interdisciplinary work through a unified system.
How it works
LLARS is structured around three modules that are tightly integrated.
The flow is sequential: prompts are co-authored collaboratively, these prompts feed into batch generation to create outputs across various models and data, and finally, these completed batches are used to define evaluation scenarios where human and LLM evaluators jointly assess the results. This end-to-end pipeline ensures that every output is traceable back to its originating prompt version, model choice, and input data item.
Key Features Summary
The platform addresses limitations found in existing tools by offering:
(1) Collaborative Prompt Engineering:
(2) Batch Generation:
(3) Hybrid Evaluation:
This unified approach allows users to move from initial prompt ideation, through systematic output scaling, to rigorous, multi-rater assessment and final model selection within a single environment. This closes the loop "from prompt development to rigorous assessment.
Improvements for AI systems
Here are the specific improvements that LLARS enables for AI systems, based on the provided research:
-
Prompts Engineering Workflow: Enable domain experts and developers to co-author prompts in a real-time, version-controlled environment with instant LLM testing. This allows for rapid iteration of complex prompts in sensitive domains (e.g., online counselling, legal analysis) by comparing prompt versions and rolling back changes without disrupting the main workflow.
-
Batch Generation Scalability: Implement configurable batch generation across the Cartesian product of prompts × models × data items, complete with cost estimation and budget control. This allows for systematically generating hundreds or thousands of diverse outputs (e.g., different email threads or legal documents) from a single prompt using various models while maintaining strict financial oversight.
-
Hybrid Evaluation Rigor: Establish structured, multi-dimensional evaluation campaigns where human and LLM evaluators jointly assess outputs using configurable Likert scales, ranking, and categorical labeling. This includes automated calculation of inter-rater reliability (e.g., Krippendorff’s α) to ensure assessment consistency across different modalities (human vs. LLM).
-
Automated Model Selection: Utilize provenance analysis to surface the optimal model–prompt combination for a given use case by tracking and analyzing the performance metrics derived from hybrid evaluations, identifying which specific model configuration best meets quality requirements for the task.
-
Seamless Interdisciplinary Collaboration: Create a unified workspace where prompts, data, and outputs reside in one place. This eliminates
translation work
between disciplines (e.g., domain experts and developers), allowing for seamless handover of tasks from prompt authoring to rigorous assessment without reliance on scattered external tools or document exchanges. -
Single-Click Scenario Conversion: Convert completed batches into formal evaluation scenarios with a single click, preserving the entire provenance chain (source item, prompt version, model used) for easy re-testing and auditability.
Sources
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection