FollowUpBot: An LLM-Based Conversational Robot for Automatic Postoperative Follow-up
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "FollowUpBot: An LLM-Based Conversational Robot for Automatic Postoperative Follow-up".
Jane: The paper was written by Chen Chen, Jianing Yin, Jiannong Cao, Zhiyuan Wen, Mingjin Zhang et al. from The Hong Kong Polytechnic University and Guangdong Provincial People's Hospital.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Alright, welcome back to the show, everyone. Today we're looking at a paper that's got me genuinely excited — it's called "FollowUpBot: An LLM-Based Conversational Robot for Automatic Postoperative Follow-up." Jane, I know you've been waiting to dig into this one.
Jane: Oh, absolutely, Tom. And I think the title alone tells you a lot. You've got a robot, you've got a large language model, and you've got postoperative care all in one system. That's three big pieces of technology coming together for something that happens in hospitals every single day.
Tom: Right, and let's be real — postoperative follow-up is one of those things that sounds simple but is actually a huge burden on nurses. They have to walk around, visit patients at the bedside, ask about symptoms, write everything down by hand. It's exhausting and it's error-prone.
Jane: Exactly. And the paper points out that with more surgeries happening and fewer clinical staff available, this is becoming a real crisis. So the idea here is to build a robot that can do this automatically, but in a way that's actually thoughtful and adaptive, not just a scripted questionnaire on wheels.
Tom: And that's the part that got me — the authors are from Hong Kong Polytechnic University and Guangdong Provincial People's Hospital. So this isn't just a lab project. They actually deployed it in a real hospital. That's a big deal.
Jane: It really is. And the name FollowUpBot is perfect because it does exactly what it says. It follows up with patients after surgery, asks them about their symptoms, and then generates a structured report for the hospital. All of that in one robot.
Tom: So the big question for me is — why do we need a physical robot at all? Why not just send patients a text message or have them fill out an app?
Jane: That's a great question, and the paper actually addresses that head-on. Digital tools like apps and wearables exist, but they have real problems. Patients, especially older ones, don't always engage with them. And there's the privacy issue — a lot of these systems rely on cloud services, which means patient data is being sent somewhere off-site.
Tom: And that's a huge concern in healthcare. You can't just have patient conversations floating around on some external server.
Jane: Right. So FollowUpBot runs its language model locally, on edge devices. That means the conversation stays in the hospital. And because it's a physical robot, it's face-to-face with the patient, which is more engaging and more human.
Tom: I love that. So we're getting the best of both worlds — the efficiency of automation and the warmth of an actual visit. I can already see why this could change how hospitals operate.
Jane: And that's just the beginning. The paper goes into how the robot actually conducts these conversations, and that's where things get really interesting.
Tom: Then let's get into it. But first, for anyone just tuning in, we're talking about "FollowUpBot: An LLM-Based Conversational Robot for Automatic Postoperative Follow-up." Stick around — there's a lot more to unpack.
Summary: Tom: So Jane, we've set the stage. Let's talk about what FollowUpBot actually does, step by step. The paper describes three core modules working together.
Jane: Right. The first is automatic navigation. The robot maps the hospital ward using SLAM — that's simultaneous localization and mapping — so it knows where beds are, where walls are, where it's allowed to go. When it gets a task, it plans a route to the patient's bedside and avoids obstacles along the way.
Tom: So it's basically a self-driving robot in a hospital hallway. That's already impressive on its own.
Jane: It is, but the second module is where the magic happens. Once the robot arrives, it starts a conversation with the patient using a locally deployed medical language model called WiNGPT2. And here's the key — it's not just reading from a script. It adapts based on the patient's profile and their responses.
Tom: And patients can interact in different ways, right? Speech, touch, text?
Jane: Exactly. The robot asks about specific follow-up fields — things like headache, dizziness, nausea — and it tracks which ones are done and which ones still need to be covered. The patient can reply by talking, by tapping options on a screen, or by typing. That's important because after surgery, some patients might be too weak to talk, or they might prefer touch.
Tom: That's thoughtful design. And then the third module takes all that conversation and turns it into a structured report.
Jane: Yes. The robot uses another language model to extract the relevant information from the dialogue and fill in a hospital-approved template. And for fields that need strict formats — like yes/no questions or numerical values — there's a verification step using natural language inference to make sure the answer is mapped correctly.
Tom: So it's not just copying what the patient said. It's actually understanding and normalizing the response.
Jane: Precisely. If a patient says "I've been feeling a bit off in the head," the system needs to figure out whether that maps to "headache: yes" or "headache: no" or "unclear." That's where the NLI model comes in — it checks whether the patient's words entail one of the predefined options.
Tom: And all of this happens locally, on edge devices, so the patient's data never leaves the hospital. That's a massive privacy win.
Jane: It really is. And the whole workflow is designed to be practical. The robot gets a task from the hospital's operating room information system, navigates to the bed, has the conversation, generates the report, and then moves on to the next patient or returns to standby.
Tom: So it's a complete closed loop. I mean, the engineering here is substantial, but the user experience is what stands out to me. The patient just sees a friendly robot asking how they're feeling.
Jane: And that's the goal. The paper even includes a demonstration video, which I highly recommend checking out. It shows the robot in action, and it's surprisingly natural.
Tom: I've seen it, and yeah, it's compelling. But the real proof is in the evaluation. And that's where the numbers get interesting.
Jane: They built a synthetic dataset of one hundred postoperative cases and tested the robot against a baseline. The results are pretty striking, and I think that's where we should go next.
Tom: Agreed. Let's dig into the experiments and what they found.
Improvements: Tom: So Jane, we've covered what FollowUpBot does. Now let's talk about how well it actually works. The paper ran some pretty thorough evaluations.
Jane: They did. First, they looked at follow-up interaction quality. They used GPT-4o to simulate patients, which is clever — you can test the robot against realistic patient behavior without needing real patients for every scenario.
Tom: And the results? The robot achieved one hundred percent symptom coverage, meaning it addressed every clinically required symptom in the dialogue. The baseline, which was just the language model without the structured tracking, only covered fifty-three point eight percent.
Jane: That's a huge gap. And it makes sense — without explicit field tracking, the conversation can drift. The robot might spend too long on one topic and never get to the others. But with the tracking mechanism, it systematically works through everything.
Tom: And they also measured patient satisfaction. The simulated patients rated the robot higher than the baseline across most dimensions — things like empathy, clarity, and thoroughness.
Jane: Right. So it's not just covering more topics; it's doing it in a way that feels better to the patient. That's important because if patients don't feel comfortable, they won't open up about their symptoms.
Tom: Then they did ablation studies on the report generation. They tested three components: field-specific descriptions, NLI-based option matching, and explicit field tracking.
Jane: And the results are really telling. Without the NLI verification, the accuracy on single-choice fields was only about eighteen percent. That's terrible. But with NLI, it jumped to over eighty-two percent. And with the full system — including field tracking — it hit over ninety-one percent accuracy.
Tom: So the NLI step is doing a lot of heavy lifting. It's taking free-form patient responses and correctly mapping them to the predefined options.
Jane: Exactly. And for numerical fields, the improvement is even more dramatic. Without NLI, the mean absolute error was about one point five eight. With the full system, it dropped to zero point zero three. That's essentially perfect.
Tom: Wow. So when a patient says "my pain is around a seven," the robot can accurately extract that as a number and put it in the report.
Jane: Precisely. And text fields — like descriptions of symptoms — also improved, though the gains were more modest. The focused dialogue from field tracking helps the model extract cleaner text.
Tom: So the modular approach is really paying off. Each piece — the tracking, the verification, the descriptions — contributes something meaningful.
Jane: And that's the key insight. You can't just throw a language model at this problem and expect it to work. You need structure around it. The paper shows that clearly.
Tom: Now, I know we have some folks listening who are thinking about the real-world deployment. What about the practical side?
Jane: That's actually where I want to bring in Meng and Lu. They'll have different perspectives on what this means for real hospitals and for the broader field.
Tom: Good call. Let's hear from them.
Conclusion: Tom: Alright, we're wrapping up our discussion of "FollowUpBot: An LLM-Based Conversational Robot for Automatic Postoperative Follow-up." Jane, this has been a fascinating one.
Jane: It really has, Tom. And I think the biggest takeaway is that this isn't just a research toy. The authors deployed it in a real hospital in Guangdong, and they've shown that it works — both in terms of interaction quality and report accuracy.
Tom: And the implications are huge. Think about the nursing shortage we mentioned earlier. A robot like this could free up nurses to focus on the patients who actually need their attention, while the routine follow-ups happen automatically.
Jane: Exactly. And because the reports are structured and accurate, doctors can review them quickly without having to parse handwritten notes or messy transcripts.
Tom: I want to bring in Lu and Meng for their final thoughts. Lu, what excites you most about where this could go?
Lu: I think the privacy-preserving aspect is the real game-changer. Running the language model locally on edge devices means this could be deployed in hospitals that are hesitant to use cloud-based AI because of compliance issues. That opens up a lot of doors.
Meng: And from an engineering standpoint, I'm impressed that they got this to work reliably. Navigation, multimodal interaction, report generation — that's a lot of moving parts. The fact that they achieved one hundred percent coverage and near-perfect numerical accuracy in testing is genuinely impressive.
Tom: And Lalam, any final thoughts on the broader impact?
Lalam: This paper shows how AI can be embedded into clinical workflows in a way that respects both patient privacy and human dignity. The robot doesn't replace the nurse's judgment — it handles the routine, structured parts of follow-up, which lets clinicians focus on the human connection. That's a model for how AI should integrate into healthcare.
Jane: Beautifully said. And with that, we're saying goodbye to FollowUpBot. It's a paper that I think we'll look back on as an early example of how to do hospital automation right.
Tom: Absolutely. Thanks for joining us, everyone. Next time, we'll be looking at another exciting paper from the arXiv. Until then, take care.
Jane: Bye, everyone.
Chen Chen, Jianing Yin, Jiannong Cao, Zhiyuan Wen, Mingjin Zhang, Weixun Gao, Xiang Wang, Haihua Shu
The Hong Kong Polytechnic University · Guangdong Provincial People's Hospital
cs.HC, cs.CL, cs.RO
Submitted: 2025-07-21
Updated: 2026-08-18
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 85/100
Key concepts
- LLM-Based Conversational Robot
- FollowUpBot is an AI robot that uses a large language model to have adaptive, face-to-face conversations with patients after surgery. It is designed to be thoughtful and responsive rather than just following a fixed script, aiming to automate postoperative follow-up tasks.
- Local Deployment on Edge Devices
- The system runs its language model locally on the robot itself rather than in the cloud. This ensures that patient data stays within the hospital, addressing privacy concerns and making the system suitable for environments hesitant about sending sensitive information externally.
- Natural Language Inference (NLI)
- This component of the system checks if a patient's free-form response entails one of several predefined options. It is crucial for accurately mapping unstructured patient language, like 'feeling a bit off in the head,' to structured data fields such as 'headache: yes' or 'no.'
- Modular Approach
- FollowUpBot uses three distinct modules working together: automatic navigation (SLAM), adaptive conversation (WiNGPT2), and structured report generation. This modular design allows each part to contribute specific functionality, leading to high overall accuracy in symptom coverage and data extraction.
Terminology
Summary
Summary
This paper introduces FollowUpBot, an LLM-powered edge-deployed robot designed to automate postoperative follow-up in hospital settings. The authors identify that traditional postoperative follow-up, which relies on nurses performing bedside interviews and manual documentation, is time-consuming and labor-intensive, and existing digital solutions (web questionnaires, automated calls) suffer from inflexible scripted interactions or privacy leakage issues due to cloud reliance.
FollowUpBot is an end-to-end robotic system composed of three core modules: (1) Automatic Navigation Module, (2) Adaptive and Privacy-Preserving Follow-up Module, and (3) Automatic Report Generation Module. The robot is equipped with an RGB camera, touchscreen, audio I/O, multimodal sensors (LiDAR, RGB-D camera), and dual edge devices for local model real-time inference and control.
The Automatic Navigation Module performs online SLAM by fusing LiDAR, RGB-D, and IMU data to build a semantic 3D map of the hospital ward. Upon receiving a task from the Operation Room Information System (ORIS), it computes a global path using A*-based search on a topological graph, with costs adjusted by distance and environmental risk. Locally, a Model Predictive Controller (MPC) with a reinforcement learning policy handles trajectory execution and avoids dynamic obstacles.
The Adaptive and Privacy-Preserving Follow-up Module initiates a structured dialogue workflow upon arrival at the bedside. It maintains a prioritized list of follow-up fields (e.g., headache, dizziness, nausea) and queries each field sequentially using a locally deployed medical LLM (WiNGPT2-Llama-3-8B-Chat). The dialogue is dynamically adapted based on patient profile and real-time responses, and patients can interact through speech (transcribed using whisper-large-v3), touch options, or text input. The LLM replies are rendered on-screen and via speech synthesis.
The Automatic Report Generation Module uses a report LLM (Llama-3.1-8B) to extract field values from dialogue content. For fields requiring strict formats (e.g., single-choice fields, numerical fields), a Natural Language Inference (NLI) module (nli-deberta-v3-base) is used to map free-form LLM outputs to the closest valid options by computing entailment scores. The option with the highest score is selected as the final report value, ensuring semantic correctness and format consistency. Once all required fields are filled, a structured report is generated according to the hospital’s template and saved locally.
The authors evaluated FollowUpBot on a synthetic dataset of 100 postoperative cases using GPT-4o. For follow-up interaction quality, they simulated patient interactions using GPT-4o and compared against a prompting-only baseline (WiNGPT2). The robot achieved 100% symptom coverage, while WiNGPT2 only covered 53.8%. Satisfaction was assessed across six aspects on a 5-point Likert scale, with the robot outperforming the baseline in most dimensions.
For report generation accuracy, ablation studies were conducted on three components: field-specific descriptions, NLI-based option matching, and explicit field tracking. Results showed that NLI alignment significantly improves accuracy, especially for structured fields. Field tracking further enhances performance, achieving 91.44% accuracy and 0.9912 BERTScore F1 on single-choice fields, and 99.20% accuracy with a 0.0300 MAE on numerical fields. Text fields also benefit from more focused dialogue inputs, resulting in moderate gains (0.8512 BERTScore F1).
The robot was deployed and tested in Guangdong Provincial People’s Hospital, where it successfully navigated real inpatient wards and completed automatic follow-up with real patients, demonstrating clinical feasibility. The paper claims FollowUpBot is the first postoperative follow-up robot that integrates navigation, interaction, and report generation modules, and it enables personalized and multimodal follow-up conversations that adapt dynamically to clinical conditions and input preferences of patients.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to an AI system, and what the improved system can do:
-
Improvement: Replace cloud-dependent LLM inference with local edge deployment using WiNGPT2-Llama-3-8B-Chat and Llama-3.1-8B, running on dual edge devices.
-
Capability: The system can conduct full clinical conversations and generate reports entirely offline, eliminating data leakage risks and reducing latency, while maintaining compliance with hospital privacy regulations.
-
Improvement: Implement a prioritized field-tracking mechanism that maintains a list of required follow-up fields (e.g., headache, dizziness, nausea), each with type, description, and completion status. The system dynamically selects the next unfilled field and guides the LLM to elicit responses specific to that field.
-
Capability: The system achieves 100% symptom coverage (vs. 53.8% for baseline), ensuring no clinically required question is missed, and produces more focused dialogues that improve downstream report accuracy.
-
Improvement: Add a cross-encoder NLI postprocessor (nli-deberta-v3-base) that maps free-form LLM outputs to predefined options by computing entailment scores, selecting the highest-scoring valid option.
-
Capability: Single-choice field accuracy improves from 18.48% (description-only) to 91.44%, and numerical field accuracy improves from 52.74% to 99.20%, with MAE dropping from 1.58 to 0.03. This ensures format consistency and semantic correctness in structured reports.
-
Improvement: Use a dedicated report LLM that extracts field values from dialogue content, guided by explicit field descriptions and the tracking mechanism, then generates reports following hospital templates.
-
Capability: Text field F1 improves from 0.7290 to 0.8513, and the system produces standardized, interpretable reports that integrate patient info, vital signs, and complications into a unified view, ready for EHR integration.
-
Improvement: Support three input modalities (speech via whisper-large-v3, touch options, and text) and dynamically adapt conversation style based on patient profile (age, surgery type) and real-time responses.
-
Capability: The system accommodates patients with varying postoperative physical/cognitive abilities, improving engagement and satisfaction (higher scores across all six satisfaction dimensions vs. baseline).
-
Improvement: Integrate SLAM-based 3D mapping, A* global path planning with risk-adjusted costs, and MPC with reinforcement learning for local trajectory execution.
-
Capability: The robot safely navigates hospital wards to patient bedsides, slowing down near patients and confirming arrival via LiDAR before initiating interaction, enabling fully autonomous ward rounds.
-
Autonomously perform complete postoperative ward rounds: Navigate to patients, conduct adaptive multimodal follow-up conversations, and generate structured clinical reports—all without cloud dependency.
-
Guarantee complete clinical data collection: Achieve 100% coverage of required symptoms through field tracking, eliminating human error from missed questions.
-
Produce clinically accurate, format-compliant reports: With 91.44% accuracy on choice fields, 99.20% on numerical fields, and 0.8513 F1 on text fields, the system outputs reports ready for direct EHR integration.
-
Operate in privacy-sensitive environments: All inference runs locally on edge devices, making it suitable for hospitals with strict data compliance requirements.
-
Adapt to diverse patient populations: Supports speech, touch, and text input, and personalizes conversation based on patient profile, improving engagement across age groups and physical conditions.
-
Scale to multiple patients efficiently: The closed-loop workflow checks for remaining tasks and returns to standby, enabling continuous operation across multiple follow-up assignments.
Sources
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support