Not All Nudges Land: Behavioral Controllability and Elaboration Quality in AI-Supported Journaling
Nadia Mehjabin, Henry Kautz, Subigya Nepal
University of Virginia
cs.HC, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-14
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 48/100
The gist: This paper analyzes 369 journal entries from an eight-week passive sensing study (the MindScape system) to determine which behaviors respond to AI-generated journaling nudges and what signals in
Terminology
Summary
This paper analyzes 369 journal entries from an eight-week passive sensing study (the MindScape system) to determine which behaviors respond to AI-generated journaling nudges and what signals in users' writing predict follow-through. The authors used an LLM-based two-stage classification pipeline to label entries as expressing an intention to change a behavior or not, then measured follow-through against 26 sensor features using a 3-day before/after comparison.
The central finding is that responsiveness depended most on whether a behavior involves other people.
Specifically, Behaviors that depend on others improved in only 15 to 22% of cases, while behaviors a person can act on alone improved more often, up to 50 to 63%, though unevenly.
The paper states: The low end is the most consistent part of this pattern: behaviors that require other people's coordination, such as calls, food-venue conversations and Greek-space visits, were uniformly among the least responsive.
For example, phone call features (25 to 31%), food venue conversations (22%), and Greek space visits (15%)
rarely improved, while incoming SMS (62.9%), walking episodes (58.0%), and time spent at home (53.6%) improved in over half of cases.
The paper emphasizes that individual controllability alone does not predict responsiveness: "The high end is noisier, and individual controllability did not by itself predict improvement. Among behaviors a person performs alone, walking improved in 58.0% of cases but gym/workout time in 20.0%, entertainment-app use in 18.8%, and cycling in 12.5% — near or below the socially dependent behaviors. The authors conclude:
We therefore read this result less as a clean controllability gradient than as a single robust asymmetry: behaviors that depend on other people are consistently unresponsive to journaling nudges, whereas behaviors under individual control vary widely and are not explained by controllability alone."
Regarding the role of writing quality, the paper reports that No single text feature separated improved from unimproved entries; writing carried signal only within specific behaviors, most clearly for text messaging and for longer, more personal intention entries.
At the aggregate level, "no single text feature predicted behavioral improvement (all p > 0.10), and intention and non-intention entries did not differ between the improved and unimproved groups (χ2 p = 0.15)."
However, within specific behaviors, intention language mattered. For incoming SMS, entries classified as INTENTION were followed by improvement in 100% of cases (8 of 8), compared to 52.2% for NON-INTENTION entries (Fisher's exact p = 0.028).
The authors note this rests on very few entries and should be read with caution
but that SMS is a plausible case for this association because it is immediately actionable: deciding to reach out to someone and sending a message can happen in the same moment as writing a journal entry.
Within intention entries specifically, elaboration quality predicted follow-through. The paper states: "Improved intention entries were longer (median word count 45 vs. 25.5, Mann-Whitney p = 0.012, r = 0.21). They also had lower type-token ratios indicating more focused vocabulary centered on a specific behavior (TTR median 0.79 vs. 0.85, p = 0.014, r = −0.26). They also used more first-person singular language (median count 5 vs. 3, p<0.01). Short, vague entries were not linked to behavioral change; longer, more personal ones were. However,
LLM semantic scoring showed that INTENTION entries differ from NON-INTENTION entries in behavioral concreteness, planning depth, and emotional engagement (all p < 0.05)... Within INTENTION entries, though, these scores did not separate improved from not-improved entries. This suggests the signal lies in how much someone elaborates and in their personal voice, not in the topic or category."
The paper frames these as exploratory findings due to the small sample: The sample is small, so we treat these as exploratory patterns that point to where AI journaling nudges are most likely to work.
The authors make two contributions: "First, a responsiveness map across the 26 sensing features, in which dependence on other people – more than individual controllability – separates responsive from unresponsive behaviors. Second, descriptive evidence that, in this sample, two things track short-term follow-through: intention language for individually controllable behaviors such as SMS, and richer elaboration within intention entries."
The design implications include: "Account for social dependence when targeting nudges... behaviors that depend on other people's presence, such as calls, in-person conversations, group settings, were consistently unresponsive to journaling nudges, so systems should not expect prompts alone to move them. Also,
Detect and act on intention language... when a user writes an intention response to a redirective prompt about a controllable behavior, that signal may be worth acting on, logging it as a commitment, triggering a follow-up, or building the next prompt on the stated plan. And
Follow up on short, vague intentions. Intention entries under ∼25 words were not associated with behavioral change in this sample."
The paper also notes that "the aggregate pattern is the most consequential design signal in this paper. Aside from the exploratory, behavior-specific patterns above, no single interaction reliably predicted behavioral change at the aggregate level; sustained engagement over weeks did. For practitioners, features encouraging regular journaling habits such as streaks, gentle reminders, low-friction entry may matter more than optimizing any individual prompt."
Improvements for AI systems
Improvements to AI systems:
- Social-Dependence-Aware Nudge Targeting
-
The AI system should classify target behaviors by social dependence (e.g., requires coordination with others vs. solo action) before generating nudges.
-
For socially dependent behaviors (calls, in-person conversations, group activities), the system should avoid relying on journaling prompts alone; instead, it should suggest concrete preparatory actions (e.g.,
text them a proposed time,
prepare a topic list
) or pair the nudge with external scheduling tools. -
For solo behaviors, the system can proceed with standard prompts but must not assume uniform success—it should monitor and adapt per behavior type.
- Intention-Language Detection and Commitment Logging
-
The system should use a lightweight LLM classifier to detect explicit intention language in user journal entries (e.g.,
I will,
I plan to,
I need to
). -
When intention is detected for a controllable behavior (e.g., SMS, walking), the system should:
-
Log it as a formal commitment with a timestamp.
-
Trigger a follow-up prompt within 24–48 hours (e.g.,
You planned to text Sarah yesterday—how did it go?
). -
Use the stated plan as the seed for the next prompt (e.g.,
You mentioned calling your mom—what's the first step?
). -
This turns passive journaling into an actionable commitment device.
- Elaboration-Quality Scoring and Adaptive Prompting
-
The system should compute real-time metrics on journal entries: word count, type-token ratio (TTR), and first-person singular pronoun frequency.
-
If an intention entry is short (<25 words) or vague (high TTR, low personal pronouns), the system should:
-
Immediately ask a clarifying follow-up question (e.g.,
What specifically will you do, and when?
). -
Offer structured prompts to elicit more concrete planning (e.g.,
Describe the first step you'll take tomorrow.
). -
If the entry is already elaborate (longer, focused, personal), the system should reinforce it positively and proceed with commitment logging.
- Behavior-Specific Prompt Personalization
-
The system should maintain a per-behavior responsiveness profile (e.g., walking: 58% improvement; gym: 20%; entertainment apps: 18.8%).
-
For low-responsiveness solo behaviors (gym, cycling, entertainment), the system should:
-
Change nudge style (e.g., from reflective to action-oriented, or add gamification).
-
Reduce frequency of prompts for these behaviors and instead suggest alternative behaviors with higher responsiveness.
-
For high-responsiveness behaviors (incoming SMS, walking), the system should prioritize these in prompt scheduling.
- Aggregate Habit-First Design
-
The system should prioritize features that drive sustained journaling engagement (streaks, gentle reminders, low-friction entry) over optimizing individual prompt content.
-
It should track user retention and journaling frequency as primary success metrics, and only secondarily measure behavior change.
-
When a user shows declining journaling frequency, the system should trigger re-engagement mechanisms (e.g., shorter prompts, positive reinforcement) before attempting any behavior-change intervention.
What the improved AI system can do:
-
It can decide when to nudge, what to nudge about, and how to phrase the nudge based on the social dependence of the target behavior and the user's current writing quality.
-
It can convert journal entries into actionable commitments and follow up on them automatically, increasing the likelihood of follow-through for controllable behaviors.
-
It can detect vague intentions and prompt users to elaborate in real time, improving the quality of plans and subsequent behavior change.
-
It can adapt its behavior-change strategy per behavior type, avoiding wasted effort on chronically unresponsive behaviors.
-
It can maintain user engagement over weeks by focusing on habit formation, which the paper identifies as the only reliable predictor of change at the aggregate level.
Sources
- Causal Stories from Sensor Traces: Auditing Epistemic Overreach in LLM-Generated Personal Sensing Explanations
- MindfulDiary: Harnessing Large Language Model to Support Psychiatric Patients' Journaling
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support