Drive-to-Music: Context-Aware Generative Audio for In-Vehicle Experiences

arXiv:2608.12615 · cs.SD, cs.LG · Submitted 2026-08-12 · Read on arXiv

Cosmin Dragoiu, Nooshin Nabizadeh

Mercedes-Benz Research & Development North America

cs.SD, cs.LG

Submitted: 2026-08-12

Updated: 2026-08-14

License: http://creativecommons.org/licenses/by-sa/4.0/

Importance score: 50/100

The gist: Drive-to-Music: Context-Aware Generative Audio for In-Vehicle Experiences Summary This paper introduces Drive-to-Music, a context-aware system that generates music in real time from multimodal

Terminology

Summary

Drive-to-Music: Context-Aware Generative Audio for In-Vehicle Experiences

Summary

This paper introduces Drive-to-Music, a context-aware system that generates music in real time from multimodal driving signals. The system uses dashcam imagery and vehicle telemetry to extract scene semantics and driving context, maps them to high-level musical descriptors, and conditions generative audio models to produce contextually aligned soundtracks. The architecture combines perception and generative components to translate visual and kinematic inputs into structured musical attributes and synthesize audio with low latency. It supports smooth transitions as driving conditions evolve, and to ensure robustness and deployment readiness, the system incorporates constraint-based controls and safety checks across the generation pipeline.

The paper notes that while AI-driven music generation has become prominent in content creation industries, the automotive industry remains an unexplored domain for this technology. Existing systems rely on selecting, reordering, or remixing existing music, and many evaluations are simulator-based prototypes rather than deployments in real vehicles. Drive-to-Music is presented as a novel AI-driven system that dynamically generates personalized music based on real-time vehicle contextual data.

System Architecture

The Drive-to-Music architecture consists of several stages that process multimodal vehicle data to produce contextually relevant music alongside complementary media elements such as cover art and song name. The main input is a front-facing camera image capturing the driving scene, with additional vehicle sensor data such as speed and outside temperature used to enhance the music generation process. All generative AI models run in the cloud, with a dedicated in-vehicle service gathering data from the vehicle, sending it to the cloud, and receiving generated music and media elements back for playback. The only requirement is an internet connection enabling periodic communications between vehicle and cloud.

The architecture includes the following stages:

  1. Image Description Stage: A vision-language model (VLM) generates an image description from the camera input. The prompt is designed to extract information like surroundings, visible landmarks, vegetation, sky conditions, and weather. This description provides a compact semantic representation of the visual environment.

  2. Music Description Stage: The scene description is passed to a large language model (LLM), which combines it with additional vehicle sensor data and generates music-oriented descriptors capturing attributes such as mood, atmosphere, and style. These descriptors condition a music-generation model.

  3. Music Payload Generation Stage: In parallel with music generation, the system produces additional media elements. A cover-art image is generated from the scene description using an image-generation model, preserving the mood and semantic content while allowing stylized visual output. A separate LLM call generates a short music title based on the music description.

  4. Dominant Color Extraction Stage: The system extracts the dominant color of the captured scene using a histogram-based approach. Muted sky regions and low-information gray areas are masked, remaining pixels are mapped to a supported in-vehicle color palette, and the resulting dominant color is extracted. This color dynamically adjusts ambient lighting in the vehicle for a more immersive environment.

  5. Safety Check Stage: All AI-generated artifacts are validated against safety and quality requirements. The generated music is checked to ensure it is natural and free from abrupt jumps or excessive noise, using music quality scoring, loudness checks, true-peak checks, sudden-jump detection, and noise-floor analysis. Cover art is screened for offensive visual content, and music titles are filtered for vulgar or disallowed language. Failed components are regenerated until they meet requirements.

Model Selection and Evaluation

Model selection across all components is guided by system-level constraints including minimizing end-to-end latency, maintaining acceptable output quality, and reducing resource usage. Metrics are estimated based on evaluating approximately ten samples per model.

For the image understanding VLM, candidates included LiquidAI LFM 2.5 VL 1.6B, Qwen 3 VL 2B Instruct, and Gemini 2.5 Flash Lite. All models provided comparable output quality with CLIP scores in the 26–27 range. Gemini 2.5 Flash Lite achieved the lowest latency at approximately 0.6 seconds per request, while LiquidAI LFM 2.5 VL 1.6B and Qwen 3 VL 2B Instruct required around 1.2 and 3.9 seconds respectively. Despite Gemini's lower latency, the authors selected LiquidAI LFM 2.5 VL 1.6B as the scene-description component.

For the music description LLM, candidates were LiquidAI LFM 2.5 1.2B Instruct, Qwen 3.5 2B, and Gemini 2.5 Flash Lite. LiquidAI LFM 2.5 1.2B Instruct and Gemini 2.5 Flash Lite achieved perfect quality scores of 5/5. LiquidAI was fastest at approximately 0.7 seconds per request, followed by Gemini at 1.2 seconds and Qwen at 2.4 seconds. LiquidAI LFM 2.5 1.2B Instruct was selected for both music description and song name generation.

For the music generation model, the system prioritized models producing instrumental music with no vocals (as vocals can distract drivers) and the ability to generate songs of 90 seconds or longer to minimize transition frequency and provide immersive experience. Candidates were ElevenLabs, Mubert, and Stable Audio 2.5. Evaluation used the MuQ-MuLan score for prompt–audio semantic alignment and a custom Technical Audio Quality Score (TAQS) summarizing engineering quality based on loudness, true peak, clipping, and sudden loudness jumps. Stable Audio 2.5 was selected with the best MuQ-MuLan score of 0.36 and a TAQS score of 78.03, despite Mubert having a higher TAQS score of 91.26.

For cover art generation, candidates were Stable Diffusion 3.5 Large Turbo, Qwen Image, and Gemini 2.5 Flash Image. Gemini 2.5 Flash Image and Qwen Image achieved the highest GenEval scores of 0.96 and 0.91 respectively, while Stable Diffusion 3.5 Large Turbo scored 0.66. Stable Diffusion 3.5 Large Turbo was the fastest at 1.4 seconds per request locally, followed by Gemini at 6.5 seconds and Qwen at over 40 seconds. Stable Diffusion 3.5 Large Turbo was selected as the cover art generation component.

Conclusions

The paper demonstrates the feasibility of real-time, context-aware music generation in automotive settings. Through systematic evaluation, components were selected that balance latency, quality, and resource constraints. Limited user studies showed positive reception, with participants appreciating the contextual relevance and personalization of the generated music. Future work will extend the system with additional signals such as traffic, navigation, and user preferences to further enhance personalization and contextual relevance.

Improvements for AI systems

Improvements to AI Systems:

  1. Multimodal Context-to-Music Mapping Pipeline

Build an AI system that ingests raw visual (camera) and telemetry (speed, temperature) data, extracts semantic scene descriptions via a VLM, and translates them into structured musical attributes (mood, style, tempo) using an LLM. This enables real-time generative music that adapts to changing environments without manual curation.

  1. Low-Latency Generative Orchestration with Safety Constraints

Implement a cloud-based pipeline that parallelizes music generation, cover art creation, and title generation while enforcing automated quality gates (loudness, true-peak, noise-floor, abrupt-jump detection) and content filtering. The system can regenerate failed components automatically, ensuring deployment-ready outputs with minimal latency.

  1. Dynamic Ambient Synchronization via Dominant Color Extraction

Enhance the system to extract the dominant color from the driving scene (masking sky and low-information areas) and map it to an in-vehicle lighting palette. This creates an immersive, synchronized environment where ambient lighting changes with the visual context, improving user experience beyond audio alone.

  1. Context-Aware Music Continuity and Transition Management

Use the system’s ability to generate long-form (90+ second) instrumental tracks with smooth transitions as driving conditions evolve. The AI can preemptively adjust musical descriptors based on predicted scene changes (e.g., entering a tunnel, highway, or scenic area), reducing abrupt audio shifts and maintaining listener immersion.

  1. Model Selection via Multi-Objective Evaluation Framework

Adopt a systematic evaluation method that scores candidate models on latency, output quality (e.g., CLIP, MuQ-MuLan, GenEval), and resource usage. The improved AI system can automatically select the best-performing lightweight models (e.g., LiquidAI for LLM/VLM, Stable Audio for music) to balance speed and quality in resource-constrained environments.

  1. Personalized and Predictive Music Generation

Extend the architecture to incorporate additional signals (navigation, traffic, user preferences) as future inputs. The improved system can learn user-specific musical tastes over time and predict preferred genres or moods based on route context, enabling fully personalized, proactive soundtracks.


What the Improved AI System Can Do:

  • Generate real-time, contextually relevant instrumental music that matches the visual and kinematic environment of a moving vehicle (e.g., calm ambient music on a quiet rural road, energetic beats on a highway).

  • Produce synchronized cover art, song titles, and ambient lighting colors that reflect the current driving scene, creating a cohesive multimedia experience.

  • Operate with end-to-end latency under 3 seconds, ensuring seamless playback without noticeable delays.

  • Self-correct by regenerating any unsafe or low-quality output (e.g., noisy audio, offensive visuals) before playback.

  • Adapt smoothly to changing conditions (weather, scenery, speed) without jarring musical transitions.

  • Scale to different vehicles and cloud environments, requiring only periodic internet connectivity.

  • Serve as a foundation for future personalization, where the system learns driver preferences and integrates navigation/traffic data to predict and generate music proactively.

Abstract

In-vehicle music can serve as an adaptive interface to enhance driver experience, attention, and well-being. We present Drive-to-Music, a context-aware system that generates music in real time from multimodal driving signals. Using dashcam imagery and vehicle telemetry, the system extracts scene semantics and driving context, maps them to high-level musical descriptors, and conditions generative audio models to produce contextually aligned soundtracks. The architecture combines perception and generative components to translate visual and kinematic inputs into structured musical attributes and synthesize audio with low latency. It supports smooth transitions as driving conditions evolve, and to ensure robustness and deployment readiness, we incorporate constraint-based controls and safety checks across the generation pipeline. Our results demonstrate the feasibility of real-time, context-aware music generation in automotive settings, providing a foundation for personalized and adaptive in-vehicle audio experiences.

Sources

Related papers