MazzikaAI: A knowledge-based performance-to-prompt compiler for real-time Arabic maqam accompaniment with a streaming text-to-music model

arXiv:2608.10360 · cs.HC, cs.AI, eess.AS · Submitted 2026-08-11 · Read on arXiv

Jiaxin Du, Boulbaba Abdeljaouad, Yong Zhuang, Haoyu Li

Grand Valley State University

cs.HC, cs.AI, eess.AS

Submitted: 2026-08-11

Updated: 2026-08-12

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: MazzikaAI is a knowledge-based system for real-time Arabic maqam accompaniment that uses natural language as the actuator of a real-time control loop.

Terminology

Summary

MazzikaAI is a knowledge-based system for real-time Arabic maqam accompaniment that uses natural language as the actuator of a real-time control loop. The system compiles live MIDI, gesture, and inferred harmony into continuously updated text prompts to steer an unmodified streaming generator (Google Lyria RealTime) without requiring model fine-tuning. The system embeds expert knowledge of six core maqamat, characteristic ornaments, and ensemble dynamics, maintaining real-time responsiveness with sub-second key-to-audible-update latency.

The paper's core methodological idea is prompt compilation as a real-time control law: a deterministic compiler C: (st, qt) 7→ pt over an estimated performance state, a four-state call-and-response policy, and a coarse control-signature gate that decides when re-steering the live stream is musically warranted. This turns an unmodified streaming text-to-music model into a responsive accompanist with no fine-tuning, and is independent of the particular generator behind it.

The system is organized as a classical knowledge-based system: an explicit, hand-authored knowledge base; a working memory holding the performance state estimated from the live event stream; and a deterministic, rule-based inference layer. What is new is the effector—where a classical expert system renders its conclusions through a symbolic synthesizer, MazzikaAI renders them through a large pretrained generator addressed in prose—a neuro-symbolic division of labour in which inspectable symbolic expertise supplies idiom and moment-to-moment intent while the foundation model supplies raw musical competence.

The architecture is a three-tier system—browser client, Python backend, cloud generative model—connected by a single bidirectional WebSocket. Four design commitments shape it: (i) Text is the sole point of control; (ii) Idiom is injected, not learned; (iii) The performer conducts; (iv) Re-steering is gated. The client captures four input modalities: note events through the Web MIDI API, hand gestures with MediaPipe Hands, spoken commands through the Web Speech API, and a Learn Mode that parses MusicXML in-browser. The backend maintains a rolling estimate of the performance—density, register, motion, dynamic arc, phrase boundaries, inferred harmony—and, on each meaningful change, compiles that state into a prompt streamed to a text-to-music model. Two generative streams run concurrently: a melodic ensemble that follows and answers the soloist, and a tempo-locked percussion bed.

The accompaniment state machine resolves the estimated performance state into one of four behaviours: Supporting (while the soloist plays, the band comps quietly beneath the line), Responding (when a phrase ends, an answer window opens in which the prompt lists the soloist's just-played pitches and instructs an expressive echo), Sustain (a single note held for >2.5 s is read as a held tone and the accompaniment stays minimal), and LongIdle (after >5 s of silence with sections active, the band assumes the floor and leads freely). Two gates bound the machine: a warmup gate requires two notes after a section is conducted in before the ensemble may sound, and a silence stop halts the melodic stream after >4 s of silence.

The knowledge base is hand-authored prose and numbers, inventoried in Table 4. Its core is the Arabic tradition: six maqamat (bayati, rast, hijaz, nahawand, saba, kurd) with quarter-tone spellings and a five-instrument takht configuration—cello, oud, strings, qanun, and nay (lead)—driven by a bespoke controller with idiom-specific harmonic and echo logic. Around that core the same structures carry eight further Western genres (15 mode descriptions, 38 ensemble sections with three role registers each, 135 generation parameters, and 36 percussion patterns in total). The knowledge base is small enough to read in an afternoon.

Maqam grounding is achieved by injecting a dense scale description into the prompt. The canonical entry for bayati is: "maqam bayati on D - scale: D E-half-flat F G A Bb C D, the half-flat second degree (E koron) is the expressive soul of bayati, ornaments: shimmer on E-half-flat, slides D->E-half-flat, descending resolution G->F->E-half-flat->D, resting tones: D (tonic), G (dominant), color: melancholic, warm, yearning - the most expressive Arabic maqam, avoid western major or minor tonality, stay in bayati modal world. Two choices matter: the microtonal degree is spelled phonetically (E-half-flat", E koron) rather than symbolically, and the closing clause supplies explicit negative guidance.

The compiler assembles the prompt as an ordered concatenation of clauses: [instrument rule / silence] ∥ [no-percussion] ∥ [specific header] ∥ [session stage] ∥ [maqam line] ∥ [harmonic context / echo] ∥ [role assignment] ∥ [register] ∥ [dynamics] ∥ [space clause]. Ordering is load-bearing—hard constraints are placed first. Because the player conducts a subset of instruments in, the prompt must actively suppress the rest—a prompt-level substitute for the stem-level control the API does not expose. The leading clause names only the active instruments and issues a hard silence directive for every inactive one.

Empirical evaluations demonstrate that dynamic prompt compilation reliably grounds generation in microtonal scales, significantly increasing off-grid quarter-tone content over baseline generation. The knowledge-based layer costs under a millisecond end to end against a generator budget three decades larger. The median key-to-audible-update latency of 263 ms lands at the front of the answer window the controller itself opens. The stream tolerates 179 re-prompts per minute without a single failure—and, in the gating ablation, 202 pushes per minute with the gate disabled entirely.

Controlled ablations over input-identical replayed sessions attribute behaviour to design choices: maqam grounding shifts generation toward the tonic and yields significantly more quarter-tone melodic content (22.4% of grounded voiced frames lie ≥35 cents off the 12-TET grid versus 14.9% ablated, and when the melodic line visits the second-degree region above D, 59.4% of grounded frames fall in the half-flat band against 24.1% ablated); a static prompt lets trigger-to-audio latency collapse to tens of seconds, quantifying what the compiler buys. The same instrumentation also marks the current limits of prompt-level control—most notably, constraint-style prohibition is no substitute for stem-level mixing.

A perceptual evaluation with two expert musicians localizes the remaining gap precisely: the compiler controls what the accompaniment plays far more accurately than when it plays it. The performers perceived the accompaniment as holding an essentially constant tempo rather than breathing with their playing—the ensemble kept its own clock. The generator's tempo is fixed for the lifetime of a stream; the percussion bed is tempo-locked by design; and the melodic engine's following is semantic rather than metric—the compiler tracks register, density, dynamics, harmony, and phrase boundaries, but performs no beat tracking on the performer's onsets, so nothing in the control loop re-anchors generation to the actual pulse.

The broader significance is that much of the intelligence a responsive creative partner needs can be located not in model weights but in a transparent, knowledge-based compiler that translates live human state into language a general model already understands—and that on current hardware this layer is cheap enough to be free. The architecture establishes a scalable paradigm for real-time human-AI co-creation, offering a generalizable blueprint for interactive accompaniment, adaptive music education, and culturally inclusive generative audio across diverse global idioms.

Improvements for AI systems

Improvements to AI systems:

  1. Add a deterministic prompt-compilation control layer between live input streams and generative models, replacing end-to-end learned controllers. This layer compiles real-time state (MIDI, gestures, inferred harmony) into structured text prompts with ordered clauses (hard constraints first), enabling sub-second re-steering without fine-tuning or model access beyond text generation.

  2. Implement a gated re-steering mechanism that decides when to update the generation prompt based on musical events (phrase boundaries, sustained notes, silence thresholds). This prevents unnecessary re-prompts, reduces latency jitter, and maintains stream stability—tolerating 180 re-prompts/minute without failure.

  3. Inject microtonal and idiom-specific knowledge via phonetic spellings and negative guidance in prompts. Use explicit textual descriptions of scales (e.g., E-half-flat instead of symbolic notation) and closing clauses like avoid western major or minor tonality to ground generation in non-Western modal systems, increasing off-grid quarter-tone content by 50% relative to baselines.

  4. Add a four-state accompaniment policy machine (Supporting, Responding, Sustain, LongIdle) that maps estimated performance state to behavior. This allows the AI to dynamically switch between comping, echoing, holding, or leading based on live input, with explicit warmup and silence-stop gates to prevent premature or unwanted output.

  5. Enable multi-modal input fusion (MIDI, hand gestures, speech, MusicXML) into a unified working-memory state, then compile that state into a single text prompt. This lets the AI respond to conducting gestures, spoken commands, and notated scores in real time, expanding control beyond keyboard/mouse.

  6. Implement a prompt-level instrument suppression system to substitute for missing stem-level control. By actively naming active instruments and issuing hard silence directives for inactive ones in the leading prompt clause, the AI can selectively mute or include ensemble parts without API-level mixing.

  7. Add a semantic-following mode that tracks register, density, dynamics, harmony, and phrase boundaries (not just pitch) to generate accompaniment that responds to expressive intent. This improves musical coherence over static or purely pitch-based following.

What the improved AI system can do:

  • Act as a real-time accompanist for Arabic maqam and other global idioms (e.g., Western genres) with sub-300ms latency, responding to live performance with idiomatically correct microtonal ornamentation and ensemble dynamics.

  • Accept control via MIDI keyboards, hand gestures, voice commands, or imported sheet music, and switch between these modalities on the fly.

  • Maintain a stable generative stream while being re-steered up to 180 times per minute, without crashing or degrading audio quality.

  • Compose and perform in maqamat (bayati, rast, hijaz, etc.) with correct quarter-tone intervals, resting tones, and characteristic ornaments, while explicitly avoiding Western tonal biases.

  • Dynamically adjust its role—supporting quietly, echoing phrases, sustaining minimal texture, or taking a solo lead—based on the performer’s activity level and silence gaps.

  • Suppress or activate specific instruments (e.g., mute cello, keep qanun) purely through text prompts, enabling sectional control without stem-level access.

  • Serve as a blueprint for adaptive music education, interactive co-creation, and culturally inclusive generative audio, with a transparent, inspectable knowledge base that can be extended to other musical traditions.

Sources

Related papers