Muyu Song, a centuries-old Chinese narrative folk tradition built on wooden-fish percussion and sung storytelling, is the latest cultural heritage art form to be handed over to an algorithm. In a study published in Discover Artificial Intelligence, researchers Kangli Chen, Fei Li, and Jing Chen of Guangdong University of Science and Technology describe a framework that treats the digital spread of this intangible cultural heritage not as a matter of intuition or fixed promotional rules, but as a formal sequential decision problem that a machine can learn to solve. Their system listens to the songs, reads the dialect-laden lyrics, predicts how audiences will respond, and then decides where, when, and how often each fragment of a performance should be published across digital platforms.
The motivation is straightforward. Muyu Song has been migrating from temple courtyards and oral transmission to short-video platforms and streaming channels, and while that shift has widened its potential audience, it has not automatically deepened its impact. The researchers point to three structural failures in current digital dissemination: content is fragmented in ways that ignore the music’s internal structure, audience feedback is scattered and slow to arrive, and publishing decisions rely on empirical rules that cannot adapt. The result is heritage content that reaches many people but often fails to hold them, engage them, or transmit the cultural meaning embedded in the performance.
What makes the new work unusual is its insistence that a folk song is not ordinary online content. Muyu Song performances contain flexible vocal pauses, beat-dependent chanting structures, dialectal expressions, rhyme-based filler words, and long narrative sequences. Generic AI pipelines slice audio into fixed-duration clips and run lyrics through standard Mandarin text encoders, which the authors show is a recipe for boundary errors and semantic dilution. Instead, their framework organizes every recording into segment-level samples whose boundaries are defined by silence, beat changes, chanting pauses, and narrative turns, so that each unit of dissemination corresponds to a coherent musical and narrative moment rather than an arbitrary time window.
Technically, the pipeline begins with a dual-branch audio encoder. A short-time Fourier transform produces a power spectrum that is mapped through a Mel filter bank into log-Mel acoustic features, while a parallel rhythm branch extracts beat period and phase information from onset intensities and tempograms. A lightweight convolutional network captures local spectral texture, a BiLSTM-Transformer combination encodes longer temporal structure, and attention pooling compresses everything into a segment-level audio embedding. On the text side, time-aligned lyrics are normalized so that function words, filler words, and rhythmic particles are annotated but preserved, with their temporal positions intact, before a Transformer encoder generates context-sensitive semantic vectors. A time-weighted attention pooling mechanism then aligns the text embedding with the audio embedding on the same timeline.
Fusion is handled by a gated cross-modal attention mechanism rather than simple concatenation. Audio features act as acoustic priors that reweight the text representation, giving more attention to lyric segments correlated with strong beats and emotional shifts, while textual semantics reweight the audio features to suppress background percussion and redundant chanting. The result is a single fused vector per segment that jointly encodes vocal style, rhythm, narrative, and dialect meaning. The gating weight is learned jointly from both modalities, so the balance between sound and sense shifts adaptively depending on the content of each segment.
On top of this representation sits a prediction model with a shared feature trunk and five separate regression heads, one each for click-through rate, first-completion rate, interaction intensity, share rate, and propagation depth. Crucially, the model also learns five cultural-quality variables: cultural understanding degree, cultural identity tendency, inheritance willingness, misinterpretation risk, and traditional-context fidelity, drawn from post-view questionnaires, semantic analysis of audience comments, and expert review of whether lyrics, dialect meanings, and performance context were correctly understood. These cultural indicators do not replace the traffic metrics; they act as constraint variables and penalty references, preventing the optimizer from degenerating into pure click-chasing.
The decision layer casts dissemination as a Markov decision process with a state combining the fused content features, platform conditions, and audience context, and an action space of 300 discrete configurations spanning four platforms, three presentation formats, five time windows, and five reach-frequency levels. A culturally constrained Actor-Critic optimizer learns a policy that maximizes a discounted long-term return, in which the weighted traffic reward is balanced against resource and disturbance costs and penalized whenever predicted cultural-quality indicators fall below predefined thresholds. Training uses epsilon-greedy exploration with the exploration rate annealed from 0.30 to 0.05, a replay buffer of 5,000 trajectories, mini-batches of 128 transitions, and an adaptive learning-rate rule that slows updates when the environment is stable and accelerates exploration when audience behavior or platform mechanisms drift.
Because the policy learns from predicted outcomes rather than always waiting for delayed platform analytics, the authors acknowledge the risk of the optimizer exploiting systematic errors in its own predictor. They counter this with three safeguards: the predictor is retrained every 500 newly observed outcomes, high-uncertainty predictions are down-weighted in the policy gradient using ensemble variance, and rewards are clipped to prevent the agent from chasing outlier estimates. The experiments themselves were run in a log-reconstructed environment built from real dissemination records rather than a live production system.
The dataset is substantial: 23,540 segment-level samples drawn from Muyu Song content published between January 2021 and December 2023 on a provincial intangible cultural heritage digital platform, cultural-institution media accounts, and public short-video and audio channels. Lyrics were transcribed by automatic speech recognition and manually corrected by annotators familiar with Muyu Song expressions, with dialect tags resolved through secondary review. To prevent data leakage, the 487 songs were split at the song level, so no segments from the same title appeared in both training and test partitions. All experiments were repeated five times with identical data partitions, seeds, and preprocessing.
The results give the domain-specific design choices concrete support. Compared with fixed-rule scheduling, the complete framework improved click-through rate by 8.2 percent, first-completion rate by 6.6 percent, interaction intensity by 11.2 percent, share rate by 9.5 percent, and propagation depth by 23.8 percent, and it also outperformed a generic multimodal Actor-Critic baseline. The ablations are the most telling part: replacing rhythm-aware boundaries with fixed windows reduced propagation depth to 1.87, removing dialect-preserving encoding lowered retention and sharing, and dropping the cultural constraints raised traffic slightly but pushed the misinterpretation risk rate from 0.073 to 0.126 while cutting traditional-context fidelity from 0.887 to 0.813. In other words, an algorithm that understands beats, dialect, and cultural meaning spreads the heritage further and distorts it less. The authors caution that performance still depends on sample size, platform data integrity, and feedback stability, and they flag cross-platform transferability and long-term cultural value as open questions, but the study makes a compelling case that keeping a folk tradition alive online may require teaching the machine to hear it the way its singers do.
Subject of Research: AI-based optimization of digital dissemination paths for Muyu Song intangible cultural heritage
Article Title: Research on optimizing the heritage transmission path of Muyu Song using AI technology
Article References: Chen, K., Li, F., & Chen, J. (2026). Research on optimizing the heritage transmission path of Muyu Song using AI technology. Discover Artificial Intelligence, 6(1), Article 1407. https://doi.org/10.1007/s44163-026-02304-x
Image Credits: AI Generated
DOI: 10.1007/s44163-026-02304-x
Keywords: Muyu Song, intangible cultural heritage, reinforcement learning, multimodal fusion, Actor-Critic, Markov decision process, dissemination prediction, digital heritage, dialect encoding, rhythm-aware segmentation, cultural fidelity, short-video platforms
News Source: Denise Maddox. (October 9, 2026). AI Learns to Keep an Ancient Chinese Folk Song Alive Online. Scienmag.



