Every time a song streams, a podcast circulates, or a voice recording travels across the internet, an invisible battle unfolds over who owns that sound and whether it has been tampered with. Digital audio is trivially easy to copy, re-encode, and redistribute, which makes proving authenticity or ownership a persistent headache for the media industry. A research team in India now reports a new audio watermarking system that hides identifying information inside sound files with unusual resilience, combining a custom deep neural network with an encryption scheme built on an unexpected mathematical object: a fractal curve that resembles a dessert. The work, published in Multimedia Tools and Applications, describes a pipeline the authors call RAWS-SBS, short for Robust Audio Watermarking System with Spatial-and-Cross Scaled Gaussian Error Linear Unit-based Deep Convolutional Neural Network and Blancmange Curve Cryptography for Security.
The team, led by Abhijit Patil of K J Somaiya Institute of Technology in Mumbai, together with Ramesh Shelke of S. S. Jondhale College of Engineering and Dilendra Hiran of Pacific Academy of Higher Education and Research University, set out to fix a specific weakness in earlier watermarking research. Many existing schemes process audio in overlapping segments when embedding a watermark, which drives up computational cost without necessarily improving security. Overlapping windows are a standard tool in signal processing because they smooth transitions between frames, but in the watermarking context the authors argue that the redundancy they introduce is largely wasted effort. Their system instead guarantees non-overlapping segmentation from the very first stage, trimming the computational overhead that has plagued previous approaches.
The pipeline begins with a technique borrowed from classical spectral analysis: the FlatTop window. Windowing functions are mathematical tapers applied to short stretches of a signal before further processing, and each shape trades off different properties. FlatTop windows are prized for their exceptionally flat passband, which means they measure the amplitude of frequency components with minimal distortion, at the cost of a wider main lobe. In RAWS-SBS, the FlatTop window slices an incoming audio signal into frames that do not overlap, establishing a clean, computationally frugal foundation for everything that follows. This choice matters because every subsequent operation, from transform to embedding to extraction, operates on these frames, so inefficiency at the segmentation stage multiplies through the whole system.
Once segmented, the frames undergo a decomposition step using what the authors term a Polar Cartesian-based Discrete Wavelet Transform, or PC-DWT. The discrete wavelet transform is a workhorse of modern signal processing, breaking a signal into components that capture both frequency content and where in time that content occurs, unlike the plain Fourier transform, which discards timing information. The polar Cartesian variant reformulates this decomposition, and the transformed coefficients are then arranged into matrices from which features are extracted. These features feed the next critical decision: where, exactly, inside the audio should the watermark bits live? Placing a watermark carelessly makes it audible to listeners or fragile under compression and filtering; placing it too timidly makes it easy to destroy.
To answer that placement question, the researchers developed an optimization method they call Gamma Functional Energy Valley Optimization, or GFEVO. Optimization algorithms of this family search a large space of candidate solutions for the one that best balances competing objectives, and here the objective is finding embedding locations that maximize robustness while preserving perceptual quality. The reported results suggest the search paid off: the GFEVO stage achieved a Perceptual Evaluation of Speech Quality score of 4.256, a metric on a scale where values near 4.5 indicate quality nearly indistinguishable from the original. In practical terms, a listener hearing the watermarked audio would be very hard pressed to notice that anything hidden inside it at all.
Security enters the pipeline through the most colorful component of the system: Blancmange Curve Cryptography. The Blancmange curve, also known as the Takagi curve, is a fractal defined as an infinite sum of sawtooth-like functions. It is famous in mathematics for being continuous everywhere yet differentiable nowhere, a jagged, self-similar shape that never smooths out no matter how closely you zoom in. The authors harness this fractal’s irregular, deterministic complexity to encrypt the watermark image before embedding, converting the encrypted bits into a vector bit stream. An attacker who extracts the embedded payload without the key would face data scrambled according to the curve’s fractal arithmetic, adding a cryptographic layer on top of the steganographic one. It is a striking example of pure mathematics, a curve named after a wobbly pudding, finding duty in applied security engineering.
Embedding, however, is only half the problem. A watermark is useless if the legitimate owner cannot reliably pull it back out of an audio file that may have been compressed, filtered, resampled, or otherwise degraded on its journey through the digital world. For extraction, the team deployed a deep convolutional neural network with two distinctive design elements. The first is a spatial-and-cross architecture, which lets the network correlate information across different spatial regions of the transformed audio representation rather than treating each location in isolation. The second is a Scaled Gaussian Error Linear Unit, or SGELU, an activation function that governs how signals pass between network layers. Activation functions shape what a neural network can learn, and the Gaussian-error variant blends a smooth probabilistic weighting into the familiar rectified linear unit, potentially giving the extractor a richer response to the subtle statistical traces a watermark leaves behind.
The experimental numbers are where the system makes its case. The proposed method achieved a minimum Bit Error Rate of 0.03232, meaning that when the hidden watermark was extracted, fewer than four bits in every hundred came back wrong, a low error rate that translates directly into more reliable ownership verification. On overall watermarking accuracy, the system reached 97 percent. The authors benchmarked this against several established deep learning baselines, reporting improvements of 3.2 percent over a standard DCNN, 4.3 percent over a conventional CNN, 7.8 percent over a deep belief network, and 11.5 percent over a plain deep neural network. Those margins may look incremental, but in a field where a few percentage points of accuracy can determine whether a copyright claim holds up, they are meaningful.
Context helps explain why this line of research is heating up. Audio watermarking has moved far beyond its original role of tagging music files for royalty tracking. Recent work in the field targets desynchronization-resilient schemes that survive time-stretching attacks, blind quantum watermarking, neural approaches with invertible dual embedding, and systems designed specifically to detect deepfake audio proactively. As generative models make synthetic speech and music indistinguishable from human-created recordings, the ability to embed a verifiable, tamper-evident signature inside audio is becoming a frontline defense. Watermarks that survive aggressive processing while remaining inaudible are the technical backbone for proving that a clip is authentic, or flagging one that is not.
The RAWS-SBS study, received by the journal in 2023 and published on 31 August 2026 after an unusually long revision cycle, reflects that broader urgency. Its authors report no funding source and no competing interests, and they state that the data generated in the study are available from the corresponding author on reasonable request. Whether fractal-based encryption and non-overlapping segmentation become standard ingredients of future watermarking systems remains to be seen, but the work illustrates a clear trend: the fight to authenticate digital audio is increasingly fought with hybrid tools, marrying signal processing classics like wavelet transforms and windowing functions to deep networks and, now, to the jagged elegance of nowhere-differentiable fractals. For an industry drowning in unattributable sound, every percentage point of extraction accuracy counts.
Subject of Research: A robust deep learning-based audio watermarking system using Blancmange curve cryptography for digital media security
Article Title: A robust audio watermarking system with SC-SGELU-DCNN and BCC algorithm for security
Article References: Patil, A., Shelke, R., & Hiran, D. (2026). A robust audio watermarking system with SC-SGELU-DCNN and BCC algorithm for security. Multimedia Tools and Applications, 85(9), Article 725. https://doi.org/10.1007/s11042-026-21865-8
Image Credits: AI Generated
DOI: 10.1007/s11042-026-21865-8
Keywords: audio watermarking, deep learning, convolutional neural network, Blancmange curve cryptography, discrete wavelet transform, bit error rate, digital security, copyright protection, signal processing, SGELU activation, optimization, multimedia forensics
News Source: Blake Davidson. (October 7, 2026). Deep Learning Meets Fractal Math: New Audio Watermarking System Hits 97% Accuracy. Scienmag.



