Machine learning has quietly become the engine behind some of the most sensitive decisions in modern society, from diagnosing diseases in hospitals to detecting fraud in financial systems and screening applications in government agencies. Yet the data that powers these models is precisely the data that must never be exposed. A new study published in Neural Computing and Applications by Dianqing Bao of Lianyungang Normal University and Wen Su of Lianyungang Technical College tackles this tension head-on, presenting a collaborative encryption-based training framework that promises to let organizations train powerful deep learning models on private data without ever revealing that data to anyone, including the parties doing the training.
The core problem the researchers address is one that has haunted privacy-preserving machine learning for years. Techniques such as fully homomorphic encryption, secure multi-party computation, and differential privacy can, in principle, guarantee that data remains usable but not visible. Homomorphic encryption, for example, allows mathematical operations to be performed directly on encrypted values, so the underlying information is never decrypted during computation. Secure multi-party computation lets several parties jointly compute a function over their combined inputs while keeping those inputs hidden from one another. Differential privacy adds carefully calibrated noise so that the contribution of any single individual cannot be reverse-engineered from the final model. In theory, these tools deliver exactly what privacy-conscious institutions need. In practice, they have struggled to keep up with the demands of modern deep learning.
The difficulty lies in overhead. Fully homomorphic encryption can slow training by orders of magnitude, turning what would take hours on unencrypted data into computations that stretch across days or weeks. Secure multi-party computation imposes heavy communication burdens, requiring participants to exchange large volumes of messages for every step of the training process. Meanwhile, the deep architectures that dominate contemporary artificial intelligence, such as ResNet convolutional networks and Transformer models, involve millions or billions of parameters whose gradients must be computed, encrypted, and aggregated at every iteration. The result is a gap between what privacy theory promises and what deployment reality allows, a gap that has kept many hospitals, banks, and agencies from adopting encrypted training at all.
Bao and Su’s framework attacks this gap from four directions at once. The first is a distributed key collaboration mechanism built on Shamir secret sharing, a classical cryptographic technique in which a secret, here an encryption key, is split into multiple shares distributed among different parties. No single participant holds the complete key; only a sufficient quorum of shares can reconstruct it. This design eliminates centralized trust dependence, meaning there is no single server or administrator whose compromise would expose the entire system. If one party is breached or behaves maliciously, the key remains safe as long as the attacker has not collected enough shares, a property that dramatically raises the bar for any would-be adversary.
The second innovation is the heart of the framework’s efficiency gains: the integration of partial homomorphic encryption with Top-k gradient sparsity performed directly in the ciphertext domain. Unlike fully homomorphic encryption, which supports arbitrary computation on encrypted data at enormous cost, partial homomorphic encryption supports a limited set of operations, such as addition, at a fraction of the computational price. Gradient sparsification exploits the observation that in most training iterations, only a small fraction of a model’s gradient components carry meaningful magnitude. By selecting and transmitting only the top-k largest gradient values, the framework slashes the amount of data that must be encrypted and exchanged. Crucially, the researchers perform this selection and aggregation on encrypted values, so the sparsification itself never exposes which parameters matter most, a detail that could otherwise leak information about the training data.
The third component is a lightweight adaptive scheduling strategy that dynamically balances three competing demands: the heterogeneity of participating devices, the desired strength of security, and the overall efficiency of training. In real-world collaborative settings, participants range from powerful data-center servers to modest edge devices, and a rigid protocol that treats them all identically will be throttled by its weakest member. The adaptive scheduler adjusts workloads and security parameters on the fly, allowing faster devices to contribute more while maintaining the cryptographic guarantees that make the whole exercise worthwhile. This kind of systems-level thinking, the authors argue, is what separates a laboratory demonstration from a deployable platform.
Finally, the team wrapped these mechanisms into a unified, end-to-end software system that provides closed-loop protection across the entire machine learning lifecycle, from the moment data enters the pipeline, through encrypted collaborative training, to the point where the finished model is served to end users. This holistic architecture matters because privacy failures often occur not during training itself but at the boundaries, when data is ingested, when intermediate results are exchanged, or when model predictions are exposed. By covering the full pipeline, the system reduces the attack surface that piecemeal solutions leave open.
The experimental results are striking. On benchmark datasets including CIFAR-10, FEMNIST, and a real-world medical dataset, the framework trained ResNet-18 and Transformer models 3.2 times faster than CryptoNets and 2.7 times faster than HE-Transformer, two well-known homomorphic encryption baselines. Communication overhead, often the hidden killer in distributed encrypted training, came in at 42.3 megabytes, dramatically lower than the homomorphic baseline schemes. Perhaps most importantly for practitioners, the accuracy penalty relative to an unencrypted federated averaging baseline was held to just 1.2 percent, a margin small enough that most applications could absorb it without noticing. The authors also report strong scalability and robustness against real-world privacy attacks, suggesting the protections hold up under adversarial pressure rather than only in idealized conditions.
Why does this matter beyond the benchmark suite? Consider a consortium of hospitals that wants to train a diagnostic model on millions of patient records scattered across institutions. Legal frameworks such as data protection regulations often prohibit sharing raw records, and even anonymized data has been re-identified in past studies. Encrypted collaborative training offers a way out: each hospital keeps its records local, gradients are encrypted before leaving the building, and the aggregated model emerges without any participant ever seeing another’s data. Similar scenarios apply to banks pooling fraud-detection intelligence, telecommunications operators improving network models, and government agencies collaborating across jurisdictions. A framework that makes such training fast enough and cheap enough to run on realistic hardware could convert privacy-preserving machine learning from a theoretical curiosity into standard practice.
The study, published in the special issue on Cognitive based Information Processing and Applications 2024, was supported by research projects in Jiangsu higher education institutions, and the authors state they have no competing interests. Its significance lies less in any single technique than in the demonstration that security and efficiency need not be opposing forces. By combining established cryptographic primitives, secret sharing and partial homomorphic encryption, with machine learning optimizations like gradient sparsification and adaptive scheduling, and by engineering the whole into a deployable software system, Bao and Su have sketched what practical privacy-preserving AI might look like. As cognitive computing continues to penetrate healthcare, finance, and government, frameworks of this kind may become the invisible infrastructure that lets intelligent systems learn from humanity’s most sensitive data while keeping that data exactly where it belongs: out of sight.
Subject of Research: Privacy-preserving collaborative machine learning using encryption-based training and data protection systems
Article Title: Collaborative encryption-based machine learning model training method and data protection software system construction
Article References: Bao, D., & Su, W. (2026). Collaborative encryption-based machine learning model training method and data protection software system construction. Neural Computing and Applications, 38(19), Article 770. https://doi.org/10.1007/s00521-026-12383-7
Image Credits: AI Generated
DOI: 10.1007/s00521-026-12383-7
Keywords: privacy-preserving machine learning, homomorphic encryption, Shamir secret sharing, federated learning, gradient sparsification, data privacy, secure multi-party computation, Transformer models, ResNet, adaptive scheduling, encrypted training, data protection
News Source: Blake Davidson. (October 4, 2026). Encrypted AI Training Gets a Speed Boost Without Sacrificing Privacy. Scienmag.



