Feature selection is one of those unglamorous tasks that quietly decides whether a machine learning model succeeds or flounders. Given a dataset with dozens of variables, an analyst must decide which ones actually carry predictive signal and which merely duplicate each other’s information. The problem is formally NP-hard: the number of possible feature subsets grows combinatorially, so classical algorithms lean on greedy heuristics or approximations. A team of researchers from the University of Valencia, Kipu Quantum, and the University of the Basque Country has now proposed a strikingly different route, described in the journal Quantum Machine Intelligence, in which the choice of features is delegated to the physics of laser-driven neutral atoms.
The method, which the authors call Quantum Feature Selection (QFS), exploits analog quantum simulation with arrays of neutral atoms excited to Rydberg states. In these processors, individual atoms are trapped by optical tweezers and act as two-level systems, with a ground state standing for zero and a highly excited Rydberg state standing for one. Crucially, atoms in Rydberg states interact through van der Waals forces that fall off with the sixth power of distance. That fixed physical law, normally a constraint, becomes the computational resource: by rearranging the atoms in a plane, the researchers can program the interaction structure of an optimization problem directly into the geometry of the array.
The mapping works in two steps. First, the relevance of each feature, quantified by its mutual information with the target variable, is encoded as a site-dependent local detuning, a laser parameter that makes certain atoms energetically more likely to be excited. Features with high relevance receive stronger detunings, biasing their atoms toward the Rydberg state. Second, pairwise redundancy between features, measured as mutual information between the features themselves, is translated into spatial distance. The team applies an inverse-sixth-root transform to the normalized redundancy matrix, so that highly redundant pairs of features are placed close together, where their strong interaction penalizes the configuration in which both are simultaneously selected. Weakly correlated features are pushed apart, where they behave nearly independently.
Because a general redundancy matrix cannot usually be embedded exactly in two dimensions, the researchers use multidimensional scaling (MDS), a classical technique that finds atomic positions best approximating the target distances. The resulting layout is then rescaled so that the closest atom pairs sit near the Rydberg blockade radius, the critical distance at which joint excitation is suppressed. An annealing protocol follows: the global Rabi frequency ramps up and down smoothly, a global detuning initializes the system, and the local detunings encoding relevance switch on in the latter half of the evolution. The dynamics steer the array toward low-energy configurations that balance relevance against redundancy, and repeated projective measurements yield bitstrings, each one a candidate feature subset.
The post-processing stage is deliberately conservative. Only the lowest-energy ten percent of measured bitstrings are retained, and features are ranked by their average Rydberg excitation density across this ensemble. A greedy construction then accepts a feature only if its normalized redundancy with all previously accepted features falls below a threshold, set to 0.7 in the reference experiments, ensuring that the final subset is not merely relevant but internally non-redundant. When the threshold rule cannot fill a requested subset size, the remaining slots are filled by density ranking, a completion step the authors flag explicitly so that large subsets are not misread as strict independent sets.
Evaluated in simulation on Amazon Braket, using a noise-free Schrödinger equation solver and 10,000 measurement shots per run, QFS was tested on three real-world binary classification benchmarks: Adult Income, Bank Marketing, and Telco Churn. The comparison set was demanding, including mutual-information ranking, Random Forest and XGBoost feature importances, L1-regularized logistic selection, and two classical optimizers of exactly the same relevance-redundancy objective, an exhaustive search called Exact-Q and a greedy heuristic called Greedy-Q. Across ten stratified train/test splits with a Random Forest downstream classifier, QFS achieved competitive, dataset-dependent performance. On Adult Income, QFS, Exact-Q, Greedy-Q, and XGBoost importance all landed within roughly three ten-thousandths of a point in mean AUC, while on Telco Churn every method was essentially tied within split-to-split variability.
The authors are candid about where the method falls short. On Bank Marketing, tree-based importance rankings achieved slightly higher mean AUC than the quantum-derived subsets, and QFS sometimes reached its best performance at larger subset sizes, a consequence of the evaluation protocol rather than an intrinsic preference of the sampler for compact sets. The geometric embedding itself introduces moderate distortion: the mean relative reconstruction error of the MDS layouts ranged from about 0.22 to 0.26 across the three datasets. Yet a stability analysis over 50 random embedding seeds showed that this distortion does not produce arbitrary subsets. Bank Marketing was invariant across seeds for subset sizes up to five, and Adult Income stabilized for sizes of three or more, with mean Jaccard similarity above 0.96.
Robustness tests of the classical extraction stage added further reassurance. Varying the retained low-energy fraction from 5 to 30 percent shifted the best AUC only marginally on a reference split, and injecting controlled noise into the measured excitation densities, mimicking readout fluctuations or calibration drift, left predictive performance largely intact even when the exact identity of the selected subsets drifted at small cardinalities. The authors interpret this as evidence that the pipeline’s predictive output is more robust than its exact subset choices, though they stress that the simulations remain noise-free: decoherence, atom loss, finite-temperature atomic motion, readout errors, and detuning miscalibration were all excluded and must be addressed in future device-level studies.
What makes the work notable is less a claim of quantum advantage than the physical interpretability of the encoding. Every term of the statistical objective has a hardware counterpart: relevance lives in local detuning amplitudes, redundancy in the sixth-power interaction law, and the relevance-redundancy trade-off emerges from the relative energy scales of the two. The main contribution, the authors write, is a neutral-atom implementation of a relevance-redundancy feature-selection objective together with a quantitative characterization of its geometric and post-processing approximations. Scaling up will require one atom per feature, site-resolved detuning control, and embedding strategies, such as sparsification or community-aware layouts, that preserve the dominant redundancy structure without forcing every pairwise relation into a single plane.
The study points toward a broader program in which analog quantum processors tackle the combinatorial core of data preprocessing rather than the models themselves. Future work, the team says, should focus on experimental validation on real neutral-atom devices, noisy analog Hamiltonian simulation, improved weighted and graph-aware embeddings, dynamic detuning optimization, and larger benchmarks where the scaling of analog sampling can be measured directly. Extensions to multiclass problems, regression, and unsupervised settings are also on the table. For now, QFS stands as a proof of concept that a laser-driven array of atoms, arranged according to the statistical anatomy of a dataset, can select features as well as established classical heuristics, and that the path from information theory to atomic physics can be traversed without losing the meaning along the way.
Subject of Research: Analog quantum feature selection using neutral-atom Rydberg processors
Article Title: Analog quantum feature selection with neutral-atom quantum processors
Article References: OrquÃn-Marqués, J. J., Flores-Garrigós, C., Gomez Cadavid, A., Simen, A., Solano, E., Hegade, N. N., MartÃn-Guerrero, J. D., & Vives-Gilabert, Y. (2026). Analog quantum feature selection with neutral-atom quantum processors. Quantum Machine Intelligence, 8(2), Article 108. https://doi.org/10.1007/s42484-026-00456-8
Image Credits: AI Generated
DOI: 10.1007/s42484-026-00456-8
Keywords: neutral atoms, Rydberg arrays, quantum feature selection, analog quantum simulation, mutual information, quantum optimization, machine learning, multidimensional scaling, Rydberg blockade, QUBO, Amazon Braket, Quantum Machine Intelligence
News Source: Katie Riggs. (October 6, 2026). Rydberg Atom Arrays Turn Feature Selection Into a Quantum Physics Problem. Scienmag.



