• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Saturday, September 12, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

AI Learns to Draw Decision Boundaries as Readable Equations

Bioengineer by Bioengineer
September 12, 2026
in Technology
Reading Time: 6 mins read
0
AI Learns to Draw Decision Boundaries as Readable Equations
Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Machine learning models are famous for their uncanny accuracy and infamous for their opacity. A random forest or a neural network can sift through thousands of patient records, financial transactions, or sensor readings and deliver a verdict in milliseconds, but when practitioners ask why the model reached its conclusion, the answer is usually buried in millions of weighted connections or hundreds of tangled decision trees. A new study published in the journal Machine Learning challenges this trade-off between power and transparency, introducing a framework that discovers a single, human-readable equation whose sign alone tells you which class a data point belongs to.

The method, called Equation Discovery for Classification, or EDC, was developed by Guus Toussaint and Arno Knobbe of Leiden University. It extends a research tradition that has long flourished in regression: symbolic regression, the automated search for analytical formulas that fit numerical data. Theauthors of the study point out that equation discovery has historically been applied almost exclusively to problems where the target is a continuous number, such as recovering physical laws like the relationship between the pressure, volume, and temperature of a gas. Their contribution is to redirect that machinery toward binary classification, where the goal is to separate two classes cleanly and explainably.

The core idea is elegantly simple in concept. Rather than learning an opaque scoring function, EDC searches for a concise mathematical expression f(x) and a threshold theta, such that a data point is assigned to the positive class whenever f(x) meets or exceeds theta. For a linear equation, this is essentially the geometry underlying logistic regression or a linear support vector machine: a hyperplane slicing the feature space into two halves. What sets EDC apart is that the search is not confined to straight lines. By allowing nonlinear building blocks such as products of features and exponential terms, the algorithm can trace curved, interaction-driven boundaries while still producing an expression short enough for a domain expert to read, critique, and even correct by hand.

Technically, the framework rests on two pillars: a structured search and a dedicated numerical optimizer. The search space of candidate equations is defined by a configurable context-free grammar, a design choice inherited from classic work on declarative bias in equation discovery. The grammar constrains equations to sums of simple summands, including linear terms, products of two features, and exponential expressions, each parameterized by constants. Crucially, the grammar is redundancy-aware: constructions that would produce syntactically different but semantically identical equations are pruned in advance, since, for example, subtraction between summands is unnecessary when a constant coefficient can simply be negative. Traversing this space exhaustively is impossible for all but the smallest problems, so the algorithm employs beam search, iteratively refining the most promising candidate equations level by level while keeping only a fixed number of survivors at each stage.

Fitting the constants inside each candidate equation proved surprisingly subtle. Although every equation in the grammar is differentiable, gradient descent performed poorly in practice, largely because of the exponential terms that pepper the search space. The authors instead adopted a tailored hill-climbing procedure that first samples a large pool of random constant configurations, then concentrates its remaining budget on refining the best ones. In a systematic comparison against off-the-shelf optimizers from the SciPy library, including Powell, Cobyqa, Cobyla, Nelder-Mead, and stochastic gradient descent, this hill climber achieved the highest mean area under the ROC curve across one hundred randomly generated test problems. Perhaps most strikingly, a simple random-sampling baseline already reached a mean AUC of 0.9974 with a budget of one thousand evaluations, revealing that the inner optimization problem is more tractable than one might fear.

The experiments on synthetic data produced one of the study’s most intriguing findings. When Gaussian noise was injected into datasets whose generating decision boundaries were known, EDC did not merely match the original boundary; beyond a certain noise level it outperformed it. The explanation is a phenomenon the authors describe carefully: noise pushes data points near a curved boundary across it, and concave sections of that boundary collect more stray points than they lose, causing the effective boundary embedded in the data to drift away from its original position and gradually straighten. EDC, fitting the data rather than the hidden formula, tracks this shifted, smoothed boundary, achieving higher AUC than the very equation that generated the data. As noise grew from negligible to substantial across seventeen hundred artificial datasets, even a restricted linear version of EDC eventually beat the original nonlinear boundary.

The framework also proved capable of reconstructing notoriously difficult structures, including XOR-like and interaction-driven boundaries, and of fitting data produced by Gaussian clusters where no explicit target equation exists at all. On these cluster-based problems, which the authors describe as closest to real-world conditions among their artificial experiments, EDC outperformed existing symbolic classification approaches while landing near state-of-the-art black-box methods such as random forests, multi-layer perceptrons, and radial-basis-function support vector machines, all of which posted mean AUC values around 0.97.

Real-world benchmarks reinforced the pattern. Across nine binary classification datasets from the UCI repository, spanning tasks from banknote authentication to income prediction, EDC achieved a higher AUC than every competing equation-discovery-based classifier on every dataset, and beat simple decision trees across the board. A critical distance analysis showed that random forests, neural networks, and SVMs did not statistically significantly outperform EDC, even though those black-box methods won on several datasets, particularly ionosphere and sonar, where feature-class relationships appear to lie outside EDC’s current set of building blocks. On the Adult income dataset, the discovered equation offered a vivid demonstration of interpretability in action: one term acted as a penalty on years of education only when the individual appeared as a child in the household, while an exponential term over marital status effectively functioned as an if-else statement, adding roughly 28,586 to the score for married individuals living with a spouse and a negligible 8 otherwise.

The method’s main weakness is computational cost. In its default configuration, with a search depth of six and a beam width of ten, EDC takes dramatically longer than the sub-second runtimes of conventional classifiers, largely because pairwise interaction terms grow quadratically with the number of features and because one-hot encoding of categorical variables inflates the feature count. The authors show, however, that the expense is largely optional. Restricting the search depth, narrowing the beam, or replacing interaction terms with simple quadratic terms produced speed-ups of up to a factor of fifty at only a marginal loss of accuracy, with no statistically significant difference in performance across the benchmark suite. The authors also note that the framework is not limited to two classes: standard one-versus-rest schemes can extend EDC to multi-class problems, with each of the resulting equations explaining when a particular class label prevails.

The work arrives amid a broader movement toward explainable machine learning, driven by domains such as medicine, finance, and criminal justice where decisions carry real consequences and regulators increasingly demand justification. EDC offers a principled bridge between symbolic regression and classification, delivering models that are simultaneously compact, expressive, and auditable. While it may not dethrone random forests or deep networks on raw accuracy, it demonstrates that the gap is small enough, and the interpretability dividend large enough, that equations once again deserve a seat at the machine learning table. The authors have released their code, experimental setup, and datasets openly so that the entire study can be reproduced and the grammar extended by practitioners in new domains.

Subject of Research: Equation discovery for binary classification using interpretable symbolic decision boundaries

Article Title: Equation Discovery for Classification: Finding Interpretable Symbolic Specifications of the Decision Boundary

Article References: Toussaint, G., & Knobbe, A. (2026). Equation Discovery for Classification: Finding Interpretable Symbolic Specifications of the Decision Boundary. Machine Learning, 115(9), Article 215. https://doi.org/10.1007/s10994-026-07156-1

Image Credits: AI Generated

DOI: 10.1007/s10994-026-07156-1

Keywords: equation discovery, binary classification, symbolic regression, decision boundary, interpretable machine learning, beam search, symbolic classification, explainable AI, context-free grammar, UCI datasets, machine learning, Equation

Cite Scienmag News
APA MLA Chicago

Denise Maddox. (September 12, 2026). AI Learns to Draw Decision Boundaries as Readable Equations. Scienmag. https://scienmag.com/ai-learns-to-draw-decision-boundaries-as-readable-equations/

Denise Maddox. “AI Learns to Draw Decision Boundaries as Readable Equations.” Scienmag, 12 September 2026, https://scienmag.com/ai-learns-to-draw-decision-boundaries-as-readable-equations/. Accessed 12 September 2026.

Denise Maddox. “AI Learns to Draw Decision Boundaries as Readable Equations.” Scienmag. September 12, 2026. https://scienmag.com/ai-learns-to-draw-decision-boundaries-as-readable-equations/

Copy citation Download RIS

Tags: AI transparency and explainabilityautomated equation discovery for data classificationbeam searchbinary classificationclassification models with analytical formulascontext-free grammardecision boundarydecision boundary equations in AIdecision tree simplificationEquationequation discoveryequation discovery in machine learningexplainable AIhuman-readable machine learning modelsinterpretable machine learninginterpretable neural networksMachine learningmachine learning model interpretabilityreadable mathematical models in AIsymbolic classificationsymbolic regressionsymbolic regression for classificationtransparent AI decision-makingUCI datasets

Share12Tweet7Share2ShareShareShare1

Related Posts

New Open-Source Platform Puts Data Maturity Self-Assessment in Every Organization’s Hands

New Open-Source Platform Puts Data Maturity Self-Assessment in Every Organization’s Hands

September 12, 2026
Indonesian Forecasters Confront Fixed Heat Thresholds and Trust Their Memories to Warn of Extreme Heat

Indonesian Forecasters Confront Fixed Heat Thresholds and Trust Their Memories to Warn of Extreme Heat

September 12, 2026

Curved Algebraic Spaces Give AI a Sharper Grasp of Multimodal Knowledge

September 12, 2026

AI Learns to Read Metal Microstructures, Unlocking Faster Additive Manufacturing Design

September 12, 2026

POPULAR NEWS

  • Green Cane Harvesting Emerges as Brazil’s Key Weapon for Soil Carbon

    29 shares
    Share 12 Tweet 7
  • Gene Mutations After Surgery Predict Lung Cancer Recurrence Risk in Large Chinese Cohort

    29 shares
    Share 12 Tweet 7
  • Hospital IT Vendors May Not Drive Digital Maturity, New Statistical Scrutiny Warns

    29 shares
    Share 12 Tweet 7
  • Sugar Cane Waste Turned Into Microbial Rhamnolipid Biosurfactants for a Circular Bioeconomy

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Green Cane Harvesting Emerges as Brazil’s Key Weapon for Soil Carbon

Gene Mutations After Surgery Predict Lung Cancer Recurrence Risk in Large Chinese Cohort

Hospital IT Vendors May Not Drive Digital Maturity, New Statistical Scrutiny Warns

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.