• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Monday, October 5, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

Self-Taught AI Reads Blockchain Money Trails to Catch Crypto Laundering

by
October 5, 2026
in Technology
Reading Time: 5 mins read
0
Self-Taught AI Reads Blockchain Money Trails to Catch Crypto Laundering

Self-Taught AI Reads Blockchain Money Trails to Catch Crypto Laundering

Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Cryptocurrency has given the world a financial system that moves value across borders in seconds, but it has also given money launderers an environment where anonymity is built into the architecture. A new study published in Discover Artificial Intelligence tackles one of the hardest problems in financial crime detection: how to identify illicit transactions when almost none of them carry labels telling investigators what they are. The research, led by Yong Shang of Henan Judicial Police Vocational College in Zhengzhou, China, introduces a model called GT-SSL, a graph-structured Transformer trained through self-supervised learning, and reports striking results on two of the most widely used benchmark datasets in the field.

The core challenge is deceptively simple to state and brutally hard to solve. In real-world blockchain data, confirmed illicit transactions represent only a tiny fraction of the network. On the Elliptic dataset used in the study, roughly 200,000 Bitcoin transactions include just 2 percent labeled as illicit and 21 percent as licit, with the vast majority unlabeled. Traditional supervised machine learning starves in such conditions, and rule-based systems that once anchored anti-money laundering compliance struggle to keep pace with laundering strategies that mutate constantly. Handcrafted features and manual annotation are expensive, and by the time a rule is written, the criminals have moved on.

Graph neural networks emerged as a promising answer because blockchain transactions naturally form networks: money flows from one transaction to another through the unspent transaction output mechanism, creating directed chains, fan-out patterns, and circular loops that are the fingerprints of laundering. But conventional graph neural networks have their own weaknesses. They can suffer from over-smoothing, in which node representations become indistinguishable after repeated aggregation, and they model long-range dependencies poorly, which matters because laundering schemes often stretch across many hops of transfers. Meanwhile, existing self-supervised methods frequently fail to exploit the graph structure itself, blunting their advantage when labels are scarce.

GT-SSL attacks the problem in three stages. First, raw blockchain records, including transaction hashes, inputs, outputs, timestamps and amounts, are converted into a directed attributed graph in which nodes are transactions and edges represent fund transfers. The model then samples local neighborhoods using a biased restart random walk, generating fixed-length sequences of transactions that serve as input to a Transformer. The walk is deliberately engineered: transition probabilities weigh transaction amount similarity, temporal proximity, edge direction and a degree penalty that prevents massive hub nodes such as exchanges and mixing services from dominating every sampled context. A restart mechanism keeps each sequence anchored near its target transaction, preserving the local laundering path while retaining neighborhood diversity.

The second stage is where the architecture departs most sharply from a standard Transformer. Attention weights are not determined solely by feature similarity in the serialized sequence; they are also constrained by the topology of the underlying fund-flow graph. The study introduces a soft multi-hop structural bias: transactions one hop apart in the graph receive strong attention guidance, two-hop and three-hop neighbors receive attenuated guidance, and distant or unreachable nodes are penalized. This design choice is grounded in how laundering actually works. Layered schemes typically involve multi-level account transfers, fund splitting and cross-node aggregation, so a strict one-hop mask would blind the model to crucial multi-hop paths, while unconstrained global attention drowns it in irrelevant noise. The ablation experiments bear this out: without structural constraints, recall fell to 88.65 percent, whereas the multi-hop soft bias achieved the best results across F1-score, AUC and Matthews correlation coefficient.

Pre-training then proceeds through two complementary self-supervised tasks that share the same encoder. In masked feature reconstruction, 15 percent of node features are replaced with a mask token and the model must reconstruct them from context, forcing it to learn fine-grained transaction attributes. In graph contrastive learning, two augmented views of the graph are generated through edge dropping and feature perturbation, and the model learns to pull the representations of the same node together while pushing different nodes apart, using a temperature-scaled InfoNCE loss with the temperature set to 0.1. The dual-task design proved essential: removing masked reconstruction dropped the F1-score to 93.64 percent, removing contrastive learning dropped it to 93.04 percent, and removing both collapsed it to 81.82 percent, confirming that self-supervised pre-training is the single most important ingredient for learning under label scarcity.

The third stage addresses the pseudo-label problem, the Achilles heel of semi-supervised detection. The model first adopts only predictions whose confidence exceeds a high threshold, then applies a second-stage filter based on graph community structure. Using the Louvain algorithm, the transaction graph is partitioned into tightly connected communities, and a medium-confidence pseudo-label is accepted only if the surrounding community, assumed to be behaviorally homogeneous, provides sufficient consensus support. The thresholds were tuned on validation data and set at 0.90 for confidence and 0.70 for consensus. Across five rounds of self-training, first-stage pseudo-label accuracy stayed above 96.58 percent, second-stage accuracy above 94.62 percent, and cumulative error propagation reached only 3.63 percent, suggesting the mechanism genuinely suppresses the noise amplification that plagues naive self-training.

The headline numbers are impressive. On the Elliptic dataset, GT-SSL achieved an F1-score of 95.80 percent and an AUC of 97.62 percent; on AML-Bitcoin, a larger dataset of roughly 500,000 transactions with about 2,100 labeled laundering cases, it reached an F1-score of 93.78 percent and an AUC of 95.91 percent. It outperformed a broad field of baselines including GCN, GAT, Skip-GCN, EvolveGCN, Inspection-L, GCAF-AML and GNN-GRU, and also beat two post-2023 competitors, Elliptic++-GNN and BERT4ETH-AML, improving F1-score by 2.72 and 2.81 percent respectively while cutting false positives and false negatives. Against recent competitive baselines, the model reduced the average false positive rate to 2.58 percent and the average false negative rate to 7.09 percent, a meaningful margin in a domain where false alarms waste investigator time and missed detections let criminals escape.

Perhaps most striking is the model’s resilience when labels are nearly absent. With only 5 percent of training labels visible, GT-SSL still achieved 91.23 percent accuracy and 82.15 percent recall, and repeated runs with different random seeds showed stable results, with a paired t-test confirming the improvement over the best baseline was statistically significant at p below 0.01. The model also performed best across four specific laundering categories: ransomware, darknet markets, fraud and Ponzi schemes, reducing false positives for fraud and Ponzi cases to 3.12 and 4.25 percent respectively and cutting the false negative rate for the highly concealed Ponzi category to 11.36 percent. An error analysis showed the remaining failures concentrated in low-frequency Ponzi transactions and small-value multi-hop transfers, cases where risk signals are inherently weak and illicit behavior closely mimics normal activity.

The author is candid about the limits. GT-SSL is a static graph model: timestamps and time-step indices are encoded into node features, and the Elliptic experiments use chronological splits to test temporal generalization, but the graph itself is not dynamically updated, communities are computed once before self-training, and no online distribution-shift adaptation is performed. When evaluated on windows progressively farther from the training period, the F1-score declined from 93.50 to 91.21 percent, evidence of partial but not unlimited temporal robustness. The computational cost is also substantial, driven by structure-aware attention, dual-task pre-training and iterative pseudo-label screening. Future work, the study suggests, lies in temporal graph encoders, dynamic community detection and near-real-time incremental inference. Even so, the framework offers regulators something they rarely get: for each flagged transaction, the model can output the high-attention neighbors, the fund-flow sequence and the community consensus score, providing traceable evidence that could survive compliance review rather than a bare, unexplainable alert.

Subject of Research: Self-supervised graph Transformer learning for detecting cryptocurrency money laundering in label-scarce blockchain transaction networks

Article Title: Digital currency money laundering identification model based on graph structure and transformer self-supervised learning

Article References: Shang, Y. (2026). Digital currency money laundering identification model based on graph structure and transformer self-supervised learning. Discover Artificial Intelligence, 6(1), Article 1288. https://doi.org/10.1007/s44163-026-02264-2

Image Credits: AI Generated

DOI: 10.1007/s44163-026-02264-2

Keywords: cryptocurrency, money laundering, blockchain, graph neural networks, Transformer, self-supervised learning, pseudo-labels, Bitcoin, financial crime, anomaly detection, machine learning, transaction graphs

News Source: Denise Maddox. (October 5, 2026). Self-Taught AI Reads Blockchain Money Trails to Catch Crypto Laundering. Scienmag.

Tags: Anomaly DetectionBitcoinblockchaincryptocurrencyfinancial crimegraph neural networksMachine Learningmoney launderingpseudo-labelsSelf-Supervised Learningtransaction graphsTransformer
Share12Tweet7Share2ShareShareShare1

Related Posts

Random Corrosion Pits Reveal Hidden Weaknesses in Bridge Cable Steel Wires

Random Corrosion Pits Reveal Hidden Weaknesses in Bridge Cable Steel Wires

October 5, 2026
Platypus-Inspired Algorithm Brings Animal Sensing to Optimization

Platypus-Inspired Algorithm Brings Animal Sensing to Optimization

October 5, 2026

Ancient Paper Gets a Modern Job: Cellulose Coating Boosts Battery Separator Endurance

October 5, 2026

AI Reads Students’ Faces to Measure Attention in Online Classes

October 5, 2026

POPULAR NEWS

  • Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    29 shares
    Share 12 Tweet 7
  • Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

    29 shares
    Share 12 Tweet 7
  • Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

    29 shares
    Share 12 Tweet 7
  • New Scale Measures How Ready Nurse Educators Really Are for the AI Era

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

Subscribe to Blog via Email

Success! An email was just sent to confirm your subscription. Please find the email now and click 'Confirm' to start subscribing.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.