• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Tuesday, September 8, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

MIMO nonlinear RTGC optimal control using Proximal Policy Optimization (PPO) and Lyapunov stability

Bioengineer by Bioengineer
September 8, 2026
in Technology
Reading Time: 7 mins read
0
MIMO nonlinear RTGC optimal control using Proximal Policy Optimization (PPO) and Lyapunov stability
Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

The world’s ports handle billions of tonnes of cargo every year, and at the heart of the container terminal sits an unglamorous workhorse: the Rubber Tyred Gantry Crane, or RTGC. These mobile giants stack and shuttle shipping containers between ships, trucks and storage yards, and the precision of their movements has an outsized influence on the speed and safety of global trade. Now, a team of researchers from Indonesia has introduced a control strategy that could fundamentally change how these cranes think. By combining the mathematical rigour of Lyapunov stability theory with the learning power of Proximal Policy Optimization, a cutting-edge reinforcement learning algorithm, the researchers have built a crane controller that not only outperforms conventional optimization techniques but can adapt its own behaviour on the fly, even when the load it is carrying is heavier or lighter than expected, or when wind and other disturbances push the crane off course.

The challenge that Steven Bandong of the Calvin Institute of Technology, Selvi Lukman of Bina Nusantara University, and Muhammad Rizalul Wahid of Universitas Pendidikan Indonesia set out to solve is deceptively simple to state but notoriously difficult in practice. An RTGC is what engineers call a multiple-input, multiple-output, or MIMO, nonlinear system. When the crane’s trolley accelerates along its gantry to move a container into position, the payload suspended from the hoisting rope behaves like a pendulum, swinging back and forth in ways that depend on the rope length, the mass of the container, and the forces applied. Lengthen the rope to stack a container high up in the yard, and the pendulum dynamics change entirely; shorten it, and the swing frequency quickens. Most existing anti-sway controllers focus narrowly on two objectives: getting the trolley to its target position and damping the sway. The new work broadens the objective set to three, treating position, sway angle and rope length as simultaneous control targets. That third dimension matters enormously in a real yard, where the crane must travel varying distances and hoist varying heights during a single loading cycle.

The core of the team’s approach rests on Lyapunov stability, one of the foundational tools of control theory. A Lyapunov function, in essence, is a mathematical description of the total energy or “distance from equilibrium” of a system. If a designer can construct a control law that forces this function to shrink monotonically over time, then stability is guaranteed: the states of the system, meaning the crane’s position, sway and rope length, will inevitably converge to their desired values. This guarantee is what makes Lyapunov methods attractive for safety-critical machinery. A controller that merely seems to work in simulation can fail catastrophically under unmodelled conditions, but a Lyapunov-certified controller carries a mathematical proof that it will not drive the system into divergence.

There is a catch, however. Lyapunov control laws contain free parameters, essentially gain coefficients, that shape how aggressively the controller corrects errors. Fixed gains work well near the operating conditions for which they were tuned, but an RTGC operates across a wide envelope of rope lengths, payload masses and travel speeds. Tuning those gains by hand, or even by classical optimization, is a compromise at best. Earlier work by the same group had explored optimizing Lyapunov parameters using established metaheuristic algorithms: Particle Swarm Optimization, which mimics the social behaviour of bird flocks; Simulated Annealing, which borrows the cooling of molten metals to escape poor solutions; and Genetic Algorithms, which evolve candidate solutions through selection and crossover. These methods produce a single, static set of optimal gains, one frozen configuration that must serve every situation the crane encounters.

The new study takes a decisive step beyond static tuning. The researchers employ Proximal Policy Optimization, or PPO, a reinforcement learning algorithm developed by OpenAI researchers and widely regarded as one of the most robust and sample-efficient methods for continuous control problems. PPO works by training a neural network policy that maps observations of the system state to control actions, improving it iteratively through interactions with the environment while using a clipped surrogate objective to prevent destructively large policy updates. Applied to the RTGC problem, PPO does not simply learn a control law from scratch. Instead, it learns to dynamically adjust the parameters of the Lyapunov control law as the crane’s state evolves. In other words, the gains that once were constants become functions of the moment-to-moment state of the system, allowing the controller to be gentle when the crane is near its target and decisive when it is far off or swinging hard, all while the Lyapunov framework preserves the underlying stability guarantee.

The team went one step further by marrying PPO to a Long Short-Term Memory network, producing a hybrid they call LSTM-PPO. Standard feedforward policies see only the current state, treating each instant as independent. But crane dynamics are inherently time-dependent: the sway at this moment carries momentum from the previous moments, and anticipating that momentum allows earlier, smoother correction. An LSTM network maintains an internal memory of past observations, capturing the time-series character of the crane’s motion. When the researchers benchmarked the two learned controllers against PSO, SA and GA-optimized Lyapunov controllers, both PPO and LSTM-PPO delivered superior control performance across the board. The LSTM-PPO variant achieved the best settling time, the duration required for the crane’s states to converge and remain within tolerance of their targets, meaning containers arrive at their destinations faster and with less residual sway.

Robustness testing formed a major part of the study’s credibility. Rather than evaluating the controller only under ideal initial conditions, the researchers subjected it to one hundred randomly drawn starting configurations, spanning the realistic range of positions, sway angles and rope lengths a yard crane might encounter. They also injected payload-mass uncertainty, modelling the common real-world situation in which the controller does not know exactly how heavy the container is, and external disturbances to emulate wind gusts and frictional forces. Across these punishing trials, the proposed controller continued to stabilize the system and drive all three states to their desired values, a result that speaks directly to the practical demands of ports where conditions are never textbook-perfect.

The team also conducted an ablation study, a systematic experiment in which components are removed or varied to isolate their contribution. In this case, the question was which set of observations fed to the LSTM-PPO agent yields the best performance: does the policy benefit from seeing rope length? From sway velocity as well as sway angle? From combinations of state derivatives? By training and evaluating controllers across different observation sets, the researchers identified the most effective sensory configuration for the LSTM-PPO-based Lyapunov controller, providing a practical recipe for future implementations rather than leaving designers to guess.

The implications for port automation are considerable. Container throughput worldwide, measured in twenty-foot equivalent units, has grown relentlessly for decades, and terminal operators face mounting pressure to move containers faster while protecting workers from the fatigue and injury risks that come with operating heavy machinery. Studies of quay crane operators have documented the physical toll of long shifts at the controls, and automating the most repetitive and precision-critical movements is an obvious route to both productivity and safety. A controller that settles faster translates directly into shorter cycle times and more containers moved per hour, while guaranteed stability and robustness to disturbance reduce the risk of the collisions and dropped loads that make crane accidents so costly.

What distinguishes this work within the broader landscape of reinforcement learning for control is its refusal to abandon classical theory in favour of pure machine learning. Deep reinforcement learning has already been applied to cranes and other pendulum-like systems, including triple-pendulum cranes with flexible payloads and unmanned aerial vehicles, but learned controllers that lack stability certificates face scepticism in safety-critical deployment. By keeping the Lyapunov structure as the backbone and using PPO only to breathe adaptive intelligence into its parameters, the researchers get the best of both worlds: the provable convergence guarantees of nonlinear control theory and the state-dependent flexibility of a learned policy. The mathematics constrains what the neural network can do, and the neural network decides how best to do it within those constraints.

The research, published in Neural Computing and Applications, arrives at a moment when ports around the world are investing heavily in automation, from remote-controlled cranes to fully autonomous stacking systems. Controllers of the kind described here, combining certified stability with learned adaptability, could shorten the path from experimental simulation to deployment, because they address the two questions every port engineer must ask: how fast can it settle, and can we trust it when conditions go wrong? On the evidence of one hundred random trials, uncertain payloads and injected disturbances, the answer the LSTM-PPO Lyapunov controller gives is encouraging on both counts. As global supply chains continue to strain under growing demand, the intelligence embedded in the machines that move the world’s containers is quietly becoming as important as the ships and boxes themselves.

Subject of Research: Optimal nonlinear control of Rubber Tyred Gantry Cranes using Lyapunov-based control laws with parameters adaptively optimized by Proximal Policy Optimization and LSTM-enhanced reinforcement learning

Subject of Research: Technology and Engineering

Article Title: MIMO nonlinear RTGC optimal control using Proximal Policy Optimization (PPO) and Lyapunov stability

Article References: Bandong, S., Lukman, S., & Wahid, M. R. (2026). MIMO nonlinear RTGC optimal control using Proximal Policy Optimization (PPO) and Lyapunov stability. Neural Computing and Applications, 38(17), Article 725. https://doi.org/10.1007/s00521-026-12459-4

Image Credits: AI Generated

DOI: 10.1007/s00521-026-12459-4

Keywords: nonlinear control, Lyapunov stability, gantry crane, RTGC, Proximal Policy Optimization, LSTM, deep reinforcement learning, port automation, anti-sway control, MIMO systems, metaheuristic optimization, container handling

Cite Scienmag News
APA MLA Chicago

Denise Maddox. (September 8, 2026). MIMO nonlinear RTGC optimal control using Proximal Policy Optimization (PPO) and Lyapunov stability. Scienmag. https://scienmag.com/mimo-nonlinear-rtgc-optimal-control-using-proximal-policy-optimization-ppo-and-lyapunov-stability/

Denise Maddox. “MIMO nonlinear RTGC optimal control using Proximal Policy Optimization (PPO) and Lyapunov stability.” Scienmag, 8 September 2026, https://scienmag.com/mimo-nonlinear-rtgc-optimal-control-using-proximal-policy-optimization-ppo-and-lyapunov-stability/. Accessed 8 September 2026.

Denise Maddox. “MIMO nonlinear RTGC optimal control using Proximal Policy Optimization (PPO) and Lyapunov stability.” Scienmag. September 8, 2026. https://scienmag.com/mimo-nonlinear-rtgc-optimal-control-using-proximal-policy-optimization-ppo-and-lyapunov-stability/

Copy citation Download RIS

Tags: adaptive crane control systemsadvanced control strategies for gantry cranesadvanced port automation technologiescontainer terminal automationcontainer terminal crane automationcrane motion optimizationdisturbance rejection in RTGCdisturbance-resistant RTGCintelligent port logisticsLyapunov stability theorymachine learning in port equipmentmachine learning in port logisticsMIMO nonlinear RTGC controloptimal control of mobile gantry cranesProximal Policy Optimization reinforcement learningreal-time crane behavior adaptationreal-time robotic crane managementreinforcement learning for industrial machinerystability analysis of crane control systemssustainable and efficient cargo handling

Share12Tweet7Share2ShareShareShare1

Related Posts

A survey of graph neural networks for network intrusion detection systems

A survey of graph neural networks for network intrusion detection systems

September 8, 2026
Exact equations discovered by computing the Gröbner basis

Exact equations discovered by computing the Gröbner basis

September 8, 2026

Graph-based federated reinforcement learning speeds service placement in mobile edge computing

September 8, 2026

Understanding AI’s societal and technical challenges through transdisciplinary research

September 8, 2026

POPULAR NEWS

  • A survey of graph neural networks for network intrusion detection systems

    29 shares
    Share 12 Tweet 7
  • Exact equations discovered by computing the Gröbner basis

    29 shares
    Share 12 Tweet 7
  • Drivers of Mosquito Microbiome Composition: Effects of Species, Locality, Season, and Plasmodium Infection

    29 shares
    Share 12 Tweet 7
  • Graph-based federated reinforcement learning speeds service placement in mobile edge computing

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

A survey of graph neural networks for network intrusion detection systems

Exact equations discovered by computing the Gröbner basis

Drivers of Mosquito Microbiome Composition: Effects of Species, Locality, Season, and Plasmodium Infection

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.