fig3

A survey on deep reinforcement learning for human-robot interaction

Figure 3. Residual RL architecture [Equation (8)], combining a nominal model-based controller with a learned correction. Only the set-invariance/conformal constraint [Equation (7)] is executable as an online filter that modifies at; the UUB bound [Equation (6)] is a property of the resulting closed loop established by analysis, not a block in the signal path, and is therefore shown separately. RL: Reinforcement learning; UUB: uniform ultimate boundedness.

Intelligence & Robotics
ISSN 2770-3541 (Online)

Portico

All published articles are preserved here permanently:

https://www.portico.org/publishers/oae/

Portico

All published articles are preserved here permanently:

https://www.portico.org/publishers/oae/