Predictive processing as a scalable computational principle for embodied multitask intelligence.

Idei, Hayato; Miyake, Tamon; Ogata, Tetsuya; Yamashita, Yuichi · Sci Adv · 2026

basic_science · Level V

Where this comes from

Abstract

Humans exhibit remarkable flexibility in adapting to diverse and uncertain environments-a hallmark arising from the brain's ability to integrate multimodal sensory streams into coherent predictive models. Drawing on this principle, we introduce a scalable hierarchical multimodal recurrent neural network grounded in predictive processing under the free-energy principle, capable of directly integrating more than 30,000-dimensional visuo-proprioceptive inputs without dimensionality reduction or handcrafted preprocessing. Using sensory data from teleoperation of a full-scale physical humanoid robot performing two caregiving-related tasks-rigid-body repositioning and flexible-towel wiping-the model learns to predict high-dimensional visuo-proprioceptive streams end to end. In open-loop adaptive inference experiments, the framework exhibits three emergent properties: (i) self-organized hierarchical latent dynamics governing task transitions, uncertainty, and occlusion inference; (ii) robustness to degraded vision via multimodal integration; and (iii) asymmetric interference in multitask learning. Although evaluated in simulations, the framework is extensible to closed-loop robot control, with proprioceptive predictions driving action, thereby establishing a generalizable computational foundation bridging brain theory, artificial intelligence, and embodied robotics.