Quadratic effects of linear interpolation between permutation-aligned neural networks.

Benkő, Beatrix · Neural Netw · 2026

basic_science · Level V

Where this comes from

Abstract

Linear weight interpolation in the parameter space has emerged as a versatile technique for regularizing deep neural networks and beyond. It can provably improve convergence speed, underpins model fusion via averaging, while also serving as a method to analyze the loss landscape-particularly in studying the mode connectivity of minima. Recent work has showcased settings where permutation alignment enables linear interpolation between final parameters of independently trained networks without a substantial loss increase. We provide insights into this phenomenon through the lens of local quadratic approximations. We demonstrate that such approximations can accurately capture the loss behavior along the path between permutation-aligned solutions, but they systematically fail between solutions not aligned to one another. Our investigation reveals that significant higher-order variations in loss arise from interpolating misaligned units within corresponding layers, and reordering them via permutation effectively reduces these variations. The remaining non-linear effects diminish with increasing network width. We also examine, from the quadratic viewpoint, differences in loss behavior and model functionality across multiple interpolation directions, beyond what is indicated by aggregated performance metrics. Moreover, we identify small but consistent classification performance gains when interpolating between solutions, and we explain them via logit-level approximations, clarifying the underlying mechanisms.