Abstract:To address the issue of degraded matching accuracy in reflective regions, a stereo matching algorithm is proposed based on a multi-kernel attention decoder and a cost volume dynamic enhancement mechanism. First, the feature extraction network employs a multi-kernel attention decoder, which utilizes parallel depth-wise convolutions to capture spatial details at different resolutions. This enhances the algorithm′s ability to fuse local details and global information, thereby improving disparity prediction accuracy in reflective regions. Next, the algorithm adopts a cost volume dynamic enhancement mechanism to optimize the cost volume. By selectively fusing multi-scale cost vol-umes based on global contextual information, it avoids redundant information accumulation, thus strengthening the cost volume′s modeling capability in reflective regions and mitigating the accuracy degradation problem. Experi-mental results on the Scene Flow dataset show that the proposed method achieves 0.45 EPE and 2.40% D1, while on reflective regions of the KITTI dataset, it attains 4.59% 3-All error, representing an 8.2% reduction com-pared to baseline methods. These results demonstrate that the proposed approach effectively improves matching accuracy in reflective regions while also reducing model parameters.