面向无人车的Transformer多模态BEV融合算法
DOI:
CSTR:
作者:
作者单位:

1.中北大学机器视觉与虚拟现实山西省重点实验室 太原 030051; 2.中北大学智能武器研究院 太原 030051

作者简介:

通讯作者:

中图分类号:

TP391;TN919.8

基金项目:

山西省应用基础研究计划(202403021211093)、山西省应用基础研究计划 (202303021221119)、机器视觉与虚拟现实重点实验室研究基金(447-110103)项目资助


Research on Transformer-based multimodal BEV fusion algorithms for autonomous vehicles
Author:
Affiliation:

1.Shanxi Key Laboratory of Machine Vision and Virtual Reality, North University of China, Taiyuan 030051, China; 2.Institute for Intelligent Weapon Research, North University of China, Taiyuan 030051, China

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    针对无人车传统感知系统模块化设计带来的信息割裂和鲁棒性不足的问题,提出基于Transformer的多模态BEV融合算法。首先,视觉分支基于BEVFormer构建图像时空协同编码机制,结合语义引导的BEV投影增强关键区域表征,缓解光照、运动模糊导致的特征失真问题;其次,雷达分支提出稀疏矩阵状态空间混合模型(Sparse-SSM),通过稀疏体素化与状态转移建模,解决点云稀疏性与不连续性问题。最后,采用Transformer交叉注意力机制实现跨模态特征精准对齐与融合,并在nuScenes数据集上进行训练和实验验证。结果表明,与图像分支和雷达点云分支相比,目标检测指标分别提升15%和23.3%;与先进基线BEVFusion方法相比,所提方法平均精度提升2.7%、目标检测指标提升2.3%,且凭借稀疏设计将浮点运算次数降低20%,推理速度更快,可为无人车提供高精度、强鲁棒性的环境感知能力。

    Abstract:

    To address the issues of information fragmentation and insufficient robustness caused by the modular design of traditional perception systems for unmanned vehicles, this paper proposes a Transformer-based multi-modal BEV fusion algorithm. Firstly, the vision branch, built upon BEVFormer, constructs an image spatio-temporal collaborative encoding mechanism. This is combined with a semantic-guided BEV projection to enhance the representation of key regions, mitigating feature distortion caused by lighting variations and motion blur. Secondly, the radar branch introduces a sparse matrix-state space hybrid model (Sparse-SSM), which tackles the sparsity and discontinuity of point clouds through sparse voxelization and state transition modeling. Finally, a Transformer cross-attention mechanism is employed to achieve precise alignment and fusion of cross-modal features. The model is trained and experimentally validated on the nuScenes dataset. Results demonstrate that compared to the image-only and point cloud-only branches, the proposed method improves the target detection score by 15% and 23.3%, respectively. Compared to the advanced baseline BEVFusion method, our approach achieves a 2.7% increase in average precision and a 2.3% increase in the target detection score, while reducing floating point operations by 20% through its sparse design and achieving faster inference speed. This provides unmanned vehicles with high-precision, highly robust environmental perception capabilities.

    参考文献
    相似文献
    引证文献
引用本文

张欣亚,李郁峰,韩慧妍,田二明,朱堉伦.面向无人车的Transformer多模态BEV融合算法[J].电子测量技术,2026,49(11):25-33

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-09-03
  • 出版日期:
文章二维码

重要通知公告

①《电子测量技术》期刊收款账户变更公告