
Editor in chief:Prof. Sun Shenghe
Inauguration:1980
ISSN:1002-7300
CN:11-2175/TN
Domestic postal code:2-369
- Most Read
- Most Cited
- Most Downloaded
Fang Yanyan , Lin Zhiling , Li Ningyu , Zhang Zhijian , Li Dan
2026, 49(12):1-9.
Abstract:To address the accuracy degradation in detecting occluded and complex stacked parts, an improved YOLO11-based object detection algorithm named YOLO11-MBM is proposed. The method integrates a multidimensional collaborative attention (MCA) mechanism and employs weighted feature fusion. The main strategies include three aspects: Firstly, an MCA module is incorporated into the C3k2 module of the backbone network to fuse channel, spatial and scale attention, which enhances the ability to distinguish texture-similar parts and effectively suppresses false positives caused by texture confusion. Secondly, a weighted feature fusion architecture BiFPN-SFF-Concat is designed in the feature fusion network. Through bidirectional feature propagation and a scale-sensitive dynamic feature weighting strategy, it improves feature complementarity and detail representation in occluded scenarios, reducing missed detections. Finally, the minimum point distance intersection over union (MPDIoU) is adopted as the loss function to improve localization accuracy and convergence speed. Experimental results on the extended Baidu Paddle industrial parts dataset show that YOLO11-MBM achieves an mAP50 of 87.8% and an mAP50:95 of 58.0%, outperforming the baseline model by 6% and 9% respectively, while maintaining realtime performance (110 fps on V100 GPU) and lightweight design (with only a 1.1×106 parameter increase). The experimental results confirm its practicality in complex industrial inspection scenarios and furnish a feasible technical solution for visual inspection systems in smart manufacturing.
Wang Shumin , Zhai Meitao , Wang Pengyu , Ma Jiankang , Yi Xiaofeng
2026, 49(12):10-16.
Abstract:To meet the practical demand for continuous monitoring of disaster-inducing water bodies in coal mine floors, an intrinsically safe STM32-based distributed resistivity monitoring system for mining is designed. The high-density resistivity method, with advantages of high resolution, intuitive results and continuous monitoring capability, is widely used for water inrush monitoring in coal mine floors. STM32 offers abundant peripheral interfaces, and its low power consumption and high programmability align with the design requirements of intrinsically safe resistivity systems. With intrinsic safety as the core premise, the system′s host and electrode converter are developed. The host handles signal transmission and data acquisition, mainly consisting of the STM32 control chip, an H-bridge regulated transmission circuit, and an ADS1256 acquisition-transmission circuit. The electrode converter is responsible for electrode state switching, comprising the STM32 control chip and a solid-state relay array circuit. Experiments verify the system′s stability and accuracy. Results indicate that the resistivity system meets intrinsic safety standards, accurately acquires electrical information, and correctly reflects geological structures.
2026, 49(12):17-24.
Abstract:To address the degradation of WiFi fingerprinting accuracy caused by multipath effects and signal obstruction in complex indoor environments, this paper proposes an RSSI-based localization algorithm integrating a convolutional neural network (CNN) and a bidirectional long short-term memory network (BiLSTM). The method leverages the spatial feature extraction capability of CNN and the temporal sequence modeling strength of BiLSTM to achieve deep spatiotemporal feature fusion. MATLAB simulation experiments compare the proposed model with CNN, CNN-LSTM, LR, KNN, RF, GBDT and SVM. Results show that the proposed CNN-BiLSTM achieves the best performance, reducing average, maximum and standard deviation of positioning errors by 52.1%, 53.2% and 55.6%, respectively. Monte Carlo trajectory experiments further demonstrate the superior accuracy and robustness of CNN-BiLSTM in continuous localization. The proposed model effectively suppresses RSSI fluctuations and provides a feasible solution for high-precision indoor WiFi fingerprinting.
Huang Jianping , Wang Haozhou , Hong Wei , Song Wenrui , Song Wenlong
2026, 49(12):25-37.
Abstract:The charge management system for test masses in space inertial sensors requires high-frequency pulsed driving of ultraviolet (UV) LEDs to control surface charge, yet traditional Howland current sources face critical bottlenecks—output impedance attenuation, limited bandwidth and instability under high-frequency square-wave driving—hindering precise charge control and space missions. To address this, we propose a dual-buffer load-included Howland (DB-LIH) current source. By designing a dual-buffer Howland (DBH) circuit with dual voltage followers isolating feedback paths and a 0.01%-matched precision resistor network LT5400 suppressing mismatch effects, the DB-LIH architecture incorporates the load into the feedback loop to block stray capacitance currents. Simulations show the output current flatness factor (FF) induced by stray capacitance drops from 18% (DBH) to 0.03% (DB-LIH), with bandwidth reaching 1 MHz. Experimental results show that output accuracy better than 4.27% (0.05~50 mA range), resolution deviation less than 1.67% and 1 mHz~1 Hz relative amplitude stability of 0.01 Hz-1/2. This work resolves high-frequency constant-current precision and stability issues, meeting stringent requirements for space gravitational wave detection.
Han Haoran , Wang Wei , Xu Bingyin , Chen Heng , Sun Zhongyu
2026, 49(12):38-49.
Abstract:Distribution networks contain numerous hybrid cable-overhead connections and branches, often causing excessive attenuation of fault-induced traveling waves, preventing traveling-wave fault location devices from activating. To address this issue, this paper proposes an placement method for device placement that ensures reliable fault location across all network segments by accounting for traveling-wave transmission attenuation. Using ground faults as an example, the study begins with a quantitative calculation of the initial voltage traveling wave. It then systematically analyzes the attenuation effects introduced by hybrid cable-overhead connection points and multi-branch T-junctions through quantitative modeling. By constructing a directed attenuation matrix for fault traveling wave propagation in a predefined network topology, the method identifies non-measurable lines in double-ended traveling-wave ranging based on the device detection threshold. Furthermore, by analyzing the transmission range of fault-induced traveling waves in non-measurable lines, additional devices are allocated at optimal nodes with the objective of minimizing the total number of deployed devices. The experimental results show that the placement method ensures the reliability of traveling wave distance measurement for the entire network due to the comprehensive consideration of various factors affecting the amplitude of fault traveling waves.
Shi Wei , Ren Wenbo , Zhang Yan , Jiang Shi , Xu Changzhong
2026, 49(12):50-57.
Abstract:A four-point perspective correction method based on random sample consensus (RANSAC) improved generalized iterative closest point (GICP) point cloud registration algorithm is proposed to solve the problem of the deviation of the grasping points caused by the machining error of the tray, the positioning error and the absolute positioning accuracy deviation of the robot during the orderly grasping process. Firstly, the positioning hole contour and the initial center coordinates are extracted by point cloud processing technology, and the accurate center coordinates are extracted by the improved GICP point cloud registration algorithm. Secondly, the perspective transformation matrix is used to correct the deviation of each grasping point on the tray. Finally, the robot is controlled to grasp according to the corrected point. The experimental results show that in the point cloud registration, the maximum deviation range of the extracted center of the circle decreases from [1.5 mm, 2.8 mm] to [1.2 mm, 1.6 mm]; in the point correction, the four-point perspective correction reduces the maximum deviation in the XY direction of different grasping points from 3.1 mm to 1.1 mm. And in the workpiece grasping experiment on a fully loaded tray, the maximum deviation in the XY direction is 1.0 mm. In summary, the method proposed in this paper can effectively correct the workpiece grasping point, improve the success rate and reliability of grasping, and meet the actual production needs.
Zhou Zhijian , Li Xiaodong , Jing Zhiwei , Wang Chao , Wang Yanzhang
2026, 49(12):58-71.
Abstract:Field source boundary identification is an essential task for interpreting field data, and normalized magnetic source strength (NSS) has become the main method for interpreting magnetic data because it is not affected by magnetization direction. This paper addresses issues related to near-field boundary identification, the divergent effects of single and multiple magnetic source boundary recognition, and the challenges of identifying the geometric boundaries of magnetic sources that are far apart and experiencing mutual coupling during multiple magnetic source boundary identification. A new magnetic target boundary identification method has been proposed, which uses second-order partial derivatives in the inclined angle boundary identification method and introduces NSS. This method is applied to the boundary identification of a single rectangular magnetic source, three rectangular magnetic sources in the same plane, as well as boundaries of deep and shallow anomalies. The simulation results indicate that the method proposed in this paper not only leverages the advantage of NSS in reducing the coupling effects between magnetic sources and the insensitivity of depth to inclination angle, but also allows the recognized magnetic source boundaries to converge more effectively by using second-order partial derivatives.
Yang Xuhong , Zhang Haohan , Zhang Nan
2026, 49(12):72-81.
Abstract:The modular multilevel matrix converter (M3C) has shown promising application prospects in medium-to high-voltage high-power scenarios such as large-scale motor drives and wind power integration, owing to its modular structure, high reliability and superior power quality. However, in traditional cascaded control strategies, the outer-loop proportional-integral(PI) controller exhibits limited response speed, while the inner-loop PI controller lacks sufficient robustness against system parameter perturbations and external disturbances, constraining the dynamic performance of the M3C under complex operating conditions. To address these issues, this paper proposes a cascaded control strategy for the M3C based on an inner-loop robust backstepping sliding mode control. This strategy retains the outer-loop PI control to ensure steady-state error-free tracking of power commands, with the core innovation lying in the inner-loop adoption of robust backstepping sliding mode control. By recursively designing the control law via the backstepping approach and incorporating a sliding mode term, it effectively integrates precise linearization with strong disturbance rejection capabilities. The global asymptotic stability of the closed-loop system is rigorously demonstrated based on Lyapunov stability theory. MATLAB/Simulink simulation results indicate that, compared to the conventional dual-PI cascaded control, the proposed strategy offers faster dynamic response, reduced overshoot and enhanced parametric robustness under transient conditions such as load sudden changes and grid voltage asymmetry, significantly improving the system′s control performance.
Wei Zhaochuan , Liu Zhiheng , Ji Yuanfa
2026, 49(12):82-89.
Abstract:To address the problems of multipath interference, high complexity and unstable performance of traditional detection algorithms in multiple-input multiple-output-orthogonal time-frequency space (MIMO-OTFS) systems under high-speed mobile scenarios, a detection algorithm (CG-MMSE-ADK-Best) for MIMO-OTFS systems based on the joint optimization of low-complexity minimum mean square error (MMSE) and K-Best is proposed.First, the conjugate gradient (CG) algorithm is introduced to iteratively solve the MMSE matrix equation, avoiding direct matrix inversion to reduce computational overhead.Second, CG-MMSE is combined and optimized with K-Best detection. The fixed path number strategy of the traditional K-Best algorithm is improved, and a dynamic K-value algorithm is proposed to further optimize complexity.Finally, the algorithm performance is verified through simulation experiments. The simulation results show that with 4-QAM modulation, path number K=8 and signal-to-noise ratio (SNR) of 20 dB, the proposed algorithm improves the bit error rate (BER) performance by 45.9% compared with the traditional K-Best algorithm. After integrating the dynamic path number algorithm, the BER curve is basically consistent with that of the fixed path number algorithm, but the computational complexity is significantly reduced due to the decrease in path number.The proposed algorithm can improve BER performance under the same path number and reduce computational complexity under the same BER, effectively solving the problems of high complexity and unstable performance of traditional algorithms. It provides technical support for the application of large-scale MIMO-OTFS systems in high-speed mobile scenarios.
An Hongyu , Wang Yaguo , Zhao Chenyang , Bai Yu
2026, 49(12):90-99.
Abstract:This paper presents a designated-time adaptive output feedback control scheme for quadrotors affected by sampling constraints. First, based on the principle of invariant manifolds, an efficient filter-based unknown system dynamics estimator is developed for overall uncertainty estimation. The unknown system dynamics estimator (USDE) design introduces a variable-threshold event-triggered state observer to replace the actual velocity measurements, maintaining estimation accuracy at low computational cost. Secondly, considering the strict response speed requirements of the quadrotor, an appointed-time funnel control (ATFC) scheme is designed to ensure that displacement errors converge within the predetermined time and to constrain error boundaries, thereby enhancing the system′s transient and steady-state performance. Finally, the stability of the closed-loop system is proven through Lyapunov analysis theory, and simulation results validate the effectiveness and superiority of the proposed method.
Hou Xinmeng , Yang Guang , Zeng Fengying , Zhao Jiyuan , Lu Sipeng
2026, 49(12):100-109.
Abstract:To address the issues of excessive noise and blurred edges in X-ray images of aero-engine rotor blades, this paper proposes a noise self-attention (NSA) dual-decoder image denoising network, which effectively enhances image quality. The network introduces a noise self-attention mechanism between the encoder and decoder to strengthen the perception of image features under noise interference. A dual-decoder structure is adopted to perform denoising and edge preservation separately, and a learnable gated fusion mechanism integrates the outputs of the two decoders, enabling the final result to simultaneously suppress noise and preserve blade edge structures. Experimental results demonstrate that the proposed method outperforms comparative algorithms in visual effects for denoising rotor blade X-ray images, achieving a peak signal-to-noise ratio (PSNR) of 33.26 dB and a structural similarity index (SSIM) of 0.883 7, proving the effectiveness of the proposed approach and providing a reliable X-ray image denoising solution for aero-engine blade inspection.
2026, 49(12):110-119.
Abstract:To address the issues of low precision, false detections and missed detections for small objects in UAV aerial images caused by large scale variations and complex backgrounds, an improved algorithm named CSFF-DETR based on RT-DETR is proposed. First, a context-aware enhancement module is designed in the backbone to improve small object feature extraction and suppress complex background interference. Second, a high-resolution P2 feature layer is introduced into the neck network, along with a novel cross-scale feature fusion module, to enhance interaction between features at different scales and boost small object detection capability. Finally, a multi-path fusion downsampling module is designed to better preserve and extract fine-grained details of small objects. Experiments on the VisDrone2019, Tinyperson and UAVVaste datasets show that compared to the baseline RT-DETR model, CSFF-DETR achieves mAP50 improvements of 3.6%, 1.0% and 2.8% and mAP50.95 improvements of 3.0%, 1.1% and 1.6%, respectively, while reducing parameters by 40.7%. This meets the requirements for both detection accuracy and lightweight deployment on UAV platforms.
Chen Zhihong , Fan Bishuang , Chen Mengyuan , Zhang Hao
2026, 49(12):120-129.
Abstract:To enhance the accuracy of photovoltaic cell defect detection in complex scenarios such as multiple categories and small targets, this paper proposes a lightweight improved model based on RT-DETR. Firstly, a lightweight dual-path feature extraction module was designed to replace the basic module of ResNet. This not only reduced the model parameters but also enabled the model to possess the ability of local and global feature modeling. Secondly, the multi-head attention module in (attention-based intra-scale feature interaction,AIFI) was improved by deleting some attention heads and performing linear transformations on the remaining attention heads to alleviate the redundancy problem in the multi-head attention mechanism. Finally, the convolutional downsampling module in the network was replaced by a downsampling module based on wavelet transformation to improve the problem of edge information loss during downsampling. The experimental results show that the number of parameters of the improved model has decreased by 47.3% compared to the original model, and the computational cost has decreased by 39.7%. On the private dataset, the improved model outperformed the baseline with increases of 2.2%, 4.1% and 2.4% in precision, recall and mAP@50, respectively, demonstrating its effectiveness. In addition, a generalization experiment was conducted on the public PVEL-AD dataset, where the improved model achieved an mAP@50 of 70.6%, which is 6.3% higher than the baseline model, providing initial evidence of its generalization ability on unseen data.
Sheng Xile , Lin Yinrui , Zhang Zhang
2026, 49(12):130-138.
Abstract:The physical separation of perception and computation in von Neumann architectures necessitates extensive data movement between storage and processing units, resulting in high latency and energy consumption that constrain the development of edge vision applications. To reduce data transmission overhead and enhance front-end computational efficiency, this study proposes a hybrid recognition scheme based on optoelectronic memristor arrays and convolutional neural networks (CNN) for low-bandwidth feature acquisition and classification of 64×64 handwritten characters. Two sets of 4k optoelectronic memristor arrays perform parallel one-dimensional convolutional compression along row and column directions respectively, yielding two 64-dimensional feature streams that are concatenated into a 128-dimensional joint feature. A lightweight CNN is employed in the backend, combined with mixed-precision quantisation. Experimental results demonstrate that the combined row-column features achieve stable recognition rates of approximately 91%~93% without quantisation. Following quantisation and fine-tuning, the model attains a peak recognition rate of approximately 88.8%.
2026, 49(12):139-145.
Abstract:Physics-informed neural network with coordinate inputs is susceptible to spectral bias in acoustic field reconstruction, making it difficult to accurately represent the high-frequency components of the sound field. This paper formulates acoustic field reconstruction under limited sound pressure measurement conditions as an image inpainting problem, and proposes a physics-informed neural network embedded with Fourier feature positional encoding. The proposed method not only embeds the Helmholtz equation into the loss function to enforce physical constraints, but also maps spatial coordinates into a high-dimensional feature space using sine and cosine functions of multiple frequencies, thereby effectively modulating the network′s ability to capture sound field variations across different spatial frequency scales. Simulation results demonstrate that compared with the coordinate-based physics-informed neural network method, the proposed method reduces the reconstruction error by more than 10 dB, and the number of iterations required for convergence is only 25% of that of the former. Loudspeaker experiments further validate the effectiveness of the proposed method in practical sound field reconstruction.
Wang Fucai , Mao Xiaoqian , Fan Chunling
2026, 49(12):146-156.
Abstract:In traditional event-related potential (ERP) research, due to the varying latencies and mechanisms of different components, studies typically focus on specific features for signal extraction and analysis, which limits the exploration of their interactions and hinders the understanding of the brain′s overall mechanism in processing visual stimuli. To address this issue, this study selects P1, N170 and P3 as target ERP components and proposes a dynamic weight assignment-based ERP feature extraction and fusion strategy. First, time-frequency domain processing is applied to ensure consistency in dimensionality across feature segments. Next, the random forest algorithm was used to evaluate the Gini importance of each component for each participant under different experimental paradigms, followed by dynamic weight assignment to the feature segments, resulting in a multi-component fused feature. Finally, convolutional neural networks are employed to classify the fused data, verifying the significant classification effect of the fused features. Experimental results show that, for the Face Perception N170 and Active Visual Oddball P3 datasets, classification accuracies reached 96.4% and 92.4%, respectively, with improvements of 11.6% and 10.2% over non-fused features. It improved by 8.4% and 5.3% compared to equal-weight fusion, and was more effective than the fusion features with attention mechanism, proving that the weighted feature fusion method proposed in this paper can enhance the relevant component signals and improve classification accuracy, providing new insights into the overall process of visual stimulus processing in the human brain.
Zhao Shuanfeng , Wang Heng , Bai Jiale , Dong Jintao
2026, 49(12):157-165.
Abstract:The measurement accuracy of dynamic weighing systems is easily affected by the coupling of multiple factors, leading to significant system errors. This paper addresses the issue that the individual responses of each quartz wafer within traditional quartz sensors are easily masked by the overall signal, making it difficult to conduct in-depth analysis of the error propagation mechanism within the sensor and thus challenging to effectively compensate for the accuracy of dynamic weighing systems. A 16-channel independent synchronous acquisition system for quartz wafers was designed and implemented, including quartz weighing sensors, charge amplifiers and synchronous acquisition devices, enabling the 16 quartz wafers that were originally connected together to be separately collected and analyzed for signals. Based on this experimental platform, a multi-source error decoupling and propagation model was proposed. By reconstructing the signals through an adaptive weighted fusion algorithm and optimizing the BP neural network with the ant colony optimization algorithm, the accuracy and robustness of the dynamic weighing system were significantly improved. Experimental results show that compared with the traditional model, the average relative error was reduced from 4.01% to 1.81%, and the maximum relative error was reduced from 10.26% to 2.22%. This provides new theoretical and practical references for high-precision dynamic weighing technology.
2026, 49(12):166-178.
Abstract:Aiming at the problem that the existing steel surface defect detection algorithms are difficult to balance resource consumption and detection accuracy, an improved lightweight steel surface defect detection algorithm based on YOLO11s (EEC-YOLO) was proposed. Firstly, the ADown downsampling module was used to alleviate the loss of fine-grained information and realize the model lightweight. Secondly, a novel enhanced pooling channel attention (EPCA) attention mechanism is proposed and embedded in the YOLO11s backbone network to strengthen multi-scale defect feature extraction from the frequency domain dimension and reduce the lack of feature information. Finally, an enhanced screening feature pyramid network (ES-FPN) was designed, and the feature fusion network of YOLO11s was reconstructed by using the enhanced local attention (ELA) mechanism and group normalization to optimize the feature selection and fusion effect, so as to improve the attention and detection ability of the model for small targets. On NEU-DET dataset, compared with YOLO11s, its mAP@0.5 is improved by 3% to 79.8%, the number of parameters is reduced by 45.1%, and the amount of calculation is reduced by 30.5%. On the GC10-DET dataset, mAP@0.5 is increased by 3.2% to 65.6%, the number of parameters is reduced by 45.0%, and the amount of calculation is reduced by 30.2%. The proposed algorithm achieves a good balance between detection accuracy, computational cost and efficiency, and provides strong support for the industrial landing of edge terminal devices.
Fu Qiang , Zhong Zhen , Ji Yuanfa , Ren Fenghua
2026, 49(12):179-188.
Abstract:To address the limitations of most visual simultaneous localization and mapping (SLAM) systems in terms of positioning accuracy and robustness within indoor dynamic environments, this paper proposes a real-time semantic SLAM visualization system based on an enhanced Photo-SLAM algorithm framework. First, a lightweight semantic segmentation network is added to the original algorithm framework to remove dynamic object features from the image. Additionally, an adaptive edge expansion method is introduced to eliminate residual dynamic feature points caused by insufficient edge detection in semantic segmentation, further enhancing the system′s localization accuracy in dynamic environments. Second, semantic information is introduced into the 3D Gaussian Splatting process to remove the Gaussian bodies of dynamic objects, constructing a static, smooth, and continuous 3D Gaussian map. Finally, a dense point cloud mapping thread is added to the Photo-SLAM algorithm framework to construct a dense point cloud map of the static background based on semantic information and key frames, and multiple filtering methods are added to further remove the influence of outliers on the point cloud map. Experimental results on the TUM dataset show that the improved algorithm achieves over 40% higher localization accuracy than the Photo-SLAM algorithm in low-dynamic scenes and over 93% higher accuracy in high-dynamic scenes. Compared to other dynamic SLAM algorithms, the proposed algorithm achieves higher localization accuracy in most scenarios and is more real-time. It can also create 3D Gaussian maps and dense point cloud maps after removing the interference of dynamic objects, enabling visualization of the static background.
Lan Zhangli , Wen Hong , Zhang Hong , Zhang Yong , Chen Xi
2026, 49(12):189-201.
Abstract:Automatic detection of traffic signs is crucial for improving the robustness of environmental perception in autonomous vehicles. In real-world scenarios, traffic signs are often occluded, resulting in issues such as complex background interference, small target scale and incomplete sign information, which degrade detection performance. To address the challenge of occluded traffic signs in complex backgrounds, we propose a feature extraction method C2f_color based on color correlogram, which leverages the prominent and stable color features of traffic signs to enhance their distinguishability in challenging scenes. For small-scale targets and low-resolution problems, we introduce a feature enhancement module C2SEAM that integrates a self-ensembling multi-scale attention mechanism (SEAM), which improves the model′s ability to extract fine-grained features across multiple scales and enhances detection accuracy for small objects and low-resolution samples. Additionally, to handle missing sign information due to occlusion, we propose contextmixing convolution(ConMix), which dynamically fuses contextual information to compensate for missing features and improve the representation of occluded traffic signs. Experiments on the CCTSDB public dataset and our custom traffic sign dataset under occlusion and complex conditions (TSOCC) show that our method improves mAP@0.5 by 8.4% and 3.5%, respectively, while P increases by 0.9% and 6.2%, respectively, compared with the baseline model.
Tang Shancheng , Wang Yan , Zhou Tong
2026, 49(12):202-213.
Abstract:To address the existing issues in image restoration methods when repairing damaged murals, such as insufficient attention to contextual information, image blurring and texture inconsistency, this paper proposes a learning based texture coherence and continuity for mural inpainting model(LTCC-MIM). Based on the theory of masked autoencoders, this paper first employs an image inpainting network with a Transformer architecture to integrate local features and global contextual information of the image, thereby enhancing the understanding of interrelationships among various parts of the image as well as its overall structure. Next, a contextual feature enhancement module (CFEM) is designed to achieve regional joint learning and multi-scale feature aggregation, strengthening local texture continuity and generating fine image details to alleviate image blurring. Structural similarity metrics are incorporated as constraints to prioritize mural-specific texture patterns, thereby harmonizing inconsistencies and optimizing restoration quality. Experimental results on digital restoration of real murals demonstrate that the proposed method outperforms comparative methods in both subjective and objective evaluations. Specifically, it achieves an average improvement of 2~8 dB in peak signal-to-noise ratio (PSNR), 2%~10% in structural similarity index (SSIM) and a reduction of 2%~9% in learned perceptual image patch similarity (LPIPS). Additionally, the fr-chet inception distance (FID) score decreased by an average of 1~5 units.
Liu Li , Zhang Chuxia , Zhang Shuo , Li Yujian , Wang Qiang
2026, 49(12):214-225.
Abstract:To address the challenges of insufficient feature representation, strong background interference and low localization accuracy in small traffic sign detection, an improved multi-scale feature enhancement algorithm TSD-YOLOs based on YOLOv11 is proposed. Firstly, a CSP-MSFPF module is designed using heterogeneous partial convolutions and cross-scale feature fusion, effectively enhances the feature expression of small targets. Secondly, a foreground enhancement feature pyramid network is constructed in the neck of the network to improve target focus and suppress background noise through contextual correlation and foreground enhancement. In addition, the Haar wavelet downsampling is adopted to retain high-frequency details while reducing computational overhead, alleviating information loss for small targets. Meanwhile, the WIoU v3 loss function is utilized to dynamically adjust gradient allocation and optimize sample distribution, improving localization accuracy. Finally, a channel pruning strategy based on the LAMP score is applied to compress the improved model, which significantly reduces model complexity and computation while maintaining detection performance. Experimental results demonstrate that compared to the baseline model, TSD-YOLOs achieves 2.6% and 2.1% improvements in mAP50 and mAP50:95 respectively on the TT100k dataset, with a detection rate of 242 fps; On the CCTSDB dataset, mAP50 and mAP50:95 improved by 1.6% and 3.9%, respectively. Furthermore, the computational cost is reduced by 27.6%, the model size is compressed by 58.7%, and the number of parameters is reduced by 59.6%. The experimental results adequately verify that TSD-YOLOs significantly improves the detection accuracy and robustness of small-target traffic signs in complex scenes while ensuring real-time performance.
Hao Jiangpeng , Niu Fanglin , Yu Ling , Han Chi
2026, 49(12):226-238.
Abstract:To address the challenges of small target sizes, significant scale variations, complex backgrounds and limited computational resources in UAV aerial images, this paper proposes an improved YOLOv11-based algorithm, termed LMH-YOLO. First, we introduce C3k2_PMDRB to enhance multi-scale semantic features generated by cascade expansion paths while reducing the model′s parameter count. Second, we design the LCFI_Net architecture to enable efficient cross-layer information fusion and accurately capture the spatial locations of small targets. Subsequently, a lightweight detection head, LGDHead, is developed to reduce computational overhead and improve efficiency. Finally, the Wv3-MPDIoU loss function is proposed to optimize model convergence and mitigate missed detections. Experimental results demonstrate that on the VisDrone2019 dataset, LMH-YOLO achieves a 4.8% improvement in mAP@50, along with a 42.3% reduction in parameter count and a 32.7% decrease in model size. On the TinyPerson dataset for extremely small targets, mAP@50 increases by 6.7%, while parameter count and model size decrease by 42.3% and 30.8%, respectively. These results indicate that LMH-YOLO attains an optimal balance between performance and model compactness, making it particularly well-suited for small target detection in UAV aerial imagery.
Mo Taiping , Chen Liangwei , Zhang Xiangwen , Sun Peng
2026, 49(12):239-249.
Abstract:Aiming at the problems of low detection accuracy of small targets in the quality inspection of bamboo chopsticks and leakage of detection when multiple targets are gathered and overlapped, this paper proposes an efficient bamboo chopstick slice detection method based on an improved real-time detection Transformer (RT-DETR), named EBCS-DETR. SDSEM module is designed to enhance the expressive capability of shallow features in ResNet18, improving the representation of fine-grained details. Second, FasterEMA module is introduced to reduce the number of parameters in deep layers and accelerate feature extraction in the backbone network. Third, the neck network is enhanced by integrating CARAFE module with Bi-FPN, which improves localization accuracy for small defective regions, strengthens discrimination at object boundaries and suppresses background noise. Experimental results on a self-constructed dataset of 5 400 images demonstrate that the proposed EBCS-DETR achieves mean average precision (mAP50) improvements of 3.3%, 3.4%, and 3.1% over the baseline model across three test sets, consistently outperforming other state-of-the-art models with comparable parameter counts. In addition, robust experiments further confirmed the reliability and application potential of EBCS-DETR for detecting defects in bamboo chopsticks slices, which is crucial for complex production environments. This work provides a novel and effective solution for high-precision automated inspection of bamboo chopsticks, offering valuable insights for intelligent quality assurance in similar manufacturing scenarios.
Chen Nan , An Ran , Zhang Liping , Hao Rudong
2026, 49(12):250-257.
Abstract:This paper presents an image acquisition and cropping technique for industrial inspection systems. Initially, it identifies the transmission bottleneck in existing industrial optical detection and elaborates on the algorithmic principle of dynamic ROI cropping based on edge points. By performing real-time Canny edge detection to localize feature points, vertical ROI windows with configurable pixel widths are generated to achieve data compression. Subsequently, modifications to the GVSP are proposed, including the addition of Index packets and adaptation of Leader packets to accommodate dynamic ROI-based image data transmission. The implementation of the dynamic ROI algorithm on an FPGA platform is then detailed, where BRAM substitution for 1 280-bit registers optimizes storage resources, and a 128-bit bus design eliminates internal bandwidth bottlenecks. Experimental verification and industrial deployment demonstrate that the system significantly reduces effective data volume while preserving raw grayscale precision, enabling reliable multi-camera data transmission.

Editor in chief:Prof. Sun Shenghe
Inauguration:1980
ISSN:1002-7300
CN:11-2175/TN
Domestic postal code:2-369