Abstract:Steel rail surface defect segmentation in complex environments is often constrained by the bulky size of imaging systems and their sensitivity to focusing, which lead to stability issues. Although lensless imaging technology can effectively address this limitation, its reconstruction-then-segmentation pipeline suffers from high computational cost and error accumulation. To address these issues, this paper focuses on rail surface defect segmentation form lensless encoded images and proposes a multi-scale separable optical-aware estimator prior-guided modulation and fusion network (MSSOE-PGMFNet). Firstly, a lensless imaging acquisition system is constructed, and a lensless version is established based on the Northeastern University rail surface defect detection dataset with augmentation (NEU RSDDS-AUG), providing fundamental data support for encoded-domain segmentation research. Secondly, considering that the existing single-scale modeling estimation methods are difficult to simultaneously cover fine cracks and large-area defects, a multi-scale separable optical-aware estimator is designed to enhance the response ability to defects at different scales, and then output more segmentation-friendly task-related representations. However, significant defect shape variations make feature fusion prone to spatial response inconsistency. Therefore, this paper designs a spatial mixed prior-gated module to generate stable and transferable spatial gating signals. Through gated modulation, the module selectively enhances and suppresses the fused information, thereby improving the focus and boundary consistency of defect regions. Finally, experimental verification is conducted on the constructed lensless dataset. The results show that the proposed method achieves the best overall performance among comparison methods including U-Net, DeepLabV3+, TransUNet, LOINet, RecSegNet, and FDTDNet. With Fωβ, MAE, Eε, Sα, Dice, and IoU reaching 0.886, 0.057, 0.905, 0.895, 0.903, and 0.842, respectively. Ablation experiments further verify that both MSSOE and PGMFNet contribute positively to performance improvement, and their combination achieves the best results. This method does not require explicit reconstruction and provides a feasible solution for steel rail surface defect segmentation under lensless encoded imaging conditions.