Abstract
X-ray baggage screening demands robust automated detection systems that can perform reliably under limited annotation. We introduce PRISM-X, a progressive semi-supervised framework that integrates pseudo-label generation, region-level multimodal grounding, and contrastive refinement into a unified detection pipeline. Evaluated on the SIXray and CLCXray benchmarks, PRISM-X consistently outperforms both vision-only and vision-language baselines. On SIXray, it achieves 66.0% mAP and 87.4% AP50, improving over the strongest vision-only method by +6.8% mAP and the leading vision-language model by +6.7% grounding accuracy. On CLCXray, PRISM-X reaches 45.7% mAP, 64.7% AP50, and 51.6% AP75, surpassing Cascade R-CNN by +3.3% mAP and GDINO by +3.4% AP75. These results demonstrate the effectiveness of PRISM-X in low-label, cluttered X-ray scenarios and its superiority over existing weakly and semi-supervised approaches.
| Original language | British English |
|---|---|
| Article number | 104640 |
| Journal | Information Processing and Management |
| Volume | 63 |
| Issue number | 4 |
| DOIs | |
| State | Published - Jun 2026 |
Keywords
- Contrastive learning
- Pseudo-labeling
- Self-supervised learning
- Semi-supervised learning
- Vision-language models
- X-ray baggage detection
Fingerprint
Dive into the research topics of 'PRISM-X: Progressive semi-supervised threat detection in X-ray scans with self-guided multimodal refinement'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver