Abstract
The problem of aligning chest radiographs to their semantically equivalent radiology reports is a first step towards zero-shot classification, cross-modal retrieval, and automatic report writing. General purpose vision–language models like CLIP do not perform well in the clinic since the pretraining distribution is often insufficient to include rare radiological lexicon, negation-dense language and nuanced pathology signals. To address the semantic alignment gap between chest X-ray images and radiology reports on MIMIC-CXR, we propose a transformer-based cross-modal contrastive framework, called DAP-CLIP (Domain-Adaptive Prompted CLIP). The three crucial components of our approach are introduced here: (i) a Vision Transformer (ViT-B/16) image encoder and a clinical BERT text encoder are trained together under a symmetric InfoNCE contrastive objective; (ii) a Domain-Adaptive Prompt Pool — a set of learnable, pathology-conditioned prompt tokens injected into both encoders to direct representations to domain-relevant sub-manifolds; and (iii) a lightweight crossattention fusion module to further fine-tune the image–text correspondence at the token level. The performance of DAP-CLIP is compared on the MIMIC-CXR v2.0.0 benchmark (377,110 studies) where it outperforms BioViL by +3.6 points for mean zero-shot AUROC on eight CheXpert pathologies, and MedCLIP by +5.2 points for image-to-report retrieval at R@10. The domain-adaptive prompt pool is the single largest factor driving semantic alignment quality as demonstrated by ablation studies and t-SNE visualization of the cross-modal clusters yield significantly tighter alignments. In conclusion, DAP-CLIP shows the feasibility and parameter efficiency of designing a prompt-based approach to domain adaptation towards clinical quality multimodal alignment.

This work is licensed under a Creative Commons Attribution 4.0 International License.
