Domain-Specific Multimodal Translation via Diffusion Models with Semantic Consistency Regularization on the ROCO Dataset
PDF

Keywords

Diffusion models
computer vision
'
natural language processing

How to Cite

[1]
Oluwafemi John Ajala, “Domain-Specific Multimodal Translation via Diffusion Models with Semantic Consistency Regularization on the ROCO Dataset”, JAISE, vol. 1, no. 1, pp. 17–23, Aug. 2026, Accessed: Oct. 02, 2026. [Online]. Available: https://kiwiresearchjournals.com/index.php/jaise/article/view/4

Abstract

Most captioning and image-synthesis systems use generic, domain-agnostic architectures, that do not take into account the presence of clinical terminology structure or the fact that diffusion-based generators tend to hallucinate non-anatomically plausible content that is not relevant to the caption's domain. We propose a domain-specific multimodal translation framework that views image-to-text and text-to-image translation in radiology as a single conditional diffusion process in a shared cross-modal latent space and adds a semantic consistency regularizer that encourages the generated caption (or image) to be close to the cycle-reconstructed version of the source modality. A regularizer is added along with a clinical-concept alignment term computed from a light-weight UMLS-style concept extractor, which helps in maintaining modality-critical entities like anatomical location, laterality, and finding type in the reverse diffusion process. We test the approach on Radiology Objects in COntext (ROCO) dataset with the standard captioning metrics (BLEU-4, ROUGE-L, CIDEr) and clinical concept F1 score, and compare to a CNN-RNN baseline and a CLIP-conditioned autoregressive baseline. The aforementioned experimental settings show that the proposed method can outperform the best baseline by 4.4 points on BLEU-4 and 6.6 points on clinical concept F1, and that the ablation on the number of reverse diffusion steps indicates that the performance gain of the semantic consistency term comes mainly from correcting early trajectory drift, instead of correcting late-stage drift. We also discuss failure modes on rare modality classes and provide a summary of implications for the deployment of diffusion-based translation in report-drafting assistance and educational tools, emphasizing the need for prospective validation on rare modality classes for clinical use. The results of this manuscript are used solely for one benchmark configuration and are not representative for third party verification.

PDF
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 International License.