Weljohn Y. Catanpatan, Zildjian Aballe, Christian V. Maderazo
Image captioning models often fail to associate generated nouns with corresponding image regions, limiting interpretability. We propose a dual-objective fine-tuning framework for spatially grounded captioning. Built on Bootstrapping Language-Image Pre-training (BLIP), the method extracts decoder cross-attention for nouns and supervises these maps via a differentiable soft Dice loss jointly optimized with cross-entropy. Training employs a stratified Microsoft Common Objects in Context (MS COCO) subset with a frozen vision encoder to isolate spatial effects. On 2933 mask-eligible test samples, the proposed method improves mean Strict Dice (0.065 vs. 0.005) and Lenient Dice (0.174 vs. 0.013) over baseline fine-tuning. Standard captioning-quality and image-text compatibility metrics remain comparable across fine-tuned tiers, demonstrating improved spatial grounding without degrading linguistic fluency. © 2026 IEEE.
Department of Computer, Information Sciences, and Mathematics, University of San Carlos, Cebu City, Philippines