Skip to content
FRFederico Raponi
All work

ProjectMar – Jun 2025

Beyond Contrastive Learning: Variational Image–Text Retrieval with RoBERTa and DINOv2

We adapt VMSST, a variational model from multilingual retrieval, to images and text. It stays within about 6.5 points of a CLIP baseline on MS-COCO, beats it in zero-shot transfer from Flickr30k to COCO, and keeps a structured latent space that a contrastive model does not have.

Contrastive models like CLIP are the standard for image–text retrieval, but they only learn to pull matching pairs together: the embedding space has no structure beyond that, and it cannot generate anything. We wanted to measure what it costs to ask for more, a latent space that separates shared meaning from modality-specific detail, and what you get in return.

Method

We adapted VMSST (Variational Multi-Modal Semantic-Style Translation), originally built for multilingual text retrieval, to the vision–language setting.

  • Frozen backbones, cached features. DINOv2 for images and RoBERTa Large for text, with all embeddings pre-computed once so the training budget goes to the alignment modules.
  • Two latent spaces per modality. A semantic latent shared across image and text, used for retrieval, and a private noise/style latent for things like texture or phrasing.
  • Deep projection heads. Residual MLPs, deeper on the semantic branch than on the noise branch, so the important information settles in the shared space.
  • A composite loss: cross-modal translation, a heavily weighted pure reconstruction term that rebuilds the input from the semantic latent alone, semantic alignment, and standard VAE regularisation. Weights and dropout were tuned with Optuna (30 trials, TPE sampler with median pruning).

Results

Text-to-image Recall@1, trained on one dataset and evaluated in-domain or zero-shot on the other:

Train → evaluate CLIP baseline VMSST
COCO → COCO 58.57 52.02
Flickr30k → Flickr30k 55.14 54.82
Flickr30k → COCO (zero-shot) 28.80 30.80

On the large COCO benchmark the generative constraint costs about 6.5 points. On the smaller Flickr30k the gap almost disappears, and when transferring from Flickr30k to COCO VMSST is ahead on every retrieval metric, which suggests the variational objective helps most when data is scarce.

Why the pure reconstruction term matters

Trained with the original VMSST settings, without that term, the model collapsed: positive and negative pairs ended up with cosine similarities of 0.99 and 0.89, everything squeezed into a narrow cone, and Recall@1 dropped to about 19%. Forcing the decoder to work from the semantic latent alone is what keeps the space spread out and usable.