Skip to content
FRFederico Raponi
All work

ProjectOct – Nov 2025

AML Challenge: Translating Text Embeddings into Image Embeddings

A course competition on model stitching: map embeddings from a text encoder into the space of a separate image encoder. Our VAE with a multi-positive contrastive loss reached an MRR of 0.868 on the public leaderboard, against a baseline of 0.462.

The task: given the embedding of a caption from one pretrained text model, predict where the matching image sits in the embedding space of a completely different image model, so that retrieval works across the two. Submissions were ranked by Mean Reciprocal Rank. We competed as the team Rofola Montain Club.

Final model

  • A variational bridge. A residual MLP encodes the 1024-dimensional text embedding into an 896-dimensional latent distribution; a decoder with residual blocks maps the latent to a normalised image embedding. At inference the latent mean is used directly.
  • Multi-positive contrastive loss. Symmetric, CLIP-style, with a learnable temperature, and aware that one image has several captions, so all of them count as positives.
  • KL regularisation with a linear warm-up of β (to 0.2 over five epochs) to avoid posterior collapse.
  • AdamW, and a learning rate cut by 0.3 when validation MRR stalled for five epochs.

Result: MRR 0.868 on the public leaderboard, baseline 0.462.

What we tried first

  • Relative representations (describing each embedding by its similarity to shared anchors): no better than the baseline.
  • Flow matching from noise, then from a projected text embedding: both failed; the loss optimises the velocity along the path, not where it ends.
  • A plain MLP with contrastive loss: a solid 0.858, but it overfitted after a few epochs. Adding the variational latent is what kept it in check.