ProjectOct – Nov 2025
AML Challenge: Translating Text Embeddings into Image Embeddings
A course competition on model stitching: map embeddings from a text encoder into the space of a separate image encoder. Our VAE with a multi-positive contrastive loss reached an MRR of 0.868 on the public leaderboard, against a baseline of 0.462.
The task: given the embedding of a caption from one pretrained text model, predict where the matching image sits in the embedding space of a completely different image model, so that retrieval works across the two. Submissions were ranked by Mean Reciprocal Rank. We competed as the team Rofola Montain Club.
Final model
- A variational bridge. A residual MLP encodes the 1024-dimensional text embedding into an 896-dimensional latent distribution; a decoder with residual blocks maps the latent to a normalised image embedding. At inference the latent mean is used directly.
- Multi-positive contrastive loss. Symmetric, CLIP-style, with a learnable temperature, and aware that one image has several captions, so all of them count as positives.
- KL regularisation with a linear warm-up of β (to 0.2 over five epochs) to avoid posterior collapse.
- AdamW, and a learning rate cut by 0.3 when validation MRR stalled for five epochs.
Result: MRR 0.868 on the public leaderboard, baseline 0.462.
What we tried first
- Relative representations (describing each embedding by its similarity to shared anchors): no better than the baseline.
- Flow matching from noise, then from a projected text embedding: both failed; the loss optimises the velocity along the path, not where it ends.
- A plain MLP with contrastive loss: a solid 0.858, but it overfitted after a few epochs. Adding the variational latent is what kept it in check.