Existing Composed Image Retrieval (CIR) datasets rely on short, single-sentence modification texts. Such texts are often ambiguous and provide insufficient descriptions of compound visual changes, resulting in limited supervision for learning fine-grained image–text composition. To address this limitation, we introduce L-CIRR, a new benchmark built on CIRR that provides long, detailed modification texts for each image pair. We further propose PACE (Progressive text Accumulation for Composed image rEtrieval), a retrieval framework designed to effectively exploit this richer supervision through two key components. First, progressive text accumulation, which incrementally integrates textual information sentence by sentence, constructs increasingly discriminative query representations. Second, multi-step hard negative mining exploits the intermediate embeddings from each accumulation step to mine hard negatives, facilitating more effective representation learning throughout training. Extensive experiments demonstrate that PACE, trained with the fine-grained supervision of L-CIRR, consistently outperforms existing methods on both fine-grained and general-purpose retrieval benchmarks while remaining robust to substantial variations in modification text length.
L-CIRR extends the existing CIRR dataset with long, detailed modification texts generated by a Vision-Language Model (Gemini-2.5-Flash). For each triplet, the reference image and the target image are jointly provided to the model along with a prompt that instructs it to produce a comprehensive, multi-sentence modification text of the visual differences between the two images. As a result, L-CIRR shares the same image pairs as CIRR while providing substantially more expressive modification texts.
Comparison of dataset characteristics and modification text statistics across different CIR benchmarks. L-CIRR texts average 113.7 tokens, 21.5 attributes, and 27.5 objects, with 76.7 unique tokens per text — a substantially larger scale across all measured dimensions than existing benchmarks.
Distribution of the number of sentences per modification text in the L-CIRR training set (left) and validation set (right). Modification texts range from 2 to 10 sentences, with the distribution peaking at 6 sentences in both splits.
PACE decomposes a long modification text into individual sentences and processes them progressively. The reference image and the full modification text are first fed into a multimodal encoder to produce a coarse query embedding, which is then progressively enriched with increasingly detailed textual information, sentence by sentence, through a stack of shared-weight text accumulators. In addition, Multi-Step Hard Negative Mining uses the intermediate query embedding of every accumulation step to score the gallery, collecting candidates that rank above the target as hard negatives. The union across all steps captures hard negatives at multiple levels of textual granularity.
Overview of the PACE framework. The reference image and full modification text are fed into the multimodal encoder to produce a coarse query embedding, which is progressively enriched with increasingly detailed textual information sentence by sentence through a stack of shared-weight text accumulators. The resulting final accumulated query embedding is optimized to distinguish the positive target from hard negative images.
PACE adapts the number of accumulation steps to the number of sentences in the input text, processing short queries directly while progressively refining its query representation for longer, more detailed ones. This length-aware processing allows PACE to maintain strong retrieval performance across both general-purpose and fine-grained retrieval scenarios, whereas models without such a mechanism remain more sensitive to variation in query length.
Qualitative comparison of top-5 retrieval results between PACE and SPRC, both trained on L-CIRR and evaluated under CIRR (top) and L-CIRR (bottom) on the same reference image. The blue-bordered image denotes the reference image, the green-bordered image denotes the ground-truth target, and the red-bordered images denote incorrectly retrieved results. PACE consistently ranks the correct target at the top-1 position under both query types.
@inproceedings{jang2026beyond,
title = {Beyond Single Sentences: Composed Image Retrieval with Long-Form Modification Texts},
author = {Jang, Junyeong and Lee, Seongwon},
booktitle = {Proceedings of the Asian Conference on Computer Vision (ACCV)},
year = {2026}
}