{"id":"62c5233b-33ef-457f-817e-1f57965fc807","arxiv_id":"2504.18158","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Adding a learnable pixel-level perturbation to the in-context pair improves MAE-VQGAN visual in-context learning on segmentation and object detection benchmarks.","lead":"E-InMeMo adds a small trainable border pattern to the example images used as prompts in visual in-context learning, and reports higher segmentation and detection accuracy than prior retrieval-based methods. The method needs only about 27,000 extra parameters while the underlying vision model stays frozen.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper never shows that the trained border perturbation acts through the in-context pair; without a blank/random-pair control, the 7.99/17.04 gains could be a task prior, not prompt enhancement.","rationale":"I read the paper in good faith: the method is simple, code is released, and the improvements are consistent across folds and datasets. The Pith Reader's conditionality is reasonable. I considered error bars, test-set hyperparameter selection, and the domain-shift inconsistency; these are real but secondary. The most load-bearing gap is causal: the paper claims the perturbation enhances in-context pairs, but never tests whether the in-context pair is actually doing the work. A task-specific border perturbation trained on the target labels could plausibly raise mIoU by itself, since the MAE-VQGAN blank-cell prediction is highly underdetermined and any consistent prior helps. The inter-class experiment (Figure 6) shows the prompt is class/task specific, which actually supports the task-prior explanation. The proposed control is cheap and decisive. If it passes, the paper's claim is substantially strengthened; if it fails, the headline should be reframed from 'enhanced prompting' to 'task-prior perturbation for visual ICL.'","tokens_in":17919,"tokens_out":15840,"duration_ms":160657,"concrete_test":"Run the published E-InMeMo pipeline but, at evaluation only, replace the retrieved in-context pair (x,y) with (a) a constant gray pair and (b) an unrelated pair from a different class, keeping the trained t_phi and the query fixed. If either control retains a large fraction of the reported gains (43.91 vs 35.92 for segmentation; 44.22 vs 27.18 for detection), the improvement is not attributable to enhancing the in-context pair. As a second arm, freeze t_phi at a random border pattern and train nothing; if a random border alone produces comparable gains, the 'learnable' component is not the active ingredient.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that the learnable perturbation t_phi in Eq. 4 improves visual ICL by enhancing the retrieved in-context pair. The training objective (Eq. 8) uses the query's ground-truth label tokens, and t_phi is shared across all pairs for a task, so the most natural alternative explanation is that t_phi simply encodes a dataset-level prior (e.g., 'produce the foreground mask distribution of this task') that would help even if the in-context pair contained no task-relevant content. The ablations in Section 4.5 (Table 4) vary where the learned prompt is placed, but every condition keeps the real in-context images in the canvas; no condition removes the pair, replaces it with a random or blank pair, or freezes t_phi at a random initialization. Without those controls, the observed gains are compatible with the perturbation acting as an unconditional task bias rather than as an enhancement of the in-context example. This matters because the paper's framing and title attribute the improvement to better prompting of the visual ICL pipeline; if the pair content is not necessary, the method is still a legitimate PEFT strategy, but the central 'enhanced in-context pair' claim is not established. Other issues (missing error bars, test-set padding selection, the domain-shift text in Section 4.4, and the parameter-count inconsistency) affect reliability, but this missing control targets the mechanism itself.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes E-InMeMo, a parameter-efficient visual prompting method for MAE-VQGAN-based in-context learning. A shared learnable pixel perturbation t_phi is added, via Eq. (4), to the retrieved in-context pair (x,y) before the pair is assembled with the query into a four-cell canvas; only t_phi is trained, using a cross-entropy loss (Eq. 8) between the MAE's predicted VQGAN tokens and the ground-truth VQGAN tokens. Retrieval is performed with a feature-map-level retriever (FMLR) using CLIP or DINOv2 features. Experiments report mean mIoU of 43.91 on Pascal-5i foreground segmentation (vs 35.92 for FMLR-DINOv2) and 44.22 on PASCAL VOC single-object detection (vs 27.18), together with results on two medical datasets, ablations of prompt placement and padding size, retrieval-set-size and dataset-size sensitivity, class-generalization matrices, and a COCO-to-Pascal domain-shift experiment. The paper frames the contribution as enhancing the in-context pair rather than simply adding a task prior.","tokens_in":18268,"tokens_out":10487,"duration_ms":94052,"significance":"The empirical core is clean: the FMLR-DINOv2 baseline isolates retrieval, and the only difference in E-InMeMo is the added learnable prompt, so the reported gains are attributable to the prompt under the stated protocol. The ablations in Table 4 show that prompt placement matters, and the method is lightweight (27,540 trainable parameters) with publicly available code. If the mechanism claim is established, this is a useful PEFT strategy for visual ICL. I do not see a circularity problem: Eq. (8) is a standard supervised cross-entropy on VQGAN tokens, and the comparison baselines are independent. The main gap is interpretational: because t_phi is shared across all pairs and trained on the task, the experiments do not yet rule out that it acts as a dataset-level prior rather than as an enhancement of the specific in-context pair; the domain-shift and near-tie conclusions also need support.","major_comments":[{"comment":"No control condition removes or replaces the in-context pair, or freezes t_phi at random initialization. Since t_phi is trained on the task dataset and shared across all pairs (Section 3.4 calls it a task-specific identifier and Section 3.7 says it captures the distribution of yq), the 7.99 and 17.04 gains are compatible with an unconditional task prior rather than with enhancing the retrieved pair. Please add a blank/random/irrelevant in-context pair condition and an untrained-random-prompt condition, and report whether the gains persist; if they do, the contribution should be described as visual PEFT rather than in-context-pair enhancement.","section":"Section 3.4, Eq. (4), Table 4"},{"comment":"The robustness conclusion is not supported by the reported numbers. The absolute drops are 2.32 (FMLR) versus 2.49 (E-InMeMo), a 0.17-point difference, and the relative drops are 6.46% versus 5.67%; these differences are negligible and no variance is reported. The sentence stating that the performance gap between the two methods under domain shifts is minimal (0.17 points) is also misleading, since the actual cross-domain gap in mIoU is about 7.8 points, essentially unchanged from the in-domain gap. The claim that E-InMeMo enhances cross-domain generalization should be removed or supported by proper statistical evidence.","section":"Section 4.4, Table 3"},{"comment":"All experiments are single runs, but several conclusions depend on near-tie or small differences. In Table 5, padding size 10 (mean 43.89) and padding size 15 (mean 43.91) are statistically indistinguishable, with fold-level reversals; in Table 3, the claimed robustness rests on a 0.17-point difference. I request mean and standard deviation over at least three seeds, or a significance test, for these comparisons, and correspondingly softer claims of optimality and robustness.","section":"Section 4.5, Table 5; Section 4.4"}],"minor_comments":[{"comment":"The reported parameter count of 27,540 is not derivable from the stated 15-pixel border on 111 by 111 images combined with Eq. (4); please specify whether t_phi is applied per image or to the combined pair, including channel and gap handling.","section":"Section 4.1"},{"comment":"The text says The full results are shown in Table 1 when describing the intra- and inter-class generalization matrix; this should refer to Figure 6.","section":"Section 4.5, Figure 6"},{"comment":"The statement that E-InMeMo requires a minimum of 32 images per class is not consistent with Figure 8, where Fold-0 reaches 37.24 with 16 images per class and the mean first exceeds the 35.92 baseline at 64 images; please reconcile the statement with the figure.","section":"Section 5, Limitations; Figure 8"},{"comment":"The notation g_l is used both for the assignment-score vector in Eq. (1) and for probability-like quantities in the cross-entropy loss in Eq. (8); please state explicitly how softmax normalization is applied.","section":"Section 3.1 and Eq. (8)"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable extension of the authors' WACV 2024 work, but the incremental novelty relative to [28] should be checked by the editor. The missing blank-pair control is the key experiment: if gains persist without real in-context pairs, the title and framing must change. I would not reject on current evidence, because the empirical comparison is clean and the ablations are useful."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper adds a learnable border perturbation to the retrieved in-context pair in the MAE-VQGAN visual ICL pipeline, and pairs this with a feature-map-level retriever (FMLR) using DINOv2. The headline numbers are real: E-InMeMo gets 43.91 mIoU on Pascal-5i segmentation and 44.22 on single-object detection, which beats the FMLR-DINOv2 baseline by 7.99 and 17.04 points respectively. The ablations in Table 4 support the central claim that adding the prompt to the in-context pair helps, and the medical dataset results are a nice extra. Code is released.\n\nWhat is genuinely new is the specific combination: applying a shared pixel-level perturbation to the in-context pair (not the query) within this particular ICL framework, plus the FMLR retrieval strategy. That is a reasonable incremental step, but note the gains over the authors' own WACV 2024 InMeMo are only 0.77 mIoU on segmentation and 1.01 on detection. The novelty is modest, not absent.\n\nThe soft spots are real but not load-bearing. No error bars or significance tests anywhere, which matters when the key claims are single-digit percentage point differences. The padding size hyperparameter (15 pixels) is chosen by looking at the evaluation folds in Table 5, so the reported number is optimistically selected. The domain-shift section (4.4) is the weakest text: the '0.17 point gap' is actually the difference in absolute drops, not a performance gap, and the robustness conclusion rests on a relative-decrease advantage that is within noise. The parameter count is also inconsistent: for 111x111 images with a 15-pixel border, the trainable pixels should be 17,280 for one image or 34,560 for two, not 27,540; either the implementation or the text is wrong.\n\nThe stress-test concern about a missing control is legitimate. Every ablation keeps the real in-context images in the canvas; there is no blank-pair or random-pair condition, and no frozen/random t_phi condition. So the gains could in principle come from t_phi acting as an unconditional task prior rather than from enhancing the in-context pair. The paper's own Section 3.7 essentially admits this ('phi captures the distribution of yq'), which is honest, but it means the title's 'enhanced prompting' claim is not directly established. This is a limitation, not a fatal flaw: the method is still a legitimate PEFT strategy, just with an oversold mechanism.\n\nOverall this is a solid, incremental paper for the visual ICL subfield. It deserves a serious referee, and I would want the authors to add the missing control, fix the parameter count, temper the domain-shift language, and report variance. My own verdict would be a weak accept or borderline reject depending on venue standards, but it should not be desk-rejected.","headline":"A useful but incremental extension of visual prompting to visual ICL: clean ablations, small gains over the authors' own InMeMo, and a missing control that leaves the 'enhanced in-context pair' mechanism underdetermined.","tokens_in":882,"tokens_out":1386,"would_cite":false,"duration_ms":42048,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adding a shared learnable pixel perturbation to the retrieved in-context pair raises mean mIoU from 35.92 to 43.91 on foreground segmentation and from 27.18 to 44.22 on single-object detection.","keywords":["visual in-context learning","learnable visual prompting","prompt enhancement","image segmentation","single object detection","parameter-efficient fine-tuning","medical image segmentation","MAE-VQGAN"],"falsifier":"A decisive experiment is to replace the learned $t_\\phi$ with a fixed random border pattern of the same shape and magnitude after training. If a random pattern gives most of the same mIoU gains, then the effect is not learned task alignment but a generic input perturbation; if it gives no gain, the learned optimization is doing the work.","tokens_in":17734,"feed_emoji":"🖼️","tokens_out":6750,"duration_ms":65458,"temperature":0.7,"pith_summary":"Visual in-context learning lets a frozen large vision model solve a new task by showing it an input-output image pair alongside the query, but the method only works when that pair is a good example. This paper tries to remove that dependency by learning a tiny shared pixel-level perturbation and adding it to the in-context pair before the model sees the canvas. The authors show that this one learned border pattern, trained with a cross-entropy loss over the model's visual tokens, improves foreground segmentation by 7.99 mIoU points and single-object detection by 17.04 points over the strongest no-prompt retrieval baseline. Because only 27,540 parameters are trained and the large model stays frozen, the approach is a cheap adapter for visual ICL. A sympathetic reader would take the paper's proof to be empirical: the perturbation helps across natural and medical images, under domain shift, and with small retrieval sets.","feed_headline":"A learnable border prompt lifts visual in-context learning by 17 points","feed_subtitle":"Adding 27,540 trainable pixels to the retrieved example pair beats prior methods on segmentation and detection.","key_machinery":"The load-bearing object is the prompt enhancer $t_\\phi$: a 15-pixel-wide border of learnable pixels (27,540 parameters) that is added, with unit strength, to both the example input and the example label in the retrieved pair. The same $t_\\phi$ is shared across all queries for a task, making it a task-specific identifier rather than a per-example edit. It is optimized by minimizing the cross-entropy between the frozen MAE's predicted VQGAN tokens in the empty output cell and the VQGAN-encoded tokens of the ground-truth canvas (Eq. 8). The paper's interpretation is that this aligns the latent distribution of the incomplete canvas with the distribution of a canvas that already contains the correct label, so the frozen decoder sees more plausible tokens. The second supporting mechanism is the feature-map-level retriever using DINOv2 features, which selects the in-context pair that the enhancer then refines.","core_discovery":"E-InMeMo's central claim is that the retrieved in-context pair can be improved, not just selected: a learnable, input-agnostic perturbation $t_\\phi$ added to both the example input and its label (Eq. 4) shifts the visual prompt so that a frozen MAE-VQGAN produces better output tokens. The perturbation is trained only on the task dataset with a cross-entropy loss between the predicted tokens in the empty canvas cell and the ground-truth tokens from a VQGAN encoder (Eq. 8). In the paper's experiments this raises mean mIoU from 35.92 to 43.91 on Pascal-5i foreground segmentation and from 27.18 to 44.22 on PASCAL VOC single object detection, with the same trend on the Kvasir and ISIC medical datasets. The authors interpret this as the prompt encoding task-specific distributional information that compensates for imperfect retrieval.","pith_inferences":["The paper leaves untested whether prompts trained on class groupings (e.g., all transportation classes together) could replace per-fold prompts; the inter-class generalization matrix suggests such grouped prompts might retain most of the gain with fewer trained parameters.","An untested corollary of the mechanism is that the same border-perturbation recipe should improve other MAE-VQGAN inpainting-style tasks such as style transfer or image inpainting, since the loss is defined purely over visual tokens.","The ablation showing that perturbing both query and example hurts implies that a single shared additive pattern is a constrained design; pair-specific or query-conditioned prompts are a natural next step and could close more of the gap to oracle in-context pairs."],"forward_implications":["Reported gains of 7.99 mIoU on Pascal-5i and 17.04 on PASCAL VOC indicate that prompt quality, not just retrieval quality, is a controllable lever in visual in-context learning.","At only 27,540 trainable parameters and a single inference per query, the method offers a parameter-efficient alternative to prompt-ensemble approaches like prompt-SelF.","The data-efficiency experiments show that the method with 10% of the retrieval set beats the no-prompt DINOv2 retriever using 100%, so the approach is useful when labeled data is scarce.","Domain-shift results (COCO to Pascal) show a smaller relative drop than the baseline, so the learned prompt partly absorbs distribution shift.","Ablations show that prompt placement matters: perturbing both the example and the query hurts, so future designs should keep the perturbation on the in-context pair."],"supporting_citations":[{"why":"Supplies the frozen MAE-VQGAN model and the four-cell canvas inpainting formulation that E-InMeMo builds on.","marker":"[15]"},{"why":"Introduces learnable input-agnostic pixel-level visual prompts, which E-InMeMo adapts to refine in-context pairs.","marker":"[19]"},{"why":"Establishes that in-context pair quality matters and provides retrieval baselines and the COCO-to-Pascal domain shift protocol.","marker":"[16]"},{"why":"Provides the DINOv2 feature extractor used by the feature-map-level retriever.","marker":"[29]"},{"why":"The authors' prior conference version that E-InMeMo extends with direct perturbation of the in-context pair and a reduced parameter count.","marker":"[28]"},{"why":"The strongest visual ICL baseline using an ensemble of eight prompt arrangements, which E-InMeMo beats with a single inference.","marker":"[18]"},{"why":"Supplies the Pascal-5i dataset used for the foreground segmentation evaluation.","marker":"[59]"}],"fun_headline_variants":["Learnable prompt border boosts visual ICL by up to 17 points","Trainable perturbation on in-context pairs lifts mIoU by 8–17","E-InMeMo: learnable prompt tweaks beat prior visual ICL","27k trainable pixels on example pairs improve vision ICL"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central assumption is that one fixed border pattern, trained once per task, can push the frozen model toward correct outputs for every query in that task, and that making VQGAN token predictions more accurate is a faithful proxy for making pixel-level segmentations more accurate.","fun_headline_variants_meta":{"raw":{"variants":["Learnable prompt border boosts visual ICL by up to 17 points","Trainable perturbation on in-context pairs lifts mIoU by 8–17","E-InMeMo: learnable prompt tweaks beat prior visual ICL","27k trainable pixels on example pairs improve vision ICL"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000772,"raw_usage":{"total_tokens":3418,"prompt_tokens":942,"completion_tokens":2476,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":558,"completion_tokens_details":{"reasoning_tokens":2395}},"tokens_in":558,"tokens_out":2476,"duration_ms":20316,"temperature":1.0,"reasoning_tokens":2395,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:22:57.157523+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive experiment is to replace the learned $t_\\phi$ with a fixed random border pattern of the same shape and magnitude after training. If a random pattern gives most of the same mIoU gains, then the effect is not learned task alignment but a generic input perturbation; if it gives no gain, the learned optimization is doing the work.","supporting_citations":[{"cited_title":"Zhang, B","cited_arxiv_id":null,"evidence_quote":"The authors' prior conference version that E-InMeMo extends with direct perturbation of the in-context pair and a reduced parameter count."}],"review_version":1}