REVIEW 3 cited by
LM4LV: A Frozen Large Language Model for Low-level Vision Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
The success of large language models (LLMs) has fostered a new research trend of multi-modality large language models (MLLMs), which changes the paradigm of various fields in computer vision. Though MLLMs have shown promising results in numerous high-level vision and vision-language tasks such as VQA and text-to-image, no works have demonstrated how low-level vision tasks can benefit from MLLMs. We find that most current MLLMs are blind to low-level features due to their design of vision modules, thus are inherently incapable for solving low-level vision tasks. In this work, we purpose $\textbf{LM4LV}$, a framework that enables a FROZEN LLM to solve a range of low-level vision tasks without any multi-modal data or prior. This showcases the LLM's strong potential in low-level vision and bridges the gap between MLLMs and low-level vision tasks. We hope this work can inspire new perspectives on LLMs and deeper understanding of their mechanisms. Code is available at https://github.com/bytetriper/LM4LV.
Forward citations
Cited by 3 Pith papers
-
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Combining multimodal understanding with future image prediction in one autoregressive model improves vision-language-action policy success rates in simulation and real-world manipulation.
-
Improving Generalization of Language-Conditioned Robot Manipulation
A two-stage fine-tuning framework with instance-level semantic fusion lets language-conditioned robots learn object-arrangement tasks from a few demonstrations and generalize to unseen environments.
-
LossAgent: Towards Any Optimization Objectives for Image Processing with LLM Agents
An LLM agent reads past loss weights and quality scores, then writes new loss weights, letting image processing models be trained toward non-differentiable objectives like IQA scores and text feedback.
Discussion (0). Continue with ORCID to comment.