Pith. sign in

Paper Citation Record · LEDGER

Continual SFT Matches Multimodal RLHF with Negative Supervision

As of 17 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2411.14797.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14797 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:58:43.424797Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 216260a6-88c1-4992-b5d7-7eb67303aaa8 · outbound

This paper cites Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond.

Continual SFT Matches Multimodal RLHF with Negative Supervision Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.874415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:58:42.711922Z digest=sha256:e65d402eaa126f861905ff1447378c645c695dce8994b0e9e37171022737cfac

Observation dd7b815f-e7c0-4a0d-b68a-e89f0ae32aad · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Continual SFT Matches Multimodal RLHF with Negative Supervision ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.716829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.716829Z digest=sha256:a592bccb718c313748156a314320e5ba8c7ebf2dbcd098cf8765199d580fb5e7

Observation 3228bcd9-5962-426c-a75d-b0cfe8b77cfc · outbound

This paper cites Self-play fine-tuning converts weak language models to strong language models.

Continual SFT Matches Multimodal RLHF with Negative Supervision Self-play fine-tuning converts weak language models to strong language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.796661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:58:42.721830Z digest=sha256:fe68f435cb304545b72b354998cfd08b81db08ad6a27f713f8d587757dd9bd94

Observation b4e405fd-7802-465a-b135-fa84d92fe807 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Continual SFT Matches Multimodal RLHF with Negative Supervision How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.726082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.726082Z digest=sha256:eb4690008ee0c9ea33f32132b0b1fa9233ea147e6aea6cc9e21604f03da24ec8

Observation d7adcc2f-a2a2-4e3f-be14-bc214a52c496 · outbound

This paper cites Scaling instruction- finetuned language models.

Continual SFT Matches Multimodal RLHF with Negative Supervision Scaling instruction- finetuned language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.731110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.731110Z digest=sha256:d2a812615f6159298f9132ff199600a55593ffa51c1f6b6f96b91cd0e0348fa0

Observation d7bd5977-be98-4522-9546-98d7487d143d · outbound

This paper cites DreamLLM: Synergistic multimodal com- prehension and creation.

Continual SFT Matches Multimodal RLHF with Negative Supervision DreamLLM: Synergistic multimodal com- prehension and creation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.676044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:58:42.735944Z digest=sha256:546f920c61bdfb65decc00bf7ac037352a802a03b2f0b31eee2e6f3b22fa32d9

Observation 6fb41568-fb5b-43c1-93ad-7c7ab2d28a3e · outbound

This paper cites Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective.

Continual SFT Matches Multimodal RLHF with Negative Supervision Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.740552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.740552Z digest=sha256:7fc7f018b3104d982fe0b9141533e5eccf38eae4704d08e9970d4a3c97b8e2a0

Observation 60aff75b-b0e4-404a-9576-bd5797658670 · outbound

This paper cites GQA: A new dataset for real-world visual reasoning and compositional question answering.

Continual SFT Matches Multimodal RLHF with Negative Supervision GQA: A new dataset for real-world visual reasoning and compositional question answering

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.592976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:58:42.781559Z digest=sha256:f1ed7c1fe7fa703a58493038638b21a92aceb3526170fe509af6fb6ac31eb639

Observation c9148c5a-5ee2-4f7d-b2e3-1c9ce62084fb · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Continual SFT Matches Multimodal RLHF with Negative Supervision Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.863703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.863703Z digest=sha256:3b25bd93384814106927a0bd483d2d45ab6b81325013dc686246d712919cd150

Observation afb2b299-70bc-4c87-a29e-5d634b8226a5 · outbound

This paper cites Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models.

Continual SFT Matches Multimodal RLHF with Negative Supervision Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.961623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.961623Z digest=sha256:02fb1a066b62f4125b1fe3d72bf15976a16f32ed98f80144d707fd74f0b954f5

Observation 67336aec-9136-451a-a5fc-e9949679a539 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Continual SFT Matches Multimodal RLHF with Negative Supervision Evaluating object hallucination in large vision-language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.966247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.966247Z digest=sha256:1677726bdf79cf834561e736f8b3766ce93eb8d833e8ce09612a4da4744a567b

Observation ca51b88d-cb91-4471-a5e1-e991a12a02ff · outbound

This paper cites Microsoft coco: Common objects in context.

Continual SFT Matches Multimodal RLHF with Negative Supervision Microsoft coco: Common objects in context

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.544074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:58:42.970537Z digest=sha256:f872ea531db0d6e1685cec61d5c81179435b3e0a21cb8d65f5881e627d00d9a7

Observation af73828a-48da-4410-ab60-3ae07928687a · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Continual SFT Matches Multimodal RLHF with Negative Supervision Improved Baselines with Visual Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.975249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.975249Z digest=sha256:70fcc482f57b87dc9aac29a2b016c840f7029c1c5a73e00bd1e87fbf36c9dfb3

Observation 6780e922-21a2-4532-b766-7fd447b1d05d · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Continual SFT Matches Multimodal RLHF with Negative Supervision Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.980350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.980350Z digest=sha256:2b65562419aae858d09347e38bb3cfb6c280ef3bdbe92688dd52e6fa50bb82a3

Observation f2f0c405-731c-449b-9866-dd977f7744f6 · outbound

This paper cites Visual instruction tuning.

Continual SFT Matches Multimodal RLHF with Negative Supervision Visual instruction tuning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.518424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:58:42.984357Z digest=sha256:e8ea47742e36ea54edd59985e67930aa8e21baba601eda472e91308c82c4a289

Observation d472d0a2-332d-4488-a741-661c935fcb61 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Continual SFT Matches Multimodal RLHF with Negative Supervision MMBench: Is Your Multi-modal Model an All-around Player?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.989080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.989080Z digest=sha256:8ec8b61e8ff3772abc3c3ac10a58e24509c3eed1fc4312ca501d1f8033a5d447

Observation a30925e4-dab7-415a-a62f-7d9533e6f59d · outbound

This paper cites Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer.

Continual SFT Matches Multimodal RLHF with Negative Supervision Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.993918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.993918Z digest=sha256:7200e93a90c5f93e06ca42d835427d987b697c589d94e99bbe7eededa3fcf101

Observation 4a149c74-bc1f-48bd-bcf7-914f3eab9554 · outbound

This paper cites Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment.

Continual SFT Matches Multimodal RLHF with Negative Supervision Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.998808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.998808Z digest=sha256:1cda01a33bbceed68a1d9975989d42085f81fbf7e1596e1496ffe731b9a3cbaa

Observation 5c36a4d7-0de0-47d1-8af0-179aa8c5939c · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Continual SFT Matches Multimodal RLHF with Negative Supervision Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.005337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.005337Z digest=sha256:62ebca54fe9b6057e772aa6fe2dee7253c37e834526777ba663063d2e08f3fe9

Observation 54b27ab6-d0a3-4c2e-8615-17d4a70fdf94 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Continual SFT Matches Multimodal RLHF with Negative Supervision Ocr-vqa: Visual question answering by reading text in images

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.009885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.009885Z digest=sha256:cc4a8dcb30e9f331de97cfacdfde5195b1db3edb7334df0786676a01d5aaf885

Observation 0073ea85-db72-4aca-bd51-d7b67566ff5c · outbound

This paper cites Strengthening Multimodal Large Language Model with Bootstrapped Preference Optimization.

Continual SFT Matches Multimodal RLHF with Negative Supervision Strengthening Multimodal Large Language Model with Bootstrapped Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.014298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.014298Z digest=sha256:3bcf04d157364bdb623665e8eabea4d84dc939f31e8b360fd171baa658258560

Observation bc6b0de0-f74a-44f4-a4ec-85938236add7 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Continual SFT Matches Multimodal RLHF with Negative Supervision Direct preference optimization: Your language model is secretly a reward model

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.482432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:58:43.018997Z digest=sha256:192291e1db4a776342e4777535e3b15afec4fa8f9b4c6afab4405b6f7041b808

Observation b85c0898-0943-4897-8bac-4a5da5a97b5c · outbound

This paper cites Object hallucination in image cap- tioning.

Continual SFT Matches Multimodal RLHF with Negative Supervision Object hallucination in image cap- tioning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.023292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.023292Z digest=sha256:5fc22e0720f11801509b36c542328086eba61492b747b07a5220a9893cfbace3

Observation c428b956-4f77-4182-9193-6431105f8112 · outbound

This paper cites Multitask prompted training enables zero-shot task generalization.

Continual SFT Matches Multimodal RLHF with Negative Supervision Multitask prompted training enables zero-shot task generalization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.455332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:58:43.061051Z digest=sha256:1a3ecfbdb0a3c3bc67c1a550b09534ae58ff419f3f4a3a64e3c0746e6f7be149

Observation bacfa353-f4bd-4eb7-8b4c-ab861d80a813 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Continual SFT Matches Multimodal RLHF with Negative Supervision Proximal Policy Optimization Algorithms

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.084563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.084563Z digest=sha256:96becbd2356d4dc30f65e0a3aef49c58dfae09a321e8bca6b31b3952ead0e448

Observation e495b535-fcf1-4180-a3d4-a1768ff997bc · outbound

This paper cites Towards vqa models that can read.

Continual SFT Matches Multimodal RLHF with Negative Supervision Towards vqa models that can read

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.438925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:58:43.090981Z digest=sha256:cc94b83a3c89f46f613c1ac6f2060f5386abc33a6bf96ccf52e154490f53313c

Observation 9360bab2-3464-492d-95d0-f3c01d58014a · outbound

This paper cites Improving Multi-modal Large Language Model through Boosting Vision Capabilities.

Continual SFT Matches Multimodal RLHF with Negative Supervision Improving Multi-modal Large Language Model through Boosting Vision Capabilities

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:58:43.881132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:58:43.096094Z digest=sha256:3f324f945ca26949ae9e90d161729d8f41a1a1b461c3f37a80dc8d419ad5aa09

Observation 0628616f-3067-4613-8c82-ca2b18bf89e4 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Continual SFT Matches Multimodal RLHF with Negative Supervision Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.100804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.100804Z digest=sha256:31836971d3a0812569e2d2be9e4795a3d975822fcad71b953bfcf2570768d391

Observation a0a32663-3ec7-471b-90af-7515e6269bf9 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Continual SFT Matches Multimodal RLHF with Negative Supervision CogVLM: Visual Expert for Pretrained Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.106019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.106019Z digest=sha256:44615439d970c5140cc172830432fafabafd5a8338b7227a132276845b1cbf1f

Observation fda87008-9453-4592-8f8a-c60a3b90b76f · outbound

This paper cites Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement.

Continual SFT Matches Multimodal RLHF with Negative Supervision Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.111187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.111187Z digest=sha256:b4968193605acf429ad440a2aa94330d2989a038022be562aea6468f7f79db09

Observation 85960aff-4a04-42e4-9bca-69303c6d15c0 · outbound

This paper cites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models.

Continual SFT Matches Multimodal RLHF with Negative Supervision Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.116549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.116549Z digest=sha256:7330377bd07082e778a9cd61458149bb6423c6ad092702eaec017c9c80770d6f

Observation ba68516c-f71c-4423-bcd3-ffdb9bd47ce9 · outbound

This paper cites Finetuned language models are zero-shot learn- ers.

Continual SFT Matches Multimodal RLHF with Negative Supervision Finetuned language models are zero-shot learn- ers

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.360844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:58:43.121904Z digest=sha256:8d6ba31c671af7897409610de656f07b8b88d5f78158d0d9bd68771a8f1c1ec9

Observation e6e2386f-5a81-483e-8f08-7af1aca2430c · outbound

This paper cites an unresolved cited work.

Continual SFT Matches Multimodal RLHF with Negative Supervision Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.127468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.127468Z digest=sha256:bfe6cbbeb4bea41308c697593804b39d318263b8f17cdc00f5a5d5651de0ebd0

Observation 2d472179-5ac1-45cd-a058-aaaf1c45f832 · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

Continual SFT Matches Multimodal RLHF with Negative Supervision Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.132410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.132410Z digest=sha256:655d701e7998de5e125bb90d202724041ea0d2e6187a265a6b7c1378529c38b1

Observation 9c9d3f91-9754-4836-83aa-61b750e7b016 · outbound

This paper cites Is dpo superior to ppo for llm alignment? a comprehensive study.

Continual SFT Matches Multimodal RLHF with Negative Supervision Is dpo superior to ppo for llm alignment? a comprehensive study

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.258968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:58:43.137030Z digest=sha256:9bc3697fbed45b987774e62a7d5710b7aabd35466947f47ef410bc652b51e197

Observation f9cbea60-90fa-4966-b69b-2fd59088d728 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Continual SFT Matches Multimodal RLHF with Negative Supervision MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.142427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.142427Z digest=sha256:661bf11b869bce50acee69e98633552c14f872775602a1eb4f28576414589c84

Observation add161f7-45d0-4945-9a34-7804831dca39 · outbound

This paper cites Token-level direct prefer- ence optimization.

Continual SFT Matches Multimodal RLHF with Negative Supervision Token-level direct prefer- ence optimization

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.192399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:58:43.147270Z digest=sha256:896e440618df8811b8ab01ab1e594bb02a882819932bc23355e518f86d952220

Observation 18c45ca8-5129-42e9-9460-eeec3061e06f · outbound

This paper cites ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning.

Continual SFT Matches Multimodal RLHF with Negative Supervision ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.151627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.151627Z digest=sha256:27fcb37797f594cd59b7e9c35144e069185e8427a6ca679a494e0bb5b9d41af4

Observation ad2f2b85-ab4b-4fc2-8a3a-d7efb8b3cadf · outbound

This paper cites Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization.

Continual SFT Matches Multimodal RLHF with Negative Supervision Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.156556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.156556Z digest=sha256:0f9cf7921ab4b554aea7bab4964ee9f73f0544b737bf5813f1c19233b68e17b3

Observation 117b7452-cfa3-4742-9915-59877dfd51ff · outbound

This paper cites Lima: Less is more for alignment.

Continual SFT Matches Multimodal RLHF with Negative Supervision Lima: Less is more for alignment

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.177098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:58:43.204931Z digest=sha256:baf063bf9844e5f38b11d5ef20d2d7fb859eb2b793ea747b609732f2e3c82f01

Observation e6314292-d25f-44ef-ba1e-b1b1035789f1 · outbound

This paper cites Calibrated Self-Rewarding Vision Language Models.

Continual SFT Matches Multimodal RLHF with Negative Supervision Calibrated Self-Rewarding Vision Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.287932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.287932Z digest=sha256:9b41629c25bb893597cb3b93efaeb6ff883f2970ba6c057a66f09cb8eadbdaf9

Observation 32679ab4-f065-4208-9c03-054790116fd6 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Continual SFT Matches Multimodal RLHF with Negative Supervision MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.351600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.351600Z digest=sha256:a6c9d970b634077223ef718e42d7a8249a5cf57edf279c9f35f850ea15632fb3

Observation 36e81391-2f15-4795-b1f8-4a5ccb92b691 · outbound

This paper cites Multi-label self- supervised learning with scene images.

Continual SFT Matches Multimodal RLHF with Negative Supervision Multi-label self- supervised learning with scene images

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.159954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:58:43.404328Z digest=sha256:3705a2c3a5d41ab55f8087504162c7498ff360401a8caa8dd1e03472449d91e6

Observation 3e34811d-7d7c-42f2-9bb1-df6772bee8f3 · outbound

This paper cites Quantized feature distillation for network quantization.

Continual SFT Matches Multimodal RLHF with Negative Supervision Quantized feature distillation for network quantization

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.143565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:58:43.409800Z digest=sha256:3b916e14ab9486bc3d0f282383cba6e049a9d0ea469deae07ddab784366a7d98

Observation d8ef4bab-6bf2-4feb-ac9e-bff7cffd3016 · outbound

This paper cites Self-Supervised Visual Preference Alignment.

Continual SFT Matches Multimodal RLHF with Negative Supervision Self-Supervised Visual Preference Alignment

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.414905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.414905Z digest=sha256:15a992e7a4370b0f1a07be28476e825b486422e93f279758310b1d4a2ba122e9

Observation 44b228c8-8180-46e1-9277-8d27eca911cc · outbound

This paper cites Llava-phi: Efficient multi-modal assistant with small language model.

Continual SFT Matches Multimodal RLHF with Negative Supervision Llava-phi: Efficient multi-modal assistant with small language model

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.127762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:58:43.420092Z digest=sha256:ab4bd83114e006e667d3e9c5dc3bcff60339dc5937a1e70f60974d183bfbf601

Observation 52625be0-87db-4cb8-a6ca-fd175383b8a7 · outbound

This paper cites Multi: Multimodal understanding leaderboard with text and images.

Continual SFT Matches Multimodal RLHF with Negative Supervision Multi: Multimodal understanding leaderboard with text and images

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.424797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.424797Z digest=sha256:dea62bdfcce9de2e164663b8f921836fa4ecf776adb94eb96958c5d6d6e583b6

Pith citing papers

No inbound Pith citation observations are available.