Pith. sign in

Paper Citation Record · LEDGER

MM-IFEngine: Towards Multimodal Instruction Following

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2504.07957.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.07957 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:34:53.141416Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T06:07:41.277547Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ce052be3-c027-4d9c-8e25-20d10f7eb5df · inbound

Generative RLHF-V: Learning Principles from Multi-modal Human Preference cites this paper.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference MM-IFEngine: Towards Multimodal Instruction Following

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:53.141416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:53.141416Z digest=sha256:f1bd85952d2e2f5d06fb8d4ba527a46fe0a01b13d8b041ffe0f82c9c209379dd

Observation e589d879-6482-434a-914b-44d8abc6423e · inbound

Enhancing LLMs' Reasoning-Intensive Multimedia Search Capabilities through Fine-Tuning and Reinforcement Learning cites this paper.

Enhancing LLMs' Reasoning-Intensive Multimedia Search Capabilities through Fine-Tuning and Reinforcement Learning MM-IFEngine: Towards Multimodal Instruction Following

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.807752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.807752Z digest=sha256:3042de48c32c4ce0c93b3b4a63d8ab64ad57caabac2e90997ba73be89f9cc8ae

Observation 7df927b7-b6e8-40ea-b6b2-e4cbc3ca62fd · inbound

SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law cites this paper.

SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law MM-IFEngine: Towards Multimodal Instruction Following

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:37:19.038128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:37:19.038128Z digest=sha256:4ede66ac88a9927303aa3be93a86bb5e564f21938cbb0eb763e427d0a75b9835

Observation 06c76bfa-a6df-405d-8cf1-291695515668 · inbound

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience cites this paper.

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience MM-IFEngine: Towards Multimodal Instruction Following

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T23:55:47.953630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:55:47.953630Z digest=sha256:b3c3df44427e7265a3bcdf0bdc0a03cc33bae343e300141bd1535112d7f8a712

Observation e5a6931a-e40b-49fe-afd6-f6fa08abcc63 · inbound

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning cites this paper.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning MM-IFEngine: Towards Multimodal Instruction Following

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.573543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.573543Z digest=sha256:4e68cb731efddfa992b9460523db6a26a4cef2d833c77c5d7bd0714fe9c715c2

Observation 14cc3989-b400-434c-afe8-a3dd3bedf4aa · inbound

MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe cites this paper.

MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe MM-IFEngine: Towards Multimodal Instruction Following

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:07:27.369232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T17:07:27.277040Z digest=sha256:e660e1f7056c462bc92080a5f68e27852bf43396b999d62a1919c52aed912e7d

Observation 57620993-b560-4e83-9480-7ca9529f824a · inbound

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction cites this paper.

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction MM-IFEngine: Towards Multimodal Instruction Following

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:46:26.483873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T09:26:00.413651Z digest=sha256:5ee1c4bfc3c833efd88e4d75a46535baeea9025ca04ec54a7b4ff9096f568710

Observation 55178eb6-937d-4822-ace7-c49303eddfbd · inbound

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs cites this paper.

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs MM-IFEngine: Towards Multimodal Instruction Following

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:23:12.304574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T10:19:11.536555Z digest=sha256:4200cc8d82063f948f6d851a9054894de4f71a2c2653534d05d0b1a0e773aec9

Observation c6c9fd09-df02-4f16-9ac9-f7f4a82ab8b6 · inbound

Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR cites this paper.

Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR MM-IFEngine: Towards Multimodal Instruction Following

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:03:03.537441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T05:02:35.271960Z digest=sha256:4ce3649c889aa5f2c3c42b31b53ce34c544f32ef9a03bda1641cf477f09d4340

Observation 88d2dbec-3d06-4db2-b481-ec5201e88200 · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models MM-IFEngine: Towards Multimodal Instruction Following

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:41.279129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T12:55:47.754632Z digest=sha256:bf767f341e0b2065d56db68af5de4576c64db10505010ec494c48e3867ee731b

Observation a6a9b865-629e-45ab-a6dd-03cf1c25bb2a · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models MM-IFEngine: Towards Multimodal Instruction Following

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T18:07:09.018997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T18:07:09.018997Z digest=sha256:566ecf10e072a1c300679acd3f699654cc5b2b99f673d1cf9cd5d2f7a8163997

Observation 910787fc-bc00-429a-aaf7-8146d37dff2f · inbound

When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation cites this paper.

When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation MM-IFEngine: Towards Multimodal Instruction Following

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:27.304749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:27.304749Z digest=sha256:5da08a29cfa510cfb1a4d5f50f68e47f78468e22fb568085a92ba823b0c3b00a