Pith. sign in

Paper Citation Record · LEDGER

M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2306.04387.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.04387 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T00:46:08.118465Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T20:28:39.249603Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c15d0f86-cbc4-4ce8-95a1-a03c5f1703d5 · inbound

Otter: A Multi-Modal Model with In-Context Instruction Tuning cites this paper.

Otter: A Multi-Modal Model with In-Context Instruction Tuning M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:43:47.865942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:43:47.775691Z digest=sha256:a216fa1faffe78dff603bef80481b83022bb554a038ebc4cbcb6bf0e124e794d

Observation d1fd2764-d939-4d45-b502-d081c97b8057 · inbound

Large Language Models are not Fair Evaluators cites this paper.

Large Language Models are not Fair Evaluators M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T12:10:42.431540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T12:10:42.248005Z digest=sha256:b562ba91875b3d320d546182ad4075a1cecc837d3c842da95d32202c640e31fa

Observation d6c9abcb-4a2e-41bc-879a-535ab6a83d07 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.695744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:9b7e0633a92799df840771f291a82b2c6de957c92603247a88f9996ee670ca3d

Observation d7316a06-7a45-4489-8abe-44752f698298 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 282

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:28:39.251990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:8b8a01b349bdf1108f64e45df51a1a3b23fc3a334cb33db70b9b7257796e59ac

Observation 149dc577-de54-4a50-a79e-d9dd3eb604de · inbound

Vision-Language Foundation Models as Effective Robot Imitators cites this paper.

Vision-Language Foundation Models as Effective Robot Imitators M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.666644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:093566a994fd463fc2ef8e9eee2ce0ed5954b9b015d5d2812d953a2a77a3f809

Observation aacdd64c-0e17-4b4c-91b4-3f6d8bd4b716 · inbound

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI cites this paper.

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:37:41.612505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T05:37:41.401736Z digest=sha256:1217282f30b7730b847418e17c598d47d37573deb59259a15ca02ce8ba77be00

Observation ffad8371-cabe-4fd4-b398-09404bcf7c7a · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.115909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:c4ec418433a7db48c550bd6ad8f5eedc06a1a876d18134625e85d9fb24e16531

Observation 2e187d83-ad34-446e-8e3e-0793fa8f2fa6 · inbound

Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations cites this paper.

Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:34:15.763258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T22:34:15.638114Z digest=sha256:9bdc1ba3359afbcc3da97ffcb41192bf26447eeec308d346121ac1a10aedfcb0

Observation a9339ec0-f478-4a61-b0ed-a0e70d112f38 · inbound

Aligning Modalities in Vision Large Language Models via Preference Fine-tuning cites this paper.

Aligning Modalities in Vision Large Language Models via Preference Fine-tuning M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 162

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:58:53.407989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T10:58:53.215887Z digest=sha256:b84111586a3e80c253744b6d11855f7e2ed70dba89054469be8d1e5172e9622d

Observation 0d84ef2d-5fec-424d-922f-c386ee15a067 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.478013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:f970788b243d081075393cc5c3f1d77b2c01794c9f744ba99a6f9b06b0e243cc

Observation ac1e12be-fce0-4996-a55a-dcd92156cbe6 · inbound

LLMs can see and hear without any training cites this paper.

LLMs can see and hear without any training M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.118465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.118465Z digest=sha256:e597cd246780e81eff5dad03f6d6007f47e1442e355ccb4f913bc1e1d846193e

Observation 9ad6a6c3-1526-49b5-ad6a-a51aaf86e13f · inbound

REASSEMBLE: A Multimodal Dataset for Contact-rich Robotic Assembly and Disassembly cites this paper.

REASSEMBLE: A Multimodal Dataset for Contact-rich Robotic Assembly and Disassembly M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T20:22:55.949459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:22:55.949459Z digest=sha256:17fe6cc1d39c8bd612b184d0268cf8e04813ececc6a3c503f4525e7ee281349f

Observation 1c3a06e3-66d0-465f-a8e0-54b4e09ee35c · inbound

From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs cites this paper.

From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T22:43:59.423255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:43:59.423255Z digest=sha256:20392c39580ce48a4416d98262c2a2cb4b6e480d7fda0650e12cfb520f139780

Observation c8440316-70c4-4b2c-aeea-5864f1099fa8 · inbound

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types cites this paper.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.629857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.629857Z digest=sha256:4a2178b38da9de3363cf40dbef2160967f353726fb693d4ed965820db17cefe2

Observation 14856309-26db-41ac-9566-3a4537ffc053 · inbound

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion cites this paper.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.620424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.620424Z digest=sha256:b988b1c98215748aba56f00cd6081a9c8ea0f482f13145a6568e7772ecff6f0d

Observation 81be836d-7602-494e-8daf-47c9171aa506 · inbound

M$^3$FinMeeting: A Multilingual, Multi-Sector, and Multi-Task Financial Meeting Understanding Evaluation Dataset cites this paper.

M$^3$FinMeeting: A Multilingual, Multi-Sector, and Multi-Task Financial Meeting Understanding Evaluation Dataset M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T11:25:16.012128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:25:16.012128Z digest=sha256:a6f913de5907a5ecca9d8b725deb8186369e2da8d09ae1aca35464b764831c74

Observation c26c49ee-07bd-446b-9729-c1ceebb6150b · inbound

VIP: Visual Information Protection through Adversarial Attacks on Vision-Language Models cites this paper.

VIP: Visual Information Protection through Adversarial Attacks on Vision-Language Models M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:29.523026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:16:29.523026Z digest=sha256:125d4ab3e28047f0fedeb4cc878d5364c62a4af7e2b46a754dea0c366c9a14d1

Observation 6b470d80-6eba-4783-a5b4-51126b5557e2 · inbound

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security cites this paper.

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:39.711544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:39.711544Z digest=sha256:5fda7be46bda9cb9bd6dbce37409f3846537131ad98a67fc320f049e6fb35c0e

Observation 4d0bd7e3-10c6-4259-acb1-704fa9e30622 · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.831964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.831964Z digest=sha256:090d4c071ac10b52cd069e4a605d7e58fcca4bcd18ccd4ec4cc8858da5b3efd7

Observation 07239a44-8713-4a3b-bcb2-c2b3efcf6a0a · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T20:28:42.331519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:28:42.331519Z digest=sha256:47044dd795fba2b24e7fdb09161773cdbfc5ef05b7fd1f3256450de27842e6e9

Observation fe1cceee-ea0c-4c22-8a3e-da732ff23488 · inbound

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model cites this paper.

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T17:56:53.246219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:56:53.246219Z digest=sha256:563117c8e3ba946ffa20c2b5c0a0be49b2791b4bf2a03db975aca6fbf9962350

Observation 821da7d8-724c-49c5-a1a6-da96264ca380 · inbound

Language-Specific Layer Matters: Efficient Multilingual Enhancement for Large Vision-Language Models cites this paper.

Language-Specific Layer Matters: Efficient Multilingual Enhancement for Large Vision-Language Models M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T16:30:12.596189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:30:12.596189Z digest=sha256:09cedf441716acb2f9574482183abea50a2d129f00b17a0703052f991e02159e

Observation b2e03eb9-624f-47a9-9882-fe769dadff86 · inbound

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning cites this paper.

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-14T22:18:56.115559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:18:56.115559Z digest=sha256:533354fed48eea4dee7a3f9f7ce6feef157762bde7d7d478c5da31a7cd6c8d91

Observation 196cf01f-10d5-4599-a8c6-801cdd4c3419 · inbound

Vision-Language Foundation Models for Comprehensive Automated Pavement Condition Assessment cites this paper.

Vision-Language Foundation Models for Comprehensive Automated Pavement Condition Assessment M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:21.206498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:30:39.410040Z digest=sha256:f4ff9bca9278eeabc178dac8c273f7489fc54bf18c35c23f73451494348d6c8a