Pith. sign in

Paper Citation Record · LEDGER

Improving Multimodal Reasoning via Worst Dimension Optimization

As of 19 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2606.07801.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.07801 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T21:55:57.691185Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact19
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e99716a8-8d0d-4069-95af-b38b06c19a9c · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Improving Multimodal Reasoning via Worst Dimension Optimization Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:37:14.804063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:d29b4e8a351444547ae8f096a1e9f1032a3c67be33b2f74e853d736b7b493cf0

Observation 42ee378d-0bf6-42a0-b4d0-4eaeb16dfbb3 · outbound

This paper cites Dense point clouds matter: Dust-gs for scene reconstruction from sparse viewpoints.

Improving Multimodal Reasoning via Worst Dimension Optimization Dense point clouds matter: Dust-gs for scene reconstruction from sparse viewpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:75e55934c70126245dd0fd9402e70fbb35ad69ba7657ad0e667b32677423d1fd

Observation 1e87c439-2e8c-479d-b3b7-fb1e0b63ebf1 · outbound

This paper cites Benchmarking multimodal cot reward model stepwise by visual program.

Improving Multimodal Reasoning via Worst Dimension Optimization Benchmarking multimodal cot reward model stepwise by visual program

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:c98eaacb846169f5e856835b30a0dae10bc7f57065914b85380194ba2aacb595

Observation 90d15e96-d409-4d4a-9b50-c4a9c5f5f5ca · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

Improving Multimodal Reasoning via Worst Dimension Optimization rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:37:14.796624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:58036efa53e0355a59c97dd4a1ffe23a206de0842b8ff562663c120df9220615

Observation 19bb96d0-9e8c-4f33-afbe-a965f5575564 · outbound

This paper cites Learning an efficient optimizer via hybrid-policy sub-trajectory balance.arXiv preprint arXiv:2511.00543, 2025a.

Improving Multimodal Reasoning via Worst Dimension Optimization Learning an efficient optimizer via hybrid-policy sub-trajectory balance.arXiv preprint arXiv:2511.00543, 2025a

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.799262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:d9db041f3849a29152cb0ef4af1d92aea20b48ebf9bde7ec7dfca623590e7227

Observation 24ae5da2-89f5-40b5-8ac2-44bdda8cd960 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Improving Multimodal Reasoning via Worst Dimension Optimization Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:37:14.801732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:4aa5420851847cbc46e8b2a7ac5d45f58359a09f76b4dab59354527fcf10d706

Observation c4b1e95f-ddcc-40fe-9b71-065d8c920170 · outbound

This paper cites RAM: Recover Any 3D Human Motion in-the-Wild.

Improving Multimodal Reasoning via Worst Dimension Optimization RAM: Recover Any 3D Human Motion in-the-Wild

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:37:14.791400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:e0a3d46908bbbdcd6e374524b97a32363dfdcb1b819f06b25cec667479bc32cf

Observation 25bd7bbc-51e2-4960-a9a7-2369903036db · outbound

This paper cites A diagram is worth a dozen images.

Improving Multimodal Reasoning via Worst Dimension Optimization A diagram is worth a dozen images

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:97388d2d4654c33216f9d7380550711fc44c22e15469158f8234c062e07a6298

Observation df5a02d2-53bd-4e78-9d48-61a9d1396adb · outbound

This paper cites Nv-embed: Improved techniques for train- ing llms as generalist embedding models.

Improving Multimodal Reasoning via Worst Dimension Optimization Nv-embed: Improved techniques for train- ing llms as generalist embedding models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:3daf1bd4b475fe6644e1c255076e8e06474125faa502d4446243423f0d925b3b

Observation b252fc26-6f34-4ea1-a859-b48504e39f71 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Improving Multimodal Reasoning via Worst Dimension Optimization LLaVA-OneVision: Easy Visual Task Transfer

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:37:14.793531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:6dbe4a3bec77565883134b11b62ffebd74a9e1150bf243f8741d3aa5a9603707

Observation d8c4b3ab-19e6-4b6c-b1d3-b51974ff04ed · outbound

This paper cites Human motion instruction tuning.

Improving Multimodal Reasoning via Worst Dimension Optimization Human motion instruction tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:6436fa778a396e8da83c50d84afdc6fe91d358e1defd3af2ee9e4e96b065eaab

Observation 2da19a14-8f2a-4170-a16d-5140d9c2529b · outbound

This paper cites Multiple human motion understanding.

Improving Multimodal Reasoning via Worst Dimension Optimization Multiple human motion understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:0f75ed95f39283085ac7f7c9787472b8a806a3a67fd252cb5f0eb0c843cee39e

Observation 5123f71f-8df4-4444-b6a4-f8e799718f51 · outbound

This paper cites Image semantic segmentation via chain- of-thought prompts.

Improving Multimodal Reasoning via Worst Dimension Optimization Image semantic segmentation via chain- of-thought prompts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:b9b3c826870caff7f7dd13f8dd81b43cdeab88ca81d3c53947fbdbc8f2996bc0

Observation dbdd4309-6d32-4d6a-a1b8-24a3a0c1e1f8 · outbound

This paper cites Graph Canvas for Controllable 3D Scene Generation.

Improving Multimodal Reasoning via Worst Dimension Optimization Graph Canvas for Controllable 3D Scene Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.788803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:57a2ae9ecf91dc3f25088bd8369fadc04e1a8992114cb5eef64c9c4f332f1839

Observation 3c249712-ad54-499f-be1f-103e69da9a96 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Improving Multimodal Reasoning via Worst Dimension Optimization Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:89566fb641aae39a2ec896d9f588d3491d7ccdfdab8d76d144b80dd1925d0069

Observation 758b02c2-9bbd-4c5b-96f8-58ff6ff3930e · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

Improving Multimodal Reasoning via Worst Dimension Optimization Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.785734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:6abb9703a3942629ba7282a7dc46c92f4df27ffe9d68c86f9df3d7e39583337a

Observation 2f0c2787-b8eb-4ac6-ac3b-af4567d8c2c1 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Improving Multimodal Reasoning via Worst Dimension Optimization Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:37:14.779823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:b48d8205ec97538633928c3e6d81f0c38c50b240a8fbdbae65f8609bde83f326

Observation d3ad855b-e1b8-4020-baf8-8e4dfab5d167 · outbound

This paper cites Ursa: Understanding and verifying chain-of- thought reasoning in multimodal mathematics.arXiv e- prints, pages arXiv–2501,.

Improving Multimodal Reasoning via Worst Dimension Optimization Ursa: Understanding and verifying chain-of- thought reasoning in multimodal mathematics.arXiv e- prints, pages arXiv–2501,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:b9805f6dcc5ba6058dcfef991c1249fe1847e2dba8314381fc9a4992de1959a5

Observation 42c34cba-879e-4744-bcf6-e8cacf0e4451 · outbound

This paper cites Chartqa: A bench- mark for question answering about charts with visual and logical reasoning.

Improving Multimodal Reasoning via Worst Dimension Optimization Chartqa: A bench- mark for question answering about charts with visual and logical reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:a1d313eb8370ee5aae628b86fa92070f26e524e78c3951c382785261c0fd9882

Observation 97037d53-e865-4a65-b893-5b4ac21661ea · outbound

This paper cites Training vision-language pro- cess reward models for test-time scaling in multimodal rea- soning: Key insights and lessons learned.

Improving Multimodal Reasoning via Worst Dimension Optimization Training vision-language pro- cess reward models for test-time scaling in multimodal rea- soning: Key insights and lessons learned

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.786548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:a148a8cdf1e42c3a062ddf141e51f557aba183d14b1aebe52010f346f422d702

Observation 33052ebb-9bef-4d81-9bd2-ead709adf165 · outbound

This paper cites Mutual reason- ing makes smaller llms stronger problem-solver.

Improving Multimodal Reasoning via Worst Dimension Optimization Mutual reason- ing makes smaller llms stronger problem-solver

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:5cd60a0693aaab033cd11d86a045dc439d62e304a0a1188497efd86937f6b49d

Observation 929facf5-a1b9-4ebb-8c0d-5f5ea592d403 · outbound

This paper cites an unresolved cited work.

Improving Multimodal Reasoning via Worst Dimension Optimization Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:26d8d6f8d84b588cda785b5ac88c282a6cd0510445aef2fbadb8bfc0e112eb5b

Observation ce4a326d-4a84-4e1d-98fb-fbab353ca368 · outbound

This paper cites Intrinsic entropy of context length scaling in llms.

Improving Multimodal Reasoning via Worst Dimension Optimization Intrinsic entropy of context length scaling in llms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:3016b46042771f6169fe0d13c920ab74646b435d54580b46a8d85cf3600bbdd2

Observation f87b0301-7f26-4395-b342-8f183e42b82f · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Improving Multimodal Reasoning via Worst Dimension Optimization Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:37:14.782921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:52982fd91e339512bc06e0b427cc5e8725bf453a810fe1da428d8e831b5fdf0c

Observation ae37ca32-ec1b-424b-84a7-296d7e773a93 · outbound

This paper cites Llamav-o1: Rethinking step-by-step visual reasoning in llms.

Improving Multimodal Reasoning via Worst Dimension Optimization Llamav-o1: Rethinking step-by-step visual reasoning in llms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:e47a20b70d64a517395d8b5333d42235a9d1bd8113cece77a9363e8e810b2607

Observation e942fe3f-4786-4ee7-9603-7cfadce43950 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Improving Multimodal Reasoning via Worst Dimension Optimization Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:37:14.791207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:e1ebf9dd76e3fc02d2aa6040644ef3623a8bbb974da8d934c6c26328314e5f53

Observation 2a6a6c27-42be-4e86-9480-70f58ec5bcb4 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

Improving Multimodal Reasoning via Worst Dimension Optimization Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:37:14.795919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:f32683b7df9571921949c063af4033420b0e70fb7ab5cda2ca4f38ebfdd79211

Observation b5cdf414-e7bc-41d9-8cf2-6a75999e1cc9 · outbound

This paper cites Multi-step problem solving through a verifier: An empir- ical analysis on model-induced process supervision.

Improving Multimodal Reasoning via Worst Dimension Optimization Multi-step problem solving through a verifier: An empir- ical analysis on model-induced process supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:d3529edac3334e43288c48bdd24d7610a272f7f4fac5831db7c3825d43898609

Observation ed2bd02a-a54d-46d2-bded-3debe6d62791 · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

Improving Multimodal Reasoning via Worst Dimension Optimization VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.806679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:d655cf2797433f33bbec48efd88216be7f8fd69883e3bb6f462520fc3e130926

Observation 8438769d-4f5a-4628-8cbe-782ad75e184d · outbound

This paper cites Beyond Examples: High-level Automated Reasoning Paradigm in In-Context Learning via MCTS.

Improving Multimodal Reasoning via Worst Dimension Optimization Beyond Examples: High-level Automated Reasoning Paradigm in In-Context Learning via MCTS

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:37:14.794057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:89f7e6b57b406f2ee7c5d37b8326aa1eae27025345968357207712beb11a582e

Observation 1fa0048c-73b5-486d-9728-3d95bf458f43 · outbound

This paper cites Llava-cot: Let vi- sion language models reason step-by-step.

Improving Multimodal Reasoning via Worst Dimension Optimization Llava-cot: Let vi- sion language models reason step-by-step

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:a149101a2e99c6ab958a059d50874435065084297e32bee6e5bab214843b41ae

Observation c7a61164-452e-48ba-9eb5-6f29d76accd6 · outbound

This paper cites 3dsceneeditor: Controllable 3d scene editing with gaussian splatting.

Improving Multimodal Reasoning via Worst Dimension Optimization 3dsceneeditor: Controllable 3d scene editing with gaussian splatting

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.777687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:166f332544e9dc9218f394367d67fdd60da2ed22ab5cfe64b539b5f21aefa933

Observation d65e27b7-5e6a-4f2d-a3cc-ef245e4b60f3 · outbound

This paper cites 3dsceneeditor: Controllable 3d scene editing with gaussian splatting.

Improving Multimodal Reasoning via Worst Dimension Optimization 3dsceneeditor: Controllable 3d scene editing with gaussian splatting

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:2b64fe7782ab2469541223ad47729c91136ace3389fc2fd0047d164b16d4b659

Observation ecabd59b-c43d-436f-9fff-09a342fea043 · outbound

This paper cites R1- onevision: Advancing generalized multimodal reasoning through cross-modal formalization.

Improving Multimodal Reasoning via Worst Dimension Optimization R1- onevision: Advancing generalized multimodal reasoning through cross-modal formalization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:ced210d765cc46d032946ea410c6b9d4c9791dbdb2188a6a8680e684272b5f60

Observation fa942ff7-cf88-43d2-829f-3c3093c309c9 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Improving Multimodal Reasoning via Worst Dimension Optimization MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:37:14.800932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:51ddab5966565694792df0b71c44315ef5d86adb50c8f9913de6e9d4a59ee703

Observation e6e906a0-c40e-4b76-b3fc-8c7fffef406a · outbound

This paper cites CountLLM: Towards Generalizable Repetitive Ac- tion Counting via Large Language Model.

Improving Multimodal Reasoning via Worst Dimension Optimization CountLLM: Towards Generalizable Repetitive Ac- tion Counting via Large Language Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:8978c4bfeff40edaf4925031ad61f2aacc9cb90c7b30c9dec4e923c43154f64b

Observation 6ac6dba7-02a7-4b3a-8b5f-4fd4f7bf9966 · outbound

This paper cites Mulberry: Empowering mllm with o1-like reasoning and reflection via collective monte carlo tree search.Advances in Neural Information Processing Systems, 38:29918–29952,.

Improving Multimodal Reasoning via Worst Dimension Optimization Mulberry: Empowering mllm with o1-like reasoning and reflection via collective monte carlo tree search.Advances in Neural Information Processing Systems, 38:29918–29952,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:6e6b943fa94aa17c54bca4edabf83105334905066095c80ae69c57fcba3fedcf

Observation 8834c047-b995-4af3-8186-267c35ef4b62 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Improving Multimodal Reasoning via Worst Dimension Optimization Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:dd1e44efc77c43e9fe421a1c4483165dbc31a847834c5627f63565e37ff9f73e

Observation 38a49562-928f-4857-97b0-29744c044a89 · outbound

This paper cites Birch: an efficient data clustering method for very large databases.ACM sigmod record, 25(2):103–114,.

Improving Multimodal Reasoning via Worst Dimension Optimization Birch: an efficient data clustering method for very large databases.ACM sigmod record, 25(2):103–114,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-27T21:55:57.691185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:3cf18a55d5037b06a3a08016b6ae949c38f8910eacb373f54ec55296f9de8471

Observation 771d9f0b-3317-4172-85eb-b6f208f24f83 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Improving Multimodal Reasoning via Worst Dimension Optimization InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:37:14.769228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:e9fa5e00aebd385592dcfa841633b7bff7e6e4a60cd635103f514afd796442af

Observation f16d2f8c-7666-4f98-9902-0b828bb51fff · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

Improving Multimodal Reasoning via Worst Dimension Optimization R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:37:14.803241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:03fe51e25193219721bf2377e792d9f3bc80effcc60fdfecfe80875b069890ad

Observation fc809c8a-8792-4516-8ed5-4ec02286e17d · outbound

This paper cites Psgs: Text-driven panorama sliding scene generation via gaussian splatting.arXiv preprint arXiv:2602.00463, 2026.

Improving Multimodal Reasoning via Worst Dimension Optimization Psgs: Text-driven panorama sliding scene generation via gaussian splatting.arXiv preprint arXiv:2602.00463, 2026

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.798473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:a4cb61b1aca24d0991e60651f98985da5e9289d667aad83482ea2379874b67b7

Pith citing papers

No inbound Pith citation observations are available.