Pith. sign in

Paper Citation Record · LEDGER

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning

As of 10 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2606.29915.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.29915 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T06:16:51.452335Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact12
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02b8d00c-4f93-4420-83ac-e36dd654eddb · outbound

This paper cites Don’t just assume; look and answer: Overcoming priors for visual question answering.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Don’t just assume; look and answer: Overcoming priors for visual question answering

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:524acdf3944b40629b7ba747f1b11cc68db7ae21e48ff88bc99bd0ba3ce8f273

Observation f2005225-0a92-4941-b183-d69a499f2df9 · outbound

This paper cites Neural module networks.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Neural module networks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:245e0c6def7ff54d1effb45dcf08501655a4f3759c8eb419d3c10766022dd0fa

Observation ee40dcdc-5370-4c95-a3d8-967bd50e467a · outbound

This paper cites Vqa: Visual question answering.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Vqa: Visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:c7368524b5cd6c47aa945dc49f91f974aed62f008892dfdfce41b5ae04c9e04a

Observation 304897c1-bffe-4823-93be-ce55427eff1c · outbound

This paper cites Qwen2.5-VL Technical Report.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Qwen2.5-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:14:22.073999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:47c93838d16cdb293f20980e3828b70c5e7f5d05b78635a9f0858d14606bdb13

Observation d7f7db7e-3c3f-49a4-9b6e-ff5569249d74 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning SAM 3: Segment Anything with Concepts

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:24:19.583033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:b3f6d615586241db90955c221264ffcac9b00b44e21f63828fa95feb11ea73a1

Observation 27fd4ac7-7e01-4009-a5d2-eb51a41c2ce3 · outbound

This paper cites Are we on the right way for evaluating large vision-language models?Advances in Neural Information Processing Systems, 37:27056–27087, 2024.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Are we on the right way for evaluating large vision-language models?Advances in Neural Information Processing Systems, 37:27056–27087, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:d3a6ce3ddf8eeaa46f3925e62754e9581baa7e09a3488bed60278b517aa46c53

Observation 381137bd-9993-4a30-837d-a17fd62a0a69 · outbound

This paper cites Beyond question-based biases: Assessing multimodal shortcut learning in visual question answering.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Beyond question-based biases: Assessing multimodal shortcut learning in visual question answering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:7e5627e66ea17cb241e7ee47ec6ebc62df600371ce631aff28f997d5034c68ab

Observation 567c85f0-f740-458e-9cdb-ab0859677745 · outbound

This paper cites Gemini 3 flash: High-efficiency agentic multimodal understanding.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Gemini 3 flash: High-efficiency agentic multimodal understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:5485ffe3f3fbad0894aef2e1ffc676823dc250e5b4756252096eeeff9da03abf

Observation 681706ef-c37d-430e-a621-b8633a67aebd · outbound

This paper cites Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081):633–638, 2025.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081):633–638, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:651bd9d77d644066b05578fe705de4297da2286a2eb2c9c1798addc46e8674f2

Observation a41934e7-2ede-4f66-a00b-f76a93870f80 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:14:22.071060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:9982b0f8a760d62aec134ba4ab79eaabbb2794e232544a954152e88d6825f762

Observation 9cb8e320-72d3-4e2f-9a86-0ee4dd2e9462 · outbound

This paper cites Hudson and Christopher D.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Hudson and Christopher D

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:991d451c4b46448dff81d78f65c19d04e247f2965d77e355246302628190d8f3

Observation a8a852af-ba1f-4046-b480-da8f5a67508c · outbound

This paper cites GPT-4o System Card.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning GPT-4o System Card

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:24:19.584388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:5f251520dc27ac0e9f11d7ea16b92f67492edd1faa79bd85b6095024e243d3fa

Observation cee35312-3f35-48a6-b8d7-35a4c954a5f3 · outbound

This paper cites Raven progressive matrices.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Raven progressive matrices

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:3e71be8e1fca57dd34d1610bdf1f98c8fdafada950005c1afe7683c82cb0b687

Observation 77debea3-3348-4396-bc2c-d0d78d85c398 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:41bf1619b96a8c54ab024a7c30c774608d0d00537c717a436c594b6ea16be218

Observation bc695b8e-3edf-47f5-b018-e49f0a68ec03 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123(1):32–73, 2017.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123(1):32–73, 2017

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:46116d50b8c43980be7e46591faf14d0f75a38d09054641d2d5d0e8e13b817e1

Observation c59b158b-a7c3-4a51-b698-9e6950b63d02 · outbound

This paper cites Imore: Implicit program-guided reasoning for human motion q&a.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Imore: Implicit program-guided reasoning for human motion q&a

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:d5fa61bcd6ddc5dfd5fdefe32e380a61aa56c003617f76be667c865cc91280ad

Observation 68e336e7-f3a2-4420-ac31-b599f80d5a74 · outbound

This paper cites Vision-sr1: Self-rewarding vision-language model via reasoning decomposition and multi-reward policy optimization.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Vision-sr1: Self-rewarding vision-language model via reasoning decomposition and multi-reward policy optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:439aac6709876abeda9dcec795583c7a928b319e7b6e077e59665e073f8ee988

Observation fccf2cf1-8ace-4b76-a479-ee35882b6513 · outbound

This paper cites Visual-rft: Visual reinforcement fine-tuning.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Visual-rft: Visual reinforcement fine-tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:652693ca07129ecdf006bcb27bd43fdd4ecc40b9be50e0b9e92e7f82c7a317cc

Observation ca8cded6-2d71-4a6c-8461-950e1b81a47f · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in neural information processing systems, 35:2507–2521, 2022.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in neural information processing systems, 35:2507–2521, 2022

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:8352c66c27a69cf5471e48d6112934071eb97047401be5124260a97e375838dd

Observation 9bfd8663-6812-4d8c-a8e8-ff796bca7233 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:191458a40db58f3f5fb9eb24aa97f8a8919d5175a4a96984c30ba66c93d3b1dc

Observation 08a5f508-d2ce-4c02-bce3-ef8ce740177d · outbound

This paper cites A computational investigation into the human representation and processing of visual information.WH San Francisco: Freeman and Company, San Francisco, 1(1):4, 1982.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning A computational investigation into the human representation and processing of visual information.WH San Francisco: Freeman and Company, San Francisco, 1(1):4, 1982

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:25244fdd6c3e908190bb818a961ff38435312f3dc947d18eb81bd0df5e1c7c0f

Observation 33942dd2-30ed-426a-acdc-5777d7bab885 · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning SmolVLM: Redefining small and efficient multimodal models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:24:19.588487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:efa89a14764b89dde4a8c0a47e868cccb69f97c6b56fae5f49a145b160b14c78

Observation 1f35ee56-26f9-4737-8384-360dc2ed56b0 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:6937750b80ff9d515968cc35457dc6cbe896793ac61c4f0cea6d056b9ee92c3f

Observation 59f4b493-5f3e-45a1-9c08-676339e7dbae · outbound

This paper cites Pkr-qa: A benchmark for procedural knowledge reasoning with knowledge module learning.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Pkr-qa: A benchmark for procedural knowledge reasoning with knowledge module learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:f9ccc579b46ba1627f41aff8d1a01ac93442fc220de5b62326e64900cc456d11

Observation bf751ef3-1174-4d11-b2c2-3d524abe4764 · outbound

This paper cites an unresolved cited work.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:c168fe171d2fe2b789719451520053a36694114ef845e5cbdad028da732e762d

Observation ccbc4a4e-de5e-42a3-a945-99785e6dda8b · outbound

This paper cites Grounding multimodal large language models to the world.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Grounding multimodal large language models to the world

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:f0559fa5c4ef233fdbe410afd150b4980248c6f6c1eb506bb8b1148a6c0559c7

Observation 5e974d39-0d41-4161-a197-8def928049a4 · outbound

This paper cites Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:24:19.585935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:a3c6f146910e5d1473a88bad0350a55cfec7a1c1629c902eb318111c0ceb1368

Observation 96d199c0-0ddf-41ce-be50-f39ca01f802b · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert-networks.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Sentence-bert: Sentence embeddings using siamese bert-networks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:0fd55058bcced02e7ff96ef0a90a51c0f56bf12265072c0dda9ac4606cd8d323

Observation b6c63f4b-5c70-4a23-962e-6d5d7467de5c · outbound

This paper cites an unresolved cited work.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:1aefc05529e939a6041865941d2cdf14f77db6b624818eb5c93296cebd3d5953

Observation c8ad17e8-4c85-44a1-bf44-cc49afe30d08 · outbound

This paper cites Grounded reinforcement learning for visual reasoning.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Grounded reinforcement learning for visual reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:17d2accd02e6ae15badcd8ff1838b1d33df41ab329f26baf2ef09d0b23da217b

Observation 324d5f17-1452-403d-810d-b9ed2a1b16ee · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning A-okvqa: A benchmark for visual question answering using world knowledge

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:472153680c36d6ea7167f80ab2f78d8e1e8aa8f29c8e5f132dc55966c1636272

Observation 1a146134-7fb7-4d46-a9fd-869d1b678983 · outbound

This paper cites an unresolved cited work.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:bc959030206063b68868a7d4566eee017a0495cf6367907379011e0c289ca6b2

Observation ee428d28-21d9-4781-98d6-00f5aed48bb6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:24:19.577448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:1305c4d7983e3ededf9b914e81e6ef14d131873742994930a38172f0a89f3391

Observation beb3ca3b-32e5-4dfb-9342-9ab25c3dc33e · outbound

This paper cites Vlm-r1: A stable and generalizable r1-style large vision-language model.arXiv e-prints, pages arXiv–2504, 2025.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Vlm-r1: A stable and generalizable r1-style large vision-language model.arXiv e-prints, pages arXiv–2504, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:127da6523ddfcf9381c16a8aa4f3d8531085beaa8f4928bf72faafbcf47b3e50

Observation 3817647c-3a84-414f-ad02-75b509cb8a2b · outbound

This paper cites Language prior is not the only shortcut: A benchmark for shortcut learning in VQA.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Language prior is not the only shortcut: A benchmark for shortcut learning in VQA

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:1eb0e787f0bfe28752caaf9e51fa7e493cb07d11cf927d97fec4b1f546cf2bbc

Observation f8492fe8-ccf3-4c5d-babe-8f2ea424a8b6 · outbound

This paper cites Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:c0dbe90bc42a87030c6912556fd39213f9704653b10f81a47bd5e26c6ad8d162

Observation 8edf08a1-565c-4f43-8536-fbd4300c7b4f · outbound

This paper cites Aligning large multimodal models with factually augmented rlhf.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Aligning large multimodal models with factually augmented rlhf

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:0a27b774232e6f5e5711df90186d5da4c0675b5fccfc0933534d7545887030e6

Observation c6c2f491-059a-47b0-bc86-75c48693c019 · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning of vision language models.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Reason-rft: Reinforcement fine-tuning for visual reasoning of vision language models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:84f64954220825ecb2363ebfe8dba1748987f2bada64b5aa6d681cb4906650b3

Observation df0e8031-8d0b-4862-becf-f94c559ca7b0 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Gemini: A Family of Highly Capable Multimodal Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:24:19.592194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:eabcda64ea6542bafbbf8ab0336abbf674748d8e8e193854ae96d0b7d9785538

Observation d76f3881-784c-40ee-96e9-31edee27a189 · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Gemini Robotics: Bringing AI into the Physical World

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:24:19.581646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:36e1ab4fc057f836356c8ad7cd73d29ecca7305980fb0990b9a34efc78d1ce52

Observation 4ef73406-d8cf-452a-bf48-4a59a9975570 · outbound

This paper cites Qwen3.5-Omni Technical Report.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Qwen3.5-Omni Technical Report

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:24:19.590819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:bda04861163f5b8a80ac5db4985556ff38016354462c61cc8b8286135e1b72e0

Observation 9f84fce8-0f5a-44ee-a14e-82980fd649b9 · outbound

This paper cites Vl-rethinker: Incentivizing self-reflection of vision-language models with reinforcement learning.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Vl-rethinker: Incentivizing self-reflection of vision-language models with reinforcement learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:36090e81c82a6d2e72c6819ae2237aa1dbbbba1c4e8f9383fa2a0e9b4560b9e9

Observation d1bc2cc3-8b94-4cda-897f-1267bc9e62b6 · outbound

This paper cites Procedures as a representation for data in a computer program for understanding natural language.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Procedures as a representation for data in a computer program for understanding natural language

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:7724875618eeb98999bbf6a3a45b3990f98de1f9518c13402976138de91764b7

Observation a06c8446-a93e-497e-abd8-c75bdbe0a6fa · outbound

This paper cites Learning structural descriptions from examples.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Learning structural descriptions from examples

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:26f38e6985e42fd25665823076436cf4a80677ec2b22e5d7e2c2578950e58371

Observation a4642844-0455-4f73-8e94-04cae7b4956a · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:24:19.569669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:b7d365571f58efe0c339ecc675a19e5958562500bceef1333fe0112f5e90173a

Observation 5ea1c24a-e558-4ac0-877b-e6af5a798835 · outbound

This paper cites Realworldqa: A benchmark for real-world spatial understanding.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Realworldqa: A benchmark for real-world spatial understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:f53c4874a088cc261c0efef9678c482fb87216c7ac721f7fda5174429f4954d6

Observation 40a64413-d951-4af7-bb72-51832351807d · outbound

This paper cites Next-qa: Next phase of question- answering to explaining temporal actions.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Next-qa: Next phase of question- answering to explaining temporal actions

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:f7d952060d44e3b1c8a9a21865c2c1c023ad9de87f8ac557fd0a5116ee3b9ae8

Observation ad88ec18-9174-4c7f-abe2-bb20bd58f2d1 · outbound

This paper cites Neural- symbolic vqa: Disentangling reasoning from vision and language understanding.Advances in neural information processing systems, 31, 2018.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Neural- symbolic vqa: Disentangling reasoning from vision and language understanding.Advances in neural information processing systems, 31, 2018

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:7e8bc2bbaf6ef64bb54605441cfbe4d5036a571efdac23a89c58b2e600a906d3

Observation deaf51d4-4060-4d61-b0d3-0f4199430f68 · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:1ed7400871ff90c8ce62ae471f164e06be06b258b5153d2fcdfac82ed5493ae5

Observation 2fa4283e-d302-4393-83fb-ed25a5e22a12 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:d5889c425a316d130b8581a42511a95c77be4b25b402d6c682ad07cbf46ad231

Observation 25ac8f50-9b0e-4b04-aeef-587f5588266e · outbound

This paper cites Raven: A dataset for relational and analogical visual reasoning.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Raven: A dataset for relational and analogical visual reasoning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:be8c02239908110756ede8c831c18b50d53066594774d2f2e19d51d3aa45b5fb

Observation 5d838940-73f2-49c8-87ed-0113bd0f5b32 · outbound

This paper cites Mitigating Easy Option Bias in Multiple-Choice Question Answering.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Mitigating Easy Option Bias in Multiple-Choice Question Answering

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:24:19.587769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:64658426f80e3465088a6505631040c75707c636c25a66a2482fc9a73cb224de

Observation 00e4f341-87a6-4d92-97a7-6bf27303ca4a · outbound

This paper cites R1-vl: Learning to reason with multimodal large language models via step-wise group relative policy optimization.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning R1-vl: Learning to reason with multimodal large language models via step-wise group relative policy optimization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:69b0ea6fd8f92e270342aa566d0aca34668dc923429f24a78d4cb4e1a1ff1250

Observation 1df64fe7-114f-4e9d-97c2-1dd1c9a1d1e2 · outbound

This paper cites Physreason: A comprehensive benchmark towards physics-based reasoning.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Physreason: A comprehensive benchmark towards physics-based reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:50a7d5e2a7835b2c9d7b73c6913b71e0696072716e72bea411f28aad77adae5b

Observation 87f9e699-cdf2-4b17-80b9-ca1a51bfb1f1 · outbound

This paper cites Multimodal chain-of-thought reasoning in language models.Transactions on Machine Learning Research, 2024, 2024.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Multimodal chain-of-thought reasoning in language models.Transactions on Machine Learning Research, 2024, 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:a9bf9f1ae9894c277402438a436cf6ee1ad4baa7008d2928a523f78dfee33ac4

Observation cb99c87c-ba04-4766-b386-b18c03a4e5dc · outbound

This paper cites Visual7w: Grounded question answering in images.

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning Visual7w: Grounded question answering in images

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-30T06:16:51.452335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T06:16:51.452335Z digest=sha256:82539a9a0efdbc3b24ff538f022c0df8421e91a9ff210e5adc97460c4b9190c9

Pith citing papers

No inbound Pith citation observations are available.