Pith. sign in

Paper Citation Record · LEDGER

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search

As of 18 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2506.11155.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11155 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:45:51.654456Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1c60edfc-0cbb-4f35-89df-1696c8185719 · outbound

This paper cites The Verifier we used is Qwen2-VL-72B.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search The Verifier we used is Qwen2-VL-72B

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.549246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:49.732140Z digest=sha256:013df37fa33514e45404789170fc68c87599fb24fadab1bbaab0abcf2dd65f00

Observation d0cb6278-ba63-4b33-bad5-f1403c087f6a · outbound

This paper cites video caption.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search video caption

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.536479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:49.769808Z digest=sha256:7a258e1e8dd14514f016a96bbe13044c4dba3e14baad71db83bf6aff52e585c7

Observation 08bb520d-f4a2-4e37-8030-ace8a2b4984e · outbound

This paper cites The defination and examples of each category is described in Figure 10.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search The defination and examples of each category is described in Figure 10

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.494812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:49.897453Z digest=sha256:f557e0f31991347a469ee880e4210ef9231b9c17e0d6310963ac73f6172e7077

Observation 43c230d3-60b0-4dda-8539-13f547e19a14 · outbound

This paper cites As shown in Figure 6, the Nature and Wildlife category has the highest number of videos, reaching 382, while the Arts and Creativity category has the fewest, with 57 videos.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search As shown in Figure 6, the Nature and Wildlife category has the highest number of videos, reaching 382, while the Arts and Creativity category has the fewest, with 57 videos

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.523671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:49.807483Z digest=sha256:94ef98f3aa560f0cbfc264245bd18e61dce67c55197e7600fae8c182d75e10f4

Observation 0edb9973-083d-4661-9ce8-bb956e41b22e · outbound

This paper cites For clarity, consider these examples: ## Example 1 ### Key Points {1.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search For clarity, consider these examples: ## Example 1 ### Key Points {1

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.128104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:51.250419Z digest=sha256:ba84a475929d01daa3a426c2162de947452c335f4bea853845c1255a753916be

Observation aca6c052-7370-41b5-9be1-b61b192335c7 · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:49.655855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:49.655855Z digest=sha256:1ba18334a7c1053e318eba9af76b717ab2e891999c3216a577de9b24989fb1ed

Observation e43c75bc-cef8-466e-b1e4-d6c27ddc085b · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.510059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:49.877197Z digest=sha256:60a2bbb01d8463f18f0b729c2b75802dbd70eb409b40b59f41523c4700f27db8

Observation 9128212f-1432-4a17-a791-81d3712fea5c · outbound

This paper cites Elaborate on the visual and narrative elements of the video in detail.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Elaborate on the visual and narrative elements of the video in detail

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.482137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:49.953882Z digest=sha256:2e2a20361fc54470247a0498f6436ee6e971d1076f759e733909939213b4d4a5

Observation 09dc07ac-d317-4bee-b07f-d469fb158832 · outbound

This paper cites Reply to me with a precise yet detailed re- sponse.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Reply to me with a precise yet detailed re- sponse

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.468746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:49.986135Z digest=sha256:41406e209dcc584a451ea3f97e337bc4123488bb41a4cb376efb1e4cf3028cc6

Observation aa5f7445-cd02-4a89-ade0-b30c3b04a451 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.455747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.040106Z digest=sha256:8667581e57edea18560d74f8821167cbc8a755ce33256a6e630efee305112584

Observation 12594294-b8ea-45a0-8d77-a4ea5371fbf2 · outbound

This paper cites If you are not sure about something, do not include it in you response.\n # Task\n Describe the background, characters and the actions in the provided video.\n”.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search If you are not sure about something, do not include it in you response.\n # Task\n Describe the background, characters and the actions in the provided video.\n”

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.442836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.063737Z digest=sha256:2f7b72da972d0af1fdfeb90a9e80a9b089a25a74aed90bfaeb39af4f8340ed7e

Observation 09a987e6-3a93-4b05-93d3-bd4e4708cb89 · outbound

This paper cites Please describe the video in detail.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Please describe the video in detail

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.429147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.091040Z digest=sha256:1f1b2116d3cbe9217895671b4fe468822a3bfad6c0eb3ce64ac0539a54c41cc6

Observation 612c6a12-7820-4ec1-8ee6-5fb423a35600 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.416554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.113506Z digest=sha256:48e79eb59d6a8bec5c2fd585fcfd05ce66d77e194c64a9ab2a5688e8aaf70b75

Observation 8ff35304-bb11-4a01-a8a6-9a32d9b4da57 · outbound

This paper cites Action Description Action description focuses on the specific behavior or activity that takes place in the video.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Action Description Action description focuses on the specific behavior or activity that takes place in the video

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.404105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.157246Z digest=sha256:74c6ae43130d66b5330cf466951922c57caa4583725e0321f432c292fe4d77d1

Observation 5e35b046-e14c-4907-a93c-8cf020039730 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.391772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.227646Z digest=sha256:51aaa2f784a9cb4951ed99055587e610c63be87eb8c210691112973694fec4c0

Observation f6407169-11de-4a84-b3c8-456171823de0 · outbound

This paper cites Environment DescriptionEnvironment description covers the background and environmental features of the scene in the video.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Environment DescriptionEnvironment description covers the background and environmental features of the scene in the video

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.380039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.272752Z digest=sha256:cf243af658b188d417fb99bfb174863ebe29f776486b31c15c7d52a85518ba8c

Observation 0a7688bf-5e33-47dc-8ebd-8049fa3c40e8 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.367175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.321552Z digest=sha256:227a178e4a94bf699e2df7720ceb08d12f56e24dbc3780cdedbe8cfd86601da8

Observation 1c4f6e0d-c3a7-4497-a5f7-082058f8b5ac · outbound

This paper cites Object Description Describes the features, characteristics, or details of inanimate objects or items present in the scene.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Object Description Describes the features, characteristics, or details of inanimate objects or items present in the scene

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.353915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.384862Z digest=sha256:47324081c9ebf57bf4860fe178d306e2b901a7971186cfeb0505240b2dd741f1

Observation 36aad59d-79c6-4bef-9b3c-afc83c6f062f · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.340986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.433451Z digest=sha256:2875ceb2289d87bb145bb7c36e5acb2e62f02aef6e53d9c68c1856ad02da28e4

Observation e9a41f06-b236-4622-b32f-2204c25f501c · outbound

This paper cites Camera Movement Describes the camera angles, movements, framing, or other cinematographic techniques used in the scene.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Camera Movement Describes the camera angles, movements, framing, or other cinematographic techniques used in the scene

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.327006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.493367Z digest=sha256:0f21617313be6165c9237266e7f7931b026b7b52da2ad96cd9af1258be5e52d4

Observation 6d106cb3-23e4-4936-81ab-e38ff7579a41 · outbound

This paper cites Overall Description.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Overall Description

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.312951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.556622Z digest=sha256:2a64a8628b022502de14f9cb04c48961b0026371ec6bcb66ebe4263c7093f193

Observation e2c286b5-8623-4035-8ed6-1564be95773f · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.299776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.602919Z digest=sha256:37dd78ea494f554a7f10f730816f657034b14358aca8a04b62ce0593619d2e43

Observation 44daad0e-240f-45f7-a52a-fefdb635de9b · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.287434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.648467Z digest=sha256:2ab078605936c590a042f9372a9982f6ead14c1f20f072d475956bbfb7b61b27

Observation 207efaf8-4853-479b-b148-e0c90cf7d968 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.273193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.685906Z digest=sha256:6f044eb779c8e832c79252f85c78de0a6381e0b31081c76152b7f034260de689

Observation 5a01bb6c-ea31-42a4-b312-5f45cfd4fdf3 · outbound

This paper cites Judgment: [yes/no] Reason: [Brief explanation].

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Judgment: [yes/no] Reason: [Brief explanation]

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.259294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.709361Z digest=sha256:d878e3dc959d407fadeb261062612afb5efeeeef2d37c882278f9a8985ece841

Observation e7a52ff5-1705-41f0-bb08-4f714e14fd67 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.246129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.751336Z digest=sha256:909c1659f5c10008794e1658819756ee3ac6a10db03f27c55d5452685dc685c7

Observation 73134215-1d39-4897-8815-04b2c4f43b6b · outbound

This paper cites For clarity, consider these examples: ## Example 1 ### Video Description: The video showcases.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search For clarity, consider these examples: ## Example 1 ### Video Description: The video showcases

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.233155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.820854Z digest=sha256:0c06c6fd9097f318abf2f50c50614e6d72f63284be42a0a40128893da4ee37c6

Observation 44eec224-6230-442b-a74a-18647b9d08b2 · outbound

This paper cites entailment.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search entailment

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.219323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.866840Z digest=sha256:a43365edee51af9d79df024b85cb010d9062516e9b474900aa4782e9ed8bfa7c

Observation e3a218d6-bbb0-476d-936c-c61e9096d38d · outbound

This paper cites contradiction.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search contradiction

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.204952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.920062Z digest=sha256:1abb78c2305bb8af66f5731a5cf8ae03947885bf2bad22c19bcd0d2285f68ea7

Observation deaf138a-b65f-4a5b-bfd3-379f268f0095 · outbound

This paper cites neutral.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search neutral

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.191475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:50.988196Z digest=sha256:e288db44bd564b685a7ec61c02bcc98455a3cb5d828c57375942dcc44455ac7a

Observation f04dc7e4-f631-4893-888d-c690782b4841 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.179059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:51.039572Z digest=sha256:e9e47f00680c40190f41dbaae5c813b22a454c52d4de9120e9ae560e5286c0fd

Observation 2dead321-8060-4d48-b8fe-099d903d3c0e · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.166708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:51.135838Z digest=sha256:1ba113b9a8c3516973b2f1938421b3910177bb409e969b55e8f56834b342796a

Observation bc4c0a8c-5d0a-4cda-b690-413ff0c6f458 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.153967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:51.205648Z digest=sha256:99929bcf6e183cc2a13b6b9ac62220aa9a9c0c00aca08e210dcd8479430db41c

Observation 6f322406-4ade-4809-85c7-01a3c1c5245d · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.141384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:51.226275Z digest=sha256:799ce7c480f000513146366da0fa1f4ca0f20a5debcb62316d56b633f24b456c

Observation 209f6f20-0529-436b-bd45-60bbf283b7a3 · outbound

This paper cites [No] Speculative.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search [No] Speculative

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.114620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:51.286328Z digest=sha256:b6b82651ed437e80d42a830743ebe4dc3ad5df5fa9c2b97ba1d6ecd2e09730b3

Observation 4e58d046-f8ca-40fd-94a3-9631b7ad3b35 · outbound

This paper cites </thought> tags.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search </thought> tags

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.101011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:51.329282Z digest=sha256:76059aa9f4682196e09591d1c59406a1572934cd3430c579295c03930e1686a1

Observation 440f3780-deb9-417e-923a-65bd5d676425 · outbound

This paper cites Overall Description.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Overall Description

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.086183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:51.449655Z digest=sha256:a75c643205a56635ba5b98bfd0e0d4bd3c9d4b59cd47974a9cfc2d2458d5cfa5

Observation ae45ca9a-7395-467c-9248-59acb4d784ae · outbound

This paper cites ## Example Input and Output: ### Input: Overall Description: {overall_description} Observation: Vehicles Key Point: There is no existence of any vehicles in the video.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search ## Example Input and Output: ### Input: Overall Description: {overall_description} Observation: Vehicles Key Point: There is no existence of any vehicles in the video

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.072730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:51.496957Z digest=sha256:3abfe85893a75976f23648a861e943a19eb44747cad52d12eddfefe388408606

Observation ba2e7ca2-4f59-4f63-a9b4-b4249dfbf04d · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.059285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:51.561541Z digest=sha256:dbb43b22347445db4bfd0d51eb2a197287b05c810e0f341d43b3b8cc67e72ce3

Observation 943bbb7a-b0a6-4ed1-97dc-7efff2e687fe · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.045858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:51.625504Z digest=sha256:70779a00459c19bb304bf6262e8a25231e7bffc87e15dfce784cec4ecd36b9f5

Observation 5cd90b8e-b737-4cb4-9a01-69db871d2438 · outbound

This paper cites Precision / Recall / F1 Score.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Precision / Recall / F1 Score

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:51.968280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:51.654456Z digest=sha256:8a82c608e62bd1f6cd4dbfaa61959efe14b0b6c48a8470e711aa47479b941b70

Observation 91c77bd8-51be-499e-a1b7-18721181226a · outbound

This paper cites CRiskEval: A Chinese Multi-Level Risk Evaluation Benchmark Dataset for Large Language Models.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search CRiskEval: A Chinese Multi-Level Risk Evaluation Benchmark Dataset for Large Language Models

Reference 2013

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T04:45:51.840389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:45:49.548866Z digest=sha256:c37ca9534faa3d119ceb5bf284a07c22e5ff58fe05c395efb70772a80d19289d

Observation 8cc72c16-a1b7-44d9-ba64-53bcd81e4d9c · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Reasoning with Language Model is Planning with World Model

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:49.431685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:49.431685Z digest=sha256:ba78e473894ce9baf2a71c31238ee456f8675e34b28b800ebe29b35721c1f00d

Observation 8b0c2260-b588-4876-b0f3-809ab39e684f · outbound

This paper cites A Survey on Data Augmentation in Large Model Era.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search A Survey on Data Augmentation in Large Model Era

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:49.695413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:49.695413Z digest=sha256:d647685855e475f68dbf62ba110cfc82de7669aeb514a21469603ee568289781

Observation d31a7ef3-db04-4ac8-b77b-04a581e6beb3 · outbound

This paper cites AugGPT: Leveraging ChatGPT for Text Data Augmentation.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search AugGPT: Leveraging ChatGPT for Text Data Augmentation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:49.338985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:49.338985Z digest=sha256:6f55a67897050b56e7cc75341f6a0298b74259e9227a84fafcc8bf59ffbf1ec7

Observation ad3cbeba-4fda-4f83-bb7d-6c5ec7d0dd94 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:49.619619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:49.619619Z digest=sha256:83ecab3bbdbf60f62ac739fa308e8c5f2ee7dc2e0eece4ecef9ebc94414f0eb7

Observation 0b957b4d-0f0a-4132-b5a7-232ce234c2fb · outbound

This paper cites ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:49.470142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:49.470142Z digest=sha256:76a17351f8df8992835e15a0511cc27d8a16a4d0ada46376b407e7e7b40d3f52

Pith citing papers

No inbound Pith citation observations are available.