Pith. sign in

Paper Citation Record · LEDGER

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search

As of 11 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2506.11155.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11155 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:45:51.654456Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1c60edfc-0cbb-4f35-89df-1696c8185719 · outbound

This paper cites The Verifier we used is Qwen2-VL-72B.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search The Verifier we used is Qwen2-VL-72B

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.549246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:49.732140Z digest=sha256:2e424f59b508165bbf74c47110af42876e50207801988d6831a4b9771291caa5

Observation d0cb6278-ba63-4b33-bad5-f1403c087f6a · outbound

This paper cites video caption.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search video caption

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.536479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:49.769808Z digest=sha256:2d0317a00cbbd849f515d96cb5bad0b36480e208cfc0a5b1c5a885e0ab2e0671

Observation 08bb520d-f4a2-4e37-8030-ace8a2b4984e · outbound

This paper cites The defination and examples of each category is described in Figure 10.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search The defination and examples of each category is described in Figure 10

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.494812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:49.897453Z digest=sha256:fc407942d970ee45386fcf3bf0d3dd4ada71673586d8a9915a60dabe14d8bb8e

Observation 43c230d3-60b0-4dda-8539-13f547e19a14 · outbound

This paper cites As shown in Figure 6, the Nature and Wildlife category has the highest number of videos, reaching 382, while the Arts and Creativity category has the fewest, with 57 videos.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search As shown in Figure 6, the Nature and Wildlife category has the highest number of videos, reaching 382, while the Arts and Creativity category has the fewest, with 57 videos

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.523671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:49.807483Z digest=sha256:70f12bc8edf5c32076b91b07d8803c2497b0169442167952c66b25a6cd6a5f09

Observation 0edb9973-083d-4661-9ce8-bb956e41b22e · outbound

This paper cites For clarity, consider these examples: ## Example 1 ### Key Points {1.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search For clarity, consider these examples: ## Example 1 ### Key Points {1

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.128104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:51.250419Z digest=sha256:4718b85406e07261b9680acb60d09c4af61e8ca7c6148f9870e65afa5665e7cf

Observation aca6c052-7370-41b5-9be1-b61b192335c7 · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:49.655855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:49.655855Z digest=sha256:137665a99ceed81ea3df3396a6c4930e2536f02a58f557b0a3f7cd2b83557c12

Observation e43c75bc-cef8-466e-b1e4-d6c27ddc085b · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.510059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:49.877197Z digest=sha256:ff5cc6dc364e9c7b459a2b8c5eb4aee74ee2157e72c819a2cb71ed7b217ceb24

Observation 9128212f-1432-4a17-a791-81d3712fea5c · outbound

This paper cites Elaborate on the visual and narrative elements of the video in detail.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Elaborate on the visual and narrative elements of the video in detail

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.482137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:49.953882Z digest=sha256:ae21d7f4d440c40c8187f7588ecfe15a73718bad7b000dcce838d3f7a04423c5

Observation 09dc07ac-d317-4bee-b07f-d469fb158832 · outbound

This paper cites Reply to me with a precise yet detailed re- sponse.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Reply to me with a precise yet detailed re- sponse

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.468746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:49.986135Z digest=sha256:03d71286bbf5fef63170074cd14d8556aeba1f7960991949199ed04eeb6cdb7e

Observation aa5f7445-cd02-4a89-ade0-b30c3b04a451 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.455747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.040106Z digest=sha256:bda2b574d5995c372e6e744b26d7d622c86be5e69d4a35ee79c2950aaf6fb994

Observation 12594294-b8ea-45a0-8d77-a4ea5371fbf2 · outbound

This paper cites If you are not sure about something, do not include it in you response.\n # Task\n Describe the background, characters and the actions in the provided video.\n”.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search If you are not sure about something, do not include it in you response.\n # Task\n Describe the background, characters and the actions in the provided video.\n”

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.442836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.063737Z digest=sha256:e8fb63697f1675fa714bb97df487ad8b4de05520eeb63b5cba1c30e2fc7eaac6

Observation 09a987e6-3a93-4b05-93d3-bd4e4708cb89 · outbound

This paper cites Please describe the video in detail.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Please describe the video in detail

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.429147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.091040Z digest=sha256:f1dce0ddf577eb065726e6ff40d0b26d3bde1e88870645eb1ff363f7438c3288

Observation 612c6a12-7820-4ec1-8ee6-5fb423a35600 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.416554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.113506Z digest=sha256:719212893b3db4d8d99423b127a7dfda33f08cd9b615579307b66bf3eeacf23d

Observation 8ff35304-bb11-4a01-a8a6-9a32d9b4da57 · outbound

This paper cites Action Description Action description focuses on the specific behavior or activity that takes place in the video.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Action Description Action description focuses on the specific behavior or activity that takes place in the video

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.404105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.157246Z digest=sha256:545f5e1cb74e8890eee49c75b19c112bb1d92167ffe4ec6777fea86ec4dec12f

Observation 5e35b046-e14c-4907-a93c-8cf020039730 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.391772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.227646Z digest=sha256:88a43864f5718db689eb0b984cbe5695ec63dc77071637fe71646ef8bc255176

Observation f6407169-11de-4a84-b3c8-456171823de0 · outbound

This paper cites Environment DescriptionEnvironment description covers the background and environmental features of the scene in the video.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Environment DescriptionEnvironment description covers the background and environmental features of the scene in the video

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.380039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.272752Z digest=sha256:7bbb04d8499405fffee582a07dbe9d8e4334098a9d961d17143be02b1b21215b

Observation 0a7688bf-5e33-47dc-8ebd-8049fa3c40e8 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.367175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.321552Z digest=sha256:77c60c28bf5ce88573376216e69f0cea607f334404dacd835d01382c941707a9

Observation 1c4f6e0d-c3a7-4497-a5f7-082058f8b5ac · outbound

This paper cites Object Description Describes the features, characteristics, or details of inanimate objects or items present in the scene.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Object Description Describes the features, characteristics, or details of inanimate objects or items present in the scene

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.353915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.384862Z digest=sha256:bdf3f2d3a6a33e9ac41dae88fc26f302e8ecedf9ed25566460cdaf88a89a7e61

Observation 36aad59d-79c6-4bef-9b3c-afc83c6f062f · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.340986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.433451Z digest=sha256:bdec325020b7f91b6646967e527759a33d208b6846504af081f388c4ff8202db

Observation e9a41f06-b236-4622-b32f-2204c25f501c · outbound

This paper cites Camera Movement Describes the camera angles, movements, framing, or other cinematographic techniques used in the scene.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Camera Movement Describes the camera angles, movements, framing, or other cinematographic techniques used in the scene

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.327006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.493367Z digest=sha256:5e1eaf3dec97b2a486b6f89424f3e64bb5e4db24e9bf271e69dcd4d29c518b34

Observation 6d106cb3-23e4-4936-81ab-e38ff7579a41 · outbound

This paper cites Overall Description.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Overall Description

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.312951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.556622Z digest=sha256:acdb7a5d9849de46607f74767a1eb77a2496275f4b625fed9bc94d5e75cc7030

Observation e2c286b5-8623-4035-8ed6-1564be95773f · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.299776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.602919Z digest=sha256:5c5f8509bcf532f6d55ab47d68e246430e017daba03ac5a2a863ab59e6901828

Observation 44daad0e-240f-45f7-a52a-fefdb635de9b · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.287434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.648467Z digest=sha256:5616d79f54791c82c3bcadce46a199c8cd9ff3e4bcadb1787d79c8fee68485bf

Observation 207efaf8-4853-479b-b148-e0c90cf7d968 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.273193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.685906Z digest=sha256:6380dce85f2c9cb559603b08b31c3e11de0861b32da464beda4f438d5ec511d4

Observation 5a01bb6c-ea31-42a4-b312-5f45cfd4fdf3 · outbound

This paper cites Judgment: [yes/no] Reason: [Brief explanation].

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Judgment: [yes/no] Reason: [Brief explanation]

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.259294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.709361Z digest=sha256:938f3523835c3dfa1bbec84845373cff71f9ae967215a2807a29b1f842c672ab

Observation e7a52ff5-1705-41f0-bb08-4f714e14fd67 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.246129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.751336Z digest=sha256:e6978ca6622b1897d841380b206bbc8c96b9fd44b244c794355032927829a48f

Observation 73134215-1d39-4897-8815-04b2c4f43b6b · outbound

This paper cites For clarity, consider these examples: ## Example 1 ### Video Description: The video showcases.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search For clarity, consider these examples: ## Example 1 ### Video Description: The video showcases

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.233155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.820854Z digest=sha256:2c91d0bc6a1dd3a7551e084412a120202dd225e4cbfa0dcc472b67f0ddd1ed15

Observation 44eec224-6230-442b-a74a-18647b9d08b2 · outbound

This paper cites entailment.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search entailment

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.219323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.866840Z digest=sha256:61068d48400b7ac49ac887a9761a540653c78ed87ff1177cb5ce158151f138a0

Observation e3a218d6-bbb0-476d-936c-c61e9096d38d · outbound

This paper cites contradiction.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search contradiction

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.204952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.920062Z digest=sha256:45f95f7b16546462baa0027325d58f873dd0bc7ff97569c6bd891149565f9812

Observation deaf138a-b65f-4a5b-bfd3-379f268f0095 · outbound

This paper cites neutral.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search neutral

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.191475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:50.988196Z digest=sha256:53932626f466ec06fc8adadb3546474f12108478e8368e7c5d398eca42bbb818

Observation f04dc7e4-f631-4893-888d-c690782b4841 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.179059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:51.039572Z digest=sha256:668734e650015be7b2639d10f4e9c014c407a9fc1bbf71b847ba77a6d4fd5ec1

Observation 2dead321-8060-4d48-b8fe-099d903d3c0e · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.166708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:51.135838Z digest=sha256:65e32c3c76925210189d802932f75550e5b5549df80a1da0687de2e313bdf33b

Observation bc4c0a8c-5d0a-4cda-b690-413ff0c6f458 · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.153967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:51.205648Z digest=sha256:a3ef8a6854823f120e292feb6452be4e7e1df2fb6dc43e555b8441dd047cfe5c

Observation 6f322406-4ade-4809-85c7-01a3c1c5245d · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.141384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:51.226275Z digest=sha256:149553d8f2e448a089a1878ef1ac27bb7747a8befa623a28ccfa0d302f05e32d

Observation 209f6f20-0529-436b-bd45-60bbf283b7a3 · outbound

This paper cites [No] Speculative.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search [No] Speculative

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.114620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:51.286328Z digest=sha256:112deae2a09447425ebe99f16edad9623de8c77c8204b860d50f01b33860d453

Observation 4e58d046-f8ca-40fd-94a3-9631b7ad3b35 · outbound

This paper cites </thought> tags.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search </thought> tags

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.101011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:51.329282Z digest=sha256:8451a6e664d42a86f178bf6fb86d466981d4838ad0eb514b7ce96f768b8e71c3

Observation 440f3780-deb9-417e-923a-65bd5d676425 · outbound

This paper cites Overall Description.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Overall Description

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.086183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:51.449655Z digest=sha256:0ceebf292784edb61282c122bcae650c05946ff717dc5786177d6df330aff9e2

Observation ae45ca9a-7395-467c-9248-59acb4d784ae · outbound

This paper cites ## Example Input and Output: ### Input: Overall Description: {overall_description} Observation: Vehicles Key Point: There is no existence of any vehicles in the video.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search ## Example Input and Output: ### Input: Overall Description: {overall_description} Observation: Vehicles Key Point: There is no existence of any vehicles in the video

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:52.072730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:51.496957Z digest=sha256:e3bfe214a82bb8752e47b7830fa974d9c08027bb3642a5168b13900fe7be2435

Observation ba2e7ca2-4f59-4f63-a9b4-b4249dfbf04d · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.059285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:51.561541Z digest=sha256:1444823e0b9158742cc43c2884aba2012b3d69bbd44a87ab5f2e0e78d57c57f0

Observation 943bbb7a-b0a6-4ed1-97dc-7efff2e687fe · outbound

This paper cites an unresolved cited work.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:45:52.045858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:51.625504Z digest=sha256:10377a6988902c85c5e5daa0d69f0ddfd260fb9bd9dfd25b7511c415d5f15bbf

Observation 5cd90b8e-b737-4cb4-9a01-69db871d2438 · outbound

This paper cites Precision / Recall / F1 Score.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Precision / Recall / F1 Score

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:51.968280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:51.654456Z digest=sha256:4f7353bb19871ad965d744c632cf9f8a7ca5a704a2e47e6cb0be555e0efad780

Observation 91c77bd8-51be-499e-a1b7-18721181226a · outbound

This paper cites CRiskEval: A Chinese Multi-Level Risk Evaluation Benchmark Dataset for Large Language Models.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search CRiskEval: A Chinese Multi-Level Risk Evaluation Benchmark Dataset for Large Language Models

Reference 2013

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T04:45:51.840389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T04:45:49.548866Z digest=sha256:921fecc7d2310c17781340c17a58a389d9517d951c7e18b85524814e1e560f88

Observation 8cc72c16-a1b7-44d9-ba64-53bcd81e4d9c · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Reasoning with Language Model is Planning with World Model

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:49.431685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:49.431685Z digest=sha256:f027cfa1c3d169d08bb39bee7fa062fbf71acbc3e307b72d5adaf984d2a73351

Observation 8b0c2260-b588-4876-b0f3-809ab39e684f · outbound

This paper cites A Survey on Data Augmentation in Large Model Era.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search A Survey on Data Augmentation in Large Model Era

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:49.695413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:49.695413Z digest=sha256:3aade97cdfaf5652d24a3645819af692e01a4564f1ec44207f45e8aed2ef1f0f

Observation d31a7ef3-db04-4ac8-b77b-04a581e6beb3 · outbound

This paper cites AugGPT: Leveraging ChatGPT for Text Data Augmentation.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search AugGPT: Leveraging ChatGPT for Text Data Augmentation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:49.338985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:49.338985Z digest=sha256:d7e1e6a65023f7214d437d06b8bcfba3fb6055505a02b65e564a167904361f18

Observation ad3cbeba-4fda-4f83-bb7d-6c5ec7d0dd94 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:49.619619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:49.619619Z digest=sha256:f6db0759287793fa336c914924b924f672df1a50709cb740fefa21c1afeefd9e

Observation 0b957b4d-0f0a-4132-b5a7-232ce234c2fb · outbound

This paper cites ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents.

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:49.470142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:49.470142Z digest=sha256:ab18780f946326552c9a917179448ae1713443a0679fde298ab911664a44154b

Pith citing papers

No inbound Pith citation observations are available.