Pith. sign in

Paper Citation Record · LEDGER

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models

As of 19 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2501.00432.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00432 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:56:23.789593Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact2
  • verified fuzzy21
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87e86ae7-c9a5-44ca-a0f7-d84752eff623 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.603998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.603998Z digest=sha256:c862b9a32853f480c900c699c48d34d3f43b0c22e1179826d23cffc3ab135ebf

Observation 38df0fac-9e36-4d14-92a5-48fe7147a714 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Gemma: Open Models Based on Gemini Research and Technology

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.610325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.610325Z digest=sha256:ac87712e36a7b5385d117290fc71c15226d4787bd626aa4805a9d36b67136cec

Observation b2a1420b-504d-4ce1-8243-97e034835968 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.472616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.615502Z digest=sha256:9ee157e5880b0b118afde8d21a4548c2f478b9e2f707794d66f431e4f9db8fa6

Observation bb9d0b30-a686-48cc-9731-7bcfe6af0515 · outbound

This paper cites Language-grounded dynamic scene graphs for interactive object search with mobile manipulation,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Language-grounded dynamic scene graphs for interactive object search with mobile manipulation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.455214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.620658Z digest=sha256:9247148e20a4b01000a87d195895f2e52e7533754fcf611dcf424d7d9637cf83

Observation 4c4088ad-c387-4a9e-86b6-dd42d59b8338 · outbound

This paper cites Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.625900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.625900Z digest=sha256:989eb760bf028e9a333f9659e65fb80e374bd3aab5c2b7ade33fa82da3d3d0cb

Observation 065d66d5-1452-4bf7-a252-c822fe8f44e6 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Code Llama: Open Foundation Models for Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.631491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.631491Z digest=sha256:86d46e031254d7212fb35972e15d6b2b25264c9be6e7e2dadf2cb22040d5a5df

Observation bc779ece-c635-45cb-bb9a-a23e874f2ed6 · outbound

This paper cites CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.637769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.637769Z digest=sha256:86539edf62bfff248b6baa376cf02bb8a604b0f3722d52f182b779bfef5b70b5

Observation 03ec1088-e661-4fce-8483-36849b6b1f6b · outbound

This paper cites Text me the data: Generating ground pressure sequence from textual descriptions for har,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Text me the data: Generating ground pressure sequence from textual descriptions for har,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.437533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.642943Z digest=sha256:f5e9389d68964ebbac2af607c7ab90a18259f755490c43f47456d890d6255966

Observation d096f72a-b7db-47f1-8beb-f690a8bd9a96 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.647629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.647629Z digest=sha256:0060668f095bdcbb00b4ce9fbd9ed244d6b32ba0ee13d6f2e6b7259ebe6c3a8d

Observation 72ec3e0b-5e59-4793-b5ea-8f0473dbf639 · outbound

This paper cites Bliva: A simple multimodal llm for better handling of text-rich visual questions,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Bliva: A simple multimodal llm for better handling of text-rich visual questions,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.421483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.652741Z digest=sha256:ca117d0d779523c101fae7910cc0830155b937a6cf773f2ca07214c067ab1479

Observation 6cdbf847-6e40-47f2-885b-ce49ad5fcced · outbound

This paper cites Infogcn: Representation learning for human skeleton-based action recognition,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Infogcn: Representation learning for human skeleton-based action recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.405639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.657569Z digest=sha256:c04b997f401e11eb2ae332cd5858eb142f71cda0a5bb6b52fa79baa4643202a4

Observation e2742733-127e-4afb-b79f-0fe78d0b7222 · outbound

This paper cites MuJo: Multimodal Joint Feature Space Learning for Human Activity Recognition.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models MuJo: Multimodal Joint Feature Space Learning for Human Activity Recognition

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:56:24.011915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.662217Z digest=sha256:ea3d27006b2b0d34208c275747fa0106a83ddf51f5762b67f9387c85072e46d3

Observation c0d010bc-eef7-4ae9-850f-5435a80c912e · outbound

This paper cites Decoupled spatial- temporal attention network for skeleton-based action-gesture recogni- tion,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Decoupled spatial- temporal attention network for skeleton-based action-gesture recogni- tion,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.387833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.667838Z digest=sha256:020c2267e377c5df59fef464edad6f88c79fed74ea448fb47fa54ba8fcd40a42

Observation bde806a5-987e-46eb-a94e-645309497de0 · outbound

This paper cites ALS-HAR: Harnessing Wearable Ambient Light Sensors to Enhance IMU-based Human Activity Recogntion.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models ALS-HAR: Harnessing Wearable Ambient Light Sensors to Enhance IMU-based Human Activity Recogntion

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.672238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.672238Z digest=sha256:a844a2a27ed944eb9fe63ca19fc84b96e5258c5e36886992679e7cdbe69b169c

Observation 4617cad2-2272-4829-9dea-f1deff563e5c · outbound

This paper cites Human- to-human interaction detection,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Human- to-human interaction detection,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.371202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.677125Z digest=sha256:013619adc8af218f191aeb08dc473339e845669f6a75c6f630e5b9d1b1c34da2

Observation 64f0c954-6907-4a33-89d6-f0f86dd997b9 · outbound

This paper cites A Two-stream Hybrid CNN-Transformer Network for Skeleton-based Human Interaction Recognition.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models A Two-stream Hybrid CNN-Transformer Network for Skeleton-based Human Interaction Recognition

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:56:23.971278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.681815Z digest=sha256:1f3e52f693bd76b91ff39e4da01bcfe6599128815b251b891b367356a6b356fc

Observation 395dfb42-dcd6-46c8-b341-a0f9637baa1d · outbound

This paper cites Hargpt: Are llms zero-shot human activity recognizers?,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Hargpt: Are llms zero-shot human activity recognizers?,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.355129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.686911Z digest=sha256:46744604177068e7548af6c27dcf140ca83030a0bc0da2e8ca7da0ebc37a1d34

Observation 657d9849-0e6d-4ebf-8902-ccfec0004185 · outbound

This paper cites Unsupervised Human Activity Recognition through Two-stage Prompting with ChatGPT.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Unsupervised Human Activity Recognition through Two-stage Prompting with ChatGPT

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.691724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.691724Z digest=sha256:c4c4db39ce1caa94beeb7addde20af9884bd88b5d8615c578a3cd289e3c2a93a

Observation 23e7e97a-c047-4ddd-9efe-647631874ae8 · outbound

This paper cites Segment anything,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Segment anything,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.338026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.696802Z digest=sha256:2fd753c4cdda2faa0d865f0a9ebec0ce297a7b3eb66583dd668019affa68720e

Observation 33c1a0c1-e98c-4e43-85b1-977c4e80d32a · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.320721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.701590Z digest=sha256:1597e803d0dc4eaed7a64ec1ab09fcea3d8a75ade8d5a54ce4e634cd8ea7b5b4

Observation 5133d8c2-772f-4b07-92b2-f6779863812e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.706672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.706672Z digest=sha256:4af2cd40ec8e6a1a60bfc6c15eade096f7eafd2b047fb207fc81383b0d9286b3

Observation bd55a662-375e-404f-a130-6b7fa18f9898 · outbound

This paper cites GPT-4 Technical Report.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models GPT-4 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.711560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.711560Z digest=sha256:0a5310a554fda01dc4f791a466c10930f3a8941bc171983e533e738b0dba656b

Observation c3bafba6-9543-4549-bc43-8a2c5ec89cca · outbound

This paper cites Track Anything: Segment Anything Meets Videos.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Track Anything: Segment Anything Meets Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.716449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.716449Z digest=sha256:e981b6fa44d98f243244158457a4904ef8292b45d1d4d5d151f95592f00f4ffd

Observation b899e3e7-ec58-4afc-83ad-3c78f8d0bb6b · outbound

This paper cites Actions in context,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Actions in context,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.304524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.722466Z digest=sha256:38b497871416b6bbc24966d09c660fc3744f0400c07487279c82be71a84b9cce

Observation f2ed5518-f78d-4e2f-aa19-02a0c6d631ca · outbound

This paper cites Sportshhi: A dataset for human-human interaction detection in sports videos,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Sportshhi: A dataset for human-human interaction detection in sports videos,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.288579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.727378Z digest=sha256:2d7fdfe39dd7c8ce45391d1b4882724fafea695631b87ec471b0c4650b5b5f08

Observation 81615868-aa76-4666-8700-2b1e3b646c9e · outbound

This paper cites High five: Recognising human interactions in tv shows.,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models High five: Recognising human interactions in tv shows.,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.272260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.732128Z digest=sha256:baff95151904a154fb2f7fc224e74ae647ccc5a757646a3f675ce5d7c41e9084

Observation 775c389b-6dff-4080-9af3-82b417bfe29c · outbound

This paper cites Two-person interaction detection using body- pose features and multiple instance learning,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Two-person interaction detection using body- pose features and multiple instance learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.255791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.736755Z digest=sha256:1e5377ad9b9b571b14846c2bf44b8d9e1739e746d064fbcc69254e4ae7a11521

Observation edb78eb0-4d28-49be-9d84-56f615613598 · outbound

This paper cites First-person activity recognition: What are they doing to me?,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models First-person activity recognition: What are they doing to me?,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.239362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.741467Z digest=sha256:93a7fd10bd3df66542b1fa65c12dac48695a8278aabe5c191ab714f45c9226c0

Observation 11898bc6-75b9-4896-8fde-d47ebddf5a3d · outbound

This paper cites Interaction relational net- work for mutual action recognition,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Interaction relational net- work for mutual action recognition,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.222094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.746258Z digest=sha256:fda8262603272c6933414ed24846b5d05d3fbddbccda5fd99aaee75785efa6a8

Observation b0c9739a-7184-421c-ac69-12b35725793c · outbound

This paper cites The Kinetics Human Action Video Dataset.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models The Kinetics Human Action Video Dataset

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.750838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.750838Z digest=sha256:709e6a5ecc8add79107f024d2450a47f09f2e5669081a41ed05de6c0cdd7f3fc

Observation aa397625-c615-4051-ad4f-7954a0ce4db9 · outbound

This paper cites Air- act2act: Human–human interaction dataset for teaching non-verbal social behaviors to robots,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Air- act2act: Human–human interaction dataset for teaching non-verbal social behaviors to robots,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.204484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.756151Z digest=sha256:ffc1c0b6563c85d17d603b6b9304e48f9d1fe0aacc558ef4a4efbf55e9023de4

Observation 579d3cd5-eaba-4fd1-82c2-27bad80a2ba5 · outbound

This paper cites Human behavior under- standing,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Human behavior under- standing,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.187189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.760835Z digest=sha256:454e7a81f3a95475bfa45eaac1279dfb31db7a3c91db7f46e519fd7649f8141c

Observation ab50a7af-e2f7-4ff8-bc9e-dc65f055ff5a · outbound

This paper cites Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.169124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.766462Z digest=sha256:11377cff3881b2538228a46908176dde4b1cd0589a905e2d34b2bfa43a42f207

Observation b619f521-e1ed-4712-b6f6-6b22246a16ba · outbound

This paper cites Caption Anything: Interactive Image Description with Diverse Multimodal Controls.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Caption Anything: Interactive Image Description with Diverse Multimodal Controls

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.771088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.771088Z digest=sha256:f1e5c9f92092511a492beb48ae7e479e0fd1d466589893e9ca39b9315d1ecc62

Observation 12207379-a3be-4412-a41d-7ea6de79d4da · outbound

This paper cites Vitpose: Sim- ple vision transformer baselines for human pose estimation,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Vitpose: Sim- ple vision transformer baselines for human pose estimation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.151800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.775670Z digest=sha256:9ab082358414122422081c05e44bfdf8451f7a91006eb45532943186476910db

Observation b58fb453-5b87-4516-9d1e-db0f902034fe · outbound

This paper cites Tokens- to-token vit: Training vision transformers from scratch on imagenet,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Tokens- to-token vit: Training vision transformers from scratch on imagenet,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.133634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:56:23.780097Z digest=sha256:e853b8eec42164c3ea221ad754df34a8aa0c450dc57ed814e46713ce66550c65

Observation a500a090-c4bf-43d6-acef-9ca4d3e8d58d · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.784686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.784686Z digest=sha256:505d22ba82d46234fffc6bc7ffff35941668053cc7372b8ccd5a8b293317aad1

Observation 91384ef2-fb44-493f-8492-a01fe3816933 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.789593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.789593Z digest=sha256:dba1115523eacbb2d2f7810fcfd341906516f14f2fa4cda1be46f12af7aede9f

Pith citing papers

No inbound Pith citation observations are available.