Pith. sign in

Paper Citation Record · LEDGER

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning

As of 10 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 0 inbound Pith citation observations for arXiv:2507.13364.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13364 v1

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:50:24.303188Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 105 outbound references displayed

  • verified exact8
  • verified fuzzy35
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d59a1872-489d-4553-9de4-41191b6890f5 · outbound

This paper cites Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.223191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.223191Z digest=sha256:8c1246583d79115f00fc6da4b5c1351393ee0ff9a1561bfe85e9f61df5cfc222

Observation 58c6bae3-fdc2-4237-8b13-d67425b5116f · outbound

This paper cites Objects that sound.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Objects that sound

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.301171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.301171Z digest=sha256:d737c8d479c0bfe8782b198ea29822cc986da1b0174f9ee647fdacd96a6493f5

Observation af1c03c2-eed2-4c63-8061-3e9de9304f1e · outbound

This paper cites 3d seman- tic parsing of large-scale indoor spaces.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning 3d seman- tic parsing of large-scale indoor spaces

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.441879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.441879Z digest=sha256:18951b6a787df07d50491c4d461b582113c25b5b9790806f6989b411d7aff9ab

Observation 39d3b6b5-846b-49fd-87e5-30dc0c0f116b · outbound

This paper cites MAE-AST: Masked Autoencoding Audio Spectrogram Transformer.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning MAE-AST: Masked Autoencoding Audio Spectrogram Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.572549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.572549Z digest=sha256:bb2956f5cb7f82060a627a1b2a23c79781dbf4619d8f50091f88f5e54f62da13

Observation a21ca1f0-fa77-4a38-b1f9-6cd59471f3d0 · outbound

This paper cites Data2vec: A general frame- work for self-supervised learning in speech, vision and lan- guage.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Data2vec: A general frame- work for self-supervised learning in speech, vision and lan- guage

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.702418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.702418Z digest=sha256:412d0cdfc77abd40abf0e1dd8e4ce82ae0c5aa4eb1043cc1dd8ce57b879e71e1

Observation 2bea425a-ac9b-44e4-8ad7-d70b29a3bf8d · outbound

This paper cites Generative adversarial networks based on transformer encoder and convolution block for hyperspectral image classification.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Generative adversarial networks based on transformer encoder and convolution block for hyperspectral image classification

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.854860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.854860Z digest=sha256:545aa6e963de48afbdd0b52ebf94047a0ec8ba441b9720a96fd6eecd56236a0b

Observation 3c461085-a472-455b-901b-a4c5fd5b9a5b · outbound

This paper cites HiP: Hierarchical Perceiver.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning HiP: Hierarchical Perceiver

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.988580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.988580Z digest=sha256:76e30a5323372625fb10443378c459c07ad4b82efff23a24a6f821b41bf0e79b

Observation 56e36f43-8424-40ea-a6a1-99d718592339 · outbound

This paper cites Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:15.091237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:15.091237Z digest=sha256:1a1c83e43de2be4e72ef722da75b9f8742db9f58f1fd03dda61330f92857a22e

Observation 54ea8e34-ffa0-4245-9962-ef434d3f3ae7 · outbound

This paper cites DialogSum: A Real-Life Scenario Dialogue Summarization Dataset.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning DialogSum: A Real-Life Scenario Dialogue Summarization Dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:15.207070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:15.207070Z digest=sha256:578f9c72d0718b254cb20b82e858cef6da9b19be11c469f92078b50233922826

Observation c5a18620-9c84-4893-8dc8-db6d1adb2641 · outbound

This paper cites Multi-Task Learning with Deep Neural Networks: A Survey.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Multi-Task Learning with Deep Neural Networks: A Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:15.356316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:15.356316Z digest=sha256:d138fa7f19310459db8e82c7157820572fcb11b5dfd0112d93b67e93cb58a4ce

Observation 56770bc8-7652-4291-a236-385ac86ae989 · outbound

This paper cites One Model, Multiple Modalities: A Sparsely Activated Approach for Text, Sound, Image, Video and Code.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning One Model, Multiple Modalities: A Sparsely Activated Approach for Text, Sound, Image, Video and Code

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:26.526503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:15.468308Z digest=sha256:b6cb4fce2e8d286c0fd22b8fe8d3f859e41c1ef533bc20e5e09941e478ca0eee

Observation 9f3f72e0-b581-414b-8320-fc49d0f25b45 · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Imagenet: A large-scale hierarchical im- age database

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:15.576775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:15.576775Z digest=sha256:96276a86de5606a637a17e2359e4d18e55a20af094271174d3f7cd1c47b3823c

Observation fd8e2161-2dab-474d-aafc-305379b80e06 · outbound

This paper cites BERT: Pre-training of deep bidirectional trans- formers for language understanding.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning BERT: Pre-training of deep bidirectional trans- formers for language understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:15.745612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:15.745612Z digest=sha256:215093fc834df0db1aab25aab1d8afefff98be4a8719944d91baca542960604a

Observation 37f95f8a-075d-45ee-b344-173b91ad14b3 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:15.911782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:15.911782Z digest=sha256:38536e5baa8a71fb2e854e4c181469040e7ff475cda14e5dbb57e33bc43d12ab

Observation 55f8eeca-e4a0-4882-b8a4-0aead2363c93 · outbound

This paper cites A generaliza- tion of transformer networks to graphs.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning A generaliza- tion of transformer networks to graphs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:16.035240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:16.035240Z digest=sha256:f1f4d3ddf20e190d729afd1f5c0925e6bb459c0c8d5cb8da0617a4281c01ebc0

Observation dc85f918-01e8-42ac-9e37-65f9d313ef58 · outbound

This paper cites Efficiently identifying task group- ings for multi-task learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Efficiently identifying task group- ings for multi-task learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:16.197878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:16.197878Z digest=sha256:be274b07b9906ff59f6c6e6fc2a002077a46505ad23e522d99c3b861fe619e17

Observation dae79af5-b80a-4199-8116-20bd3ce75789 · outbound

This paper cites End-to-End Audio Strikes Back: Boosting Augmentations Towards An Efficient Audio Classification Network.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning End-to-End Audio Strikes Back: Boosting Augmentations Towards An Efficient Audio Classification Network

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:16.324991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:16.324991Z digest=sha256:41f9a9a07d370ba3f530c7781f8ffa30885f123a102492de05ce14f8822ae3cb

Observation d9b07949-7982-4622-b307-34c03a5d24c0 · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Audio set: An ontology and human- labeled dataset for audio events

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:16.490217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:16.490217Z digest=sha256:cef662e4302db17c452bd56ea18a57b58ec74c44fd9a0ff9126d71801d56fd24

Observation 932aad78-8e47-4db1-aeef-9cf77236476d · outbound

This paper cites OmniMAE: Single Model Masked Pretraining on Images and Videos.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning OmniMAE: Single Model Masked Pretraining on Images and Videos

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:26.285243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:16.598120Z digest=sha256:d775e111c4cf8d5f3c4f794a5cc0802f81e17dd8b3d97498bee73d3a1682c8ae

Observation f09ffe82-9b83-4d02-aef9-ae2198364411 · outbound

This paper cites Omni- vore: A single model for many visual modalities.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Omni- vore: A single model for many visual modalities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:16.711982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:16.711982Z digest=sha256:2d7822d3bbb1bd7feb0dc9654d535a97e66d66b1756994883679a3ec0b974c72

Observation 475ea48a-1140-4d20-9ed5-3805e5e0fc1e · outbound

This paper cites SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:16.915486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:16.915486Z digest=sha256:418b17ff04ec1f5d8dc6da5db7786c7504b317082f23c721a120eb7a70f17121

Observation c5ee46f4-7d1b-482f-8cad-1e106d61c0a3 · outbound

This paper cites AST: Audio Spectrogram Transformer.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning AST: Audio Spectrogram Transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.039560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.039560Z digest=sha256:ef936d0809bf050a3798f9f05cb56f5a8fe72ff2a499ee7b8b83e63f1bd7b1dc

Observation e59a712b-21b1-4fb7-81b0-22fb95d4f43f · outbound

This paper cites Uavm: Towards unifying audio and visual models.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Uavm: Towards unifying audio and visual models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.202842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.202842Z digest=sha256:1bf674bf93397894e1acf06ea4801beec37a105902ff9aa3d328ec09ae7ed7c9

Observation 1ef8eb6e-cdb5-41bb-9fbe-0567c56ac70e · outbound

This paper cites The” something something” video database for learning and evaluating visual common sense.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning The” something something” video database for learning and evaluating visual common sense

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.322899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.322899Z digest=sha256:704ae393e57f83195991a252ef6e1a246486b7998cdaeaf9ff858da2f6b36b5a

Observation 4684aabb-05a1-438c-b075-9b2859961c47 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Ego4d: Around the world in 3,000 hours of egocentric video

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.420096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.420096Z digest=sha256:732371328c1a767e66730776a0770d5e2be33dc089129f60c312be8bc6049ece

Observation 718fe2f5-fe32-4d5c-84fc-5ac5709e6bdf · outbound

This paper cites Dynamic task prioritization for multitask learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Dynamic task prioritization for multitask learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.507498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.507498Z digest=sha256:32ce5c25eac8bfc199c951c6c2762972a449fc7e071bedf9d89e282575007041

Observation 87e16809-6d0e-4eda-b546-ff98c30ed626 · outbound

This paper cites MaskViT: Masked Visual Pre-Training for Video Prediction.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.617134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.617134Z digest=sha256:b45c1d5ddc82796fd9262b68b52a6c620414d8398ccecbd0fa1cd38e10b10c1e

Observation e0b9e3d5-d3ef-4db0-92b5-c89935dbf407 · outbound

This paper cites Masked autoencoders are scal- able vision learners.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Masked autoencoders are scal- able vision learners

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.712917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.712917Z digest=sha256:daaa1946c965a7508792223015d19b356f7960a9a0835bcf89af1d4e89d0a4c6

Observation 0a79934a-741c-4e41-b197-0f059c137274 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Gaussian Error Linear Units (GELUs)

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.826839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.826839Z digest=sha256:529575eabdc8a7cf77629cfbac51f6d15bf303b27968667bab49be718e871fcf

Observation 1f89a59d-b2a3-4f03-bbb4-dad5dfd4add6 · outbound

This paper cites Spectral- former: Rethinking hyperspectral image classification with transformers.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Spectral- former: Rethinking hyperspectral image classification with transformers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.932549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.932549Z digest=sha256:0611af02bdcad83e2e056570ab8b420f9d38971e94737f0ed5de75dc0ba7271d

Observation 621186a5-ba8c-416d-8ff9-bd93cdde9a52 · outbound

This paper cites Unit: Multimodal multitask learning with a unified transformer.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Unit: Multimodal multitask learning with a unified transformer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.021377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.021377Z digest=sha256:45451a4e0390077aa64aef17727c4961229c03d2a41cc8854daa28a862db2aaf

Observation 84660675-7b4b-4f5c-82f4-da2d19e23c2c · outbound

This paper cites OGB-LSC: A Large-Scale Challenge for Machine Learning on Graphs.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning OGB-LSC: A Large-Scale Challenge for Machine Learning on Graphs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.119583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.119583Z digest=sha256:bce13b6e5fe5d665e630f3b50cebc785ffcc10e327d3352619ee3e32dd1eb11f

Observation 342a32dc-600d-4380-a433-3c3e8138eded · outbound

This paper cites Perceiver IO: A General Architecture for Structured Inputs & Outputs.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Perceiver IO: A General Architecture for Structured Inputs & Outputs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.251387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.251387Z digest=sha256:bc7c3bfbb99f560bc3fd92aacb9736718737774a8af805e523efe0eeb66522db

Observation a5d0eb73-3086-4f59-aead-67f1a0556924 · outbound

This paper cites Perceiver: General perception with iterative attention.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Perceiver: General perception with iterative attention

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.379108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.379108Z digest=sha256:58037d186519714be8d98015b45d44df3bc117f895de88086a187b0af99020d9

Observation 9d9f546e-fc24-4b2c-89fc-3a190f2f92bb · outbound

This paper cites A review of multimodal image matching: Methods and applications.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning A review of multimodal image matching: Methods and applications

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.514139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.514139Z digest=sha256:135493eb1ddfde55caa59f8d5a85b52b8528d808831769254705f53b47d9b1bb

Observation 9cd1fe4c-4482-434f-b3e5-f327b37fc8f1 · outbound

This paper cites One Model To Learn Them All.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning One Model To Learn Them All

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.598502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.598502Z digest=sha256:074d7cf5e62aec1f7f74a4fd95dff2252376084538826ea6276d4d2c30dce16a

Observation 2c9ea7e5-0131-4774-9b13-11500b736714 · outbound

This paper cites The Kinetics Human Action Video Dataset.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning The Kinetics Human Action Video Dataset

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.711195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.711195Z digest=sha256:c6e2e87bf9e9cfb6b799390d712e66e4f1528ac2300d55c39911cc5581d5f79b

Observation 30d3f7e6-cc88-430e-8730-dc40862b5df6 · outbound

This paper cites Mind the Gap! Injecting Commonsense Knowledge for Abstractive Dialogue Summarization.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Mind the Gap! Injecting Commonsense Knowledge for Abstractive Dialogue Summarization

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:25.963938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:18.839182Z digest=sha256:8e0521a3c66d8490a403cbc8f27e0d3291b4dccc6b57a778bebfdb69a0a12fe7

Observation 0b31472a-8eda-49b8-9476-252fb415548f · outbound

This paper cites Re- former: The efficient transformer.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Re- former: The efficient transformer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.937345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.937345Z digest=sha256:adde09efb5408f459b0d003e7edc89caea97cecd06253eca5c7bc76e0188629f

Observation e7b145c0-c1a0-4ea9-b71c-f24d22135333 · outbound

This paper cites Hmdb: a large video database for human motion recognition.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Hmdb: a large video database for human motion recognition

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.051884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.051884Z digest=sha256:760c5d2a32fc4764ba62820e0ef774743b5655610ca653e5237d4a07400e6446

Observation 1345fdee-e667-447e-8af6-ab88c4dd7d26 · outbound

This paper cites Modeling long-and short-term temporal patterns with deep neural networks.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Modeling long-and short-term temporal patterns with deep neural networks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.166703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.166703Z digest=sha256:5a0391881c766001d4e97f69ec7a8a0438675ef03842f8663020e4b730c14b82

Observation 2b5de6d7-497d-47c5-9c91-c72f11041bf1 · outbound

This paper cites Stratified trans- former for 3d point cloud segmentation.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Stratified trans- former for 3d point cloud segmentation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.290419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.290419Z digest=sha256:a7ae510a8eab5af06c5e666eb2a9f0e03ad4df665cfa8b99c103346a92681e4d

Observation 6230888b-edf8-4947-8121-8823afc854e5 · outbound

This paper cites Regu- larization strategy for point cloud via rigidly mixed sample.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Regu- larization strategy for point cloud via rigidly mixed sample

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:32.311366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:19.428323Z digest=sha256:a117637b4dd9dd9fbb97547a2246a1e2dbcab7b412ff70b7a4fef40770e4764a

Observation 1d5913ce-11ba-4c0d-8c9e-28202a97a75c · outbound

This paper cites Uni-perceiver v2: A generalist model for large-scale vision and vision-language tasks.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Uni-perceiver v2: A generalist model for large-scale vision and vision-language tasks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:32.222632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:19.541910Z digest=sha256:ceb33f47f783168a9acb23b40030a8425cce2be18948dbd5b9ac66810b0692fa

Observation f1a474c0-6c44-45d3-9420-44286c7b83b3 · outbound

This paper cites UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.646948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.646948Z digest=sha256:2e54457775bacc1cf62066793b841bc141f20d4364878e886f0d2265a23c50c7

Observation 941ae255-0a70-4627-8651-0196319e1ce4 · outbound

This paper cites Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:32.074110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:19.768028Z digest=sha256:b9b19e5ec2fe99f6888f0759d13a273438816b7aca3eb7f8b5c9a3518af02664

Observation 9a8f0407-498c-4677-afa6-5800b0e5c46b · outbound

This paper cites Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.855755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.855755Z digest=sha256:4336f504bd60c0831ccc83a34739e2a9fef767775ad66e9a7280f4bbb1f1152a

Observation df00d131-ea42-45f1-bd8c-76ec203b2e1c · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.892880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:19.953525Z digest=sha256:d64323b629ccd2b835e1a92bb4b5e09ec053a0158e720967ebb78498204df686

Observation 56008d9f-43a8-4667-b043-2ab6eeb4e977 · outbound

This paper cites OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:25.761552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:20.051811Z digest=sha256:35ab0cd67152602300ab2d8168f13078b2cba45f2ab37b816eee081abd538c90

Observation 8f652af2-877d-4b29-90cc-df16bbcac204 · outbound

This paper cites Pyraformer: Low- complexity pyramidal attention for long-range time series modeling and forecasting.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Pyraformer: Low- complexity pyramidal attention for long-range time series modeling and forecasting

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.679062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:20.168295Z digest=sha256:a2ba9ff09fece4841f5fbb06ef87a7c38a43fa5b28f7f4a478a4533e5db4d0f8

Observation 08f78ae6-58ba-42ce-b471-ac7950235867 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.304795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.304795Z digest=sha256:35c4e560b4542a9d4034821eb00cfaf507a0693842326024be0028524dcfb213

Observation 194acbc9-711a-48bb-9adc-f519ce36393a · outbound

This paper cites Moments in time dataset: one million videos for event understanding.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Moments in time dataset: one million videos for event understanding

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.539943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:20.398454Z digest=sha256:e74c51274a9bc8bdb02695984c2a3dc6bacd07378057a19e03f3fd79ff341908

Observation 43b4d318-5029-4061-b742-6bc3205a9dad · outbound

This paper cites Person recognition system based on a combination of body images from visible light and thermal cameras.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Person recognition system based on a combination of body images from visible light and thermal cameras

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.423488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:20.525721Z digest=sha256:ee77c60475ab73532ecf3ebe866f89045d01e1f30d0a77f6ac6fe48f051ca0da

Observation 43fa27db-de94-4685-827b-56ae6cb30588 · outbound

This paper cites N-BEATS: Neural basis expansion analysis for interpretable time series forecasting.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning N-BEATS: Neural basis expansion analysis for interpretable time series forecasting

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.661670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.661670Z digest=sha256:01d27fd37eb6104fac5f80e0bbd2b7df381f3c725c4416e1cbfd70469474d661

Observation 60e8e3d8-5162-4883-a951-b6a1b9b0e488 · outbound

This paper cites Cats and dogs.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Cats and dogs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.767129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.767129Z digest=sha256:d5ac4448d25e35f2add8e0e6da188f9016eeb04f3d9e8e5f23e9e07ff780496b

Observation ed785505-9347-4868-91d0-3a3ae50cc0af · outbound

This paper cites Esc: Dataset for environmental sound clas- sification.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Esc: Dataset for environmental sound clas- sification

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.311763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:20.883668Z digest=sha256:aec23245a83117e2d10d62f0356ca12e626b8ecb3a23067ef807493dd5ac5b25

Observation 6b0d0f90-f56e-4ab7-8257-a0b3f156ec28 · outbound

This paper cites Re- thinking video vits: Sparse video tubes for joint image and video learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Re- thinking video vits: Sparse video tubes for joint image and video learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.184763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:20.990939Z digest=sha256:5e491b91a502a2111f1436db36f798313b8e6856ac749560714aed237e8e4bc8

Observation 8a5f51c5-9667-473c-9240-5e22df4fd717 · outbound

This paper cites OmniNet: A unified architecture for multi-modal multi-task learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning OmniNet: A unified architecture for multi-modal multi-task learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.121011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.121011Z digest=sha256:d665fc70c5ddcdf378f119fbd2cc4e7c5dd94fe73444ef9e3bba692272455494

Observation 960b3122-f46a-4655-8e8c-c6780eec6014 · outbound

This paper cites Point- net++: Deep hierarchical feature learning on point sets in a metric space.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Point- net++: Deep hierarchical feature learning on point sets in a metric space

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.040826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:21.235366Z digest=sha256:4d1f4b6792fbf8c9156c2610cf9f6e297f06da3a69e67294aa2fc4d35d5c98a9

Observation 85a502b9-ace2-44b5-b02a-a1445bda232b · outbound

This paper cites Improving language understanding by gen- erative pre-training.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Improving language understanding by gen- erative pre-training

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.380262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.380262Z digest=sha256:fffc0e2b996e13e6f2716b6f714a40b0af8e8973f4a6830a63e703df29bc36a7

Observation cebe27b6-d809-402e-aa63-1b51e2012ace · outbound

This paper cites Reliable tuberculosis de- tection using chest x-ray with deep learning, segmentation and visualization.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Reliable tuberculosis de- tection using chest x-ray with deep learning, segmentation and visualization

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.946155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:21.519423Z digest=sha256:b677926640648159dadcbaf2b75e2f6771881c2930d9cbd563e3381efff90d89

Observation 5f88d135-92c1-4203-823c-dee0e0d6a492 · outbound

This paper cites Zorro: the masked multimodal transformer.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Zorro: the masked multimodal transformer

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:25.551900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:21.672452Z digest=sha256:e5c27cfd57ec69ef48fadfbab1b70ec6e1858d4ef873b81d2e967c0cba7b0088

Observation 3f24eccf-3cc9-461e-a080-e9ab0c57190f · outbound

This paper cites Indoor segmentation and support inference from rgbd images.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Indoor segmentation and support inference from rgbd images

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.817984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:21.776607Z digest=sha256:41b0eff5867833f361b1ff393d6af28be85ec46a7c48bceba66a6e0b9d738891

Observation 21993166-f068-41a5-84ad-b57e3fca20bc · outbound

This paper cites Mpnet: Masked and permuted pre-training for lan- guage understanding.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Mpnet: Masked and permuted pre-training for lan- guage understanding

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.698263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:21.861378Z digest=sha256:a5ffd4efd08ff4a61adf588a7e4e564531d0aa1d9b3eca4fc68e570336770429

Observation 9a540032-3568-429f-af77-4e6fe5f3d6a0 · outbound

This paper cites Sun rgb-d: A rgb-d scene understanding benchmark suite.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Sun rgb-d: A rgb-d scene understanding benchmark suite

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.580490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:21.997522Z digest=sha256:f86619b58fe0d24dc78fad8f0133ad51e99ff26744700034b9ea6ada7a0c6fb0

Observation f8a1fc2e-e727-48e9-b9d7-b1eae07b61c2 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:22.073655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:22.073655Z digest=sha256:7a2eb06b16712f4b0acc8f0637c90cfa2915823610f37ecb7106bff339236e97

Observation 3fed73df-e32b-4259-aa76-b7be8cddd991 · outbound

This paper cites OmniVec: Learning robust representations with cross modal sharing.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning OmniVec: Learning robust representations with cross modal sharing

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:22.147490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:22.147490Z digest=sha256:529c2ca29db8565037a411113e74f0c6bfd03691afc2eccce5a3838d7e16a7ab

Observation 8049192c-b68b-414c-a96b-c586b6cd28ba · outbound

This paper cites Hierarchical multi-task learning via task affin- ity groupings.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Hierarchical multi-task learning via task affin- ity groupings

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.455919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:22.210014Z digest=sha256:0c53941eab374f983201e167ad26ef076be3e2f4cb08ca391cfa8c0c3279e004

Observation f3edae28-32b7-4164-8523-4d5b73918175 · outbound

This paper cites Benchmarking Robustness of 3D Point Cloud Recognition Against Common Corruptions.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Benchmarking Robustness of 3D Point Cloud Recognition Against Common Corruptions

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:22.270903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:22.270903Z digest=sha256:16b8b814ea9a8956dfad3a7b909ccff357f8e16beb66601b06eda377669c5b3b

Observation edf6e8df-359f-4bad-8d02-a2c8f3fbc02c · outbound

This paper cites Efficientnet: Rethinking model scaling for convolutional neural networks.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Efficientnet: Rethinking model scaling for convolutional neural networks

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:22.333424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:22.333424Z digest=sha256:f3acf7dbc2dd15ef1bf5dc5c8fbb3538e24f4d54f1e898f2b6b08ad438e9013b

Observation c367eb61-daef-4882-96a4-71cc92cb0c7f · outbound

This paper cites Contrastive boundary learning for point cloud segmentation.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Contrastive boundary learning for point cloud segmentation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.308893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:22.389608Z digest=sha256:184a0c723693095a4d7c4965cb5f602d689cb856470820de1e4b92043f40b3e4

Observation 3271b91b-af12-4545-bb64-f5f2e23434bb · outbound

This paper cites Small sample hyper- spectral image classification based on the random patches network and recursive filtering.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Small sample hyper- spectral image classification based on the random patches network and recursive filtering

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.150417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:22.462428Z digest=sha256:68ae27c8e9d01818ff4eab2bedf650d7c0d0e6c92b5a6124942b74c41f1ea4f6

Observation f1b25de1-73cf-4f19-b089-6357805ff965 · outbound

This paper cites Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:29.959279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:22.537912Z digest=sha256:d5bf0a65c445156fe0e8cb785e188ba1222e8e7a2d58119722b3b60672caa133

Observation 327fdeff-ef5c-489b-9c7f-53dee17d507c · outbound

This paper cites The inaturalist species classification and detection dataset.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning The inaturalist species classification and detection dataset

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:29.779706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:22.632705Z digest=sha256:974f30b887b26a87332a091d97abf72b6a3c0ae53c57309d9eeb129d31715956

Observation b6988d14-15b5-46a2-98e9-f821a035c5a3 · outbound

This paper cites Attention is all you need.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Attention is all you need

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:29.627959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:22.716235Z digest=sha256:65d1ed2d8db4067edcf9d2200c35cce50dd960e433ed678e30739cd30b83e3d0

Observation 11b94fd8-c192-4958-8fbc-4900b5be83e0 · outbound

This paper cites Internimage: Exploring large-scale vi- sion foundation models with deformable convolutions.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Internimage: Exploring large-scale vi- sion foundation models with deformable convolutions

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:29.467629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:22.775789Z digest=sha256:9629173e7462eb07510ddd67896981d155ed79618f98fa9f889c2e2201ef9bd9

Observation 90af2435-5042-4201-ad4c-2e1abd6c5506 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:22.840001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:22.840001Z digest=sha256:e7f5ebfee2a6614f4735b2de6e1582f1376373e8de2dac06e98f12b39ec25e73

Observation f9b85ae5-edfd-4db3-9b20-6bd5903666be · outbound

This paper cites Masked feature pre- diction for self-supervised visual pre-training.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Masked feature pre- diction for self-supervised visual pre-training

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:29.260075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:22.894679Z digest=sha256:ce4ecfe27cf169d61a9bee0ba5948c74ca99c816e38cc74a2f40fa11614c1718

Observation 349ba8bd-f659-4ebf-8a76-276997f08069 · outbound

This paper cites Syn- cretic modality collaborative learning for visible infrared person re-identification.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Syn- cretic modality collaborative learning for visible infrared person re-identification

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:29.084751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:22.941403Z digest=sha256:3ac22ab1dc77fbf3f32bc7182215d6ce9bbc2df2ad3dae7841bcd14df6f46bf4

Observation 0d42cd48-9c56-4dc2-846e-3514925e8c9a · outbound

This paper cites Controllable Abstractive Dialogue Summarization with Sketch Supervision.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Controllable Abstractive Dialogue Summarization with Sketch Supervision

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:25.321727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:22.994077Z digest=sha256:b88c37800c8b1825e3f28b8dc49d1e92acb26f0483b56fa002156026950dba6d

Observation cecf0c2d-d737-4047-a8ac-88955a12dafb · outbound

This paper cites Tinyvit: Fast pretraining distillation for small vision transformers.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Tinyvit: Fast pretraining distillation for small vision transformers

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.942665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:23.044939Z digest=sha256:9fbcf2697168b0f68239d0278a35e7612f3d9959d9418500b7944d757ee3a8e3

Observation 825288bf-78c8-4094-8cf0-1eaa9b17c048 · outbound

This paper cites Point transformer v2: Grouped vector at- tention and partition-based pooling.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Point transformer v2: Grouped vector at- tention and partition-based pooling

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.793382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:23.101262Z digest=sha256:d0844d1de79446a717669f748d206e9ade7987786527f413aac4019dd4d5eb12

Observation 4c15fc0b-37a7-42c0-a0ca-8f888624e3ab · outbound

This paper cites 3d shapenets: A deep representation for volumetric shapes.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning 3d shapenets: A deep representation for volumetric shapes

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.654139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:23.180431Z digest=sha256:ac40e676d98216e0baed9181d1013c15300adc7eba3aa96aa31619bc39cd3ef9

Observation d4f48735-2087-4877-8af9-04073b46d0bd · outbound

This paper cites Audiovisual SlowFast Networks for Video Recognition.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Audiovisual SlowFast Networks for Video Recognition

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:23.240129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:23.240129Z digest=sha256:c71784ff8b6c398d31564e7ae08a6ed7c390f1f39d71fcf6c8d8b6e56321d32b

Observation bfb374f0-7601-4427-96df-5ee5faf71b0e · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Msr-vtt: A large video description dataset for bridging video and language

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.456028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:23.286212Z digest=sha256:6aef85552b312ddc32536d0b771f353c77ee72db436998b72b8200e5162ba123

Observation e0c2cd51-f150-447c-a73f-a643a7c81994 · outbound

This paper cites Multimodal Learning with Transformers: A Survey.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Multimodal Learning with Transformers: A Survey

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:25.148397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:23.399732Z digest=sha256:8c2e83e33751f0f541c66834852058fd6b211999bb792f457ffe925b56a243e2

Observation cfa2cd38-10c1-42a4-8f1e-7a24a37a3490 · outbound

This paper cites Multi-modal masked pre-training for monocular panoramic depth completion.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Multi-modal masked pre-training for monocular panoramic depth completion

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.330526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:23.499703Z digest=sha256:3013c284a9c99a514c2d7074edaa5a9c8883e0a848b6c05a666d9690e15cc2dd

Observation f955fc72-ae65-4e44-bf53-cb3a6af81b53 · outbound

This paper cites Swin3D: A Pretrained Transformer Backbone for 3D Indoor Scene Understanding.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Swin3D: A Pretrained Transformer Backbone for 3D Indoor Scene Understanding

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:23.580652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:23.580652Z digest=sha256:344541f19cbcfb18b8204fee1a04cc7f4632b6367b6ab3cff3658c04d066996c

Observation b008ece1-7827-4814-802a-93820de787d8 · outbound

This paper cites XLNet: Generalized Autoregressive Pretraining for Language Understanding.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning XLNet: Generalized Autoregressive Pretraining for Language Understanding

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:23.682264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:23.682264Z digest=sha256:76b910ef5977176d04e0cbee2299c56c16f9d1f12c0b97efc087557704798376

Observation 25e42e8f-0a25-4e26-a6fa-ea9b67d18dbd · outbound

This paper cites Deep Learning for Person Re-identification: A Survey and Outlook.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Deep Learning for Person Re-identification: A Survey and Outlook

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:24.969127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:23.740334Z digest=sha256:453562d79dfa648ea61758698084aa6011edc57ae5c66b78bc3a905a192209d9

Observation 240ad846-4d80-4e93-a896-f29ed0221efb · outbound

This paper cites Do transformers really perform badly for graph representation? In Thirty-Fifth Conference on Neural Information Process- ing Systems, 2021.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Do transformers really perform badly for graph representation? In Thirty-Fifth Conference on Neural Information Process- ing Systems, 2021

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.189988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:23.790277Z digest=sha256:c41e105465b37fb6e023ed7f5926f5ffd283d6f03a311cd5e34928d9cc8c8b92

Observation 7aceffcd-fbf6-4549-99b2-da442e37cdde · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:23.861350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:23.861350Z digest=sha256:6ffb8ae0143444f7b60b5c2292b4bdbaf116024f961ea6a4d2e8182c9394472e

Observation e1e514a6-2fda-4a17-8cac-dda7c93bbe04 · outbound

This paper cites Metaformer is actually what you need for vision.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Metaformer is actually what you need for vision

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.047216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:23.921175Z digest=sha256:ded27a90e763f493d3ee425e3b602ebedd3e7d20b681031cbbc7b87efbed072d

Observation b4baed03-004d-4915-9ad5-4c3216d2b102 · outbound

This paper cites Point-bert: Pre-training 3d point cloud transformers with masked point modeling.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Point-bert: Pre-training 3d point cloud transformers with masked point modeling

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:27.928246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:23.991722Z digest=sha256:fc06ca31d8bde8648a1e62a84779a097752f44ef603d98fc9dd1382333cdb58a

Observation 4ffff764-3d5f-4d5f-8564-951c54b457e6 · outbound

This paper cites Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:24.066566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:24.066566Z digest=sha256:fb7ee1998fc5c97005b342367f7aa756c8a131c9f21cc44752a0889bc480c5db

Observation b2ebe9a9-9854-4749-93c2-f4b57288dbc8 · outbound

This paper cites Point- cutmix: Regularization strategy for point cloud classifica- tion.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Point- cutmix: Regularization strategy for point cloud classifica- tion

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:27.750612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:24.114385Z digest=sha256:63b79cce2e1c4f9d917f6a0ed0689f6d375e1e88eea3749a7fc4fc7bc6d59653

Observation 521fa700-e1e6-4b00-8163-3c2a9f01e95c · outbound

This paper cites An overview of multi-task learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning An overview of multi-task learning

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:27.594862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:24.164336Z digest=sha256:7612b8ae01a1981d6665307564a656039e775981a9f7c80c270a26e50981f608

Observation 3953969f-63e7-4d49-91f8-48982b077969 · outbound

This paper cites Modality synergy complement learning with cascaded aggregation for visible-infrared person re- identification.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Modality synergy complement learning with cascaded aggregation for visible-infrared person re- identification

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:27.512561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:24.215744Z digest=sha256:6b5cee983f1a7f1e8569c623dba7aef615f5024d539f70986030a22f14e5ebb5

Observation af7aa99d-b892-47ac-a506-e0154c640fe8 · outbound

This paper cites Meta-Transformer: A Unified Framework for Multimodal Learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:24.267156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:24.267156Z digest=sha256:ef305178a65d440a97f6ff4533f7f49bad81ad3fbf3cc5a769f9219cc5b58143

Observation a316fe08-b859-461a-853f-a88b28f6f36c · outbound

This paper cites Places: A 10 million image database for scene recognition.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Places: A 10 million image database for scene recognition

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:27.367186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:50:24.303188Z digest=sha256:3581bd742d6ca7e5a0fe314e5e17d68532dce4727d6b6f63f70c4a73890259db

Pith citing papers

No inbound Pith citation observations are available.