Pith. sign in

Paper Citation Record · LEDGER

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark

As of 21 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2411.13056.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13056 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:57:23.947375Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact6
  • verified fuzzy19
  • unresolved20
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 78de6edd-6ba5-4416-a0f8-b5a9de3591b8 · outbound

This paper cites write newline.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.782447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.782447Z digest=sha256:966f86da6980ce05a6ef76d403189c05e46098bb7601ac2509391a2f26d5e752

Observation c2614720-692f-476a-8d3f-41d5899fd900 · outbound

This paper cites Arteta, V.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Arteta, V

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:25.052902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.787353Z digest=sha256:d8636fe9bea33f819b5c88250a735970753568aa6c9747cfa06de240555b4393

Observation 9336d576-683b-4a9e-a2a2-e54644fc83db · outbound

This paper cites A spatio-temporal attentive network for video-based crowd counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark A spatio-temporal attentive network for video-based crowd counting

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.791186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.791186Z digest=sha256:8543b5a91ffa396b78df55eda03aabe4577f3720cb1073cdab75832f13da7cad

Observation 52a8ef27-0ddc-4792-b827-281bb1618396 · outbound

This paper cites MultiMAE : Multi-modal multi-task masked autoencoders.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark MultiMAE : Multi-modal multi-task masked autoencoders

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:25.042712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.795090Z digest=sha256:58a9168afa672d88e8cea5d27106fd8d8e8dd1808fd555bce6c3982164489879

Observation 78f9d0d9-06a9-4221-9460-8fd303c44804 · outbound

This paper cites an unresolved cited work.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:57:25.032456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.798655Z digest=sha256:d39c425223478bdab1546a5934a864d5b15d2f893e254e423fce81384bde98a0

Observation b1d05056-85ce-473d-ae0e-94583a8d7f67 · outbound

This paper cites BE it: BERT pre-training of image transformers.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark BE it: BERT pre-training of image transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.802121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.802121Z digest=sha256:6f5e41f275e134ce2c290a32b7045395dee9fdfc69cda6f18052aa58b186500e

Observation ab34073d-7ad5-4ec8-85b8-cc4932a56fcc · outbound

This paper cites Generative pretraining from pixels.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Generative pretraining from pixels

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:25.016749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.805406Z digest=sha256:590bfc52e5d5b85bb6bb32a993bf48d5ca3c4f4124c35af33813085f085d9fb0

Observation 7559e6d6-2689-4933-babd-80b49dffba1c · outbound

This paper cites BERT : Pre-training of deep bidirectional transformers for language understanding.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark BERT : Pre-training of deep bidirectional transformers for language understanding

Reference 8

Resolution
malformed identifier
no resolver link, observed 2026-08-12T16:57:23.809064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.809064Z digest=sha256:51b6492e8a428e2c7a6f37a5393edc2f83a78aea66c570aa22637607ecff8906

Observation 440ed128-4068-4d88-aa34-4a7999eff820 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.812413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.812413Z digest=sha256:ab081eacea689ce72b8450be126cfa0dad3c6525f73dd7008d5c1d4428f678e1

Observation 9685b215-d4aa-48d5-a5d6-15a924bd2266 · outbound

This paper cites Redesigning multi-scale neural network for crowd counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Redesigning multi-scale neural network for crowd counting

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:25.007105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.815966Z digest=sha256:273670b2b73e62220cb551c9774ac21759764f08c5da52495e737328a3c6e9e5

Observation 626e37d2-a91a-4636-9903-0addb52abe50 · outbound

This paper cites Locality-constrained Spatial Transformer Network for Video Crowd Counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Locality-constrained Spatial Transformer Network for Video Crowd Counting

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.819308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.819308Z digest=sha256:e7c9e4971bd15c49c36ee85fad38b13a4d40a281019b2da87bf69826c5690082

Observation 371c6409-ac7b-4fe8-9441-4d53923a1442 · outbound

This paper cites Multi-level feature fusion based locality-constrained spatial transformer network for video crowd counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Multi-level feature fusion based locality-constrained spatial transformer network for video crowd counting

Reference 12

Resolution
verified exact
doi, observed 2026-08-12T16:57:24.018963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.823305Z digest=sha256:d7b93d78e2d16d31b41c24ba1afe94e746956c53df5933b114d30bd225bd97a6

Observation d4e7bb18-c7e9-4865-be09-3b03414112da · outbound

This paper cites Masked Autoencoders Are Scalable Vision Learners.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Masked Autoencoders Are Scalable Vision Learners

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.826976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.826976Z digest=sha256:567de428a056854d3a1216da1f292004bb3f0e7d00dbbba542b8ac42f9d4608e

Observation dd01dfb4-1de7-4279-a4ee-3027a3802743 · outbound

This paper cites Video-based crowd counting using a multi-scale optical flow pyramid network.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Video-based crowd counting using a multi-scale optical flow pyramid network

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.996512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.830574Z digest=sha256:72a5cd57632b499e8a8c23e60ab02079397271c64b2856ff78e3d186be7459e5

Observation 9db75f98-6b03-47ea-a154-dc0b13e0e245 · outbound

This paper cites Frame-recurrent video crowd counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Frame-recurrent video crowd counting

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.833699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.833699Z digest=sha256:1c2efbe06c636995e2ba2ac79a9166135863099db666ac040402e8fd1460f7fa

Observation d57eee47-f21d-4ffc-81d4-41e1418f6ed8 · outbound

This paper cites Clip-count: Towards text-guided zero-shot object counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Clip-count: Towards text-guided zero-shot object counting

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.837157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.837157Z digest=sha256:e744790ebe7a7881b8d4406b414d1f4c9c4c18651aa4e05f2c795f61292c1627

Observation b07f420c-f39f-4c4b-bbdc-2760dec31ff1 · outbound

This paper cites Vlcounter: Text-aware visual representation for zero-shot object counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Vlcounter: Text-aware visual representation for zero-shot object counting

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.985644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.840443Z digest=sha256:d76dd9fff5d22c2ea93fa0a3582c588b75393a8db66535d10c3f557338d9eeb3

Observation 44abdf7a-f553-4568-9b82-d0e3f0096172 · outbound

This paper cites Video crowd localization with multifocus gaussian neighborhood attention and a large-scale benchmark.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Video crowd localization with multifocus gaussian neighborhood attention and a large-scale benchmark

Reference 18

Resolution
metadata mismatch
raw_fallback, observed 2026-08-12T16:57:24.609996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.843597Z digest=sha256:fbc31619fa6225b259b0a742a5194ccff03529c7e8312c6c8c53d7b4a07565ad

Observation 43b5dbcc-1851-4bbe-b4aa-145e236db61d · outbound

This paper cites Csrnet: Dilated convolutional neural networks for understanding the highly congested scenes.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Csrnet: Dilated convolutional neural networks for understanding the highly congested scenes

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.846966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.846966Z digest=sha256:798c9971f8db6c5c0d99266b91988300f73bf604e16a6988ce657f771545e6cb

Observation 7e7bb7f6-0872-48bc-9c94-cd381b828bad · outbound

This paper cites Transcrowd: weakly-supervised crowd counting with transformers.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Transcrowd: weakly-supervised crowd counting with transformers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.975681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.853888Z digest=sha256:364cad3a1eb97515773c5478bf54f95d439b880d8a66f878276e1aa7e735e05e

Observation f8569ae5-1d0d-48be-a570-e427f6b725be · outbound

This paper cites Boosting crowd counting via multifaceted attention.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Boosting crowd counting via multifaceted attention

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.965341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.857463Z digest=sha256:0c6620f8aab74260e3109d376249d3be43df8f94225219d5ec721b37f429b0b7

Observation cd0653db-3a7d-43e6-9cdd-0a97e685dace · outbound

This paper cites Gramformer: Learning Crowd Counting via Graph-Modulated Transformer.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Gramformer: Learning Crowd Counting via Graph-Modulated Transformer

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:57:24.474037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.860634Z digest=sha256:7b21aa83eff588551889e6d6318567746b9025b33a2a3ae2206a5eb93fe36ee4

Observation d159f02c-9c61-4f5d-b4a5-f0a345eddb72 · outbound

This paper cites Point-query quadtree for crowd counting, localization, and more.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Point-query quadtree for crowd counting, localization, and more

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.955272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.863958Z digest=sha256:6509d95432ca6b068a92cc8abfc786e31b628ab0e93d5a87604e3ae60a21a0b4

Observation 7a0b16c9-afd0-4c71-9896-e1652db2a2ee · outbound

This paper cites Context-aware crowd counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Context-aware crowd counting

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.945253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.867202Z digest=sha256:47d0e09ecc7d6d238fe9baf3a6602aef6d4c2f9ad5235d8ad8c11b70455b3f05

Observation 664e97d6-0f59-46a2-990b-63d94dcb3050 · outbound

This paper cites Estimating people flows to better count them in crowded scenes.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Estimating people flows to better count them in crowded scenes

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.934936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.870428Z digest=sha256:eaa911b0f5b45d08aaab12dbca49da0e8cdb7e8f99c527158a9da6260009e315

Observation aa9eff38-6cd3-4e1a-aec1-b52ec02577b8 · outbound

This paper cites From semi-supervised to transfer counting of crowds.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark From semi-supervised to transfer counting of crowds

Reference 26

Resolution
verified exact
doi, observed 2026-08-12T16:57:24.008980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.873558Z digest=sha256:db088ed75a3e894f530fa9bd1e16d0297fa7149236360d4ff68ed9bcbcb3fdf2

Observation 166a18b5-fc40-4574-9de1-af9f9e8641b0 · outbound

This paper cites Bayesian loss for crowd count estimation with point supervision.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Bayesian loss for crowd count estimation with point supervision

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.924155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.876933Z digest=sha256:9b745ffd3286994dba9cb5fcf75e087f95f6e2b7e466fd63e3522b786d7d4660

Observation ecad169a-ae56-4f8c-bd91-ec3133ddc848 · outbound

This paper cites Phnet: Parasite-host network for video crowd counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Phnet: Parasite-host network for video crowd counting

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.880503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.880503Z digest=sha256:46fbc540d629cb7ed2a6231d64e28ca51c223f287a08bc78d7ca3446f69c8943

Observation e3544976-f0bd-4017-a5b4-3356690211c3 · outbound

This paper cites Rudin, Stanley Osher, and Emad Fatemi.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Rudin, Stanley Osher, and Emad Fatemi

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.883550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.883550Z digest=sha256:4a8afea1b14996baa02f4b7e1de83ba764a34cfb79ac0d25cf01389ee2ebd198

Observation 5bb89ac6-e73a-40be-8d23-01b8fa385da9 · outbound

This paper cites Convolutional lstm network: a machine learning approach for precipitation nowcasting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Convolutional lstm network: a machine learning approach for precipitation nowcasting

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.913583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.886796Z digest=sha256:25b174ec56c121814567c2f45792c016df56d8e60fd83853bf00fbbef60f850f

Observation 91c7fe4f-fb45-4a2d-be96-89aa511175b8 · outbound

This paper cites Crowd counting in the frequency domain.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Crowd counting in the frequency domain

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.903338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.889842Z digest=sha256:92b31660d16689f66e7b4ab3b1cd6bec62591379f6f678d4f46b5e3a3451b6ca

Observation 42c4479e-a51b-402e-9b9b-5cd1d2c1b31a · outbound

This paper cites PWC-Net : CNNs for optical flow using pyramid, warping, and cost volume.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark PWC-Net : CNNs for optical flow using pyramid, warping, and cost volume

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.893026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.893078Z digest=sha256:2ed1d1dcfa963db0c32d9e25e40d73e98e2256b6ddafab7f1e795f2e7d614280

Observation d51332c6-d2ce-48cf-a815-15be8225e664 · outbound

This paper cites CCTrans: Simplifying and Improving Crowd Counting with Transformer.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark CCTrans: Simplifying and Improving Crowd Counting with Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.896271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.896271Z digest=sha256:7e362e39ec0a8a661e82c421f805f6bd110a22d8343dbcefcaa737c99e8fe4fc

Observation 3e5a2c63-1336-4c00-9265-d6ff21e9ba3b · outbound

This paper cites Video MAE : Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Video MAE : Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.882248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.899774Z digest=sha256:c991d5cfbf00a3e2e41714f590899cb7b49435a9bc7b48dbeb5d993153a8989f

Observation 8c05006d-6a9c-4f66-a70d-d29040192035 · outbound

This paper cites Attention is all you need.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Attention is all you need

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.903168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.903168Z digest=sha256:77fd20e88015625b006f15e31e45fd7208ccc6944df87e94b8ea63f044774db9

Observation 0b6566b4-245c-4135-9a4b-6423cc651fa1 · outbound

This paper cites Extracting and composing robust features with denoising autoencoders.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Extracting and composing robust features with denoising autoencoders

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.906726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.906726Z digest=sha256:f1bbd0aa695ec2bbb56e28c17f52f935ddceee665d013391c98325e386f3a70a

Observation c431bf60-4e15-4b29-accb-0a01d9669159 · outbound

This paper cites Bird-count: a multi-modality benchmark and system for bird population counting in the wild.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Bird-count: a multi-modality benchmark and system for bird population counting in the wild

Reference 37

Resolution
verified exact
doi, observed 2026-08-12T16:57:23.998950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.910115Z digest=sha256:60398570decebc0cbec5e3187f6d821610060bcc9405d0ec71e1ddb22e441c20

Observation a2bd0bf0-a4b8-4aa7-a698-fce5016d0455 · outbound

This paper cites Fast video crowd counting with a temporal aware network.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Fast video crowd counting with a temporal aware network

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.866031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.913415Z digest=sha256:6cd9ad7216fac87561ebeec12bd1fa7179b6acf85947e586ad7d5b0735b1c723

Observation c3b0ec67-e2ae-441f-b134-c66cba37238b · outbound

This paper cites Spatial-temporal graph network for video crowd counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Spatial-temporal graph network for video crowd counting

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.916569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.916569Z digest=sha256:2cb63d440f800c118e57fbf71120b27ba73b903188a6abb20ce8404188ed4422

Observation 8c06c77d-2b8a-43bc-a1d8-7d2e5070bc3a · outbound

This paper cites Spatiotemporal modeling for crowd counting in videos.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Spatiotemporal modeling for crowd counting in videos

Reference 40

Resolution
verified exact
doi, observed 2026-08-12T16:57:23.985587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.919832Z digest=sha256:e368d718f480f2d84d9acffee23576203e2d5f554f2de78db0feb3e7bf0f554b

Observation 470da1b4-0df2-4d25-9f9c-e3fa5f45fd4f · outbound

This paper cites Reverse perspective network for perspective-aware object counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Reverse perspective network for perspective-aware object counting

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.855306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.923135Z digest=sha256:6554ac391252fb47d0a8595202fe9640d75dd94c7c0a430a4007481d65944322

Observation 7a3fbda3-32c5-486a-9fdc-48b6ff5a7707 · outbound

This paper cites Single-image crowd counting via multi-column convolutional neural network.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Single-image crowd counting via multi-column convolutional neural network

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.844758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.926246Z digest=sha256:7914c06d1019fecbf89479f71da3a4bb451d109247d05196a19f293453495bd1

Observation c75526b9-4f58-4993-b322-03008a703bc4 · outbound

This paper cites Locality-aware crowd counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Locality-aware crowd counting

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.929315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.929315Z digest=sha256:a60e93e47b8a400256faa000c9542f49fdc4bd2d5a672d4f325ffdf306cb33d4

Observation b7e0ffea-fd1d-4f8a-bf2b-0649ab05a780 · outbound

This paper cites Graph regularized flow attention network for video animal counting from drones.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Graph regularized flow attention network for video animal counting from drones

Reference 44

Resolution
metadata mismatch
raw_fallback, observed 2026-08-12T16:57:24.135086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.932906Z digest=sha256:9388b2c9d51cfdebc042d0f0f07257498217ce2e70bec28ab7e73e8a4c7f9a81

Observation f90697e8-e4f4-4dab-9989-1d5d94f0adc5 · outbound

This paper cites Enhanced 3D convolutional networks for crowd counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Enhanced 3D convolutional networks for crowd counting

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:57:24.039410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.936167Z digest=sha256:b130385176f2773c00e15e22f921d038a15228ce6e4003502f2ac1393031ce1f

Observation 7d9bc860-b41b-481b-bf69-a368c78472d8 · outbound

This paper cites @esa (Ref.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark @esa (Ref

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.939779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.939779Z digest=sha256:9c3c628513ca21f3942b2c5708a29a27ad5935a5df238d53d68fc6de13a728bb

Observation 2c6a71cc-5246-406f-9fea-608bfe245980 · outbound

This paper cites an unresolved cited work.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.943688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.943688Z digest=sha256:8ae0ab91ee809f2f46d2cc4d7bdc1f1e434a47a7fd2c5e79bd8ca663de0ec13f

Observation 0fe7fb19-91b2-479c-842f-0ee17427ea7c · outbound

This paper cites an unresolved cited work.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.947375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.947375Z digest=sha256:6da3aa2d29a2c480e6e0e6058ae74db657f4e1435140bfb31d59371a56435095

Pith citing papers

No inbound Pith citation observations are available.