Pith. sign in

Paper Citation Record · LEDGER

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025

As of 14 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 4 inbound Pith citation observations for arXiv:2506.05856.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05856 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:18:59.503626Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T20:34:35.048469Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T04:08:59.885995Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f443e88-6f5d-4d08-ad89-862e47251202 · outbound

This paper cites Temporal-aware Hierarchical Mask Classification for Video Semantic Segmentation.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Temporal-aware Hierarchical Mask Classification for Video Semantic Segmentation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:56.249004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:56.249004Z digest=sha256:1d1a0b5e05a1f8545ac964ccf9582ebe19c5d55850b9ceaefa6412d828eb3b28

Observation 725b2662-33f0-4a09-a40b-44a60a7ca21b · outbound

This paper cites Cafuser: Condition-aware multimodal fusion for robust semantic perception of driving scenes.IEEE Robotics and Automation Letters, 2025.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Cafuser: Condition-aware multimodal fusion for robust semantic perception of driving scenes.IEEE Robotics and Automation Letters, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:04.443359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:18:56.403870Z digest=sha256:80df3b19160b40fb00c835c1685ca874db6c5fcb11324b1e26c9f43aae7d4391

Observation 005689ea-7fa6-4dd9-a440-cf2077ddbf35 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Masked-attention mask transformer for universal image segmentation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:56.608701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:56.608701Z digest=sha256:3a0dc2fb873b3d0c4260be2bf9c2206853bd081eb303254ac2c14aa28cf37773

Observation 9a9bd9f0-db89-474f-9dbb-3059afafdf53 · outbound

This paper cites ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:56.887838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:56.887838Z digest=sha256:08463bea982b991ff89eb511983fc832dbce65ba02789a9592395962f718b396

Observation e9ede233-0518-4a9b-adb7-71326dc6b4df · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:04.082575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:18:57.082792Z digest=sha256:414f3bbe49f8841fa5b658a1c40a4966479a737cf6a8097a3d9318ae81399a05

Observation dcc6bcdd-1821-4e1d-bdbc-b6f208832efe · outbound

This paper cites X-prompt: Multi-modal visual prompt for video object segmentation.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 X-prompt: Multi-modal visual prompt for video object segmentation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:03.750415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:18:57.211327Z digest=sha256:2ae07efc9c53c9139fa03d7b6e01e09c8dbb17b9a33908d415270d16e0f83383

Observation 2248e8af-93c6-461e-8043-9a5845b92469 · outbound

This paper cites Mask r-cnn.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Mask r-cnn

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:57.352999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:57.352999Z digest=sha256:852986d31310f933313d7dbb30735b4483829702fc4252154f22b10cefb9496e

Observation 8c715af6-5be4-49c1-9adf-03527e671253 · outbound

This paper cites Segment any- thing.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Segment any- thing

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:03.388111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:18:57.574611Z digest=sha256:db9379e3a59f0947af4799c09d42bacccf27504167e1d52c5a01644c8c74e33f

Observation 1eec0d54-d72c-48c0-9cef-dd27f930a10e · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Lisa: Reasoning segmentation via large language model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:57.715692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:57.715692Z digest=sha256:80aae3480d9c36bdd2c031ebbe9180b05a7fa51610cc5c5a22d9f53ee3b3123c

Observation dffa37fb-3d66-4632-aa6e-62a3288a89c9 · outbound

This paper cites Onevos: unifying video object segmentation with all-in-one transformer framework.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Onevos: unifying video object segmentation with all-in-one transformer framework

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:03.076345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:18:57.915378Z digest=sha256:a31cca0598f12a2dbe56d96e6769e5229d61b0816c5138e0c6ddcc764f10befe

Observation bb2f6463-df76-41e6-a09a-21e7517ed54e · outbound

This paper cites Omg-seg: Is one model good enough for all segmentation? InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 27948–27959, 2024.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Omg-seg: Is one model good enough for all segmentation? InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 27948–27959, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:02.653871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:18:58.062171Z digest=sha256:e81aca035c1ad57541c4fccdfddbc5a93168f3587cce7b61fb3b67f899ab6ccb

Observation b17b2a62-90ab-4414-8a7e-aeedd7231407 · outbound

This paper cites Visual instruction tuning, 2023.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Visual instruction tuning, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:02.276457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:18:58.185928Z digest=sha256:d19b4bf6b8727c4a5b917ae3ececd1106ced61f920827bf708e4bb26dfd04c27

Observation 8f8f8368-3aaf-454a-9730-17e3c90ee25b · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Glamm: Pixel grounding large multimodal model

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:01.929591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:18:58.377127Z digest=sha256:54962c3d71ae66e5e18497be344ef81d3465e422ecfd4d02e8a8078ba4aeac1a

Observation c39d921c-cc6a-44fd-9681-6a09baecdfa2 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Pixellm: Pixel reasoning with large multimodal model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:58.523241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:58.523241Z digest=sha256:e14034605596af276a757e94de21b9b1aa71d4da019cf60845578fda04c842bb

Observation dd7159c9-5b50-43aa-8975-22f8c1e711db · outbound

This paper cites Object segmentation by mining cross-modal se- mantics.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Object segmentation by mining cross-modal se- mantics

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:01.625200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:18:58.644474Z digest=sha256:ec989dd2318be651104ebb4b499bf43166f5fdc7a82c2cfe48ba0f8e53e6bbbf

Observation c747ee1b-86f7-4769-86c8-34b97e4cbe48 · outbound

This paper cites Universal instance percep- tion as object discovery and retrieval.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Universal instance percep- tion as object discovery and retrieval

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:01.197110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:18:58.788169Z digest=sha256:76f0696619483018b3757b2debb4f943b83be99d4ba164ad08385c85acaac5b1

Observation 759a649d-20e8-4891-9aef-ac73f5c4a62c · outbound

This paper cites Psalm: Pixelwise segmentation with large multi-modal model.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Psalm: Pixelwise segmentation with large multi-modal model

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:00.878470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:18:58.981985Z digest=sha256:84c80cdb3067d64de2f80cc146b830839006695a07ff3a51502a58714f92db8f

Observation 19814800-e546-41d5-9444-5c9be91d9830 · outbound

This paper cites Learning modality-agnostic representation for semantic segmentation from any modalities.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Learning modality-agnostic representation for semantic segmentation from any modalities

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:00.648495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:18:59.093530Z digest=sha256:d929c937c9a6765f5ea8e2c4302e9f23d9b79e42a858e3402285fad2fee0708e

Observation 4410f5b4-1a5e-42e0-a987-d38127f4db51 · outbound

This paper cites Distilling efficient vision transformers from cnns for seman- tic segmentation.Pattern Recognition, 158:111029, 2025.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Distilling efficient vision transformers from cnns for seman- tic segmentation.Pattern Recognition, 158:111029, 2025

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:00.321486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:18:59.200838Z digest=sha256:19457605e51e549b2e98f50f0aab002587416b59c8dfd7f45d94836825419f1a

Observation 78db28db-7858-4696-b071-c70b71236434 · outbound

This paper cites Camsam2: Segment any- thing accurately in camouflaged videos.arXiv preprint arXiv:2503.19730, 2025.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Camsam2: Segment any- thing accurately in camouflaged videos.arXiv preprint arXiv:2503.19730, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:59.275393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:59.275393Z digest=sha256:3846a07fa210dccd9963b63c110fc9e225b5eaa92e862de32b9eee49ffb91944

Observation c941f7fe-cc7c-4db7-95a0-8ab2d87d253f · outbound

This paper cites Segment everything everywhere all at once.Advances in Neural Information Processing Systems, 36, 2024.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Segment everything everywhere all at once.Advances in Neural Information Processing Systems, 36, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:00.032209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T10:18:59.503626Z digest=sha256:3924fdca3223d3892bae1433a6b64abf07023952168d6e5a2755e2abd8ec7571

Pith citing papers

Observation 45f5f125-8168-42f0-a724-759311ab1d31 · inbound

V$^{2}$-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence cites this paper.

V$^{2}$-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:08:59.888437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-17T04:08:32.413555Z digest=sha256:4a3bac6dc330c81e0a73b7727c835fbbe68109ed7a78a4d89f825c40fba8c187

Observation 378ac56e-6c33-4603-adf4-bd3f9b767014 · inbound

EgoSound: Benchmarking Sound Understanding in Egocentric Videos cites this paper.

EgoSound: Benchmarking Sound Understanding in Egocentric Videos Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:46:42.453337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T21:44:47.636912Z digest=sha256:f66bff24e821b22446a3f7e443af76c92512cdedc9f63e9978744e71d22839f5

Observation e1f06317-7c39-4021-8e94-02a06f54a6ca · inbound

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation cites this paper.

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:25:26.135752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T13:23:35.738229Z digest=sha256:89d372b63505293a8f3a22ace6ca243068e5f71b850df03789ab1b3e38f98b7f

Observation bead3897-5736-46bc-bd36-2ff7e272972b · inbound

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation cites this paper.

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T20:34:35.048469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:34:35.048469Z digest=sha256:81d10730a7f25c8affaf6f99bd29c7a861ff6a92e96729cb15c8fee04865e5fb