Pith. sign in

Paper Citation Record · LEDGER

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency

As of 8 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2512.01008.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.01008 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T19:22:04.169020Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee625a1d-d930-4b74-9f6f-9b4ed33aef11 · outbound

This paper cites ReferIt3D: Neural Listen- ers for Fine-Grained 3D Object Identification in Real-World Scenes.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency ReferIt3D: Neural Listen- ers for Fine-Grained 3D Object Identification in Real-World Scenes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.068354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.068354Z digest=sha256:968a67df68929e4a61f2d7a8c47a669246ceab69145f8e011d0669b690e51ae5

Observation 59d774a7-62da-4978-9f45-16185531320c · outbound

This paper cites Qwen2.5-VL Technical Report, 2025.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Qwen2.5-VL Technical Report, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.073150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.073150Z digest=sha256:258cecda8190d49a8be24b12d8d3ceade21e6b19f83331adef980d8a05738be2

Observation e9d708dc-b5f9-4f7c-867a-d211eaae0c0b · outbound

This paper cites From thousands to billions: 3d visual language ground- ing via render-supervised distillation from 2d VLMs.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency From thousands to billions: 3d visual language ground- ing via render-supervised distillation from 2d VLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.076635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.076635Z digest=sha256:b1c1e77c904d700911aa8821c76cc7729f9fb3c4b6f1ba8e67d696345c47ca82

Observation 1f688f06-5710-4471-828a-053e172b4c09 · outbound

This paper cites SAM 3: Segment Anything with Concepts, 2025.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency SAM 3: Segment Anything with Concepts, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.080543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.080543Z digest=sha256:a57ba372165a607dae891c55af0d538ec487ce870ea43bdb87db46fa8daa4149

Observation 6ed007e5-8c81-49ac-adc2-f42b51050920 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.084353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.084353Z digest=sha256:be6c54734a1d554f257a81cb83f33970a64dd945073651b9fd81aff7efe01921

Observation 9bb644bd-eed4-4f30-a43a-53a1277142c0 · outbound

This paper cites Sdfusion: Multimodal 3d shape completion, reconstruction, and generation.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Sdfusion: Multimodal 3d shape completion, reconstruction, and generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.088236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.088236Z digest=sha256:18d43aa3a17477cff50144791379e5b67df2bce64a8a670c8532efeadeb6bfb1

Observation cc2bb181-470e-4cfa-8d68-c55480adc5a8 · outbound

This paper cites Lam3d: Large image-point clouds align- ment model for 3d reconstruction from single image.Ad- vances in Neural Information Processing Systems, 37:4454– 4480, 2024.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Lam3d: Large image-point clouds align- ment model for 3d reconstruction from single image.Ad- vances in Neural Information Processing Systems, 37:4454– 4480, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.092329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.092329Z digest=sha256:17b0266f4008f0f138882a465a45349d608871e00554be5a9fb7f5c7f515909d

Observation f447d6a5-bab7-42b5-b350-ff14773c9387 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.095972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.095972Z digest=sha256:e9f28c14d35aad565bca9ccc5d3e43566822a15097a65de64c3eea9add238daa

Observation 99c161c3-e72b-4eaf-a6b1-ce0fddbfbc87 · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.099573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.099573Z digest=sha256:41e14025f5bd60b509e9abcf87864d9dd6f5c5760da4c58b0e16e0449ee4280c

Observation 5ec97523-7de9-4617-9329-5c491c15786a · outbound

This paper cites SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.103070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.103070Z digest=sha256:b73de380dbc8f60007cf608465e2ef20dc4554d10fa766dbb5b8ade8c1c8c610

Observation d2565d1d-0aab-4bf0-a4b1-d374143ef46d · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.ACM Trans.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency 3d gaussian splatting for real-time radiance field rendering.ACM Trans

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.107066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.107066Z digest=sha256:8c7329719b40757bd1d2484792ed25f0b7960f86f035dd26a866d7815b2af08f

Observation 143fcfd2-9e46-45a1-92cc-8e347d6e4ef9 · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.111126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.111126Z digest=sha256:56a47bf21ae3ba6eda8bfb1032838675086018bdf85ba5f3e7b814a0b4c335d4

Observation 0c6be40f-b30c-41f9-9bb0-a6a6d6a64354 · outbound

This paper cites LISA: Reasoning Segmen- tation via Large Language Model, 2024.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency LISA: Reasoning Segmen- tation via Large Language Model, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.114579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.114579Z digest=sha256:680bc3f0c036d23703b6fef4fea820f867b95e2171bb75873916b69b872e5832

Observation 99dca848-b166-4097-a0a1-c8e8aa0820c1 · outbound

This paper cites Part123: part-aware 3d reconstruction from a single-view image.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Part123: part-aware 3d reconstruction from a single-view image

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.117895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.117895Z digest=sha256:c601a863299c3c443f1574893917fad20fd1f79b1551aa6463bc360f82a6bec4

Observation 88db3e1a-84d6-4151-a57e-b0d6ae85998b · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.121349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.121349Z digest=sha256:3e3ab3dd8d29f9856f9c21c206e6416abfb34b8539ef452fe36cd35de808a75d

Observation fd11c686-cc90-433f-a5e6-a0a8125de3d9 · outbound

This paper cites Decoupled Weight Decay Regularization.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Decoupled Weight Decay Regularization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.125440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.125440Z digest=sha256:c817b4ffe376ce7f924105c3e51d0148a90e47671a429d5d0da2add72897ce27

Observation 98f6dc14-9a7c-40ce-88c4-044f2359a9cf · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Learning transferable visual models from natural language supervi- sion

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.129035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.129035Z digest=sha256:f7adefcb597fefbf5918278f8ace33952920ff663674794f76748d2d72139433

Observation 59a19e38-e36d-4cc6-9bbe-5ded861c5608 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos,.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency SAM 2: Segment Anything in Images and Videos,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.132279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.132279Z digest=sha256:5a5e00d52c233cd877f99d3df0bdfa21ce6d25d3d561bdd4a7e65b0bda09aed6

Observation 1bcf6dc4-b341-4eeb-9ffe-29b565129a99 · outbound

This paper cites A survey of language-grounded mul- timodal 3d scene understanding.Knowledge-Based Systems, 321:113650, 2025.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency A survey of language-grounded mul- timodal 3d scene understanding.Knowledge-Based Systems, 321:113650, 2025

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.135636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.135636Z digest=sha256:e0c921f85e7c4a6563f9b3e6ddc3507a254467a7e90b33ba355987df961c9e4f

Observation 1f29bdc8-0905-4165-94ac-ee4b37488acc · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.138576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.138576Z digest=sha256:0ae4d6934fa231e71a673ae1f412075365094b94185f18488c8c1a8a1b0af1e7

Observation 2241c7dc-e393-4d51-b749-64b7fcaa0f37 · outbound

This paper cites Anything-3D: Towards Single-view Anything Reconstruction in the Wild.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Anything-3D: Towards Single-view Anything Reconstruction in the Wild

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.142399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.142399Z digest=sha256:1c307065f4f68b1742c881a81eaba9345288774ad89c33f0174749d504160dfa

Observation addd49e9-abbb-4e9a-9fe3-b4f0933fc794 · outbound

This paper cites Point-gnn: Graph neural net- work for 3d object detection in a point cloud.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Point-gnn: Graph neural net- work for 3d object detection in a point cloud

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.145796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.145796Z digest=sha256:d2ff3d35215764652a82964655ca818304d842da7bc58f30c4bbab9cb89ea65e

Observation e89dc7f5-0491-427d-a054-7ffc7845727a · outbound

This paper cites Splat-MOVER: Multi-Stage, Open- V ocabulary Robotic Manipulation via Editable Gaussian Splatting.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Splat-MOVER: Multi-Stage, Open- V ocabulary Robotic Manipulation via Editable Gaussian Splatting

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.148641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.148641Z digest=sha256:2c146cd26088d95b1dda0c6d335ad899b47264330fcd5eddd22ce39aca24fc5a

Observation 021f976c-f0b5-47ba-948b-2c75e928d469 · outbound

This paper cites Natural Language Guided Goals for Robotic Manipulation.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Natural Language Guided Goals for Robotic Manipulation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.151842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.151842Z digest=sha256:478bce84d4a114fb9f1999b0d4066419536591cc551f882870af0f239e79c1d3

Observation e388a48c-70e8-48a6-94dd-43d8215c9c91 · outbound

This paper cites SAM 3D: 3Dfy Anything in Images.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency SAM 3D: 3Dfy Anything in Images

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.155120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.155120Z digest=sha256:5180c3cb92747292bfa9f625dcbe311bc02d26b5c8f81b9aa1b7e45b5915b8e3

Observation 01114c86-0576-4c11-9c4d-97d0a888b6b4 · outbound

This paper cites Sa2V A: Marrying SAM2 with LLaV A for Dense Grounded Understanding of Images and Videos, 2025.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Sa2V A: Marrying SAM2 with LLaV A for Dense Grounded Understanding of Images and Videos, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.159081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.159081Z digest=sha256:9953c7d3d988d9f45da5c686aacfcef8e0d73caf523b898e027b84a5cd852f7b

Observation 4022fd66-06cf-4a6e-9034-bcf8553acea4 · outbound

This paper cites Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.162258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.162258Z digest=sha256:0ab61e6ae48daa2f6269bfa4b9fc13ed6f7e7b86b334c438dde3f8095a8a101b

Observation 91e3baa7-afcb-44f5-84ed-f9500e5b0f16 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.165316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.165316Z digest=sha256:750079ba5a306d01ae5619a1b7d9fa2eee004e9b9ea067e46b34fe7c08d76bcf

Observation 6d03deb5-8ba9-48af-8377-2da9c989cb31 · outbound

This paper cites EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing,.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.169020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.169020Z digest=sha256:eaf9269dc9bd49419172de2edd05200b0efc32224288af5fa2c9826a116af26d

Pith citing papers

No inbound Pith citation observations are available.