{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:AILP7NVM4B333VBOIBXD44JVH5","short_pith_number":"pith:AILP7NVM","schema_version":"1.0","canonical_sha256":"0216ffb6ace077bdd42e406e3e71353f5d0297a4d67ec868a91f30fe20b67c11","source":{"kind":"arxiv","id":"2504.13180","version":3},"attestation_state":"computed","paper":{"title":"PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.LG"],"primary_cat":"cs.CV","authors_text":"Andrea Madotto, Andrew Westbury, Babak Damavandi, Christoph Feichtenhofer, Daniel Bolya, Effrosyni Mavroudi, Hanoona Rasheed, Huiyu Wang, Jang Hyun Cho, Kristen Grauman, Lorenzo Torresani, Miguel Martin, Muhammad Maaz, Nikhila Ravi, Peize Sun, Philipp Kr\\\"ahenb\\\"uhl, Piotr Doll\\'ar, Po-Yao Huang, Salman Khan, Shane Moon, Shashank Jain, Shuming Hu, Suyog Jain, Tammy Stark, Tengyu Ma, Triantafyllos Afouras, Tushar Nagarajan, Vivian Lee, Yale Song","submitted_at":"2025-04-17T17:59:56Z","abstract_excerpt":"Vision-language models are integral to computer vision research, yet many high-performing models remain closed-source, obscuring their data, design and training recipe. The research community has responded by using distillation from black-box models to label training data, achieving strong benchmark results, at the cost of measurable scientific progress. However, without knowing the details of the teacher model and its data sources, scientific progress remains difficult to measure. In this paper, we study building a Perception Language Model (PLM) in a fully open and reproducible framework for"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2504.13180","kind":"arxiv","version":3},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CV","submitted_at":"2025-04-17T17:59:56Z","cross_cats_sorted":["cs.AI","cs.LG"],"title_canon_sha256":"e111b2ce5ecee17b406fb6cb259fdd8bbedd52338f3a6760ee40bf5acf7efec2","abstract_canon_sha256":"378acb7191d4e8c28c46947291d2ba9f1a2636d7a22ee53017441c0140ab6ace"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:42:28.733393Z","signature_b64":"4fUdnJNkeZrDHF1V8nvVdbIkELkXw6dTs3NB2c7/qlxJvW/sGaDTr5DRnJMERUia8MhgUDQ1jsBZCGGLld+5AA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"0216ffb6ace077bdd42e406e3e71353f5d0297a4d67ec868a91f30fe20b67c11","last_reissued_at":"2026-07-05T11:42:28.732833Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:42:28.732833Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.LG"],"primary_cat":"cs.CV","authors_text":"Andrea Madotto, Andrew Westbury, Babak Damavandi, Christoph Feichtenhofer, Daniel Bolya, Effrosyni Mavroudi, Hanoona Rasheed, Huiyu Wang, Jang Hyun Cho, Kristen Grauman, Lorenzo Torresani, Miguel Martin, Muhammad Maaz, Nikhila Ravi, Peize Sun, Philipp Kr\\\"ahenb\\\"uhl, Piotr Doll\\'ar, Po-Yao Huang, Salman Khan, Shane Moon, Shashank Jain, Shuming Hu, Suyog Jain, Tammy Stark, Tengyu Ma, Triantafyllos Afouras, Tushar Nagarajan, Vivian Lee, Yale Song","submitted_at":"2025-04-17T17:59:56Z","abstract_excerpt":"Vision-language models are integral to computer vision research, yet many high-performing models remain closed-source, obscuring their data, design and training recipe. The research community has responded by using distillation from black-box models to label training data, achieving strong benchmark results, at the cost of measurable scientific progress. However, without knowing the details of the teacher model and its data sources, scientific progress remains difficult to measure. In this paper, we study building a Perception Language Model (PLM) in a fully open and reproducible framework for"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2504.13180","kind":"arxiv","version":3},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2504.13180/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2504.13180","created_at":"2026-07-05T11:42:28.732890+00:00"},{"alias_kind":"arxiv_version","alias_value":"2504.13180v3","created_at":"2026-07-05T11:42:28.732890+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2504.13180","created_at":"2026-07-05T11:42:28.732890+00:00"},{"alias_kind":"pith_short_12","alias_value":"AILP7NVM4B33","created_at":"2026-07-05T11:42:28.732890+00:00"},{"alias_kind":"pith_short_16","alias_value":"AILP7NVM4B333VBO","created_at":"2026-07-05T11:42:28.732890+00:00"},{"alias_kind":"pith_short_8","alias_value":"AILP7NVM","created_at":"2026-07-05T11:42:28.732890+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":24,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.25842","citing_title":"Graph it first! Enabling Reasoning on Long-form Egocentric Videos through Scene Graphs","ref_index":7,"is_internal_anchor":false},{"citing_arxiv_id":"2606.25634","citing_title":"SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2606.24888","citing_title":"DiffusionBench: On Holistic Evaluation of Diffusion Transformers","ref_index":82,"is_internal_anchor":false},{"citing_arxiv_id":"2606.19100","citing_title":"AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model","ref_index":6,"is_internal_anchor":false},{"citing_arxiv_id":"2606.19100","citing_title":"AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model","ref_index":6,"is_internal_anchor":false},{"citing_arxiv_id":"2606.28551","citing_title":"DataComp-VLM: Improved Open Datasets for Vision-Language Models","ref_index":44,"is_internal_anchor":false},{"citing_arxiv_id":"2607.01117","citing_title":"MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models","ref_index":6,"is_internal_anchor":false},{"citing_arxiv_id":"2606.00390","citing_title":"Zamba2-VL Technical Report","ref_index":48,"is_internal_anchor":false},{"citing_arxiv_id":"2606.28551","citing_title":"DataComp-VLM: Improved Open Datasets for Vision-Language Models","ref_index":44,"is_internal_anchor":false},{"citing_arxiv_id":"2606.19100","citing_title":"AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model","ref_index":6,"is_internal_anchor":false},{"citing_arxiv_id":"2606.30288","citing_title":"VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2606.02569","citing_title":"AdaCodec: A Predictive Visual Code for Video MLLMs","ref_index":72,"is_internal_anchor":false},{"citing_arxiv_id":"2605.21625","citing_title":"Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly","ref_index":6,"is_internal_anchor":false},{"citing_arxiv_id":"2605.21988","citing_title":"Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning","ref_index":5,"is_internal_anchor":false},{"citing_arxiv_id":"2511.16719","citing_title":"SAM 3: Segment Anything with Concepts","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2601.10611","citing_title":"Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2605.13080","citing_title":"Learning to See What You Need: Gaze Attention for Multimodal Large Language Models","ref_index":155,"is_internal_anchor":false},{"citing_arxiv_id":"2504.13181","citing_title":"Perception Encoder: The best visual embeddings are not at the output of the network","ref_index":21,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08560","citing_title":"ZAYA1-VL-8B Technical Report","ref_index":64,"is_internal_anchor":false},{"citing_arxiv_id":"2604.24317","citing_title":"Don't Pause! Every prediction matters in a streaming video","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2604.21718","citing_title":"Building a Precise Video Language with Human-AI Oversight","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2604.12896","citing_title":"Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2604.08762","citing_title":"InstrAct: Towards Action-Centric Understanding in Instructional Videos","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2506.09985","citing_title":"V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning","ref_index":15,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/AILP7NVM4B333VBOIBXD44JVH5","json":"https://pith.science/pith/AILP7NVM4B333VBOIBXD44JVH5.json","graph_json":"https://pith.science/api/pith-number/AILP7NVM4B333VBOIBXD44JVH5/graph.json","events_json":"https://pith.science/api/pith-number/AILP7NVM4B333VBOIBXD44JVH5/events.json","paper":"https://pith.science/paper/AILP7NVM"},"agent_actions":{"view_html":"https://pith.science/pith/AILP7NVM4B333VBOIBXD44JVH5","download_json":"https://pith.science/pith/AILP7NVM4B333VBOIBXD44JVH5.json","view_paper":"https://pith.science/paper/AILP7NVM","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2504.13180&json=true","fetch_graph":"https://pith.science/api/pith-number/AILP7NVM4B333VBOIBXD44JVH5/graph.json","fetch_events":"https://pith.science/api/pith-number/AILP7NVM4B333VBOIBXD44JVH5/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/AILP7NVM4B333VBOIBXD44JVH5/action/timestamp_anchor","attest_storage":"https://pith.science/pith/AILP7NVM4B333VBOIBXD44JVH5/action/storage_attestation","attest_author":"https://pith.science/pith/AILP7NVM4B333VBOIBXD44JVH5/action/author_attestation","sign_citation":"https://pith.science/pith/AILP7NVM4B333VBOIBXD44JVH5/action/citation_signature","submit_replication":"https://pith.science/pith/AILP7NVM4B333VBOIBXD44JVH5/action/replication_record"}},"created_at":"2026-07-05T11:42:28.732890+00:00","updated_at":"2026-07-05T11:42:28.732890+00:00"}