Pith. sign in

Paper Citation Record · LEDGER

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models

As of 8 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 2 inbound Pith citation observations for arXiv:2506.18421.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18421 v2

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:04.057694Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T22:18:45.189576Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T14:05:46.699261Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4b3a2a05-58ce-4e41-97c9-b4b49af9becd · outbound

This paper cites GPT-4 Technical Report.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:00.709832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:00.709832Z digest=sha256:40974b79427fb77cc41864d49c709ea5f527cd9474948076c1fcf36c664fee76

Observation 338be5f1-e943-4a6f-8e59-060ada6d38a0 · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:01.001340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:01.001340Z digest=sha256:cc1606022a0cbf7051f8ae0464b9b496abe30d6c16f6faa31d4a31f0441a5945

Observation f5a0899b-0996-452e-a24a-0004dc95ce6a · outbound

This paper cites R., et al.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models R., et al

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:05.462141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:20:01.051627Z digest=sha256:390e36d5435237a566ff61627fd0b8b59d7d28d8b0aa50f8a57756b82c7a3d9b

Observation 4744b0a9-e1cc-423d-be1d-4ab347163d09 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:01.103215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:01.103215Z digest=sha256:e7c09954f3ca6afeaa830297cf83e4611bf84588ca1e14421af0ddfd197982f8

Observation 37bf6e34-0550-4b47-9a1c-d55322eeacbc · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models RLHF Workflow: From Reward Modeling to Online RLHF

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:01.350485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:01.350485Z digest=sha256:35dbd874824c29fe4602a3d0ea6aefe88bf7d215c75cac5215ace49f08b2688e

Observation da391928-f427-42e1-b027-8a07c6e7404f · outbound

This paper cites The Llama 3 Herd of Models.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:01.408209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:01.408209Z digest=sha256:a71ab3b2f225052ed27ea7ed62c948f21fc2e5ff56721f44bb88dec8fd77d5e0

Observation 27e5dc92-2636-4092-8a11-e867b2856365 · outbound

This paper cites A Survey on LLM-as-a-Judge.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models A Survey on LLM-as-a-Judge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:01.489710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:01.489710Z digest=sha256:cbfa2f4943884ce3a4a16bdb8b3f26488906f0406aa3e7d5627c5dcff5a5a840

Observation 8735e25e-e6f3-4e3b-9d82-d8a27040ef97 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:01.630884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:01.630884Z digest=sha256:4355dbe729f20856da5b982e69b4d12cfab6d179db036a70b8497e07354ad610

Observation 28ca5b0e-bb76-4375-bb69-a5af0144a6fe · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:01.729656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:01.729656Z digest=sha256:af6afd149b98c37b9e45a1983b7669e33ca3e8f47a1526631050a3ed52307818

Observation d4799446-eba0-46b9-abc3-7b18fee89f36 · outbound

This paper cites Qwen2.5-Coder Technical Report.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Qwen2.5-Coder Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:01.826117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:01.826117Z digest=sha256:87847f5d2a403d77cfbf9295af4ac6a10c1a153aa518bd48a8522e817ad9a403

Observation 80cd3bb1-90be-4ef7-bb44-8c12aa87e8be · outbound

This paper cites FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:02.097569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:02.097569Z digest=sha256:e8a73b16cfa1d05e5e20298fb099c274dfe3ea21208ab23da1b33b498a3bb15c

Observation 8d00a86d-a6ac-4f38-ac47-b5ce33ac0601 · outbound

This paper cites Ait-qa: Question answering dataset over complex tables in the airline industry.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Ait-qa: Question answering dataset over complex tables in the airline industry

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:05.113736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:20:02.229374Z digest=sha256:1fac4d7a5dd8c1a9696e01dede9a069bd2e9aaa72708fbdbefdbe9e6159479d6

Observation 12d33cae-f6e2-4ee8-8fb0-29985d374f4c · outbound

This paper cites TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:02.470277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:02.470277Z digest=sha256:ff6b14d5c6f195e95aa75e5d66583d375f37097235ab7ddb209effe48c3b250c

Observation d897cdce-52af-46b2-ac59-5c62e1a1d68f · outbound

This paper cites TableQAKit: A Comprehensive and Practical Toolkit for Table-based Question Answering.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models TableQAKit: A Comprehensive and Practical Toolkit for Table-based Question Answering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:02.598781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:02.598781Z digest=sha256:b504e37b718df0fc15fbfb8e70c4bf8b0bb86963893b5eedf3958f7f79112e85

Observation fa4b1cff-1354-4f94-a1b6-c2ffd55a2a6d · outbound

This paper cites UHGEval: Benchmarking the Hallucination of Chinese Large Language Models via Unconstrained Generation.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models UHGEval: Benchmarking the Hallucination of Chinese Large Language Models via Unconstrained Generation

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:20:04.474776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:20:02.713992Z digest=sha256:7c9fe4b64e5b9b4a9bd669f750e7a49cdb233918a4f55e98c99e17442025874a

Observation 86b4b131-021c-4617-80dd-67208fb4602b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:03.205192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:03.205192Z digest=sha256:edfd28f28ba3995754d07c4d73727a2236ad5622d106f0dc5a9c2a5a2432b823

Observation bdf2d175-1d2b-4802-868a-b97e70998156 · outbound

This paper cites TableGPT2: A Large Multimodal Model with Tabular Data Integration.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models TableGPT2: A Large Multimodal Model with Tabular Data Integration

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:03.333517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:03.333517Z digest=sha256:126c4fc5982aba9205a9c25cecbfbf68c310db4316fda36f62a776aac90201ba

Observation 6f6c6654-9fed-443f-8b78-24f80ffffc80 · outbound

This paper cites Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:03.461424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:03.461424Z digest=sha256:7467e20936e847f63563c96462ce2b155aa6e041cee339110f7eddfcdcc945c3

Observation 30c8885d-9fdf-40f1-b468-994f82b39700 · outbound

This paper cites MAC-SQL: A Multi-Agent Collaborative Framework for Text-to-SQL.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models MAC-SQL: A Multi-Agent Collaborative Framework for Text-to-SQL

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:03.588865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:03.588865Z digest=sha256:d4301f7f28e4522bfe3c9f6e8710ed7ce19e27933aecb0437875fd02ec39c406

Observation c82b98e5-bc45-4474-9226-1faf9aee2f94 · outbound

This paper cites Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:03.663026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:03.663026Z digest=sha256:a8b8ebeb9d24dd5be1e99647eea7a90f2fc38aad493b7e6e16b34452961fe03d

Observation b8ee7cea-3bc6-4461-85a4-b7c960ad444e · outbound

This paper cites Qwen2 Technical Report.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Qwen2 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:03.756297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:03.756297Z digest=sha256:3186d07cf26a4435a9a680e2fabdbc26cd3b817a185b134cd603f0d0f046a4c4

Observation d44e477d-a053-445e-b0d9-c49b8f7aa476 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Yi: Open Foundation Models by 01.AI

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:03.872378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:03.872378Z digest=sha256:a8db2cb5d8f74a2e5682f820eb7a004cca4a4090c05990d82c1bbd512a97a8aa

Observation 5424cf12-9e9c-4ac5-996c-feb3e76872a5 · outbound

This paper cites Spider: A large-scale human-labeled dataset for complex and cross- domain semantic parsing and text-to-sql task.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Spider: A large-scale human-labeled dataset for complex and cross- domain semantic parsing and text-to-sql task

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:04.802096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:20:03.958511Z digest=sha256:d7abd627283b40323c25b464682179c3c766edff9c2212740b98f53bb53ca6cf

Observation 409a6220-26df-4c16-b46e-06bbdf72dbe6 · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:04.057694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:04.057694Z digest=sha256:38309b0754db1c6bc2646da89ce371d9f98123d07cd053d178c1a66bf019e389

Observation 85074de9-9cc1-4f2f-a4cd-dd4e197a6f2a · outbound

This paper cites Totto: A controlled table- to-text generation dataset.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Totto: A controlled table- to-text generation dataset

Reference 2002

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:04.964808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:20:02.940097Z digest=sha256:da737e9424b8eedb7baa9d079fa7c5ec227917953d7f105f9e902c51e4a8315f

Observation 193dd8cc-9d80-4f51-8211-c7db35f3a691 · outbound

This paper cites MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:02.838439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:02.838439Z digest=sha256:135464d63ac2dec001da3ab66c75f6864ab4612d3f80563424239c348616029c

Observation 111542bc-8bb5-4529-a60e-5f500ef79af2 · outbound

This paper cites TQA-Bench: Evaluating LLMs for Multi-Table Question Answering.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models TQA-Bench: Evaluating LLMs for Multi-Table Question Answering

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:03.071359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:03.071359Z digest=sha256:9df7240b87106a13d977ee190bd897f7ce6332c9646213950e3e40b63f3c6ec8

Observation f7b13eb3-f7c8-4ef6-aa61-1394f2fbae1f · outbound

This paper cites Mistral 7B.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Mistral 7B

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:01.965467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:01.965467Z digest=sha256:3c5af3c18cf5564087b7e30bd35beb24079e220046be46ade245e0df87cca37b

Observation 0f7daa07-7220-4925-a7c1-5faf4bcceee0 · outbound

This paper cites McEval: Massively Multilingual Code Evaluation.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models McEval: Massively Multilingual Code Evaluation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:00.861149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:00.861149Z digest=sha256:a2aa84b0f9821e874ff3e6a4e6da22492fc1fc35cdd25d448ce74fe09b670d57

Observation bf0866df-e83f-417e-ba0d-bdbccb941bb4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:01.190008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:01.190008Z digest=sha256:284f07154464c457e3048050d1a5da14f540fdaecdd8ac60a0c75a4dc60dc9bc

Observation 48d2a58e-7701-44fd-a6bd-4891431660d6 · outbound

This paper cites Table Foundation Models: on knowledge pre-training for tabular learning.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Table Foundation Models: on knowledge pre-training for tabular learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:02.370073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:02.370073Z digest=sha256:3d7c9ce3749c53daf8338abe6340effebad78a414c7cff1ee2b35393ce566a40

Observation 1d4d8fcb-ca5b-4c36-b977-d8630df5f8ee · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:00.771769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:00.771769Z digest=sha256:92c8b97a5a79b95b31e65a486d25b9d2e53eba25d234ccf1347b1361f0096cf8

Observation 84d11dd7-80f8-4fce-929c-67c01f7d8945 · outbound

This paper cites an unresolved cited work.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:20:05.669312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:20:00.940491Z digest=sha256:0c1b39fcea14eaff0aed2b9554e58d2081395132eb42b29670b909bd1d44a88d

Observation c8c641df-fc1f-43b5-ae47-b30e5d7f0afa · outbound

This paper cites Tables as texts or images: Evaluating the table reasoning ability of llms and mllms.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Tables as texts or images: Evaluating the table reasoning ability of llms and mllms

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:05.289155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:20:01.257151Z digest=sha256:43133708c3746f7100e7bed2a89b28bcf0d9229fe612cdb146cb2ef67ceb5c83

Pith citing papers

Observation 014cfda9-bca4-4b8a-98b3-9f9f3905be6e · inbound

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows cites this paper.

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-22T00:22:12.404705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T05:56:36.312877Z digest=sha256:fd9cf7b0f32147337274b3a4917036026973bb26aeb240617977010db44d1d57

Observation 5dad171a-f708-44a9-9e65-1a3ab37af9d1 · inbound

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows cites this paper.

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-22T00:22:12.404705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T22:18:45.189576Z digest=sha256:4947dfcf9feacc41d1bdc5e7bb10ca65dd037246359103f90f05d684c3bf5576