Pith. sign in

Paper Citation Record · LEDGER

TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2303.11897.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.11897 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:35:00.351522Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T14:57:03.744199Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f3b4525a-3f97-4bed-9eb8-9dd8ca0791b8 · inbound

GAIA: a benchmark for General AI Assistants cites this paper.

GAIA: a benchmark for General AI Assistants TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 168

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:46:03.602428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T15:46:03.247029Z digest=sha256:402e3770f1816c6d65250110b3723a371123ae8dc9fae691c174fff05a791ae6

Observation 532dd131-098b-434e-a28d-6f80c7074153 · inbound

BLINK: Multimodal Large Language Models Can See but Not Perceive cites this paper.

BLINK: Multimodal Large Language Models Can See but Not Perceive TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.689615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:a820feb683711e0373a9404653e60aa46f4e55613f8e3c4023e67247a94ba315

Observation 338c2a94-95c9-43ff-94cb-622dd6a851e1 · inbound

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP cites this paper.

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:35:00.351522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:35:00.351522Z digest=sha256:e330c583a1edcbf36f8c877b96974ba86d2bd43829b289a3bfffc42a157ebfac

Observation ac972a64-6cde-41a2-b6f0-db927337c602 · inbound

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning cites this paper.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.101775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.101775Z digest=sha256:ff545a321e2d8567e031a2524bc881a094e1dee5c069ee5c8c300059d338b7ec

Observation ca6781f0-1d25-4fe7-b785-4916d580ee93 · inbound

DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative Models cites this paper.

DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative Models TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 2020

Resolution
malformed identifier
no resolver link, observed 2026-08-07T10:30:27.996361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:30:27.996361Z digest=sha256:ddd8b5d5e1cfb630848f864a18affc5126c74739182d2f050a6e9527884ea04c

Observation ea392ca7-4baf-476d-9aa9-5cd84cd6d310 · inbound

AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation cites this paper.

AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:03:00.879353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:03:00.879353Z digest=sha256:0fda62c40e28f12256e4ec7ac0948aaeb5e9535884982b614f87d4cc2603994a

Observation 676ddf00-af40-400e-bdb1-fad7f263a898 · inbound

Towards Evaluating Robustness of Prompt Adherence in Text to Image Models cites this paper.

Towards Evaluating Robustness of Prompt Adherence in Text to Image Models TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:50:42.666253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:50:42.666253Z digest=sha256:780b89437e3a7f50fdbc3752dd12209c869318a86e18fec07be293d93f6e2068

Observation 67850597-ba20-4148-bcdb-b7fd79c80a67 · inbound

Disentanglement in T-space for Faster and Distributed Training of Diffusion Models with Fewer Latent-states cites this paper.

Disentanglement in T-space for Faster and Distributed Training of Diffusion Models with Fewer Latent-states TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T18:41:32.084239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:41:32.084239Z digest=sha256:ae10143f648cb92475e520c8d9917535751a25d09ca3877e40013f3a5a33b51b

Observation 5fa3a9d5-e248-4f60-a87a-216b5424476f · inbound

Understanding and evaluating computer vision models through the lens of counterfactuals cites this paper.

Understanding and evaluating computer vision models through the lens of counterfactuals TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 208

Resolution
unresolved
no resolver link, observed 2026-08-05T14:49:31.846334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:49:31.846334Z digest=sha256:c9e5162a36e10211d2aede2a1b0e8aa34a252fd60781fba1d4bb19a3d0f536af

Observation e75287e3-5e9d-47c4-a4c7-88e450328cb7 · inbound

Discovering Divergent Representations between Text-to-Image Models cites this paper.

Discovering Divergent Representations between Text-to-Image Models TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:32.044233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:32.044233Z digest=sha256:f0f57223470700cf012a7277bc03f0c3fcefa9d3d57206df927b97529df33652

Observation 58469a3f-45db-4005-a7a3-ef05c61946c9 · inbound

Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs cites this paper.

Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:45:30.714518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T23:43:43.201888Z digest=sha256:c2c5639278bc6bd4a6d4a62dede64da353ecc3e78279b5e11c2188d1551375f4

Observation 63356f5d-cb41-4586-97de-a4ed4b9b54bd · inbound

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics cites this paper.

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:03:19.982943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T20:59:52.448832Z digest=sha256:f894ef91e0a90fb6bc1eda28fe3c36ef2f32edfbc9fdc788adce6516ed01b87e

Observation 737195c3-2cc8-435c-bc32-5f5e19954836 · inbound

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics cites this paper.

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:40:00.755494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T10:35:39.269869Z digest=sha256:5f0e0b14439d250345d001473a9c0060f1a02177cf9911bd41ce43b2d6dfba0d

Observation 71710eff-027c-49da-8616-035590152c32 · inbound

QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering cites this paper.

QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:36:16.404118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T04:51:18.203826Z digest=sha256:57b8b4c16784890204868db6f7e65dbc673b553b74c3d3770d1d780610ee728f

Observation 89b589b9-67e5-432f-b3c4-ee3559a019dc · inbound

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment cites this paper.

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.092852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T12:11:23.775843Z digest=sha256:3974b9fb1ae741f40c98ead7b349fdd5034c71d2a29dfec24850e3159299d005

Observation 7b2dff6e-02cf-4253-90af-4f3cb4d3f9a3 · inbound

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment cites this paper.

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:21:21.500312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T09:19:39.848194Z digest=sha256:0dcb3ee08918cfbf37cc9c2df99d81d06237a0550affed2fc42c477b80fbd28d

Observation fd596e6c-e198-4a01-b9fb-ad3df36462e8 · inbound

The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models cites this paper.

The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:57:03.745541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T14:56:40.860766Z digest=sha256:d88b1146b63fe20579f92ae7d754af02401d8e0d2d98b2a8de8a0aa5183eae80

Observation b557d234-f48b-4829-8f35-10d73131098d · inbound

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions cites this paper.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.867139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.867139Z digest=sha256:b263056508e76a4772a64de4e1efab0282b5eefbca6221199a5511a78eab4082

Observation 6dee82f8-7c2c-4f64-8a3c-7490ed143fdd · inbound

Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis cites this paper.

Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T01:00:04.666098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T01:00:04.666098Z digest=sha256:e726bad2f1f38b6b12a3a01889350a0409116d36635ede4f32550e66ae841791