Pith. sign in

REVIEW 1 cited by

Attentive Explanations: Justifying Decisions and Pointing to the Evidence

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1612.04757 v2 pith:75Z74U3B submitted 2016-12-14 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords decisionsvisualdecisionmodelsevidencepointinganswerchallenging
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep models are the defacto standard in visual decision models due to their impressive performance on a wide array of visual tasks. However, they are frequently seen as opaque and are unable to explain their decisions. In contrast, humans can justify their decisions with natural language and point to the evidence in the visual world which led to their decisions. We postulate that deep models can do this as well and propose our Pointing and Justification (PJ-X) model which can justify its decision with a sentence and point to the evidence by introspecting its decision and explanation process using an attention mechanism. Unfortunately there is no dataset available with reference explanations for visual decision making. We thus collect two datasets in two domains where it is interesting and challenging to explain decisions. First, we extend the visual question answering task to not only provide an answer but also a natural language explanation for the answer. Second, we focus on explaining human activities which is traditionally more challenging than object classification. We extensively evaluate our PJ-X model, both on the justification and pointing tasks, by comparing it to prior models and ablations using both automatic and human evaluations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generative Visual Commonsense Answering and Explaining with Generative Scene Graph Constructing

    cs.CV 2025-01 conditional novelty 5.0 of 10

    G2 generates a location-free scene graph from an image and feeds it, with confidence-based token weighting, into an LLM to produce visual commonsense answers and explanations.

Pith tools