Pith. sign in

REVIEW 2 cited by

Hierarchical Memory Decoding for Video Captioning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.11886 v1 pith:R3YCT4AU submitted 2020-02-27 cs.CV

classification cs.CV
keywords decoderinformationcaptioningmemorynetworkvideodecodinglong-term
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances of video captioning often employ a recurrent neural network (RNN) as the decoder. However, RNN is prone to diluting long-term information. Recent works have demonstrated memory network (MemNet) has the advantage of storing long-term information. However, as the decoder, it has not been well exploited for video captioning. The reason partially comes from the difficulty of sequence decoding with MemNet. Instead of the common practice, i.e., sequence decoding with RNN, in this paper, we devise a novel memory decoder for video captioning. Concretely, after obtaining representation of each frame through a pre-trained network, we first fuse the visual and lexical information. Then, at each time step, we construct a multi-layer MemNet-based decoder, i.e., in each layer, we employ a memory set to store previous information and an attention mechanism to select the information related to the current input. Thus, this decoder avoids the dilution of long-term information. And the multi-layer architecture is helpful for capturing dependencies between frames and word sequences. Experimental results show that even without the encoding network, our decoder still could obtain competitive performance and outperform the performance of RNN decoder. Furthermore, compared with one-layer RNN decoder, our decoder has fewer parameters.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A nested MLMC framework for efficient simulations on FPGAs

    q-fin.CP 2025-02 conditional novelty 6.0 of 10

    A nested multilevel Monte Carlo framework with approximate normal random numbers and per-variable fixed-point precision optimization predicts 5-7x cost savings at coarse levels versus standard MLMC, pending FPGA validation.

  2. Hardware-Aware Data and Instruction Mapping for AI Tasks: Balancing Parallelism, I/O and Memory Tradeoffs

    cs.AR 2025-09 conditional novelty 4.0 of 10

    A message-driven mapping framework for VGG-19 inference on the MAVeC accelerator is claimed to generate over 97% of messages on-chip and sustain 88-92% SiteO utilization in simulation.

Pith tools