Pith. sign in

Paper Citation Record · LEDGER

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy

As of 20 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2608.07427.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07427 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:51:16.342069Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact2
  • verified fuzzy8
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e86c4175-46a7-4286-b62e-f555d045ec46 · outbound

This paper cites Can US infrastructure keep up with AI economy?.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Can US infrastructure keep up with AI economy?

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:51:17.682845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T04:51:16.250855Z digest=sha256:44b782e87e7e60804c1f7907d0fc6703db781547c39b01627d12d0796157a909

Observation 7d0ca9c2-e7cd-4c01-917e-2319287b237d · outbound

This paper cites The future of datacenters,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy The future of datacenters,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:51:17.669592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T04:51:16.255689Z digest=sha256:bd566856268b42a08a58708a981dd61e6ad7f8916c3a779f23de6c11a6b392d3

Observation fc289954-9359-4e9d-9304-67464b88dc4a · outbound

This paper cites How hungry is AI? Benchmarking energy, water and carbon footprint of LLM inference,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy How hungry is AI? Benchmarking energy, water and carbon footprint of LLM inference,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.260016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.260016Z digest=sha256:d9d8ec300047404368ca2ad23fa7f9528e6252456821265fda7879c1eee774e3

Observation dd8b3348-8555-4d0f-bf7d-735982898ac6 · outbound

This paper cites Artificial intelligence and energy overcon- sumption: Data center electricity demand, cooling burdens, and regional sustainability constraints,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Artificial intelligence and energy overcon- sumption: Data center electricity demand, cooling burdens, and regional sustainability constraints,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:51:17.655713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T04:51:16.264462Z digest=sha256:5235d25388a4f064e198ab2d44398394e0ea70efb81d1d9d5448ee3af1a56a40

Observation bbc5bbef-a87a-45da-b37f-69c669e01afd · outbound

This paper cites Comparative analysis on developed optimization techniques for reducing energy consumption in AI training and inference,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Comparative analysis on developed optimization techniques for reducing energy consumption in AI training and inference,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:51:17.640829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T04:51:16.268990Z digest=sha256:238adc5c3fe236f4fe15d22cb27406b4b2166c441d5df707c3fc2674bb3bed25

Observation bbb05f71-675b-423b-a571-ba990905c72a · outbound

This paper cites From prompts to power: Measuring the energy footprints of LLM inference,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy From prompts to power: Measuring the energy footprints of LLM inference,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.273744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.273744Z digest=sha256:b12de3360ae697d48bfd5362aa8d48af54653a39826511ce8776eb29447b621e

Observation 1aeb037a-8a26-4336-9779-2a5426fc744a · outbound

This paper cites Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.278459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.278459Z digest=sha256:1f2311ae30df33fa94905e544c2d6bf8a2c34252184685e9099a75982c455609

Observation e30abeef-7af0-4cce-90e5-89d990b2da18 · outbound

This paper cites How different tokenization algo- rithm impact LLMs and transformer models for binary code analysis,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy How different tokenization algo- rithm impact LLMs and transformer models for binary code analysis,

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-08-10T04:51:17.217901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T04:51:16.283034Z digest=sha256:f5b14f02cd8d4ce45aa96a4007aecf02eb4e2ff06dc5dba2561da37cadb02ac8

Observation b4287e23-e609-4ad9-9b7d-6c048b1df8f8 · outbound

This paper cites Harnessing vision-language models for time series anomaly detection,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Harnessing vision-language models for time series anomaly detection,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.286983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.286983Z digest=sha256:136a00bb68e1bb19bcd9cebd09c2fcaf4e51463dea01862cf0ed42bd7db77a58

Observation 0d2f52fe-fbe6-4f04-a526-cada29527f67 · outbound

This paper cites TokenPowerBench: Benchmarking the power consumption of LLM inference,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy TokenPowerBench: Benchmarking the power consumption of LLM inference,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.291049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.291049Z digest=sha256:68c454b07937bb5083e575b88cf354a59ca74deded453204b68037debc8a9b00

Observation cf8957cc-a73d-4287-94c5-1b4f3a91685f · outbound

This paper cites The ML.ENERGY benchmark: Toward automated inference energy measurement and optimization,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy The ML.ENERGY benchmark: Toward automated inference energy measurement and optimization,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.294955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.294955Z digest=sha256:261e892e792cec3278274ef7dac952170ee1874fbaf1ad72a9e653fc743acb0a

Observation fc361cb8-1430-4d56-8148-8ea9cdff5c19 · outbound

This paper cites An energy-efficient vision language model inference with importance-aware token pruning,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy An energy-efficient vision language model inference with importance-aware token pruning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:51:17.626613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T04:51:16.298990Z digest=sha256:46d26672bb688a016962e9ba5bf784f7d096445186f8b3152eeb48df555a3024

Observation 92c823a5-1470-42e2-8e75-c50bdc24a828 · outbound

This paper cites Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.303359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.303359Z digest=sha256:a7a892daa40944c22c6b605c4f1e0ec884c9f63f23c8c1f24a5661504d35119c

Observation f3381d6d-fe4e-4a4e-9902-ea80c3f64b56 · outbound

This paper cites A Picture is Worth A Thousand Numbers: Enabling LLMs Reason about Time Series via Visualization.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy A Picture is Worth A Thousand Numbers: Enabling LLMs Reason about Time Series via Visualization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.308177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.308177Z digest=sha256:af9e69ea21b09140b774c2e8da95ca9a615322e0bd9d4860eb328e169603ad7a

Observation 23f19057-c6ed-4369-8bdb-6a68b50353a2 · outbound

This paper cites Edge-Deployable LLM Fine-Tuning on a Single GPU for Telecom Network Troubleshooting.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Edge-Deployable LLM Fine-Tuning on a Single GPU for Telecom Network Troubleshooting

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-10T04:51:16.426717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T04:51:16.312840Z digest=sha256:04a6dd26e980067a9e7c2f82ec8cf81a3d4b8b90a59496f586ccf8e69e3cef4a

Observation f42fc48d-43db-4062-b758-6986cc746f6d · outbound

This paper cites Zeus: Understanding and optimizing GPU energy consumption of DNN training,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Zeus: Understanding and optimizing GPU energy consumption of DNN training,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:51:17.612703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T04:51:16.317365Z digest=sha256:646691cbcb44fc65728ef10c37d5944fa5f7d71c30e5dc8d9b389466cd292ae5

Observation d5f6ba40-85f2-42cf-acd7-1412f693b57d · outbound

This paper cites The Llama 3 Herd of Models.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy The Llama 3 Herd of Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.323125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.323125Z digest=sha256:026d16bd92c2410a322c978909efbc7aadfb48b82d623dc345615759989c561d

Observation be40f393-fb5b-4a06-a150-58204cb18b2e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.328390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.328390Z digest=sha256:2512865a5887ad262249aea0606b21f12cdb7528529b31492ab11b629fa1bf98

Observation 8c5fe114-94a8-444f-93aa-c622da1dfad6 · outbound

This paper cites Pixtral 12B.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Pixtral 12B

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.333215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.333215Z digest=sha256:2625ac38ed151d62ac45c7f46525200f5399eb1eee32de1a4834b0f0159966b9

Observation a8f2d798-5472-4240-b21b-e5888612a969 · outbound

This paper cites Causal inter- vention sequence analysis for fault tracking in radio access networks,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Causal inter- vention sequence analysis for fault tracking in radio access networks,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:51:17.598284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T04:51:16.337747Z digest=sha256:b2a1f8535692d1eb0999118063de02ea77559efd94f8c3d3204f39c0a6166fdb

Observation bda76dd2-3446-4364-821f-8e760aa0dadd · outbound

This paper cites Cost efficient GPU cluster management for training and inference of deep learning,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Cost efficient GPU cluster management for training and inference of deep learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:51:17.583501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T04:51:16.342069Z digest=sha256:5f2e7c4b5b53654b7e365cabb0b34d16427c6523404b99dfd5ce7570b85d612a

Pith citing papers

No inbound Pith citation observations are available.