Pith. sign in

Paper Citation Record · LEDGER

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making

As of 16 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2509.08500.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08500 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:11:01.051591Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87b29f39-d4d4-4a3c-9040-a70258a82d27 · outbound

This paper cites Imitating Interactive Intelligence.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Imitating Interactive Intelligence

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.766240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.766240Z digest=sha256:fe5bb5891da49ba3c4d7b3ad97c3b6c315f678a1653e9796e8c30a420460c225

Observation 9f56c574-3654-403d-b367-4c78d77d4259 · outbound

This paper cites GPT-4 Technical Report.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.772456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.772456Z digest=sha256:3fa3d576e9d17ae4ab967cdc4ad9460c2ffcd9f8f0c0496790287dc8bd5923a4

Observation 0b1041d3-34d7-40fc-bf73-871067a26eeb · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.777561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.777561Z digest=sha256:6229b403f397175f2cba39e5ecbfc73457edaa0d27dd3bc636892a436aca1d72

Observation b49d3562-4528-44df-90d7-1ffdc44cbe49 · outbound

This paper cites Language Models are Few-Shot Learners.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Language Models are Few-Shot Learners

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.782636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.782636Z digest=sha256:0fd8e16782189435ae29d57e54bd2bc1e0b622e870b465aac4f7938c4833d44a

Observation f4869960-63a6-49f1-8ab0-3854af2b47e6 · outbound

This paper cites an unresolved cited work.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:11:01.905661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:00.788505Z digest=sha256:ebef3521afc33c4dbc97b04fbd7fb1032794b196c5ab84ffa9196e417616180d

Observation 4a2ffb51-0dff-4223-a3e4-9c42d215cd99 · outbound

This paper cites OPTune: Efficient Online Preference Tuning.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making OPTune: Efficient Online Preference Tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.792899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.792899Z digest=sha256:04b98253f5590b95199800f54cc0fa544a3704f4b0b887cfc1fb839c8ed67b90

Observation bb2a0d2b-9cdb-42fe-94bf-bcc6da695259 · outbound

This paper cites BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.798499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.798499Z digest=sha256:76924f43ba383b569d849671090f573ed4a73842dff81e490b84ac40e677b8ed

Observation 562355af-eb32-4f28-9f2b-398964c294d2 · outbound

This paper cites an unresolved cited work.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.803929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.803929Z digest=sha256:3c2841a65d42597524a517385ac583a4ca7604fe9f0021a3ccf784a04299f738

Observation c2e33a4e-4900-42bc-b1f6-96d16332c448 · outbound

This paper cites Collaborating with language models for embodied reasoning.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Collaborating with language models for embodied reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.808308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.808308Z digest=sha256:ba4f4415e2d48f22c03453d4ebfe89c1a17bcdc45fa0b7555bbb82a22f8f9dd4

Observation a80fc14e-5391-41a6-8fe1-80fad7649b1b · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making PaLM-E: An Embodied Multimodal Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.813141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.813141Z digest=sha256:8e50714b98836f9f3a8bd835afdd3da5e8b676dbc5101370a0937a0e454e9164

Observation 1d392e73-e90f-45dc-bbda-ae13166e4318 · outbound

This paper cites an unresolved cited work.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:11:01.880014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:00.819258Z digest=sha256:a6d8ad3cec0d7804e2cb1a2b0c2f9388a3cc74c6b9b6ee3a225bbf2ed1d3ea17

Observation 129c308b-f37e-4701-b451-6d762e4964a6 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making LoRA: Low-Rank Adaptation of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.823923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.823923Z digest=sha256:94c82a2c29d397ae3a0861f22c726cdbd3d1b019371e1e72fa3c78c4bd8dbbe1

Observation 821b8223-5c76-44e9-b805-3a8b540e4131 · outbound

This paper cites an unresolved cited work.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:11:01.865597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:00.829817Z digest=sha256:b7f3535997e7f6dcc72cae44a3d64d4e23f73f45d1fe6fe26c16fd4fa2dd3983

Observation 69624f41-b3f1-4e93-837c-429c3e74e991 · outbound

This paper cites an unresolved cited work.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.834186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.834186Z digest=sha256:f319e73e0fc32676a4ddd25b2c8b42cc1d56a8ba11fa7dba6c77fecbc98c77d5

Observation db55326e-a261-47c8-8ad4-8a1368be8587 · outbound

This paper cites Inner Monologue: Embodied Reasoning through Planning with Language Models.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Inner Monologue: Embodied Reasoning through Planning with Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.838943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.838943Z digest=sha256:f6313fda8b4e2c8e174d2443e1dc95b70b8b6bb913344b82884db68dcc48d6c3

Observation 3e5ec5ce-5a4a-47af-b0bc-f55ec56631be · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making VIMA: General Robot Manipulation with Multimodal Prompts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.845835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.845835Z digest=sha256:839100f02181f81b200d979c07861d005d8284418368cb0d5eb16372abc32e2a

Observation 813a334c-acd1-4940-ae98-0427289f4bdb · outbound

This paper cites an unresolved cited work.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:11:01.841417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:00.850449Z digest=sha256:5abd9bbd7fb4c3902982bd083f142eb8a33b30f4a03d8c03c67b8b6303a55d00

Observation 27ac079e-be50-4e82-8777-a8dfe79a87cd · outbound

This paper cites an unresolved cited work.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:11:01.828672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:00.855488Z digest=sha256:34c4173e8542f8ba214ea883737d75fde294d9558bedd8faaa130397b412b213

Observation 97e20964-8a30-4e2b-a6d1-0fce6829b058 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making OpenVLA: An Open-Source Vision-Language-Action Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.860569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.860569Z digest=sha256:3ae919f6395d7c430ca44238955c29031fb9447bb49c8a7dabe2773f06e21afb

Observation 10d2c1ef-7518-47a2-9077-539b32c76d44 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.865781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.865781Z digest=sha256:1d2b1eb2ddbac9aa6bbcc88aeddf81d87048a6b679f22c014ba365c86f3fa90d

Observation 1eacf870-1e5d-4716-954b-f79a449c30aa · outbound

This paper cites an unresolved cited work.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:11:01.812953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:00.870822Z digest=sha256:f48bee53a76f8f6059ab39c9ac6918bcc65b451543b991b2ab995a0c50034a7e

Observation afcb89d8-51ab-47a0-bee5-448108f9d5ed · outbound

This paper cites an unresolved cited work.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.877005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.877005Z digest=sha256:39805394a22bf1731953a993b2c194d0b415733a6f6a30cf3cf16ff7cc920c4c

Observation 144c4876-7ef7-4261-ab6c-4d31d09730e7 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Improved Baselines with Visual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.885762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.885762Z digest=sha256:10c0eb315add9e476835cda82ca247f16a8561453b774099c0662174b6c3af17

Observation dc5e1e8a-3c45-46ba-b23e-0a1633568b5d · outbound

This paper cites Visual Instruction Tuning.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Visual Instruction Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.890802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.890802Z digest=sha256:5ba4b6c3511af29fa84fb3e1e733b34f74616d2bb31c23d250c1b669264dd650

Observation 206b4537-d745-432b-ac5d-f7fdcfb3567c · outbound

This paper cites an unresolved cited work.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:11:01.787061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:00.896012Z digest=sha256:213dcb810966444da724f2abaf2a6d57428179351a67dbda5a15b31b5611e788

Observation 2d127f8f-ae53-4e0d-9bbc-45d8ec811fc5 · outbound

This paper cites an unresolved cited work.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:11:01.755003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:00.901797Z digest=sha256:f73f0deac2156cff9e2ee59f1ed6694a2fd8c6e7e78bd8cae35d1e44588cd040

Observation 36e3c82d-b90b-4dab-a148-3541925de2a7 · outbound

This paper cites EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.908527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.908527Z digest=sha256:d4b9d7c6519394a610fb19f6990bfa305e411419e42e0e4386fcb4b7a96b1d04

Observation 010c77d7-6462-407e-98ee-01b1524edbee · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.913731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.913731Z digest=sha256:d2044fda9e9c56c94b5cfecffe30028d021f0a7a1a595f25ba6c7c3e8646515c

Observation 4d35374c-99af-484a-8f67-3e30b3e6d4a6 · outbound

This paper cites an unresolved cited work.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:11:01.741379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:00.919964Z digest=sha256:e897ec4511a0a74c464a92ea34b7d90717cd7d91b3e23b94f8e5d43a425d9fbb

Observation c2c3ef01-97be-43d3-8910-f41be6b0d160 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.924355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.924355Z digest=sha256:0e0ac9c24c3cff601b717bc6764364118a4b71c891c9ea9f2868abe377feb64f

Observation 6cea1050-7299-4ed4-802b-33a7fa550d43 · outbound

This paper cites an unresolved cited work.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:11:01.728816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:00.929926Z digest=sha256:a5bf6a0422e30d9085cab92fbd9dd309cb4cc1d29f38691b18be173d31329d31

Observation 42b97753-2795-4b45-90eb-9c2bad98b315 · outbound

This paper cites Proximal Policy Optimization Algorithms.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Proximal Policy Optimization Algorithms

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.934472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.934472Z digest=sha256:5e47002744e1f6e42ce56eaaf9e6e5ecb96e83c806d4e70929af1579e3adf55b

Observation 869b56a2-1051-4a2d-a67b-e1b60b722ecd · outbound

This paper cites an unresolved cited work.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.939250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.939250Z digest=sha256:447226d42a31b8b8ea24316d2e5b3822dddd725ea030e44ab14033f68c4e39c6

Observation d8631351-4987-4ff4-80dd-ee00d54ad4ce · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.943781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.943781Z digest=sha256:55ced6d697c3910128b8524b67300f6ee5e3f6674600007225f4a1dbafbc3ae5

Observation fd26df8c-2402-496d-8b7b-a916052f58b8 · outbound

This paper cites an unresolved cited work.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:11:01.705195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:00.949828Z digest=sha256:8b14efb927f922977d8457de586f66d3d6d2aa46790e047f05d6d802d89a7dd0

Observation d454170c-4997-47d0-b31b-38b5d23ecbbc · outbound

This paper cites ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.953945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.953945Z digest=sha256:d3194ea96d55a1d73a9c47a507986a1d0070dbbec4b3238e4494686a63937941

Observation e15cd55d-153e-4c08-b6b3-9aad3c3e3d65 · outbound

This paper cites an unresolved cited work.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.959117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.959117Z digest=sha256:89f30f6abc5a401fdd64a794c6084bcc9b4c9da5a59631a1b417cce09af9fa30

Observation a9ebf648-e27e-4221-b1c5-ac409db280b6 · outbound

This paper cites LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.964019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.964019Z digest=sha256:abb92367b466c4a739ad649d94adcaf99d99d4063011d83e97b581fd9189ee6e

Observation 379e5ebb-931f-46da-896f-3cc959faf73a · outbound

This paper cites an unresolved cited work.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:11:01.680347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:00.968877Z digest=sha256:882732d8c99371689044b032b87639aabf4961de8aa2663001866c9baa8cad8f

Observation 48887b85-f7a7-48e1-92a8-9a168ceba31c · outbound

This paper cites Large Language Models as Generalizable Policies for Embodied Tasks.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Large Language Models as Generalizable Policies for Embodied Tasks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.974344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.974344Z digest=sha256:c0bf28f53d4be8af42567c68ee943a5eb89a64c6f6dbdfae050834f9d3cd5137

Observation f35751e6-0ba2-405d-a205-e6b8a554c955 · outbound

This paper cites True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.980528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.980528Z digest=sha256:c721e10fe42d99e0d6c207ef8c26444befdbefdbc9b62c026ed421dcdac163ff

Observation b0284552-175a-4dcd-b962-09625b13b4aa · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Gemini: A Family of Highly Capable Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.986012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.986012Z digest=sha256:8f39f746e43071f2505220d23b07b996d38e4b2f8095ee84a1d41257698cc354

Observation 7312b00f-c7f5-42d3-a04d-ac2c832597c7 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.993052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.993052Z digest=sha256:47dc3dad2c98a734e83336b8c3b08c6853d4dbc437b2d8ba367ae1cbc8e53fa4

Observation 867aaa16-e278-4b50-9f21-71f28701bf5c · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:00.998779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:00.998779Z digest=sha256:3495b4d0d4630e29a98b7703a02e27b2f225483d05a2b0d93ae41b2ed901fefa

Observation e85970ca-6342-4352-93bd-9e30cc75a47d · outbound

This paper cites A Survey on Robotics with Foundation Models: toward Embodied AI.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making A Survey on Robotics with Foundation Models: toward Embodied AI

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:01.004013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:01.004013Z digest=sha256:e4b2e97e8e6c55a06c6fc169afdccc6242b7315fa65f1e1348c105040ed445d0

Observation e660ba09-3afa-4256-85b3-35384d12acd8 · outbound

This paper cites an unresolved cited work.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:01.010642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:01.010642Z digest=sha256:9f0171744ec84092d6b04b34bbf09747b960cdd3ca6c459a5554c5203093299a

Observation 8b5ca6b3-37ac-4d39-8647-1e157589c1c6 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:01.017570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:01.017570Z digest=sha256:1ec0ff8afd5d50d04c858ba8df4136803f9a31e64c207f9c68be0a1df1c4f6fd

Observation 2a65a84a-288d-4b3b-adef-a8f45a76fec9 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making ReAct: Synergizing Reasoning and Acting in Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:01.027185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:01.027185Z digest=sha256:82fc3caa607daf0e73fdd6dc9db4038b0a0aa7d0829417f0a418703877499f73

Observation 03ab1afe-3796-4add-bf1c-ac130e4bf1ec · outbound

This paper cites Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:01.031589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:01.031589Z digest=sha256:05e9202169755b65a33535046bba84a4ff7c5fc38b8d1d85bb0f0500c59abd14

Observation b3a8be4e-9ebd-45f0-a26e-e792535251c5 · outbound

This paper cites Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:01.036940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:01.036940Z digest=sha256:fd022ed886b333cd57695e99baf74b4d7df5e54efb236a4e29003610d232bc87

Observation b000d9bf-5e07-480e-ba36-748488e8d227 · outbound

This paper cites an unresolved cited work.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:11:01.666476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T16:11:01.041510Z digest=sha256:0f78411218bdb3701c1ef43b116ecac1c9985a736e7e221472ace72eebe4d5b9

Observation 002e3b0c-c23a-41aa-9a25-bad7182398c5 · outbound

This paper cites online" 'onlinestring :=.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making online" 'onlinestring :=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:01.045998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:01.045998Z digest=sha256:b53fb9e16b3b5fb6ae82c74d6fc6b20c398825894d9c04a6aefffedbce9ac176

Observation a8342e27-d07c-486c-af36-f53be54b84fc · outbound

This paper cites write newline.

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making write newline

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T16:11:01.051591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:11:01.051591Z digest=sha256:297afb279646b51d06e9494ca6b59b05505b99994ea2c741c5b14adf6c8ecce5

Pith citing papers

No inbound Pith citation observations are available.