Pith. sign in

Paper Citation Record · LEDGER

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection

As of 23 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2502.09829.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09829 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:24:08.254315Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:15:34.412634Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T22:15:35.405013Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact1
  • verified fuzzy46
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26c6ca63-413b-481a-93ed-bf280aed8ed2 · outbound

This paper cites Sim-to-real transfer for vision-and-language navigation.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Sim-to-real transfer for vision-and-language navigation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:09.084157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.005910Z digest=sha256:aadc88221408d88b03acba911e9fbba1d95f786d13a7aa4c2cbf48d1095eddd4

Observation ed285a56-c073-4fc5-92c2-1360f2f1e3d6 · outbound

This paper cites Con- trast sets for evaluating language-guided robot policies.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Con- trast sets for evaluating language-guided robot policies

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:09.069010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.011688Z digest=sha256:cdd98c6b03688e6f21ba44c917586eed917a1a5d39ef64aecb2dac6d7fdd7809

Observation a94521aa-8179-4cf4-a5a7-c1e53b0b35b7 · outbound

This paper cites Remembr: Building and reasoning over long-horizon spatio-temporal memory for robot navigation.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Remembr: Building and reasoning over long-horizon spatio-temporal memory for robot navigation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:09.053902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.016879Z digest=sha256:ff9a5f4643576b4ddf277da158f7d6d65a74fe028c2d0c9c1371ba2c83dc383a

Observation b6c14229-43af-4668-9123-170bb3a08b9d · outbound

This paper cites Surrogate assisted generation of human-robot interaction scenarios.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Surrogate assisted generation of human-robot interaction scenarios

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:09.039893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.022672Z digest=sha256:22d98ee3938872ecc36e7ec06c74152598d60572270555e39a4b5a2d4a6d0046

Observation cc92b23e-6d74-4a9e-a3e9-948afe2cdf3c · outbound

This paper cites Mixture density networks.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Mixture density networks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:09.025105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.027904Z digest=sha256:fea44c71c79f08ee56dc524fa1db9f8ebeadd5712baec06e528a0c52d92a92e9

Observation db870277-d89b-4a27-89cc-73048e4422ed · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T20:24:08.032884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:24:08.032884Z digest=sha256:8a0961557e8084c36019d8e761d4a37ff591985430609f1cd5a11c9265a1c682

Observation f22198aa-5c07-48af-a897-8d3fffdc7860 · outbound

This paper cites A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T20:24:08.038303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:24:08.038303Z digest=sha256:cd6d35a1f741ea1a0a149179f650521be5fccaa7518d546a3f571ef2e6b83823

Observation a89ec4cd-1752-49f3-9edd-5370b68f058a · outbound

This paper cites A survey on evaluation of large language models.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection A survey on evaluation of large language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:09.010063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.043968Z digest=sha256:63ff6650b8e894d017b142ae3587387f915a41335aae08122822c8257fd67ecc

Observation 7af56f50-e073-4895-9362-7eaeeb9e849c · outbound

This paper cites Learning surrogate models for simulation-based opti- mization.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Learning surrogate models for simulation-based opti- mization

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.995807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.048745Z digest=sha256:8cf878cd75711dc9bc388c8a4c479b1fb977030a2559f164aae8f62683f6ed71

Observation 2559ab78-111d-49eb-900b-57a59fde8450 · outbound

This paper cites RoboTHOR: An Open Simulation-to-Real Embodied AI Platform.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection RoboTHOR: An Open Simulation-to-Real Embodied AI Platform

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.982018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.053466Z digest=sha256:ecee111f6b17bde60e180b7e829107d46e2549a0cd843ca0121b7cf760c12ef3

Observation 7cf5f8fa-e923-4c70-9727-6bfae3fff535 · outbound

This paper cites Efficient benchmarking of hyper- parameter optimizers via surrogates.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Efficient benchmarking of hyper- parameter optimizers via surrogates

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.964796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.058304Z digest=sha256:5e393840d333bd51ff4657b374f8219b13c5527ccc06ab9b51811053d8b8ccff

Observation 55e5f08b-1cff-4d78-a5ab-059acbd87550 · outbound

This paper cites Dropout as a bayesian approximation: Representing model uncertainty in deep learning.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Dropout as a bayesian approximation: Representing model uncertainty in deep learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.950854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.063294Z digest=sha256:eea3cf001d9a8293351be533cf58924b4d95653399b62a28e4babc6e11ce7b63

Observation fadbd227-7ddb-47a9-b67b-8ce59014fc0e · outbound

This paper cites Efficient Data Collection for Robotic Manipulation via Compositional Generalization.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Efficient Data Collection for Robotic Manipulation via Compositional Generalization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.936588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.067842Z digest=sha256:3de1e7cdf5ada7c66f484f123d8c7feca4be86b1ca1bc2bef5e8ca024995cacf

Observation 663b53e1-19c5-4ca3-abaf-0fcec1a5f52b · outbound

This paper cites Liu, Phoebe Mul- caire, Qiang Ning, Sameer Singh, Noah A.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Liu, Phoebe Mul- caire, Qiang Ning, Sameer Singh, Noah A

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.922623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.072828Z digest=sha256:7240938f5180f029764d12bffd7a0c19e14aaa45d742729630e5ab764f30b10f

Observation f9ebe5db-34ae-4c54-8d89-1eb480d460c8 · outbound

This paper cites Navigating to objects in the real world.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Navigating to objects in the real world

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.908659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.077336Z digest=sha256:96a136cad43ac89dc62bd2208c3adffb66f471beddb6dc038752ae614de0d180

Observation 7b9e03f1-767c-4715-9a1e-7908ecd69939 · outbound

This paper cites Minillm: Knowledge distillation of large language mod- els.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Minillm: Knowledge distillation of large language mod- els

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.894585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.081928Z digest=sha256:098535b169dca8264b5f9a9087007534c67713298519d204c33705af1a09e762

Observation 40d46ed5-1ef9-438f-b1ce-3ad86553ddaa · outbound

This paper cites World models.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection World models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.880292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.086506Z digest=sha256:dd90afad81232851df76ba63c59fcdc04ec6b63496951369d4a11bcf784e515f

Observation ae60c01c-38bb-4eab-bda8-933256a9e198 · outbound

This paper cites Benchmarking neural network robustness to common corruptions and perturbations.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Benchmarking neural network robustness to common corruptions and perturbations

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.866591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.091677Z digest=sha256:c2814d7abf5dc81684e6bc8240f445a11897a5b405ba5442d57fe0669cea3529

Observation 3c0329c7-1abc-44ab-923c-bd8c08fd8085 · outbound

This paper cites Pretrained transformers improve out-of-distribution robustness.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Pretrained transformers improve out-of-distribution robustness

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.852914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.096785Z digest=sha256:b5def448bfc0d5d9056a1d368db656d8a3c75749cb36ba4ef5464a902f5d7888

Observation c923236e-2c6a-4d33-9749-448d587481be · outbound

This paper cites Bayesian Active Learning for Classification and Preference Learning.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Bayesian Active Learning for Classification and Preference Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T20:24:08.101849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:24:08.101849Z digest=sha256:c16b09183f6d163241e1c01f666893103b4dccd1b043003d7479b0211a10130e

Observation 60d4b5d4-e7a8-4930-b747-99cbbc76109a · outbound

This paper cites Deploying and Evaluating LLMs to Program Service Mobile Robots.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Deploying and Evaluating LLMs to Program Service Mobile Robots

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.838641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.107159Z digest=sha256:d00959ab37757d15a8399ea42d2927b7bdf4b668b724c2fac669abeaf5100a23

Observation fd54f599-94bd-4af1-a767-60d7402d995f · outbound

This paper cites Sim2Real Predictivity: Does Evaluation in Simulation Predict Real- World Performance? IEEE Robotics and Automation Letters (RA-L), 2020.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Sim2Real Predictivity: Does Evaluation in Simulation Predict Real- World Performance? IEEE Robotics and Automation Letters (RA-L), 2020

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.823846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.111866Z digest=sha256:eabd9423fda2791656f094cfb2c495a7a7af66742bd1b3346c56751c80447baf

Observation 0220736c-de73-4e23-a708-e1c0eed7aebe · outbound

This paper cites Openvla: An open-source vision-language-action model.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Openvla: An open-source vision-language-action model

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.808906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.116672Z digest=sha256:ebd38ea32c07b7a2fcc98f71a3ec6b543f7d0d2bce29567f5550048c57e64097

Observation 7254eecc-6edb-4afd-b687-2409a6c5afb6 · outbound

This paper cites Active testing: Sample-efficient model eval- uation.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Active testing: Sample-efficient model eval- uation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.794038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.121494Z digest=sha256:17ae2289f608a76218ca802ef1f488fff818029e1d33744e54e1ad743e8cc463

Observation 345be716-df4f-4671-a31c-23085459e484 · outbound

This paper cites Robot learn- ing as an empirical science: Best practices for policy evaluation.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Robot learn- ing as an empirical science: Best practices for policy evaluation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.778709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.126153Z digest=sha256:28dfc86e644222bfe196edf6b8d3153066a0c7c1c95e997266444887e2487ab5

Observation f273a139-e0bd-4c76-be64-b1bd2191dfd9 · outbound

This paper cites Dropout injection at test time for post hoc uncertainty quantification in neural networks.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Dropout injection at test time for post hoc uncertainty quantification in neural networks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.763886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.132137Z digest=sha256:778123a7e7eba5d4656a565b2588b0e5d92ad6be0bc8c8f5b76cbcf1dae981c2

Observation b966978c-e2d1-485a-9ba5-56afc1409e74 · outbound

This paper cites Cost-aware Bayesian Optimization.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Cost-aware Bayesian Optimization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T20:24:08.136868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:24:08.136868Z digest=sha256:3bc974b7a534bf5daa95797a8b26d41899b582302c7942c2e3bf426f218bae2f

Observation f9782aab-e785-462e-9e64-ba8d3e0c759b · outbound

This paper cites Evaluating real-world robot manipulation policies in sim- ulation.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Evaluating real-world robot manipulation policies in sim- ulation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.749435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.141947Z digest=sha256:aa721a741f055c06329ace0a7223032ae6a7023260c4f0463fa0b51e13724c5d

Observation e6f2b0d9-07de-42b5-b5a7-47db44b4efdd · outbound

This paper cites Ham- ster: Hierarchical action models for open-world robot manipulation.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Ham- ster: Hierarchical action models for open-world robot manipulation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.735234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.146913Z digest=sha256:ba24c75c6c70593bbe8ef608fcc3caa8b043cf5feba2fa8fef031bdc46112b03

Observation b9cab6f8-e7db-4cfb-96e6-2392c58d760a · outbound

This paper cites Holistic evaluation of language models.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Holistic evaluation of language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.720648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.151565Z digest=sha256:bfc188842a91fd08dbcbaf8be17ac16245650b45193f3e3b8172feee678e757a

Observation e4d9510a-19ce-4c7c-b63d-7146c9900ee1 · outbound

This paper cites A general framework for uncertainty estimation in deep learning.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection A general framework for uncertainty estimation in deep learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.706028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.155969Z digest=sha256:1a31c452d014d080605a99711c7dc32b39db4b2cc1e61b857e0bfc13ea516c7e

Observation ec8a53fe-460d-453b-9e82-7eda7d2072d0 · outbound

This paper cites Probabilistic matrix factorization.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Probabilistic matrix factorization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.691595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.160717Z digest=sha256:e80d33ee3fe81c52021541ad9feec1996a6680b63604fb1f0a58625563c5696f

Observation 0510ec4a-e5a2-4c0a-8277-218380b6ea71 · outbound

This paper cites Differential assessment of black-box ai agents.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Differential assessment of black-box ai agents

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.677074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.165429Z digest=sha256:706680162fec68b218d1200119998bd3bb29fd0ef3e2c6a8b7cabdcfbfccc130

Observation 8d361def-f9ec-4b91-bbb9-e2cf162c2127 · outbound

This paper cites Octo: An open-source generalist robot policy.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Octo: An open-source generalist robot policy

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.661880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.170063Z digest=sha256:4cda673657b28fd14b370899a14eca375d51cf1adb9b87bb94138934ee1fa60a

Observation 485d3982-4785-4bc9-9df6-b9ca7e4476ac · outbound

This paper cites Investigating the Role of Instruction Variety and Task Difficulty in Robotic Manipulation Tasks.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Investigating the Role of Instruction Variety and Task Difficulty in Robotic Manipulation Tasks

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.646307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.174469Z digest=sha256:2ed4379a0f9955a9a52362e778cc4f1256c7489fcad425776617136e075b39b2

Observation 6485b935-0ae8-457b-b4bb-7a710568aac7 · outbound

This paper cites Cost-aware bayesian optimization via information directed sampling.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Cost-aware bayesian optimization via information directed sampling

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.630414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.179124Z digest=sha256:22e2bbe7db8f8982ed4c5c946e612e1ddad8abefa3ae1cc7de7fb8791af69ece

Observation c4a3c133-140d-4825-a647-ecf73f7f4213 · outbound

This paper cites THE COLOS- SEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection THE COLOS- SEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.613142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.183537Z digest=sha256:6050eaf2506b27baa1861750caeff65793eac3cad85bb51741d6b610c8650ba1

Observation c1c848a8-6bb5-4efd-bc03-dc74cf668a9a · outbound

This paper cites Building surrogate models based on detailed and approximate simulations.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Building surrogate models based on detailed and approximate simulations

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.596637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.187971Z digest=sha256:8685a8e926985b51b8456091d870f0f265f6fe86cab0ba0624578dec218f7c60

Observation 5fcfeaf2-47ed-4286-8fa4-f33c7eec3645 · outbound

This paper cites Modern bayesian experimental design.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Modern bayesian experimental design

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.581857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.192597Z digest=sha256:d29f25a9dae3c2f5c18befa81289746d9d12bf00d852924427658d15bb3bac53

Observation 7bb24d94-39b3-4b04-bf42-c0f64301c6d9 · outbound

This paper cites Do imagenet classifiers generalize to imagenet? In International Conference on Machine Learning (ICML), 2019.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Do imagenet classifiers generalize to imagenet? In International Conference on Machine Learning (ICML), 2019

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.567419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.197577Z digest=sha256:94c2f3be5a6833a3986c7e2ba3bd95500465b1e6d192953bf3a6bc6e7ac70836

Observation 3baddab0-c4cb-423c-bb0a-c2b21ac47fdf · outbound

This paper cites Active risk estimation.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Active risk estimation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.550901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.202262Z digest=sha256:3a3490dc786efc49e3547bcaf2d186f443e9471d2e0154dc5f1873c73c6c09ac

Observation 35f8747c-1dc2-4f35-b18a-4273c8bf4a8f · outbound

This paper cites Vint: A foundation model for visual navigation.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Vint: A foundation model for visual navigation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.535117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.206920Z digest=sha256:356bf0348cd9ec82875cd5371fce67880b2927b9e9b612b9fbd49f87c9c1f223

Observation 6e72152f-c96f-4b56-89ea-b38137f396d2 · outbound

This paper cites Lm- nav: Robotic navigation with large pre-trained models of language, vision, and action.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Lm- nav: Robotic navigation with large pre-trained models of language, vision, and action

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.518129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.211508Z digest=sha256:d86ea2fefd3f8839a0042c68a489f605a0044f8d6c5998fe1b0ebe264526370e

Observation dc026d25-b7dd-4277-90d1-d17e9346d0c1 · outbound

This paper cites Taking the human out of the loop: A review of bayesian optimization.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Taking the human out of the loop: A review of bayesian optimization

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.502523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.216126Z digest=sha256:b5da8d11485a8ac56c668b4d024f2c3ac9c6f6a37d6105bfef53673a28509daf

Observation 53efac60-5367-47f7-b1ca-6fe513fef398 · outbound

This paper cites Targeted active learning for probabilistic models.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Targeted active learning for probabilistic models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-07T20:24:08.321102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.220721Z digest=sha256:58072f535481f6c6439ad03b99a180a3767f5b95378808ca529e8aefddc14049

Observation fe7e84a9-ab7e-4e51-acb8-099f44cdbdb4 · outbound

This paper cites Discovering user-interpretable capabilities of black-box planning agents.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Discovering user-interpretable capabilities of black-box planning agents

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.486723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.225896Z digest=sha256:b9b3a2ac97f198516674d4b4d8a6bc3e81eec3a09db6ebf8e2865a6dc6ef8b70

Observation 899f423e-4ef5-44af-82b2-1b9f420132ad · outbound

This paper cites Autonomous capability assessment of sequen- tial decision-making systems in stochastic settings.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Autonomous capability assessment of sequen- tial decision-making systems in stochastic settings

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.469447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.230274Z digest=sha256:2a0242688e03f97ef313db51b9e1c1c7cad484fa9f5a3de27f863cc1f6353c08

Observation b8e12e8f-c18d-43fb-bbde-97d163b158a2 · outbound

This paper cites How Generalizable Is My Behavior Cloning Policy? A Statis- tical Approach to Trustworthy Performance Evaluation.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection How Generalizable Is My Behavior Cloning Policy? A Statis- tical Approach to Trustworthy Performance Evaluation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.452435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.234744Z digest=sha256:c335a38c2beebab3bbfa4f97799906019416608f1e8afd3a055d4e630460bd71

Observation e85d0b1b-4eb4-4f05-8848-d96ae0fd9673 · outbound

This paper cites CLINE: Contrastive Learning with Semantic Negative Examples for Natural Language Understanding.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection CLINE: Contrastive Learning with Semantic Negative Examples for Natural Language Understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.436050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.239673Z digest=sha256:494bb6493bb15e1b5884624b138895abc0e49c3f3ab01f5ad3d617efeff1f886

Observation b65b5839-f159-4f26-a717-7a1f269544a3 · outbound

This paper cites Decomposing the generalization gap in imitation learning for visual robotic manipulation.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Decomposing the generalization gap in imitation learning for visual robotic manipulation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.419828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.244153Z digest=sha256:434c71c1ab22e422cffc9e1bd12a3b537b9518df4bf223abcb93db3d53b29fa1

Observation 27c5d9c3-fc00-4aff-b267-6a8fa47ed7a5 · outbound

This paper cites Sample Efficient Model Evaluation.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Sample Efficient Model Evaluation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T20:24:08.248881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:24:08.248881Z digest=sha256:9c38b6e56280b83cb8193da08bfe748cd1dd54e4e30e232023e5d692dc520736

Observation b321e13e-317e-4693-8189-37ef7d9bdfca · outbound

This paper cites Meta- world: A benchmark and evaluation for multi-task and meta reinforcement learning.

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection Meta- world: A benchmark and evaluation for multi-task and meta reinforcement learning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:24:08.403255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T20:24:08.254315Z digest=sha256:6679c648f94ac4dbda1c9e5db5b5b342f9c89d763a24c7adf69e21367a8e19c5

Pith citing papers

Observation af482dac-a339-45dc-8ba8-3eb3f78cb76b · inbound

Guiding Data Collection via Factored Scaling Curves cites this paper.

Guiding Data Collection via Factored Scaling Curves Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-15T22:15:35.412151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T22:15:34.412634Z digest=sha256:5a9b7faa93af2e8f26d86086e9245316afa4eab02257a8ccf3263a4d071d946b