Pith. sign in

Paper Citation Record · LEDGER

Decision-Focused On-Policy Learning for Contextual Linear Optimization with Partial Feedback

As of 10 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2606.01081.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.01081 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T17:49:58.484710Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 23d07746-e462-4689-94c4-33a6fe27e6a9 · outbound

This paper cites an unresolved cited work.

Decision-Focused On-Policy Learning for Contextual Linear Optimization with Partial Feedback Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T17:49:58.484710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:49:58.484710Z digest=sha256:33d33adc5d0e336695d42b7e2745489796a5592d2184883d5d7d315aec33d4e1

Observation 7dbad068-a05e-4407-827d-57adf908cc41 · outbound

This paper cites predict, then optimize.

Decision-Focused On-Policy Learning for Contextual Linear Optimization with Partial Feedback predict, then optimize

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:52:27.362704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T17:49:58.484710Z digest=sha256:28b20b553f1a5a8cdfda0f5afce39179b4ff2a1f6b6c2fe694c6146e941c15ed

Observation 5f93699a-b264-4f2d-9468-2f3a5e4e89eb · outbound

This paper cites an unresolved cited work.

Decision-Focused On-Policy Learning for Contextual Linear Optimization with Partial Feedback Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T17:49:58.484710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:49:58.484710Z digest=sha256:77d0cb8c0593be980251e4213263f93663f664dee4d7d8235d6a91a5effcb32c

Observation 0f9ea7ff-0009-4df9-8a73-86cba540bcb0 · outbound

This paper cites Applying the triangle inequality,∥w ⋆(ˆc)∥ ≤BS, and (12) yields ∥∇θw⋆ θ(x)− ∇ θw⋆ θ′(x)∥ ≤B S Z ∥∇θpθ(ˆc|x)− ∇ θpθ′(ˆc|x)∥dˆc≤B S M∇ ∥θ−θ ′∥.

Decision-Focused On-Policy Learning for Contextual Linear Optimization with Partial Feedback Applying the triangle inequality,∥w ⋆(ˆc)∥ ≤BS, and (12) yields ∥∇θw⋆ θ(x)− ∇ θw⋆ θ′(x)∥ ≤B S Z ∥∇θpθ(ˆc|x)− ∇ θpθ′(ˆc|x)∥dˆc≤B S M∇ ∥θ−θ ′∥

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T17:49:58.484710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:49:58.484710Z digest=sha256:bb8e20ae1e6dd51004aea399e539f7b5363a270378c8646d929006f147121c10

Observation 7523f714-ed8a-4da0-a637-4dbe55a06f1c · outbound

This paper cites At the beginning of period t, the context is drawn as xt ∼ N(0, I p), with default p= 25.

Decision-Focused On-Policy Learning for Contextual Linear Optimization with Partial Feedback At the beginning of period t, the context is drawn as xt ∼ N(0, I p), with default p= 25

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T17:49:58.484710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:49:58.484710Z digest=sha256:e53d4d90d7daf71e7e070d22785102f3feefd1c4eca9f1a42b27b4cc4a648f88

Observation ea3b30ac-1ccf-4324-a9eb-f51c10c8ac23 · outbound

This paper cites Algorithm 2:Greedy contextual bandit (GREEDYCB) Input:Initial parametersθ 1 ∈R dθ; stepsizeη t >0 fort= 1,2,.

Decision-Focused On-Policy Learning for Contextual Linear Optimization with Partial Feedback Algorithm 2:Greedy contextual bandit (GREEDYCB) Input:Initial parametersθ 1 ∈R dθ; stepsizeη t >0 fort= 1,2,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T17:49:58.484710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:49:58.484710Z digest=sha256:7aaeb30222abd84953e554b43c268b45fa3b1bf6b96ef73123cd38ccfa148967

Observation f57c3e71-037d-4593-9295-5723987dcddb · outbound

This paper cites Figure 4 reports the same comparison on the three remaining benchmarks: top-k selection, shortest path, and energy scheduling.

Decision-Focused On-Policy Learning for Contextual Linear Optimization with Partial Feedback Figure 4 reports the same comparison on the three remaining benchmarks: top-k selection, shortest path, and energy scheduling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T17:49:58.484710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:49:58.484710Z digest=sha256:a738a740e0ec732820c6c1527494d8cbe974e2621dcd32ce4d2164805313d8a4

Observation 0c66c33e-a458-4317-bd7c-a8a604fc3a0c · outbound

This paper cites an unresolved cited work.

Decision-Focused On-Policy Learning for Contextual Linear Optimization with Partial Feedback Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T17:49:58.484710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:49:58.484710Z digest=sha256:5400c440722512d75fa7219049160cf16348313a71b4746cacde86979324133f

Observation a38bccb6-9b4d-4f17-87ec-6ae1500a5354 · outbound

This paper cites an unresolved cited work.

Decision-Focused On-Policy Learning for Contextual Linear Optimization with Partial Feedback Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T17:49:58.484710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:49:58.484710Z digest=sha256:a0a1e0123f5e298d5180f8f6620aadea0d63490c57a334d5ef590e68c899381b

Observation 4b5f4a9c-b729-430d-a8af-2ad911bee0b2 · outbound

This paper cites The negative signs on the softmax arguments convert cost minimization to score maximization for the listwise ranking step.

Decision-Focused On-Policy Learning for Contextual Linear Optimization with Partial Feedback The negative signs on the softmax arguments convert cost minimization to score maximization for the listwise ranking step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T17:49:58.484710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:49:58.484710Z digest=sha256:a844abe7bd673b6d77e51d71eb6e64c0b802fbfc80da7a90689be55383eaa846

Observation 3624af3b-4922-40b4-9270-3bb0004cde83 · outbound

This paper cites an unresolved cited work.

Decision-Focused On-Policy Learning for Contextual Linear Optimization with Partial Feedback Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T17:49:58.484710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:49:58.484710Z digest=sha256:eff583fee846adff603408e5d7fc744665d6631c8aea97f0b97d8e853fa0ab83

Observation 2ee09b42-e634-4d28-84d8-6a44786fd1b6 · outbound

This paper cites On top- k the gap is more modest at low degrees (around 2× at deg = 2 ) and widens to roughly 4–5× at deg≥4.

Decision-Focused On-Policy Learning for Contextual Linear Optimization with Partial Feedback On top- k the gap is more modest at low degrees (around 2× at deg = 2 ) and widens to roughly 4–5× at deg≥4

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T17:49:58.484710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:49:58.484710Z digest=sha256:b619446f4befb6461b36245035ac8d9b5963cdfe6ad362fac912f56e7510cdaf

Pith citing papers

No inbound Pith citation observations are available.