Pith. sign in

Paper Citation Record · LEDGER

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL

As of 17 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 2 inbound Pith citation observations for arXiv:2506.05501.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05501 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:23:12.587345Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T05:12:44.268199Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:46:26.764949Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 96a77d7f-fa98-41a8-9e07-edf007569884 · outbound

This paper cites Qwen Technical Report.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.478364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.478364Z digest=sha256:83155a1be8b7736210652a62f418b85a99e74431d67b2e148d404a76395a383e

Observation 200d7dc9-908e-44b5-9d48-9045e77bda97 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.481259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.481259Z digest=sha256:683972329c8567c0111679f276aa1ceac0fa7cfb459afd2a83d7333d4bc8fcb5

Observation aa50303b-20ca-4208-8b1d-4821dec5f712 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.483699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.483699Z digest=sha256:8e6844dcab84c8f87a525738290c73b5f9e6bfd6fd629d5373b9c6f72536f6e0

Observation aa006ef1-21b6-45e0-857e-fbe5e29c86db · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:23:12.956200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T10:23:12.485922Z digest=sha256:6ad313edcad2d7c58e605a97d5b5022bb20d66b812bfdc265370650ff25f142f

Observation 5b7d0096-7748-4a23-9d69-4ac1d3a6712c · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.488547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.488547Z digest=sha256:c17a8984cf3ef2ea80d8cd8ba372819bfb4fa58d96c1c1dd65a959aa48961015

Observation 7ef8ac61-7805-45dd-938f-6a7e989dbefb · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.491082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.491082Z digest=sha256:f9142384e626a674d9a1f34890e1e1d70da53ef3dcc62b0dcf2d3bae032e239e

Observation 8078d485-d69e-4472-96ea-3820844c2e71 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.494017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.494017Z digest=sha256:2604ad6ab808c08a30a6b583d765980c8c5a1b8edfdf731f6e1aca2d5dfed338

Observation cb7208e1-e5af-4525-994d-e3bacfddec81 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.496519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.496519Z digest=sha256:4e778915e13dc08ad45f1f62919e7e9f733787dca313aa16397139e76aa92478

Observation d38f3067-beb7-4847-bf6f-4fa8939bcc51 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.498502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.498502Z digest=sha256:2a92bc9442ba31eb3546cf439dfbbd2dc260ecdd624971948e05b5b5a7ed8dc6

Observation f2ff5d2d-5d0c-45a8-b643-68d21febdb40 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.500546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.500546Z digest=sha256:9ec2cc8bfc0f635cba122f576d3eca612a82f607e725488ae35cecfcab9c059c

Observation d3487a38-dc62-4b86-bca4-cf93de2dbdb1 · outbound

This paper cites Making LLaMA SEE and Draw with SEED Tokenizer.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Making LLaMA SEE and Draw with SEED Tokenizer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.502691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.502691Z digest=sha256:adab50222069aaf7a7ff5bca57d9a724f429fa1d5f791f064113fd6ec582c1c8

Observation 43cabedf-93fb-4bc3-95f8-3054e75a5a97 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.504995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.504995Z digest=sha256:d5be5d1c3aefb75eb63febf68cb178eb4bc4a3d4b8e23a592ce1eedf6f22b34a

Observation 729a06c6-ce8a-4ff9-b2b8-40ac076a70a3 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:23:12.939100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T10:23:12.507148Z digest=sha256:f879d36bec709f06d66ccd772c23ab8f34eea3d5beaf883830ac4d860e365741

Observation 335c3dbc-5520-48c4-9bd2-b300b359f201 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.509070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.509070Z digest=sha256:fc811b5ddbbcc8772ad9d3868674bacbb16b0917b427627f174b159fbc572be9

Observation 96b75dea-f716-4ad6-8b6c-2b525eb43572 · outbound

This paper cites Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.511095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.511095Z digest=sha256:5465bc1381efe8ea4be19ac0673b08778a9fcf95ab6203a56220f4e366d3eb25

Observation e2855e21-d9a8-4b73-972d-110d7f90ec7f · outbound

This paper cites Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.513374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.513374Z digest=sha256:a1fb52faf4189b5bef898bb3fb4abc5f90c938359641462def345cc7f4e91a0f

Observation b80772bd-6c4d-4a8a-a5b8-b41a05f3c324 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.515817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.515817Z digest=sha256:edf4059bb5395f871cc7ab0ad384737875feee9ec6241e67a81a383c27f5dad6

Observation d3300954-a378-45ed-bfd4-cba46133dae8 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.518502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.518502Z digest=sha256:fa975b8ae284b8160a237d3862e23355bc0ed1d15ad138f6392179e8e06e7a01

Observation 38bfa069-c6de-4947-9281-2cf7158039a3 · outbound

This paper cites T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.520489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.520489Z digest=sha256:57667713d57ab111f0c0b0afe8a770121607f5607a71723099171be70272ba13

Observation ac842250-4125-431c-9591-e5c9565c852e · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.522483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.522483Z digest=sha256:41a9df90c8a89a242dacb4e2be9ce8ba44d793d676cf4c5dfd5fe289decde3a9

Observation 2702fb33-86bf-49bd-aac5-964a85abd772 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL LLaVA-OneVision: Easy Visual Task Transfer

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.524294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.524294Z digest=sha256:832a7f429fa6ad047c8cd877d340720d926ce2c592feeefc2a35f93d14292267

Observation a5d08ef2-8fb8-46da-aa02-8e3f1342f3db · outbound

This paper cites Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.526429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.526429Z digest=sha256:6980bcaf85dfdc18a3c9b86f76898ef90e0ce99ba9931499279bf08cd6326da2

Observation 103a01cb-0bb3-487f-acbc-489b357ee2d8 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.528420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.528420Z digest=sha256:9c904356fd392547122ee561f07c36e73f7467cb52f1cf296011b0e276d63295

Observation b0172b25-e168-43c2-a262-98253bd749ec · outbound

This paper cites Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.530267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.530267Z digest=sha256:27e0bae98b81b4480a7c260adc38e8e3cd864a8d1997ed5544dcaf3c2c94d520

Observation 3b20debb-3f5e-4a67-b117-b03d9d4b5172 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.532475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.532475Z digest=sha256:1b8d1466706b4467c638d9d8428e97988337140188af7754f8909567d3e728ea

Observation 9b68318a-909b-42d0-915e-5c207f2ee8c7 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.534533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.534533Z digest=sha256:b8c27fe48a05fc5fd6625aa01d7c5f6a77cb40efa825d8db8da7d7b0a2750fdc

Observation 9c5eef34-f8f8-4ee5-a798-8c6b7248206b · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.536598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.536598Z digest=sha256:94fc3552d2071eabd9c64bce9e286ebc1e4f8fd3ca53223ba546a09ae36a3ecb

Observation 9ffdc315-6711-481d-8abc-328255e8f321 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:23:12.913662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T10:23:12.538467Z digest=sha256:befca723f6f81de912bab0ab4f25b75599dd3fd431820537df634ca65cf8d66e

Observation 1d87d19a-e1ae-4303-9f62-5672c95aadb0 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:23:12.907359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T10:23:12.540420Z digest=sha256:13cdabb815051764f66e03aa837a066bc5615275cbfa6514a5ef71b4185e9579

Observation 7332d79e-57f8-4174-85d8-8b70c25bb5a2 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:23:12.901377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T10:23:12.542459Z digest=sha256:3fe76af28f97a931c84e57282b6756ea1b717ef8f07b78c8cebd5e7ef619c567

Observation 9b2e2319-32ed-461f-b9e9-6439c909aa14 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:23:12.895261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T10:23:12.544337Z digest=sha256:a5673b9f15cea9a2715d16e67e296222dbf3f5fff3e3314669e98224730c63a1

Observation 2774d770-059e-47a3-b1ad-0b65c07fcf89 · outbound

This paper cites Self-supervised Meta-Prompt Learning with Meta-Gradient Regularization for Few-shot Generalization.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Self-supervised Meta-Prompt Learning with Meta-Gradient Regularization for Few-shot Generalization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.546258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.546258Z digest=sha256:8df79029ca63d7e6085cb1bb6a3efafda5d99d34b5408969326d57bc1b789685

Observation ba6c5649-32ce-4b33-8167-eb10939e377a · outbound

This paper cites Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.548885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.548885Z digest=sha256:ae2c58beb72cec308f6fde3e1388f495f4a9121b52cc1aa2109952099cc2eb32

Observation dca5abc9-c343-408a-8b5e-7b94d32c7641 · outbound

This paper cites Auto-Encoding Morph-Tokens for Multimodal LLM.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Auto-Encoding Morph-Tokens for Multimodal LLM

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.551112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.551112Z digest=sha256:fc5d61b17759034541828dcd9f2181b7db6f6005a458d295ec4bdcfa631160d7

Observation 847cb11b-3ecf-415b-86f8-11836b42ae8c · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.553747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.553747Z digest=sha256:995784cb2112ca953da6ce5e8c67b146c4edfcbf449693ffc522dcd50475aee6

Observation c139af4a-b154-427e-9360-3ae217bb71dd · outbound

This paper cites STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.555613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.555613Z digest=sha256:cc16c7ca6215c3486f73de7b18a5a38af8f1d43a60e9cc0933f765ff7e1c18bf

Observation d5ab22ed-e6c7-4935-acc6-2f424e9e5885 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.558043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.558043Z digest=sha256:8191aa5f0bff08b9b352afefe49933893a04d0e9fb82b385f71926fb33313aeb

Observation 80576249-4419-4000-b45b-25bfa0c04620 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.560144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.560144Z digest=sha256:d982f06a21592c33b00577b8290a5938bdb9660ef39f472a226d6fb63357a1a0

Observation d0eaa143-c41d-4162-9181-63929e916336 · outbound

This paper cites Exploring Bias in over 100 Text-to-Image Generative Models.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Exploring Bias in over 100 Text-to-Image Generative Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.562389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.562389Z digest=sha256:f79c62859bd07fce4a720ad6d68c52d1e00da6e227056fc58181a5c2589d390d

Observation 4577579a-4ee3-4c6b-a3c7-d3abc8a12f2a · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Emu3: Next-Token Prediction is All You Need

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.564317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.564317Z digest=sha256:99bdde86df0f04d4f2d18b7132e51f68e0006d2393cf1b064ff42f637dcaff0e

Observation 6f51ef8b-170d-4fd7-9587-a084cd6aca56 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.566657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.566657Z digest=sha256:310eafce0fc3fa0dac3713858d158fa369ad45567e94bd9524e2d8485dd5f6fa

Observation 06cef3f4-e53e-4366-a48a-ab0590a4e052 · outbound

This paper cites SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.568772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.568772Z digest=sha256:48b5edc9b43349c4ea5d841e949e006f64d385bac5b61f074c65e49358c8bc6c

Observation cc0bb740-0c9c-45ab-97af-f5ac02e04aae · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.571060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.571060Z digest=sha256:b9c0803f8cb926982cb8850530e15c66acf761a6fdd5999f7dcc00f64266b844

Observation 002e6126-6d0f-4cb9-9fd6-a1e4c5f649ed · outbound

This paper cites SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.573432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.573432Z digest=sha256:9050896d32e767cd6b37f0f443b33ff5dcb7acc4e4506e5d27bb8df5d932792e

Observation f5bfb895-cb14-434c-94c6-f3ca527b0271 · outbound

This paper cites AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.575466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.575466Z digest=sha256:8e448a366581b90114311216977f97949a55db70436c7d1d6ecaa3b5e0a1a5a5

Observation b21326d8-46de-41c1-a09d-972fc86c1262 · outbound

This paper cites an unresolved cited work.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:23:12.888567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T10:23:12.578159Z digest=sha256:6d49f50104fff14139c4836ea486f94b68ca3837133b9c21b2a7d70bd2f0dfeb

Observation 806798e9-e5f9-4bf7-bf90-b96e1a959eaa · outbound

This paper cites Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:23:12.617976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T10:23:12.580211Z digest=sha256:4b6d31bdbd637ba4d21d77c9ae8e077721e9e24aee00c4a14e45ca1b574765ee

Observation f2de3076-eb6e-40b3-a718-a37c94f2152c · outbound

This paper cites VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.582550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.582550Z digest=sha256:b20d7f6746a42a44f6cf8eaf59375dc5d007286904ff09b6fd9c621fe187d8d6

Observation f1cca2d8-79c8-4148-a7e6-e175d8903d09 · outbound

This paper cites online" 'onlinestring :=.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL online" 'onlinestring :=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.584717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.584717Z digest=sha256:9ecdc5d2f2c8fdefbe402d315e670a3372bf07b45a84edfc720e9fdd88ed555a

Observation 92668bde-dfec-4198-a361-bb225f2e4309 · outbound

This paper cites write newline.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL write newline

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.587345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.587345Z digest=sha256:2f013e28ae0c50edffdb5b81c46196b0c536ee17eea04c29ff7cb78b760d5e4a

Pith citing papers

Observation c012480a-14d9-4fdf-8d47-e429a5d16559 · inbound

Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs cites this paper.

Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T05:12:44.268199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:12:44.268199Z digest=sha256:57a7e4c0ae53e88360a79bdad1c41b27e2a2ce8df26405e7d5ab093f21b78f4d

Observation 334b2604-7141-46d0-96bc-011d1b7f2b03 · inbound

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness cites this paper.

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:26.767451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T13:45:53.346402Z digest=sha256:bd2a6c4443d5a3a20e86993cedb723c7e8f6a7916eb58db03f0a96d6112c1fba