Pith. sign in

Paper Citation Record · LEDGER

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection

As of 19 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2509.06524.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06524 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:33:43.787531Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact2
  • verified fuzzy24
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 36611695-46bc-4fbd-84b8-32bb852895e0 · outbound

This paper cites write newline.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:38.928351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:38.928351Z digest=sha256:dadc0c1afade19605e8386cba61489af1b9ed0f822859e50b62cd426aebfea9b

Observation b8a1c987-f9d3-4830-97bd-5cac28f49bee · outbound

This paper cites Program Synthesis with Large Language Models.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Program Synthesis with Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:39.032615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:39.032615Z digest=sha256:02604fa3141b9653be4c5919034681d3d8a72775e14cf0c4a5c18c3f4c534f55

Observation 1f02c21e-0bd0-43f5-8ed4-61aa24933c99 · outbound

This paper cites Color-filter: Conditional loss reduction filtering for targeted language model pre-training.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Color-filter: Conditional loss reduction filtering for targeted language model pre-training

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:47.896802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:39.142597Z digest=sha256:68a96bd2b769e992963deeb776cfe4ca47b3a665b1228ad292b8fd740fa0fdb3

Observation 7f3a2181-16a7-4984-a9b9-c71cb163f229 · outbound

This paper cites Statistical inference.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Statistical inference

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:39.203281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:39.203281Z digest=sha256:295d009336e4467557007b7eb0bb5b320561bc456279b30f09f44aa52362da3f

Observation 31f8ba62-0d74-4f2a-93fa-901086584fb9 · outbound

This paper cites Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:39.307095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:39.307095Z digest=sha256:d3af6cf18b0218f04089b8bb4529dc1367d78728ab22ab1cfd5388ae8cc14e40

Observation d394646a-6661-4c17-8887-8e4894c4c885 · outbound

This paper cites an unresolved cited work.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:39.420873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:39.420873Z digest=sha256:4e654a144c3693341895c7fbb2704f8668cf6d1ad2c6a4c4f28deeffeacf9cd0

Observation 704d50ee-96cd-4719-aea2-56c47465eb94 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:39.555558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:39.555558Z digest=sha256:d75aba49a7c8b425ca45390c042d7b01f723814515d18abe2bde94da707c24be

Observation 287cdd5e-d610-4783-8cf1-ef9e6da14912 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:39.625275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:39.625275Z digest=sha256:6d1510a1435b746f5182e3ca332b2ac8759b3fcd55a74d5a175ba2153e429c26

Observation 2ce12f0a-0236-4e3e-919b-87c55dcfe2ae · outbound

This paper cites Sketchy moment matching: Toward fast and provable data selection for finetuning.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Sketchy moment matching: Toward fast and provable data selection for finetuning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:47.785803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:39.744619Z digest=sha256:171d1a2b05ba6b17b3bd29ff8a58f7105ab1ed05669b124e2b740ad7cc161c0f

Observation 14b74413-3450-492a-a832-84dcac04a4c6 · outbound

This paper cites Dsdm: Model-aware dataset selection with datamodels.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Dsdm: Model-aware dataset selection with datamodels

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:47.679914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:39.805459Z digest=sha256:2379268c2a05dc6bef0de9baf1914a8849eae11fb60d04434bd36e965b06d253

Observation b41baaee-b249-45ce-80bb-435e46fd4c17 · outbound

This paper cites Dsdm: Model-aware dataset selection with datamodels.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Dsdm: Model-aware dataset selection with datamodels

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:47.552421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:39.929141Z digest=sha256:ae512b671bb432500ab77dc60aa650a8bf3e38bc7c229f0c558c129585660140

Observation 9e408eab-4345-4f8c-ba54-82c54f1b0b66 · outbound

This paper cites Evans, Gordon B.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Evans, Gordon B

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:47.419687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:40.029161Z digest=sha256:976af80c2f7a28599adb3fc4ca3eabf3719b0073c1dd9d673c0539fa17367023

Observation 7768ed40-0b9f-49ac-a8a0-5c6e23406c02 · outbound

This paper cites GIO : Gradient information optimization for training dataset selection.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection GIO : Gradient information optimization for training dataset selection

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:47.268358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:40.125862Z digest=sha256:2fb4db79c7163b3bf47ff9ca3f4a72d91e7a8f90410ff9b05bf4d6dba6279023

Observation 76aea87e-84af-4f9d-90f8-6dbdca0f4488 · outbound

This paper cites CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:40.212176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:40.212176Z digest=sha256:ace17b70cc6058a75ab3402ac91480673be63f6aaf8a18f58bce956861ecd840

Observation cafc6fa2-8a3c-48f4-a098-066c6a9741c4 · outbound

This paper cites Data selection via optimal control for language models.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Data selection via optimal control for language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:47.145456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:40.321974Z digest=sha256:bac72b459655f7b4f3436c2b0f2d3c59805d2e0602a0f9931534043f59d36732

Observation 46d76be7-5a67-48c0-98a5-fc5bab34c857 · outbound

This paper cites SHED : Shapley-based automated dataset refinement for instruction fine-tuning.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection SHED : Shapley-based automated dataset refinement for instruction fine-tuning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:47.000442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:40.435948Z digest=sha256:573cb00b9714eda5041b0eddffccbddc5094c260dcbac9a58811492e2723b4a2

Observation f797b552-f111-48b8-87c1-5ee5e7140c79 · outbound

This paper cites Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:40.574514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:40.574514Z digest=sha256:d4a5c26e791ffeb3e468657dc2604c3ed8a53a844c45cda031f7bb450a697f25

Observation f75f2856-375b-4f60-9c50-be598e2a3354 · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:46.868735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:40.678681Z digest=sha256:a30939798f442a4df06e47705807cfe611aabb459f9f069c6302caae5a4c6123

Observation b51f0d1c-0cd2-4268-b5a4-3e4c2f94796f · outbound

This paper cites Scaling Laws for Neural Language Models.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Scaling Laws for Neural Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:40.766727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:40.766727Z digest=sha256:d2fa3ec4a7fc22edaf36008d99c1aa070b18780ca1ba3bc0562e591aae22f791

Observation 64f03e7b-26a4-4318-a00c-338e4acd4453 · outbound

This paper cites Rule-based data selection for large language models.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Rule-based data selection for large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:40.848381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:40.848381Z digest=sha256:c0da00d0e2c0279c85ffdf886b5deeb2a198c9de65343388309ef81b26a54264

Observation e6b10f93-d954-4d7b-b8c1-9a101d5fd801 · outbound

This paper cites One shot learning as instruction data prospector for large language models.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection One shot learning as instruction data prospector for large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:46.794695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:40.963618Z digest=sha256:5d909fd157b44ec84e3c08f0b90e8fd89340463372b44b830be62eea05e3201c

Observation c561e07e-59f4-4c65-a8cd-535feab56a53 · outbound

This paper cites D 2 LLM : Decomposed and distilled large language models for semantic search.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection D 2 LLM : Decomposed and distilled large language models for semantic search

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:41.043044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:41.043044Z digest=sha256:b6cecb79df9fa810f38c8d6ed57cf42e63df14834a918884098e50ee09be66bb

Observation bbd39cec-346f-458f-8e51-2b755225f82e · outbound

This paper cites Let's Verify Step by Step.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Let's Verify Step by Step

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:41.158139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:41.158139Z digest=sha256:b7509bd197cc063c036f11adf0edeb1ba6e63ae01a893598f547e93910880ad3

Observation b9cfc409-4cc6-4d3b-be00-7e8eea0a4f66 · outbound

This paper cites What makes good data for alignment? a comprehensive study of automatic data selection in instruction tuning.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection What makes good data for alignment? a comprehensive study of automatic data selection in instruction tuning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:46.617318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:41.278710Z digest=sha256:f58302a9d1ca75e6a6d97a7eca637ba3139f9a76ed0a88732f5026e451e35869

Observation 033ccbfb-96c2-4282-93a0-05f3e839a323 · outbound

This paper cites Take the essence and discard the dross: A Rethinking on Data Selection for Fine-Tuning Large Language Models.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Take the essence and discard the dross: A Rethinking on Data Selection for Fine-Tuning Large Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-04T23:33:44.238868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:41.391669Z digest=sha256:74971d70377af0bad8016a3d1343e0442872ebb23503c3adcc2b8ac98f92a293

Observation 0fab0a4b-2e75-415d-83e8-751c43aebed6 · outbound

This paper cites TSDS : Data selection for task-specific model finetuning.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection TSDS : Data selection for task-specific model finetuning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:46.481496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:41.467553Z digest=sha256:8fc04af3f6aa8721f74ade99990a517dfb5fa7b67b266ed07c5379918f13e1bb

Observation 5af94284-6847-4d3b-a7ab-6e956edfb3e9 · outbound

This paper cites When less is more: Investigating data pruning for pretraining llms at scale.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection When less is more: Investigating data pruning for pretraining llms at scale

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:46.335102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:41.587282Z digest=sha256:1fa172459f5521b5fa41bdc580c999b12a4f5b473a72aa7e90e066ce28e928e6

Observation 5e28398d-e475-4374-a303-2403924068b6 · outbound

This paper cites SGPT: GPT Sentence Embeddings for Semantic Search.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection SGPT: GPT Sentence Embeddings for Semantic Search

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:41.675399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:41.675399Z digest=sha256:f3801fb814ca4aa0e12d96510acbecff33ea2c2c016917883961714b70d3986e

Observation 68f05c8d-a0b9-43a0-9d3e-a26c6679a32c · outbound

This paper cites Scaling data-constrained language models.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Scaling data-constrained language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:46.156191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:41.773607Z digest=sha256:702cf06fd85311fecea8c41895d8ff26c03b5e4f98cd228164de9e5e8d27cab9

Observation 9631d39c-eeeb-4309-90f6-e80fb02a35b2 · outbound

This paper cites Trak: Attributing model behavior at scale.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Trak: Attributing model behavior at scale

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:46.032466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:41.873846Z digest=sha256:b7b13ac2fbfae6b773178892a94e589dc94f93d44448e61549dcdf1da8c08793

Observation acc2dabf-d3ea-449f-9358-ffe2e6ef2a66 · outbound

This paper cites O1 Replication Journey: A Strategic Progress Report -- Part 1.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection O1 Replication Journey: A Strategic Progress Report -- Part 1

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:42.000535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:42.000535Z digest=sha256:613ae4f747364894f35b59c81b167f3bafe75dc90ea38ea4aeee08d0a2d67851

Observation 7f810919-8077-4187-bde1-1256f67ea180 · outbound

This paper cites Learning to retrieve prompts for in-context learning.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Learning to retrieve prompts for in-context learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:42.093912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:42.093912Z digest=sha256:e7c20ed42f284f4f3c2b404423d2d424f938751de4a64486f20bba3a02168459

Observation 3a86b5a3-61a3-4750-a132-2cb73815b823 · outbound

This paper cites How to Train Data-Efficient LLMs.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection How to Train Data-Efficient LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:42.206386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:42.206386Z digest=sha256:388c469cafb07282563aee681aa4f4613bb43719797297ebef1bafd826583e8d

Observation 2796cf24-d707-4247-87ea-8d61144007e7 · outbound

This paper cites Improving dense retrieval models with llm augmented data for dataset search.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Improving dense retrieval models with llm augmented data for dataset search

Reference 34

Resolution
metadata mismatch
raw_fallback, observed 2026-08-04T23:33:44.112021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:42.268814Z digest=sha256:ef23222948c9262bca5ce587dbbeed435141cc68a8decf8f811049c31db637c7

Observation 46a246d7-5e3a-422f-9a78-9df8a2838166 · outbound

This paper cites Improving pretraining data using perplexity correlations.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Improving pretraining data using perplexity correlations

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:45.862471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:42.396817Z digest=sha256:ba41fc075e5665d54f54da3f5ef0713243accac216de89d3de0a0db6c87d045e

Observation fa571434-966d-4901-b7b2-d3dd0024c090 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection LLaMA: Open and Efficient Foundation Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:42.513074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:42.513074Z digest=sha256:f71ec0b4a6eec748ad0ca1f5172a7fb7d399e0a240f38804a90223583ab608ee

Observation f0f48a32-4456-4e6a-b220-8ef69370507d · outbound

This paper cites How do your code LLM s perform? empowering code instruction tuning with really good data.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection How do your code LLM s perform? empowering code instruction tuning with really good data

Reference 37

Resolution
verified exact
doi, observed 2026-08-04T23:33:45.737973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:42.576550Z digest=sha256:c9e4007714785295baec81317cb247efbba7120dd448923a2e86f38aa72a1887

Observation 4c1b800b-8ca5-4467-8ea0-988b5cf2a74d · outbound

This paper cites QuRating : Selecting high-quality data for training language models.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection QuRating : Selecting high-quality data for training language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:45.605862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:42.689126Z digest=sha256:ba1c037ada5338f00010b5dbb8e18482d7042efe7bb1b99df3c4c0964bb2b561

Observation cb22c5e6-6e93-49f8-998b-a50e84f659d8 · outbound

This paper cites LESS : Selecting influential data for targeted instruction tuning.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection LESS : Selecting influential data for targeted instruction tuning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:45.464921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:42.789506Z digest=sha256:5c56726afd593e44470c3b9800fd3f773ec025ba52f1efb9c6ac4bb157029b08

Observation d58a49f6-035b-4227-b3bd-5532c68c0beb · outbound

This paper cites Data selection for language models via importance resampling.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Data selection for language models via importance resampling

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:45.284526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:42.880975Z digest=sha256:a77c7ed54ba21303518deeb79b0c62f9670cd537a7a9e05a0b8ae6832f325b51

Observation f481929d-97f7-4e8b-9e38-5883038d6f20 · outbound

This paper cites Wizard LM : Empowering large pre-trained language models to follow complex instructions.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Wizard LM : Empowering large pre-trained language models to follow complex instructions

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:45.103528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:42.976695Z digest=sha256:97be0d756fe10342f204d076bf9b2ccde0d69dda7660fb994cf7ecd689f48b9e

Observation 46a96f41-d786-4306-8a1b-96158de5e3c9 · outbound

This paper cites Diffusion models: A comprehensive survey of methods and applications.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Diffusion models: A comprehensive survey of methods and applications

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:44.964591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:43.091247Z digest=sha256:879c50cb7ccefe9a7af20c8c7e27d01c9d99d3e68ea451cf7ae11c0ed07812e0

Observation 2a0ed0ab-e21f-40d8-a1c9-3fc0cb07deef · outbound

This paper cites LIMO: Less is More for Reasoning.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection LIMO: Less is More for Reasoning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:43.168667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:43.168667Z digest=sha256:90f24c09815344e680da0fe8b80cb0586986886ac7bc69de1f32390e76760805

Observation 101acc96-721f-4334-ba95-5ca3454b2b14 · outbound

This paper cites Mates: Model-aware data selection for efficient pretraining with data influence models.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Mates: Model-aware data selection for efficient pretraining with data influence models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:44.836188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:43.283879Z digest=sha256:153e0f129361c2cf1bd78144f37abac74a19c86ba7deeeb8aa6a625d0e5a4a02

Observation c9202663-b1d7-44aa-a5fc-987e70b7257a · outbound

This paper cites Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:43.405092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:43.405092Z digest=sha256:ee54aafb3c2ff5367e88581b71b189e6499c5762a84e1f05f5a6b01500813757

Observation c5e85840-c6e5-4cad-959e-3d8b7ab4e27a · outbound

This paper cites Beyond similarity: A gradient-based graph method for instruction tuning data selection.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Beyond similarity: A gradient-based graph method for instruction tuning data selection

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:44.679324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:43.483174Z digest=sha256:04cd5d23c8a796d454cbffecc008ce16b5f1383cc4d109dc167aa3b886c383ea

Observation d5965d9f-92de-4616-bf9d-b506567a1a1e · outbound

This paper cites OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:43.578607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:43.578607Z digest=sha256:d31b62f859824a2e8f0a55db66b9e1904a9a5b4b2efa7f57b9b7e173e4426bb3

Observation f21446f6-c20a-4297-929b-296eea4bcc5a · outbound

This paper cites LIMA : Less is more for alignment.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection LIMA : Less is more for alignment

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:33:44.510160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-04T23:33:43.635829Z digest=sha256:fbd3249e19eef64ba8632c5a78e2fc4ec4628948d33b11ec19f5cb84ca238089

Observation aa2630f7-9d58-403b-a007-d0b16e3eb506 · outbound

This paper cites @esa (Ref.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection @esa (Ref

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:43.685397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:43.685397Z digest=sha256:e6e808aea2d3117958bc4c0d34e4134e04defa7a02b629dc5a7ded66473f949f

Observation 973bf3af-5dd0-49ab-99f3-8b02a3d4ca37 · outbound

This paper cites an unresolved cited work.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:43.748037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:43.748037Z digest=sha256:ddd0ad8dfe27d59d9ee92057580cae2797870c57106955afe6fc6d4c1ca5e8de

Observation 8358d436-b2da-4c60-8d74-37ec1cbb45bc · outbound

This paper cites an unresolved cited work.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:43.787531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:43.787531Z digest=sha256:4670bfcc8b16ab396e00296fea4799c59d9f1db32acaec5f76f5bf1b3bafca85

Pith citing papers

No inbound Pith citation observations are available.