Pith. sign in

Paper Citation Record · LEDGER

Adapting Multilingual LLMs to Low-Resource Languages using Continued Pre-training and Synthetic Corpus

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2410.14815.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.14815 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T06:05:23.792357Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T18:04:09.703530Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8488089c-fb94-4c00-ad0e-5b9baf711201 · inbound

BgGPT 1.0: Extending English-centric LLMs to other languages cites this paper.

BgGPT 1.0: Extending English-centric LLMs to other languages Adapting Multilingual LLMs to Low-Resource Languages using Continued Pre-training and Synthetic Corpus

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T15:35:55.350490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:35:55.350490Z digest=sha256:5fe70ab071995d34da7080254133f80bdfabe04fc3768178a42b7ba8a6bda6cb

Observation d3fcca7c-9bd1-4b46-9439-fa5e102b20c8 · inbound

Sample-Efficient Language Model for Hinglish Conversational AI cites this paper.

Sample-Efficient Language Model for Hinglish Conversational AI Adapting Multilingual LLMs to Low-Resource Languages using Continued Pre-training and Synthetic Corpus

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T06:05:23.792357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T06:05:23.792357Z digest=sha256:9a4a2bf80a7d221f3aea3b75a5c5e80caeb6022d26cf0601195a4cdca516cc25

Observation 6a4e2f34-fa7a-43e2-88ab-bcc4174f669e · inbound

Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages cites this paper.

Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages Adapting Multilingual LLMs to Low-Resource Languages using Continued Pre-training and Synthetic Corpus

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T18:41:13.825977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:41:13.825977Z digest=sha256:574c95b4e111613150c10c7aa23e50663fd031bf96139b9c4a0ef4483a2370d9

Observation 61758689-51d3-4526-85a8-a1000599c8d8 · inbound

FormosanBench: Benchmarking Low-Resource Austronesian Languages in the Era of Large Language Models cites this paper.

FormosanBench: Benchmarking Low-Resource Austronesian Languages in the Era of Large Language Models Adapting Multilingual LLMs to Low-Resource Languages using Continued Pre-training and Synthetic Corpus

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:31:31.516463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:31:31.516463Z digest=sha256:e0d0565dd7806335165b7d1433ab526402bdaff11a5166a0fb5f7927e5d191d3

Observation 6d093ded-a63c-4d53-bf92-d470225f94f1 · inbound

Teaching a Language Model to Speak the Language of Tools cites this paper.

Teaching a Language Model to Speak the Language of Tools Adapting Multilingual LLMs to Low-Resource Languages using Continued Pre-training and Synthetic Corpus

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:45.406764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:48:45.406764Z digest=sha256:f36a7e90990a06c454da47d1efb90049b55c93ed13bcbb0cea73c5deb40e2160

Observation 0b486fc2-8896-4499-9c9c-47ce8ae31084 · inbound

PARAM-1 BharatGen 2.9B Model cites this paper.

PARAM-1 BharatGen 2.9B Model Adapting Multilingual LLMs to Low-Resource Languages using Continued Pre-training and Synthetic Corpus

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:04:01.967065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:04:01.967065Z digest=sha256:00d4ea17dbfb8e0b2dfb2db55afebc3904832c0ca6eabeb39c7b6eb51dcfbedb

Observation 5be86d29-97a4-4b1b-bf61-66cb84817111 · inbound

A Case Against Implicit Standards: Homophone Normalization in Machine Translation for Languages that use the Ge'ez Script cites this paper.

A Case Against Implicit Standards: Homophone Normalization in Machine Translation for Languages that use the Ge'ez Script Adapting Multilingual LLMs to Low-Resource Languages using Continued Pre-training and Synthetic Corpus

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:16.141138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:44:16.141138Z digest=sha256:de2c88f89ffad6e68f6913a846aa035de57c1801fe8130517807ed53c8724454

Observation ef2cb03c-07e0-4b09-a6ca-f8fe6c7ad425 · inbound

Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary Extension cites this paper.

Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary Extension Adapting Multilingual LLMs to Low-Resource Languages using Continued Pre-training and Synthetic Corpus

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T18:04:09.748618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T18:04:07.964375Z digest=sha256:17cedfbb47cebc0457a4e62cb940d2f5fd843d2cc2db53218d6075dbe9aa4595