Pith. sign in

REVIEW 4 cited by

Airavata: Introducing Hindi Instruction-tuned LLM

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.15006 v2 pith:RIQB5KFF submitted 2024-01-26 cs.CL cs.AI

classification cs.CLcs.AI
keywords airavatahindidatasetsdiverseindicinstruction-tunedinstruction-tuningtasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We announce the initial release of "Airavata," an instruction-tuned LLM for Hindi. Airavata was created by fine-tuning OpenHathi with diverse, instruction-tuning Hindi datasets to make it better suited for assistive tasks. Along with the model, we also share the IndicInstruct dataset, which is a collection of diverse instruction-tuning datasets to enable further research for Indic LLMs. Additionally, we present evaluation benchmarks and a framework for assessing LLM performance across tasks in Hindi. Currently, Airavata supports Hindi, but we plan to expand this to all 22 scheduled Indic languages. You can access all artifacts at https://ai4bharat.github.io/airavata.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Faster Machine Translation Ensembling with Reinforcement Learning and Competitive Correction

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A DQN-based candidate selection and a competitive correction block improve MT ensembling quality while reducing inference cost on English-Hindi and Hindi-English tasks.

  2. Krutrim LLM: Multilingual Foundational Model for over a Billion People

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Krutrim LLM is a 7B parameter multilingual model trained on 2T tokens with the claimed largest Indic corpus, reporting strong Indic benchmarks and English scores near Llama-2.

  3. IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding

    cs.CL 2025-01 conditional novelty 4.0 of 10

    A machine-translated version of MMLU-Pro in nine Indic languages is released as a benchmark, with baseline accuracy scores for multilingual LLMs.

  4. Analysis of Indic Language Capabilities in LLMs

    cs.CL 2025-01 conditional novelty 4.0 of 10

    A desk-research review finds that LLM performance is strongest for Hindi, Bengali, Marathi, Telugu, and Tamil, and recommends prioritizing these five languages for safety benchmarks.

Pith tools