Pith. sign in

Paper Citation Record · LEDGER

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications

As of 19 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2607.06411.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06411 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:21:50.175416Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 25153bba-7766-48ff-8da2-350b24c0482f · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:49.708376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:49.708376Z digest=sha256:1ecbe0070714dea357cb14a7cf999ec58c6c74cb70bc9458da92190b04f96016

Observation 7f097b81-7b4e-4618-9fa7-f37081d6ea79 · outbound

This paper cites SWE-Bench+: Enhanced Coding Benchmark for LLMs.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications SWE-Bench+: Enhanced Coding Benchmark for LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:49.775844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:49.775844Z digest=sha256:61eeddd7b6285f5823db44f8c93055a2256829f170c01cbdd041bfab74866a2e

Observation 20828afc-0fec-467e-9933-0bd1187ff184 · outbound

This paper cites Liang, S.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Liang, S

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:49.976954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:49.976954Z digest=sha256:e87324f4bcc7a9abf210aaa8c37ffe5cc30459197af99ad869e38b1bf236ad50

Observation c2414804-da3f-4c9a-a437-59199db19a10 · outbound

This paper cites Prathifkumar, N.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Prathifkumar, N

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.107696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.107696Z digest=sha256:293ff554c98da030f5f041b167e8b73582fb1c44cf2603a2088e2fe0202bd32b

Observation 49f5b076-22b7-43ab-a7d4-47a40943f5e4 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.110905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.110905Z digest=sha256:0e82b91ac35515ca76f4790ff3f4b55be1139685d75928bf6b14e407363daf32

Observation 67a67033-2161-4465-93c2-71fcfe096b3f · outbound

This paper cites Badertdinov, A.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Badertdinov, A

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.113607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.113607Z digest=sha256:3bd2fe4892285b8f47916d3b47153f9b58e6b65d555c13dcb05fcaa805a1d070

Observation ca14c0ea-849c-4e29-b863-c20cd752f562 · outbound

This paper cites SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.116198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.116198Z digest=sha256:9d133e461f4bde5d160060193702d85a77d9a5fc4e3fbe76875239507d92fe0a

Observation ba49c9f0-5eed-44f6-afb9-69243e087812 · outbound

This paper cites Chervyakov, A.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Chervyakov, A

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.119591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.119591Z digest=sha256:931665bf9216516c408c4bad122abc7b0e785c9020659e9a3590f8a0f77335cc

Observation 1bb4a7d5-bffc-4c28-9bb3-313c05f67bdb · outbound

This paper cites an unresolved cited work.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.122122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.122122Z digest=sha256:6972b52e67917b6315d080fa18f483e6775c1bde1b16f12d51847a29bb7b4cf3

Observation 398a7918-8eb2-47f4-bcc1-6394e362bd83 · outbound

This paper cites Chinese Language Is Not More Efficient Than English in Vibe Coding: A Preliminary Study on Token Cost and Problem-Solving Rate.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Chinese Language Is Not More Efficient Than English in Vibe Coding: A Preliminary Study on Token Cost and Problem-Solving Rate

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.124329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.124329Z digest=sha256:694e3248fd9e82cf835f53209576f29c90f6940a02aa87cc4937eb0c7f917f23

Observation ffc6dabd-4f26-488a-8066-685959e86027 · outbound

This paper cites Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.127140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.127140Z digest=sha256:1981e5db998b447fcb297e2519d56c632fc8cb228ae9491cec6704bd6472ad90

Observation 1327ca94-c8d9-4bf7-a999-1f1c80b399dd · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.129688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.129688Z digest=sha256:e68cca82196b01a87d96e3e7188bbc969acd3152129b73431612795250c05ade

Observation 448f7704-aa64-4cd9-b2ed-553bc86daad9 · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications GAIA: a benchmark for General AI Assistants

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.132378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.132378Z digest=sha256:ef50d156888942088231df70b8b32aebf8717d3d0e833847404a8e1eaa95ed7a

Observation 961877a6-14a1-4ed2-aa27-9d8b3536c034 · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.134928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.134928Z digest=sha256:2a1d7d9197ff28eb07ed561e12af75bb030664bd698d2fc35a796ee16e7c4fef

Observation 666f9baf-c066-44ac-af9a-50f1e1eb8987 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Evaluating Large Language Models Trained on Code

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.137021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.137021Z digest=sha256:0c50819007ef6867c6824ddc2d52492c48f1d1145fa5e7a1b398c71825177161

Observation dbf58742-749a-4db9-b16e-c2eb66ffed38 · outbound

This paper cites Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.139165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.139165Z digest=sha256:68dbc016800e3e2437824adf7937b3a56fedc64560b3b2a861ffa786fb5f4185

Observation 3b95bbdd-6e60-4b0f-9f58-9c40fb52cc0f · outbound

This paper cites AI Agents That Matter.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications AI Agents That Matter

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.141580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.141580Z digest=sha256:b2183285539fe9bffd9e9bbea9d72524ecd36096dce508ff3cf67e8b3d7827c4

Observation fc84386f-7bd2-4ece-b298-087048b31f8a · outbound

This paper cites Kapoor, B.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Kapoor, B

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.143937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.143937Z digest=sha256:e6fe4c80dc719aa948118118e71fc2478a5845777ac6226c63be6ab2eae88347

Observation 087ac713-6a1f-45db-ba43-78014aef5628 · outbound

This paper cites Measuring AI Ability to Complete Long Software Tasks.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Measuring AI Ability to Complete Long Software Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.145877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.145877Z digest=sha256:843323802e4eef445fc83aaaa0b0b9dfa52fbd45f7df7b190e904ea6ec5a7c41

Observation 58fcc192-7e31-4da4-944f-af8b18859fb3 · outbound

This paper cites RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.147998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.147998Z digest=sha256:4649f4d91cdacaa00f99b8d9dcc3e5df4ca2eaa2f8d9405c571bd85b8a9511da

Observation dc4bb9bc-feab-490c-b45e-35719505f9ec · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.150187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.150187Z digest=sha256:a83118e750e94eda43ca4db9ac745236fd1a354d16c5ff5ccdbfe15292aee46f

Observation 220659b0-5cb8-4906-bd48-13c981ac6111 · outbound

This paper cites Matching Matters: A Fair Quality-Efficiency Benchmark for Command-Line Agents.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Matching Matters: A Fair Quality-Efficiency Benchmark for Command-Line Agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.152279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.152279Z digest=sha256:a3890a1cbef1fb115b4e87ad0b65591098194453b5b2ac3b8878485e45f7eb64

Observation 036e3ce2-2e83-4ef4-820a-aec381f84720 · outbound

This paper cites ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.154596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.154596Z digest=sha256:a5226a1bcdeb2fbc43c3869a979178e7280a6c457fe984fb51cf41b4586f08f9

Observation 1e3db29c-191f-47ba-8277-1711100b9253 · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.156961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.156961Z digest=sha256:d1d47c452d99ac4c3b22da7aea9afab8379a78684b1cc23a9f73c9157b0f0ec2

Observation 3877d2cc-b2d2-4b87-acf0-9b9500240cf6 · outbound

This paper cites The Leaderboard Illusion.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications The Leaderboard Illusion

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.159277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.159277Z digest=sha256:af46d6bc0e1e60c66d4506cd391d73b654d831d358f25a139e91bc97668b2a5d

Observation 9fcad652-77e6-4baf-95bc-178d169639c5 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.161820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.161820Z digest=sha256:b73cd68ed068162e3517ff728b886953ed64e9841c7f9ec6ab2ff221fdcb9b74

Observation d9069d21-4c4a-469a-ad8c-0bd1f78a22aa · outbound

This paper cites MCoNaLa: A Benchmark for Code Generation from Multiple Natural Languages.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications MCoNaLa: A Benchmark for Code Generation from Multiple Natural Languages

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.163874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.163874Z digest=sha256:0399d66b9cddab941fcc0d9199ab7128e81843edc1c73d400c4b2391a8ce3895

Observation aa951d9d-e589-41f3-9cc1-f482a6729cc1 · outbound

This paper cites Execution-Based Evaluation for Open-Domain Code Generation.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Execution-Based Evaluation for Open-Domain Code Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.166018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.166018Z digest=sha256:790546fa59ddfbff8eb665a2cfa080f9c68ae0f664d780c91b5c0a2c9ca5db80

Observation 05b12bbe-234b-4b37-b526-95b20b9c02b4 · outbound

This paper cites HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.168183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.168183Z digest=sha256:faef8dca5f76bfd11455a4a18c301daa58fef318e9f74a8fcfd36542984f2cc6

Observation 1e75f657-3fea-452f-a6db-d3c2d022cfea · outbound

This paper cites Если FSM-контекст для апдейта недоступен — сценная обвязка спокойно пропускает апдейт дальше, а не падает.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Если FSM-контекст для апдейта недоступен — сценная обвязка спокойно пропускает апдейт дальше, а не падает

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.171250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.171250Z digest=sha256:7f3e55c80b10350019fdc9d10dd867262ad631bef182b343d816e286bd397851

Observation 7413c14b-3663-49a1-8334-82ff3405278a · outbound

This paper cites Для стратегий FSM, которые и так работают в разрезе чата (CHAT и CHAT_TOPIC), контекст должен определяться и без пользователя — по самому ча- ту/каналу.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Для стратегий FSM, которые и так работают в разрезе чата (CHAT и CHAT_TOPIC), контекст должен определяться и без пользователя — по самому ча- ту/каналу

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.173335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.173335Z digest=sha256:207ac8e2f076e003c1392592b480dae2066bd2746ed67d335209dbaa7e541812

Observation 8adc0628-f9d6-4f33-b356-ab6f51a01607 · outbound

This paper cites Существующее поведение для личек и групп ломать нельзя: боты со сценами и без, которые сейчас работают, должны работать ровно как раньше.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Существующее поведение для личек и групп ломать нельзя: боты со сценами и без, которые сейчас работают, должны работать ровно как раньше

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.175416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.175416Z digest=sha256:7fa602e5df99a920b3fdc3b1c40d083db8b61071bce5c499c2cdb43812ca5283

Pith citing papers

No inbound Pith citation observations are available.