Pith. sign in

Paper Citation Record · LEDGER

Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training

As of 17 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 4 inbound Pith citation observations for arXiv:2501.06658.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.06658 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:58:30.036313Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:15:28.739449Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T04:05:58.158813Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1f264e67-4753-4fe4-8278-6ca87417590e · outbound

This paper cites Chatbot interaction with artificial intelligence: human data augmentation with t5 and language transformer ensemble for text classification.

Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training Chatbot interaction with artificial intelligence: human data augmentation with t5 and language transformer ensemble for text classification

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:30.749445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T20:58:29.458758Z digest=sha256:dd588139ecea09c82ced1a57792ed14675d530bd7c0c82c9b87d72ec888ab54e

Observation 82e7386e-aca7-488a-affe-7e91b23cc2f7 · outbound

This paper cites an unresolved cited work.

Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:58:30.639667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T20:58:29.495813Z digest=sha256:da6f695f7d6b86b48de372edbf47662a9a617e834d18849aeee98e1834ec1fdf

Observation d5ad787d-caa9-43d5-b405-31a805dc592d · outbound

This paper cites Language Models are Few-Shot Learners.

Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:29.519913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:58:29.519913Z digest=sha256:71e11afeb04c1d2c44add7e221e05a80d15a6c67357e216dabedc9a23b10e8d0

Observation 33a6939c-fc9e-4c2a-9887-2f40b07e6707 · outbound

This paper cites PTT5: Pretraining and validating the T5 model on Brazilian Portuguese data.

Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training PTT5: Pretraining and validating the T5 model on Brazilian Portuguese data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:29.622345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:58:29.622345Z digest=sha256:1fd836277a29b678b44efcb28c87b1033de199cc8b46d4247cb7d72dc96562c3

Observation c19625db-c59f-4029-b76a-b2ce230c1228 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:29.675464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:58:29.675464Z digest=sha256:f0c1be030ba54853b0643bd7b226fc93e82823d4b675040bd5ef8ace8c4ff4b7

Observation d870744f-6dec-4e4b-b645-780a9e9bca4d · outbound

This paper cites A systematic review on machine learning models for online learning and examination systems.

Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training A systematic review on machine learning models for online learning and examination systems

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:30.623257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T20:58:29.680566Z digest=sha256:4ff2557717b23628f2bb00598ab14e47656ed516eefd5a056c2c0f7888191916

Observation 78cfabf8-9081-4f71-af78-b94fc09138ba · outbound

This paper cites Using Large Language Models to Assess Tutors' Performance in Reacting to Students Making Math Errors.

Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training Using Large Language Models to Assess Tutors' Performance in Reacting to Students Making Math Errors

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:58:30.100981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T20:58:29.685401Z digest=sha256:805794e0b496cdca84c42745b26889298e42f7a2c587213d375b121f3db4d652

Observation 73b9b660-10c5-4a95-8a1f-25a80311afa2 · outbound

This paper cites A survey of gpt-3 family large language models including chatgpt and gpt-4.

Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training A survey of gpt-3 family large language models including chatgpt and gpt-4

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:30.607334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T20:58:29.689945Z digest=sha256:eebd50442aef91675a2bfb22621fbed00750d80467531f438783d34d930748b4

Observation fca00ae9-53fa-48ed-8619-9886df56ea2d · outbound

This paper cites An improved aspect-category sentiment analysis model for text sentiment analysis based on roberta.

Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training An improved aspect-category sentiment analysis model for text sentiment analysis based on roberta

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:30.593657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T20:58:29.693957Z digest=sha256:2ca43035e63efbb3c4490f67d437ff001ac82bfa478de444511e892b2c7fd060

Observation 1e63a663-d6a1-49d1-a430-da7c94596f2b · outbound

This paper cites How can i get it right? using gpt to rephrase incorrect trainee responses.

Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training How can i get it right? using gpt to rephrase incorrect trainee responses

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:30.575617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T20:58:29.698107Z digest=sha256:e5575c04aebe784b8f53d703f44ad502b6ce61bf7e7e52ffc4958cd6683b0475

Observation c032cf83-59cd-454c-82f4-bfe332baa400 · outbound

This paper cites When the tutor becomes the student: Design and evaluation of efficient scenario-based lessons for tutors.

Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training When the tutor becomes the student: Design and evaluation of efficient scenario-based lessons for tutors

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:30.443424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T20:58:29.702728Z digest=sha256:4958e9f8b076c4f2deec455a60d77077145d8ccb5daa63c730ef7d73028eef54

Observation 205fccf9-e361-4279-8652-9ffb57ba90f3 · outbound

This paper cites Do Tutors Learn from Equity Training and Can Generative AI Assess It?.

Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training Do Tutors Learn from Equity Training and Can Generative AI Assess It?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:58:29.707015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:58:29.707015Z digest=sha256:b4117bf531332dc22c145ca05cdbe7155b3d79b124feaad48bd78c6d3933c1ea

Observation f53be87c-17be-4134-8095-dbf12c2a0288 · outbound

This paper cites Beyond the rubric: Classroom assessment tools and assessment practice.

Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training Beyond the rubric: Classroom assessment tools and assessment practice

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:30.343776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T20:58:29.761055Z digest=sha256:465582bfe27ffbbcd583890c537ebcc14dba13fe29cd69f4120715b69f69291c

Observation bbd248e6-9de7-4a3f-b1b6-f167885eeb6e · outbound

This paper cites Five ways to look at cohen's kappa.

Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training Five ways to look at cohen's kappa

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:30.328882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T20:58:29.852167Z digest=sha256:5004c494c2d8191c094d9043f0444e4f87a517b3aacb6e24efb7a4446e748818

Observation 61f4e4cd-7bc1-4887-9416-c9c35887c29e · outbound

This paper cites Ai and machine learning for next generation sci-ence assessments.

Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training Ai and machine learning for next generation sci-ence assessments

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:30.314858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T20:58:29.954850Z digest=sha256:15c553de482a3be126f8f13d1050449d08c98c328ca3bdd5028fdb4675028ecf

Observation ad4e0375-dc6f-4107-ada9-c2075910d820 · outbound

This paper cites Applying machine learning in science assessment: a systematic review.

Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training Applying machine learning in science assessment: a systematic review

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:30.300891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T20:58:30.030866Z digest=sha256:18aa3901256a377550e75d418c4f6a3e4e097d48d18b31c983ae31d6d7a88bc9

Observation 1e5decda-6452-4d66-9253-36229a57a8de · outbound

This paper cites Using large language models to detect self-regulated learning in think-aloud protocols.

Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training Using large language models to detect self-regulated learning in think-aloud protocols

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:58:30.221386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T20:58:30.036313Z digest=sha256:81e173197c0952214101c137acf83b6845c33064a3658231a7190a9bf8b0edfa

Pith citing papers

Observation f03c13f5-171d-4d6a-beef-365ede511696 · inbound

Exploring LLM-Generated Feedback for Economics Essays: How Teaching Assistants Evaluate and Envision Its Use cites this paper.

Exploring LLM-Generated Feedback for Economics Essays: How Teaching Assistants Evaluate and Envision Its Use Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:16:00.585500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:16:00.585500Z digest=sha256:dcd1e410b316b3e7f2dbe88fb27ca2cde15effc168a6734fb0286604fee39762

Observation de3fccfa-96fc-487b-9d92-1fa5ee7efbd7 · inbound

Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation cites this paper.

Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:05:58.166064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-11T01:56:51.209842Z digest=sha256:dddfaf51b0a6921f7e7a5945f9d4a4d35c80dde95e0e1aa3860895ebb999cee4

Observation 430478d8-6d21-4a35-8404-42feb41de2fa · inbound

Assessment in Team Problem-Solving Exercises in Computing Education cites this paper.

Assessment in Team Problem-Solving Exercises in Computing Education Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T13:10:36.620280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:10:36.620280Z digest=sha256:bcc6ef2e12d0550b23d550d5d25597610a55ac469f6d0b12f8a87288183d8bed

Observation aba8af69-4b1e-46e1-a80f-17f35040d5cf · inbound

Fine-Tuning Large Language Models for Codebook-Guided Coding of Students' Mathematics Metaphor Responses cites this paper.

Fine-Tuning Large Language Models for Codebook-Guided Coding of Students' Mathematics Metaphor Responses Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:28.739449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:28.739449Z digest=sha256:415cdaf2517dcbc2964d79563fa6d7d8b7cfa6a6df2ec40f5489d24316f92f35