REVIEW 3 cited by
LLMs as Workers in Human-Computational Algorithms? Replicating Crowdsourcing Pipelines with LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
LLMs have shown promise in replicating human-like behavior in crowdsourcing tasks that were previously thought to be exclusive to human abilities. However, current efforts focus mainly on simple atomic tasks. We explore whether LLMs can replicate more complex crowdsourcing pipelines. We find that modern LLMs can simulate some of crowdworkers' abilities in these ``human computation algorithms,'' but the level of success is variable and influenced by requesters' understanding of LLM capabilities, the specific skills required for sub-tasks, and the optimal interaction modality for performing these sub-tasks. We reflect on human and LLMs' different sensitivities to instructions, stress the importance of enabling human-facing safeguards for LLMs, and discuss the potential of training humans and LLMs with complementary skill sets. Crucially, we show that replicating crowdsourcing pipelines offers a valuable platform to investigate 1) the relative LLM strengths on different tasks (by cross-comparing their performances on sub-tasks) and 2) LLMs' potential in complex tasks, where they can complete part of the tasks while leaving others to humans.
Forward citations
Cited by 3 Pith papers
-
Linting is People! Exploring the Potential of Human Computation as a Sociotechnical Linter of Data Visualizations
Crowd-sourced Community Notes on social media can function as a sociotechnical linter for data visualizations, extending the linting metaphor from code to human computation.
-
Redefining Research Crowdsourcing: Incorporating Human Feedback with LLM-Powered Digital Twins
A study of an LLM-powered 'digital twin' system for crowd workers shows modest accuracy on Likert-scale surveys, with caveats around threshold tuning and evaluation contamination.
-
Are Large Language Models the future crowd workers of Linguistics?
In two replicated linguistics experiments, GPT-4o-mini's zero-shot responses matched or beat published human performance, but the study lacks statistical validation and relies on only two tasks.
Discussion (0). Continue with ORCID to comment.