Amazon says it will close Mechanical Turk on 30 September. The announcement itself is one banner across the top of mturk.com. It reads: "Following an assessment, we've made the decision to close Amazon Mechanical Turk, effective September 30, 2026."
Anthony Ha of TechCrunch dates the launch to 2005 and describes the service as a marketplace where people were paid tiny amounts for simple tasks that resisted automation. The Next Web's account of the closure lists the work as labelling data, transcribing audio and video, and answering surveys, at cents a task. I loaded the page on 28 August 2026 and found the closure banner sitting on top of a sales pitch still written in the present tense, offering content moderation and product categorisation to anyone who wants them.
The closure FAQ is specific where the banner is vague. Job submission stops on 30 September. Buyers keep their standard 30 days to approve or reject what was submitted, so bonuses can be paid until 30 October, and anything nobody acts on is approved automatically. Unsubmitted jobs expire. Prepaid balances are refunded within 30 days.
Three write-ups reached for the same irony. The Next Web headlined it "the human workforce it sold as AI". Alfonso Maruccia at TechSpot wrote that the platform was "clearly being left behind" as AI-focused rivals emerged. Ha supplied the historical rhyme back in July: the original Mechanical Turk was "itself a hoax, with a hidden human chess player pretending to be a chess-playing machine". It is a good story.
Amazon has published no reason for the closure. I counted the sentences of explanation on the banner: one, and it is "We regularly evaluate our programs, tools, and services and make adjustments based on those assessments." The FAQ adds nothing. So the causal story is guesswork for everyone outside the company. What can be worked out is narrower: what changed between the contamination this platform survived and the one researchers measured on it five years later, and how the vendors still selling human data now describe what they sell.
What buyers were relying on
Amazon's own sales copy, still live under the closure banner as I write, puts it plainly: "While technology continues to improve, there are still many things that human beings can do much more effectively than computers, such as moderating content, performing data deduplication, or research." What Amazon sold was that gap. Cheap labour was the delivery mechanism for it.
Amazon's copy names both of the very different groups that ended up in the same queue. On one side, "simple data validation and research" and "survey participation", the survey work social scientists buy because human responses are the object of study. On the other, what the same page calls human-in-the-loop, "where human feedback is used to help validate and retrain your model". The machine learning buyer wants a label it can check a model against. The survey buyer wants a response produced by a person rather than by a model on that person's behalf. In both cases part of the value sits in who made it.
In June 2023, Veniamin Veselovsky, Manoel Horta Ribeiro and Robert West published a paper called "Artificial Artificial Artificial Intelligence" (arXiv:2306.07899). They re-ran an abstract summarisation task on Mechanical Turk and, using keystroke detection combined with a classifier trained to spot synthetic text, estimated that 33 to 46 per cent of the crowd workers on that task had used a large language model to produce their answers.
Their finding is not that the work was bad. The summaries may have been perfectly good. Their finding is that on that task, the gap Amazon was selling had partly closed from the inside, in a way a buyer reading the delivered summaries would not have seen.
The number is narrower than the coverage made it
Before anyone builds a thesis on 33 to 46 per cent, including me: Veselovsky, Horta Ribeiro and West write in their own abstract that generalisation to other, less LLM-friendly tasks is unclear. LLM-friendly is their term, and abstract summarisation is squarely one of those tasks. Whether the same rate would show up in, say, drawing boxes around pedestrians in dashcam footage is precisely what they say is unclear. One text task, one platform, one study, one moment in the middle of that year.
Maruccia's TechSpot piece this week gives the figure as "a significant number of Turkers (46 percent) were using generative AI and chatbots to complete the platform's tasks", without the task limit and without the lower bound. That is the version now circulating. It is not the version in the paper.
The neglect explanation
The strongest case against reading anything into the closure is that it needs no reading at all. Ha reported on 5 July 2026 that it would close to new customers on 30 July, and quoted AWS saying the decision came after "careful consideration" and that it "continues to invest in security and availability improvements for Mechanical Turk, but we do not plan to introduce new features." That is a wind-down, stated in public in July, before any of this week's coverage existed.
One of the organisers who works on the platform says much the same. The Next Web's report quotes Krista Pawloski, a data worker and organiser with the advocacy group Turkopticon, telling CNBC that Mechanical Turk had been in decline for years, that Amazon put fewer resources into improving it while rivals arrived, and that many workers moved to competing platforms. I could not open CNBC's own article to check the wording against the original, so treat that as second-hand. Her account of neglect, from someone that report identifies as a data worker on the platform, is better evidence about the closure than anything I can offer. It also leaves the interesting question open, because neglect only bites relative to what the buyers started demanding.
Why the earlier crisis was survivable
Researchers hit a data-integrity problem on this platform in 2018. On 8 August that year, Max Hui Bai published a note describing a wave of nonsense responses in his survey data and a way to find them. The bad responses shared repeating GPS coordinates, with one latitude ending in 88639831 turning up across multiple unrelated surveys. Bai wrote that around 90 studies appeared to be affected, with contamination traceable back to that March. Amazon's banner puts the closure at 30 September 2026, eight years after Bai's note.
The difference between Bai's problem and Veselovsky's is what a buyer could do about it alone. Bai's left a fingerprint in the delivered data. You could search your own results for the coordinate, drop those rows, add attention checks, and keep buying at the same price. Quality problems have prices: buy redundancy, take three answers instead of one, move on.
A model-written summary is different for the buyer in a specific way. It is a competent summary, and it arrives looking like the thing you ordered. Veselovsky and his co-authors used two instruments to estimate how often it was happening: a classifier trained on synthetic text, which reads what was delivered, and keystroke detection, which requires watching the worker type. The paper reports a range. Note where those two signals come from. The keystroke data existed because the authors designed the task to capture it; a requester buying finished work off the marketplace receives the submission. Run a classifier over a batch yourself and you get scores, which support a threshold and an error rate before they support a verdict on the person whose work you are about to reject. Rejection is what the marketplace's workflow runs on: the FAQ's whole treatment of the wind-down is about who gets to approve or reject what, and by when.
What Amazon's product pages do not offer the buyer, as far as I could find on them, is any record of how a submission was produced. I went looking for a current worker count on the same pages while I was there, and could not find one of those either.
How the surviving vendors sell it
Look at how two of the surviving vendors sell themselves. I loaded Prolific's home page on 28 August 2026. The first line is "Real human data. Ready in minutes." The pitch to AI teams underneath offers "human feedback from representative populations" for "preference tuning, safety evals, and benchmarks you can defend", from "verified participants". Each of those words describes the participants. None of them describes how many responses you get.
Mercor's home page, same day, leads with figures. An average contracted rate of US$111 an hour. 415,000 roles created. More than US$4 million in daily payouts. Experts across more than 300 professional fields.
I can't show you buyers moving from Mechanical Turk to either, and the markets are not the same shape. What those two pages do show is what each is now advertising: screening and expertise rather than volume and price. Prolific's phrase "benchmarks you can defend" is the one I keep coming back to, because defending a benchmark means answering a question about the people who produced it.
The bet
The Next Web, TechSpot and TechCrunch all have this as automation closing a loop on itself. My claim is smaller and does not depend on knowing Amazon's reasons. On the task Veselovsky and his co-authors measured, the contamination arrived in a form a buyer could chase by running a classifier over a batch and living with its error rate. A requester could build keystroke capture into a task as well, though instrumenting a cents-a-job task to a research standard is its own expense. Five years earlier, Bai's note had told buyers exactly what string to search their own data for. How far that reaches beyond one summarisation task is the open question the authors themselves flag, and I am not going to close it for them.
Here is the falsifiable version of my bet. Buying a human judgment you can defend gets more expensive from here and stays that way, because the buyer is paying for screening and process alongside the answer. If between now and the end of 2027 someone runs an anonymous, unscreened, cents-per-item labelling marketplace at scale and frontier labs buy from it for evaluation and preference work rather than bulk training filler, then screening was not what the buyers were paying for, and I have read this wrong.