header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

You Helped Google Train AI for 15 Years for Free, Just to Be Kept in the Dark

Read this article in 11 Minutes
You have proven that you are human, only to end up making yourself replaceable.
Original Title: You've been training Google's AI for 15 years. You had no idea.
Original Author: Sharbel, Co-founder of Unfungible
Original Translation: Lila, BlockBeats


Editor's Note: CAPTCHA, the familiar challenge of clicking the numbers or images when logging into a website, is well-known to every Internet user. However, each time you click "I'm not a robot," thinking you are just verifying your identity, you are actually participating in the world's largest and most secretive data production. Luis von Ahn's reCAPTCHA has aggregated scattered human behaviors into a data cornerstone that supports Google and its subsidiary, the self-driving company Waymo, among other core businesses.


Beneath the facade of "free" and "secure," the Internet has quietly reshaped a new kind of labor relationship: you spend time proving you are human but contribute to AI training, and once AI learns, this labor is completely replaced. This article has garnered over 9.5 million views on Twitter in less than 20 hours. The following is the original content:


Approximately 500,000 hours of human labor are utilized by Google for free every day. And the individuals making these contributions simply want to log into online banking.


reCAPTCHA is the most successful invisible data operation in Internet history. At its peak, 200 million people completed the verification process daily. However, almost no one realized the implications behind each click.


Google's self-driving car company, Waymo, is now valued at $45 billion. Yet most of its core training data has been freely provided by you as you accessed various websites.


Here is the full story:


Origin: A Clever Idea


In the year 2000, spam bots were wreaking havoc on the Internet. Forums were inundated, inboxes were overflowing, and websites needed a way to distinguish between humans and machines.


Carnegie Mellon University's Professor Luis von Ahn solved this problem. He invented CAPTCHA: distorted text that only humans could decipher, thereby preventing bots from passing through.


But von Ahn saw more than just that. Millions of people spent their energy on these challenges. What if this energy could do two things at once?


In 2007, he introduced reCAPTCHA. Its brilliance was this: instead of showing random garbled text, it displayed two words. One word was known to the system, and the other was from a real scanned book that computers could not recognize yet. And your answer helped in the digitization of these books.


These books came from The New York Times archives and Google Books, totaling up to 130 million.


You thought you were just logging into a regular website, but you were actually doing OCR (Optical Character Recognition) for the world's largest digital library.


In 2009, Google officially acquired reCAPTCHA.



Later, Google changed the game


The era of "distorted text" ended around 2012.


Google faced a new challenge: Street View cars captured every road globally, but the photos were just raw data. To enable AI, it needed to understand what it saw: street signs, crosswalks, traffic lights, storefronts.


So Google redesigned reCAPTCHA v2. Instead of distorted text, there were photo grids. "Click on all squares with traffic lights." "Select every crosswalk." "Identify storefronts."


These images came directly from Google Street View. Your clicks were the labels.


Every selection was telling Google's computer vision model: these cluster of pixels are traffic lights, that shape is a crosswalk. You're not taking a test; you're building a dataset.



Beyond Imagination Scale


At its peak, 200 million reCAPTCHAs were solved each day. Each challenge took 10 seconds, meaning 2 billion human labor seconds were generated daily. That's 500,000 hours a day.


The cost of paid data annotation is around $10 to $50 per hour. Calculated at the lowest rate: the daily value of the free labor extracted reaches as high as $5 million.


Moreover, reCAPTCHA is not limited to a specific app. It is present in every bank, every government portal, every e-commerce website. You have no choice: Want to log in to your account? First, help label a dataset. Google has never asked for your opinion, paid you a cent of salary, or even informed you of this fact.



What has all this led to?


This data is directly fed into two products:


-Google Maps: The most widely used navigation tool in the world. Its ability to recognize road signs, stores, and city geography is partly thanks to the billions of times humans have labeled data while logging into websites.


-Waymo: Google's self-driving car project. For safe navigation, autonomous vehicles need to almost perfectly identify thousands of visual patterns.


The ground truth training data for that identification work is precisely what millions of people labeled through reCAPTCHA unknowingly. Waymo completed over 4 million paid trips in 2024, valued at $45 billion. Its cornerstone, laid by those "unpaid internet users" who just wanted to check their email.


Why can't anyone replicate this model?


Data labeling is extremely expensive. Companies like Scale AI, Appen, and Labelbox exist to solve this problem, hiring hundreds of thousands of workers, sometimes paying less than $1 per hour.


Google took a different approach: they turned labeling into a mandatory task. No payment required, no consent needed, but rather as a "ticket" to enter every corner of the internet. The result is: billions of tagged images, global coverage, all-weather conditions, every city in the world. No labeling company can achieve this. The internet itself is a factory, and every netizen is an undocumented employee.



You are still participating


reCAPTCHA v3, launched in 2018, no longer even displays a challenge. It observes how you move your mouse, your scrolling speed, and dwell time. Your behavioral fingerprint informs it whether you are human. These behavioral data also feed back into Google's AI systems.


You never actively chose to join, never had a checkbox to tick. But right now, on most websites you visit, you are still doing so.


Disturbing Irony


Luis von Ahn's original vision was brilliant: to harness human energy that was otherwise being wasted and turn it into useful output. However, what Google did with this vision was a different story altogether. They leveraged a security mechanism that users were forced to use, deployed it across the web, reaped the output to build a business product worth hundreds of billions of dollars. Users got nothing in return, not even awareness.


The deepest irony lies in the fact that: you spent years proving you were human by performing visual recognition tasks that AI couldn't do at the time. Yet once AI learned to do these tasks, human visual annotations became obsolete.


You proved you were human, only to end up making yourself replaceable.


Original Post Link


Welcome to join the official BlockBeats community:

Telegram Subscription Group: https://t.me/theblockbeats

Telegram Discussion Group: https://t.me/BlockBeats_App

Official Twitter Account: https://twitter.com/BlockBeatsAsia

Choose Library
Add Library
Cancel
Finish
Add Library
Visible to myself only
Public
Save
Correction/Report
Submit