Kled Raises $10M in Funding
Learn More
Kled Raises $10M in Funding
Learn More
09.30.26
•
Avi Patel

AI is spending $700 billion on the wrong problem
AI is spending $700 billion on the wrong problem
AI is spending $700 billion on the wrong problem
5 MIN READ
INSIGHTS

ai has a data problem. it is spending like it has a chip problem.
four companies will spend more than $700 billion on ai infrastructure in 2026. amazon alone is at $220 billion. that is nearly double what the same four spent in 2025.
the entire human data industry, every vendor combined, is about $15 billion a year.
47 to 1.
the biggest problem in ai is not the $700 billion. it is the $15 billion.
there is one internet and we used it
epoch ai counted it. about 300 trillion tokens of usable human text. at current training rates this will be fully consumed between 2026 and 2032. earlier if models are overtrained.
they are.
synthetic data does not escape this. someone still writes the rubric, ranks the outputs, and decides what good looks like.
so every lab trained on the same finite pile, and is now buying $700 billion of machinery to squeeze it harder.
no matter how many gpus are purchased, model improvement will stall.
new data is the only solution.
compute is a purchase order. data is a coordination problem.
buying compute is easy. you wire nvidia the money and the gpus show up in a truck. the only constraint is how much money you have. that is why the labs are raising hundreds of billions. it is a solved problem.
data is not this simple.
the data that trains the next model is not in a warehouse. it is in the pockets of most of humanity, in 170 countries, behind consent, behind identity, behind payment rails that do not reach the people who have it.
a mechanic's hands. a tax return. a radiology scan. a kitchen in jakarta at 6am.
to get this data you need, at minimum:
a reason for a stranger to give it to you
proof the stranger is one human and not a script
proof the file is real, owned, and licensed
quality control across billions of files
a way to pay that stranger in a country your bank does not serve
a way to ask for the specific thing you need, not the thing lying around
a way to do all of it again tomorrow
this is the hardest coordination problem. probably ever.
the incumbents are solving the wrong problem
scale is valued at $29 billion. mercor is in talks at $20 billion. surge was talking to investors at $15 to $25 billion. micro1 raised at $4 billion this month. every one of them is priced on the same sentence: human data is the bottleneck.
they are right about the sentence. they are wrong about the product.
all four run the same machine. a web dashboard, a task queue, a contractor found by resume or ai interview, sitting at a laptop, adding opinions to data that already exists.
that is processing. it is not collection.
none of them can capture anything that is not already on a screen. none of them have a single user who opens the app in the morning without being told to.
the other half of each company is a sales floor.
my friends get calls from mercor reps on any given day, and the pitch is always the same: i know the labs, i can sell your data.
they do not make the data and they do not own it. they route it and keep roughly a third of every dollar.
this is glorified dropshipping.
these are not bad companies. in fact most of them are quite good. but they are no more than staffing agencies with sales teams in front.
i can confidently say that this product has a ceiling and we are approaching it soon.
how do you get a million people to give you exactly what you want
this is the only question in human data. not how do you label it. how do you get it.
the answer has been in your hand since 2007.
tesla has more than 10 billion miles of fsd data. tesla's customers bought the sensors themselves, drove them everywhere, and paid tesla for the privilege. the fleet is the collection strategy. the car is just the sensor.
in analogy the phone is the same.
5.65 billion people carry one. 68% of the species. americans check it 186 times a day, 85% of them within ten minutes of waking up. it carries a 48 megapixel camera, a lidar, a gps and a secure enclave. a $1,000 sensor package that people bought themselves and carry everywhere.
this is the answer. not a web recruiting service, but a mobile application.
what a phone can do that a web dashboard cannot
a phone is different in five ways.
1. it is there when things happen.
image of a doctor's note in a waiting room. pov video of a mess you just made. audio recording of a genuine conversation. the most valuable and authentic data is simply not possible to capture with a computer, the form factor does not work.
2. it can be asked.
a lab needs 20,000 photos of wet roads at night in southeast asia. on a web dashboard this is a project. on a phone it is a notification to the people who are already there. the photos come back the same day.
3. it already holds the data.
2.1 trillion photos were taken last year, 94% of them on a phone, most of them never posted anywhere. no model has trained on them. every one has a single owner who can license it. no other device holds a dataset that size, already collected, with the rights attached.
4. it is a sensor pointed at the world.
ego4d, the academic benchmark for first person video, took 13 universities across 9 countries to collect 3,670 hours. a phone in a kitchen does the same job, and there are billions of kitchens.
5. it knows who is holding it.
up to 46% of mechanical turk workers were using chatgpt to do their tasks. this is not possible on a phone. yesterday amazon shut down mechanical turk after 21 years.
more web platforms will continue to die. real human input data can only come from a mobile application. there is no other solution that will survive.
elegantly put: i say fuck em all
i launched kled in early 2025. this has been my thesis on data for quite some time.
we made many contrarian bets.
we believed model improvement is a function of scale, not just experts.
we believed the coordination problem needed a simple mobile app, not a complicated website.
we believed that a global network needed to be incentivized not only by fiat payment but also a single reward currency.
we believed that the working person would willingly participate in data collection.
we were rejected by over 200 vcs before our $14 million seed round.
i understand why. everything in ai is consensus now. every lab has the same money, buys the same chips from the same company, and buys the same data from the same four vendors.
but when everyone buys the same input, input stops being an edge.
nvidia sells compute to everyone but data is the only input that can be different. data depends on who you got it from and how. so the only edge left in this industry is data nobody else collected, and that comes from a method nobody else used.
i believe the next real jump in models comes from a collection method the consensus rejected. and the consensus will always reject it, because capital funds what it recognizes. the rejection is not a signal that the idea is wrong. the rejection is what keeps the data non consensus.
if two hundred vcs had said yes, two hundred companies would have the same data.
so being contrarian here is not a personality. it is the win condition. it cannot be done halfway, because a contrarian who hedges cannot create data different enough to matter.
elegantly put: i say fuck em all.
ai has a data problem. it is spending like it has a chip problem.
four companies will spend more than $700 billion on ai infrastructure in 2026. amazon alone is at $220 billion. that is nearly double what the same four spent in 2025.
the entire human data industry, every vendor combined, is about $15 billion a year.
47 to 1.
the biggest problem in ai is not the $700 billion. it is the $15 billion.
there is one internet and we used it
epoch ai counted it. about 300 trillion tokens of usable human text. at current training rates this will be fully consumed between 2026 and 2032. earlier if models are overtrained.
they are.
synthetic data does not escape this. someone still writes the rubric, ranks the outputs, and decides what good looks like.
so every lab trained on the same finite pile, and is now buying $700 billion of machinery to squeeze it harder.
no matter how many gpus are purchased, model improvement will stall.
new data is the only solution.
compute is a purchase order. data is a coordination problem.
buying compute is easy. you wire nvidia the money and the gpus show up in a truck. the only constraint is how much money you have. that is why the labs are raising hundreds of billions. it is a solved problem.
data is not this simple.
the data that trains the next model is not in a warehouse. it is in the pockets of most of humanity, in 170 countries, behind consent, behind identity, behind payment rails that do not reach the people who have it.
a mechanic's hands. a tax return. a radiology scan. a kitchen in jakarta at 6am.
to get this data you need, at minimum:
a reason for a stranger to give it to you
proof the stranger is one human and not a script
proof the file is real, owned, and licensed
quality control across billions of files
a way to pay that stranger in a country your bank does not serve
a way to ask for the specific thing you need, not the thing lying around
a way to do all of it again tomorrow
this is the hardest coordination problem. probably ever.
the incumbents are solving the wrong problem
scale is valued at $29 billion. mercor is in talks at $20 billion. surge was talking to investors at $15 to $25 billion. micro1 raised at $4 billion this month. every one of them is priced on the same sentence: human data is the bottleneck.
they are right about the sentence. they are wrong about the product.
all four run the same machine. a web dashboard, a task queue, a contractor found by resume or ai interview, sitting at a laptop, adding opinions to data that already exists.
that is processing. it is not collection.
none of them can capture anything that is not already on a screen. none of them have a single user who opens the app in the morning without being told to.
the other half of each company is a sales floor.
my friends get calls from mercor reps on any given day, and the pitch is always the same: i know the labs, i can sell your data.
they do not make the data and they do not own it. they route it and keep roughly a third of every dollar.
this is glorified dropshipping.
these are not bad companies. in fact most of them are quite good. but they are no more than staffing agencies with sales teams in front.
i can confidently say that this product has a ceiling and we are approaching it soon.
how do you get a million people to give you exactly what you want
this is the only question in human data. not how do you label it. how do you get it.
the answer has been in your hand since 2007.
tesla has more than 10 billion miles of fsd data. tesla's customers bought the sensors themselves, drove them everywhere, and paid tesla for the privilege. the fleet is the collection strategy. the car is just the sensor.
in analogy the phone is the same.
5.65 billion people carry one. 68% of the species. americans check it 186 times a day, 85% of them within ten minutes of waking up. it carries a 48 megapixel camera, a lidar, a gps and a secure enclave. a $1,000 sensor package that people bought themselves and carry everywhere.
this is the answer. not a web recruiting service, but a mobile application.
what a phone can do that a web dashboard cannot
a phone is different in five ways.
1. it is there when things happen.
image of a doctor's note in a waiting room. pov video of a mess you just made. audio recording of a genuine conversation. the most valuable and authentic data is simply not possible to capture with a computer, the form factor does not work.
2. it can be asked.
a lab needs 20,000 photos of wet roads at night in southeast asia. on a web dashboard this is a project. on a phone it is a notification to the people who are already there. the photos come back the same day.
3. it already holds the data.
2.1 trillion photos were taken last year, 94% of them on a phone, most of them never posted anywhere. no model has trained on them. every one has a single owner who can license it. no other device holds a dataset that size, already collected, with the rights attached.
4. it is a sensor pointed at the world.
ego4d, the academic benchmark for first person video, took 13 universities across 9 countries to collect 3,670 hours. a phone in a kitchen does the same job, and there are billions of kitchens.
5. it knows who is holding it.
up to 46% of mechanical turk workers were using chatgpt to do their tasks. this is not possible on a phone. yesterday amazon shut down mechanical turk after 21 years.
more web platforms will continue to die. real human input data can only come from a mobile application. there is no other solution that will survive.
elegantly put: i say fuck em all
i launched kled in early 2025. this has been my thesis on data for quite some time.
we made many contrarian bets.
we believed model improvement is a function of scale, not just experts.
we believed the coordination problem needed a simple mobile app, not a complicated website.
we believed that a global network needed to be incentivized not only by fiat payment but also a single reward currency.
we believed that the working person would willingly participate in data collection.
we were rejected by over 200 vcs before our $14 million seed round.
i understand why. everything in ai is consensus now. every lab has the same money, buys the same chips from the same company, and buys the same data from the same four vendors.
but when everyone buys the same input, input stops being an edge.
nvidia sells compute to everyone but data is the only input that can be different. data depends on who you got it from and how. so the only edge left in this industry is data nobody else collected, and that comes from a method nobody else used.
i believe the next real jump in models comes from a collection method the consensus rejected. and the consensus will always reject it, because capital funds what it recognizes. the rejection is not a signal that the idea is wrong. the rejection is what keeps the data non consensus.
if two hundred vcs had said yes, two hundred companies would have the same data.
so being contrarian here is not a personality. it is the win condition. it cannot be done halfway, because a contrarian who hedges cannot create data different enough to matter.
elegantly put: i say fuck em all.
The Leading Data Marketplace.
A Nitrility Inc. Company
Kled AI © 2026

The Leading Data Marketplace.
A Nitrility Inc. Company
Kled AI © 2026

The Leading Data Marketplace.
A Nitrility Inc. Company
Kled AI © 2026
