Every working day, millions of professionals sit in front of a model and tell it that it is wrong. A tax accountant reads a generated summary of a filing and spots the deduction it has misclassified. A radiologist looks at a draft report and rewrites the sentence about a shadow at the base of the left lung. A lawyer scrolls through a contract the model has marked as standard and flags the one clause that is not. An engineer reads a generated patch, rejects it, and writes the version that actually handles the edge case. Each of these is a small act of teaching. Almost none of it reaches the people building the next model.

That is the paradox we started benchturn to resolve. The most valuable training signal in the world is produced every day by the people best qualified to produce it. And it is thrown away.

Using AI is not the same as teaching it

When a professional corrects a model, the correction lives in a chat window, a code review, or a document that gets overwritten. The model does not learn from it. The lab that trained it does not see it. The expert who made the correction gets nothing for it beyond a slightly better draft. The loop that should connect daily use to the next round of training is open at both ends.

This is not because labs do not want the signal. They want it badly enough to spend heavily on annotation, synthetic data and evaluation suites that try to approximate what an expert would say. It is because there is no channel through which an expert's ordinary working judgement can travel to a lab in a form that is usable, consented to, and paid for. Using AI and training AI have become two industries that barely speak to each other, even though they depend on the same people.

The signal is the process, not the answer

It is tempting to think of training data as a pile of correct answers. Give the model enough well-drafted clauses or clean patches and it will learn to produce them. Up to a point that works. But a correct answer hides the thing that made it correct. It does not show which alternatives were considered and rejected, which rule was checked against which exception, or why the obvious approach was abandoned halfway through.

Real work is made of those moments. The accountant's value is not in knowing the standard treatment of a deduction but in knowing when it does not apply. The radiologist's value is in the scan that looks normal and is not. The engineer's value is in recognising that the generated patch passes the tests and still introduces a race condition. This is the hard case, and it is where models still struggle, because it is precisely what is missing from the material they were trained on.

Scraped text cannot supply it. The internet records finished outputs, polished and stripped of the reasoning that produced them, and most of it has already been read. Synthetic tasks cannot supply it either, because a task written by someone who does not do the job tends to test what an outsider imagines the job to be. Models trained on outputs plateau. Models trained on how people actually work keep improving. The difference is whether the decisions, tradeoffs and corrections are in the data, or only their residue.

The party that is never in the room

Look at how expert signal reaches a lab today and one absence stands out. The annotation industry hires people by the hour to perform work on demand: answer this prompt, rank these responses, write this example. Some of it is done by genuine specialists, but it is performed work, staged for the purpose, and paid as labour rather than as the asset it becomes. The person is compensated for their time. The judgement they carry, built over a career, is priced at zero.

Scraping is simpler still: it pays nobody. The text is taken, the model is trained, and the people whose expertise is embedded in that text are not part of the transaction. Between these two channels, the practitioner whose judgement is the actual product is either rented briefly or bypassed entirely. Every other party in the chain, the lab, the annotation vendor, the platform that hosted the text, is paid. The expert is the one exception.

We think this is not only unfair but inefficient. A market in which the source of value is unpaid will underproduce the thing it needs most. The best practitioners have no reason to share their working judgement, and no way to do so if they wanted to. So the signal stays in their heads and their inboxes, and the labs make do with what they can buy or take.

Consent, provenance and payment as infrastructure

Most attempts to fix this treat consent and payment as compliance, bolted on after the data has been gathered. We think they need to be the substrate. When they are, the shape of the whole system changes.

Contributors choose what to share and for what purpose. A lawyer might share how they reviewed a lease but not how they advised a particular client, and might permit use for training but not for a public demonstration. Provenance travels with the material: a lab can see where each item came from, what was stripped from it, and what the contributor agreed to, without the contributor being identified. Payouts are itemised, so a contributor can see which pieces of their work were licensed, to whom in aggregate, and what each earned. Work that has not yet shipped in a licensed set can be withdrawn. And nothing that enters the pipeline overlaps with the public evaluations that labs use to measure progress, because data that leaks into a benchmark stops being useful to anyone.

None of this is exotic. It is what any honest supply chain looks like, applied to the one supply chain that has so far run almost entirely on taking.

What we are building

benchturn sits between the people who do real work with AI and the labs that train models. We take the work people choose to share, verify that it is what it claims to be, strip anything personal or proprietary, and structure it so a model can learn not just from the answer but from the path to it. We license that material to labs with clean provenance and informed consent, and we pay the contributors when their work is licensed. Experts teach models how to work. We make sure they get paid for it.

We are early, and we are deliberate about it. The pipeline that turns a working day into training material has to be trustworthy on both sides before it is large. If you are a practitioner who corrects models as part of your job and would like that judgement to count for something, we want to hear from you. If you are a lab that has hit the ceiling of scraped text and performed annotation, we want to hear from you too. Write to us at hello@benchturn.com.