Research
Notes on real work, training data, and who gets paid.
Essays from benchturn on why the best training material is in people's workdays, and how it should reach the labs.
- How often do people correct models in the wild? In a 5,000-conversation sample of WildChat, 8.3% of multi-turn English chats contain a turn where the user tells ChatGPT it got something wrong.
- A provenance record for contributed work The per-item record benchturn keeps so that a lab can audit what it licensed and a contributor can see what was used, without asking anyone.
- Process versus outcome: what the evidence says A review of the published evidence on process supervision, data efficiency, data limits and benchmark contamination, and where the gaps are.
- The benchturn thesis Why the best training data lives in professionals' workdays, why the experts who hold it are never paid, and what benchturn is building to change that.