A packed Ars Technica forum thread read aloud passages from Cory Doctorow's new book, and the mood was combative. In the Ars OpenForum conversation about The Reverse Centaur's Guide to Life After AI, Doctorow argued that the simplest way to puncture the market glamour around large language models is to cut off their raw inputs and expose the plumbing. He urged stricter enforcement of existing copyright rules, limits on commercial web scraping and public disclosure of the human labour behind outputs, and warned that staged demos and unregulated scraping are propping up inflated valuations. If enforced through litigation and penalties, those measures would raise training costs, slow model upgrades and change the incentives that have fuelled speculative investment.

In the Ars thread Doctorow read scenes from his book and then walked the audience through what he described as tactical, legally grounded interventions. He framed the problem in physical terms: the models run on a steady diet of internet text and their market value depends on uninterrupted access to that feedstock. That, he said, is a vulnerability rather than an inevitability.

Expose the demos, choke the feed

Doctorow began by attacking public demos and marketing rituals. "Many of the demos for AI have just turned out to be people in India pretending to be robots," he said during the Ars OpenForum discussion, using the example to argue that staged demonstrations inflate public awe and investor belief in autonomous capability. By insisting on independent verification and by publicising fakery, Doctorow argued, companies would find it harder to sell novelty as margin.

His second, and central, line of attack focused on training data. Doctorow emphasised that modern neural networks often memorise and reproduce large chunks of source text rather than inventing wholly novel content. From that premise he built a practical playbook: enforce existing copyright law against unlicensed training uses, limit commercial web scraping, and ensure websites can opt out of being harvested for model training.

Those aren't abstract proposals in his telling. Doctorow urged immediate legal remedies aimed at firms that ignore site-level refusals, such as robots.txt. He argued for prosecuting willful, for-profit scraping that flouts those refusals and for using litigation to set precedents that apply current copyright rules to model outputs that verbatim reproduce copyrighted material. The mechanics are straightforward: if scrapers must secure licences or pay creators, the economics of model training change.

Doctorow mapped the effects in concrete terms. Cutting off broad, free access to internet text would raise the marginal cost of training. That makes faster iteration more expensive and slows the cadence of model upgrades, removing one lever that has helped newly funded labs iterate past incumbents.

Requiring recurring payments to content owners turns a one-off data grab into a continuing operating cost, which in turn erodes the financial case for valuations that rely on cheap, abundant corpora.

He tied those mechanics to the gig economy with a rhetorical image that recurs in the book. Doctorow defines a "reverse centaur" as "a machine head on a human body, a person who's serving as a squishy meat appendage for an uncaring machine." The phrase captures how gig workers who create, curate or label data are converted into appendages for automated systems. Making platforms disclose the human labour embedded in outputs would alter the public story about where value is created and who should be paid for it.

Those interventions, Doctorow said, target the data advantage that lets small, well-funded labs iterate quickly. If that advantage is constrained, the valuations that depend on uninterrupted, low-cost access to massive web corpora will be harder to defend in private financing rounds and public markets. In his account, these moves operate at the roots of the hype cycle rather than at the surface where safety checklists and feature flags tend to gather attention.

He framed the enforcement path as legal work rather than new statute writing. The conversation at Ars OpenForum emphasised using existing copyright law and prosecutorial tools, plus statutory penalties against commercial scrapers that ignore site-level restrictions, to create a set of judicial and administrative precedents. Doctorow didn't point to a specific regulator or pending bill; he presented litigation and enforcement as the near-term route to effect.

He also argued that public exposure matters. If users come to expect that demos are staged and that outputs often rest on copied passages or paid gig work, appetite for novelty-driven features could fall. Lower uptake would reduce product traction, which in turn would affect valuations that assume rapid user growth driven by impressive-sounding demos.

The conversation left open the timeline for such change. Doctorow's book and the Ars discussion offered no specific legislative schedule, and the proposals depend on courts and enforcement agencies taking up cases that set clear precedents about scraping and model outputs. Still, he presented enforcement as a pragmatic lever that doesn't require bespoke AI statutes to begin shifting incentives.

Related Articles

Doctorow's clearest call was legal: pursue enforcement actions and statutory penalties against commercial scrapers that ignore site-level restrictions, and use existing copyright law to challenge models that reproduce copyrighted text verbatim. The next fights will play out in courtrooms and at enforcement agencies.

This article was created with AI assistance.