The 100,000 whys of AI
Now with more of the same!
One of the most painful arguments I keep having with fellow techies is the question of whether you can distinguish between human-written and AI-generated text.
Their skepticism is rooted in reason: at their core, LLMs are state-of-the-art statistical models of how humans talk. If so, the output from the model should be almost by definition indistinguishable from human language under any statistical test.
I don’t think this is always argued in good faith; at least some of the debates are started by folks who wish to maintain deniability for their own underhanded use of the tech. But if you sincerely hold this belief, I present you the following collage:
The image shows about 220 Amazon book covers that appear if you search the site for “100000 whys” (link). Some of these books are category bestsellers in children’s literature. You can view a zoomable, full-resolution version here.
There’s nothing inhuman about any of these titles or covers. At the same time, I probably don’t need to convince you that you’re staring at the purest form of AI slop that now fills up many nonfiction book categories on Amazon. More specifically, it’s the artifact of the tools being quasi-deterministic: if a hundred “authors” give their favorite AI tool a similar prompt — say, “generate a reference book for children” — the model will produce functionally identical output perhaps 80% of the time.
The similarities in the collage go far beyond the choice of titles: for example, more than thirty covers in the two top rows feature a roaring T-Rex on the left:
There are many other visual clusters in the data, too. Look for glowing books, red-and-white cartoon rockets, golden retrievers, lions, blue-eyed robots, and so forth. The similarities extend even to author names: Ethan Bright, Nolan Bright, Pamela Bright, Daniel Bright, Thomas Bright, Andrew W. Bright, Mayan Bright, Mary Bright, Levi Bright — the Brights must be a big and exceptionally talented family.
This is precisely what makes LLM writing distinctive: it’s not that the models’ individual mannerisms are different from ours. It’s that they resort to the same, complex set of mannerisms in response to almost any normal prompt. This is a fuzzy signal, so you shouldn’t fire your intern when they say “it’s not this — it’s that”. But in more casual settings, it’s OK to trust your gut. In fact, these instincts are becoming increasingly important because traditional models of online interactions fall apart if it takes much less effort to produce content than to engage with it.
PS. If you’re using an LLM to automate blogging or social media commentary: yes, the tech is amazing, but chances are, your publication could be renamed to “100,000 Whys”.
For a followup article, click here:
And if you like horrors beyond our comprehension, please subscribe.




Some postscripts:
1) Beware of the "why" inflation: https://lcamtuf.coredump.cx/blog/why100k2.jpg
2) No, it's not just this title: https://lcamtuf.coredump.cx/blog/whys-vale.jpg
3) Yes, covers aside, these books are about what you'd expect: https://lcamtuf.substack.com/p/ai-childrens-books-body-horror-edition
4) Someone on Mastodon pointed out that the title is probably coming from the following 1929 book: https://en.wikipedia.org/wiki/One_Hundred_Thousand_Whys - apparently largely unknown in the West, but popular in China.
5) Presented without comment: https://lcamtuf.coredump.cx/blog/whys-make.jpg
And to the person on HN arguing that maybe these books aren't AI: my friend, you have a beautiful mind.
Discussing AI writing with a friend recently, I came up with 3 points that can be applied to virtually all AI use situations - 1) we need to remind people who ask AI to do the entire task that this behavior pretty much guarantees that they will be the first person replaced by AI, 2) AI is a tool as much as a piece of heavy equipment or a doctor's scalpel and equally dangerous, and we absolutely need to decide how to train and certify people to use AI properly, and 3) perhaps more realistic pricing on AI use will encourage people to use it more intelligently.
I still don't think we need AI, but it is here, so let's try to use it better.