Forgive me, I beg, the odd, breezy posting schedule this week. I reasoned that many would still be on vacation and wanted to save the (for a word) denser material for next week. Like the article on deriving objective morality: not simple, not short, not easy.
So I took this week to finish what I could of the second edition of a book that did surprisingly well (at least I thought so; it sold about 5,000 copies). Working title: Son of Everything You Believe Is Wrong. I got to the point I wanted an index, and since I wrote in LaTeX, I thought this would be particularly simple using AI.
I was wrong.
I should have expected it, since it was I who wrote about AI limitations: here and here (among others). These language models excel where they have all the causes, like in images and moving images (movies). They do okay, but err often enough, with code or math. As long as the code or math is common, it works fine. But once you move into new areas, well, you had better know the right answers, or you could fool yourself easily.
My book, I told both Grok and ChatGPT, was a book on logical arguments. It therefore assumed it would be like all those other books, which it is not. You can follow the hilarity in the video below.
Briefly, since I already blew two days on this and want to purge the bad taste, this:
I asked Grok for index tags, which in LaTeX are simple: “\index{word}” next to the words. It kept swearing, and promising larger and larger numbers, of index tags were being inserted into my text. Best I could get was two tags. Mostly none. Worst was my own idiocy for not checking and continuing on as if it was working.
Because I had to keep uploading chapters, I eventually ran out of my weekly credits. I had heard ChatGPT does better with language, so I created an account there (I never had one) and asked the model if it could do LaTeX index tags. It could!
I uploaded the whole book, as requested, and what followed was the best comedy. Over and over and over ChatGPT said something along the lines of “This is good. Here is our plan. We’ll first create a Table of Concepts. We’ll then review the entries and I’ll do the tags.”
I kept answering “Great, let’s do it.” Then it said, in effect, “Agreed. Here is our plan. We’ll first create a …”
Over and over again. Then it turned out it said it couldn’t read the book it was quoting back to me. It told me to re-load chapters each time I asked a question, because of some internal limitation. I did so.
Promises by the acre were provided. But not one damned index tag did I get.
Then I ran out of the daily credits or whatever they use to track usage.
My own fault, all of it. I spent all yesterday afternoon doing the tagging by hand. Which I should have done in the first place.
Anyway, to tease the book, I created that seven second video with Grok. And since it knows the causes of images, it did a reasonable job. I thought it was cute. So I posted it on YouTube (and here yesterday).
I immediately lost a lot of followers especially on YouTube, doubtless on the suspicion that I had turned into yet another slop AI account.
I have learned my lesson. I promise.
Video
Here are the various ways to support this work:
- Subscribe at Substack (paid or free)
- Cash App: $WilliamMBriggs
- Zelle: use email: matt@wmbriggs.com
- Buy me a coffee
- Paypal
- Other credit card subscription or single donations
- Hire me
- Subscribe at YouTube
- PASS POSTS ON TO OTHERS
Discover more from William M. Briggs
Subscribe to get the latest posts sent to your email.


I had similar experiences with Grok.
I asked it to extract some data from graphs in a published paper whose paper I had already extracted the graphical data myself. I got a response that was completely wrong for the data extraction and a comment that the mathematical model in the paper was exactly the same as the one that I was working on. It was not. When I pointed out the discrepancy Grok complained that the resolution of the graph was too low. I had used exactly the same graph.
On another occasion, I asked Grok to summarise several published papers. It returned the abstracts of these papers pretty much verbatim. Not very helpful.
It also made some very basic errors in programming when I asked it to write some MatLab code for me. To be fair though, it has been very helpful to me in helping me understand the correct syntax for Matlab.
I guess that some of my problems may have been the result of poor prompting but it is still necessary to check any responses that it makes.
For what it is worth, I use a quite-good human indexer.
zurain.shahzad@gmail.com
He is way over-qualified for this work, but seems to appreciate side-gig money.
https://www.linkedin.com/in/zurain-shahzad3a367a204/?lipi=urn%3Ali%3Apage%3Ad_flagship3_profile_view_base_contact_details%3BwWh6GLW1TEiu3yxvhfgz9g%3D%3D
I hope you’re promote your new book here when available, as it sounds quite intriguing. I only discovered your website a few months ago, and have enjoyed your content, so I’d certainly be interested in supporting your work that way.
Many items I find on line seem redolent of AI; they contain a great deal of superfluous verbigeration. Perhaps the programers mistake that for erudition.
AI’s logical arsenal also seems incapable of a certain (my favorite) logical ploy, to wit, “Reductio ad absurdum”.
Artificial Intelligence is no better at recognizing it’s own faults than the common, garden, variety.
Tagging by hand! I would have spent at least the same amount of time writing a perl script or three to “automate” the task probably involving make files, likely after having spent at least as long trying to make it perl one-liner . Where’s the fun of doing it by hand. Using perl and make means that it’s as unmaintainable as AI slop code, of course, but then again you know you’ll “do it better” next time.
Of course, indexing is the last task an LLM would be capable of because it works using next token prediction i.e fancy autocomplete, and has no understanding of meaning so you were foolish to even try using an LLM in the first place. As ever, Briggs is at fault. An LLM could make you a concordance I guess but that’s definitely a one liner in shell.
A couple of things I have noticed about how LLMs have developed in the last year or so:
1.) Somehow they have become even more overconfident. I’ve tried a bunch of tests on ones equipped with search capabilities to find information on somewhat obscure books, software, TV shows, websites, etc. To the credit of the LLMs, sometimes they do get detailed and accurate information. However, they more often will pull information from a completely unrelated work or from what you’d expect from a generic example and output pages of detailed (but erroneous) information based on that. I don’t think I’ve seen an output like “I’m sorry, I can’t find enough information on that subject to give you an answer” for something like half a year.
2.) They get really hung up on initial conditions. One test I run is to have them summarize websites which are obscure, but which come up on search results (so if they really can summarize searches they should be able to describe them.) What will happen is that the LLM search will find one or two pages on that site (seemingly at random) and pull all data from that page. So for example, suppose I ask it to analyze the opinions of a blog where the writer discusses both science fiction and philosophy. I might ask it what the author thinks of the new Star Wars movie, and success! The LLM finds that page and gives me an accurate summary. I then ask what the author thinks of Plato and I get a weird essay which connects Plato’s as a precursor to the Jedi and Sith, because the LLM isn’t looking up the blog author’s essays on Plato at all but is instead using the same Star Wars review and trying to predict how Plato could appear in it.
For an example of it going for a generic take, if I ask it to review a site discussing the Rules as Written (RAW) approach to 1st Edition, with the prompt of “which edition of AD&D/D&D does this author prefer” I will get a long essay about how the author thinks that everything before 3rd edition was a mess and that the author advocates simply ignoring core rules from earlier editions for the sake of having fun. That is, the exact opposite of what actually appears on the site I asked about, but probably the most common opinion in the general D&D community.
It’s definitely reaching the point where these things are only good for asking about things you already know about (or at least which you can verify and are willing to double check.) The formal structure of the accurate and inaccurate outputs are the same.
I’ve also experienced that thing where it claims to not be able to do something it just did. I see this all the time when one of the LLMs hyperfixates on a specific webpage and tries to use it answer every question about a larger site. (Like trying to obtain all of an author’s opinions on every topic from a single Star Wars review). I will point out that it is taking information from one page that is irrelevant, and that it should really be looking at an essay on another page. The response I will get is always that the LLM can’t look at ANY webpage, only what was in its training data, so it can’t help. But it clearly found the Star Wars review by a websearch, and if I do a hard reset on the model and ask it about another essay it will find that through a websearch (and base all of its later outputs on that one new essay.)
I also like when they tell me that I can easily fix the problems I am having by using settings options which do not exist (with the suggestions probably being pulled from unrelated tech manuals.)
Lol, I remember “Iron Eyes Cody,” the spaghetti Indian!