In-Depth

Why AI Writing Has a Long-Term Memory Problem

Over the last few years, we have seen users increasingly use generative AI for relatively simple tasks. A user might, for instance, use an AI tool to compose an email or to turn a Word document into a PowerPoint presentation. All of those things work great. But what happens when users inevitably begin to use generative AI for lengthier and more complex tasks? It's one thing to ask an AI tool to create a 700-word essay. It's quite another thing for an AI tool to competently handle a 100,000-word document.

Recently, I got a firsthand feel for how well AI works with large documents. Recently, I have been spending a lot of my time working on a rerelease of a book that I published several years ago. A lot has changed since the time that the book was first created, and so I wanted to bring the book up to date and add some new material. I didn't use AI to write the book, obviously, but I wanted to try something new. I wanted to create a series of podcasts to go along with the book. My idea was to create one podcast for each chapter, and I decided to let AI help with creating the podcast scripts. Ultimately, however, this simple task became way more interesting than I ever would have expected.

Initially, my approach to this problem was really straightforward. I provided AI with a basic description of what I wanted. Then I gave AI a chapter, and I let it produce a podcast script. At first, the results seemed to be exactly what I was after. As I dug a little bit deeper, however, I started realizing that the AI had omitted some key details and left out entire sections, presumably in an effort to avoid exceeding a maximum token count.

But then I noticed something a lot more interesting -- output drift. I gave the AI tool another chapter from which to build a podcast script. This time, the script included the same types of errors and omissions as before, but there were also some minor deviations from the rules governing what I wanted the podcast script to look like.

I tried giving the AI a third chapter, and this time the results were completely unusable. The AI had changed the entire podcast format. Whereas before, the podcast script had been based on two anonymous participants casually discussing my book, the newly generated script was built around the idea that someone would be interviewing me during the podcast. There were also numerous other style rules that were violated. The question is, what happened?

It would be easy to look at a situation like this and blame a rolling context window. Every generative AI tool has context limits, so in a long conversation, older exchanges are purged in order to make way for new text. However, that is not what happened. In fact, I quizzed the AI about the rules, and it had no trouble remembering the rules that had been established, yet those rules were being flagrantly violated. So again, the question is why?

The answer to this question appears to have less to do with forgetting and more to do with how AI constantly reinterprets an ever-expanding context. Let's suppose for a moment that I am working with a human editor and my goal is to come up with a series of podcast scripts stemming from the 18 chapters in my book. Establishing the baseline rules for and the overall style of the podcast is going to involve quite a bit of work. And after all of that work, I would expect the editor to remember what we came up with. When it comes to an AI, however, hard-established rules can blend into the rest of the conversation rather than being treated as authoritative. The problem might best be described as context pollution.

Here is how it works. Imagine that I do the same thing as before. I give the AI a long list of rules, a short synopsis of the book and an explanation of what I am trying to generate. Although this might sound trivial, in my case, this prompt is about 3,000 words in length. Now, suppose that I give the AI a chapter from my book. The chapters vary in length, but let's pretend that the chapter is 5,000 words. Now, the AI responds with the requested podcast script. On average, these have been about 8,000 words in length. Just to keep things simple, let's pretend that the script that was generated is perfect and there are no problems with it.

So, at this point, there are about 16,000 words in the AI's context. So let's suppose that I give the AI another 5,000-word chapter and ask it to create a podcast script. Now, the AI has an interesting problem to deal with. It has to filter through the 21,000 words that are now in its context and figure out which of those words matter and which do not.

More importantly, the first podcast script that the AI produced is now a part of the conversation context. Remember, that script was AI-generated. It is not authoritative, but there is now a risk that it might be treated as such. This is when things get dangerous. Not everything in the context window should be treated with the same authority. The rules are authoritative. My book is authoritative. But the AI-generated content is not authoritative, and yet there is a risk of it being treated as just as authoritative as the book itself.

There was an old game that the teachers occasionally made us play in elementary school. The game was called Telephone. Maybe you've heard of it. The idea was that someone whispers something to the person on their left. That person whispers whatever they heard to the next person. By the time the message gets all the way across the room, it has completely changed. That is basically the same thing that was happening with the AI in this case. AI didn't forget the instructions; it just reinterpreted them alongside everything else that had been added to the conversation.

The bottom line is that when working on short projects, AI is extremely good at maintaining the illusion of continuity. However, longer projects reveal that there is a huge difference between conversational continuity and project continuity.

Of course, this raises the question of what you can do in these situations in an effort to get better results. When it comes to my specific project, there were three things that I did to vastly improve the AI output.

First, I never asked AI to deal with more than one chapter within a single session. When it was time to move on to a new chapter, I created a new session.

Second, I broke the lengthier chapters into segments and treated those segments as though they were chapters (never discussing more than one segment in a session). Honestly, the AI had no trouble understanding the complete chapter, but the podcast scripts tended to omit a lot of details because the AI was trying to create its response using no more than a specific number of tokens. Breaking longer chapters into segments allowed me to create lengthier podcast scripts.

Third, and this is the important one, I created a source of truth that lives outside of the conversation. Rather than expecting the AI to remember all of my criteria, I created a Word document containing the rules for the script, a synopsis of the book, a list of recurring themes and terminology and a list of recurring characters. I used this document as the basis for each new session, thereby ensuring that the chapter (or chapter segment) associated with that session would be treated in a way that was consistent with the other chapters.

About the Author

Brien Posey is a 22-time Microsoft MVP with decades of IT experience. As a freelance writer, Posey has written thousands of articles and contributed to several dozen books on a wide variety of IT topics. Prior to going freelance, Posey was a CIO for a national chain of hospitals and health care facilities. He has also served as a network administrator for some of the country's largest insurance companies and for the Department of Defense at Fort Knox. In addition to his continued work in IT, Posey has spent the last several years actively training as a commercial scientist-astronaut candidate in preparation to fly on a mission to study polar mesospheric clouds from space. You can follow his spaceflight training on his Web site.

Featured

Subscribe on YouTube