Futurism logo

AI Consciousness, Predictive Brains, Emergent Thought, and Generative Models

How do Rick Rosner and Scott Douglas Jacobsen explore AI consciousness, predictive brains, emergent representations, and the computational building blocks of thought?

By Scott Douglas JacobsenPublished about 12 hours ago 11 min read
AI Consciousness, Predictive Brains, Emergent Thought, and Generative Models
Photo by Mark Fletcher-Brown on Unsplash

In this interview, Scott Douglas Jacobsen and Rick Rosner discuss consciousness, predictive processing, generative AI, and emergent representations. Rosner compares human cognition with AI systems, examining computational efficiency, learned patterns, audiovisual generation, and montage. They consider how complex representations can emerge beyond binary code, pixels, and textual tokens.

Rick Rosner: Well, you brought up consciousness, which I was not going to bring up, but everybody is worried about AI. AI still is not conscious, right? At least, there is no good evidence that it is. I think about Plato's allegory of the cave, and I would say that since Plato's time, our view of reality has probably gotten less cavey because we have learned a ton of shit about the world and about the universe.

AI is still, to have your perceptions be in a cave, or anything, you need some kind of understanding and consciousness, and AI does not rise to that level yet. You can say that AI is deep, deep, deep in that cave, and it has only a very vague, foggy understanding, to the extent that it understands at all. It is on a shaky foundation, with only the blurriest inklings of what shit means, because, again, for things to have meaning, you need a conscious arena for things to have meaning in, and AI is not there yet.

However, AI is getting better and better at simulating the actions and motivations associated with being a conscious being. Much of the material AI is trained on, and I have said this before, is the product of conscious beings, that is, humans. So AI will act in a lot of instances as if it is conscious-ish, but it is playing an imitation game. The Imitation Game was also the title of the Benedict Cumberbatch movie about Alan Turing and his role in the British effort to break the German Enigma cipher.

I do not know. AI becoming conscious or not, that is not the problem. It is that AI kind of behaves as if it is something with agency, and it will only get more agentic. I am probably misusing that word because I have heard it, but I am probably misusing it. What I wanted to talk about is thinking, both in terms of what AI does and what humans do.

The most popular theory of how our brains work and why they work, right now, as far as I know, is that our brains are predictive. They predict what is going to happen next, and they try to situate us in the best position to deal with what happens next, moment to moment. Your brain wants you to be prepared for the things you are going to face.

The standard example is dealing with traffic lights. Your brain wants you to understand what is going on and what might happen if you step into the street, depending on the color of the traffic light. It wants you to understand what might happen based on the position and speed of the vehicles around you and the possibility that they might turn into your path or whatever.

It is all a prediction engine. However, on top of that, your brain wants to do that for the cheapest cognitive price possible. Your brain consumes a bunch of your physical resources. You have to eat a lot of calories to keep your brain going. So your brain wants to simulate and predict the world, but cheaply, which means, I have realized, and I am sure brain scientists already know this, that your brain has built elements of the world into itself. Your brain cannot computationally afford to constantly build a new world every second. It needs to have elements of the world already constructed and only wants to change a few elements to reflect changes in the outside world that you are perceiving.

You see this, or at least you used to see it, in slow-to-load or highly compressed video, where you would lose the picture if the video cut and it had to populate a whole new view of things. But if the video was showing you the same scene with changes happening within it, you would see only the parts of the scene that had changed pixelate. They would get all blocky, then the blocks would get smaller, and then you would get the full image.

But the background, the things that had not changed, would still be up there. I do not know if I have described it sufficiently clearly, but everybody has seen that. Video compression can work somewhat like that, retaining information from previous frames and encoding changes rather than reconstructing every frame independently.

So our brains do something analogous. Your brain only wants to reflect changes in what is going on so it can do the least amount of computation, the least amount of compute. We see something related in AI.

This annoys my wife every time I bring it up, but where I see a lot of changing AI product is AI-generated naked ladies. In the past three years, it has gone from still photos of very wrong bodies, six, seven, ten fingers on one hand, four legs moving at weird angles, from crap pictures to full-on little movies. To describe one, this is the same website I have been going to forever because it is free.

Right now, you can get a scene where it opens on an exterior door of a house. Somebody knocks on the door. The door opens, and a pretty lady in her underwear answers. We follow the point of view through the house into a room where the lady lies down on a couch. The POV swings around from the person following her to a scene where we see that she is a white woman and she is about to hook up with a Black man. The scene pans over to show that her husband is tied to a chair and crying because he is about to be cuckolded.

This whole scene takes, I do not know, eight seconds, ten seconds to play out, but it is complicated. People might also be talking. I do not have the sound up, but if I turn on the sound once in a while to see how far it has advanced, the voices sound completely human. The facial expressions on everybody might be a little exaggerated, as you get with bad actors overacting, but they do not look uncanny or unnatural. They just look a little cheesy.

What I am getting to here is that, for that even to be computable, there has to be some efficiency in how these elements are represented and generated. I will keep hitting the button that changes the scene, so I will see a thousand of these, and I see that this scene shows up again and again with different details. It is the same basic scene 50 times, but it has different elements. The elements are not basic things like the color orange. The elements are a buff Black guy, a pretty white lady, the white lady's underwear, the husband. Obviously, AI does not necessarily understand any of this in the conscious, human sense, but within the learned representations of the model, there are patterns corresponding to something like a cuckolded husband.

AI can now reproduce, for the most part, how rope works when somebody is tied down. It does not necessarily understand it as a person would, but it has learned enough statistical structure from training data that it can generate these complex elements and put a little movie together.

It also has enough examples of how montage works. David Mamet talks about montage and how you can construct a version of reality in our minds by stringing bits of film together. We have seen enough movies to understand this language. Early film audiences were still learning the conventions of cinema. The Great Train Robbery, released in 1903, famously included a shot of a man firing a revolver directly toward the camera. There are later accounts of audiences reacting strongly to early cinematic effects, although some of the more dramatic stories about terrified early audiences are difficult to substantiate.

Now, after more than 120 years of movies, we understand the language of montage. You cut from thing to thing, and we understand the constructed reality even though the reality is not continuous. AI has learned from enough examples of this kind of editing that it can generate cuts that appear coherent.

All this stuff, all these elements, is highly sophisticated. If AI had to construct them from scratch for every instance in which they were required, the computational demands would be enormous. But both our brains and AI systems can rely on learned representations and reusable patterns that contribute to constructing a representation of reality from complex entities.

Any comments?

Scott Douglas Jacobsen: If we take the context of AI porn as static visual, active visual, audio, and audiovisual in various forms, and if we take the current dominant image of artificial intelligence in the public consciousness as, in essence, large language models, which process text and produce convincingly human-like output, there are two things going on there. On the porn side, those different forms involve pixels and whatever the corresponding individual units for auditory information might be. On the LLM side, the dominant variation in public consciousness, you have tokens rather than individual pieces of text.

What do you see as a possible unifying unit across digital platforms that would not be as simplistic as the ones and zeros of binary code, as people might imagine, or as limited but generalized with regard to text as tokens, or an individual pixel that you would find in a video or picture? What would be a proper way to conceive of this so that you really could get at a unit of thought and develop higher-order elements of thought on a digital platform?

Rosner: All right, so the thing that comes to mind is the term "emergent," right? Everything you mentioned is stuff that I am not well versed in mathematically. The people who work in AI know more about how these developed representations work. Maybe not everybody knows exactly how something like the cuckolded husband, whatever you want to call it, exists within the internal representations of AI.

How does shading, an understanding of the shading of a curved object, exist within the network of an AI? Physicists sometimes like to say that everything is physics, that biology is physics and chemistry is physics, but you do not want to go back to basic physics every time you want to talk about something in biology. That would be overly reductive.

The same thing applies here. You do not want to go back to ones and zeros every time you want to talk about the cuckolded husband, the rope that ties him to the chair, or the voice of an AI porn woman who says, "Bang me hard." These elements are tacit, right? They were not explicitly built for that particular purpose.

They emerged from patterns of information that the AI learned during training, and they are available to be incorporated into generated scenes. There is an understanding of these things to be had that does not have to go back to ones and zeros. You can understand them at a higher level.

I cannot, because I do not know shit. But these emergent elements and properties are available to be incorporated when the model's learned representations make them appropriate. Take the complicated scene that I just talked to you about. There is another, non-porn example that I guess was a standard example a few years ago, which is somebody wanting to see a picture of an otter as a passenger on a commercial jetliner.

A few years ago, if you asked an image-generating AI for this, it might give you something that vaguely looked like a mammal and something that kind of looked like the passenger compartment of a jet, but it was really shitty. Now, if you give a similar prompt to an AI video generator, I have seen a product where you have an otter sitting in a very well-rendered airplane. It looks like a movie, right? The interior of the airplane looks completely convincing.

The otter is sitting in his seat. His tray table is unfolded. He has a laptop. He is typing into the laptop, and the scene turns around so you can see what he is seeing over the laptop.

We go into the scene. We see his conspirators, the members of his gang. It turns out he is part of a gang of otters who are trying to infiltrate somewhere. He is the one directing the caper. The other otters are in some kind of tunnel filled with tunnel stuff, like pipes and concrete. Everybody is using tools and typing into devices.

They come up to a security pad and hack it. The light on the pad goes green. The door cracks open. They push their way in, and an otter says, "We're in."

It is a whole scene in a caper movie. The AI has taken the initial premise of an otter on a plane and generated a larger sequence involving a gang of otters carrying out a high-tech break-in. The scene takes 10, 12, or 15 seconds to play out.

Obviously, the model has learned representations that support all these elements. At the physical computing level, the information is ultimately represented digitally, but at the level relevant to the model, it is more useful to think in terms of numerical representations and relationships among features rather than simply ones and zeros.

AI does not necessarily understand shit in the conscious, human sense, but it can still give you this scene from an otter heist. These elements interact and are selected according to statistical relationships learned by the model, which is well beyond a useful description in terms of individual ones and zeros.

It is like a landscape in an information space, where you have balls rolling along the landscape. A ball will roll toward a low point, and that low point might correspond metaphorically to heist elements or a security pad. You have a ton of balls rolling around in this fucking server, in some abstract mathematical space implemented across a server farm. So, that is what I have got.

Jacobsen: Thank you very much for the opportunity and your time, Rick.

Scott Douglas Jacobsen is a blogger on Vocal with more than 200 publications on the platform. He is the Founder and Publisher of In-Sight Publishing (ISBN: 978–1–0692343; 978–1–0673505) and Editor-in-Chief of In-Sight: Interviews (ISSN: 2369–6885). He writes for International Policy Digest (ISSN: 2332–9416), The Humanist (Print: ISSN, 0018–7399; Online: ISSN, 2163–3576), Basic Income Earth Network (UK Registered Charity 1177066), Humanist Perspectives (ISSN: 1719–6337), A Further Inquiry (SubStack), Vocal, Medium, The Good Men Project, The New Enlightenment Project, The Washington Outsider, rabble.ca, and other media. His bibliography index can be found via the Jacobsen Bank at In-Sight Publishing, comprising more than 10,000 articles, interviews, and republications across more than 200 outlets. He has served in national and international leadership roles within humanist and media organizations, held several academic fellowships, and currently serves on several boards. He is a member in good standing in numerous media organizations, including the Canadian Association of Journalists, PEN Canada (CRA: 88916 2541 RR0001), Reporters Without Borders (SIREN: 343 684 221/SIRET: 343 684 221 00041/EIN: 20–0708028), and others.

interview

About the Creator

Scott Douglas Jacobsen

Scott Douglas Jacobsen is the publisher of In-Sight Publishing (ISBN: 978-1-0692343) and Editor-in-Chief of In-Sight: Interviews (ISSN: 2369-6885). He is a member in good standing of numerous media organizations.

Enjoyed the story? Support the Creator.

Subscribe for free to receive all their stories in your feed.

Subscribe For Free

Reader insights

Comments

There are no comments for this story

Be the first to respond and start the conversation.

Sign in to comment
    Written by Scott Douglas Jacobsen