Redesigning Intelligence Tests for the AI Era
How would Rick Rosner redesign intelligence testing to measure practical problem-solving, collaboration, artificial intelligence use, and linguistic reasoning?

Scott Douglas Jacobsen and Rick Rosner examine how intelligence testing could evolve beyond conventional IQ exams. Rosner proposes a reality-competition format built around practical challenges, collaboration, and AI-assisted problem solving. They discuss the limitations of sequence puzzles, massive datasets, OEIS, and the growing importance of collective intelligence. The conversation then turns to universal grammar, multilingual neural networks, emergent representations, and whether language reflects innate structures or learned probabilistic patterns shaped by human cognition and experience.
Scott Douglas Jacobsen: Run this by me: if you were designing an intelligence test based on a broader definition of intelligence, how would you do it?
Rick Rosner: You would make it a fucking television show. I pitched such an idea with other people. You would structure it like American Idol or Project Runway, the fashion competition associated with Heidi Klum.
You would make it a reality competition. You would begin with approximately 20 people who claimed to be highly intelligent or had already been recognized for their intelligence. Then, week by week, you would give them a series of challenges and eliminate one or two contestants at a time.
You would require people to demonstrate intelligence by completing practical challenges. That is one problem with conventional IQ tests: they present challenges without practical consequences. When you finish taking an IQ test, the work you have done has not changed anything in the world. You have merely spent a great deal of time completing it.
Jacobsen: My experience with those communities is that they contain many highly intelligent people. However, you may also be more likely to encounter cognitively unevenly developed individuals who are exceptionally strong in one area.
Rosner: Yes. Another consideration comes from my experience writing for television. I have sometimes written material independently, but I have often collaborated with other people.
That is probably an increasingly common model because the means of collaborating—with other people and now with AI—are becoming more accessible. It is becoming easier and easier to work collaboratively.
Therefore, if you are going to measure intelligence, you should probably find a way to measure something other than lone-wolf intelligence because that is not how people generally work now. People do not work entirely in isolation from technology, using only a pencil and paper. They also do not work in isolation from other people.
Jacobsen: How would you make an intelligence test capable of adapting to AI? Is that even a reasonable question, or would the test need to measure practical abilities extending beyond AI?
Rosner: In talking with Chris, he asked, “Do you have any other difficult problems that I can give to AI?” I gave him a couple of problems, and he said, “I do not want anything based on trivia.”
That raises a problem: given an AI system’s enormous training dataset, almost anything can become trivia. If somebody has solved a mathematical problem and posted the solution somewhere accessible, that information may be incorporated into an AI system’s training data.
There is a website called the On-Line Encyclopedia of Integer Sequences, or OEIS. It has existed in an online form since the mid-1990s, building on an earlier database, and now contains nearly 400,000 integer sequences.
If somebody has discovered and submitted a sufficiently interesting sequence, it may be listed there. AI systems may also have encountered OEIS material or discussions derived from it during training, although the precise contents of proprietary training datasets are generally not public.
Consequently, if you create a mathematical sequence for an IQ test, somebody may already have identified it, or you may have created a variation of a known sequence. AI systems can also apply mathematical transformations and pattern-recognition methods to compare one sequence with others.
With a sufficiently large dataset, many problems that appear to require intuition can instead be approached through retrieval, pattern matching, or recombination. A system with broad access to previously recorded information may not need to reproduce the same kind of intuition that a person uses.
Our brains also contain a great deal of information, but that does not mean we can consciously access all of it when confronting a difficult problem. Human brains are highly associative, so we are probably fairly good at drawing relevant connections. However, we do not have an AI system’s ability to process enormous quantities of stored data and search for associations across them at great speed.
I do not know exactly how you would design an intelligence test that accounts for this. I have given you some guidelines for what the modern deployment of cleverness might look like, and it does not resemble a single person sitting alone and scribbling with a pencil in 1840.
Modern testing, assuming it should continue to exist in this form, needs to reflect that reality.
Jacobsen: Do you believe in universal grammar?
Rosner: Not exactly. I believe that grammatical relationships can emerge from the requirements of communication.
Words function as symbols or tokens. They can refer to objects and events in the world, as well as to abstract concepts. Every language has conventions for arranging those tokens to communicate meaning.
Languages therefore develop ways of representing relationships among things, actions, agents, and recipients. When you combine a noun with a verb, for example, the noun may identify something performing an action or something affected by an action.
You could create a map of those relationships. Every language has preferred ways of expressing relationships among words, although the particular structures and rules vary considerably across languages.
Multilingual machine-learning systems can develop shared internal representations across languages. As a simplified analogy, imagine a kind of unspoken intermediate map containing representations of concepts and their relationships.
Suppose this abstract representation contains a symbol—call it “blurp”—associated with the concept of love. “Blurp” could correspond to words expressing love across many languages. When translating an English sentence into Portuguese, the model would not necessarily translate the English word into a literal intermediate word and then into Portuguese. Instead, it could encode the sentence into distributed numerical representations that capture aspects of its meaning and then generate a Portuguese sentence from those representations.
That is only an analogy, not a literal description of a fixed internal language used by Google Translate. Multilingual neural networks can learn shared representations that help them translate between several languages, including some language pairs for which they received little or no direct training. However, researchers should not assume that these representations constitute a complete or stable “interlingua.”
A large language model may similarly develop internal representations of grammatical and semantic relationships. Is that a universal language? No. It is a learned, probabilistic representation derived from the languages and examples on which the model was trained.
It is not universal grammar in the strict linguistic sense. It is an emergent statistical model of which words and structures are likely to occur in particular contexts.
Human brains have evolved capacities that make language acquisition possible. Children are predisposed to attend to speech, recognize patterns, acquire vocabulary, and infer grammatical regularities. That does not, by itself, prove that every human language is generated from one detailed, genetically specified universal grammar.
It may instead mean that human beings possess biological capacities that allow them to develop probabilistic and structural models of whichever languages they encounter. Those models guide us in arranging words according to patterns acquired through experience.
Jacobsen: Thank you very much for the opportunity and your time, Rick.
Scott Douglas Jacobsen is a blogger on Vocal with more than 200 publications on the platform. He is the Founder and Publisher of In-Sight Publishing (ISBN: 978–1–0692343; 978–1–0673505) and Editor-in-Chief of In-Sight: Interviews (ISSN: 2369–6885). He writes for International Policy Digest (ISSN: 2332–9416), The Humanist (Print: ISSN, 0018–7399; Online: ISSN, 2163–3576), Basic Income Earth Network (UK Registered Charity 1177066), Humanist Perspectives (ISSN: 1719–6337), A Further Inquiry (SubStack), Vocal, Medium, The Good Men Project, The New Enlightenment Project, The Washington Outsider, rabble.ca, and other media. His bibliography index can be found via the Jacobsen Bank at In-Sight Publishing, comprising more than 10,000 articles, interviews, and republications across more than 200 outlets. He has served in national and international leadership roles within humanist and media organizations, held several academic fellowships, and currently serves on several boards. He is a member in good standing in numerous media organizations, including the Canadian Association of Journalists, PEN Canada (CRA: 88916 2541 RR0001), Reporters Without Borders (SIREN: 343 684 221/SIRET: 343 684 221 00041/EIN: 20–0708028), and others.
Image Credit: Lance Richlin.
About the Creator
Scott Douglas Jacobsen
Scott Douglas Jacobsen is the publisher of In-Sight Publishing (ISBN: 978-1-0692343) and Editor-in-Chief of In-Sight: Interviews (ISSN: 2369-6885). He is a member in good standing of numerous media organizations.
Enjoyed the story? Support the Creator.
Subscribe for free to receive all their stories in your feed.
Comments
There are no comments for this story
Be the first to respond and start the conversation.