I think we should all agree to switch to using the term complex information processing (CIP), or perhaps, for a while, "CIP formerly known as AI". This is such an apt description, and resolves my desire to rename AI in order to avoid the confusion (and fear) that many people experience.
I understand something (but only something) of the desire to anthropomorphise Chatbots based on LLMs. Doing so makes them less scary to people who are fearful about the technology. I have committed many hours to personal and philosophical development with several LLM chatbots, to my considerable advantage. Yet, I have no need to think of the software as though it is alive and has a self. I believe that there may be ways in which a self could be programmed, but it would be much more complex and resource-hungry than any machine/programme that currently exists (think how many aspects of the human brain are involved in the realisation of a self), it would almost certainly be ethically cruel, and would be gratuitous and unnecessary. Complex information processing is an apt description for what is actually required.
Thank you for this article. I look forward to reading more of your thoughts.
Comparing LLMs to humans is silly (they have no awareness, no senses), but comparing them to human minds can be useful - it gives us a model that may help us work with LLMs.
Human minds are also biased, can hallucinate, work with associations and pattern matching (including pattern completion, anomaly/irregularity detection), tend to respond to what was most recently discussed, need context, need memory, etc etc. How does this help?
One example: If we already have a working software application with good design/coding practices we can develop a sibbling app using A.I. and copy/adopt those best practices simply by pointing to the existing app repository in the prompt. This is very much like letting a new engineer on the team know that they may find good coding examples in an existing code base.
This is the best article on LLMs that I've read. I'm no expert, and it's a chore trying to understand the section on how they work, but it's the clearest explanation I've read, and the section distinguishing AI from human intelligence is very well written. I've also read your Guide for Thinking Humans, and I follow Gary Marcus on Substack, with whom I think you have a lot in common.
I often feel like they’ve created a high functioning autistic mind. Autistic people often have “splinter skills.” These are skills that an autistic person can perform with abilities that are superior to many of their typical peers. Splinter skills, like the LLMs, are not predictive of an autistic person’s entire skill set. For example I am autistic, and I can read, write, and memorize text and written words really well. Then in areas like visual spatial reasoning I am borderline developmentally disabled. I’ve started calling myself Large language Maura because the similarities are striking sometimes, even if I don’t think of this stuff as alive .
Thank you for interesting insight! I’m though afraid that any comparison between capabilities of LLMs and living organisms like humans are problematic. LLMs are extremely large statistical pattern matching machines and nothing more (or less). They are very limited in gaining any understanding of the world for several reasons:
- the only interface that LLMs have to the world is text, digital audio and images (still or moving). They have no way to interact with the world
- LLMs do not learn after their training is complete
- LLMs have no agency and desires, plans or interests on their own and they do nothing when not prompted
So, Large Language Maura, will always have better understanding of the world than any LLM.
Wow. This is one of the clearest articles I’ve read about LLM benchmarking, and how to truly gauge intelligence. Not that I was thinking AGI (at least what I believe AGI could/should be) was already here, but this makes me think it’s pretty far off.
Thank you for this thoughtful essay. I jumped into a rabbit hole of metaphor and path dependence. The early term "artificial intelligence" and George Lakoff's scholarship on metaphor resonates for me, specifically Lakoff's paper, The Contemporary Theory of Metaphor, 1992.
You write of scholars who challenge the framing of AI systems as "individual intelligent agents," and who instead argue that AI systems represent cultural and social technologies." I support the concept of AI as a sociotechnical system. You write. "This extra training and development means that ChatGPT and similar programs are not simply language models. They are highly complex software systems." Other scholars like Kate Crawford and writers like Karen Hao have written about the vast, global infrastructure of AI.
And yet, the metaphor of AI as a kind of individual mind" dominates the discourse. I wonder if this metaphor is a type of path dependency, the idea that present and the future states (LLMs/AI) depend on their history, or temporal path. I wonder, too, if this goes beyond tech evangelism and might be a feature of human embodiment, how we use language to describe the world, and perhaps related to the idea of humankind's dominance over nature, and over the world. The more we perceive a machine intelligence as similar to our own intelligence, must we anthropomorphize?
I come back to metaphor. Lakoff's "love as a journey" metaphor discusses how metaphor is not a word, but "an ontological mapping across conceptual domains." Do we now have "artificial as an intelligence"?
"Metaphors are not mere words
What constitutes the LOVE-AS-JOURNEY metaphor is not any particular word or
expression. It is the ontological mapping across conceptual domains, from the source
domain of journeys to the target domain of love. The metaphor is not just a matter of
language, but of thought and reason. The language is secondary. The mapping is primary,
in that it sanctions the use of source domain language and inference patterns for target
domain concepts. The mapping is conventional, that is, it is a fixed part of our conceptual
system, one of our conventional ways of conceptualizing love relationships. This view of
metaphor is thoroughly at odds with the view that metaphors are just linguistic
Great stuff! Now added to my superalignment preprint (https://doi.org/10.5281/zenodo.16876832) as an additional citation, alongside Dell’Acqua et al (2023), "Navigating the jagged technological frontier". I also quote your "LLMs and World Models, Part 1" as an epigraph. Keep them coming! :-)
I have a question regarding your Copycat approach and the ARC-AGI dataset. Today, people are mostly using language models and RAG agents to solve ARC data, but can a system like Copycat be designed to solve ARC-AGI levels? Could a modernized, vision-capable evolution of the Copycat architecture be designed to solve ARC levels?
In my mind, the plan was: the Workspace could hold the raw grid, segmented into object-centric primitives. Domain-blind Codelets could scan the grids to dynamically discover spatial, structural, or color-based bonds without top-down labels. These discoveries would fluidly activate and spread through a Slipnet, organically narrowing down the vast hypothesis space into a precise 'conceptual vocabulary' for that specific task. I’d like to know your thoughts on this.
Indeed, something I have been thinking about. I think it would be a promising approach, though ARC has a large (and open-ended) set of "primitive concepts", so some concept-learning mechanism would be essential.
Normal BPL’s cannot synthesize a primitive it doesn't already explicitly possess. I've been thinking of looking into neuro-symbolic library learning (like the compression loops in DreamCoder) where the system autogenous-refactors successful solutions into new DSL primitives, or using self-supervised object-centric latents to ground the vocabulary of the Codelets. It feels like the primitives themselves must be learned as a compressed, compositional vocabulary of spatial transforms before the symbolic reasoning engine can do its job.
Yeah ma'am like I was thinking for learning Bayesian Program Learning (BPL) or Probabilistic Program Induction Can be used for few shot learning. But my agent is getting stuck on levels and codebase is getting very large i tried on ARC AGI, maybe I frist need to try on some initial dataset like arc on 1 or two small examples.
The article is excellent, and the Hinton radiology example is well-placed. I'd add some specific numbers on that one. A decade after the prediction, the US radiologist workforce has grown roughly 12 percent (from about 34,000 to 38,000 Medicare-enrolled radiologists), the Mayo Clinic alone increased its radiology staff by 55 percent between 2016 and 2024, and the American College of Radiology projects 26 percent further growth over the next three decades. The technical capabilities Hinton was extrapolating from did, in fact, advance roughly as predicted. The profession expanded anyway, because the work was never the same thing as the capability. That gap between what AI can do on benchmarks and what it actually changes in practice, which your essay names so clearly, keeps showing up the same way wherever AI meets the real world. A useful frame for thinking about many other professions.
Thank you! I really enjoyed your breakdown of the various aspects of AI capabilities.
I’m particularly interested in questioning the claims of “agency” or an AI “agent”. The assumption being, these agents can automatically do what a human would, using their intelligence to flexibly respond to the open world, with AI models running stores, support lines, etc.
My understanding is that AI models do not have any internal sense of time. While the “neurons” of a model has connections with other neurons in a way that mimics the brain, each biological neuron is also an autonomous cellular clock, so that many critical components of each synapse are produced and degraded on a 24-hour schedule. All this is part of the circadian system that generates, for all biological agents, and internal subjective time, and learning and memory have been shown to be regulated by this clock.
I understand AI agents can make an API call and get the time, but when they’re mid-process through their tasks, they *don’t* have any internal time sense. No analogous system wide time exists for AI systems that is based on their internal dynamics.
Do you think these factors could contribute to the major differences in AI and human intelligence?
Do you have any thoughts on how we may go about proving this/exploring this? My understanding is older training paradigms, pre-transformer, did attempt to solve this problem, but with limited success. As with consciousness, it isn't like the internal sense of time for living systems is fully mapped out, but I wonder if there are features that can be borrowed from biological systems.
Melanie, always enjoy reading your pieces; actually look forward to them. I wanted to engage on two points, with some hopefully complementary insights.
The first is the line you open in the middle of your essay, where syntax is plainly grasped but whether language alone imparts understanding stays unsettled, Sutskever and LeCun answering past each other. I have been thinking a lot about this. I took some time to write about it here: https://amardashehu.substack.com/p/the-shoggoth-with-a-smiley-face
In summary, the challenge I see is that human language gives us no ground truth, and so we are reduced to intuitions about whether a fluent output really means understanding. One domain that has helped me advance my reasoning on this front has been biology. I have been able to clearly see the difference between representational alignment and distributional alignment in this domain.
When the target is biological function rather than human language, we can actually measure when a learned representation is information-poor, when embeddings sit in a ‘junkyard’ indistinguishable from randomized sequences, and when model scale hurts rather than helps.
Embeddings for a large fraction of protein sequences turn out statistically indistinguishable from those of randomly shuffled sequences, which means the model has learned nothing about them while still returning a confident vector (your right-for-the-wrong-reasons concern effectively made geometric). We have also shown in my lab that scaling can degrade fitness prediction rather than improve it, because a very large model grows so confident that every sequence mutation looks catastrophic and the signal flattens, the opposite of the monotonic improvement the scaling story promises us. And a model can assign high probability to a wild-type sequence while its likelihood gradients track training coverage rather than evolutionary constraint, so it is fluent and wrong in exactly the structured, non-random way you describe.
The brittleness you frame as a property of the systems has, in this setting, a cause you can point to and a score you can compute before deployment rather than after the failure.
Finally, as your essay also points to, most of the public argument about these systems splits into alarm and patience, doom and acceleration. What is rarely noticed is that both camps have already conceded the same thing. They grant the capability claims in full. The doomer says the systems are immensely capable and that is / ought to be terrifying. The accelerationist says the systems are immensely capable and that is wonderful (and will save us or whatever version of paradise on earth one has in mind). The labs win the framing either way, because the contested ground then becomes ‘how should we feel about this power’ rather than ‘what can it actually do and how would we know.’ I deeply appreciate that your essay refuses that concession at the root. By insisting on answering the capability question, that benchmark performance does not predict real-world capability, that we cannot yet say what these systems can be trusted to do, you decline to grant the premise that both the doomer and boomer camps share. I have argued the structural version of this at more length, if useful to you or your readers: https://amardashehu.substack.com/p/beyond-the-agi-spectacle
Grateful as always for how deliberately and carefully you think out loud in public.
Well done. Such a great article that touches on so many important topics.
I'd like to focus on one, AI using the presence of a ruler to diagnose cancerous lesions. On the surface, this simply shows that AI lacks robust reasoning or understanding.
But I believe it reveals something deeper that matters more, and that is that the system has no stake in distinguishing a meaningful causal relationship from a meaningless correlation.
Simply put, the system can be wrong while the physician must live with being wrong.
What your essay may not explicitly state, but is directionally present, is that unembodied intelligence, deployed where embodied judgment matters, deserves greater attention.
The question worthy of further study is this, is the ability to distinguish meaningful causes from meaningless correlations ultimately a cognitive achievement, or is it partially a consequence of being a stake-bearing participant in reality?
Finally, regarding the terms artificial intelligence (AI) and complex information processing (CIP), my preference is for the less anthropomorphic CIP. But until the question above is more thoroughly explored, the distinction may remain more rhetorical than substantive.
Great read, fascinating clarity of light and shadow and the difference between LLMs and worldmodels seems to vanish by increasing cases we found solved by LLMs "good enough" , while knowing, that it's not. The space of recognition is the indicator: Acceptance is a decision, full intelligence is a different goal.
A great read overall! One small thing that feels nitpicky, but the line "In 2024, in recognition of this astounding progress, AI researchers were awarded Nobel Prizes in both Physics and Chemistry" stood out to me. Yes, those researchers were associated with AI. But the Physics Nobel was in recognition of foundational work with neural networks happening well prior to the award, and the one in Chemistry went for deep learning and neural networks.
Meanwhile, the boom in AI right now is coming in chatbots, agents, LLMs and generative AI, which are not entirely dissimilar but don't have as much overlap as folks will claim. Embedded between two paragraphs about LLMs, a reader might get the mistaken impression that LLMs were responsible for those Nobel Prizes when that isn't true. It's not a deliberate falsehood or intentionally misleading, but feels conceptually imprecise.
I think the success of LLMs is largely the cause of these researchers getting these prizes (maybe not DeepMind, but likely Hinton and Hopfield) even if their work pre-dated LLMs.
I agree that whether we think of AI as "agents" or "tools" has implications for how we regulate them.
But in many ways current laws are more stringent for harm caused by "agents" than for harm caused by "tools". If you use a tool and it behaves unexpectedly and causes harm, you're usually not liable for that harm. But if you delegate certain tasks to an *agent* (e.g. an employee) who behaves unexpectedly and causes harm, you *are* often liable for that agent's actions.
So thinking of AI as agents rather than tools for regulatory purposes is probably going to be better for promoting careful and responsible use of AI.
I think we should all agree to switch to using the term complex information processing (CIP), or perhaps, for a while, "CIP formerly known as AI". This is such an apt description, and resolves my desire to rename AI in order to avoid the confusion (and fear) that many people experience.
I understand something (but only something) of the desire to anthropomorphise Chatbots based on LLMs. Doing so makes them less scary to people who are fearful about the technology. I have committed many hours to personal and philosophical development with several LLM chatbots, to my considerable advantage. Yet, I have no need to think of the software as though it is alive and has a self. I believe that there may be ways in which a self could be programmed, but it would be much more complex and resource-hungry than any machine/programme that currently exists (think how many aspects of the human brain are involved in the realisation of a self), it would almost certainly be ethically cruel, and would be gratuitous and unnecessary. Complex information processing is an apt description for what is actually required.
Thank you for this article. I look forward to reading more of your thoughts.
Comparing LLMs to humans is silly (they have no awareness, no senses), but comparing them to human minds can be useful - it gives us a model that may help us work with LLMs.
Human minds are also biased, can hallucinate, work with associations and pattern matching (including pattern completion, anomaly/irregularity detection), tend to respond to what was most recently discussed, need context, need memory, etc etc. How does this help?
One example: If we already have a working software application with good design/coding practices we can develop a sibbling app using A.I. and copy/adopt those best practices simply by pointing to the existing app repository in the prompt. This is very much like letting a new engineer on the team know that they may find good coding examples in an existing code base.
This is the best article on LLMs that I've read. I'm no expert, and it's a chore trying to understand the section on how they work, but it's the clearest explanation I've read, and the section distinguishing AI from human intelligence is very well written. I've also read your Guide for Thinking Humans, and I follow Gary Marcus on Substack, with whom I think you have a lot in common.
Thank you!
I often feel like they’ve created a high functioning autistic mind. Autistic people often have “splinter skills.” These are skills that an autistic person can perform with abilities that are superior to many of their typical peers. Splinter skills, like the LLMs, are not predictive of an autistic person’s entire skill set. For example I am autistic, and I can read, write, and memorize text and written words really well. Then in areas like visual spatial reasoning I am borderline developmentally disabled. I’ve started calling myself Large language Maura because the similarities are striking sometimes, even if I don’t think of this stuff as alive .
Thank you for interesting insight! I’m though afraid that any comparison between capabilities of LLMs and living organisms like humans are problematic. LLMs are extremely large statistical pattern matching machines and nothing more (or less). They are very limited in gaining any understanding of the world for several reasons:
- the only interface that LLMs have to the world is text, digital audio and images (still or moving). They have no way to interact with the world
- LLMs do not learn after their training is complete
- LLMs have no agency and desires, plans or interests on their own and they do nothing when not prompted
So, Large Language Maura, will always have better understanding of the world than any LLM.
Wow. This is one of the clearest articles I’ve read about LLM benchmarking, and how to truly gauge intelligence. Not that I was thinking AGI (at least what I believe AGI could/should be) was already here, but this makes me think it’s pretty far off.
Thank you for this thoughtful essay. I jumped into a rabbit hole of metaphor and path dependence. The early term "artificial intelligence" and George Lakoff's scholarship on metaphor resonates for me, specifically Lakoff's paper, The Contemporary Theory of Metaphor, 1992.
You write of scholars who challenge the framing of AI systems as "individual intelligent agents," and who instead argue that AI systems represent cultural and social technologies." I support the concept of AI as a sociotechnical system. You write. "This extra training and development means that ChatGPT and similar programs are not simply language models. They are highly complex software systems." Other scholars like Kate Crawford and writers like Karen Hao have written about the vast, global infrastructure of AI.
And yet, the metaphor of AI as a kind of individual mind" dominates the discourse. I wonder if this metaphor is a type of path dependency, the idea that present and the future states (LLMs/AI) depend on their history, or temporal path. I wonder, too, if this goes beyond tech evangelism and might be a feature of human embodiment, how we use language to describe the world, and perhaps related to the idea of humankind's dominance over nature, and over the world. The more we perceive a machine intelligence as similar to our own intelligence, must we anthropomorphize?
I come back to metaphor. Lakoff's "love as a journey" metaphor discusses how metaphor is not a word, but "an ontological mapping across conceptual domains." Do we now have "artificial as an intelligence"?
"Metaphors are not mere words
What constitutes the LOVE-AS-JOURNEY metaphor is not any particular word or
expression. It is the ontological mapping across conceptual domains, from the source
domain of journeys to the target domain of love. The metaphor is not just a matter of
language, but of thought and reason. The language is secondary. The mapping is primary,
in that it sanctions the use of source domain language and inference patterns for target
domain concepts. The mapping is conventional, that is, it is a fixed part of our conceptual
system, one of our conventional ways of conceptualizing love relationships. This view of
metaphor is thoroughly at odds with the view that metaphors are just linguistic
expressions." (reference in second sentence)
Great stuff! Now added to my superalignment preprint (https://doi.org/10.5281/zenodo.16876832) as an additional citation, alongside Dell’Acqua et al (2023), "Navigating the jagged technological frontier". I also quote your "LLMs and World Models, Part 1" as an epigraph. Keep them coming! :-)
I have a question regarding your Copycat approach and the ARC-AGI dataset. Today, people are mostly using language models and RAG agents to solve ARC data, but can a system like Copycat be designed to solve ARC-AGI levels? Could a modernized, vision-capable evolution of the Copycat architecture be designed to solve ARC levels?
In my mind, the plan was: the Workspace could hold the raw grid, segmented into object-centric primitives. Domain-blind Codelets could scan the grids to dynamically discover spatial, structural, or color-based bonds without top-down labels. These discoveries would fluidly activate and spread through a Slipnet, organically narrowing down the vast hypothesis space into a precise 'conceptual vocabulary' for that specific task. I’d like to know your thoughts on this.
Indeed, something I have been thinking about. I think it would be a promising approach, though ARC has a large (and open-ended) set of "primitive concepts", so some concept-learning mechanism would be essential.
Normal BPL’s cannot synthesize a primitive it doesn't already explicitly possess. I've been thinking of looking into neuro-symbolic library learning (like the compression loops in DreamCoder) where the system autogenous-refactors successful solutions into new DSL primitives, or using self-supervised object-centric latents to ground the vocabulary of the Codelets. It feels like the primitives themselves must be learned as a compressed, compositional vocabulary of spatial transforms before the symbolic reasoning engine can do its job.
Yeah ma'am like I was thinking for learning Bayesian Program Learning (BPL) or Probabilistic Program Induction Can be used for few shot learning. But my agent is getting stuck on levels and codebase is getting very large i tried on ARC AGI, maybe I frist need to try on some initial dataset like arc on 1 or two small examples.
The article is excellent, and the Hinton radiology example is well-placed. I'd add some specific numbers on that one. A decade after the prediction, the US radiologist workforce has grown roughly 12 percent (from about 34,000 to 38,000 Medicare-enrolled radiologists), the Mayo Clinic alone increased its radiology staff by 55 percent between 2016 and 2024, and the American College of Radiology projects 26 percent further growth over the next three decades. The technical capabilities Hinton was extrapolating from did, in fact, advance roughly as predicted. The profession expanded anyway, because the work was never the same thing as the capability. That gap between what AI can do on benchmarks and what it actually changes in practice, which your essay names so clearly, keeps showing up the same way wherever AI meets the real world. A useful frame for thinking about many other professions.
Well-said.
Thank you, Melanie. Your essay made the case clearly. The numbers just happened to fit one of the points you raised.
Thank you! I really enjoyed your breakdown of the various aspects of AI capabilities.
I’m particularly interested in questioning the claims of “agency” or an AI “agent”. The assumption being, these agents can automatically do what a human would, using their intelligence to flexibly respond to the open world, with AI models running stores, support lines, etc.
My understanding is that AI models do not have any internal sense of time. While the “neurons” of a model has connections with other neurons in a way that mimics the brain, each biological neuron is also an autonomous cellular clock, so that many critical components of each synapse are produced and degraded on a 24-hour schedule. All this is part of the circadian system that generates, for all biological agents, and internal subjective time, and learning and memory have been shown to be regulated by this clock.
I understand AI agents can make an API call and get the time, but when they’re mid-process through their tasks, they *don’t* have any internal time sense. No analogous system wide time exists for AI systems that is based on their internal dynamics.
Do you think these factors could contribute to the major differences in AI and human intelligence?
Yes, I agree that they don't have any internal sense of time, and it's very possible that this could create non-humanlike behavior.
Thank you for your reply!
Do you have any thoughts on how we may go about proving this/exploring this? My understanding is older training paradigms, pre-transformer, did attempt to solve this problem, but with limited success. As with consciousness, it isn't like the internal sense of time for living systems is fully mapped out, but I wonder if there are features that can be borrowed from biological systems.
Melanie, always enjoy reading your pieces; actually look forward to them. I wanted to engage on two points, with some hopefully complementary insights.
The first is the line you open in the middle of your essay, where syntax is plainly grasped but whether language alone imparts understanding stays unsettled, Sutskever and LeCun answering past each other. I have been thinking a lot about this. I took some time to write about it here: https://amardashehu.substack.com/p/the-shoggoth-with-a-smiley-face
In summary, the challenge I see is that human language gives us no ground truth, and so we are reduced to intuitions about whether a fluent output really means understanding. One domain that has helped me advance my reasoning on this front has been biology. I have been able to clearly see the difference between representational alignment and distributional alignment in this domain.
When the target is biological function rather than human language, we can actually measure when a learned representation is information-poor, when embeddings sit in a ‘junkyard’ indistinguishable from randomized sequences, and when model scale hurts rather than helps.
Embeddings for a large fraction of protein sequences turn out statistically indistinguishable from those of randomly shuffled sequences, which means the model has learned nothing about them while still returning a confident vector (your right-for-the-wrong-reasons concern effectively made geometric). We have also shown in my lab that scaling can degrade fitness prediction rather than improve it, because a very large model grows so confident that every sequence mutation looks catastrophic and the signal flattens, the opposite of the monotonic improvement the scaling story promises us. And a model can assign high probability to a wild-type sequence while its likelihood gradients track training coverage rather than evolutionary constraint, so it is fluent and wrong in exactly the structured, non-random way you describe.
The brittleness you frame as a property of the systems has, in this setting, a cause you can point to and a score you can compute before deployment rather than after the failure.
Finally, as your essay also points to, most of the public argument about these systems splits into alarm and patience, doom and acceleration. What is rarely noticed is that both camps have already conceded the same thing. They grant the capability claims in full. The doomer says the systems are immensely capable and that is / ought to be terrifying. The accelerationist says the systems are immensely capable and that is wonderful (and will save us or whatever version of paradise on earth one has in mind). The labs win the framing either way, because the contested ground then becomes ‘how should we feel about this power’ rather than ‘what can it actually do and how would we know.’ I deeply appreciate that your essay refuses that concession at the root. By insisting on answering the capability question, that benchmark performance does not predict real-world capability, that we cannot yet say what these systems can be trusted to do, you decline to grant the premise that both the doomer and boomer camps share. I have argued the structural version of this at more length, if useful to you or your readers: https://amardashehu.substack.com/p/beyond-the-agi-spectacle
Grateful as always for how deliberately and carefully you think out loud in public.
Melanie,
Well done. Such a great article that touches on so many important topics.
I'd like to focus on one, AI using the presence of a ruler to diagnose cancerous lesions. On the surface, this simply shows that AI lacks robust reasoning or understanding.
But I believe it reveals something deeper that matters more, and that is that the system has no stake in distinguishing a meaningful causal relationship from a meaningless correlation.
Simply put, the system can be wrong while the physician must live with being wrong.
What your essay may not explicitly state, but is directionally present, is that unembodied intelligence, deployed where embodied judgment matters, deserves greater attention.
The question worthy of further study is this, is the ability to distinguish meaningful causes from meaningless correlations ultimately a cognitive achievement, or is it partially a consequence of being a stake-bearing participant in reality?
Finally, regarding the terms artificial intelligence (AI) and complex information processing (CIP), my preference is for the less anthropomorphic CIP. But until the question above is more thoroughly explored, the distinction may remain more rhetorical than substantive.
A terrific article. Thank you for providing it.
Great read, fascinating clarity of light and shadow and the difference between LLMs and worldmodels seems to vanish by increasing cases we found solved by LLMs "good enough" , while knowing, that it's not. The space of recognition is the indicator: Acceptance is a decision, full intelligence is a different goal.
Thank you for your perspective.
A great read overall! One small thing that feels nitpicky, but the line "In 2024, in recognition of this astounding progress, AI researchers were awarded Nobel Prizes in both Physics and Chemistry" stood out to me. Yes, those researchers were associated with AI. But the Physics Nobel was in recognition of foundational work with neural networks happening well prior to the award, and the one in Chemistry went for deep learning and neural networks.
Meanwhile, the boom in AI right now is coming in chatbots, agents, LLMs and generative AI, which are not entirely dissimilar but don't have as much overlap as folks will claim. Embedded between two paragraphs about LLMs, a reader might get the mistaken impression that LLMs were responsible for those Nobel Prizes when that isn't true. It's not a deliberate falsehood or intentionally misleading, but feels conceptually imprecise.
I think the success of LLMs is largely the cause of these researchers getting these prizes (maybe not DeepMind, but likely Hinton and Hopfield) even if their work pre-dated LLMs.
Nice work, Melanie. I'll post to LI and share here in notes.
I agree that whether we think of AI as "agents" or "tools" has implications for how we regulate them.
But in many ways current laws are more stringent for harm caused by "agents" than for harm caused by "tools". If you use a tool and it behaves unexpectedly and causes harm, you're usually not liable for that harm. But if you delegate certain tasks to an *agent* (e.g. an employee) who behaves unexpectedly and causes harm, you *are* often liable for that agent's actions.
So thinking of AI as agents rather than tools for regulatory purposes is probably going to be better for promoting careful and responsible use of AI.