by Haytham ElFadeel - hfadeelm@gmail.com
2018
Abstract
A recurring failure mode in AI research is task isolation: we optimize narrowly scoped problems (either: language, vision, planning, or control) as if these faculties were separable in the real world. This produces systems that are statistically competent on benchmarks yet brittle under distribution shift, weak at commonsense inference, and poor at connecting perception to action.
A more scientifically grounded path toward generalist AI is to treat intelligence as an integrated perception–action–language loop: learning representations that preserve the structure of the environment, support counterfactual/causal reasoning, and ground symbols in sensorimotor experience.
Discarding world structure before learning begins
Many ML pipelines begin by factorizing reality into academically convenient subproblems, then training models inside those boundaries:
- Vision → “recognize objects from pixels”
- NLP → “model token sequences”
- Planning/control → “search in abstract state spaces”
- Knowledge → “retrieve text facts”
This decomposition is productive for engineering (it yields tractable benchmarks and modular systems), but it is scientifically risky if the long-term goal is robust, general intelligence. In natural cognition, semantics, perception, and action are mutually constraining: language refers to objects/events in the world; perception is shaped by physics and affordances; and planning depends on causal dynamics and embodied constraints. Treating these as independent encourages systems that learn surface correlations rather than useful structure.
Intelligence as inference + control under partial observability
A technical way to express “holism” is to model an agent interacting with an environment, e.g., as a (PO)MDP. The research question becomes:
Can we learn a representation that is (i) predictive of future observations under interventions, (ii) suitable for planning/control, and (iii) supports grounded semantics for communication?
This framing makes the cost of siloing explicit:
if you learn language without grounding, or vision without dynamics, you learn a representation that is not sufficient for planning and causal inference.
Language without grounding
Language is not independent from the rest of the world. Language arises from the species’ need to communicate about the world around them, languages got more expressive and complex once there was a need for that. Language is a way to represent thought and information about the world, it’s highly dependent on the structure and features of the world it aims to describe. An alien species that isn’t capable of perceiving pain will not have a word for pain in their language.
A well-known critique of purely symbolic or purely distributional treatments of language is the symbol grounding problem: manipulating tokens by formal rules does not, by itself, produce intrinsic meaning—symbols can remain “about” other symbols without any connection to the world. In practical ML terms:
- Next-token prediction (or sequence modeling) can yield powerful statistical regularities.
- But reference (linking “pain,” “weight,” “fragile,” “inside,” “cause,” “intention” to perceivable/controllable reality) is underdetermined from text alone.
- As a result, models can be eloquent yet fail on tasks requiring grounded constraints (e.g., physical plausibility, affordance reasoning, or causal explanations), especially when phrasing shifts or when the task requires interacting with an environment rather than describing it.
Grounding does not mean “attach an image to every word.” It means a learning process where linguistic representations are constrained by multimodal perception, embodiment, and intervention—so that words and phrases become compressed handles on predictive, actionable world models.
Vision without physics
Classic computer vision benchmarks incentivized mapping pixels labels. But real-world understanding requires moving beyond recognition to concepts like:
- Affordances: what actions the environment enables for an agent (e.g., “sit-on-able,” “graspable,” “supporting,” “pourable”).
- Scale and material inference: a “car” in an image could be a toy, a real vehicle, or a rendering; the pixels alone often underspecify mass, compliance, and dynamics.
- Counterfactual reasoning: what would happen if the agent pushed, lifted, or collided with an object?
If a vision system is trained primarily as a static classifier, it can be excellent at naming objects yet weak at predicting how those objects behave under actions. The gap shows up as failures in “obvious” human judgments that rely on embodied priors: stability, support relations, occlusion, friction, reachability, and so on.
The technical takeaway is that “vision” should often be treated as state estimation for control, not just categorization.
Commonsense is (mostly) interaction data
Commonsense knowledge is rarely written down explicitly; it is largely learned through:
- Observation of other agents.
- Causal interventions (“try and see”).
- Abstraction across contexts.
This is why commonsense is hard to bolt on as a static database. It is more naturally captured as predictive structure in a learned world model: given context , predict likely futures and rule out impossible ones.
Recent work on world models formalizes this idea: learn compact latent dynamics that support prediction, imagination, and planning—often via self-supervised objectives and latent-state transition models.
More recent surveys emphasize that actionable world models for embodied agents must capture controllable dynamics (not merely generate plausible images).
Why academic divisions persist (and why they might be costly)
Silos exist for good reasons:
- Specialization accelerates local progress.
- Benchmarks offer measurable iteration loops.
- Modularity helps productization.
But the costs compound when we mistake benchmark competence for general competence:
- Representation mismatch: features optimized for static datasets can be insufficient statistics for downstream control and reasoning.
- Spurious shortcuts: models learn correlations that hold in the dataset but not under intervention.
- Evaluation blind spots: tasks rarely require joint success across modalities plus long-horizon planning.
In short: we get systems that are highly capable within the scaffold of the benchmark, yet fragile when required to integrate perception, language, memory, and action.
Selected References
- Stevan Harnad, “The Symbol Grounding Problem” (1990).
- James J. Gibson, The Ecological Approach to Visual Perception (1979).
- David Ha & Jürgen Schmidhuber, “World Models” (2018).