Towards Data-and Knowledge-Driven Artificial Intelligence: A Survey on Neuro-Symbolic Computing

Wenguan Wang, Yi Yang, Fei Wu

Introduction

Current advances in Artificial Intelligence (AI), especially large AI models, have caused significant changes in numerous research fields, and had profound impacts on every nook and cranny of societal and industrial sectors. At the same time, there is also growing concern in the public and scientific communities regarding the trustworthiness, safety, interpretability, and accountability of the modern AI techniques . This leads to a natural question: What could be the key enabler for the next generation of AI?

AI has historically been dominated by two paradigms: symbolism and connectionism. Symbolism conjectures that symbols representing things in the world are the fundamental units of human intelligence, and that the cognitive pro- cess ⁣{}_{\!} can ⁣{}_{\!} be ⁣{}_{\!} accomplished ⁣{}_{\!} by ⁣{}_{\!} the ⁣{}_{\!} manipulation ⁣{}_{\!} of ⁣{}_{\!} the ⁣{}_{\!} sym- bols, ⁣{}_{\!} through ⁣{}_{\!} a ⁣{}_{\!} series ⁣{}_{\!} of ⁣{}_{\!} rules ⁣{}_{\!} and ⁣{}_{\!} logic ⁣{}_{\!} operations ⁣{}_{\!} upon ⁣{}_{\!} the symbolic representations . Many ⁣{}_{\!} early ⁣{}_{\!} AI ⁣{}_{\!} systems, from ⁣{}_{\!} the ⁣{}_{\!} mid-1950s ⁣{}_{\!} to ⁣{}_{\!} the ⁣{}_{\!} late ⁣{}_{\!} 1980s, ⁣{}_{\!} were ⁣{}_{\!} built ⁣{}_{\!} upon ⁣{}_{\!} sym- bolistic ⁣{}_{\!} models. ⁣{}_{\!} Symbolic ⁣{}_{\!} methods ⁣{}_{\!} have ⁣{}_{\!} several ⁣{}_{\!} virtues: ⁣{}_{\!} they require only a few input samples, use powerful declarative languages for knowledge representation, and have conceptually clear internal functionality. It soon became apparent, however, that such a rule-based, top-down strategy demands substantial hand-tuning and lacks true learning. Moreover, as discrete symbolic representations and hand-crafted rules are intolerant of ambiguous and noisy data, symbolic methods typically fall short when solving real-world problems.

Connectionism, represented by its most successful technique, deep neural networks (DNNs) , serves as the architecture behind the majority of recent successful AI systems. Inspired by the physiology of the nervous system, connec- tionism ⁣{}_{\!} explains cognition by interconnected networks of simple and often uniform units. Learning happens as weight modification, in a data-driven manner; the network weights are adjusted in the direction that minimises the cumulative error from all the training samples, using techniques such as gradient back-propagation . Connectionist models are fault-tolerant, since they learn sub-symbolics, i.e., continuous embedding vectors, and compare these vectorized represen- tations instead of the literal meaning between entities and relations by discrete symbolic representations. Moreover, by learning statistical patterns from data, connectionist models ⁣{}_{\!} enjoy the advantages of inductive learning and generaliza- tion capabilities. Yet, like every coin has two sides, such ap- proaches ⁣{}_{\!} also ⁣{}_{\!} suffer ⁣{}_{\!} from ⁣{}_{\!} several ⁣{}_{\!} fundamental ⁣{}_{\!} problems ⁣{}_{\!} . ⁣{}_{\!} First, ⁣{}_{\!} connectionist ⁣{}_{\!} models ⁣{}_{\!} fall ⁣{}_{\!} significantly ⁣{}_{\!} short ⁣{}_{\!} of ⁣{}_{\!} compositional generalization, the robust ability of human cognition to correctly solve any problem that is composed of fami- liar ⁣{}_{\!} parts ⁣{}_{\!} . ⁣{}_{\!} Second, ⁣{}_{\!} such ⁣{}_{\!} bottom-up ⁣{}_{\!} approaches ⁣{}_{\!} are ⁣{}_{\!} known to ⁣{}_{\!} be ⁣{}_{\!} data ⁣{}_{\!} inefficient. ⁣{}_{\!} Third, ⁣{}_{\!} connectionist models are logically opaque, lacking comprehensibility. It is almost impossible to understand why decisions are made. In the absence of any kind of identifiable or verifiable train of logic, people are left with systems that make potentially catastrophic decisions that are difficult to understand, arduous to correct, and therefore hard to trust. These shortcomings hinder the adoption of connectionist systems in decision-critical applications and reasoning-heavy tasks, such as medical diagnosis, autonomous driving, and mathematical reasoning, and lead to the increasing concern about contemporary AI techniques.

Against this background, neural-symbolic computing (NeSy) ⁣{}_{\!} , as a hybrid of symbolism and connectionism, is widely recognized as an enabler of the next generation of AI . NeSy essentially looks for the integration of two fundamental cognitive abilities : learning (the abili- ty ⁣{}_{\!} to ⁣{}_{\!} learn ⁣{}_{\!} from ⁣{}_{\!} experience), ⁣{}_{\!} and ⁣{}_{\!} reasoning ⁣{}_{\!} (the ⁣{}_{\!} ability ⁣{}_{\!} to ⁣{}_{\!} rea- son from what has been learned), so as to exploit the major strengths ⁣{}_{\!} and ⁣{}_{\!} circumvent ⁣{}_{\!} the ⁣{}_{\!} inherent ⁣{}_{\!} deficiencies ⁣{}_{\!} of ⁣{}_{\!} the ⁣{}_{\!} two paradigms. However, building such an integrated machinery is challenging – one has to conciliate the methodologies of distinct areas ⁣{}_{\!} , for example, statistical inductive learn- ing ⁣{}_{\!} based ⁣{}_{\!} on ⁣{}_{\!} distributed ⁣{}_{\!} representations ⁣{}_{\!} vs ⁣{}_{\!} logical ⁣{}_{\!} deductive reasoning based on localist representations. Though challenging, NeSy has attracted soaring research attention in the recent past, and has demonstrated its superiority in many application scenarios, including visual relationship unders- tanding ⁣{}_{\!} , visual question answering ⁣{}_{\!} , visual scene parsing ⁣{}_{\!} , and commonsense reasoning ⁣{}_{\!} .

In order to facilitate readers to catch up on the rapidly-developing evolution of this field, this paper offers a systematical and timely collection of recent important literature on NeSy, with a focus on the past five years. The surveyed papers are those works published in the flagship repositories for machine learning and related areas, such as computer vision, natural language processing (NLP), and knowledge graph, or have been widely cited. This survey is expected to offer an exhaustive and up-to-date literature overview to re- searchers of interest, and nourish the exploration of open and ⁣{}_{\!} developmental ⁣{}_{\!} issues. ⁣{}_{\!} We ⁣{}_{\!} also ⁣{}_{\!} remark ⁣{}_{\!} that ⁣{}_{\!} this ⁣{}_{\!} survey ⁣{}_{\!} is inevitably a biased view, since there is a broad spectrum of research in this fast-growing area, but we do attempt to identify and analyze common and critical properties of land- mark ⁣{}_{\!} practices ⁣{}_{\!} in ⁣{}_{\!} order ⁣{}_{\!} to ⁣{}_{\!} cover ⁣{}_{\!} major ⁣{}_{\!} research ⁣{}_{\!} threads. ⁣{}_{\!} Rea- ders ⁣{}_{\!} are ⁣{}_{\!} also ⁣{}_{\!} encouraged ⁣{}_{\!} to ⁣{}_{\!} refer to discussions ⁣{}_{\!} in ⁣{}_{\!} , ⁣{}_{\!} among ⁣{}_{\!} others, ⁣{}_{\!} to ⁣{}_{\!} gain ⁣{}_{\!} a ⁣{}_{\!} sense ⁣{}_{\!} of ⁣{}_{\!} the ⁣{}_{\!} breadth ⁣{}_{\!} of ⁣{}_{\!} this ⁣{}_{\!} area. ⁣{}_{\!}

A summary of the structure of this article can be found in Fig. ⁣{}_{\!} 1, which is presented as follows: Sec. ⁣{}_{\!} 2 gives a brief review of early research results of NeSy, which shape the latest effort in this area. Sec. ⁣{}_{\!} 3 introduces the general concepts of mind in psychology and cognitive science, which underpin the theoretical foundations of NeSy, and discusses the recent debate on the necessary and sufficient building blocks of AI, which promotes the advance of this area. Sec. ⁣{}_{\!} 4 presents our taxonomy of NeSy, which classifies recent important NeSy literature according to four dimensions: neural-symbolic interrelation, knowledge representation, knowledge embedding, and functionality. Sec. ⁣{}_{\!} 5 elaborates on popular and emerging application areas of NeSy. Sec. ⁣{}_{\!} 6 conducts performance evaluation and analysis. Finally, Sec. ⁣{}_{\!} 7 and 8 suggest potential valuable directions for further research and conclude the survey. We hope that this survey will help newcomers and practitioners to navigate in this massive field that gained significant momentum in the past few years, as well as provide AI community with background information for generating future research.

History

This section offers a historical perspective of NeSy, prior to its recent acceleration in activity. NeSy aims to provide a unifying view for symbolism and connectionism, advance the modelling of cognition and further behaviour, and build preferable computational methodologies for integrated machine learning and logical reasoning . NeSy has a long-standing tradition that can be traced back to McCulloch and Pitts in 1943 , even before AI was recognized as a new scientific field. For readers who are eager to obtain a more particular overview of the primitive works, we recommend consulting previous review articles, such as .

Although in the seminal work McCulloch and Pitts established strong connection between finite automata (boolean logic) and artificial neural networks, by interpreting simple logical connectives such as conjunction, disjunction ⁣{}_{\!} and ⁣{}_{\!} negation ⁣{}_{\!} as ⁣{}_{\!} binary ⁣{}_{\!} threshold ⁣{}_{\!} units ⁣{}_{\!} in neural networks , NeSy only began to be a formalized field of study since the 1990s and gained systematic research in the early 2000s . For instance, Towell et al. compiled hand-coded symbolic rules into a neural network, and the approximately correct knowledge can be further corrected by empirical learning. Based on some landmark efforts , researchers developed various neural systems for logical inference ⁣{}_{\!}  ⁣{}_{\!} and ⁣{}_{\!} knowledge ⁣{}_{\!} representation ⁣{}_{\!} . As their neural architectures are mainly meticulously designed for hard logic reasoning, they lacked the ability to learn representations and to reason over large-scale, heterogeneous, and noisy data . Nevertheless, these early NeSy systems laid the groundwork for today’s research.

During the 2010s, NeSy received relatively less attention, as DNN-based connectionist techniques achieved remarkable success across a variety of AI tasks. However, as the shortcomings of DNNs became evident, NeSy has recently ushered in its renaissance in the research community.

Background and Context

This section elucidates the two main driving forces behind the field of NeSy: The first one is the theoretical aspiration to understand and model human cognition (Sec. ⁣{}_{\!} 3.1), while the second one is the practical value of combining connectionism and symbolism paradigms in AI application scenario (Sec. ⁣{}_{\!} 3.2). Sec. ⁣{}_{\!} 3.3 further summarizes the recent AI debate among influential thinkers, which motivates a broad range of AI researchers to recognize the significance of NeSy.

∙\bullet Symbols vs Neurons. What is the essence of human cognition? Many researchers agree that symbolic facility is what distinguishes humans from other animals. The prosperity of human sociology and technology is closely concerned with the co-evolution of human brain with symbolic thinking, making us the “symbolic species” . Many cognitive scientists hold the view that human thinking relies on symbol manipulation. From this perspective, human mind is undisputedly symbolic. Therefore, symbolism was conceived in the attempts to structurally code knowledge and logic reasoning into machines. However, human cognition has a physical basis in the brain, which is composed of numerous mostly homogenous neurons. The neurons, together with the connections, or synapses, as well as diverse firing patterns among them, support different cognitive processes, such as attention, problem-solving, memory, learning, decision-making, language, perception, imagination, and logic reasoning. So it seems reasonable to assume if we can simulate the anatomy and physiology of the nervous system with artificial neurons, intelligence will be developed in computers. This belief leads to the emergence of connectionism.

Spontaneously, in order to advance the understanding of the human mind, it appears to be reasonable to seek ways of integration of symbolic and connectionist approaches, instead of focusing on the dichotomy. In this context, artificial neural networks can be regarded as an abstraction of the physical workings of the brain, while the symbolic logic can be viewed as an abstraction of what we introspect, when we engage in explicit cognitive reasoning . Therefore, it is of necessity to ask how these two abstractions can be related or even unified, or how symbol manipulation can emerge from a neural substrate .

∙\bullet Deduction vs Induction. Deductive reasoning and inductive learning arguably constitute two indispensable building blocks of human thinking, helping human to develop knowledge of the world (even though there are yet other building blocks, such as abductive reasoning) . However, their tension might be the most fundamental issue in areas such as philosophy, cognition, and, of course, AI ⁣{}_{\!} . The ⁣{}_{\!} deduction ⁣{}_{\!} camp ⁣{}_{\!}  ⁣{}_{\!} is ⁣{}_{\!} aware ⁣{}_{\!} of ⁣{}_{\!} the ⁣{}_{\!} expressiveness ⁣{}_{\!} of formal languages for representing knowledge about the world, along with proof systems for reasoning from such knowledge bases. The learning camp ⁣{}_{\!} attempts to generalize from examples about partial descriptions of the world . Historically, the dichotomy between the two camps roughly divided the development of AI. Symbolic techniques clearly stand on the side of deductive reasoning; symbolic logic emphasizes high-level reasoning, and sticks to structure the world in terms of objects, attributes, and relations .

By contrast, neural networks are in the statistical learning camp; they learn statistical patterns, i.e., distributed representations of entities, from data. Nevertheless, humans make extensive use of both deduction and induction in everyday life as well as scientific investigation. We cannot precisely determine which part of human cognition is essentially symbolic, and which part is essentially statistical. Consequently, it is imperative to rethink the relationships between deductive reasoning and inductive learning, necessitating robust computational models that are able to coordinate the symbolic essence of reasoning with the statistical nature of learning.

∙\bullet Compositionality vs Continuity. Smolensky et al. proposed to simultaneously exploit two scientific principles, which can explain the way the human brain works, for machine intelligence, from the viewpoint of the underlying computation mode of human cognition. Neurophysiological measurements suggest that information is encoded in the brain through the numerical activation levels of massive neurons, and is processed by spreading this activation through myriad synapses of varying strengths and permanence . Hence it seems evident that human cognition deploys neural computing , which conforms to the Continuity Principle: “the encoding and processing of information are formalized with real numbers that vary continuously” . However, modern scientific studies in philosophy and cognition suggested that all aspects of human intelligence, from language and perception to reasoning and planning, rely on a different type of computing: compositional-structure processing . This type of computing follows the Compositionality Principle : complex information is encoded in large structures which are systematically composed from smaller structures that encode simpler information. Compositionality is widely acknowledged as a core of human intelligence . Our knowledge representation is naturally compositional. For example, we understand the world as a sum of its parts: objects can be broken down into pieces, events are a sequence of actions, and sentences are a series of words. Human cognition exhibits strong compositional generalization – the ability of reorganizing familiar knowledge components in novel ways to solve new problems, so as to handle the potentially infinite number of states of the world . Historically, compositional-structure processing is formalized in the form of discrete symbolic computing, like using words to make sentences. Thus, to some extent, the nature of computation in our brains is both neural and compositional-structure. How can this be? Smolensky called this the Central Paradox of Cognition . Resolving this paradox inevitably calls for a new computing mechanism, that addresses both the Continuity and Compositionality Principles simultaneouslyNote that the solution – neurocompositional computing – proposed by Smolensky et al. is slightly different from NeSy. NeSy broadly refers to any possible hybrid systems that couple, loosely or tightly, neural and symbolic approaches. Neurocompositional computing, instead, is to directly realize compositional-structure processing through continuous neural computing, which can be viewed as a compact, neural network based NeSy system. However, in spite of such difference, both NeSy and neurocompositional computing share the same motivation..

∙\bullet System 1 vs System 2. Kahneman’s ‘fast and slow thinking’ theory, which explains the machinery of human thought, also motivated recent research interest in NeSy . In , Kahneman argued that humans’ decisions are supported by the cooperation of two different kinds of capabilities, called system 1 (‘fast thinking’) and system 2 (‘slow thinking’). Specifically, system 1 thinking is a near-instantaneous and experience-driven process for intuitive, imprecise, quick, and largely unconscious decisions, accounting for 98% of thinking. System 1 thinking, for example, can be in the form of knowing how to zip your jacket without a second thought. Differently, system 2 thinking is slower, deliberative, and conscious, often associated with the subjective experience of agency, choice, and concentration; it provides a powerful tool for solving more complicated problems, where logical, sequential, algorithmic thinking is needed. For example, system 2 thinking is used when working on math problems. It is also worth mentioning that compositional generalization is exhibited in both system 1 thinking and system 2 thinking . Interestingly, system 2 can be viewed as a “slave” of system 1: when system 1 runs into difficulty, it is system 1 that decides to initiate system 2. Even during the execution of system 2, system 1 is ultimately in charge . In addition, solutions discovered by system 2 can be readily available for later use by system 1. Thus, after a while, some problems, initially solvable only by resorting to system 2, can become manageable by system 1 . The consistent and effective use of system 2 can calibrate system 1, which, in turn, promotes system 2, leading to a feedback loop. As the characteristics of system 1 and system 2 are strikingly similar to those of the connectionist approach and the symbolic approach to AI, more and more AI researchers began to rethink the relation between the two traditions and recognize the value of NeSy.

2 NeSy: Best-of-Both Worlds

Rather than taking the motivation from the objective of achieving rational understanding and modeling of human cognition, the study of NeSy is also driven by a more technically motivated perspective – combining numerical connectionist and symbolic logic approaches in order to construct more powerful reasoning and learning machines for computer science applications. The second motivation is based on the observation that connectionist techniques, especially modern DNNs, and symbolic approaches complement each other with respect to their strengths and weaknesses. In particular, connectionist techniques are good at discovering statistic patterns from raw data and are robust against noisy data. Hence they are effective in intuitive judgements, such as image classification. On the other hand, connectionist techniques are data hungry, and black boxes – it is especially challenging to understand their decision-making processes. Alternatively, symbolic approaches are excellent at principled judgements, such as logical reasoning; they exhibit inherently high explainability and provide the ease of using powerful declarative languages for knowledge representation. Nevertheless, symbolic approaches are far less trainable and susceptible to out-of-domain brittleness. As a result, the integration of neural and symbolic approaches seems to be a natural step toward more powerful, trustworthy, and robust AI.

3 Current Debate on AI

Recent years have witnessed remarkable breakthroughs in AI, brought by connectionist approaches and deep learning in particular. But researchers are also coming to realize that contemporary AI systems suffer from serious deficiencies in terms of, for example, data efficiency, comprehensibility, and compositional generalization . This led to influential debates between famous researchers, which are about the underlying principles of AI. As a result, NeSy research gained renewed importance.

Specifically, the 2019 Montreal AI Debate between Yoshua Bengio and Gary Marcus , and the AAAI-2020 fireside conversation with Economics Nobel Laureate Daniel Kahneman and the 2018 Turing Award winners and deep learning pioneers Geoff Hinton, Yoshua Bengio, and Yann LeCun, brought new perspectives and concerns on the future of AI. In the debate between Yoshua Bengio and Gary Marcus, Marcus emphasizes the importance of hybrid systems: “⋯\cdots in order to get to robust artificial intelligence, we need to develop a framework for building systems that can routinely acquire, represent, and manipulate abstract knowledge, with a focus on building systems that use that knowledge in the service of building, updating, and reasoning over complex, internal models of the external world.” Though Hinton agreed that “we need those higher-level concepts to be grounded and have a distributed representation to achieve generalization”, he also addressed that “(numerical connectionist approaches) can get many of the attributes of symbols without the kind of explicit representations of them which has been the hallmark of classical AI” and that “The reason why connectionists really wanted to depart from symbolic processing is because they thought that is wasn’t a sufficiently rich kind of representation.” At AAAI-2020, Kahneman highlighted the importance of symbol manipulation in system 2: “⋯\cdots as far as I’m concerned, system 1 certainly knows language ⋯\cdots system 2 does involve certain manipulation of symbols.” Although there are disagreements about, for example, how to represent symbols in DNNs and how to achieve the hybrid of connectionism and symbolism, the thinkers, in broad strokes, are in agreement that new-generation AI systems ought to be able to handle high-level abstract concepts and to conduct sound reasoning.

NeSy: Taxonomy and State of the Art

As already alluded to in the introduction, this section is devoted to a structured and comprehensive review of state-of-the-art ⁣{}_{\!} NeSy ⁣{}_{\!} algorithms. ⁣{}_{\!} Sec. ⁣{}_{\!} 4.1 ⁣{}_{\!} details ⁣{}_{\!} our ⁣{}_{\!} taxonomy ⁣{}_{\!} for NeSy, based on which we survey recent major research results in this area from four perspectives: neural-symbolic integration (Sec. ⁣{}_{\!} 4.2), knowledge representation (Sec. ⁣{}_{\!} 4.3), knowledge embedding (Sec. ⁣{}_{\!} 4.4), and functionality (Sec. ⁣{}_{\!} 4.5).

Our overall taxonomy for NeSy AI is mainly built upon the classification scheme proposed by Sebastian and Hitzler in 2005 , but modified according to our specific focus and recent development tend in this field. Basically, our scheme has four main dimensions, namely neural-symbolic integration, knowledge representation, knowledge embedding, and functionality. Each dimension contains elements representing the notable properties of NeSy approaches.

The first dimension – neural-symbolic integration – categorizes NeSy systems according to the combination mode – how the symbolic and neural parts are integrated as a hybrid. Along this dimension, we further adopt the classification schema recently introduced by Henry Kautz at AAAI-2020 , which is influential and insightful. More details of this dimension will be given in Sec. 4.2.

For the second dimension – knowledge representation, we focus on the symbolic aspect of the NeSy AI system. Depending on how the knowledge is represented, i.e., symbolic vs logic, we can distinguish the systems, as discussed in Sec. 4.3.

The third dimension – knowledge embedding – considers at which component of the neural machine the symbolic knowledge is integrated into. We find that the integration can be made at every key part of the connectionist pipeline, namely data preprocessing, network training, network architecture, as well as final inference. Based on this insight, in Sec. 4.4, we make the categorization along this dimension.

The forth dimension refers to the functionality of the NeSy system, namely whether it focuses more on machine learning or automated symbolic reasoning. More detailed discussions can be found in Sec. 4.5.

Note that the four dimensions are proposed to comprehensively describe the key characteristics of a NeSy system; they are independent and non-exclusive. Along with these four dimensions of our taxonomy, we summarize the key features of recent remarkable works in this field in Table I and give detailed review below.

2 Neural-Symbolic Integration

With a good understanding of the reasons behind the need for integrating symbolic and connectionist approaches, it should turn next to the integration mode. Following the rationale laid out in , we distinguish six types of NeSy AI systems:

∙\bullet Type 1. Symbolic Neuro Symbolic (Fig. 3): This is, in Kautz’s words, the current standard operating procedure of deep learning methods in some application tasks where the input and output are symbols. For example, most current NLP systems, including large language models like GPT-3 , fall under this category (see Table I); the input symbols are converted to vector embeddings by word2vec , GloVe , etc., and then processed by the neural models, whose output embeddings are further converted to the required symbolic category or sequence of symbols via a softmax operation. This type, though some may argue is a stretch to refer to as NeSy, is included by Kautz to emphasize that the input and output of a neural network can be made of symbols , e.g., in the case of language translation, or graph classification.

∙\bullet ⁣{}_{\!} Type 2. Symbolic ⁣ ⁣{}_{\!\!} [Neuro] (Fig. 4): Type 2 refers to hybrid but overall symbolic systems, where neural modules are internally used as subroutine within a symbolic problem solver. In a nutshell, the symbolic and neural parts are only loosely-coupled (see Table ⁣{}_{\!} II). Kautz includes DeepMind’s AlphaGo ⁣{}_{\!}  ⁣{}_{\!} as ⁣{}_{\!} an ⁣{}_{\!} example: ⁣{}_{\!} the ⁣{}_{\!} problem ⁣{}_{\!} solver ⁣{}_{\!} is ⁣{}_{\!} the ⁣{}_{\!} Monte Carlo ⁣{}_{\!} Tree ⁣{}_{\!} Search ⁣{}_{\!} algorithm ⁣{}_{\!} and ⁣{}_{\!} its ⁣{}_{\!} heuristic ⁣{}_{\!} evaluation ⁣{}_{\!} func- tion is a neural network. In ⁣{}_{\!} , a NeSy model is designed to learn and generalize compositional rules. Based on a sequence-to-sequence (seq2seq) generation network, a symbolic stack machine is adopted for supporting recursion and sequence manipulation, and the execution trace is produced by a neural network. In ⁣{}_{\!} , a rule-based system, which uses abstract concepts captured by a neural perception module as I/O specifications, is introduced for program synthesis from raw visual observations. Another example is recent large language model (LLM) based AI agents, e.g., VisProg ⁣{}_{\!} , HuggingGPT ⁣{}_{\!} , and ViperGPT ⁣{}_{\!} , which leverage LLMs to decompose a complex task into a sequence of sub-tasks that can be solved using off-the-shelf AI models.

∙\bullet Type 3. Neuro∣|Symbolic (Fig. 5): Type 3 is also a hybrid system where the neural and symbolic parts focus on different but complementary tasks in a big pipeline. Kautz mentioned Type 3 and Type 2 “differ in that the neuro part is a coroutine rather than a subroutine.” To eliminate ambiguity, here we further restrict Type 3 as the hybrid systems where the interaction between neural and symbolic parts can boost both the individual and collective performance. Therefore, the relation between neural and symbolic parts in Type 3 systems is collaboration, rather than only functional dependency in Type 2 (see Table II). For instance, presents an abductive learning framework, which conducts sub-symbolic perception learning and symbolic logic reasoning separately but interactively. Many deep

learning based program synthesis algorithms that leverage deep learning techniques to generate symbolic programs/rule systems satisfying high-level task specifications also fall in this category. Another notable case is , where a neural perception module learns visual concepts and a symbolic reasoning module executes symbolic programs on the concept representations for question answering. The symbolic reasoning module provides feedback signals that support gradient-based optimization of the neural perception module. Recent efforts in neural-symbolic reinforcement learning (RL) also belong to Type 3. For example, in , symbolic planning are integrated into reinforcement learning (RL) for robust decision-making. Symbolic plans are used to guide task execution, and the task experiences are fed back for improving symbolic planning. Some other examples include Neural Theorem Provers (NTPs) , Conditional Theorem Provers (CTPs) , NLProlog , DeepProbLog , NeuroLog , DiffLog . Among them, a notable case is DeepProbLog , which adopts neural networks as predicates to compute the probabilities of probabilistic facts, and hence uses the inference mechanism of ProbLog, a probabilistic logic programming language based on first-order logic, to compute the gradient of the desired loss.

∙\bullet Type 4. Neuro:​ Symbolic​ →\rightarrow​ Neuro (Fig. ⁣{}_{\!} 6): In a TYPE 4 NeSy system, symbolic rules/knowledge are compiled into the architecture or training regime of neural networks (see Table ⁣{}_{\!} III). For instance, there is a recent surge of interest in learning vector based representations of symbolic knowle- dge so as to naturally incorporate symbolic domain knowledge into ⁣{}_{\!} connectionist ⁣{}_{\!} architectures ⁣{}_{\!} . ⁣{}_{\!} A few ⁣{}_{\!} neural-symbolic ⁣{}_{\!} mathematics ⁣{}_{\!} systems ⁣{}_{\!} for ⁣{}_{\!} equation ⁣{}_{\!} solving ⁣{}_{\!} and verification represent mathematical expressions as trees, which are used as training data. A family of (visual) question answering models generate and execute symbolic programs for answering questions, where the programs are implemented as fully differentiable operations and/or neural networks. A huge body of recent algorithms leverage graph neural networks (GNNs) to embed entities and relations in external knowledge bases, so as to boost the performance in various applications tasks in computer vision and NLP. Broadly speaking, these methods fall into this type as suggested by Kautz, while some may argue that the reasoning ability of GNNs is rather weak.

∙\bullet Type 5. Neuro\textscSymbolic{}_{\textsc{Symbolic}} (Fig. 7): This type of NeSy systems turns symbolic knowledge into additional soft-constraints in the loss function used to train DNNs. Thus, the knowledge is compiled into the weights of DNNs (see Table IV). Some other recent efforts in this direction include . Logic Tensor Networks (LTNs) , as a prominent example, translate first-order logic formulae as fuzzy relations on real numbers for neural computing, so as to allow gradient based sub-symbolic learning. The core idea is to relax boolean first-order logic as soft fuzzy logic, which deals with reasoning that is approximate instead of fixed and exact. In fuzzy logic, variables have a truth degree that ranges in : zero and one meaning that the variable is false and true with certainty, respectively . LTNs approximate non-differentiable logic connectives (i.e., ∧,∨,¬,⇒\land,\lor,\neg,\Rightarrow), and quantifiers (i.e., ∃\exists, ∀\forall) with differentiable fuzzy logic operators . In this way, logic rules can be embedded into network learning objective for end-to-end training. A trend of approaches consider class hierarchies when designing classifiers , where the class hierarchies act as both classification targets and background knowledge. They design different training objective functions to encourage the coherence between the prediction and the class hierarchy. For instance, in , compositional relations over semantic hierarchies are cast as extra training targets for hierarchical scene parsing.

∙\bullet Type 6. Neuro[Symbolic] (Fig. 8): Type 6 system, which Kautz ⁣{}_{\!} believes ⁣{}_{\!} “has ⁣{}_{\!} the ⁣{}_{\!} greatest ⁣{}_{\!} potential ⁣{}_{\!} to ⁣{}_{\!} combine ⁣{}_{\!} the ⁣{}_{\!} strengths of logic-based and neural-based AI”, are fully-integrated systems that directly embed a symbolic reasoning engine inside a neural engine. By imitating logical reasoning with tensor calculus, a line of approaches learn the execution of symbo- lic ⁣{}_{\!} operations ⁣{}_{\!} through ⁣{}_{\!} neural ⁣{}_{\!} networks ⁣{}_{\!} , ⁣{}_{\!} which, ⁣{}_{\!} to some extent, can be classified into Type ⁣{}_{\!} 6 (see Table ⁣{}_{\!} IV). Yet, their ⁣{}_{\!} symbolic ⁣{}_{\!} reasoning ⁣{}_{\!} ability ⁣{}_{\!} is ⁣{}_{\!} still ⁣{}_{\!} relatively ⁣{}_{\!} weak. ⁣{}_{\!} Kautz views Type 6 methods as computational models of Kahneman’s system 1 and system 2 and further addresses that Type 6 methods should be capable of combinatorial reasoning. From Kautz’s viewpoint, it seems that there is no NeSy approach to-date that can truly meet the standard of Type 6.

3 Knowledge Representation

After clarifying and categorizing the main ways in which symbolic and deep learning approaches are integrated together in this area, we turn next to symbolic knowledge, based on which symbol manipulation/logical calculus can be carried out. Understanding symbolic knowledge serves as the cornerstone of a NeSy system. Hence a new categorization dimension for NeSy approaches emerges purely from the perspective of how symbolic knowledge is represented. As illustrated in Fig. 9, the representation approaches for symbolic knowledge can be classified into five main groups: knowledge graph, propositional logic, first-order logic, programming language, and symbolic expression. Hence this section is structured according to such these different categories of knowledge representations.

∙\bullet Knowledge Graph: Knowledge graphs, as a popular and effective tool for knowledge representation, contain a large amount of entities and the relationships between them. Knowledge graphs are typically directed labeled graphs, formed by representing entities – e.g., people, places, things – as nodes, and relations between entities – e.g., “is a friend of”, “is located in”, “is a” – as edges. They contain facts that are represented as “SPO” triples: (Subject, Predicate, Object) where Subject and Object are entities and Predicate is the relation between them. Edges are directed from subject to object, and edge labels represent different types of relations. In an unweighted graph, all edges have the same weight. In a weighted graph, each edge is associated with a number representing its weight. The edge weight quantifies the strength and the sign of the corresponding relationship between nodes. Please refer to the first two examples in Fig. 9. A considerable body of works in computer vision and NLP fields build (weak) NeSy systems upon knowledge graphs. Their knowledge graphs are frequently built upon our world knowledge. Since we humans understand the world by components, graphs are naturally used to represent relations between visual entities in the field of computer vision. Many famous computer vision datasets, such as ImageNet , and Cityscapes , are released with structured/hierarchical annotation. For example, the ImageNet labels are organized according to WordNet . A family of visual parsing algorithms are developed for interpreting the part-of (compositional) relations in common visual scenes or human-/object-centric visual stimuli . These methods fall into a broader field of machine learning, called hierarchical classification, which is devoted to class taxonomy aware classification .

∙\bullet Propositional Logic: Logical statements provide a flexible declarative language for formalizing knowledge about facts and dependencies, hence playing an important role for the integration of prior knowledge into connectionist architectures. Propositional logic, also known as boolean logic or sometimes zeroth-order logic, is the simplest form of logic where all the statements are made by propositions. A pro- position is a declarative statement that is either true or false. Propositional logic studies the logical relationships between propositions which are connected via logical connectives. Typically, logical connectives (or operators), including Conjunction (“∧\wedge”), Disjunction (“∨\vee”), Negation (“¬\neg”), and Implication (“⇒\Rightarrow”), are used to create compound propositions or represent a sentence logically. Propositional Logic allows for translating ordinary language statements (i.e., ⁣{}_{\!} IF ⁣{}_{\!}​ AA​ THEN BB) into formal logic rules (A ⁣⇒ ⁣BA\!\Rightarrow\!B). In propositional logic, simple statements – statements that contain no other statement as a part – are treated as indivisible wholes. Hence, propositional logic does not deal with logical relationships and properties that involve smaller parts of statements, such as ⁣{}_{\!} the ⁣{}_{\!} subject ⁣{}_{\!} and ⁣{}_{\!} predicate ⁣{}_{\!} of ⁣{}_{\!} a ⁣{}_{\!} statement. ⁣{}_{\!} Due ⁣{}_{\!} to ⁣{}_{\!} the ⁣{}_{\!} simp- licity ⁣{}_{\!} of ⁣{}_{\!} propositional ⁣{}_{\!} logic, ⁣{}_{\!} many ⁣{}_{\!} early ⁣{}_{\!} NeSy ⁣{}_{\!} systems, ⁣{}_{\!} such as , consider the symbolic knowledge in the form of propositional logic. Recent work in this direction includes . For instance, derives differentiable semantic loss from constraints expressed in propositional logic, for improving the performance in semi-supervised classification. In , GNNs are adopted to embed symbo- lic knowledge, represented as propositional formulae, into DNNs for visual reasoning. For example, given an image which contains a person and a pair of glasses, exploits logic rules to reason about questions like whether the person is ⁣{}_{\!} wearing ⁣{}_{\!} the ⁣{}_{\!} glasses. ⁣{}_{\!} The ⁣{}_{\!} corresponding ⁣{}_{\!} propositional ⁣{}_{\!} logic rule is given as: wear(person, glasses)⇒\Rightarrowin(glasses, person) ∧\wedgeexist(person)∧\wedgeexist(glasses), ⁣{}_{\!} which ⁣{}_{\!} states ⁣{}_{\!} that ⁣{}_{\!} if the per- son is wearing the glasses, the image of the glasses should be “inside” the image of the person . In , domain knowledge regarding monotonic relationships between process variables are incorporated into DNN’s training. Considering a function h(x) ⁣= ⁣yh(x)\!=\!y such that x1 ⁣>x2 ⁣⇒ ⁣h(x1) ⁣> ⁣h(x2)x_{1}\!>x_{2}\!\Rightarrow\!h(x_{1})\!>\!h(x_{2}). Then x1,x2x_{1},x_{2} and h(x1),h(x2)h(x_{1}),h(x_{2}) are said to share a monotonic relationship.

∙\bullet First-Order Logic: Propositional logic is a finitary system that only involves a finite number of propositions and does not require sophisticated symbol manipulation operations, i.e., substitution and unification, which are needed for nested terms. Thus, it is relatively easy to implement propositional logic programs using DNNs . However, the expressive power of propositional logic is rather limited, since it cannot express assertions about elements of a structure. First-order logic, also called quantified logic or predicate logic, is an extension to propositional logic and more powerful. First-order logic can express the relationship between objects by allowing variables in predicates bound by quantifiers. Specifically, first-order logic augments propositional logic with two new linguistic features, viz. variables and quantifiers. Variables are introduced to refer to objects of a certain type (i.e., domain of discourse) and can be substituted by a specific object. The universal quantifier (“∀\forall”) and existential quantifier (“∃\exists”) allow us to quantify over objects (see examples in Fig. 9). A few solutions have emerged to enable DNNs to represent first-order logic. However, most of these solutions can only handle restricted fragments of first-order logic. For example, some NeSy systems turn to Datalog for logic reasoning, or leverage GNNs to reason over local subgraph structure for inductive relation prediction , or regularize distributed representations via domain-specific logic rules . To capture the full expressive power of first-order logic, some approaches use fuzzy logic to translate prior knowledge, expressed as a set of first-order logic clauses, into extra training objectives. In , first-order logic rules are compiled into differentiable operations. Another group of approaches use first-order logic to generate a random field, based on Markov logic networks . Some other approaches adopt Prolog , a logic programming language, for knowledge representation. It shall be noted here that, due to the conflict between the infinitary nature of first-order logic – allowing the use of function symbols as language primitives, and the finiteness of DNNs , it is much harder to model first-order logic in a connectionist setting compared with propositional logic.

∙\bullet Programming Language: Programming language, such as logic language Prolog and action language BC\mathcal{BC} , is a family of formal language used for writing computer pro- grams and communicating with machines. Typically, they consist ⁣{}_{\!} of ⁣{}_{\!} syntax ⁣{}_{\!} and ⁣{}_{\!} semantics, ⁣{}_{\!} where ⁣{}_{\!} syntax ⁣{}_{\!} represents ⁣{}_{\!} rules that define the combinations of symbols and semantics assigns computational meaning to valid strings formulated with respect to the syntax. A set of NeSy methods store knowledge in programs to execute. For ⁣{}_{\!} example, ⁣{}_{\!}  ⁣{}_{\!} formulate ⁣{}_{\!} domain ⁣{}_{\!} knowledge ⁣{}_{\!} in ⁣{}_{\!} ac- tion language BC\mathcal{BC} to perform long-term planning, performs a type-directed search over the library composed of parameterized programs defined in the HOUDINI language for concept reusing in other tasks. Note that for the NeSy systems that adopt Datalog or Prolog – a subset of first-order logic ⁣{}_{\!} – ⁣{}_{\!} for ⁣{}_{\!} knowledge ⁣{}_{\!} representation, ⁣{}_{\!} they ⁣{}_{\!} are ⁣{}_{\!} classified ⁣{}_{\!} into the group of first-order logic, along the dimension of know- ledge representation.

∙\bullet Symbolic Expression: Symbolic expression here roughly refers to other types of knowledge representation other than those mentioned above. Representative examples include mathematical expressions and specific symbolic sequences generated from some informal symbolic systems with self-defined rules. For instance, in , the source symbolic strings can be arbitrary forms of algebraic or logic expressions. In , learning and reasoning are conducted in conjunction with mathematical equations, which are typically translated into a syntax tree according to the grammatical or structural knowledge. Apart from that, decompose the generation of complex program into multiple predefined operators and combine them together afterwards, which improves the accuracy and can be applied to different domains by simply extending the set of symbolic operator. This kind of compositionality is also a notable case of NeSy in symbolic reasoning.

4 Knowledge Embedding

After studying how the knowledge is represented in NeSy, we next focus on the dimension of knowledge embedding, which addresses the question of where in the neural network based connectionist solutions the symbolic knowledge is embedded. Answering this question renders us a more profound understanding of the integration of symbolic knowledge and neural networks in modern NeSy systems. Our literature survey revealed that modern NeSy solutions are able to embed symbolic knowledge into training data, sub-symbolic representation, connectionist architecture, and neural inference, corresponding to the key element of the connectionist pipeline. Note that Type 1, Type 2, and Type 3 NeSy systems are not discussed here, as they either do not take symbolic knowledge into consideration (Type 1) or adopt an independent symbolic model for exploiting knowledge (Type 2 and Type 3). Whereas for Type 4, Type 5, and Type 6 NeSy systems, the knowledge can be simultaneously integrated into different components of the neural pipeline.

∙\bullet Data: A natural strategy to embed knowledge into connectionist approaches is to straightforwardly embed it in the ⁣{}_{\!} structure ⁣{}_{\!} of ⁣{}_{\!} data. ⁣{}_{\!} A ⁣{}_{\!} prominent ⁣{}_{\!} approach ⁣{}_{\!} is ⁣{}_{\!} to ⁣{}_{\!} translate ⁣{}_{\!} symbolic ⁣{}_{\!} expressions, ⁣{}_{\!} such ⁣{}_{\!} as ⁣{}_{\!} computer ⁣{}_{\!} programs ⁣{}_{\!} , molecular ⁣{}_{\!} structures ⁣{}_{\!} , ⁣{}_{\!} mathematical ⁣{}_{\!} expressions , ⁣{}_{\!} logic ⁣{}_{\!} formulae ⁣{}_{\!} , ⁣{}_{\!} into ⁣{}_{\!} a structured (typically tree-/graph-organized) symbolic sequence, with respect to the ⁣{}_{\!} corresponding ⁣{}_{\!} grammars, ⁣{}_{\!} semantics, ⁣{}_{\!} and/or ⁣{}_{\!} the ⁣{}_{\!} relational structure of the knowledge. The advantage of this knowledge embedding strategy is the tremendous relief of burden on the engineering of network architecture and training objective ⁣{}_{\!} – ⁣{}_{\!} off-the-shelf ⁣{}_{\!} seq2seq ⁣{}_{\!} networks ⁣{}_{\!} and ⁣{}_{\!} GNNs ⁣{}_{\!} can be directly applied. A few Type 4 NeSy systems adopt this strategy . However, such simple strategy has its limits for embedding complex knowledge.

∙\bullet Sub-Symbolic ⁣{}_{\!} Representation: ⁣{}_{\!} Type ⁣{}_{\!} 5 ⁣{}_{\!} NeSy ⁣{}_{\!} systems ⁣{}_{\!} embed symbolic knowledge into the distributed representation, by means of training objectives that are specialized to the knowledge. A common way of building such knowledge-specialized training objectives is to make discrete symbolic operations differentiable . Designing appropriate loss functions for distributed encoding of symbolic knowledge is appealing as it does not require architectural change to the connectionist models or extra load of preprocessing the input data. However, it appears to be particular challenging, for example, when encoding highly abstract symbolic knowledge, such as compositional generalization . Moreover, there is no guarantee that distributed knowledge embedding can always lead to valid outputs that are coherent with the symbolic knowledge.

∙\bullet Network Architecture: Another common way to integrate knowledge into DNNs is to design the network architecture to reflect the structure of the knowledge. For instance, GNNs are widely adopted for capturing the complex relations in knowledge graphs and graph-structured symbolic expressions . In and , composi- tional ⁣{}_{\!} relations ⁣{}_{\!} between ⁣{}_{\!} visual ⁣{}_{\!} entities ⁣{}_{\!} are ⁣{}_{\!} explicitly ⁣{}_{\!} encoded into differentiable network layers/models. Although this strategy is adopted by many Type 3 systems and is believed as the key building block for Type 6 systems, it requires significant engineering efforts in neural architecture design.

∙\bullet Neural Inference: Embedding symbolic knowledge into network feedforward inference is also a feasible way, which imposes explicit constraints to force the final hypothesis to agree with the knowledge. For instance, generate both syntactically and semantically correct predictions of molecular structures by parsing a feasible path from a tree-structured knowledge space. packages logical constraints into an iterative process and injected into the DNNs in a form of several matrix multiplications, so as to bind logic reasoning into network feed-forward prediction.

5 Functionality

It is clear that the ultimate goal of NeSy is to implement a powerful ⁣{}_{\!} AI ⁣{}_{\!} system ⁣{}_{\!} with ⁣{}_{\!} combined ⁣{}_{\!} capabilities ⁣{}_{\!} of ⁣{}_{\!} both ⁣{}_{\!} data-driven learning and knowledge-driven reasoning. However, most existing NeSy systems are either good at learning or good at reasoning, but rarely both . To better understand the strengths and weaknesses of the systems, we examine the current NeSy systems from another dimension – core functionality. This dimension reflects whether the systems focus more on statistical learning or on symbolic reasoning.

∙\bullet Learning: Some Type 3 NeSy systems, like those neural-symbolic RL approaches , and the vast majority of Type 4 and Type 5 systems usually exhibit strong learning ability, but are relatively weak at logic reasoning. For example, neural-symbolic RL approaches and visual reasoning algorithms of Type 4 are typically limited to a small set of pre-defined and simple programs/operations, and the sequences of the programs are usually generated through DNNs. For those Type 4 systems based on knowledge embedding or training regime modification and Type 5 systems that integrate logical knowledge as additional constraints in the loss function, they pay more attention to symbolic knowledge embedding, rather than performing logic reasoning. As the symbolic knowledge is only implicitly encoded into the weights of DNNs, they struggle with explicit reasoning and their explainability is also weak.

∙\bullet ⁣{}_{\!} Reasoning: ⁣{}_{\!} Generally ⁣{}_{\!} speaking, ⁣{}_{\!} most ⁣{}_{\!} Type ⁣{}_{\!} 2 ⁣{}_{\!} NeSy ⁣{}_{\!} models  ⁣{}_{\!} and ⁣{}_{\!} a ⁣{}_{\!} few ⁣{}_{\!} Type ⁣{}_{\!} 3 ⁣{}_{\!} systems ⁣{}_{\!} , ⁣{}_{\!} that are built upon statistical relational learning and logic program- ming, ⁣{}_{\!} retain ⁣{}_{\!} the ⁣{}_{\!} main ⁣{}_{\!} focus ⁣{}_{\!} on ⁣{}_{\!} the ⁣{}_{\!} manipulation ⁣{}_{\!} of ⁣{}_{\!} the ⁣{}_{\!} sym- bols and thus yield relatively strong reasoning ability rather than statistical learning. In particular, for the Type 2 NeSy systems , the neural part is only involved as a submodule, while the whole system acts as a symbolic model. For those logical-programming-based Type 3 systems , they allow for (differentiable) logical inference over probabilistic evidence from neural networks; however, their scalability is typically limited.

∙\bullet Reasoning and Learning: Despite the recent progress, it is still hard to achieve a compact NeSy system that has both strong logic reasoning and expressive statistic learning abilities. In the sense of Kautz’s vision, Type 6 symbolic systems could have such combined abilities. However, there are only a ⁣{}_{\!} few ⁣{}_{\!} models ⁣{}_{\!}  ⁣{}_{\!} can ⁣{}_{\!} be ⁣{}_{\!} barely ⁣{}_{\!} recognized ⁣{}_{\!} as ⁣{}_{\!} Type ⁣{}_{\!} 6.

Application Areas and Tasks

With the ambitious goal and recent rapid progress of NeSy research, various novel applications and tasks have emerged across different domains (Fig. 10), such as computer vision, natural language processing, robotics, and even other scientific disciplines. In this section, we showcase some of the prominent application examples that illustrate the potential and impact of NeSy. However, we note that due to the high diversity of the application scenarios and the significant difference among different NeSy systems, it is not feasible to provide a comprehensive and fair evaluation framework for all the NeSy systems under a unified setting.

∙\bullet Scientific ⁣{}_{\!} Discovery: ⁣{}_{\!} Scientific ⁣{}_{\!} discovery ⁣{}_{\!} typically ⁣{}_{\!} requires

algorithms that discover scientific hypotheses or concepts from data, and respect physical constraints and domain-specific knowledge that are known to hold in the world. Moreover, the algorithms should better interpret and explain how they come up with their solutions and convey their insights to human scientists . As a result, scientific discovery poses great challenges for current pure data-driven AI techniques, yet serves as a good testbed for NeSy.

Several ⁣{}_{\!} recent ⁣{}_{\!} studies ⁣{}_{\!} showed ⁣{}_{\!} the ⁣{}_{\!} extraction ⁣{}_{\!} of ⁣{}_{\!} symbolic models from experimental data of mechanical systems ⁣{}_{\!} and ⁣{}_{\!} in ⁣{}_{\!} astronomy ⁣{}_{\!} . ⁣{}_{\!} For ⁣{}_{\!} instance, ⁣{}_{\!} the ⁣{}_{\!} authors ⁣{}_{\!} of ⁣{}_{\!}  ⁣{}_{\!} apply a symbolic regression technique to a GNN model that is trained on cosmological dark matter data, and demonstrate that explicit physical relations can be discovered in the form of analytic formula. For protein structure prediction, applies the gradient descent algorithm to uncover the graph structure of proteins in 3D space, where the edges of the graph are determined by the proximity of residues. Moreover, some recent works employ neuro-symbolic programming to analyze the behavior of laboratory animals, such as classifying sequential animal behaviors, clustering animal behaviors in an interpretable way, and representing expert knowledge in a reusable domain-specific language and more general domain-level labeling functions. In addi- tion, ⁣{}_{\!} NeSy ⁣{}_{\!} techniques ⁣{}_{\!} are ⁣{}_{\!} applied ⁣{}_{\!} to ⁣{}_{\!} retrosynthesis ⁣{}_{\!} and ⁣{}_{\!} reac- tion ⁣{}_{\!} prediction ⁣{}_{\!} in ⁣{}_{\!} organic ⁣{}_{\!} chemistry . ⁣{}_{\!} For ⁣{}_{\!} example, ⁣{}_{\!}  ⁣{}_{\!} combines ⁣{}_{\!} Monte ⁣{}_{\!} Carlo ⁣{}_{\!} tree ⁣{}_{\!} search ⁣{}_{\!} with ⁣{}_{\!} an ⁣{}_{\!} expansion ⁣{}_{\!} policy ⁣{}_{\!} network to discover retrosynthetic routes. adds syntax and semantics checking during molecule synthesis. devises a conditional graph logic network to learn when to apply rules ⁣{}_{\!} from ⁣{}_{\!} reaction ⁣{}_{\!} templates, ⁣{}_{\!} implicitly ⁣{}_{\!} considering ⁣{}_{\!} both ⁣{}_{\!} the chemical and strategic feasibility of the resulting reaction.

∙\bullet Programming Systems: Program synthesis is another important application domain of NeSy. The goal is to automati- cally generate programs from high-level task specifications. The ⁣{}_{\!} specifications ⁣{}_{\!} are ⁣{}_{\!} typically ⁣{}_{\!} hard ⁣{}_{\!} logical ⁣{}_{\!} constraints, ⁣{}_{\!} for example, ⁣{}_{\!} test ⁣{}_{\!} cases ⁣{}_{\!} that need to be satisfied exactly, pre-postcondition pairs, or temporal logic formulas. The pro- grams ⁣{}_{\!} are ⁣{}_{\!} structured, ⁣{}_{\!} symbolic ⁣{}_{\!} expressions ⁣{}_{\!} that ⁣{}_{\!} follow ⁣{}_{\!} the ⁣{}_{\!} syntax of a domain-specific language. NeSy tools are more suitable for this domain than purely neural alternatives, by virtue of their modularity and use of symbolic primitives.

To combine neural learning with the formal, logic constraints of programming languages, NeSy based programmers typically learn input and context-specific heuristics from large-scale data and use such heuristics to guide search-based symbolic methods to guarantee soundness. For example, given some tokens that indicate the desired program functionality, such as API calls, types, or keywords, generates strongly typed Java-like source code in two steps. First, it learns neural models to produce sketches of programs, which ⁣{}_{\!} are ⁣{}_{\!} abstract ⁣{}_{\!} representations ⁣{}_{\!} of program syntax that omit low-level details. Second, it concretizes the sketches into type-safe programs using a combinatorial search procedure. utilizes a combination of DNNs and stochastic search to parse drawings into symbolic specifications; these specifications are then fed into a general-purpose program synthesis engine to infer a structured graphics program. introduces the concept of neurosymbolic attribute grammars, which combine a stochastic context-free grammar with semantic attributes computed by static program analysis. The neural network learns to condition its generation actions on these attributes, which provide useful semantic clues and long-distance dependencies.

∙\bullet Question-Answering: Question-answering (QA) is a long-standing AI task that aims to build intelligent systems that can automatically answer questions from humans in natural ⁣{}_{\!} language, ⁣{}_{\!} typically ⁣{}_{\!} with ⁣{}_{\!} the ⁣{}_{\!} aid ⁣{}_{\!} of ⁣{}_{\!} a ⁣{}_{\!} knowledge ⁣{}_{\!} source ⁣{}_{\!} com- posed of unstructured text corpora and/or structured concepts. Answering complex questions that involve ⁣{}_{\!} multiple ⁣{}_{\!} subjects, ⁣{}_{\!} compound ⁣{}_{\!} relations, ⁣{}_{\!} and ⁣{}_{\!} numerical ⁣{}_{\!} operations is a grand challenge in QA. To address this challenge, NeSy based QA systems have been recently developed, typically following either a semantic parsing paradigm or a knowledge embedding paradigm. Semantic parsing-based methods learn to translate a question into a symbolic logic form by conducting semantic and syntactic analysis and then derive the answer by executing the parsed logic form against ⁣{}_{\!} the ⁣{}_{\!} knowledge ⁣{}_{\!} source. ⁣{}_{\!} For ⁣{}_{\!} example, ⁣{}_{\!} adopts a neural seq2seq model that maps language utterances to programs and utilizes a key-variable memory to save and reuse intermediate ⁣{}_{\!} execution ⁣{}_{\!} results ⁣{}_{\!} for ⁣{}_{\!} supporting ⁣{}_{\!} language ⁣{}_{\!} compositionality and complex semantics. Then a symbolic Lisp interpreter is used to perform program execution over the knowledge source, and helps find good programs by pruning the search space. Knowledge embedding-based methods learn and store neural representations of the knowledge source and then retrieve the answers from the stored neural form of the knowledge source considering the information conveyed in the questions. For example, builds a fact memory that encodes the entities of a knowledge source as numerical vectors and provides a contextualized reference for a neural language model to create answers. Overall, semantic parsing based NeSy QA systems can produce a more interpretable reasoning process by generating expressive logic forms. However, they heavily rely on the design of the logic form and parsing algorithm, which turns out to be the bottleneck of performance improvement. As a comparison, knowledge embedding based NeSy QA systems enjoy more benefits of end-to-end training but lack traceable reasoning.

∙\bullet Vision-Language Analysis and Reasoning: NeSy techni- ques have been also successfully applied to vision-language analysis and reasoning tasks (e.g., visual question-answering (VQA) and visual grounding), which often require comprehensive understanding and reasoning over both the visual and linguistic modalities. VQA is a challenging task that is concerned with answering questions based on visual content. Existing NeSy based VQA systems focusing on parsing questions and visual scenes into structured representations for cross-modality reasoning. For textual semantic parsing, Andreas et al. proposed a Neural Module Network that interprets questions as executable programs composed of learnable neural modules that can be directly applied to images. A module is typically implemented by the neural attention operation and corresponds ⁣{}_{\!} to ⁣{}_{\!} a ⁣{}_{\!} certain ⁣{}_{\!} atomic ⁣{}_{\!} rea- soning ⁣{}_{\!} step, ⁣{}_{\!} such ⁣{}_{\!} as ⁣{}_{\!} recognizing ⁣{}_{\!} objects, ⁣{}_{\!} classifying ⁣{}_{\!} colors, ⁣{}_{\!} etc. ⁣{}_{\!} This ⁣{}_{\!} pioneering ⁣{}_{\!} work ⁣{}_{\!} inspired ⁣{}_{\!} many ⁣{}_{\!} subsequent ⁣{}_{\!} studies ⁣{}_{\!} . In the context of visual semantic parsing, a few NeSy based VQA systems utilize scene graphs as structured, symbolic representations of visual scenes, and derive answers by graphical reasoning. As for visual grounding, which studies how to localize objects in visual scenes based on language descriptions, Hsu et al. combine large language-to-code models with modular neural networks to parse natural language into symbolic programs for 3D visual reasoning.

∙\bullet Robotics ⁣{}_{\!} and ⁣{}_{\!} Control: ⁣{}_{\!} Robots ⁣{}_{\!} are ⁣{}_{\!} complex ⁣{}_{\!} systems ⁣{}_{\!} with mechanical elements and controllers. To build an autono- mous embodied ⁣{}_{\!} system, ⁣{}_{\!} we ⁣{}_{\!} are ⁣{}_{\!} supposed ⁣{}_{\!} to ⁣{}_{\!} design ⁣{}_{\!} suitable policies that ensure the system operates within reasonable mechanical constraints. Moreover, safety and data efficiency are also crucial for constructing the system. So far, many NeSy based autonomous systems have been developed, where high-level, symbolic planning is generated to guide low-level RL for task execution and ⁣{}_{\!} learning. ⁣{}_{\!} They ⁣{}_{\!} recycle ⁣{}_{\!} the ⁣{}_{\!} common ⁣{}_{\!} practice ⁣{}_{\!} in ⁣{}_{\!} this ⁣{}_{\!} field that ⁣{}_{\!} decision-making ⁣{}_{\!} in ⁣{}_{\!} robotics ⁣{}_{\!} environments ⁣{}_{\!} can ⁣{}_{\!} be ⁣{}_{\!} decomposed into a high level (i.e., what to do) and a low level (i.e., how to do it). Another major source of their idea can be traced back to the notion of hierarchical RL which seeks to impose the task structure onto the learned policy. In particular, learns parameterized polices in combination with operators and samplers, which are packaged into modular neuro-symbolic skills and easily reused in new tasks. combines geometric and symbolic scene graphs as a two-level abstraction of manipulation scenes and leverages GNNs for predicting high-level task plans and low-level motions. In , a program search method is proposed for autonomous driving decision module design, where differentiable neuro-symbolic programs that specify all the behaviors for reactive and deliberative autonomous driving are synthesized and can be end-to-end trained with the whole system. As a result, more interpretable and transparent decision-making process can be delivered.

∙\bullet Visual Scene Understanding: In the context of interpreting high-level semantics from visual perception, there are a set of NeSy models that seek to exploit external symbolic knowledge regarding the relations between visual semantics and structured properties of novel objects to improve the robustness and performance. For example, in LTN , part-of relations between objects are formalized in the form of first-order logic, and converted into differentiable training objectives for end-to-end object classification learning. In HSS , semantic concepts and their complex meronymy relations are organized into a tree/directed acyclic graph, from which a set of con- strains are derived for boosting network training for hierar- chical semantic segmentation. In , symbolic knowledge about visual relations are expressed as propositional formu- lae and embedded onto a manifold via a GNN. The authors ⁣{}_{\!} also introduce semantic regularization and heterogeneous node embedding to enhance the semantic fidelity and expressiveness of the embeddings for visual relation understanding. Li et al. proposed LogicSeg, a NeSy based visual semantic parser that formalizes the complex meronymy and exclusion relations among symbolic concepts as first-order logic rules. After fuzzy logic-based continuous relaxa- tion, the logical formulae are grounded onto data and neural computational graphs for end-to-end network training and encapsulated into an iterative optimization process for network feed-forward inference. Recently, LLM based AI agents are developed for solving complex ⁣{}_{\!} vision ⁣{}_{\!} tasks. ⁣{}_{\!} They ⁣{}_{\!} employ ⁣{}_{\!} LLMs ⁣{}_{\!} to ⁣{}_{\!} automatically ⁣{}_{\!} crate ⁣{}_{\!} plans and ⁣{}_{\!} execute ⁣{}_{\!} the ⁣{}_{\!} plan ⁣{}_{\!} by ⁣{}_{\!} systematically invoking external tools (e.g., off-the-shelf specialized models) to get the solution.

∙\bullet Mathematical Reasoning: As a distinct and specialized capability inherent in humans, mathematical reasoning has garnered substantial attention within AI community. This multifaceted skill encompasses linguistic reasoning, visual reasoning, common sense reasoning, logical reasoning, numerical reasoning, and symbolic reasoning . Human approach to understanding and solving mathematical problems is not primarily rooted in experience and evidence, but on the basis of learning, inferring, and applying laws, axioms, and symbolic manipulation rules . The structured and reasoning-heavy nature of mathematical problems enables the construction of NeSy based solvers . For mathematical problem solving, introduced tree structured decoder to explicitly explore the abstract syntax tree of mathematical expressions, and stimulated many follow-up efforts . For theorem proving, considers syntax trees of formulas as graphs and apply message-passing for higher-order proof search. For handwritten formula evaluation, models the symbol solution states as a Boltzmann distribution, avoiding expensive state searching and facilitating mutually beneficial interactions between network training and symbolic reasoning. For discovering faster matrix multiplication algorithms, makes use of a Monte Carlo tree search planning pro- cedure, aided by DNNs.

Performance Comparison

To offer more empirical insights, in this section we tabulate

the performance of some of the NeSy algorithms discussed before. As NeSy has become a quite broad research field that covers various application tasks, and the algorithmic design is often highly customized for each task, it is infeasible to compare all the NeSy algorithms on a common task or dataset. Therefore, we select three representative tasks of NeSy, based on our review in Sec. 5, for performance evaluation. The performance scores are either obtained from our own implementation or collected from the original papers.

As a fundamental task in organic synthesis, retrosynthesis prediction aims to predict the reactants given a core product.

∙\bullet Dataset: ⁣{}_{\!} We adopt the widely-used retrosynthesis prediction dataset USPTO-50K for evaluation. USPTO-50K comprises about 50,00050,000 reactions with precise atom mappings between reactants and products. The 80%/10%/10% of the total 5050K reactions are set as train/val/test data. For fair comparison, all the experiments are conducted without knowing the reaction class in advance.

∙\bullet Benchmarking Algorithms: For thorough assessment, we involve four Nesy based retrosynthesis algorithms (i.e., NSR ⁣{}_{\!} , ⁣{}_{\!} GLN ⁣{}_{\!} , ⁣{}_{\!} MEGAN ⁣{}_{\!} , ⁣{}_{\!} Graph2Edits ⁣{}_{\!} ), as well as five purely neural methods (i.e., Seq2seq ⁣{}_{\!} , ⁣{}_{\!} GTA , ⁣{}_{\!} RetroPrime ⁣{}_{\!} , ⁣{}_{\!} Dual-TF ⁣{}_{\!} , ⁣{}_{\!} GraphRetro ⁣{}_{\!} ).

∙\bullet Evaluation ⁣{}_{\!} Metric: ⁣{}_{\!} As ⁣{}_{\!} standard, ⁣{}_{\!} Top-k ⁣{}_{\!} exact ⁣{}_{\!} match ⁣{}_{\!} accuracy is used as the evaluation metric. It is computed as the ratio that one of the Top-k predicted results exactly match the ground truth, where k ranges from {1, 3, 5, 10, 20, 50}.

∙\bullet Result: As shown in Table ⁣{}_{\!} V, the newly proposed Nesy based solution (i.e., Graph2Edits ⁣{}_{\!} ) has a clear advantage over conventional neural models such as Dual-TF and GraphRetro ⁣{}_{\!} , yielding improvements of 1.8% and 1.4% on Top-1 Exact match accuracy. This confirms the efficacy of data-and knowledge-driven methods in organic chemistry.

Visual semantic parsing, i.e., interpreting high-level semantic concepts of visual stimuli at pixel level, is a fundamental and challenging task in the field of computer vision.

∙\bullet Dataset: ⁣{}_{\!} We ⁣{}_{\!} select ⁣{}_{\!} PASCAL-Person-Part ⁣{}_{\!} , ⁣{}_{\!} a ⁣{}_{\!} widely-used dataset for visual semantic parsing, to evaluate the performance. PASCAL-Person-Part consists of 1,7161,716/1,8171,817 images for train/test. It provides dense annotations for 2020 fine-grained human parts (e.g., head, left-arm) from which a three-layer label hierarchy can be derived: the fine-grained parts belong to two superclasses, upper-body and lower-body, which are further merged into full-body.

∙\bullet Benchmarking Algorithms: For performance comparison, we involve four NeSy based structured visual parsers (i.e., CNIF ⁣{}_{\!} , ⁣{}_{\!} HHP ⁣{}_{\!} , ⁣{}_{\!} HSSN ⁣{}_{\!} , ⁣{}_{\!} LogicSeg ⁣{}_{\!} ) ⁣{}_{\!} which ⁣{}_{\!} ex- ploit the three-level human semantic hierarchy. For a comprehensive evaluation, we also include a group of hierarchy-agnostic segmentation algorithms ⁣{}_{\!} (i.e., ⁣{}_{\!} DeepLabV3+ ⁣{}_{\!} , PCNet ⁣{}_{\!} , CrossSeg ⁣{}_{\!} , ProtoSeg ⁣{}_{\!} , Mask2Former , GMMSeg ⁣{}_{\!} , ClustSeg ⁣{}_{\!} ), whose segmentation results on coarse-grained semantics are simply obtained by merging the predictions of the corresponding subclasses.

∙\bullet Evaluation Metric: As customary, we employ the mean intersection-over-union (mIoU) for evaluation. As in , we further report the average score for each hierarchy level ll (denoted as mIoUl) for detailed analysis.

∙\bullet Result: Table VI demonstrates that, the newest NeSy based visual semantic parser, i.e., LogicSeg ⁣{}_{\!} , achieves superior performance over ClustSeg , the current top-leading purely neural solution, by 0.63%/1.11%/0.32% over the three semantic levels, in terms of mIoU. This suggests the great potential of integrating symbolic reasoning and sub-symbolic learning in large-scale machine perception.

3 Performance Benchmarking: Math Word Problem Solving

The task of solving math word problems (MWPs) is to automatically answer a mathematical question that is described in natural language. MWP solving is an important natural language understanding task that requires logical reasoning over the quantities presented in the context to compute the numerical answer.

∙\bullet Dataset: ⁣{}_{\!} Math23K ⁣{}_{\!} , ⁣{}_{\!} a ⁣{}_{\!} large-scale ⁣{}_{\!} MWP ⁣{}_{\!} dataset, ⁣{}_{\!} is ⁣{}_{\!} used in ⁣{}_{\!} our ⁣{}_{\!} experiments. ⁣{}_{\!} Math23K ⁣{}_{\!} has ⁣{}_{\!} a ⁣{}_{\!} total ⁣{}_{\!} of ⁣{}_{\!} 23,161 ⁣{}_{\!} real math word problems for elementary school students with prob- lem ⁣{}_{\!} descriptions, ⁣{}_{\!} structured ⁣{}_{\!} equations ⁣{}_{\!} and ⁣{}_{\!} answers. ⁣{}_{\!} The ⁣{}_{\!} pro- blems are crawled from multiple online education websites and solved by one-unknown-variable linear expressions.

∙\bullet Benchmarking Algorithms: We compare the performance

of ⁣{}_{\!} ten ⁣{}_{\!} famous ⁣{}_{\!} MWP ⁣{}_{\!} solvers; ⁣{}_{\!} four ⁣{}_{\!} of them are purely based on neural networks, namely DNS ⁣{}_{\!} , Math-EN ⁣{}_{\!} , T-RNN ⁣{}_{\!} , GROUP-ATT ⁣{}_{\!} . The remaining six solvers, namely TSD , GTS , Graph2Tree , NSS , HMS , BERT-Tree , are NeSy based.

∙\bullet Evaluation Metric: Here the standard evaluation metric in MWP solving, namely answer accuracy, is adopted.

∙\bullet Result: Table VII shows that NeSy based solvers generally outperform the four neural competitors. This proves the efficacy of NeSy in dealing with reasoning-heavy symbolic problems. Moreover, with the aid of pre-trained BERT, BERT-Tree provides impressive performance, suggesting the power of combining NeSy with LLMs.

Open Challenges

While recent years have witnessed remarkable progress in NeSy, there still exist several open challenges to overcome.

∙\bullet Scalability: Current NeSy systems are still struggling with large-scale ⁣{}_{\!} symbolic/logic ⁣{}_{\!} reasoning. ⁣{}_{\!} First, ⁣{}_{\!} the ⁣{}_{\!} increasing ⁣{}_{\!} ex- pressivity of symbolic/logic rules , such as the inclusion of universal quantification over variables , and the complex syntax in higher-order logic, typically comes with growing computational complexity. Second, the frequent use of symbolic knowledge also impedes the applicability of NeSy systems in large-scale applications in the wild , since grounding massive symbolic knowledge on real-world examples is time-consuming. Third, collecting large-scale symbolic knowledge, especially in specific domains, is often difficult and expensive. Finally, while recent NeSy systems are relatively easy to make a full use of rich data with the aid of modern connectionist tools, it is less clear whether they can indeed exhibit the desirable features, such as sound reasoning, out-of-distribution generalization, data-efficient learning, transparency, and transferability to new domains, at a large scale. These features are promised by NeSy’s symbolic aspect, but their realization in the face of real-world complexity warrants further investigation.

∙\bullet Compositional ⁣{}_{\!} Generalization: ⁣{}_{\!} As ⁣{}_{\!} discussed ⁣{}_{\!} in ⁣{}_{\!} Sec. ⁣{}_{\!} 3.1, compositionality, a central aspect of human intelligence, is among the most desirable characterizations that NeSy sys- tems ⁣{}_{\!} are ⁣{}_{\!} expected ⁣{}_{\!} to ⁣{}_{\!} offer. ⁣{}_{\!} It ⁣{}_{\!} requires ⁣{}_{\!} systematical ⁣{}_{\!} decom- position ⁣{}_{\!} and ⁣{}_{\!} recombination ⁣{}_{\!} of ⁣{}_{\!} the ⁣{}_{\!} learned ⁣{}_{\!} knowledge, ⁣{}_{\!} so as to ⁣{}_{\!} generalize ⁣{}_{\!} to ⁣{}_{\!} novel ⁣{}_{\!} reasoning ⁣{}_{\!} problems. ⁣{}_{\!} Whilst ⁣{}_{\!} a ⁣{}_{\!} few attempts have been made for algorithmic implementation of compositional generalization within the NeSy framework, they are primarily specialized for toy language games. As a result, whether NeSy is able to achieve human-

like ⁣{}_{\!} compositionality, ⁣{}_{\!} such ⁣{}_{\!} as ⁣{}_{\!} the ⁣{}_{\!} comprehensive ⁣{}_{\!} use ⁣{}_{\!} various forms of logics (e.g., modal, temporal, commonsense, epistemic, etc.) and different types of knowledge (e.g., declarative, procedural, causal, and relational, etc.) for generalization and reasoning, and apply the compositional skills to solve real-world problems, still remains a grand challenge.

∙\bullet  ⁣{}_{\!}Automated ⁣{}_{\!} Knowledge ⁣{}_{\!} Acquisition: ⁣{}_{\!} Symbolic knowledge ⁣{}_{\!} serves ⁣{}_{\!} as the foundation for developing a NeSy system; it influences the quality and scope of the system’s reasoning capability. ⁣{}_{\!} Nevertheless, ⁣{}_{\!} most ⁣{}_{\!} modern ⁣{}_{\!} NeSy ⁣{}_{\!} systems ⁣{}_{\!} simply take ⁣{}_{\!} the ⁣{}_{\!} knowledge ⁣{}_{\!} for ⁣{}_{\!} granted, ⁣{}_{\!} ignoring ⁣{}_{\!} two ⁣{}_{\!} crucial ⁣{}_{\!} issues: i) ⁣{}_{\!} how ⁣{}_{\!} to ⁣{}_{\!} acquire ⁣{}_{\!} domain-specific ⁣{}_{\!} knowledge ⁣{}_{\!} that ⁣{}_{\!} is ⁣{}_{\!} required for the system; and ii) how to ensure the completeness and adequacy ⁣{}_{\!} of ⁣{}_{\!} the ⁣{}_{\!} knowledge ⁣{}_{\!} for ⁣{}_{\!} supporting ⁣{}_{\!} the ⁣{}_{\!} system’s ⁣{}_{\!} fun- ctionality. ⁣{}_{\!} Given ⁣{}_{\!} the ⁣{}_{\!} aforementioned ⁣{}_{\!} challenge ⁣{}_{\!} of ⁣{}_{\!} scalability, ⁣{}_{\!} knowledge acquisition seems a bottleneck in the process of developing ⁣{}_{\!} NeSy ⁣{}_{\!} systems ⁣{}_{\!} in ⁣{}_{\!} large-scale ⁣{}_{\!} and ⁣{}_{\!} real-world ⁣{}_{\!} application scenarios. This calls for the automatic acquisition of knowledge (preferably, from different data sources). This is also closely relevant to the concept of learning to reason , which studies the entire process of learning a knowledge base representation from examples, and then reasoning with that knowledge base by querying with similar examples. In fact, automated knowledge acquisition, which is essentially a problem in modeling a human expert’s introspective capabilities, has experienced a long-period history of development in the field of AI . Despite recent efforts in this direction are mainly restricted to the construction of concep- tual knowledge graphs , we believe the integration of automated knowledge acquisition, symbolic reasoning, and sub-symbolic ⁣{}_{\!} learning ⁣{}_{\!} will ⁣{}_{\!} lead ⁣{}_{\!} to ⁣{}_{\!} more ⁣{}_{\!} complete ⁣{}_{\!} and ⁣{}_{\!} powerful NeSy ⁣{}_{\!} systems. ⁣{}_{\!} Such ⁣{}_{\!} an ⁣{}_{\!} integration touches many fundamental challenges and aspects of NeSy and even AI, such as: i) which kind of knowledge representation is more favored for building NeSy systems; ii) how to abstract knowledge from data; and iii) how to build a close and mutual-feedback loop between deductive reasoning and inductive learning, i.e., grounding knowledge onto data to guide the practice, and updating knowledge according to the practical results.

∙\bullet Recursive Neuro[Symbolic] Engine: Another invaluable research direction is the construction of a Neuro[Symbolic] engine (i.e., the Type 6 NeSy system elaborated in Sec. 4.2) that can deeply embed a symbolic reasoning engine inside a neural sub-symbolic engine. Unlike existing Type 1-5 NeSy systems and modern connectionist machines, such a Neuro[Symbolic] engine explores the mechanism of human intelligence more deeply: how neural activations, which are sub-symbolic and widely distributed in the human brain, give rise to complex behaviors that are symbolic, such as language and logical reasoning. In such a Neuro[Symbolic] engine, the neural part shall be trained with the guidance of the symbolic component’s reasoning results, which are derived from the symbolic knowledge, and recursively, the symbolic component shall be evolved by updating its knowledge according to the neural component’s feedback, which are induced from data. This loop is closely related to the aforementioned challenge of automated knowledge acquisition. In addition, the Neuro[Symbolic] engine provides a computational realization of Kahneman’s System 1 and System 2 theory 1 of cognition, which distinguishes between fast, intuitive, and unconscious System 1 thinking and slow, deliberate, and conscious System 2 thinking. Hence it can achieve both types of thinking and leverage their strengths.

∙\bullet Testbed for Metacognitive Skills of NeSy: From a practical perspective, though NeSy systems is widely regarded as one of the most promising avenues towards human-like AI , the main strands of NeSy’s applications are still limited to a handful of tasks (c.f., Sec. 5). Many of the application tasks are placed in simulators, designed around limited proof-of-concept settings, or with small examples, in contrast to the large vision we hold onto the metacognitive capabilities of human beings, such as productivity, systematicity, compositionality and inferential coherence of mental thought , causal and counterfactual thinking , deductive reasoning , interpretability , etc., as well as their extensive, daily use. To advance NeSy towards this vision, we need more challenging and appropriate playgrounds that seek fundamental progress of NeSy in mastering human metacognitive skills. Some promising domains for such benchmarks include social robotics, health informatics, hardware/software specification, and scientific problems in genomics, chemistry, and astronomy, where both large amounts of data and knowledge are available and the discovery of scientific hypotheses is needed.

∙\bullet NeSy ⁣{}_{\!} in ⁣{}_{\!} the ⁣{}_{\!} Big ⁣{}_{\!} Model ⁣{}_{\!} Era: ⁣{}_{\!} The ⁣{}_{\!} community ⁣{}_{\!} has ⁣{}_{\!} recently witnessed remarkable progress fueled by large AI models. Large AI models exhibit emergent abilities (e.g., in-context learning, chain-of-thought reasoning), and can accomplish diverse tasks in a zero-shot fashion or with the aid of a few examples, akin to human beings. Albeit these astonishing abilities, it is becoming increasingly clear that large AI models still suffer several deficiencies, such as their pronounced opacity, insatiable demand for data and computational resources, and tendency to generate nonsensical or unfaithful content, known as “hallucination”. These drawbacks reveal their inherent biases, lack of real-world understanding, and weakness in generalizing or reasoning beyond their scope. These are intrinsic limitations of connectionist models and exacerbated by the heightened sophistication and scale of ⁣{}_{\!} large ⁣{}_{\!} models. ⁣{}_{\!} With ⁣{}_{\!} regard ⁣{}_{\!} to ⁣{}_{\!} this, ⁣{}_{\!} it ⁣{}_{\!} is ⁣{}_{\!} appealing ⁣{}_{\!} to ⁣{}_{\!} explore the integration of large AI models and symbolic techniques. Such integration can address the limitations of big neural models and empower the symbolic part with the massive implicit knowledge encoded by the big models, hence stepping closer towards artificial general intelligence. The recent emerge of LLM based AI agents that can automatically compose external tools for real-world task solving supports this view, although they only achieve loose neural-symbolic integration (see Sec. 4.2 Type 2). In short, developing NeSy with big AI models is a promising and challenging direction that requires dense collaboration across different AI fields.

Conclusions

Though having a long history, NeSy remained a rather niche topic until recently when landmark advances in machine learning – pushed by the wave of deep learning – caused increasing interest in forming the bridge between neural and symbolic methods. In this work, we conducted a large-scale and up-to-date survey of the rapidly growing area, from five perspectives: i) A historical point of view – we provide a brief review of early research results of NeSy; ii) A motivation point of view – we clarify two major driving forces behind the field as well as the recent AI debate which promotes the research activity in NeSy; iii) A methodological point of view – we classify and analyze the contemporary NeSy systems from four dimensions: neural-symbolic integration, knowledge representation, knowledge embedding, and functionality; iv) An application point of view – we outline several key application areas including scientific discovery, programming systems, question-answering, vision-language analysis and reasoning, robotics and control, visual scene understanding, and mathematical reasoning; and v) An experimental point of view – we providing a performance benchmarking of several NeSy methods on three representative application tasks including retrosynthesis prediction, visual semantic parsing, and math word problem solving. In the end, we discuss outstanding challenges and areas for future research. Although a strong NeSy system is still far from achieved, given the significant progress in AI over the past decade, we remain optimistic about the future and believe NeSy is a promising direction for the development of the next generation of AI.

References