Analyzing Differentiable Fuzzy Implications

Emile van Krieken, Erman Acar, Frank van Harmelen

Introduction

In recent years, integrating symbolic and statistical approaches to AI gained considerable attention (?; ?). This research line has gained further traction due to recent influential critiques on purely statistical deep learning (?; ?). While deep learning has brought many important breakthroughs (?; ?; ?), there have been concerns about the massive amounts of data that models need to learn even a simple concept. In contrast, traditional symbolic AI could easily reuse concepts e.g., using only a single logical statement one can already express domain knowledge conveniently.

However, symbolic AI also has its weaknesses. One is scalability: dealing with large amounts of data while performing complex reasoning tasks. Another is not being able to deal with the noise and ambiguity of e.g. sensory data. The latter is also related to the well-known symbol grounding problem which (?) defines as how \saythe semantic interpretation of a formal symbol system can be made intrinsic to the system, rather than just parasitic on the meanings in our heads. In particular, symbols refer to concepts that have an intrinsic meaning to us humans, but computers manipulating these symbols cannot understand (or ground) this meaning. On the other hand, a properly trained deep learning model excels at modeling complex sensory data. Therefore, several recent approaches (?; ?; ?; ?; ?) aimed at interpreting symbols that are used in logic-based systems using deep learning models. These implement (?) \saya hybrid nonsymbolic/symbolic system […] in which the elementary symbols are grounded in […] non-symbolic representations that pick out, from their proximal sensory projections, the distal object categories to which the elementary symbols refer.

In this article, we introduce Differentiable Fuzzy Logics (DFL) which aims to integrate reasoning and learning by using logical formulas expressing background knowledge. In order to ensure loss functions are differentiable, DFL uses fuzzy logic semantics (?). Moreover, predicate, function and constant symbols are interpreted using a deep learning model. By maximizing the degree of truth of the background knowledge using gradient descent, both learning and reasoning are performed in parallel. The loss function can be used for weakly supervised learning (?), like detecting noisy or inaccurate supervision (?), or semi-supervised learning (?; ?). Under this interpretation, DFL corrects the predictions of the deep learning model when it is logically inconsistent.

Next, we present an analysis of the choice of fuzzy implication. A fuzzy implication generalizes the Boolean implication, and it is usually differentiable, which enables its use in DFL. Interestingly, the derivatives of the implications determine how DFL corrects the deep learning model when its predictions are inconsistent with the background knowledge. We show that the qualitative properties of these derivatives are integral to both the theory and practice of DFL.

More specifically, the main contribution of this article is to answer the following question: Which fuzzy logic implications have convenient theoretical properties when using them in gradient descent? To this end,

we introduce several known implications from fuzzy logic (Section 2) and the framework of Differentiable Fuzzy Logics (Section 3) that uses these implications;

we analyze the theoretical properties of fuzzy implications and introduce a new family of fuzzy implications called sigmoidal implications (Section 4);

we perform experiments to compare fuzzy implications in a semi-supervised experiment (Section 5).

we conclude with several recommendations for choices of fuzzy implications.

Background

We will denote predicates using cushion, variables by x,y,z,x1,...x,y,z,x_{1},... and objects by o1,o2,...,o_{1},o_{2},...,. For convenience, we will be limiting ourselves to function-free formulas in prenex normal form. Formulas in prenex normal form start with quantifiers followed by a quantifier-free subformula. An atom is P(t1,...,tm)\textsf{{P}}(t_{1},...,t_{m}) where t1,...,tmt_{1},...,t_{m} are terms. If t1,...,tmt_{1},...,t_{m} are all constants, we say it is a ground atom.

Fuzzy logic is a real-valued logic where truth values are real numbers in $where0denotescompletelyfalseand1denotescompletelytrue.Wewillbelookingatpredicatefuzzylogicsinparticular,whichextendpropositionalfuzzylogicswithuniversalandexistentialquantification.Inthistext,welimitourselvestotheclassicfuzzynegationwhere 0 denotes completely false and 1 denotes completely true. We will be looking at predicate fuzzy logics in particular, which extend propositional fuzzy logics with universal and existential quantification. In this text, we limit ourselves to the classic fuzzy negationN(a)=1-a$.

To properly introduce fuzzy implications, we require the notions of t-norms that generalize boolean conjunction, and t-conorms that generalize boolean disjunction. A t-norm is a function T:2→T:^{2}\rightarrow that is commutative, associative, increasing, and for all a∈a\in, T(1,a)=aT(1,a)=a. A t-conorm is a function S:2→S:^{2}\rightarrow that is commutative, associative, increasing, and for all a∈a\in, S(0,a)=aS(0,a)=a. T-conorms are constructed from a t-norm using S(a,b)=1−T(1−a,1−b)S(a,b)=1-T(1-a,1-b).

Fuzzy implications are used to compute the truth value of p→qp\rightarrow q. pp is called the antecedent and qq the consequent of the implication. We follow (?) and refer to it for details and proofs.

A fuzzy implication is a function I:2→I:^{2}\rightarrow so that for all a,c∈a,c\in, I(⋅,c)I(\cdot,c) is decreasing, I(a,⋅)I(a,\cdot) is increasing and for which I(0,0)=1I(0,0)=1, I(1,1)=1I(1,1)=1 and I(1,0)=0I(1,0)=0.

From this definition follows that I(0,1)=1I(0,1)=1. We next introduce several optional properties of fuzzy implications that we will use in our analysis.

left-neutrality (LN) if for all c∈c\in, I(1,c)=cI(1,c)=c (generalizes (1→p)≡p(1\rightarrow p)\equiv p);

the exchange principle (EP) if for all a,b,c∈a,b,c\in, I(a,I(b,c))=I(b,I(a,c))I(a,I(b,c))=I(b,I(a,c)) (generalizes p→(q→r)≡q→(p→r)p\rightarrow(q\rightarrow r)\equiv q\rightarrow(p\rightarrow r));

the identity principle (IP) if for all a∈a\in, I(a,a)=1I(a,a)=1 (generalizes the tautology p→pp\rightarrow p);

contrapositive symmetry (CP) if for all a,c∈a,c\in, I(a,c)=I(1−c,1−a)I(a,c)=I(1-c,1-a) (generalizes p→q≡¬q→¬pp\rightarrow q\equiv\neg q\rightarrow\neg p);

left-contrapositive symmetry (L-CP) if for all a,c∈a,c\in, I(1−a,c)=I(1−c,a)I(1-a,c)=I(1-c,a) (generalizes ¬p→q≡¬q→p\neg p\rightarrow q\equiv\neg q\rightarrow p);

right-contrapositive symmetry (R-CP) if for all a,c∈a,c\in, I(a,1−c)=I(c,1−a)I(a,1-c)=I(c,1-a) (generalizes p→¬q≡q→¬pp\rightarrow\neg q\equiv q\rightarrow\neg p).

Using a common construction, we find R-implications. They are the standard choice for implication in t-norm fuzzy logics.

Let TT be a t-norm. The function IT:2→I_{T}:^{2}\rightarrow is called an R-implication and defined as IT(a,c)=sup⁡{b∈∣T(a,b)≤c}I_{T}(a,c)=\sup\{b\in|T(a,b)\leq c\}.

The supremum of a set AA, denoted sup⁡{A}\sup\{A\}, is the lowest upper bound of AA. All R-implications are fuzzy implications, and all satisfy LN, IP and EP. Note that if a≤ca\leq c then IT(a,c)=1I_{T}(a,c)=1.

S-Implications

In classical logic, the (material) implication is defined using p→q=¬p∨qp\rightarrow q=\neg p\vee q. Generalizing this definition, we can use a t-conorm SS to construct a fuzzy implication.

Let SS be a t-conorm. The function IS:2→I_{S}:^{2}\rightarrow is called an S-implication and is defined for all a,c∈a,c\in as IS(a,c)=S(1−a,c)I_{S}(a,c)=S(1-a,c).

All S-implications ISI_{S} are fuzzy implications and satisfy every property from Definition 2 but IP.

Table 1 shows some common differentiable S-implications and R-implications.

Differentiable Fuzzy Logics

Differentiable Fuzzy Logics (DFL) are fuzzy logics with differentiable connectives for which differentiable loss functions can be constructed that represent logical formulas. Examples of logics in this family (?; ?; ?; ?; ?) will be discussed in Section 6. They use background knowledge to deduce the truth value of statements in unlabeled or poorly labeled data to be able to use such data during learning. This can be beneficial as unlabeled, poorly labeled and partially labeled data is cheaper and easier to come by.

We motivate the use of DFL with the following scenario: Assume we have an agent MM whose goal is to describe the scene on an image. It gets feedback from a supervisor SS, who does not have an exact description of these images available. However, SS does have a background knowledge base K\mathcal{K}, encoded in some logical formalism, about the concepts contained on the images. The intuition behind DFL is that SS can correct MM’s descriptions of scenes when they are not consistent with its knowledge base K\mathcal{K}.

‘Agent MM has to describe the image II in Figure 1 containing two objects, o1o_{1} and o2o_{2}. MM and the supervisor SS only know about the unary class predicates {chair,cushion,armRest}\{\textsf{{chair}},\textsf{{cushion}},\textsf{{armRest}}\} and the binary predicate {partOf}\{\textsf{{partOf}}\}. Since SS does not have a description of II, it will have to correct MM based on the knowledge base K\mathcal{K}. MM describes the image as follows, where the probability indicates the confidence in an observation:

Suppose that K\mathcal{K} contains the following logic formula which says objects that are a part of a chair are either cushions or armrests:

SS might now reason that since MM is relatively confident of chair(o1)\textsf{{chair}}(o_{1}) and partOf(o2,o1)\textsf{{partOf}}(o_{2},o_{1}) that the antecedent of this formula is satisfied, and thus cushion(o2)\textsf{{cushion}}(o_{2}) or armRest(o2)\textsf{{armRest}}(o_{2}) has to hold. Since p(cushion(o2)∣I,o2)>p(armRest(o2)∣I,o2)p(\textsf{{cushion}}(o_{2})|I,o_{2})>p(\textsf{{armRest}}(o_{2})|I,o_{2}), a possible correction would be to tell MM to increase its degree of belief in cushion(o2)\textsf{{cushion}}(o_{2}).

We would like to automate the kind of supervision SS performs in the previous example. Therefore, we next formally introduce DFL, in which truth values of ground atoms are in $$, and logical connectives are interpreted using fuzzy operators. DFL defines a new semantics using vector embeddings and functions on such vectors in place of classical semantics. In classical logic, a structure consists of a domain of discourse and an interpretation function, and is used to give meaning to the predicates. Similarly, in DFL a structure consists of a probability distribution defined on an embedding space and an embedded interpretation:

To address the symbol grounding problem (?), objects in the domain of discourse are dd-dimensional vectors of reals. Their semantics come from the underlying semantics of the vector space as terms are interpreted in a real (valued) world (?). Predicates are interpreted as functions mapping these vectors to a fuzzy truth value. Embedded interpretations are implemented using a neural network model with trainable network parameters θ{\boldsymbol{\theta}}. Different values of θ{\boldsymbol{\theta}} will produce different embedded interpretations ηθ\eta_{{\boldsymbol{\theta}}}. The domain distribution is used to limit the size of the vector space. For example, pp might be the distribution over images representing only the natural images.

Next, we define how to compute the truth value of formulas of DFL, which generalizes the computation of Real Logic (?). An aggregation operator is a function A:n→A:^{n}\rightarrow that is symmetric and increasing with respect to each argument, and for which A(0,...,0)=0A(0,...,0)=0 A(1,...,1)=1A(1,...,1)=1. A variable assignment μ\mu maps variable symbols xx to objects o∈Oo\in O. μ(x)\mu(x) retrieves the object o∈Oo\in O assigned to xx in μ\mu.

Let ⟨p,ηθ⟩\langle p,\eta_{{\boldsymbol{\theta}}}\rangle be a DFL structure, TT a t-norm, SS a t-conorm, II a fuzzy implication and AA an aggregation operator. Then the valuation function eηθ,p,T,S,I,Ae_{\eta_{{\boldsymbol{\theta}}},p,T,S,I,A} (or, for brevity, eθe_{{\boldsymbol{\theta}}}) computes the truth value of a formula φ\varphi in L\mathcal{L} given a variable assignment μ\mu. It is defined inductively as follows:

Equation 1 defines the fuzzy truth value of an atomic formula. μ\mu finds the objects assigned to the terms x1,...,xmx_{1},...,x_{m} resulting in a list of dd-dimensional vectors. These are the inputs to the interpretation of the predicate symbol ηθ(P)\eta_{{\boldsymbol{\theta}}}(\textsf{{P}}) to get a fuzzy truth value. Equations 2 - 5 define the truth values of the connectives using the operators T,ST,S and II.

Equation 6 defines the truth value of universally quantified formulas ∀x ϕ\forall x\ \phi. This is done by enumerating the domain of discourse o∈Oo\in O, computing the truth value of ϕ\phi with oo assigned to xx in μ\mu, and combining the truth values using an aggregation operator AA. When enumerating the objects is not viable, we can choose to sample a batch of objects to approximate the computation of the valuation. It is commonly assumed in Machine Learning (?)(p.109) that a dataset contains independent samples from the domain distribution pp and thus using such samples approximates sampling from pp. Unfortunately, by relaxing quantifiers in this way we lose soundness of the logic.

In DFL, the parameters θ{\boldsymbol{\theta}} are learned using fuzzy maximum satisfiability (?), which finds parameters that maximize the valuation of the knowledge base K\mathcal{K}.

Let K\mathcal{K} be a knowledge base of formulas, ⟨p,ηθ⟩\langle p,\eta_{{\boldsymbol{\theta}}}\rangle a DFL structure for the predicate symbols in K\mathcal{K} and eηθ,p,T,S,I,Ae_{\eta_{{\boldsymbol{\theta}}},p,T,S,I,A} a valuation function. Then the Differentiable Fuzzy Logics loss LDFL\mathcal{L}_{DFL} of a knowledge base of formulas K\mathcal{K} is computed using

where wφw_{\varphi} is the weight for formula φ\varphi which denotes the importance of the formula φ\varphi in the loss function. The fuzzy maximum satisfiability problem is the problem of finding parameters θ∗{\boldsymbol{\theta}}^{*} that minimize Equation 7:

This optimization problem can be solved using a gradient descent method. If the operators T,S,IT,S,I and AA are all differentiable, we can repeatedly apply the chain rule, i.e. reverse-mode differentiation, on the DFL loss LDFL(θn;O,K)\mathcal{L}_{DFL}({\boldsymbol{\theta}}_{n};O,\mathcal{K}), n=0,...,Nn=0,...,N. This procedure finds the derivative with respect to the truth values of the ground atoms ∂LDFL(θn;O,K)∂ηθn(P)(o1,...,om)\frac{\partial\mathcal{L}_{DFL}({\boldsymbol{\theta}}_{n};O,\mathcal{K})}{\partial\eta_{{\boldsymbol{\theta}}_{n}}(\textsf{{P}})(o_{1},...,o_{m})}. We can use these partial derivatives to update the parameters θn{\boldsymbol{\theta}}_{n} using the chain rule, resulting in a different embedded interpretation ηθn+1\eta_{{\boldsymbol{\theta}}_{n+1}}.

One particularly interesting property of Differentiable Fuzzy Logics is that the partial derivatives of the subformulas with respect to the satisfaction of the knowledge base have a somewhat explainable meaning. For example, turning back to Example 1, the computed partial derivatives reflect whether we should increase p(cushion(o2))p(\textsf{{cushion}}(o_{2})), that is, increase the agents belief in cushion(o2)\textsf{{cushion}}(o_{2}).

Differentiable Fuzzy Implications

A significant proportion of background knowledge is written as universally quantified implications of the form ∀x ϕ(x)→ψ(x)\forall x\ \phi(x)\rightarrow\psi(x), like ‘all humans are mortal’. The implication is used in two well known rules of inference. Modus ponens inference says that if ∀x ϕ(x)→ψ(x)\forall x\ \phi(x)\rightarrow\psi(x) and we know that ϕ(x)\phi(x), then also ψ(x)\psi(x). Modus tollens inference says that if ∀x ϕ(x)→ψ(x)\forall x\ \phi(x)\rightarrow\psi(x) and we know that ¬ψ(x)\neg\psi(x), then also ¬ϕ(x)\neg\phi(x), as if ϕ(x)\phi(x) were true, ψ(x)\psi(x) should also have been.

When the learning agent predicts a scene in which an implication is false, the supervisor has multiple choices to correct it. Consider the implication ‘all ravens are black’. There are 4 categories for this formula: black ravens (BR), non-black non-ravens (NBNR), black non-ravens (BNR) and non-black ravens (NBR). Assume our agent observes an NBR, which is inconsistent with the background knowledge. There are four options to consider.

Modus Ponens (MP): The antecedent is true, so by modus ponens, the consequent is also true. We trust the agent’s observation of a raven and believe it was a black raven (BR).

Modus Tollens (MT): The consequent is false, so by modus tollens, the antecedent is also false. We trust the agent’s observation of a non-black object and believe it was not a raven (NBNR).

Distrust: We believe the agent is wrong both about observing a raven and a non-black object and it was a black object which is non-raven (BNR).

Exception: We trust the agent and ignore the fact that its observation goes against the background knowledge. Hence, it has to be a non-black raven (NBR).

The distrust option seems somewhat useless. The exception option can be correct, but we cannot know when there is an exception from the agent’s observations alone. In such cases, DFL would not be very useful since it would not teach the agent anything new.

We can safely assume that there are far more non-black objects which are not ravens than there are ravens. Thus, from a statistical perspective, it is most likely that the agent observed an NBNR. This shows the imbalance associated with the implication, which was first noted in (?) for the Reichenbach implication. It is quite similar to the class imbalance problem in Machine Learning (?) in that the real world has far more ‘negative’ (or contrapositive) examples than positive examples of the background knowledge.

This problem is closely related to the Raven paradox (?; ?) from the field of confirmation theory which ponders what evidence can confirm a statement like ‘ravens are black’. It is usually stated as follows:

Premise 1: Observing examples of a statement contributes positive evidence towards that statement.

Premise 2: Evidence for some statement is also evidence for all logically equivalent statements.

Conclusion: Observing examples of non-black non-ravens is evidence for ‘all ravens are black’.

The conclusion follows from the fact that ‘non-black objects are non-ravens’ is logically equivalent to ‘ravens are black’. Although we are considering logical validity instead of confirmation, we note that for DFL a similar thing happens. When we correct the observation of an NBR to a BR, the difference in truth value is equal to when we correct it to NBNR. More precisely, representing ‘ravens are black’ as I(a,b)I(a,b), where, for example, I(1,1)I(1,1) corresponds to BR:

as I(0,0)=I(1,1)=1I(0,0)=I(1,1)=1. Furthermore, when one agent observes a thousand BR’s and a single NBR, and another agent observes a thousand NBNR’s and a single NBR, their truth value for ‘ravens are black’ is equal. This seems strange, as the first agent has actually seen many ravens of which only a single exception was not black, while the second only observed many non ravens which were not black, among which a single raven that was not black either. Intuitively, the first agent’s beliefs seem to be more in line with the background knowledge. We will now proceed to analyse a number of implication operators in light of this discussion.

We define two functions for a fuzzy implication II:

dIcd_{Ic} is the derivative with respect to the consequent and dI¬ad_{I\neg a} is the derivative with respect to the negated antecedent. We choose to take the derivative with respect to the negated antecedent as it makes it easier to compare them.

A fuzzy implication II is called contrapositive differentiable symmetric if dIc(a,c)=dI¬a(1−c,1−a)d_{Ic}(a,c)=d_{I\neg a}(1-c,1-a) for all a,c∈a,c\in.

A consequence of contrapositive differentiable symmetry is that if c=1−ac=1-a, then the derivatives are equal since dIc(a,c)=dI¬a(1−c,1−a)=dI¬a(1−(1−a),c)=dI¬a(a,c)d_{Ic}(a,c)=d_{I\neg a}(1-c,1-a)=d_{I\neg a}(1-(1-a),c)=d_{I\neg a}(a,c). This could be seen as the ‘distrust’ option in which it increases the consequent and negated antecedent equally.

If a fuzzy implication II is contrapositive symmetric, it is also contrapositive differentiable symmetric.

Say we have an implication II that is contrapositive symmetric. We find that dIc(a,c)=∂I(a,c)∂cd_{Ic}(a,c)=\frac{\partial I(a,c)}{\partial c} and dI¬a(1−c,1−a)=−∂I(1−c,1−a)∂1−cd_{I\neg a}(1-c,1-a)=-\frac{\partial I(1-c,1-a)}{\partial 1-c}. Because II is contrapositive symmetric, I(1−c,1−a)=I(a,c)I(1-c,1-a)=I(a,c). Thus, dI¬a(1−c,1−a)=−∂I(a,c)∂1−c=∂I(a,c)∂c=dIc(a,c)d_{I\neg a}(1-c,1-a)=-\frac{\partial I(a,c)}{\partial 1-c}=\frac{\partial I(a,c)}{\partial c}=d_{Ic}(a,c). ∎

In particular, by this proposition all S-implications are contrapositive differentiable symmetric. This says that there is no difference in how the implication handles the derivatives with respect to the consequent and antecedent.

If an implication II is left-neutral, then dIc(1,c)=1d_{Ic}(1,c)=1. If, in addition, II is contrapositive differentiable symmetric, then dI¬a(a,0)=1d_{I\neg a}(a,0)=1.

First, assume II is left-neutral. Then for all c∈c\in, I(1,c)=cI(1,c)=c. Taking the derivative with respect to cc, it turns out that dIc(1,c)=1d_{Ic}(1,c)=1. Next, assume II is contrapositive differentiable symmetric. Then, dIc(1,c)=dI¬a(1−c,1−1)=dI¬a(1−c,0)=1d_{Ic}(1,c)=d_{I\neg a}(1-c,1-1)=d_{I\neg a}(1-c,0)=1. As 1−c∈1-c\in, dI¬a(a,0)=1d_{I\neg a}(a,0)=1. ∎

All S-implications and R-implications are left-neutral, but only S-implications are all also contrapositive differentiable symmetric. The derivatives of R-implications vanish when a≤ca\leq c, that is, on no less than half of the domain. Note that the plots in this section are rotated so that the smallest value is in the front. In particular, plots of the derivatives of the implications are rotated 180 degrees compared to the implications themselves.

Implications based on the Gödel t-norm (TG(a,b)=min⁡(a,b)T_{G}(a,b)=\min(a,b)) make strong discrete choices and increase at most one of their outputs. The two associated implications are shown in Figure 2. The Gödel implication is a simple R-implication with the following derivatives:

The Gödel implication increases the consequent whenever a>ca>c, and the antecedent is never changed. This makes it a poorly performing implication in practice. For example, consider a=0.1a=0.1 and c=0c=0. Then the Gödel implication increases the consequent, even if the agent is fairly certain that neither is true. Furthermore, as the derivative with respect to the negated antecedent is always 0, it can never choose the modus tollens correction, which, as we argued, is actually often the best choice.

The derivatives of the Kleene-Dienes implication are

Or, simply put, if we are more confident in the truth of the consequent than in the truth of the negated antecedent, increase the truth of the consequent. Otherwise, decrease the truth of the antecedent. This decision can be somewhat arbitrary and does not take into account the imbalance of modus ponens and modus tollens.

Łukasiewicz Implication

The Łukasiewicz implication is both an S- and an R-implication. It has the simple derivatives

Whenever the implication is not satisfied because the antecedent is higher than the consequent, it simply increases the negated antecedent and the consequent until it is lower. This could be seen as the ‘distrust’ choice as both observations of the agent are equally corrected, and so does not take into account the imbalance between modus ponens and modus tollens cases. The derivatives of the Gödel implication IGI_{G} are equal to those of ILKI_{LK} except that IGI_{G} always has a zero derivative for the negated antecedent.

Product-based Implications

The product t-norm is given as TP(a,b)=a⋅bT_{P}(a,b)=a\cdot b. The associated R-implication is called the Goguen implication. We plot this implication in Figure 3. The derivatives of IGGI_{GG} are

We plot these in Figure 4. This derivative is not very useful. First of all, both the modus ponens and modus tollens derivatives increase with ¬a\neg a. This is opposite of the modus ponens rule as when the antecedent is low, it increases the consequent most. For example, if raven is 0.1 and black is 0, then the derivative with respect to black is 10, because of the singularity when aa approaches 0.

The derivatives of the Reichenbach implication are given by:

These derivatives closely follow modus ponens and modus tollens inference. When the antecedent is high, increase the consequent, and when the consequent is low, decrease the antecedent. However, around (1−a)=c(1-a)=c, the derivative is equal and the ‘distrust’ option is chosen. This can result in counter-intuitive behaviour. For example, if the agent predicts 0.6 for raven and 0.5 for black and we use gradient descent until we find a maximum, we could end up at 0.3 for raven and 1 for black. We would end up increasing our confidence in black as raven was high. However, because of additional modus tollens reasoning, raven is barely true.

Furthermore, if the agent mostly predicts values around a=0, c=0a=0,\ c=0 as a result of the modus tollens case being the most common, then a majority of the gradient decreases the antecedent as dIRC¬a(0,0)=1d_{I_{RC}\neg a}(0,0)=1. We next identify two methods that counteract this behavior.

Log product aggregator

The first method for counteracting the ‘corner’ behavior notes that different aggregators change how the derivatives of the implications behave. Note that the aggregator based on the product t-norm is AP(x1,...,xn)=∏i=1nxiA_{P}(x_{1},...,x_{n})=\prod_{i=1}^{n}x_{i}. As formulas are in prenex normal form, maximizing this aggregator is equivalent to maximizing the logarithm of this aggregator, which gives Alog⁡P(x1,...,xn)=∑i=1nlog⁡(xi)A_{\log P}(x_{1},...,x_{n})=\sum_{i=1}^{n}\log(x_{i}) that is reminiscent of the cross-entropy loss function. Using the chain rule, we find that the negated antecedent derivative becomes:

As this divides by the truth value of the implication, implications that do not have a high truth value get stronger derivatives. We plot the negated antecedent derivative for the Reichenbach implication when using the log-product aggregator in Figure 5. Note that the derivative with respect to the negated antecedent in ai=0a_{i}=0, ci=0c_{i}=0 is still 1. By differentiable contrapositive symmetry, the consequent derivative is 0. Therefore, when using the log-product aggregator, one of antecedent and consequent will still have a gradient.

Sigmoidal Implications

For the second method for tackling the corner problem, we introduce a new class of fuzzy implications formed by transforming other fuzzy implications using the sigmoid function and translating it so that the boundary conditions still hold.The derivation, along with several proofs of properties, can be found at https://github.com/HEmile/differentiable-fuzzy-logics/blob/master/appendix_sigmoidal_implications.pdf.

If II is a fuzzy implication, then the II-sigmoidal implication σI\sigma_{I} is given for some s>0s>0 as

where σ(x)=11+ex\sigma(x)=\frac{1}{1+e^{x}} denotes the sigmoid function.

Here ss controls the ‘spread’ of the curve. σI\sigma_{I} is the function σ(s⋅(I(a,c)−12))\sigma\left(s\cdot\left(I(a,c)-\frac{1}{2}\right)\right) linearly transformed so that its codomain is the closed interval $..\sigma_{I}isafuzzyimplicationinthesenseofDefinition1.Furthermore,is a fuzzy implication in the sense of Definition 1. Furthermore,\sigma_{I}satisfiestheidentityprincipleifsatisfies the identity principle ifIdoes,andiscontrapositive(differentiable)symmetricifdoes, and is contrapositive (differentiable) symmetric ifIis.WeplottheReichenbach−sigmoidalimplicationis. We plot the Reichenbach-sigmoidal implication\sigma_{I_{RC}}inFigure6fortwovaluesofin Figure 6 for two values ofs.Notethatfor. Note that fors=0.01$, the plotted function is indiscernible from the plot of the Reichenbach implication in Figure 3 as the interval on which the sigmoid acts is extremely small and the sigmoidal transformation is almost linear. The derivative is computed as

The derivative keeps the properties of the original function but smoothes the gradient for higher values of ss. As the derivative of the sigmoid function (that is, σ(x)⋅(1−σ(x))\sigma(x)\cdot(1-\sigma(x))) cannot be zero, this derivative vanishes only when ∂I(a,c)∂¬a=0\frac{\partial I(a,c)}{\partial\neg a}=0 or ∂I(a,c)∂c=0\frac{\partial I(a,c)}{\partial c}=0.

We plot the derivatives for the Reichenbach-sigmoidal implication σIRC\sigma_{I_{RC}} in Figure 7. As expected, it is clearly differentiable contrapositive symmetric. Compared to the derivatives of the Reichenbach implication it has a small gradient in all corners. In Figure 8 we compare the consequent derivative of the normal Reichenbach implication with the Reichenbach-sigmoidal implication when using the log⁡\log product aggregator. A significant difference is that the sigmoidal variant is less ‘flat’ than the normal Reichenbach implication. This can be useful, as this means there is a larger gradient for values of cc that make the implication less true. In particular, the gradient at the modus ponens case (a=1, c=1a=1,\ c=1) and the modus tollens case (a=0, c=0a=0,\ c=0) are far smaller, which could help balancing the effective total gradient by solving the ‘corner’ problem of the Reichenbach implication. These derivatives are smaller for for higher values of ss.

Experiments

To get an idea of the practical behavior of these implications we now perform a series of simple experiments to analyze them in practice. In this section, we discuss experiments using the MNIST dataset of handwritten digits (?) to investigate the behavior of different fuzzy operators introduced in this paper.

To investigate the performance of the different configurations of DFL, we first introduce several useful metrics. In this section, we assume we are dealing with formulas of the form φ=∀x1,...,xm ϕ(x1,...,xm)→ψ(x1,...,xm)\varphi=\forall x_{1},...,x_{m}\ \phi(x_{1},...,x_{m})\rightarrow\psi(x_{1},...,x_{m}).

Given a labeling function ll that returns the truth value of a formula according to the data for instance μ\mu, the consequent and antecedent correctly updated magnitudes are the sum of partial derivatives for which the consequent or the negated antecedent is true:

That is, if the consequent is true in the data, we measure the magnitude of the derivative with respect to the consequent. The correctly updated ratios quantify what fraction of the updates are going in the right direction. When they approach 1, DFL will always increase the truth value of the consequent or negated antecedent correctly. When it not close to 1, we are increasing truth values of subformulas that are wrong, thus ideally, we want these measures to be high.

2 Experimental Setup

We use a knowledge base K\mathcal{K} of universally quantified logic formulas. There is a predicate for each digit, that is zero, one,...,eight\textsf{{zero}},\ \textsf{{one}},...,\textsf{{eight}} and nine. For example, zero(x)\textsf{{zero}}(x) is true whenever xx is a handwritten digit labeled with 0. Secondly, there is the binary predicate same that is true whenever both its arguments are the same digit. We next describe the formulas we use.

∀x,y zero(x)∧zero(y)→same(x,y)\forall x,y\ \textsf{{zero}}(x)\wedge\textsf{{zero}}(y)\rightarrow\textsf{{same}}(x,y), …, ∀x,y nine(x)∧nine(y)→same(x,y)\forall x,y\ \textsf{{nine}}(x)\wedge\textsf{{nine}}(y)\rightarrow\textsf{{same}}(x,y). If both xx and yy are handwritten zeros, for example, then they represent the same digit.

∀x,y zero(x)∧same(x,y)→zero(y)\forall x,y\ \textsf{{zero}}(x)\wedge\textsf{{same}}(x,y)\rightarrow\textsf{{zero}}(y), …, ∀x,y nine(x)∧same(x,y)→nine(y)\forall x,y\ \textsf{{nine}}(x)\wedge\textsf{{same}}(x,y)\rightarrow\textsf{{nine}}(y). If xx and yy represent the same digit and one of them represents zero, then the other one does as well.

∀x,y same(x,y)→same(y,x)\forall x,y\ \textsf{{same}}(x,y)\rightarrow\textsf{{same}}(y,x). This formula encodes the symmetry of the same predicate.

We split the MNIST dataset so that 1% of it is labeled and 99% is unlabeled. We use two models.Code is available at https://github.com/HEmile/differentiable-fuzzy-logics. Given a handwritten digit x{\boldsymbol{x}}, the first model pθ(y∣x)p_{\boldsymbol{\theta}}(y|{\boldsymbol{x}}) computes the distribution over the 10 possible labels. We use 2 convolutional layers with max pooling, the first with 10 and the second with 20 filters, and two fully connected hidden layers with 320 and 50 nodes and a softmax output layer, which is trained using cross entropy. The probability that same(x1,x2)\textsf{{same}}({\boldsymbol{x}}_{1},{\boldsymbol{x}}_{2}) for two handwritten digits x1{\boldsymbol{x}}_{1} and x2{\boldsymbol{x}}_{2} holds is modeled by pθ(same∣x1,x2)p_{\boldsymbol{\theta}}(\textsf{{same}}|{\boldsymbol{x}}_{1},{\boldsymbol{x}}_{2}). This takes the 50-dimensional embeddings of x1{\boldsymbol{x}}_{1} and x2{\boldsymbol{x}}_{2} of the fully connected hidden layer ex1e_{{\boldsymbol{x}}_{1}} and ex2e_{{\boldsymbol{x}}_{2}}. These are used in a Neural Tensor Network (?) with a hidden layer of size 50. It is trained using binary cross entropy on the cross product of the labeled dataset. As there are far more negative examples than positive examples, we undersample the negative examples. The DFL loss is weighted by the DFL weight wDFLw_{DFL} and added to the other two losses.

For all experiments, we use the aggregator to the product aggregator with DFL weight of wdfl=10w_{dfl}=10, and optimize the logarithm of the truth value. For conjunction, we use the Yager t-norm with p=2p=2, defined as TY(a,b)=max⁡(1−((1−a)p+(1−b)p)1p,0)T_{Y}(a,b)=\max(1-((1-a)^{p}+(1-b)^{p})^{\frac{1}{p}},0).

3 Results

In Table 2, we compare different fuzzy implications. The Reichenbach implication and the Łukasiewicz implication work well, both having an accuracy around 97%. Using the Kleene Dienes implication surpasses the baseline as well.

As hypothesized, the Gödel implication and Goguen implication have worse performance than the supervised baseline. While the derivatives of ILKI_{LK} and IGI_{G} only differ in that IGI_{G} disables the derivatives with respect to negated antecedent, ILKI_{LK} performs among the best but IGI_{G} performs among the worst, suggesting that the derivatives with respect to the negated antecedent are required to successfully applying DFL. Note that all well performing implications are S-implications, which inherently balance derivatives with respect to the consequent and negated antecedent by being contrapositive differentiable symmetric.

Reichenbach-Sigmoidal Implication

The newly introduced Reichenbach-sigmoidal implication σIRC\sigma_{I_{RC}} is a promising candidate for the choice of implication.

Influence of Individual Formulas

Finally, we compare what the influence of the different formulas are in Table 3. Removing the reflexivity formula (3) does not largely impact the performance. The biggest drop in performance is by removing formula (1) that defines the same predicate. Using only formula (1) gets slightly better performance than only using formula (2), despite the fact that no positive labeled examples can be found using formula (1) as the predicates zero to nine are not in its consequent. Since 95% of the derivatives are with respect to the negated antecedent, this formula contributes by finding additional counterexamples. Furthermore, improving the accuracy of the same predicate improves the accuracy on digit recognition: Just using the reflexivity formula (3) has the highest accuracy when used individually, even though it does not use the digit predicates.

Analysis

Although DFL significantly improves on the supervised baseline and is thus suited for semi-supervised learning, it is currently not competitive with state-of-the-art methods like Ladder Networks (?) which has an accuracy of 98.9% for 100 labeled pictures and 99.2% for 1000.

Related Work

Differentiable Fuzzy Logics falls into the discipline of Statistical Relational Learning (?), which concerns models that can reason under uncertainty and learn relational structures like graphs. Special cases of DFL have been researched in several papers under different names. Real Logic (?) implements function symbols and uses a model called Logic Tensor Networks to interpret predicates. It uses S-implications. Real Logic is applied to weakly supervised learning on Semantic Image Interpretation (?; ?) and transfer learning in Reinforcement Learning (?). Semantic-based regularization (SBR) (?) applies DFL to kernel machines. They use R-implications, like (?) which simplifies the satisfiability computation and finds generalizations of common loss functions can be found. In (?), which employs the Goguen implication, DFL is applied to image generation. By using function symbols that represent generator neural networks, they create constraints that are used to create a semantic description of an image generation problem. The Reichenbach implication is used in (?) for relation extraction by using an efficient matrix embedding of the rules.

The regularization technique used in (?) is equivalent to the Łukasiewicz implication. Instead of using existing data, it finds a loss function which does not iterate over objects, yet can guarantee that the rules hold. A promising approach is using adversarial sets (?), which is a set of objects from the domain that do not satisfy the knowledge base, which are probably the most informative objects. Adversarial sets are applied to natural language interpretation in (?). Both papers use the Łukasiewicz implication.

Some approaches use probabilistic logics instead of fuzzy logics and interpret predicates probabilistically. DeepProbLog (?) and Semantic Loss (?) are probabilistic logic programming languages with neural predicates that compute the probabilities of ground atoms. They support automatic differentiation which can be used to back-propagate from the loss at a query predicate to the deep learning models that implement the neural predicates, similar to DFL.

Conclusion

We analyzed fuzzy implications in Differentiable Fuzzy Logics in order to understand how reasoning using implications behaves in a differentiable setting. We have found substantial differences between the properties of a large number of fuzzy implications, and showed that many of them, including some of the most popular implications, are highly unsuitable for use in a differentiable learning setting.

The Reichenbach implication has derivatives that are intuitive and that correspond to inference rules from classical logic. The Łukasiewicz implication is the best R-implication in our experiments. The Gödel and Goguen implications, on the other hand, were much less successful, performing worse than the supervised baseline. The newly introduced Reichenbach-sigmoidal implication performs best on the MNIST experiments. The spread of sigmoidal implications can be tweaked to decrease the imbalance of the derivatives with respect to the negated antecedent and consequent.

We noted an interesting imbalance between derivatives with respect to the negated antecedent and the consequent of the implication. Because the modus tollens case is much more common, we conclude that a large part of the useful inferences on the MNIST experiments are made by decreasing the antecedent, or by ‘modus tollens reasoning’. Furthermore, we found that derivatives with respect to the consequent often increase the truth value of something that is false as the consequent is false in the majority of times. Therefore, we argue that ‘modus tollens reasoning’ should be embraced in future research.

References