July 12, 2026
Neuro-symbolic AI based on \(IFOL_B\) is a way to combine neural learning and symbolic reasoning to overcome limitations of purely neural systems (like lack of interpretability and logical structure) with formal logical
machinery for self-reference [1].
In this paper we expand the cognitive power of \(IFOL_B\) by using the probability computation for the currently unknown sentences, based on Nilsson’s probability structure for the \(IFOL_B\). We introduce the global symmetry transformation that preserves the current knowledge database \(\mathcal{K}\) and logical deduction, and the local one used for real-time decisions
about concrete (sub)problems that involve only a very strict subset of \(IFOL_B\) predicates. The computation of probability density function \(KI\) in both cases, based on the Shannon’s
maximum information entropy, is provided by neural networks of this probabilistic neuro-symbolic AGI.
This theory on robot self-awareness is rooted in the development of Intensional First-Order Logic (IFOL) [2]. We argue that true "Strong AI" or "autoepistemic" robots require a symbolic architecture that allows them to reason about their own internal states as distinct from the external objects they perceive. This introduction is a short presentation of the my approach to AGI (Strong-AI) for a new generation of intelligent robots, recently published in the papers [3] and [4]. Neuro-symbolic AI attempts to integrate neural and symbolic architectures in a manner that addresses strengths and weaknesses of each, in a complementary fashion, in order to support robust strong AI capable of reasoning, learning, and cognitive modeling. In this approach to AGI I considered the Intensional First Order Logic [2] as a symbolic architecture of modern robots, able to use natural languages to communicate with humans and to reason about their own knowledge with self-reference and abstraction language property. In what follows we will consider the 4-valued typed \(IFOL_B\) based on Belnap’s bilattice of truth-values \(X = \mathcal{B}_4 = \{f,t,\bot,\top\}\) introduced in [1], [5] and denoted by \(IFOL_B\).
Intensional entities (or concepts) are such things as Propositions, Relations and Properties (PRP). What make them "intensional" is that they violate the principle of extensionality; the principle that extensional equivalence implies identity. All (or most) of these intensional entities have been classified at one time or another as kinds of Universals [6], in the case of many-valued \(IFOL_B\), \(D_I = D_1+D_2+D_3+...\) (with propositions (L-concepts) \(D_1\), and relational concepts \(D_n\), \(n\geq 2\)), and particulars \(D_{0}\) [7], which define the PRP domain \(\mathcal{D}= D_0+D_I\), with intensional mapping from the set of FOL formulae \(\mathcal{L}\) to these intensional concepts \[I:\mathcal{L}\rightarrow \mathcal{D}\] If \(\phi(\boldsymbol{x})\) is an open formula (virtual predicate with a list (a tuple) of free variables in \(\boldsymbol{x} =(x_1,...,x_n)\)), then \(I(\phi(\boldsymbol{x})) \in D_n\) is a n-ary concept. This concept mapping can be extended to the homomorphism \(I:\mathcal{A}\mathfrak{B}_{FOL}\rightarrow \mathcal{A}\mathfrak{B}_{int}\), between the FOL syntax algebra to the algebra of concepts. Thus, for a given 4-valued Herbrand base \(H\) of \(IFOL_B\) and \(X =\mathcal{B}_4 = \{f,t,\bot,\top\}\) the set of Belnap’s truth-values (false, true, unknown and inconsistent, respectively) and a Herbrand interpretation that must satisfy all built-in predicates as well) \[\label{eq:Herbrand} v:H\rightarrow X\tag{1}\] and its unique extension to all sentences \(v^*:\mathcal{L}_0 \rightarrow X\), with \(v^* \in \mathcal{I}_{MV}\) of the set of all well-defined Herbrand interpretations (that respects the built-in predicates in Definition 4 in [5]). Note that the set of all well-defined Herbrand interpretations \(\mathcal{I}_{H}\) is only a subset of functions in \(X^H\), because we have the built-in predicates as well (from Definition 4 in [5]), \[\label{eq:Herbrand2} \mathcal{I}_{H} \subset X^H ~~~with ~~~\mathcal{I}_{MV} = \{v^*:\mathcal{L}_0 \rightarrow X | v \in \mathcal{I}_H\}\tag{2}\] where \(v^*\) is the extension of Herbrand interpretation \(v\) to all sentences \(\mathcal{L}_0\) of our typed \(IFOL_B\).
Each extensional interpretation \(h\) assigns to the intensional elements of \(\mathcal{D}\) an appropriate extension: in the case of particulars \(u \in D_{0}\), \(h_0(u) \in D_0\), such that for each logic value \(a\in X\subset D_0\), \(h_0(a) = a\). Thus, we have the particular’s mapping \(h_0:D_0\rightarrow D_0\) and more generally (here \('+'\) is considered as disjoint union) from [5], \[\label{eq:dueM6} h = \sum_{i\in \mathbb{N}}h_i:\mathcal{D}\longrightarrow D_0+\sum_{i\geq 1}\mathfrak{Rm}_i\tag{3}\] where \(\mathfrak{Rm}_i\) are i-ary relations and where \(h_1\) assigns to each L-concept \(u \in D_1\) (of a sentence with truth-value \(a\in X\)), a relation composed by the single tuple \(h_1(u) = \{a\}\) if \(a \neq \bot\)), \(~\emptyset\) otherwise, and \(h_i:D_i\rightarrow \mathfrak{Rm}_i\), for \(i\geq 2\), that assigns a m-extension to non-sentence concepts (obtained, for example, an \((i-1)\)-ary predicate, so that the last \(i\)-th column of \(\mathfrak{Rm}_i\) is a truth-value of a ground atom of this predicate. In this way, the relation \(\mathfrak{Rm}_i\) represents the set of tuples of ground atoms of a given predicate both with their truth-values \(a \in X\).
The two-steps interpretation of \(IFOL_B\) based on two homomorphisms, fixed intensional \(I\), and extensionalization mapping \(h\), is provided by commutative diagram in Corollary 3 in [5], which defines the MV-interpretation \[I^*_{B} = h\circ I\].
So we can define the following set of different "possible worlds" for this many-valued \(IFOL_B\) [5] based on Herbrand interpretations (1 ) above: \[\label{eq:SetExtens} \mathcal{W}= \{I^*_{B} ~|~ v^* \in \mathcal{I}_{MV}\}~~\tag{4}\] such that for each Herbrand interpretation \(v\in \mathcal{I}_{H} \subset X^H\), we have a unique MV-interpretation \(I^*_{B} = h\circ I\), that is, the bijections \[\label{eq:SetExtens2} is_H:\mathcal{I}_{H} ~\simeq \mathcal{W}~~~~and~~~~ is_{MV}:\mathcal{I}_{MV}~\simeq \mathcal{W}\tag{5}\] with \(I^*_{B} = is_{MV}(v^*) = is_H(v)\) and \(v^* = is^{-1}_{MV}\circ is_H(v)\), that is, from the fact that \(I\) is fixed intensional interpretation, each possible world is fundamentally an extensionalization function \(h\), and we denote by \(\mathcal{E}_{in}\) the set of these extensionalization functions that respects all built-in predicates, with bijection [5] \[\label{eq:SetExtens3} is_{in}:\mathcal{W}\simeq \mathcal{E}_{in}\tag{6}\] The extensions of the IFOL concepts change in time (the robot’s knowledge), so that we can use for specification of \(h\) the time-index as their ordered representation.
In reflective languages, reification data is causally connected to the related reified aspect such that a modification to one of them affects the other; by using intensional FOL the robots can formalize also the natural language expressions "I see the blue color" by a predicate "See(I,blue color)" where the sense of the ground term "I" (Self, me)1 for a robot is the name of the main working coordination program which activate all other algorithms (neuro-symbolic AI subprograms) like visual recognition of color of the object in focus. But also the auto-conscience sentence like "I know that I see the blue color" by using abstracting operators "\(\lessdot\_\gtrdot\)" of intensional FOL, expressed by the predicate "Know(I,\(\lessdot\) See(I, blue color)\(\gtrdot\))", etc... If \(\phi(\boldsymbol{x})\) is a virtual predicate with a list (a tuple) of free variables in \(\boldsymbol{x} =(x_1,...,x_n)\) and \(\alpha\) is its subset of distinct variables, then \(\lessdot \phi(\boldsymbol{x}) \gtrdot_{\alpha}^{\beta}\) is a term, where \(\beta\) is the remaining set of free variables in \(\boldsymbol{x}\). The externally quantifiable variables are the free variables not in \(\alpha\). When \(n =0,~ \lessdot \phi \gtrdot\) is a term which denotes a proposition, for \(n \geq 1\) it denotes a n-ary concept.
Definition 1. Intensional abstraction convention:
From the fact that we can use any permutation of the variables in a given virtual predicate, we introduce the convention that \[\label{eq:abstrctConv} \lessdot \phi(\boldsymbol{x})\gtrdot_{\alpha}^{\beta}~~ is~a~ term~ obtained~ from~ virtual ~ predicate ~~\phi(\boldsymbol{x})\qquad{(1)}\] if \(\alpha\) is not empty* such that \(\alpha\bigcup\beta\) is the set of all variables in the list (tuple of variables) \(\boldsymbol{x} = (x_1,...,x_n)\) of the virtual predicate (an open logic formula) \(\phi\), and \(\alpha\bigcap\beta = \emptyset\), so that \(|\alpha|+|\beta| = |\boldsymbol{x}| = n\). Only the variables in \(\beta\) (which are the only free variables of this term), can be quantified. If \(\beta\) is empty then \(\lessdot \phi(\boldsymbol{x})\gtrdot_{\alpha}\) is a ground term. If \(\phi\) is a sentence and hence both \(\alpha\) and \(\beta\) are empty, we write simply \(\lessdot \phi \gtrdot\) for this ground term.*
More about this general definition of abstract terms can be find in [2]. In this paper we will use the most simple cases of ground terms \(\lessdot \phi \gtrdot\), where \(\phi\) is a sentence.
By using the intensional mapping \(I\), we are able to extend the simple assignment to variables to all (also abstracted) terms:
Definition 2. An assignment \(g:\mathcal{V}\rightarrow \mathcal{D}\) forb variables in \(\mathcal{V}\) is applied only to free variables in terms and formulae. Such an assignment \(g \in \mathcal{D}^{\mathcal{V}}\) can be recursively uniquely extended into the assignment \(g^*:\mathcal{T}\rightarrow \mathcal{D}\), where \(\mathcal{T}\) denotes the set of all terms (here \(I\) is an intensional interpretation of this FOL, as explained in what follows), by :
\(g^*(t) = g(x) \in \mathcal{D}\) if the term \(t\) is a variable \(x \in\mathcal{V}\).
\(g^*(t) = I(c) \in \mathcal{D}\) if the term \(t\) is a constant (nullary functional symbol) \(c\in P\).
If \(t\) is an abstracted term obtained for an open formula \(\phi_i\), \(\lessdot \phi_i(\boldsymbol{x}_i) \gtrdot_{\alpha_i}^{\beta_i}\), then we must restrict the assignment to \(g\in \mathcal{D}^{\beta_i}\) and to obtain recursive definition (when also \(\phi_i(\boldsymbol{x}_i)\) contains abstracted terms: \[\label{eq:assAbTerm} g^*(\lessdot \phi_i(\boldsymbol{x}_i)\gtrdot_{\alpha_i}^{\beta_i}) =_{def} \left\{ \begin{array}{ll} I(\phi_i(\boldsymbol{x}_i))~~ \in D_{|\alpha_i|+1}, & if \beta_i is empty\\ I(\phi_i(\boldsymbol{x}_i)[\beta_i /g(\beta_i)])~~ \in D_{|\alpha_i|+1}, & otherwise \end{array} \right.\qquad{(2)}\] where \(g(\beta) = g(\{y_1,..,y_m\}) = \{g(y_1),...,g(y_m)\}\) and \([\beta /g(\beta)]\) is a uniform replacement of each i-th variable in the set \(\beta\) with the i-th constant in the set \(g(\beta)\). Notice that \(\alpha\) is the set of all free variables in the formula \(\phi[\beta /g(\beta)]\).
If \(~t = \lessdot \phi_i\gtrdot\) is an abstracted term obtained from a sentence \(\phi_i\) then
\(g^*(\lessdot \phi_i\gtrdot) = I(\phi_i) \in D_0\).2
By introduction of the abstraction operators with autoepistemic capacities, suported by the \(Know\) (meta)predicate, we do not use more a pure logical deduction of the standard FOL, but a kind of autoepistemic deduction [8], [9] with a proper set of new axioms for \(IFOL_B\) provided in [5]. It has been demonstrated that in such a minimal intensional enrichment of standard (extensional) FOL, we obtain exactly the Montague’s definition of the intension (see Proposition 5 in [2]).
We recall that each robot’s extensionalitation function \(h\) in (3 ) is indexed by the time-instance. Clearly, the robots knowledge changes in time and hence determines the extensionalization function \(h\) in any given instance of time, based on robots experiences. Thus, as for humans, also the robot’s knowledge and logic is a kind of temporal logic, and evolves with time. Note that the explicit (conscious) robot’s knowledge in actual world \(\hbar\) here is represented by the ground atoms of the \(Know\) predicate, for a given assignments of variables in \(\mathcal{V}\), \(g:\mathcal{V}\rightarrow \mathcal{D}\), \[\label{eq:esem2} Know(y_1,y_2,\lessdot \psi(\boldsymbol{x})\gtrdot^\beta_\alpha)/g = Know(g^*(y_1),g^*(y_2), g^*(\lessdot \psi(\boldsymbol{x})\gtrdot^\beta_\alpha))\tag{7}\] with \(\{y_1,y_2\}\bigcup \beta \bigcup \alpha \subseteq \mathcal{V}\), such that \(g^*(y_1) = in ~ present\) and \(g^*(y_2) = \boldsymbol{I}\) (the robot itself), for the extended assignments \(g^*:\mathcal{T}\rightarrow \mathcal{D}\) (from Definition 17 in [2]), where the set of terms \(\mathcal{T}\) of \(IFOL_B\) is composed by the set \(\mathcal{V}\) of variables used in the set of predicates of \(IFOL_B\), by the set of constants and abstracted terms \(\lessdot \psi(\boldsymbol{x})\gtrdot^\beta_\alpha)\), so that in the actual world \(\hbar\), the known fact (7 ) for robot becomes
\(Know(y_1,y_2,\lessdot \psi(\boldsymbol{x})\gtrdot^\beta_\alpha)/g = Know(in ~ present,\boldsymbol{I}, I(\psi[\beta/g(\beta)]))\)
which is true in actual word, that is, from proposition (intensional concept)
\(u=I(Know(in ~ present,\boldsymbol{I}, I(\psi[\beta/g(\beta)])))\in D_0\), we obtain the truth value
\(\hbar(u) = \hbar(I(Know(in ~ present,\boldsymbol{I}, I(\psi[\beta/g(\beta)])))) = t\). Note that for the assignments \(g:\mathcal{V}\rightarrow \mathcal{D}\), such that \(g(y_1)= in ~future\) and \(g(y_2)\) we consider robot’s hypothetical knowledge in future, while in the cases when \(g(y_1) = in ~past\) we consider what was
robot’s knowledge in the past. So, the robots current knowledge (ground atoms of predicate \(Know\)) is directly derived from its experiences (based on its neuro-system processes that robot is using) in an analog way as
human brain does:
As an activation (under robot’s attention) of its neuro-system process, as a consequence of some human command to execute some particular job.
As an activation of some process under current attention of robot, which is part of some complex plan of robot’s activities connected with its general objectives and services.
Remark: We consider that only robot’s experiences (under robot’s attention) are transformed into the ground atoms of the many-valued \(Know\) predicate, and the required (by robot) deductions from them
(by using general many-valued deduction [1] extended by the three epistemic axioms) are transformed into ground atoms of \(Know\) predicate, and hence are saved in robot’s temporary memory as a part of robot’s conscience. Some background process (unconscious for the robot) would successively transform these temporary memory knowledge into
permanent robot’s knowledge as it happen for humans.
By such fixing by humans of robot’s unconciseness part with active semantics (which can not be modified by robots and their live experience) of all significant for human robot’s concepts and their properties, we will obtain ethically confident and socially
safe and non danger robots (controlled by public human ethical security organizations for the production of robots with general strong-AI capabilities).
\(\square\)
IFOL-based approach is part of a general neuro-symbolic AI paradigm, where researchers aim to blend: Neural methods (e.g., deep learning) for pattern recognition and learning from data, and symbolic logic for structured
knowledge representation and reasoning. This broader field (which includes work at IBM Research, MIT, and in academic surveys) seeks to build AI systems that can both learn from experience and reason abstractly. These ideas contribute to
ongoing discussions in AI about symbol grounding, logical inference, and self-referential reasoning — all important for Strong AI research.
Key aspects of this theory include:
Formalizing the "Self" (I):
We propose that for a robot, the ground term "I" (or "me", Self) acts as the name of its main working coordination program. This master program activates all other sub-programs, such as visual recognition or motor control, allowing the
robot to represent expressions like \(See(\boldsymbol{I}, blue color)\) in its logical framework.
Neuro-Symbolic Integration:
We advocate for a "dual-process" model inspired by Daniel Kahneman’s System 1 and System 2.
Intensional Abstraction and Self-Reference:
Using intensional abstraction, a robot can treat its own knowledge (propositions) as "individuals" within its logic. This allows the robot to perform self-referential reasoning—literally thinking about its own thoughts—which we identify as a prerequisite
for self-awareness.
Grounding through Experience:
We argue that robots must "ground" their language concepts by associating them with their own sensory-motor experiences. For example, the sense of the word "blue" isn’t just a label, but is grounded in the robot’s specific internal neural experience of
processing that color.
We addresses logical omniscience—the unrealistic assumption that an agent automatically knows all logical consequences of its beliefs—by using intensional abstraction to shift from "extensional" to "intensional" reasoning. This work is detailed in the book, Intensional First-Order Logic [2]: From AI to New SQL Big Data, which outlines these principles as a path toward a new generation of Strong-AI robots. In standard modal logics (like S5), knowledge is closed under logical implication, meaning if an agent knows \(A\), it must instantly know all \(B\) where \(A\rightarrow B\). We solve this by re-engineering how propositions are represented:
Reification of Propositions: Through intensional abstraction, predicates and sentences are "reified"—treated as individual objects (intensional entities) within the same domain as physical objects.
Decoupling Truth from Meaning: In this framework, two propositions can be extensionally equivalent (true in all the same possible worlds) but intensionally distinct. For example, a robot might know "Triangle A is equilateral" without yet knowing "Triangle A is equiangular," because those two concepts are different intensional "individuals" that require a specific computational inference step to link.
This approach allows for Strong-AI robots that can reason about their own knowledge (autoepistemic reasoning) as a finite, step-by-step process, mirroring human cognitive limitations rather than possessing infinite mathematical foresight.
In our framework, autoepistemic reasoning is the ability of a robot to reason about its own state of knowledge and belief as if they were objects in the world. Unlike standard AI, which might "know" a fact without "knowing that it knows," our Strong-AI Autoepistemic Robots use specialized logical structures to achieve formal self-reflection.
In this research, bridging the gap between neural networks (sub-symbolic) and symbolic logic is not just about making them work side-by-side; it is about creating a formal translation layer where the robot’s "internal feelings" (neuron firings) become "meaningful concepts" (logical terms). We achieve this via a Dual-System Architecture grounded in Intensional First-Order Logic:
Summary Table: The neuro-symbolic’s Bridge
| Sistem 1 (Neural) | The Bridge (Intensional) | System 2 (Symbolic) | |
|---|---|---|---|
| Function | Fast, reactive sensing | Abstraction \(\&\) Grounding | Slow, logical planning |
| Data Type | Vectors / Tensors | Intensional Entities | Logic Formulas |
| Role | Perceives the world | Maps neurons to symbols | Reasons about perceptions |
| Self-awareness | Unconscious | The "I" links the two | Conscious autoepistemic |
| thought |
In this paper we introduce more expressive 4-valued \(IFOL_B\), based on Belnap’s bilattice, by probabilistic features as well, in order to make neuro-symbolic learning and many-valued deductions (as humans) also in the
presence of unknown and inconsistent (contradictory) sentences [1], [5], because is such cases the
standard 2-valued FOL is unable to work well (deduces absolutely anything). In this paper we will extend this truth-many-valued neuro-symbolic framework also to the probabilistic learning inside this \(IFOL_B\) useful also
for the simulation, planing and verification of different strategies in order that the AGI-robot’s would be able to achieve their objectives and adaptive goal-directed behaviour under varying circumstances. With this extension, the robots will be able to
reason both with algebraic many-valued method (as epistemic logic) and also to manage uncertainty by probabilistic logic and symbol-guided common-sense obtained by deep neural learning in the same formal \(IFOL_B\) cognitive model. Such AGI-robots cognitive model is a general highly unified neuro-symbolic model.
It is passed twenty years from my proposed a fairly ambitious revision in that time of the earlier probabilistic logic programming (PLP) frameworks developed by V. S. Subrahmanian and collaborators [10], [11], where I tried to address exactly the kinds of their semantic weaknesses: problematic fixpoint semantics, interval ambiguity, mismatch between syntax and possible-world semantics, temporal inconsistency, and many-valued logical complications. In particular, I criticized implicit temporal semantics, unclear Herbrand interpretation structure, interval annotations without explicit world semantics, overly syntactic fixpoint constructions. I argued that time should be explicitly represented, probabilistic worlds should remain classical possible worlds, probability should be defined over sets of temporal Herbrand interpretations. This is actually a fairly deep semantic correction which effectively reduces temporal probabilistic logic to classical logic over expanded temporal predicates with classical model theory, accepted by the next developments of this revision, especially in today’s main-stream of distributive semantics for PLP (modern PLP research mostly evolved toward distribution semantics, weighted logic, probabilistic graphical models, and differentiable probabilistic inference).
Earlier systems often treated time as metadata, annotation, external modal structure, while in my revision I instead internalized time directly into Herbrand models, predicates and possible worlds, by restoring semantic clarity: each possible world becomes an ordinary temporal interpretation, and probability distributions are defined over classical logical models. Rather than relying primarily on interval probability propagation and many-valued fixpoint operators, this revision has been philosophically important: probability should remain measure-theoretic and logic should remain 2-valued at the world level, which aligns more closely with standard probability theory, modal semantics, Kripke semantics, and probabilistic model theory. I explicitly criticized the idea that probabilistic truth should be treated as ordinary many-valued logic, by adopting Nils Nilsson probability distributions over possible worlds, instead of embedding probabilities directly as many-valued truth values: The probabilities are not truth values themselves, they are properties of propositions across worlds. This direction is actually philosophically closer to Nilsson, Sato, and important researchers: Fabrizio Riguzzi, Luc De Raedt, Angelika Kimmig and Terrance Swift, with key systems ProbLog, PRISM and LPADs. A major semantic consolidation paper was by Riguzzi \(\&\) Swift on well-defined distribution semantics [12].
My revision reflected a deeper foundational issue: Is probability a generalized truth value, or a measure over classical worlds? This is a profound distinction, with the answer: worlds remain classical and probability is meta-level structure over worlds. That position is philosophically closer to Kolmogorov probability, modal realism and stochastic semantics, than, fir example, to fuzzy logic. This resembles modal logic, intensional semantics [13], [14] with possible-world semantics and Montague-style semantics. Here I will extend it [2] also to 4-valued (Belnap’s bilattice) logic for AGI robots.
The probability theory is a well-studied branch of mathematics, in order to carry out formal reasoning about probability. Thus, it is important to have a logic, both for computation of probabilities and for reasoning about probabilities, with a well-defined syntax and semantics. Both current approaches, based on Nilsson’s probability structures/logics, and on linear inequalities in order to reason about probabilities, have some weak points (Section 2.2 in [2]).
We will show that the logic for reasoning about probabilities can be naturally embedded into a 4-valued intensional FOL with intensional abstraction, by avoiding current ad-hoc system composed of two different 2-valued logics: one for the classical propositional logic at lower-level, and a new one at higher-level for probabilistic constraints with probabilistic variables.
The main motivation for an introduction of the intensionality in the probabilistic-theory of the propositional logic is based on the desire to have the full logical embedding of the probability into the First-Order Logic (FOL), with a clear difference from the classic concept of truth of the logic formulae and the concept of their probabilities. In this way we are able to replace the ad-hoc syntax and semantics, used in current practice for Probabilistic Logic Programs [10], [15], [16] and probabilistic deduction [17], by the standard syntax and semantics used for the FOL where the probabilistic-theory properties are expressed simply by the particular constraints on their interpretations and models.
In this section we will consider the probabilistic semantics for the propositional logic only (it can be easily extended to predicate logics as well) [18], [19] with a fixed finite set \(P = \{p_1,...,p_n\}\) of primitive propositions, which can be thought of as corresponding to basic probabilistic events. The set \(\mathcal{L}(P)\) of the propositional formulae is the closure of \(P\) under the Boolean operations for conjunction and negation, \(\wedge\) and \(\neg\), that is, it is the set of all formulae of the propositional logic \((P,\{\wedge, \neg\})\).
In order to give the probabilistic semantics to such formulae, we first need to review briefly the probability theory (see, for example, [20], [21]):
Definition 3. A probability space \((S,\mathcal{X}, \mu)\) consists of a set \(S\), called the sample space, a \(\sigma\)-algebra \(\mathcal{X}\) of subsets of \(S\) (i.e., a set of subsets of \(S\) containing \(S\) and closed under complementation and countable union, but not necessarily consisting of all subsets of \(S\)) whose elements are called measurable sets, and a probability measure \(\mu:\mathcal{X} \rightarrow [0,1]\) where \([0,1]\) is the closed interval of reals from 0 to 1. This mapping satisfies Kolmogorov axioms [22]:
A.1 \(~\mu(Y) \geq 0\) for all \(Y \in \mathcal{X}\).
A.2 \(~\mu(S) = 1\).
A.3 \(~\mu(\bigcup_{i\geq 1} Y_i) = \sum_{i\geq 1} \mu(Y_i)\),
\(~~\)if \(~Y_i\)’s are nonempty pairwise disjoint members of \(~\mathcal{X}\). We define a probability density function, \(KI=
\mu\circ in\), where \(in:S \hookrightarrow \mathcal{P}(S)\) is an inclusion such that \(in(s) = \{s\}\).
The \(\mu(\{s\}) = KI(s)\) is the value of probability in a single point of space \(s\).
The property A.3 is called countable additivity for the probabilities in a space \(S\). In the case when \(~\mathcal{X}\) is finite set, then we can simplify property A.3 above to
A.3’ \(~\mu(Z \bigcup Y) = \mu(Z) + \mu(Y)\),
if \(Z\) and \(Y\) are disjoint members of \(~\mathcal{X}\), or, equivalently, to the following axiom:
A.3” \(~\mu(Z ) = \mu(Z\bigcap Y) + \mu(Z \bigcap \overline{Y})\),
where \(\overline{Y}\) is the compliment of \(Y\) in \(S\), so that \(\mu(
\overline{Y}) = 1 - \mu(Y)\).
In what follows we will consider only finite sample space \(S\), so that \(~\mathcal{X} = \mathcal{P}(S)\). Thus, in our case of a finite set \(S\) we obtain, form A.1 and A.2, that for any \(Y \in \mathcal{P}(S)\), \[\label{eq:muKI} ~\mu(Y) = \sum_{s \in Y}\mu(\{s\}) = \sum_{s \in Y} KI(s)\tag{8}\] Based on the work of Nilsson in [18] we can define for a given propositional logic with a finite set of primitive propositions \(P\) the sample space \(S = \boldsymbol{2}^{P}\), where \[\label{eq:muKIbb} \boldsymbol{2} = \{f,t\} \subset X = \mathcal{B}_4 = \{f,t,\bot,\top \}\tag{9}\] is the set of truth values of standard 2-valued logic, so that the probability space is equal to the Nilsson’s structure \(N = (\boldsymbol{2}^{P}, \mathcal{P}(\boldsymbol{2}^{P}), \mu)\).
In his work (page 2, line 4-6 in [18]) Nilsson considered a Probabilistic Logic "in which the truth values of sentences can range between \(0\) and \(1\). The truth value of a sentence in probabilistic logic is taken to be the probability of that sentence in ordinary first-order logic." That is, he considered this logic as a kind of a many-valued (shown in Section 2.2.1 in [2]), but not a compositional truth-valued, logic. But in his paper he did not defined the formal syntax and semantics for such a probabilistic logic, but only the matrix equations where the probability of a sentence \(\phi \in \mathcal{L}(P)\) is the sum of the probabilities of the sets of possible worlds (equal to the set \(S = \boldsymbol{2}^{P}\)) in which that sentence is true. So that he assigns two different logic values to each sentence \(\phi\): one is its probability value and another is a classic 2-valued truth value in a given possible world \(v \in \mathcal{W}= S = \boldsymbol{2}^{P}\).
The logic inadequacy of this seminal work [18] of Nilsson is also considered in [23], by extending this Nilsson’s structure into a more general probability structure \(M = (\boldsymbol{2}^{P}, \mathcal{P}(\boldsymbol{2}^{P}), \mu, \pi)\), where \(\pi\) associates with each \(s \in S = \boldsymbol{2}^{P}\) the truth assignment \(\pi(s):P \rightarrow \boldsymbol{2}\). However, in our case when \(S = \boldsymbol{2}^{P}\), \(\pi\) is just an identity, so not necessary, and we consider each \(s\) as a truth valuation \(s=v:P\rightarrow \{f,t\}\) which can be uniquely extended to the truth assignment \(v^*\) to all formulae in \(\mathcal{L}(P)\), by taking the usual rules of propositional logic (the unique homomorphic extension to all formulae), and we can associate to each propositional formula \(\phi \in \mathcal{L}(P)\) the set \(\phi^M\) consisting of all states \(s \in S\) where the sentence \(\phi\) is true, so that \[\label{eq:muKIaa} ~\|\phi\| = \{v \in \mathcal{W}= \boldsymbol{2}^{P} ~|~ v^*(\phi) =t\}~~\tag{10}\] But, differently from Nilsson, in [23] the authors did not define a many-valued propositional logic, but a kind of 2-valued logic based on probabilistic constraints. They denoted by \(w_N(\phi)\) the weight or probability of \(\phi\) in Nilsson structure \(N\), correspondent to the value \(\mu(\|\phi\|)\), so that the basic probabilistic 2-valued constraint can be defined by expressions \(c_1 \leq w_N(\phi)\) and \(w_N(\phi) \leq c_2\) for given constants \(c_1,c_2 \in [0,1]\). They expected their logic to be used for reasoning about probabilities. But, again, they did not defined a unique logic, but two different logics: one for the classical propositional logic \(\mathcal{L}(P)\), and a new one for 2-valued probabilistic constraints obtained from the basic probabilistic formulae above and Boolean operators \(\wedge\) and \(\neg\). They did not consider the introduced symbol \(w_N\) as a formal functional symbol for a mapping \(w_N:\mathcal{L}(P) \rightarrow [0,1]\), such that for any propositional formula \(\phi \in \mathcal{L}(P)\), with \(\mathcal{W}= S = \boldsymbol{2}^P\), the probability to be true of sentence \(\phi\) is
\(w_N(\phi) = \mu(\|\phi\|) = \sum_{v\in \|\phi\|} KI(v)\).
Instead of this intuitive meaning for \(w_N\) they considered each expression \(w_N(\phi)\) as a particular probabilistic term (more precisely, as a structured probabilistic
variable over the domain of values in \([0,1]\)).
It seams that such a dichotomy and difficulty to have a unique 2-valued probabilistic logic, both for an original propositional formulae in \(\mathcal{L}(P)\) and for the probabilistic constraints, is based on the fact that if we consider \(w_N\) as a function with one argument then it has to be formally represented as a binary predicate \(w_N(\phi,a)\) (for the graph of this function) where the first argument is a formula and the second is its resulting probability value. Consequently, a constraint "the probability of \(\phi\) to be true is less or equal to \(c\)", has to be formally expressed by the logic formula \(w_N(\phi,a) \wedge \leq (a,c)\) (here we use a symbol \(\leq\) as a built-in rigid binary predicate where \(\leq (a,c)\) is equivalent to \(a \leq c\)), which is a second-order syntax because \(\phi\) is a logic formula in such a unified logic language. That is, the problem of obtaining the unique logical framework for probabilistic logic comes out with the necessity of a reification feature of this logic language, analogously to the case of the intensional semantics for RDF data structures [2], [24].
Consequently, we need a logic [13] which is able to deal directly with reification of logic formulae, that transforms a propositional formulae \(\phi \in \mathcal{L}(P)\) into an abstracted term, denoted by \(\lessdot\phi \gtrdot\). By this approach the expression \(w_N(\lessdot\phi \gtrdot,a) \wedge
\leq (a,c)\) remains to be an ordinary first-order formula. In fact, if \(\lessdot\phi
\gtrdot\) is translated into non-sentence "that \(\phi\)", then the first-order formula above corresponds to the sentence "the probability that \(\phi\) is true is less
than or equal to \(c\)".
Such approach has been used [2], [25] for the reduction of temporal probabilistic databases into constraint logic programs and to apply the interval PSAT in order to find the models of such interval-based probabilistic programs. In our case, for selfindependent neuro-symbolic AGI robots, we do not need to write the probabilistic programs for them, but only to provide a general self-reasoning about the probabilities of their 4-valued sentences. So, in next Section we will show how the probabilistic reasoning can be embedded in the 4-valued \(IFOL_B\) based on Belnap’s bilattice of truth-values in \(X = \mathcal{B}_4 = \{f,t,\bot,\top\}\).
There are numerous proposals for probabilistic logics. Very roughly, they can be categorized into two different classes: those logics that attempt to make a probabilistic extension to logical entailment, such as Markov logic networks, and those that attempt to address the problems of uncertainty and lack of evidence (evidentiary logics).
For AGI robots we considered the \(IFOL_B\) with 4-valued Belnap’s bilattice of truth-values in \(X = \mathcal{B}_4 = \{f,t,\bot,\top\}\) with knowledge ordering as well [5], where the value "unknown" is the bottom value, the sentences with this value are indeed unknown facts, that is, the missed knowledge in the AGI robots.
Thus, these unknown facts are not part of the robot’s knowledge database, and by learning through input and experiences, the robot’s knowledge would be naturally expanded over time. Consequently, this phenomena has been represented by the Closed Knowledge Assumption and Logic Inference [1], here for unknown facts, because of lack of evidence inside non-probabilistic \(IFOL_B\) and its autoepistemic deductive system enable to establish some of the known logic states in \(\{t,f,\top\} \subset X\) of the facts (to be true, false or inconsistent) the robot has to assign to such facts the "absolute uncertainty", that is the truth-value \(\bot\) (unknown logic value). So, the minimization of this uncertainty can be obtained by introducing the probabilistic computation (by using Nilsson’s structures with Kolmogorov’s axioms) for these unknown facts, to obtain what is the probability of a currently unknown fact to be true, false or inconsistent.
We recall that the extensionalization functions \(h \in \mathcal{E}_{in}\) in \(IFOL_B\), are given in the disjoint union from (3 ), \[h = \sum_{i\in \mathbb{N}}h_i:\mathcal{D}\longrightarrow D_0+\sum_{i\geq 1}\mathfrak{Rm}_i\] Thus, the intensions can be seen as names of abstract or concrete entities, while the extensions correspond to various rules
that these entities play in different worlds we use the bijective mapping (6 ), \(is_{in}:\mathcal{W}\rightarrow \mathcal{E}_{in}\), where \(\mathcal{E}_{in}\) is the set of all extensionalization functions that respect all built-in predicates, and the "set of possible worlds" \(\mathcal{W}= \{h\circ I | h \in
\mathcal{E}_{in}\}\).
So, from the commutative truth-diagram for the set \(\mathcal{L}_0\) of all sentences of this logic \(\mathcal{L}_{in}\), provided by Theorem 1 in [5], we obtain that \(I^*_{B} = h\circ I\), so that from Definition 15 in [5], given an assignment \(g:\mathcal{V}\rightarrow \mathcal{D}\), for any sentence \(\phi/g\in \mathcal{L}_0\), \[\label{eq:mvR} (h\circ I)(\phi/g) =I^*_{B}(\phi/g) = \left\{ \begin{array}{ll} \{v^*(\phi/g)\} \in \mathfrak{Rm}, & if ~~v^*(\phi/g) \neq \bot\\ \emptyset, & otherwise \end{array} \right.\tag{11}\] Consequently,
we are able to represent the whole AGI robot’s knowledge Database of its n-ary concepts, \(n \geq 1\), in each fixed instance of time (it is an analog to traditional relational Database with standard 2-valued FOL)
as follows:
Definition 4. AGI Robot’s Current Atomic Knowledge Database3:
For a given instance of time, the current AGI robot knowledge is defined by its current MV-interpretation \(\boldsymbol{I}^*_{B}\) (that represents robot’s current world \(w \in
\mathcal{W}\)) defines the extension of robot’s knowledge Database \(\mathcal{K}\) by: \[\label{eq:two-know}
~~ \mathcal{K}~= ~ \{R_{p_i^k} = \boldsymbol{I}^*_{B}(p_i^k(x_1,...,x_k) ~|~p_i^k \in P, ~for~ x_i \in \mathcal{V}, 1\leq i\leq k\}\qquad{(3)}\] where \(\boldsymbol{I}^*_{B}\) is the current MV-interpretation
(a function from \(\mathcal{L}\) to \(\mathfrak{Rm}\)) and \(\mathcal{V}\) the set of variables and \(R_{p_i^k}\) is the
\((k+1)\)-ary relation obtained for the k-ary predicate \(p_i^k\).
Consequently, the atomic knowledge database is the subset of current metaknowledge of a robot is the set of relations, that is, \[\label{eq:two-metaknow} \mathcal{K}\subset~~ ~ \{R ~|~R \in Im(\boldsymbol{I}^*_{B}), ~for~ ar(R) \geq 2\}\tag{12}\] where for each tuple of ground terms, \(\boldsymbol{d} = (t_1,...,t_k,a) \in R_{p_i^k} \in \mathcal{K}\), of relation with arity \(k+1\geq 2\), the current truth-value of this, for robot known, fact is equal to last value of this tuple, \(a =\pi_{k+1}(\boldsymbol{d}) \in \{f, \top, t\}\) with truth-ordering \(f < \top < t\). So, for any ground instance of each robot’s (real or virtual) predicate (and corresponding intensional concept), robot is able to know the level of truth in a given instance of time. Thus, given current atomic knowledge database \(\mathcal{K}\), we are able to derive from it the current Herbrand model \(v:H \rightarrow X\) and its extension \(v^*\) to all sentences, for robot’s knowledge by, for any ground atom \(p_i^k(t_1,...,t_k)\in H\), with the ground terms \(t_i\), for \(1\leq i\leq k\), \[\label{eq:two-metaknowH} v(p_i^k(t_1,...,t_k)) = \left\{ \begin{array}{ll} a\in \{f, \top, t\}, & if ~~(t_1,...,t_k,a) \in R_{p_i^k} \in \mathcal{K}\\ \bot, & otherwise \end{array} \right.\tag{13}\] Thus, from current atomic knowledge database \(\mathcal{K}\), we are able to derive current Herbrand model \(v\), its extension to all sentences \(v^*\) and thence the current possible world (MV-model in Definition 4) \(\boldsymbol{I}^*_B = h\circ I \in \mathcal{W}\).
It was demonstrated that the knowledge database \(\mathcal{K}\) satisfies the following assumption4:
Definition 5. Closed Knowledge Assumption(CKA):
The CKA for the many-sorted Intensional FOL based on Belnap’s 4-valued bilattice of truth-values is defined, for every Herbrand interpretation \(v:H \rightarrow X\) and assignment \(g:\mathcal{V}\rightarrow \mathcal{D}\), as follows:
For each k-ary virtual predicate \(\phi(x_1,...,x_k)\), \(k \geq 1\),
\(v^*(\phi(x_1,...,x_k)/g) \neq \bot~~\) iff \(~~(g(x_1),...,g(x_k))\in \pi_{-i}(I_B^*(\phi(x_1,...,x_k)))\).
The intensional abstract terms are "that-clauses" used for reification features [26] of \(IFOL_B\) so that, for a logic formula \(\phi(\boldsymbol{x})\) and assignment \(g\in \mathcal{D}^\mathcal{V}\) of variables in \(\mathcal{V}\), "that \(\phi\)", from Definition 1 is denoted by the ground abstracted term \(\lessdot \phi(\boldsymbol{x})[\beta/g(\beta)] \gtrdot_\alpha\) (if \(\phi\) is a sentence then both \(\alpha\) and \(\beta\) are empty). Hence, the sentence "the probability that a sentence \(\phi\) has a truth-value \(a\) is less then or equal to \(c_1\)" can be expressed by the first-order logic ground atom \(w_N(\lessdot \phi \gtrdot,a,c) \wedge \leq(c,c_1)\), where \(\leq\) is the binary built-in predicate with standard denotation \(c\leq c_1\), while "the probability that \(\phi\) has a truth-value \(a\) is equal to \(c\)" is denoted by the ground atom \(w_N(\lessdot \phi \gtrdot,a,c)\) with the special ternary predicate \(w_N\) which first argument is from an sentence abstracted term, the second argument is for the truth-value of this sentence and third argument is the probability that it is so. So, we introduce the following built-in concepts (for Definition 9 in [5]) of typed \(IFOL_B\):
"truth values:\(~s_X\)", introduced in [5], that is a concepts in \(D_2\) such that for \(X \subset D_0\), its extension is \(h_2(\boldsymbol{truth values}:~s_X) = \{(a,t)~|~a,t \in X\}\). We will use the special typed variable \(x_X\) for the set of truth values in \(X = \mathcal{B}_4 = \{f,t,\bot,\top\}\), so that for each assignment \(g\), \(g(x_X) \in X\).
"probability:\(~s_p\)", that is a concepts in \(D_2\) such that its extension is \(h_2(\boldsymbol{probability}:~s_p) = [0,1]\times \{t\}\), where \([0,1]\) is interval of positive reals used for the probabilities and \(t\in X\) the true logic value. We will use the special typed variable \(x_p\) for the probabilities, so that for each assignment \(g\), \(g(x_p) \in [0,1]\).
for each finite ordered calendar interval of time \([\boldsymbol{t}_I,\boldsymbol{t}_F]_\tau\) with initial \(\boldsymbol{t}_I\) time-instance and final \(\boldsymbol{t}_F\) time-instance, of sorts \(s_\tau\)) with a given granularity \(\tau \in \{year, year:month, year:month:day, year:month:day:our,...\}\), we can
have a particular calendar-concept in \(D_2\), "calendar: \(s_\tau\)", such that its finite extension is \(h_2(\boldsymbol{calendar}:~s_\tau) =
[\boldsymbol{t}_I,\boldsymbol{t}_F]_\tau \times \{t\}\) with \(t\in X\) the true logic value. We will use the special typed variables \(x_\tau\) for the time-instances of these
calendars, so that for each assignment \(g\), \(\boldsymbol{t} = g(x_\tau) \in [\boldsymbol{t}_I,\boldsymbol{t}_F]_\tau\).
Each predicate \(p_i \in P\) that has exactly one attribute of sort \(s_\tau\) is called an "atomic event" predicate.
Thus, in this version of \(IFOL_B\) we will have the following three meta-predicates (two-valued predicates that represent’s the semantically very specific knowledge about properties of standard domain-based predicates in \(P\) (the predicates in \(P\) have no the typed variable \(x_s\) of sort ‘nested sentence’, provided in Definition of static sorts in [27]):
4-valued Knowledge predicate \(Know\) used for autoepistemic deduction [1].
2-valued (true or false) Probabilistic predicate \(w_N\) with the first argument has the sort of ‘nested sentence’ and with free-variable atom \(w_N(\lessdot \psi(\boldsymbol{x})\gtrdot^\beta, x_X, x_p)\) where the formula \(\psi\) is composed by predicates in \(P\). If \(\psi\) is composed also by "atomic event" predicates, then \(w_N(\lessdot \psi(\boldsymbol{x})\gtrdot^\beta, x_X, x_p)\) is called Event probabilistic predicate".
so that \(P \bigcap \{Know, w_N, \boldsymbol{e}\}\) is empty set. This separation of standard and meta-predicates is based on the fact that the truth-value of meta-predicate ground atoms depends on the truth-values of
the another sentence used in the abstract term of the typed variable \(x_s\). In effect, from Definition 14 in [1],
we have that the computation of the truth-value of the 4-valued \(Know\) ground atoms is done as this atom is a particular logic sentence (formula) and not ground atom of Herbrand base:
\[\begin{gather} \label{eq:esem3s} v^*(Know(t_1,t_2,\lessdot \psi(\boldsymbol{x})\gtrdot^\beta)/g ) = v^*(\psi(\boldsymbol{x})/g)\in X, ~~~~and~ for ~~I_B^* = is_{MV}(v^*)\\
I_B^*(Know(t_1,t_2,\lessdot \psi(\boldsymbol{x})\gtrdot^\beta)/g ) = I_B^*(\psi(\boldsymbol{x})/g)~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
\end{gather}\tag{14}\] We have shown by Lemma 1 in [1] that, during time-evolution for time-instances \(\boldsymbol{t}_i \leq \boldsymbol{t}_{i+1}\) of robot’s knowledge, it is always satisfied that \[\label{eq:knowledge} v^*_i\preccurlyeq_k
v^*_{i+1}\tag{15}\] that is, we have a monotonic increments of knowledge, initially most (non built-in predicate) ground atoms are unknown, and that during learning processes in future, the (non built-in predicate) ground atoms can change
their truth-value in this knowledge-monotonic way:
1. \(\bot \mapsto f\) of \(\bot \mapsto t\);
2. \(f \mapsto \top\);
3. \(t \mapsto \top\);
4. \(\top \mapsto \top\).
Thus, for the Herbrand base (from Definition 4 in [1])
\(~~H = \{p_i^k(t_1,..,t_k)~|~p_i^k \in P\) and \(t_1,...,t_k\) are ground terms \(\}\)
and the current Herbrand interpretation \(\boldsymbol{v}:H\rightarrow X\), for still unknown atoms \(A\in H\) with \(\boldsymbol{v}(A) = \bot\)
(note that such atoms can note be of built-in predicates that have invariant truth-value true or false) we can use the probabilistic methods to establish with which probability they can have any truth-value in \(X\), and to
use this knowledge for another logical deductions.
Consequently, in order that AGI robots be able to reason about probabilities of the sentences using the syntax of the FOL with the set of predicate symbols \(P\) of current robot’s knowledge (robot would be able to increment this set with new predicates as well, based on its extended in time knowledge about external world) we have the sample space of Definition 3, from (2 ), \(S \subseteq \mathcal{I}_H \subset X^{H}\) where \(H\) is Herbrand base of this logic \(IFOL_B\) with Belnap’s bilattice of truth-values in \(X\). And we impose the following constraints for the probability density \(KI = \mu\circ in:S \rightarrow [0,1]\):
Definition 6. Probability Density Constraints:
For any current Herbrand interpretation \(\boldsymbol{v}:H\rightarrow X\), the robot’s current knowledge database \(\mathcal{K}\) of \(IFOL_B\) (from
Definition 4 ) must remain invariant* by introducing the probability space \((S,\mathcal{P}(S),\mu)\), in Definition 3.
That is, for every atom \(A \in H\) such that \(u = \boldsymbol{v}(A) \neq \bot\), the probability that \(A\) has the truth-value \(u\) must be equal to 1, i.e., \[\label{eq:KIconstraint}
\forall A\in H.(~if ~~u = \boldsymbol{v}(A) \neq \bot ~~~~then ~~\sum_{v\in S, v(A) = u} KI(v) = 1)\tag{16}\] *
Thus, the action of introducing the probability computation in \(IFOL_B\) will have the effects only to diminish the uncertainty of unknown facts, for which we can now compute the probability to have any
truth-value in \(X\).
Remark: This is a particular case of Noether’s conservation theorem applied to robot’s cognitive system: the (global) symmetry transformation of robot’s cognitive power by action of introduction of the probability space \((S,\mathcal{P}(S),\mu)\) preserves its knowledge database \(\mathcal{K}\).
This symmetry is global because is valid for all current Herbrand base (that is to all predicates in \(P\) of robot’s knowledge base: so computationally it has relevant cost to be applied in real-time robot’s
processes. However, we are able to consider less computationally expensive symmetry transformation, that can be done dynamically by robot to resolve some decisions about concrete problems that involve a very restricted subset of it predicates, relevant to
such problems: such dynamic symmetries will be called a local symmetries in what follows.
\(\square\)
So, in what follows we will extend the globally consistent and empirically satisfactory unification of classic probability theory and standard (two-valued) first-order logic that is suitable for inductive reasoning, developed by Gaifman and Snir [28], [29] , to our more sophisticated Belnap’s based 4-valued typed \(IFOL_B\) used for humanoid AGI robots.
Definition 7. Global Current Nilsson’s Structure for \(IFOL_B\):
Let \(\boldsymbol{v}:H\rightarrow X\) be the current Herbrand interpretation of AGI robot with Belnap’s based 4-valued typed \(IFOL_B\), and corresponding current possible world \(\boldsymbol{I}_B^* = is_H(\boldsymbol{v}) \in \mathcal{W}\). Then we define the current Nilsson’s probability structure \((S,\mathcal{P}(S),\mu)\) satisfying Kolmogorov axioms in Definition 3, such that the sample set \(S\) is defined as a subset of \(\mathcal{I}_H\) in (2 ), by \[\label{eq:IFOLpstructure}
S =_{def} \{v \in \mathcal{I}_H \subset X^H | ~~for~ every ~~A \in H, v(A) = \boldsymbol{v}(A) ~~if~ \boldsymbol{v}(A) \neq \bot\}\qquad{(4)}\] so that for any \(v\in S\), and ground atom \(A\in H\) such that \(\boldsymbol{v}(A) = \bot\), we can have that \(v(A)\in X\) can have the value different from \(\bot\) as
well.
Note that this sample set \(S\) in (?? ) is defined only for current possible world (and current knowledge base \(\mathcal{K}\), and does not modify the logic truth-value of ground atoms
\(A\in H\) for which \(\boldsymbol{v}(A) = \bot\) or currently unknown sentences \(\phi\) with \(\boldsymbol{v}^*(\phi) =
\bot\), but only offers the capacity to compute their probability to have any truth value in \(X\).
This global current Nilsson’s probability structure has the following important properties:
Corollary 1. The global current Nilsson’s probability structure with sample set \(S\) in (?? ) satisfy the probability density constraints in Definition 6.
Thus, it is consistent with current robot’s knowledge database \(\mathcal{K}\) (in Definition 4), consistent with Kolmogorow probability
axioms and logical deduction, and allows inductive reasoning and confirmation of universally quantified hypotheses.
Proof: It is enough to show the consistency with current robot’s knowledge database \(\mathcal{K}\). In fact, for each ground atom \(A \in H\) for which the current truth-value in \(\mathcal{K}\) is \(u = \boldsymbol{v}(A) \neq \bot\), we have that the probability density constraint (16 ) is satisfied. That is, the probability \(p_u\) that this atom has the truth-value \(u\) is equal to 1:
\(p_u = \sum_{v\in S, v(A) = u} KI(v)\)
\(= \sum_{v\in S} KI(v)\) from (?? )
\(= \mu(S)\)
\(= 1\).
which preserves robot’s current knowledge database and all logical deductions from them.
\(\square\)
Consequently, for new predicate \(w_N\) (not in \(P\)), we obtain that for any assignment \(g\) and formula \(\psi(\boldsymbol{x})\) composed by predicates (only) in \(P\), if, for a given assignment \(g\), in current world \(\boldsymbol{v}^*(\psi(\boldsymbol{x})/g) = \bot\), then we can compute in this current world (current knowledge database \(\mathcal{K}\)) \[\begin{gather} \label{eq:esem3s2} \boldsymbol{v}^*(w_N(\lessdot \psi(\boldsymbol{x})\gtrdot^\beta/g, g(x_X),g(x_p) )) =\\ =\left\{ \begin{array}{ll} t, & if ~~ ~g(x_p) = \sum_{v\in S, v^*(\psi(\boldsymbol{x})/g) =g(x_X)} KI(v) \\
f, & otherwise \end{array} \right. ~~~~~~~~~~~~~~~~~~~~~
\end{gather}\tag{17}\] That is,
The \(w_N(\lessdot \psi(\boldsymbol{x})\gtrdot^\beta/g, g(x_X),g(x_p))\) is true iff the probability that \(\psi(\boldsymbol{x})/g\) has the truth-value \(g(x_X)\) is equal to \(g(x_p)\).
In the case when the formula \(\psi\) contains also "atomic event" predicates, the event \(w_N(\lessdot \psi(\boldsymbol{x})\gtrdot^\beta/g, g(x_X),g(x_p))\), with \(\tau \in \beta\) is a free variable in \(\psi\), is true iff the probability that \(\psi(\boldsymbol{x})/g\) has the truth-value \(g(x_X)\) at time-instance \(g(x_\tau)\) is equal to \(g(x_p)\). Note that from the fact that \(x_\tau\) is typed variable, the instance of time \(g(x_\tau)\) is always inside the finite interval of granular calendar specified by this type \(\tau\). Obviously, \(\psi\) can be composed by a number of different "atomic events" as well, each of them with particular granular calendars time-variables \(x_\tau\) and their intervals.
Remark: Thus, for the current possible world of robot (its current knowledge database \(\mathcal{K}\)), for any unknown fact (which is not in \(\mathcal{K}\)), we can
have additional knowledge of what is the probability that this fact is true, false or inconsistent. In this way, by probabilistic computation based on Nilsson’s probability stricture in Definition 7 we diminish generally the robot’s knowledge uncertainty by preserving (from Corollary 1) its consistency and logical
deduction.
That is, for any sentence \(\phi \in \mathcal{L}_0\) such that in current Herbrand interpretation \(\boldsymbol{v}:H\rightarrow X\), \(\boldsymbol{v}^*(\phi) =
\bot\), robot would be able to compute the probability \(p_u\) that this sentence has the truth value \(u \in X\), \[\label{eq:prob}
p_u = \sum_{v\in S, v^*(\phi) = u} KI(v)\tag{18}\] which the robot can use for probability-based decisions and actions, by using the reasoning about the probabilities (as in Section 2.2.2 in [2]) by using the true atoms \(w_N(\lessdot \phi\gtrdot, u, p_u)\) with also "atomic events" that tells to robot that "the probability at time-instance \(\boldsymbol{t}_i\) that \(\phi\) has the truth-value \(u\) is equal to \(p_u\)".
These true facts about the probabilities of the unknown sentences to have a particular truth-value can be also inserted in the conscious part of robot’s cognitive system (see (7 ) for more details) by,
\(Know(in ~presence, \boldsymbol{I}, w_N(\lessdot \phi\gtrdot, u,p_u))\),
and hence to be used by robot’s autoepistemic deductions and explanations as well.
\(\square\)
Example 1. For example, let us consider an "atomic event" predicate \(p_i\), with its atom of free variables \(p_i(...,x_\tau,...)\), such that there exists the
time-instance \(\boldsymbol{t} \in [\boldsymbol{t}_I,\boldsymbol{t}_F]_\tau\) and assignment \(g\) such that \(g(x_\tau) = \boldsymbol{t}\), and in current
possible world \(\boldsymbol{v}(p_i(...,x_\tau,...)/g) = \boldsymbol{v}(p_i(...,\boldsymbol{t},...)/g) = t\), which means that this event of ground atom \(p_i(...,\boldsymbol{t},...)/g\)
will surely happen, and so we have also that \(\boldsymbol{v}^*((\exists x_\tau.p_i(...,x_\tau,...))/g) = t\).
Now, suppose that in this current world it is unknown if this event will happen, that is we have that \(\boldsymbol{v}^*((\exists x_\tau.p_i(...,x_\tau,...))/g) = \bot\) (it can happen only iff for each \(\boldsymbol{t} \in [\boldsymbol{t}_I,\boldsymbol{t}_F]_\tau\), \(\boldsymbol{v}(p_i(...,\boldsymbol{t},...)/g) = \bot\)). So, in this currently completely uncertain event-state, a robot would be
able able to compute the probability that it will happen, by using (17 ), so that for \(x_\tau \notin \beta\) (\(\beta\) is the set of only free variables),
\[\begin{gather} \label{eq:esem3s5} \boldsymbol{v}^*(w_N(\lessdot \exists x_\tau.p_i(...,x_\tau,...))\gtrdot^\beta/g, t,g(x_p) )) =\\ =\left\{ \begin{array}{ll} t, & if ~~ ~g(x_p) = \sum_{v\in S,
v^*(\exists x_\tau.p_i(...,x_\tau,...))/g = t} KI(v) \\ f, & otherwise \end{array} \right. ~~~~~~~~~~~~~~~~~~~~~
\end{gather}\qquad{(5)}\] that is, \(g(x_p)\) is the probability that this event will happen.
\(\square\)
In this way, by using the current Nilsson’s probability structure, we obtain the logical system of robots that are probabilistically-open (are not constrained by the pure-logical Closed Knowledge Assumption [1]) by minimizing the uncertainty of unknown sentences in most prudent way.
This is obtained by the this neuro-symbolic AGI of robots based on symbolic \(IFOL_B\) and this conservative probabilistic extension. The neural component of robot’s AGI cognitive system has to be used for computation of the probability density function \(KI: S \rightarrow [0,1]\) based on the principle of maximum information entropy. The concept of information entropy was introduced by Claude Shannon [30] and is also referred to as Shannon entropy or Information entropy , which is a mathematical measure of the average uncertainty, randomness, or "surprise" inherent in a set of data or events. Higher entropy means the outcome is less predictable, while lower entropy indicates a more certain and predictable outcome. In neuroscience and physics this concept has been adapted to model the flow of information in neural pathways and shares deep mathematical roots with thermodynamic entropy in physics.
The principle of maximum entropy states that, among all probability distributions consistent with a given set of constraints (16 ) in Definition 6, the distribution that maximizes Shannon entropy should be selected. This yields the least committal distribution (most prudent minimization of the uncertainty) compatible with the known constraints,
introducing no structure beyond what is logically implied by the available information, which corresponds to the following general principle in physics:
The Principle of Least Action (committal distribution) KI:
The least committal distribution is the probability distribution KI that maximizes information entropy subject to whatever constraints you currently know about the data. In statistics and information theory, this is formally known as the Principle
of Maximum Entropy. This principle minimizes the absolute uncertainty of any unknown fact in current internal model \(\boldsymbol{v}:H\rightarrow X\), by giving the information about the probability that these
facts are true, false or inconsistent. Thee sense of this Principle of Least Action is that it does not modify the knowledge database \(\mathcal{K}\) (that is, does not modify the robot’s internal model \(\boldsymbol{v}\)).
We can consider it as the free energy principle (of Karl Friston’s Active Inference) which is a mathematical principle of information physics. Its application to robot’s "brain" reduces surprise or uncertainty by making predictions based on
internal models and uses "sensory input" (in our case the committal distribution KI) to extend its models so as to improve the accuracy of its predictions (in our case of unknown, and hence totally uncertain, sentences). In effect, with this Principle of
Least Action, we apply a particular case of free energy principle of Karl Friston, which he used in Bayesian approaches to human brain function and some approaches to artificial intelligence (introduced it as an explanation for embodied perception-action
loops in neuroscience [31]).
More information is provided at the end, in Appendix, Section 6.
\(\square\)
The justification is that entropy measures the expected information content (or log-surprise) of outcomes relative to a specified reference measure. Maximizing entropy ensures that no additional structure is imposed beyond the stated constraints. Principle
of maximum entropy may be taken to compute degrees of belief of formulae [32], and it is shown in [33] for the consistent probabilistic inference. This method applied to probabilistic logic programming [11], [34], based on conditional probabilistic clauses, has shown that reduces the original entropy maximization to relatively small optimization problems, This entropy of \(KI\), for current Nilsson’s structure in Definition 7, is defined by, \[\label{eq:KIentropy}
H(KI) = - \sum_{v\in S}KI(v) \cdot logKI(v)\tag{19}\] We can use robot’s neural networks to calculate maximum-entropy probability distributions. Instead of solving complex mathematical equations, we use a dedicated robot’s neural network as a
numerical optimizer. The network learns to generate the correct probabilities by adjusting its weights to maximize entropy while respecting data constraints. This approach is particularly useful when dealing with variables that have many possible states
(high dimensionality). The standard architecture for this task involves a neural network in which the last layer applies a function called Softmax:
Softmax Output: The Softmax function takes the numbers generated by the network and transforms them into a set of probabilities. This automatically solves the normalization constraint.
Constraints: they are directly inserted into the rule with which the network learns.
Loss Function: To train the network, we don’t use classic classification errors. Instead, we create a custom loss function based on Lagrange multipliers.
Moreover, the scoring function can be also approximated by a nonlinear architecture, namely, a feedforward neural network (FFNN). Also know as deep neural network, artificial neural network or multilayer perceptron, it is a simple model where several linear combinations of inputs are passed through nonlinear activation functions called nodes. A set of nodes is called a layer, and the output of ones layer’s node becomes the input of the next layer in multilayer architectures. Finally, the output of the last layer is combined linearly to produce a cost-to-go [35] and [36] (Section 4.2 about Multilayer and Deep Neural Networks).
In this section we will considered the dynamic real-time probabilistic actions, baaed on local symmetry transformations of robot’s cognitive power, to resolve, as human do, the probabilistic decision on very focused (non-global) problems in real-time. Instead of time-expensive global symmetry transformation, described in previous section, that has the computational consequences over all predicates of the whole robot’s current knowledge, each practical problem to resolve does not require the involvement of all predicates (knowledge), but of only a very restricted subset of these predicates relevant to describe this problem by a particular sentence \(\phi\). This is a robot’s analogy with human brain localized neural activations (activating the "focused attention") to resolve different cognitive functions, depending on the current cognitive problem by using the principle of least robot’s knowledge action (AGI analog to human brain free energy principle, Active Inference of Karl Friston).
So, decision process can be also divided in a number of subproblems, each of which described in relatively small number of predicates relevant (by focusing attention) to this subproblem. In each of these subproblems our probability Nillson’s space \(S\) will be enormously smaller w.r.t. the global space provided by Definition 7, and hence the neural networks would be able to calculate the
maximum-entropy probability distribution for such smaller probability space in real-time, providing the dynamic robot’s probabilistic decisions about these subproblems.
Thus, we introduce this efficient local symmetry transformations, based on a (sub)problem-sentence \(\phi\) describing this (sub)problem.
Definition 8. Local Current Nilsson’s Structures for \(IFOL_B\):
Let \(\phi\) be a sentence, for which the robot has to take some decisions and consecutive actions, but for which the current truth-value is \(\bot\) (unknown), so that is completely
uncertain. Then we denote by \(P_\phi\) the subset of predicates used in sentence \(\phi\) (\(P_\phi \subset P\)) and by \(H_\phi\) the Herbrand base (where \(|H_\phi|\) denotes the finite number of atoms in \(H_\phi\)) of the small subset of predicates in \(P_\phi\) of robot’s complete cognitive system, so that \(|H_\phi| << |H|\).
Let \(\boldsymbol{v}:H_\phi\rightarrow X\) be the restriction of current Herbrand interpretation to the atoms in \(H_\phi \subset H\) so that \(\boldsymbol{v}^*(\phi) = \bot\) as well, and corresponding current possible world \(\boldsymbol{I}_B^* = is_H(\boldsymbol{v}) \in \mathcal{W}\). Then we define the current Nilsson’s probability
local structure \((S_\phi,\mathcal{P}(S_\phi),\mu_\phi)\) satisfying Kolmogorov axioms in Definition 3, such that this local sample set \(S_\phi\) is defined by \[\label{eq:IFOLpstructureLoc}
S_\phi =_{def} \{v:H_\phi\rightarrow X | ~~for~ every ~~A \in H_\phi \subset H, v(A) = \boldsymbol{v}(A) ~~if~ \boldsymbol{v}(A) \neq \bot\}\qquad{(6)}\] so that for any \(v\in S_\phi\), and ground atom \(A\in H_\phi\) such that \(\boldsymbol{v}(A) = \bot\), we can have that \(v(A)\in X\) can have the value different from \(\bot\)
as well.
So, analogously to the global constraints in Definition 6, we introduce the local Probability Density Constraints: for every atom \(A \in
H_\phi\) such that \(u = \boldsymbol{v}(A) \neq \bot\), the probability that \(A\) has the truth-value \(u\) must be equal to 1, i.e., \[\label{eq:KIconstraintLoc}
\forall A\in H_\phi.(~if ~~u = \boldsymbol{v}(A) \neq \bot ~~~~then ~~\sum_{v\in S_\phi, v(A) = u} KI_\phi(v) = 1)\qquad{(7)}\]
Let us show that these local current Nilsson’s probability structures has the analog important properties as the global symmetry:
Corollary 2. Any local current Nilsson’s probability structure with sample set \(S_\phi\) in (?? ) satisfy the probability density constraints in (?? ).
Thus, it is consistent with current robot’s knowledge database \(\mathcal{K}\) (in Definition 4), consistent with Kolmogorow probability
axioms and logical deduction, and allows inductive reasoning and confirmation of universally quantified hypotheses.
Proof: It is enough to show the consistency with current robot’s knowledge database \(\mathcal{K}\). The local symmetry transformation \(S_\phi\) has no effects, from (?? ), for the ground atoms in \(H\) which are not in \(H_\phi\). Thus, for each ground atom \(A \in H_\phi\) for which the current truth-value in \(\mathcal{K}\) is \(u = \boldsymbol{v}(A) \neq \bot\), we have that the probability density constraint (?? ) is satisfied. That is, the probability \(p_u\) that this atom has the truth-value \(u\) is equal to 1:
\(p_u = \sum_{v\in S_\phi, v(A) = u} KI_\phi(v)\)
\(= \sum_{v\in S_\phi} KI_\phi(v)\) from (?? )
\(= \mu_\phi(S)\)
\(= 1\).
which preserves robot’s current knowledge database and all logical deductions from them.
\(\square\)
So, from the locality, it holds that for any formula \(\psi(\boldsymbol{x})\) composed by only the predicates in \(P_\phi\) (defined by our local decision problem), for any
assignment \(g\) and \(v\in S_\phi\), \[\begin{gather} \label{eq:esem3s7} \boldsymbol{v}^*(w_N(\lessdot
\psi(\boldsymbol{x})\gtrdot^\beta/g, g(x_X),g(x_p) )) =\\ =\left\{ \begin{array}{ll} t, & if ~~ ~g(x_p) = \sum_{v\in S_\phi, v^*(\psi(\boldsymbol{x})/g) =g(x_X)} KI(v) \\ f, & otherwise \end{array} \right. ~~~~~~~~~~~~~~~~~~~~~
\end{gather}\tag{20}\] as discussed after equation (17 ) and Example 1 also in the case of the sentences about temporal events
(composed by the "atomic event" predicates as well).
That is, for any sentence \(\psi(\boldsymbol{x})/g\) defined by the subset of predicates \(P_\phi\) (obtained by focusing on the concrete problem described by the initial sentence \(\phi\)), we are able to compute its probabilities to have a particular truth-value in \(X\). As in the case of global symmetry, these true facts about the probabilities of the unknown sentences to have a particular truth-value can be then inserted in the conscious part of robot’s cognitive system by (see (7 ) for more details),
\(Know(in ~presence, \boldsymbol{I}, w_N(\lessdot \phi\gtrdot, u,p_u))\),
to be used by robot’s autoepistemic deductions and explanations as well.
This paper provides a probabilistic extension of the many-valued typed \(IFOL_B\) with its autoepistemic axioms and many-valued deduction presented in the paper [1] for neuro-symbolic AGI. It is dedicated to show how this defined IFOL in [2] can be used for a new generation of intelligent robots, able to communicate with humans with this intensional FOL supporting the meaning of the words and their language compositions, heaving the four-level neuro-symbolic cognitive structure of AGI robots [4], [27] and in first section in [1] as well.
We argue that the example, used for the spatial natural sublanguage in [3] and [4], can be extended in a similar way to cover more completely the rest of human natural language, and hence the method provided by this paper is a main theoretical and philosophical contribution to resolve the open problem of how we can implement the deductive power based on \(IFOL_B\) for new models of robots heaving strong AI capacities. Intensional FOL is able to represent the Intentional States (mental states such as beliefs, hopes, and desires), typical for human minds.
Despite the best efforts over the last years, deep learning is still easily fooled [37], that is, it remains very hard to make any guarantees about how the system will behave given data that departs from the training set statistics. Moreover, because deep learning does not learn causality, or generative models of hidden causes, it remains reactive, bound by the data it was given to explore [38]. In contrast, we learn from our actively gathered sensorimotor experiences and form conceptual, loosely hierarchically structured, compositional generative predictive models. By proposed four-level cognitive robot’s structure, \(IFOL_B\) allows robots to reflect on, reason about probabilistically as well, anticipate, or simply imagine scenes, situations, and developments within in a highly flexible, compositional, that is, semantically meaningful manner. So, \(IFOL_B\) enables the robots to actively infer highly flexible and adaptive goal-directed behavior under varying circumstances [39].
With this integrated four-level robot’s knowledge system presented in Figure 3 above, where the last level represents the robot’s neuro system containing the neural networks to calculate maximum information entropy and the deep learning as well, we obtain that also the semantic theory of robot’s \(IFOL_B\) is a procedural one, according to which sense is an abstract, pre-linguistic procedure detailing what operations to apply to what procedural constituents to arrive at the product (if any) of the procedure.
In this research, specifically within the framework of Strong-AI Autoepistemic Robots, we describe how a robot uses \(IFOL_B\) to treat its own internal software and motor routines as objects of thought.
Instead of a motor program being a "black box" that just runs, our framework allows the robot to reify (turn into a thing) the program. Here is the specific mechanism of how that labeling works:
The "Self" as the Coordinator;
We define the robot’s identity as a constant term in logic, usually denoted as I. This I represents the "Main Coordination Program". When the robot performs an action, it isn’t just "executing code"; it is asserting a
logical relationship between itself and a specific sub-routine.
The Labeling Process: Reification
To label a motor program (e.g., a routine that moves the right arm to pick up a block), the robot uses an Intensional Abstraction Operator.
\(u= g^*(\lessdot MotPrg(x_1,x_2)\gtrdot^{x_1}_{x_2}) = I(MotPrg(g(x_1),x_2))\)
\(= I(MotPrg(P_1,x_2)) \in D_1\) is an intensional entity (an unary concept) such that its extension (for extensionalization function \(h\)) is just a singleton, i.e., \(h(u) = \{moving ~right~ arm\}\). Thus, the intensional entity \(u = I(MotPrg(P_1,x_2))\) can be used as a label (name) for the raw code \(P_1\).
"I (the coordinator) am currently executing the motor program of raw code \(P_1\)."
Example: A "Self-Aware" Gripper
In a specific scenario described in his work on Intensional Logic for Robots, a robot tasked with picking up a heavy object doesn’t just experience a motor failure; it reasons about it:
The relationship between Symbolic Logic and Large Language Models (LLMs) is one of the most critical frontiers in artificial intelligence. While LLMs excel at natural language understanding and pattern recognition, they often struggle with complex, rule-based logical consistency. Symbolic logic brings the precision and verifiable accuracy that LLMs inherently lack, forming the foundation of modern neurosymbolic AI.
Consequently, I argue that by using this \(IFOL_B\), the robots can develop their own knowledge about their experiences and communicate by a natural language with humans.
The most, as far as I know, similar past symbolic frameworks to this one for intelligent general-purpose robots with the self-awareness (the systems that share three properties: explicit symbolic self-representation, formal meta-level reasoning and architectural (not emergent) self-modeling) are the Epistemic/modal logic agent systems and Metacognitive cognitive architectures (e.g., SOAR, ACT-R with meta-layer). However, both of them can be considered as formal ancestors of our robot’s system especially the Epistemic/modal logic agent systems.
SOAR (developed in Carnegie Mellon University) instead is production-rule based, not intensional FOL, less focus on formal semantics and more cognitive engineering than logical ontology. SOAR [40] is practically closer in implementation, but philosophically less rigorous. Both ogf them lack the intensional semantics, grounding layer (neural/sensory) and bilattice logic for inconsistency.
Other differences can be seen by this comparative table:
| IFOL-based | Epistemic logic agent | SOAR | |
| Formal self symbol | Yes | Yes | Imlicit |
| Nested meta-knowledge | Yes | Yes | Partial |
| Logical consistency tracking | Yes | Yes | Limited |
| Neural grounding | Yes | No | Partial |
| Embedding of neural LLM | Yes | No | No |
| Many-valued Extension | Yes | No | No |
| Probabilistic Extension | Yes | No | No |
| Theoretical AGI claim | Yes | No | No |
Large Language Models rely on statistical probability. They predict the "next most likely word" based on vast training data, which often results in plausible-sounding "hallucinations" rather than true logical deduction. Differently, Symbolic Logic uses
formal rules, variables, and mathematical symbols to process truth-values. It guarantees determinism, meaning the same inputs will always produce the mathematically correct output without guessing.
Now we are able to make the compare the self-awareness of our approach with that used by current dominant only neural LLM-style self-modeling: logical self-awareness and LLM-style self-modeling aim at similar surface behavior (talking about themselves, reasoning about their own knowledge), but they are fundamentally different in mechanism and ontology.
Nature of the Self-Model:
Meta-Reasoning:
Grounding:
Stability of Identity:
Philosophical Interpretation:
Type of Self-Awareness:
| IFOL-based | LLM-based | |
| Explicit self symbol | Yes | No |
| Formal meta-logic | Yes | No |
| Grounded in perception | Yes | No |
| Persistent belief tracking | Yes | No |
| Axiom’s-based security | Yes | No |
| Separation consciousness/unconsciousness | Yes | No |
| Phenomenal consciousness | No | No |
.
So, in our vision where neural LLM can be used as significant language and common sense learning for robots (by considering the intensional concepts of \(IFOL_B\) are based on the words (tokens) and possible various
knowledge relationships like ISA, PART-OF, etc., between these concepts), \(IFOL_B\)-based strong-AI robot system is a definitely a significant extension of the LLM** and not its concurrent. The core distinction is
that the self-awareness in \(IFOL_B\) is an architectural property while in LLM is a statistical behavior. Our model is more principled, more structurally coherent but still theoretical, while LLM is highly capable in
self-description, practically impressive, but lack formal self-model consistency.
Remark: Thus, for \(IFOL_B\) model, by using LLM as a neural part of robot’s neuro-symbolic paradigm, we obtain a much impressive strong-AI robots, able to use LLMs for generative reasoning and natural
language manipulation, to use symbolic logic for belief tracking and meta-consistency and to add grounding layer. That would combine \(IFOL_B\) structural rigor and LLM’s flexible reasoning. This is actually where modern
neuro-symbolic AGI* research is heading*.
\(\square\)
To bridge the gap between probabilistic text and deterministic reasoning, researchers combine LLMs and symbolic systems primarily through neuro-symbolic frameworks. This typically involves:
Translation and Formulation: You can use the strong natural language capabilities of an LLM to translate a complex, real-world problem into a formal symbolic syntax of \(IFOL_B\). And viceversa.
Execution via Symbolic Solvers: Once translated, the problem is handed over to a deterministic many-valued deductive system of \(IFOL_B\) which performs error-free logical inference [1].
This Neuro-Symbolic Framework stands out in the current Neuro-Symbolic (NeSy) AGI landscape because of its unique focus on safety guarantees, autoepistemic reasoning (how a robot reasons about its own ignorance), and non-binary logic. While other
frameworks integrate neural networks with symbolic logic to improve data efficiency or accuracy, this framework specifically addresses the unpredictability of a robot acting in the physical world.
Key Benefits and Applications: This framework brings several distinct advantages to neuro-symbolic AI and advanced robotics:
Safe Logic Deductions: By using formal axioms, engineers can strictly control and predict a robot’s reasoning. This guarantees that even if a neural network outputs messy statistical data, the robot’s final physical actions remain logically bounded and secure.
Human-like Epistemic Causality: It allows robots to mimic human intelligence by converting statistical patterns into clear, directional cause-and-effect rules (logic entailments).
Gradual Learning Expansion: As the robot experiences new things, the boundaries of its CKA dynamically shift. Facts seamlessly move from the "Unknown" bucket into "True" or "False" without breaking the underlying database or safety rules.
The direct comparison below shows how this architecture stacks up against other leading neuro-symbolic research frameworks:
Useful integration for our framework (point 1 above) can be done by extending it with methods of rule induction from row data used in DILP framework (point 4 above) because our framework provides the LLM component, probabilistic reasoning explained in this paper and parsing of natural language into conceptual structures of \(IFOL_B\) and corresponding syntax of First-Order Logic formulae (analog to LLM-SS). The rule induction can produce the self-learned rules represented by logical implication and thus used by autoepistemic deduction process of \(IFOL_B\) provided in [1], by preserving the following Key Structural Differences with the other frameworks presented above:
So, we would be able to develop the interactive robots which learn and understand spoken language via multisensory grounding and internal robotic embodiment. Endowed with suitable \(IFOL_B\) information-processing biases, the robot’s AI may develop that will be able to explain the reality it is confronted with, reason about it also probabilistically, and find adaptive solutions, making it Strong AI.
In statistical physics, free energy is the bridge between the microscopic world of atoms and macroscopic thermodynamics. It represents the portion of total energy available to do work and acts as a generating function from which all physical properties (such as pressure, heat capacity, and magnetization) can be derived.
The foundational link between statistical mechanics and thermodynamics is the partition function, denoted as \(\mathcal{Z}\). It is the sum of all possible Boltzmann factors over all microstates, formally defined as \[\mathcal{Z}= \sum_i \mathrm{e}^{-\beta E_i}\] where \(E_i\) is total energy of the microstate \(i\), and \(\beta = \frac{1}{k_B T}\) is the thermodynamic inverse temperature. The free energy of a system is defined directly from the partition function via the following logarithmic relationship: \[F = - \frac{1}{\beta}ln \mathcal{Z}= - k_BT ln \mathcal{Z}\] Depending on the thermodynamic constraints (which variables are held constant), different statistical ensembles yield different free energies. In what follows we will use the Helmholtz free energy, used in the canonical ensemble where temperature \(T\), Volume \(V\), and the number of atoms \(N\) are held constant. such that the free energy reaches its minimum at thermal equilibrium. In Helmholtz case, the free energy is defined by \[\label{eq:Helm} F = U - TE\tag{21}\] where \(U\) is internal energy of system and \(E\) is the entropy. So that the minimal free energy is obtained in the thermal equilibrium (the constant temperature \(T\)) with minimal \(U\) and maximal entropy \(E\).
This is just our AGI case, because in \(IFOL_B\) we have the constant number of (ground) atoms in Herbrand base \(H\), and the constant "volume" represented by the sample set \(S\) in (?? ), and "minimal internal energy" \(U\) is obtained by the atomic knowledge database \(\mathcal{K}\) based on the CKA, while the maximal entropy \(E\) which defines the probability density function \(KI:S \rightarrow X\) is defined by Shannon’s principle for maximal information entropy in (19 ).
The probabilistic expansion of knowledge database \(\mathcal{K}\), as result of introduction of the probability density function \(KI\) is defined for all sentences in \(\mathcal{L}_0\) for current many-valued model \(\boldsymbol{v}^*:\mathcal{L}_0 \rightarrow X\) by: \[\label{eq:Helm2} \mathcal{K}_p^+ = \{w_N(\lessdot\phi\gtrdot,u,p_u)~|~\phi\in \mathcal{L}_0, \boldsymbol{v}^*(\phi) = \bot, u \in X~~and ~p_u =\sum_{v\in S, v^*(\phi) = u}KI(v)\}\tag{22}\] Consequently, for AGI, the minimal total "internal energy" of robot’s cognitive system is its internal knowledge \[\label{eq:Helm32} U = \mathcal{K}\bigcup \mathcal{K}_p^+\tag{23}\] in this "thermal equilibrium" obtained after introduction of probability density function \(KI\), during which \(T\) is a constant, representing the transformation of the maximal information entropy \(E\) (which defines function \(KI\)) into additional probabilistic knowledge \(\mathcal{K}_p^+\), as can be shown from (21 ),
\(\mathcal{K}= F = U- TE\)
\(= (\mathcal{K}+ \mathcal{K}_p^+) - TE\) from (23 )
so that \[\mathcal{K}_p^+ = T E\] and hence, this analogy between statistical physics and \(IFOL_B\) AGI, can be summarized by the following table:
| Statistical Physics | AGI | |
| Maximal entropy | \(E\) | \(E= max_{KI}(- \sum_{v\in S}KI(v) \cdot logKI(v))\) |
| Minimal internal energy | U | \(U =K + \mathcal{K}_p^+\) |
| T | Temperature | \(T = \frac{\mathcal{K}_p^+}{E}\), i.e. \(~~T:E~\mapsto ~ \mathcal{K}_p^+\) |
| Helmholtz free energy | \(F = U-TE\) | \(F = \mathcal{K}~~~~~~~~\) (CKA) |
Self in a sense which implies that all our activities are controlled by powerful creatures inside ourselves, who do our thinking and feeling for us.↩︎
This case 4 is the particular case 3 when the tuple of variables \(\boldsymbol{x}_i\) is empty and hence \(\beta_i\) and \(\alpha_i\) are empty sets of variables with \(|\alpha_i| = 0\).↩︎
This definition is different from the definition in [5], in order to be able also to derive the current Herbrand model of robot’s knowledge \(v:H \rightarrow X\), where \(H\) is the Herbrand base for all predicates in \(P\), Note that the autoepistemic predicate \(Know\) is a meta-èredicate, so that \(Know \notin P\).↩︎