By Emile van Krieken
Review Details
Reviewer has chosen not to be Anonymous
Overall Impression: Good
Content:
Technical Quality of the paper: Good
Originality of the paper: Yes
Adequacy of the bibliography: Yes, but see detailed comments
Presentation:
Adequacy of the abstract: Yes
Introduction: background and motivation: Limited
Organization of the paper: Needs improvement
Level of English: Satisfactory
Overall presentation: Weak
Detailed Comments:
# Review
## Summary of paper
The authors extend a previously derived formalisation for neurosymbolic methods called NeSyCat to the "NeSy design space" to 'survey' the literature on neurosymbolic AI. They consider 4 dimensions: Syntax, Semantics, Inference and Integration. They use the NeSyCat formalisation to show that all these 4 dimensions can be put into a single formal framework for comparison, and give worked out examples.
## Main review
I should preface this review that I read the NeSyCat paper fairly thoroughly ~half a year ago. I think it has a lot of promise and I liked reading it, but a significant rewrite is needed to make it more generally accessible and clarify and focus the contribution.
Strengths:
- The paper reads, to me, quite well, and I thought it was a good application of NeSyCat. It englithened me on some of the more practical aspects of this framework and how category theory could actually be used for obtaining practical systems and for analysis.
- The authors have a strong commitment to formalisation, which sometimes pays off quite well
- The paper does a lot for a single manuscript: The framework is very ambitious in that it aims to not only compare NeSy methods on a high-level, but even goes down to the computational and semantic level.
Weaknesses:
- It is not clear who the intended audience is or what the core contribution is for them that the authors intended. This makes for a really... curious paper that cannot really be called a proper survey, but also not really a new contribution (at least not in the way it's currently written).
- Although the paper focusses on formalisation, I do not always believe the right level of formalisation is applied. Whether it should be more or less depends on the section and intended audience, IMO.
- The paper is quite riddled with AI assisted language and structures, which often is at the detriment of the paper.
- Many sections are far too dense for a survey, and furthermore miss citations to existing papers.
- Comparison to existing works is not entirely complete
Expanded weaknesses, major comments and questions:
1. Audience and intended core contribution:
- This paper presents itself as a survey in the introduction and the cover letter. It compares to many existing surveys on how it is different. But a survey tends to be for people _new_ to the field to get an idea of what the current research is about, what papers to read, and what gaps there are. I do not believe the paper succeeds in this aspect. I suspect the paper is a very challenging read without having studied the NeSyCat paper, and furthermore requires quite some intuitions and previous knowledge about the field of NeSy to truely grasp.
- So then I wonder, who really was the intended audience and what was the intended goal of the paper? I do not think this is really a survey, but rather a (high-level) formalisation of many aspects of the NeSy literature, meant for active NeSy researchers. Is this correct? If so, the fact that it is a challenging read (esp for not having read NeSyCat) is more acceptable.
- If instead it's really meant as a survey, I would like to know what are the selection criteria for what papers to review and include, as throughout the paper this seems quite arbitrary.
2. Formalisation:
- I find the level of formalisation very inconsistent throughout the paper. The semantics section I think strikes a good balance, but is also most close to NeSyCat. I found the other three sections are however in a weird space that the writing isn't really all that formal, and a lot of jargon is introduced without clear definitions.
- I think the paper would be significantly improved if it centralised and formalised the monadic interpretation framework as a first section before going into the four dimensions. This allows the different dimensions to be read independently, and makes it easier to see the larger picture. Finally, it allows the sections to focus on what I think was their core task: Comparing existing systems across that dimension.
3. One common writing issue is many paragraphs are somewhat 'performative' for lack of a better word. They contain a lot of quick ideas that relate to work done in NeSy(-adjacent fields), but do not contribute to the overall story and are hard to follow without already being familiar with the content. And if familiar, they are not actually teaching the reader new ideas. I'll give examples below.
4. The paper misses a related work section where it compares thoroughly to existing surveys and work.
- The title of the paper is far too similar to [1]... But this (very strong, IMO) survey is not cited at all, even though it's probably the closest comparison point as it's quite well formalised.
- Other works that aim to unify (such as ULLER and Dickens et al 2025) should be more thoroughly compared to
- NeSyCat is of course a large reference point, but is not properly compared to either. How much is the paper taking from it, and what is new?
5. The AI language (I think!) that is most hurting the paper
## Comments per section
I will refer to the main comments above as evidence.
Introduction
- The huge wall of citations to surveys in the intro seems unnecessary
- (4) I think the introduction almost reads like a (somewhat defensive) related work section, mostly focusing on comparing to existing frameworks. A bit more focus on the core contributions and motivation of the paper itself in the first paragraphs would help new readers.
- The semantics section says "we show x y z are instances of a single structure...", but this was already done in NeSyCat afaik.
- The paper discusses 4 "dimensions", which seems to suggest the four sections can be captured in a simple way. But each of these 4 is seriously complicated and cannot be captured in a single "dimension". Maybe a better word could be found?
Syntax
- (2) I found this section a bit strange at times. As the first content section, it doesn't really talk about anything novel or the monadic interpretation. I wonder if it is really doing a
- (2) The section entirely focusses on vocabuary and not on the grammar rules for combining them into language statements. This is an odd choice to me, as for example 2-variable fragment of FOL is weaker than full FOL itself yet has the same sets of symbols, technically.
- It also doesn't consider the difference between logic programming and standard logics thoroughly. 2.3 seems to suggest DeepProbLog is a full FOL, for example.
- (2, 3, 5?) That said, of course it's also not possible to capture everything about syntax. I think this section might be doing a bit too much. Very little is actually done with the differences in syntax, and almost nowhere in the monadic interpretation is there a payoff. In particular the HOL and temporal sections are far too short to really get anything out of, yet are not referred to anywhere after (except in the future work saying it should be done better :-)). I'd drop them.
- Nit: Model operators aren't quantifiers :-)
- (3, 5?) The second and third paragraphs in page 6 are instances of 3., I'd say. Without having discussed inference/semantics, the second paragraph here is far too soon. And it's not clear to me what it's doing for the story you want to tell. Maybe better at the end of this section?
- I didn't get the argument why DeepProbLog is a relational language
- Scallop is based on Datalog, but I always thought datalog doesn't contain negations and disjunctions, but it does have implications.
Semantics
- (5) "If syntax is the contract a NeSy system exposes, Semantics is its fulfillment" is some serious AI speak :-)
- "Received the least systematic attention" is not my experience? Also [1] is a good source.
- I was very lost on the formalisation of Tensors and how it relates to the other two. Especially how its basis is a truth value.
- (1) I suspect this section is not easy to follow for people who have not read NeSyCat (or aren't familair with CT/Kleisli/Monads. I would not put the table so early, and introduce much slower the ideas of the monadic interpretation.
- (1) In particular page 14 is a serious challenge. I think you will lose a lot of readers here, which would be a shame! I would start with examples and go much slower.
- The Axioms column is not well defined here.
- (5?) "Category theory provides the right language to resolve this uniformly" is a strong claim without considering alternatives
- (2) It is not clear how the parameters come into play
- (2) Definition 3 requires cartesian closed categories but this is only quickly mentioned below that
- Nit: 2Mon-BLat is sometimes written without -
- Also the 2Mon-Blat should have a citation to NeSyCat
- The example with the 'law of total probability' is a Markov chain or something? It would be nice to make the probabilistic interpretation more explicitly.
- Why is the commutativity requirement of T not stated? It's not clear to me also why it's necessary. Also introduce p and q
- It seems wrong to me that TY is the space of effectful compute over X?
- It's not clear to me why the distributional semantics uses the Product-BL axioms rather than a boolean one
- Nit: LTN doesn't necessarily fix the fuzzy operators
- Nit: "Godel negation" is notdefined as 1-a, 1-a is the standard negation instead
- Footnote 3: What is 'fuzzy equality'?
- The DeepProbLog and DeepSeaProbLog examples are very hard to follow. Even though I think these could be very useful examples! Also, proof theoretic at this point is not yet introduced.
- For deepSeaProbLog, how is temp~N() modelled in our syntax? And what is a Kleisli bind?
- Not sure about this, but the Product BL algebra uses the probabilistic sum also for \exists, yet in DeepProbLog it's interpreted as a sum.
Inference
- Again the table here could appear a lot later
- (2) The table is... very dubious to me. I am not sure how it was obtained and it contains no citations. I don't think the three axes are formalised well enough to have a clear idea how these are instantiated. For instance, what does proof-theoretical probabilistic mean? What is hybrid?
- Nit: Viterbi is exact, not approximate
- (2) In general, where the previous section is very formal, this one is very informal, and it makes it quite unclear. Page 19 and 20, for instance, uses much more philosophical intrductions, which is a stark contrast.
- (2) Then definition 7 and 8 give the impression of being more formal, but it's not clear to me how these are applicable to the non-logical categories. For instance, what do \proves and \models mean in these categories? What is a theory? What are deduction systems, formally?
- I also have no idea how the table at bottom 20 is created
- (2) Definitions 9 and 10 again are not so clear either. What is 'an inference' formally? What is an 'estimate'? Should there be any guarantees on the accuracy of this estimate? Or can it be any computation :-)?
- (2) Because of the limited formalisation, I could only follow the examples from an intuitive point of view (and from knowing the existing papers)
- Why is LTN called approximate? I thought the point was that (model-theoretically) the fuzzy semantics are exactly computed? But now it says "fuzzy semantics only approximates classical logical validity" - in what sense?
Integration
- I think this is really where the framework has to pay off. I think it does to some degree. But I'm not sure about structuring it around Kautz' taxonomy (which seems to be a common... interesting... choice... in the NeSy survey sphere). Instead the three patterns introduced here are much more interesting.
- Nit: "N calls L/P" is said to be rare, but in current NeSy applications it's the most common (think: LLMs calling tools)
- (2) Although I like the overall setup of this section, I found the formalisation on a strange leel. For isntance, definitions 11 and 12 are equivalent, except f is either called a 'neural-parameterized symbol" or a "symbol with a fixed L/P interpretation". It is not clear to me what these terms formally mean.
- Nit: Example 15 now uses LTn with Lukasiewicz connectives while earlier it was defined with the Product BL.
- "every NeSy architecture in the literature can be written as a Kleisli term over them" is far too strong a claim with the limited surveyed literature.
- Example 13: LTN does not generalise semantic loss at all, they compute very different quantities
- Nit: Example 14: In A-NeSI, it can be seen as compilation, but the approximation neural network can also be trained jointly online :) (Btw, you used the wrong initial of my name for the citation! :))
Conclusions
- Semantics: "How compositionality is achieved within the neural component itself" This is an active area of research which desrves a cite [2]
Appendix (I did not read the appendix; comments here are nits claude found and I checked)
- Belnap and Kleene conjunction is not correct as then {0} and {1} = {}, but it should be {0}
- Why is 2Mon-Blat the 'minimal' algebraic requirement?
Further nits
- In some places eg in semantics, it's stated there is no 'fundamental' difference between fuzzy and probabilistic semantics, yet appendix E2 says 'the two semantics are fundamentally different'. I would agree with the latter and the paper doesn't make the earlier claim strong imo. Also, "Belle and Marcus (2025) note that whether fuzzy or probabilistic semantics is better suited for NeSy systems remains a matter of debate … The Kleisli-categorical framework presented here resolves this by showing both are instantiations of the same mathematical basis." I don't think this is resolved at all, just showing it is a common structure says nothing about the relevant practical aspects that are commonly discussed in the NeSy literature.
- The abstract says the systems differ along 'independent dimensions', but in limitations the opposite is mentioned
A comment Claude had that I don't have the brain capacity to understand :-)
- Apparently BorelMeas is not cartesian closed, which motivated quasi-Borel spaces in Heunen-Kammer-Staton-Yang 2016. Aumann 1961 supposedly proved this.
## References
1. Feldstein, Jonathan, et al. "Mapping the neuro-symbolic AI landscape by architectures: A handbook on augmenting deep learning through symbolic reasoning." arXiv preprint arXiv:2410.22077 (2024).
2. Gavranović, Bruno, et al. "Position: Categorical deep learning is an algebraic theory of all architectures." Forty-first International Conference on Machine Learning. 2024.