Safety
Neuro-Symbolic Meta-Policies for Temporal Knowledge-Graph Memory under Partial Observability
arXiv:2607.18368v2 Announce Type: replace Abstract: Partially observable reinforcement learning requires deciding what to retain, retrieve, and forget over time. We introduce a neuro-symbolic meta-pol
arXiv:2607.18368v2 Announce Type: replace Abstract: Partially observable reinforcement learning requires deciding what to retain, retrieve, and forget over time. We introduce a neuro-symbolic meta-policy that learns which symbolic memory heuristic to apply at each decision point while keeping execution symbolic. Our setting uses temporal knowledge-graph memory in RoomKG, where hidden state and observations are represented as Resource Description Framework (RDF) graphs and memory is augmented with temporal RDF triple annotations. The model combines knowledge-graph encoding of memory contents with value heads for question answering, exploration, and forgetting, yielding a controller that is both adaptive and inspectable. This gives the work a direct Semantic Web grounding through RDF-based representation, annotation-compatible graph semantics, and graph-based symbolic operations over explicit memory state. On train/test room splits at long-term memory capacity of 512, the qualifier-aware StarE-GNN configuration achieves the best held-out performance among the compared symbolic, neural, and neuro-symbolic systems while preserving step-level traceability of memory-management decisions.
Related
- Neuro-Symbolic Injection of LTLf Constraints in Autoregressive Reinforcement Learning Policies
- Why Linear Recurrent Memory Works in Partially Observable Reinforcement Learning
- Synthesizing POMDP Policies: Sampling Meets Model-checking via Learning
- SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents
- From Passive Reuse to Active Reasoning: Grounding Large Language Models for Neuro-Symbolic Experience Replay
Source: arXiv cs.AI | 2026-07-28