Safety
LLM-Guided Contextual Action Evaluation for Operational Decisions in Industrial Processes
arXiv:2608.24156v1 Announce Type: cross Abstract: Industrial actor--critic methods usually represent continuous actions as anonymous numerical coordinates. They must therefore learn from limited inter
arXiv:2608.24156v1 Announce Type: cross Abstract: Industrial actor--critic methods usually represent continuous actions as anonymous numerical coordinates. They must therefore learn from limited interactions which process variables each action affects, in which direction, and after what delay. Fixed industrial documents already describe part of these relations, but their open-text statements neither represent the current operating condition nor directly fit a numerical policy. This article presents LLM-Guided Contextual Action Evaluation for Operational Decisions in Industrial Processes (LCAE), which uses a large language model before training to normalize fixed documents into a frozen action--observation--direction--delay relation basis. Recent numerical action--response history then modulates the current strength of each relation, while the evaluated action forms a state-conditioned nonlinear action-effect field in the same basis. The critic evaluates actions through this field, and the actor uses the same relation gains to generate actions, making document semantics part of maximum-entropy policy learning. Neither the LLM nor the embedding model runs online during training or deployment; the deployed policy uses only frozen semantic artifacts and visible numerical history. The method states a falsifiable hypothesis: when documented relations are correct and recent history reflects their contextual strength, this action representation should provide a more useful decision bias than raw action coordinates.
Related
- Gauging, Measuring, and Controlling Critic Complexity in Actor-Critic Reinforcement Learning
- Topology-Guided Modular Actor-Critic Learning for Continuous Systems under Temporal Objectives
- ART for Diffusion Sampling: Continuous-Time Control and Actor-Critic Learning
- Second-Order Actor-Critic Methods for Discounted MDPs via Policy Hessian Decomposition
Source: arXiv cs.AI | 2026-08-26