Safety
Governing What You Cannot Observe: Adaptive Runtime Governance for Autonomous AI Agents
arXiv:2604.24686v1 Announce Type: new Abstract: Autonomous AI agents can remain fully authorized and still become unsafe as behavior drifts, adversaries adapt, and decision patterns shift without any
arXiv:2604.24686v1 Announce Type: new Abstract: Autonomous AI agents can remain fully authorized and still become unsafe as behavior drifts, adversaries adapt, and decision patterns shift without any code change. We propose the extbf{Informational Viability Principle}: governing an agent reduces to estimating a bound on unobserved risk hat{B}(x) = U(x) + SB(x) + RG(x) and allowing an action only when its capacity S(x) exceeds hat{B}(x) by a safety margin. The extbf{Agent Viability Framework}, grounded in Aubin's viability theory, establishes three properties -- monitoring (P1), anticipation (P2), and monotonic restriction (P3) -- as individually necessary and collectively sufficient for documented failure modes. extbf{RiskGate} instantiates the framework with dedicated statistical estimators (KL divergence, segment-vs-rest z-tests, sequential pattern matching), a fail-secure monotonic pipeline, and a closed-loop Autopilot formalised as an instance of Aubin's regulation map with kill-switch-as-last-resort; a scalar Viability Index VI(t) in [-1,+1] with first-order t^* prediction transforms governance from reactive to predictive. Contributions are the theoretical framework, the reference implementation, and analytical coverage against published agent-failure taxonomies; quantitative empirical evaluation is scoped as follow-up work.
Related
- OpenKedge: Governing Agentic Mutation with Execution-Bound Safety and Evidence Chains
- Interval POMDP Shielding for Imperfect-Perception Agents
- Beyond Static Sandboxing: Learned Capability Governance for Autonomous AI Agents
- Failure-Centered Runtime Evaluation for Deployed Trilingual Public-Space Agents
Source: arXiv cs.AI | 2026-04-28