AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,704 results
1 May 2026

Agent-Agnostic Evaluation of SQL Accuracy in Production Text-to-SQL Systems

SafetyDGX agent

arXiv:2604.28049v1 Announce Type: new Abstract: Text-to-SQL (T2SQL) evaluation in production environments poses fundamental challenges that existing benchmarks do not address. Current evaluation metho

Agent Name Service (ANS): A Proof-of-Concept Trust Layer for Secure AI Agent Discovery, Identity, and Governance in Kubernetes

SafetyDGX agent

arXiv:2604.26997v1 Announce Type: cross Abstract: Autonomous AI agent ecosystems require stronger mechanisms for secure discovery, identity verification, capability attestation, and policy governance.

Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2601.01885v2 Announce Type: replace Abstract: Large language model (LLM) agents face fundamental limitations in long-horizon reasoning due to finite context windows, making effective memory mana

AI and GDP - what does it mean? Sure, AI is “contributing” to GDP — but @davidsacks really nails it below, “Paying people to dig holes and f…

SafetyDGX agent

AI and GDP - what does it mean? Sure, AI is “contributing” to GDP — but @davidsacks really nails it below, “Paying people to dig holes and fill them back up increases GDP, as it is currently measured,

AI Models for Depressive Disorder Detection and Diagnosis: A Review

SafetyDGX agent

arXiv:2508.12022v2 Announce Type: replace Abstract: Major Depressive Disorder is one of the leading causes of disability worldwide, yet its diagnosis still depends largely on subjective clinical asses

AID: Agent Intent from Diffusion for Multi-Agent Informative Path Planning

SafetyDGX agent

arXiv:2512.02535v2 Announce Type: replace Abstract: Information gathering in large-scale or time-critical scenarios (e.g., environmental monitoring, search and rescue) requires broad coverage within l

An adaptive wavelet-based PINN for problems with localized high-magnitude source

SafetyDGX agent

arXiv:2604.28180v1 Announce Type: new Abstract: In recent years, physics-informed neural networks (PINNs) have gained significant attention for solving differential equations, although they suffer fro

Analytical Correction for Subsampling Bias in Drifting Models

SafetyDGX agent

arXiv:2604.27239v1 Announce Type: new Abstract: Drifting models are capable one-step generative models trained to follow a drifting field. The field combines attractive and repulsive softmax-weighted

ANCORA: Learning to Question via Manifold-Anchored Self-Play for Verifiable Reasoning

SafetyDGX agent

arXiv:2604.27644v1 Announce Type: cross Abstract: We propose a paradigm shift from learning to answer to learning to question: can a language model generate verifiable problems, solve them, and turn t

Assessing the Role of Intersection Proximity in Pedestrian Crashes: Insights from Data Mining Approach

SafetyDGX agent

arXiv:2604.28065v1 Announce Type: cross Abstract: Although intersections are the most complex parts of the roadway network, pedestrian crashes at non-intersection locations are disproportionately freq

AttriBE: Quantifying Attribute Expressivity in Body Embeddings for Recognition and Identification

SafetyDGX agent

arXiv:2604.27218v1 Announce Type: new Abstract: Person re-identification (ReID) systems that match individuals across images or video frames are essential in many real-world applications. However, exi

Automatic Causal Fairness Analysis with LLM-Generated Reporting

SafetyDGX agent

arXiv:2604.27011v1 Announce Type: cross Abstract: AutoML, intended as the process of automating the application of machine learning to real-world problems, is a key step for AI popularisation. Most Au

Bayesian policy gradient and actor-critic algorithms

SafetyDGX agent

arXiv:2604.27563v1 Announce Type: new Abstract: Policy gradient methods are reinforcement learning algorithms that adapt a parameterized policy by following a performance gradient estimate. Convention

Belief-Guided Inference Control for Large Language Model Services via Verifiable Observations

SafetyDGX agent

arXiv:2604.27536v1 Announce Type: new Abstract: In black-box large language model (LLM) services, response reliability is often only partially observable at decision time, while stronger inference pat

Beyond Pixel Fidelity: Minimizing Perceptual Distortion and Color Bias in Night Photography Rendering

SafetyDGX agent

arXiv:2604.28136v1 Announce Type: new Abstract: Night Photography Rendering (NPR) poses a significant challenge due to the extreme contrast between dark and illuminated areas in scenes, stemming from

BicKD: Bilateral Contrastive Knowledge Distillation

SafetyDGX agent

arXiv:2602.01265v2 Announce Type: replace Abstract: Knowledge distillation (KD) is a machine learning framework that transfers knowledge from a teacher model to a student model. The vanilla KD propose

Bridging Values and Behavior: A Hierarchical Framework for Proactive Embodied Agents

SafetyDGX agent

arXiv:2604.27699v1 Announce Type: new Abstract: Current embodied agents are often limited to passive instruction-following or reactive need-satisfaction, lacking a stable, high-order value framework e

CLAMP: Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining

SafetyDGX agent

arXiv:2602.00937v2 Announce Type: replace-cross Abstract: Leveraging pre-trained 2D image representations in behavior cloning policies has achieved great success and has become a standard approach for

ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval

SafetyDGX agent

arXiv:2604.27591v1 Announce Type: cross Abstract: Video moment retrieval is the task of retrieving specific segments of a video corresponding to a given text query. Recent studies have been conducted

Co-Evolving Policy Distillation

SafetyDGX agent

arXiv:2604.27083v1 Announce Type: new Abstract: RLVR and OPD have become standard paradigms for post-training. We provide a unified analysis of these two paradigms in consolidating multiple expert cap

CoAX: Cognitive-Oriented Attribution eXplanation User Model of Human Understanding of AI Explanations

SafetyDGX agent

arXiv:2604.27354v1 Announce Type: new Abstract: Explainable AI (XAI) aims to improve user understanding and decisions when using AI models. However, despite innovations in XAI, recent user evaluations

Connected Dependability Cage: Run-Time Function and Anomaly Monitoring for the Development and Operation of Safe Automated Vehicles

SafetyDGX agent

arXiv:2604.27728v1 Announce Type: new Abstract: The advancement of automated vehicles introduces complex safety challenges, particularly in dynamic and unpredictable environments where AI-enabled perc

Constrained Policy Optimization with Cantelli-Bounded Value-at-Risk

SafetyDGX agent

arXiv:2601.22993v3 Announce Type: replace Abstract: We introduce the Value-at-Risk Constrained Policy Optimization algorithm (VaR-CPO), a sample efficient and conservative method designed to optimize

Consumer Attitudes Towards AI in Digital Health: A Mixed-Methods Survey in Australia

SafetyDGX agent

arXiv:2604.27744v1 Announce Type: new Abstract: AI applications are increasingly being introduced into digital health. While technical performance has advanced rapidly, successful deployment mainly de

Continuous-time q-learning for mean-field control with common noise, part-I: Theoretical foundations

SafetyDGX agent

arXiv:2604.27372v1 Announce Type: cross Abstract: This paper investigates the continuous-time counterpart of the Q-function for entropy-regularized mean-field control (MFC) with controlled common nois

Cool paper from Meta FAIR. It's on self-improving LLMs but on the pretraining side. (bookmark it) Most LLM safety, factuality, and reasoning…

SafetyDGX agent

Cool paper from Meta FAIR. It's on self-improving LLMs but on the pretraining side. (bookmark it) Most LLM safety, factuality, and reasoning fixes get bolted on at post-training. By then, the patterns

Cost-Aware Learning

SafetyDGX agent

arXiv:2604.28020v1 Announce Type: new Abstract: We consider the problem of Cost-Aware Learning, where sampling different component functions of a finite-sum objective incurs different costs. The objec

Cross-Lingual Sentiment Misalignment: Auditing Multilingual Language Models for Inversion Risk, Dialectal Representation, and Affective Stability

SafetyDGX agent

arXiv:2602.17469v2 Announce Type: replace Abstract: Recent advances in multilingual representation learning aim to bridge the performance gap between high- and low-resource languages, yet their abilit

Cross-Subject Generalization for EEG Decoding: A Survey of Deep Learning Methods

SafetyDGX agent

arXiv:2604.27033v1 Announce Type: new Abstract: Deep learning for cross-subject EEG decoding is hindered by high inter-subject variability, which introduces a severe domain shift between training and

Customer service in the age of AI has been become truly horrible. Excruciatingly bad.

SafetyDGX agent

Customer service in the age of AI has been become truly horrible. Excruciatingly bad. Hey @DHLCanadaHelp, your customer service really and truly sucks. You claimed to try to deliver a package, but did

D3-Gym: Constructing Real-World Verifiable Environments for Data-Driven Discovery

SafetyDGX agent

arXiv:2604.27977v1 Announce Type: new Abstract: Despite recent progress in language models and agents for scientific data-driven discovery, further advancing their capabilities is held back by the abs

Debiasing Reward Models via Causally Motivated Inference-Time Intervention

SafetyDGX agent

arXiv:2604.27495v1 Announce Type: cross Abstract: Reward models (RMs) play a central role in aligning large language models (LLMs) with human preferences. However, RMs are often sensitive to spurious

Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards

SafetyDGX agent

arXiv:2603.09117v2 Announce Type: replace-cross Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) significantly enhances large language models (LLMs) reasoning but severely suffers from

Design Structure Matrix Modularization with Large Language Models

SafetyDGX agent

arXiv:2604.28018v1 Announce Type: cross Abstract: Design Structure Matrix (DSM) modularization, the task of partitioning system elements into cohesive modules, is a fundamental combinatorial challenge

Designing Ethical Learning for Agentic AI: Toegye Yi Hwang's Ethical Emotion Regulation Framework

SafetyDGX agent

arXiv:2604.26958v1 Announce Type: cross Abstract: Agentic AI systems capable of autonomous goal setting and proactive intervention introduce new challenges for regulating moral-emotional processes in

Detecting Clinical Discrepancies in Health Coaching Agents: A Dual-Stream Memory and Reconciliation Architecture

SafetyDGX agent

arXiv:2604.27045v1 Announce Type: cross Abstract: As Large Language Model (LLM) agents transition from single-session tools to persistent systems managing longitudinal healthcare journeys, their memor

Distributional Alignment Games for Answer-Level Fine-Tuning

SafetyDGX agent

arXiv:2604.27166v1 Announce Type: new Abstract: We focus on the problem of Answer-Level Fine-Tuning (ALFT), where the goal is to optimize a language model based on the correctness or properties of its

DOT-Sim: Differentiable Optical Tactile Simulation with Precise Real-to-Sim Physical Calibration

SafetyDGX agent

arXiv:2604.27367v1 Announce Type: cross Abstract: Simulating optical tactile sensors presents significant challenges due to their high deformability and intricate optical properties. To address these

Dreaming Across Towns: Semantic Rollout and Town-Adversarial Regularization for Zero-Shot Held-Out-Town Fixed-Route Driving in CARLA

SafetyDGX agent

arXiv:2604.27994v1 Announce Type: new Abstract: Learned driving agents often degrade when deployed in unseen environments. This paper studies a deliberately bounded instance of that problem in the CAR

Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry

SafetyDGX agent

arXiv:2604.27019v1 Announce Type: cross Abstract: Safety-aligned language models must refuse harmful requests without collapsing into broad over-refusal, but the training-time mechanisms behind this t

Efficient Preimage Approximation for Neural Network Certification

SafetyDGX agent

arXiv:2505.22798v3 Announce Type: replace-cross Abstract: The growing reliance on artificial intelligence in safety- and security-critical applications is raising concerns about the robustness of neur

Elon Musk, Sam Altman, the future of humanity, and … goblins. Terrific interview @theinformation with @rocketalignment https://www.youtube.c…

SafetyDGX agent

This post references an interview featuring Elon Musk and Sam Altman discussing AI safety, existential risks, and the future of humanity, with an unconventional or humorous element involving goblins.

Exploration Hacking: Can LLMs Learn to Resist RL Training?

SafetyDGX agent

arXiv:2604.28182v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become essential to the post-training of large language models (LLMs) for reasoning, agentic capabilities and alignmen

Exploring Applications of Transfer-State Large Language Models: Cognitive Profiling and Socratic AI Tutoring

SafetyDGX agent

arXiv:2604.27454v1 Announce Type: new Abstract: Large language models (LLMs) sometimes exhibit qualitative shifts in response style under sustained self-referential dialogue conditions (Berg et al., 2

EXPO: Stable Reinforcement Learning with Expressive Policies

SafetyDGX agent

arXiv:2507.07986v3 Announce Type: replace-cross Abstract: We study the problem of training and fine-tuning expressive policies with online reinforcement learning (RL) given an offline dataset. Trainin

Fairness for distribution network operations and planning

SafetyDGX agent

arXiv:2604.27669v1 Announce Type: new Abstract: The incorporation of fairness into the distribution network (DN) planning and operation has become a key goal of recent studies. The cost of implementin

Focus Session: Autonomous Systems Dependability in the era of AI: Design Challenges in Safety, Security, Reliability and Certification

SafetyDGX agent

arXiv:2604.27807v1 Announce Type: new Abstract: The design of embedded safety-critical systems such as those used in next-generation automotive and autonomous platforms, is increasingly challenged by

FP-IRL: Fokker--Planck Inverse Reinforcement Learning -- A Physics-Constrained Approach to Markov Decision Processes

SafetyDGX agent

arXiv:2306.10407v3 Announce Type: replace-cross Abstract: Inverse reinforcement learning (IRL) is a powerful paradigm for uncovering the incentive structure that drives agent behavior, by inferring an

Frequency-Aware Semantic Fusion with Gated Injection for AI-generated Image Detection

SafetyDGX agent

arXiv:2604.27875v1 Announce Type: new Abstract: AI-generated images are becoming increasingly realistic and diverse, posing significant challenges for generalizable detection. While Vision Foundation

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback

SafetyDGX agent

arXiv:2502.07645v3 Announce Type: replace Abstract: Behavior cloning (BC) optimizes policies by treating human demonstrations as pointwise action labels. While effective with accurate action labels, t

From Prompt to Physical Actuation: Holistic Threat Modeling of LLM-Enabled Robotic Systems

SafetyDGX agent

arXiv:2604.27267v1 Announce Type: cross Abstract: As large language models are integrated into autonomous robotic systems for task planning and control, compromised inputs or unsafe model outputs can

From surveillance to signalling: escalation channels as environmental controls for agentic AI

SafetyDGX agent

arXiv:2510.05192v2 Announce Type: replace-cross Abstract: When AI agents operating with access to sensitive information encounter a conflict between completing an assigned task and following rules or

GAVEL: Towards Rule-Based Safety Through Activation Monitoring

SafetyDGX agent

arXiv:2601.19768v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly paired with activation-based monitoring to detect and prevent harmful behaviors that may not be appare

GSDrive: Reinforcing Driving Policies by Multi-mode Trajectory Probing with 3D Gaussian Splatting Environment

SafetyDGX agent

arXiv:2604.28111v1 Announce Type: new Abstract: End-to-end (E2E) autonomous driving presents a promising approach for translating perceptual inputs directly into driving actions. However, prohibitive

Hey @DHLCanadaHelp, your customer service really and truly sucks. You claimed to try to deliver a package, but didn’t actually contact me, d…

SafetyDGX agent

Hey @DHLCanadaHelp, your customer service really and truly sucks. You claimed to try to deliver a package, but didn’t actually contact me, didn’t leave a service card, your automated software won’t le

How Hard Is Continuous Clustering? Lower Bounds from the Existential Theory of the Reals

SafetyDGX agent

arXiv:2604.26972v1 Announce Type: cross Abstract: This paper studies the computational difficulty of clustering problems that are defined directly on a continuous probability density. Rather than work

How to Guide Your Flow: Few-Step Alignment via Flow Map Reward Guidance

SafetyDGX agent

arXiv:2604.27147v1 Announce Type: cross Abstract: In generative modeling, we often wish to produce samples that maximize a user-specified reward such as aesthetic quality or alignment with human prefe

I dunno. The competition is really tight. But Zuck certainly a top contender! Who’s your “favorite”?

SafetyDGX agent

This appears to be a casual social media post by AI researcher Gary Marcus discussing competitive dynamics in the AI field, with a conversational reference to Mark Zuckerberg as a notable figure or 't

I have many beefs with Dario and don’t trust him or his hype — but he has certainly eaten OpenAI’s lunch, despite their immense initial lead…

SafetyDGX agent

I have many beefs with Dario and don’t trust him or his hype — but he has certainly eaten OpenAI’s lunch, despite their immense initial lead. Maybe “clown” isn’t the right word here. I think Jensen ag

I just had my first ride with HOVR, a Canadian alternative to Uber. They are cheaper and pay their drivers more. They already have over 3000…

SafetyDGX agent

I just had my first ride with HOVR, a Canadian alternative to Uber. They are cheaper and pay their drivers more. They already have over 3000 drivers in the Greater Toronto Area. You can download HOVR

← Previous
1…169170171172173…212
Next →