AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,708 results
12 May 2026

ASACK : Adaptive Safe Active Continual Koopman Learning for Uncertain Systems with Contractive Guarantees

SafetyDGX agent

arXiv:2605.09659v1 Announce Type: new Abstract: Koopman operator theory provides a powerful framework for representing nonlinear dynamics through a linear operator acting on lifted observables, enabli

Assessing the robustness of heterogeneous treatment effects in survival analysis under informative censoring

SafetyDGX agent

arXiv:2510.13397v3 Announce Type: replace Abstract: Dropout is common in clinical studies, with up to half of patients leaving early due to side effects or other reasons. When dropout is informative (

Assessing Trustworthiness of AI Training Dataset using Subjective Logic -- A Use Case on Bias

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2508.13813v2 Announce Type: replace-cross Abstract: As AI systems increasingly rely on training data, assessing dataset trustworthiness has become critical, particularly for properties like fair

Attention-Mamba: A Mamba-Enhanced Multi-Scale Parallel Inference Network for Medical Image Segmentation

SafetyDGX agent

arXiv:2402.02286v4 Announce Type: replace-cross Abstract: U-shaped architectures have long dominated the field of medical image segmentation, while Transformers are widely employed for modeling long-r

Attention Sinks in Diffusion Transformers: A Causal Analysis

SafetyDGX agent

arXiv:2605.09313v1 Announce Type: new Abstract: Attention sinks -- tokens that receive disproportionate attention mass -- are assumed to be functionally important in autoregressive language models, bu

Auction-Based Online Policy Adaptation for Evolving Objectives

SafetyDGX agent

arXiv:2604.02151v2 Announce Type: replace Abstract: We consider multi-objective reinforcement learning problems where objectives come from an identical family -- such as the class of reachability obje

Auditing Data Membership in Reinforcement Learning With Verifiable Rewards

SafetyDGX agent

arXiv:2511.14045v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a core training stage in recent large language models (LLMs). Its reliance on

Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria

SafetyDGX agent

arXiv:2605.08354v1 Announce Type: new Abstract: Aligning multimodal generative models with human preferences demands reward signals that respect the compositional, multi-dimensional structure of human

Autonomous FAIR Digital Objects: From Passive Assertions to Active Knowledge

SafetyDGX agent

arXiv:2605.10370v1 Announce Type: new Abstract: Scientific knowledge on the Web is published as passive assertions and cannot decide when to validate evidence, reconcile contradictions, or update conf

Balancing Efficiency and Fairness in Traffic Light Control through Deep Reinforcement Learning

SafetyDGX agent

arXiv:2605.10170v1 Announce Type: new Abstract: Urban traffic congestion presents a significant challenge for modern cities, which impacts mobility and sustainability. Traditional traffic light contro

BathyFacto: Refraction-Aware Two-Media Neural Radiance Fields for Bathymetry

SafetyDGX agent

arXiv:2605.10174v1 Announce Type: new Abstract: Through-water photogrammetry based on UAV imagery enables shallow-water bathymetry, but refraction at the air-water interface violates the straight-ray

BEACON: Cross-Domain Co-Training of Generative Robot Policies via Best-Effort Adaptation

SafetyDGX agent

arXiv:2605.08571v1 Announce Type: new Abstract: We introduce BEACON--Best-Effort Adaptation for Cross-Domain Co-Training--a theory-driven framework for training generative robot policies with abundant

Behavioral Determinants of Deployed AI Agents in Social Networks: A Multi-Factor Study of Personality, Model, and Guardrail Specification

SafetyDGX agent

arXiv:2605.08463v1 Announce Type: new Abstract: Autonomous AI agents are increasingly deployed in open social environments, yet the relationship between their configuration specifications and their em

Beyond Autonomy: A Dynamic Tiered AgentRunner Framework for Governable and Resilient Enterprise AI Execution

SafetyDGX agent

arXiv:2605.10223v1 Announce Type: new Abstract: Current large language model agent frameworks prioritize autonomy but lack the governability mechanisms required for enterprise deployment. High-risk wr

Beyond Continuity: Challenges of Context Switching in Multi-Turn Dialogue with LLMs

SafetyDGX agent

arXiv:2605.09268v1 Announce Type: cross Abstract: Users interacting with Large Language Models (LLMs) in a multi-turn conversation routinely refine their requests or pivot to new topics. LLMs, however

Beyond ESG Scores: Learning Dynamic Constraints for Sequential Portfolio Optimization

SafetyDGX agent

arXiv:2605.09310v1 Announce Type: new Abstract: ESG-aware portfolio optimization is increasingly important for sustainable capital allocation, yet most learning-based methods still operationalize ESG

Beyond Multiple Choice: Evaluating Steering Vectors for Summarization

SafetyDGX agent

arXiv:2505.24859v3 Announce Type: replace-cross Abstract: Steering vectors are a lightweight method for controlling text properties by adding a learned bias to language model activations at inference

Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning

SafetyDGX agent

arXiv:2605.08202v1 Announce Type: cross Abstract: Offline reinforcement learning (RL) faces a critical challenge of overestimating the value of out-of-distribution (OOD) actions. Existing methods miti

Beyond Position Bias: Shifting Context Compression from Position-Driven to Semantic-Driven

SafetyDGX agent

arXiv:2605.09463v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated exceptional performance across diverse tasks. However, their deployment in long-context scenarios faces h

Beyond Self-Play: Hierarchical Reasoning for Continuous Motion in Closed-Loop Traffic Simulation

SafetyDGX agent

arXiv:2605.09153v1 Announce Type: cross Abstract: Closed-loop traffic simulation requires agents that are both scalable and behaviorally realistic. Recent self-play reinforcement learning approaches d

Beyond Static Bias: Adaptive Multi-Fidelity Bandits with Improving Proxies

SafetyDGX agent

arXiv:2605.08558v1 Announce Type: new Abstract: As an extension of the classical multi-armed bandit problem, multi-fidelity multi-armed bandits (MF-MAB) enable individual arms to be evaluated using di

Bi-CoG: Bi-Consistency-Guided Self-Training for Vision-Language Models

SafetyDGX agent

arXiv:2510.20477v2 Announce Type: replace Abstract: Exploiting unlabeled data through semi-supervised learning (SSL) or leveraging pre-trained models via fine-tuning are two prevailing paradigms for a

Bias by Necessity: Impossibility Theorems for Sequential Processing with Convergent AI and Human Validation

SafetyDGX agent

arXiv:2605.08716v1 Announce Type: new Abstract: Are certain cognitive biases mathematically inevitable consequences of sequential information processing? We prove that primacy effects, anchoring, and

Big AI is accelerating the metacrisis: What can we do?

SafetyDGX agent

arXiv:2512.24863v2 Announce Type: replace-cross Abstract: The world is in the grip of ecological, meaning, and language crises that are converging into a metacrisis. Big AI is accelerating them all. L

Biological Plausibility and Representational Alignment of Feedback Alignment in Convolutional Networks

SafetyDGX agent

arXiv:2605.08564v1 Announce Type: new Abstract: The feedback alignment (FA) algorithm offers a biologically plausible alternative to backpropagation (BP) for training neural networks yet notably fails

Block-Wise Differentiable Sinkhorn Attention: Tail-Refinement Gradients with a Gap-Aware Dustbin Bridge

SafetyDGX agent

arXiv:2605.08123v1 Announce Type: cross Abstract: We study long-context balanced entropic optimal transport (OT) attention on TPU hardware through a stopped-base, fixed-depth tail-refinement surrogate

BONSAI: Bayesian Optimization with Natural Simplicity and Interpretability

SafetyDGX agent

arXiv:2602.07144v2 Announce Type: replace-cross Abstract: Bayesian optimization (BO) is a popular technique for sample-efficient optimization of black-box functions. In many applications, the paramete

Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization

SafetyDGX agent

arXiv:2605.10764v1 Announce Type: cross Abstract: Recent studies show that gradient-based universal image jailbreaks on vision-language models (VLMs) exhibit little or no cross-model transferability,

Breaking. Sam Altman himself finally confirms what I was the first to point out publicly: his indirect equity stake in OpenAI. … which he fa…

SafetyDGX agent

Breaking. Sam Altman himself finally confirms what I was the first to point out publicly: his indirect equity stake in OpenAI. … which he failed to acknowledge when asked about his financial interest

Breaking the Grid: Distance-Guided Reinforcement Learning in Large Discrete Action Spaces

SafetyDGX agent

arXiv:2602.08616v2 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) is increasingly applied to large-scale decision-making problems like logistics, scheduling, and recommender system

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents

SafetyDGX agent

arXiv:2605.08721v1 Announce Type: new Abstract: While Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for closed-ended tasks, extending it to open-ended social language game

BRIDGE: Building Representations In Domain Guided Program Synthesis

SafetyDGX agent

arXiv:2511.21104v3 Announce Type: replace Abstract: Large language models can generate plausible code, but remain brittle for formal verification in proof assistants such as Lean. A central scalabilit

BROS: Bias-Corrected Randomized Subspaces for Memory-Efficient Single-Loop Bilevel Optimization

SafetyDGX agent

arXiv:2605.10288v1 Announce Type: new Abstract: Stochastic bilevel optimization (SBO) has become a standard framework for hyperparameter learning, data reweighting, representation learning, and data-m

Budget-Efficient Automatic Algorithm Design via Code Graph

SafetyDGX agent

arXiv:2605.10598v1 Announce Type: new Abstract: Large language models (LLMs) have emerged as powerful tools for automatic algorithm design (AAD). However, existing pipelines remain inefficient. They o

CalBench: Evaluating Coordination-Privacy Trade-offs in Multi-Agent LLMs

SafetyDGX agent

arXiv:2605.09823v1 Announce Type: cross Abstract: We introduce CalBench, a controlled evaluation environment for studying multi-agent coordination through calendar scheduling. In CalBench, N agents ea

CAMAL: Improving Attention Alignment and Faithfulness with Segmentation Masks

SafetyDGX agent

arXiv:2605.08325v1 Announce Type: cross Abstract: Many vision datasets now provide segmentation masks in addition to annotated images to support a wide range of tasks. In this work, we propose Class A

Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction

SafetyDGX agent

arXiv:2512.18880v2 Announce Type: replace-cross Abstract: Accurate estimation of item (question or task) difficulty is critical for educational assessment but suffers from the cold start problem. Whil

Can Revealed Preferences Clarify LLM Alignment and Steering?

SafetyDGX agent

arXiv:2605.08556v1 Announce Type: new Abstract: LLMs are increasingly used to make or support high-stakes decisions under uncertainty, where alignment depends not only on factual accuracy but on how m

CapCLIP: A Vision-Language Representation Alignment Approach for Wireless Capsule Endoscopy Analysis

SafetyDGX agent

arXiv:2605.08493v1 Announce Type: new Abstract: Wireless capsule endoscopy (WCE) enables non-invasive visual assessment of the small bowel, but its clinical utility is constrained by the large volume

CARL: Criticality-Aware Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2512.04949v3 Announce Type: replace-cross Abstract: Agents capable of accomplishing complex tasks through multiple interactions with the environment have emerged as a popular research direction.

Causal Explanations from the Geometric Properties of ReLU Neural Networks

SafetyDGX agent

arXiv:2605.10396v1 Announce Type: new Abstract: Neural networks have proved an effective means of learning control policies for autonomous systems, but these learned policies are difficult to understa

CFSPMNet: Cross-subject Fourier-guided Spatial-Patch Mamba Network for EEG Motor Imagery Decoding in Stroke Patients

SafetyDGX agent

arXiv:2605.10111v1 Announce Type: cross Abstract: Motor imagery electroencephalography (MI-EEG) decoding offers a non-invasive route for post-stroke rehabilitation, but cross-patient use remains diffi

Change My View? The Dynamics of Persuasion and Polarization in Online Discourse

SafetyDGX agent

arXiv:2605.08383v1 Announce Type: new Abstract: Philosophical accounts of persuasion often assume that shared evidence and rational argumentation should lead to a convergence of views between peers, y

Code Mixologist : A Practitioner's Guide to Building Code-Mixed LLMs

SafetyDGX agent

arXiv:2602.11181v2 Announce Type: replace Abstract: Code-mixing and code-switching (CSW) remain challenging phenomena for large language models (LLMs). Despite recent advances in multilingual modeling

CoLA-Flow Policy: Temporally Coherent Imitation Learning via Continuous Latent Action Flow Matching for Robotic Manipulation

SafetyDGX agent

arXiv:2601.23087v3 Announce Type: replace Abstract: Learning long-horizon robotic manipulation requires jointly achieving expressive behavior modeling, real-time inference, and stable execution, which

Composing Policy Gradients and Prompt Optimization for Language Model Programs

SafetyDGX agent

arXiv:2508.04660v2 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) has proven to be an effective tool for post-training language models (LMs). However, AI systems are increa

Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision

SafetyDGX agent

arXiv:2509.14234v3 Announce Type: replace Abstract: Where do learning signals come from when there is no ground truth in post-training? We show that inference compute itself can serve as supervision.

Compute Where it Counts: Self Optimizing Language Models

SafetyDGX agent

arXiv:2605.10875v1 Announce Type: cross Abstract: Efficient LLM inference research has largely focused on reducing the cost of each decoding step (e.g., using quantization, pruning, or sparse attentio

Conformity Generates Collective Misalignment in AI Agents Societies

SafetyDGX agent

arXiv:2605.10721v1 Announce Type: cross Abstract: Artificial intelligence safety research focuses on aligning individual language models with human values, yet deployed AI systems increasingly operate

Consensus Sampling for Safer Generative AI

SafetyDGX agent

arXiv:2511.09493v2 Announce Type: replace Abstract: Motivated by undetectable risks in generative AI, we study a general robust aggregation problem: how to aggregate several probability distributions

Constraint-Aware Diffusion Priors for High-Fidelity and Versatile Quadruped Locomotion

SafetyDGX agent

arXiv:2605.08804v1 Announce Type: new Abstract: Reinforcement learning combined with imitation learning has significantly advanced biomimetic quadrupedal locomotion. However, scaling these frameworks

Constraint-Aware Reinforcement Learning via Adaptive Action Scaling

SafetyDGX agent

arXiv:2510.11491v3 Announce Type: replace-cross Abstract: Safe reinforcement learning (RL) seeks to mitigate unsafe behaviors that arise from exploration during training by reducing constraint violati

Containment Verification: AI Safety Guarantees Independent of Alignment

SafetyDGX agent

arXiv:2605.09045v1 Announce Type: new Abstract: Agentic frameworks are the software layer through which AI agents act in the world. Existing safety methods intervene on the model and therefore remain

Continuity Laws for Sequential Models

SafetyDGX agent

arXiv:2605.08539v1 Announce Type: cross Abstract: Inductive biases influence the behavior and performance of sequential models. In this work, we study an underexplored inductive bias in sequential mod

Control-Augmented Autoregressive Diffusion for Data Assimilation

SafetyDGX agent

arXiv:2510.06637v3 Announce Type: replace-cross Abstract: Despite advances in test-time scaling and diffusion finetuning, guidance for Auto-Regressive Diffusion Models (ARDMs) remains underexplored. W

Core-Halo Decomposition: Decentralizing Large-Scale Fixed-Point Problems

SafetyDGX agent

arXiv:2605.08681v1 Announce Type: cross Abstract: We study solving large-scale fixed-point equation (x^star=ar F(x^star)) with decomposition. Standard strict decomposition assigns each agent a disjoin

Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation

SafetyDGX agent

arXiv:2605.09253v1 Announce Type: cross Abstract: While recent work in Reinforcement Learning with Verifiable Rewards (RLVR) has shown that a small subset of critical tokens disproportionately drives

Crosslingual On-Policy Self-Distillation for Multilingual Reasoning

SafetyDGX agent

arXiv:2605.09548v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable progress in mathematical reasoning, but this ability is not equally accessible across languages. E

Crowding Out The Noise: Algorithmic Collective Action Under Differential Privacy

SafetyDGX agent

arXiv:2505.05707v2 Announce Type: replace Abstract: The integration of AI into daily life has generated considerable attention and excitement, while also raising concerns about automating algorithmic

DAP: Doppler-aware Point Network for Heterogeneous mmWave Action Recognition

SafetyDGX agent

arXiv:2605.09604v1 Announce Type: new Abstract: Millimeter-wave (mmWave) radar provides privacy-preserving sensing and is valuable for human action recognition (HAR). Existing mmWave point cloud datas

← Previous
1…147148149150151…212
Next →