AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,485 results
19 May 2026

AI Agents May Always Fall for Prompt Injections

SafetyDGX agent

arXiv:2605.17634v1 Announce Type: cross Abstract: Prompt injection is the most critical vulnerability in deployed AI agents. Despite recent progress, we show that the prevailing defense paradigm (data

AI of the People, by the People, for the People: A Social Choice Approach to Collective Control of Artificial Intelligence

SafetyDGX agent

arXiv:2605.16291v1 Announce Type: cross Abstract: With the growing adoption of AI systems, reasoning about how society can exert control over AI becomes an increasingly urgent problem. Existing work o

AIM: Adversarial Information Masking for Faithfulness Evaluation of Saliency Maps

SafetyDGX agent

arXiv:2605.16905v1 Announce Type: cross Abstract: Post-hoc saliency methods are widely used to interpret deep neural networks, but their faithfulness is difficult to evaluate reliably. Existing evalua

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Algorithmic Cultivation: How Social Media Feeds Shape User Language

SafetyDGX agent

arXiv:2605.17010v1 Announce Type: cross Abstract: Algorithmic feeds have become primary environments for encountering information online, yet while they shape what people see, less is known about how

ALIGN: A Vision-Language Framework for High-Accuracy Accident Location Inference through Geo-Spatial Neural Reasoning

Local AiDGX agent

arXiv:2511.06316v3 Announce Type: replace Abstract: In low- and middle-income countries, public safety and urban planning initiatives frequently face a critical shortage of accurate, location-specific

Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework

SafetyDGX agent

arXiv:2605.16516v1 Announce Type: cross Abstract: Long-term interaction with LLM-based systems may produce alignment drift: a gradual process in which system outputs become less constrained by the use

AMATA: Adaptive Multi-Agent Trajectory Alignment for Knowledge-Intensive Question Answering

SafetyDGX agent

arXiv:2605.17352v1 Announce Type: new Abstract: Despite substantial advances in large language models (LLMs), generating factually consistent responses for knowledge-intensive question answering remai

AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment

SafetyDGX agent

arXiv:2605.18529v1 Announce Type: new Abstract: The alignment of Large Language Models (LLMs) for complex reasoning heavily relies on Reinforcement Learning with Verifiable Rewards (RLVR). However, st

An Assessment of Human vs. Model Uncertainty in Soft-Label Learning and Calibration

SafetyDGX agent

arXiv:2605.18648v1 Announce Type: cross Abstract: Central to human-aligned AI is understanding the benefits of human-elicited labels over synthetic alternatives. While human soft-labels improve calibr

An Efficient Streaming Video Understanding Framework with Agentic Control

SafetyDGX agent

arXiv:2605.17921v1 Announce Type: new Abstract: Streaming video requires handling dynamic information density under strict latency budgets. Yet, existing methods typically employ static strategies, su

An Empirical Study of Privacy Leakage Chains via Prompt Injection in Black-Box Chatbot Environments

SafetyDGX agent

arXiv:2605.18133v1 Announce Type: cross Abstract: LLM-based chatbot agents increasingly process user requests by combining natural-language reasoning with external tools such as web browsing. These ca

AnchorDiff: Topology-Aware Masked Diffusion with Confidence-based Rewriting for Radiology Report Generation

SafetyDGX agent

arXiv:2605.17071v1 Announce Type: new Abstract: Radiology report generation (RRG) aims to automatically produce clinically accurate textual reports from medical images. Existing methods predominantly

Anytime and Difficulty-Adaptive PAC-Bayes for Constrained Density-Ratio Network with Continual Learning Guarantees

SafetyDGX agent

arXiv:2605.17212v1 Announce Type: new Abstract: A unified framework for learning under covariate shift is presented, in which a constrained density-ratio network approximates the Radon-Nikodym derivat

Are Multimodal LLMs Ready for Surveillance? A Reality Check on Zero-Shot Anomaly Detection in the Wild

SafetyDGX agent

arXiv:2603.04727v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have demonstrated impressive general competence in video understanding, yet their reliability for rea

ARROW: Augmented Replay for RObust World models

SafetyDGX agent

arXiv:2603.11395v2 Announce Type: replace-cross Abstract: Continual reinforcement learning challenges agents to acquire new skills while retaining previously learned ones with the goal of improving pe

Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews Large Language Model Decision-Making

SafetyDGX agent

arXiv:2605.17228v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in high-stakes domains such as clinical decision support and medical documentation. However, the

As a Soviet historian who has spent years writing about the extreme, repressive control Soviet Communism exercised over its unfortunate citi…

SafetyDGX agent

As a Soviet historian who has spent years writing about the extreme, repressive control Soviet Communism exercised over its unfortunate citizens, I find it really hard to bring a similar accusation ag

AURORA: Contextual Orthogonalization for Geometric Representation Learning in Healthcare Foundation Models

SafetyDGX agent

arXiv:2605.17765v1 Announce Type: new Abstract: Recent healthcare foundation models have achieved strong predictive performance through large scale self supervised learning, yet their latent represent

Automatic Generation of High-Performance RL Environments

SafetyDGX agent

arXiv:2603.12145v2 Announce Type: replace-cross Abstract: Translating complex reinforcement learning (RL) environments into high-performance implementations has traditionally required months of specia

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment

SafetyDGX agent

arXiv:2605.17602v1 Announce Type: new Abstract: Aligning Text-to-Image (T2I) generation models with human preferences increasingly relies on image reward models that score or rank generated images acc

Avoiding Structural Failure Modes in Tabular Fair SSL: Online Primal-Dual Allocation under Confidence Gating

SafetyDGX agent

arXiv:2605.16446v1 Announce Type: cross Abstract: Semi-supervised learning (SSL) enables prediction with limited labels, but high-stakes tabular applications (medical, credit, recidivism) require stat

Benchmarking transferability of SSL pretraining to same and different modality segmentation tasks

SafetyDGX agent

arXiv:2605.18491v1 Announce Type: new Abstract: Methods: Nine SSL methods spanning four pretext-task families were pretrained from scratch using the same 10{,}412 3D CT scans (1.89~M 2D axial slices)

Beyond Compliance: How AI Could Help Creative Writers by Refusing Them

SafetyDGX agent

arXiv:2605.16272v1 Announce Type: cross Abstract: Mainstream creativity support design prioritizes compliant AI for seamless writing interactions, but concerns over inappropriate AI reliance highlight

Beyond Policy Optimization: A Data Curation Flywheel for Sparse-Reward Long-Horizon Planning

SafetyDGX agent

arXiv:2508.03018v2 Announce Type: replace Abstract: Large Language Reasoning Models have demonstrated remarkable success on static tasks, yet their application to multi-round agentic planning in inter

Beyond RLHF: A Unified Theoretical Framework of Alignment

SafetyDGX agent

arXiv:2506.01523v2 Announce Type: replace Abstract: Alignment via reinforcement learning from human feedback (RLHF) has become the dominant paradigm for controlling the quality of outputs from large l

Beyond Scaling: Agents Are Heading to the Edge

SafetyDGX agent

arXiv:2605.18535v1 Announce Type: new Abstract: The bottleneck of useful agentic intelligence has shifted from compressing world knowledge into a single model to executing a coordinated system. This p

Beyond the Final Actor: Modeling the Dual Roles of Creator and Editor for Fine-Grained LLM-Generated Text Detection

SafetyDGX agent

arXiv:2604.04932v3 Announce Type: replace Abstract: The misuse of large language models (LLMs) requires precise detection of synthetic text. Existing works mainly follow binary or ternary classificati

Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational Speech

SafetyDGX agent

arXiv:2605.17652v1 Announce Type: new Abstract: There are not enough established benchmarks for the task fo speech summarization. Creating new benchmarks demands human annotation, as LLMs could embed

BIDO: A Biometric Identity Online Authentication Framework

SafetyDGX agent

arXiv:2605.16908v1 Announce Type: cross Abstract: Security systems demand continuous, cryptograph- ically robust identity verification without requiring subjects to carry physical tokens, smart cards,

Body-Grounded Perspective Formation and Conative Attunement in Artificial Agents

SafetyDGX agent

arXiv:2605.16728v1 Announce Type: new Abstract: This paper proposes a minimal architecture for body-grounded perspective formation in artificial agents. Extending prior work, the model introduces an i

Bridging the Intention-Expression Gap: Aligning Multi-Dimensional Preferences via Hierarchical Relevance Feedback in Text-to-Image Diffusion

SafetyDGX agent

arXiv:2603.14936v3 Announce Type: replace Abstract: Users often possess a clear visual intent but struggle to articulate it precisely in language. This intention-expression gap makes aligning generate

Building Reliable Arithmetic Multipliers Under NBTI Aging and Process Variations

SafetyDGX agent

arXiv:2605.18444v1 Announce Type: cross Abstract: Hardware aging poses a significant challenge for integrated circuits (ICs), leading to performance degradation and eventual failure. In this work, we

Canonical Regularisation of Wide Feature-Learning Neural Networks

SafetyDGX agent

arXiv:2605.18180v1 Announce Type: cross Abstract: Wide neural networks in the feature-learning regime drive modern deep learning, and yet they remain far less studied than their kernel-regime counterp

CatalyticMLLM: A Graph-Text Multimodal Large Language Model for Catalytic Materials

SafetyDGX agent

arXiv:2605.17254v1 Announce Type: new Abstract: Property prediction and inverse structural design of catalytic materials are typically modeled as two independent tasks: the former predicts target prop

ChartDesign: Towards LLM Designer of Data Visualization

SafetyDGX agent

arXiv:2605.16274v1 Announce Type: cross Abstract: Charts are the dominant medium for visualizing data, discovering patterns and trends, and communicating data driven insights, yet designing them still

ChemVA: Advancing Large Language Models on Chemical Reaction Diagrams Understanding

SafetyDGX agent

arXiv:2605.17214v1 Announce Type: new Abstract: While Large Language Models (LLMs) have revolutionized scientific text processing, they exhibit a significant capability gap when interpreting chemical

ClaHF: A Human Feedback-inspired Reinforcement Learning Framework for Improving Classification Tasks

SafetyDGX agent

arXiv:2605.17458v1 Announce Type: new Abstract: Text classification models are typically trained via supervised fine-tuning (SFT). However, SFT essentially performs behavior cloning from instance-wise

CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook

SafetyDGX agent

arXiv:2605.18257v1 Announce Type: cross Abstract: Multimodal representation alignment is pivotal for large language models and robotics. Traditional methods are often hindered by cross-modal informati

Collaborative Learning for Semi-Supervised LiDAR Semantic Segmentation

SafetyDGX agent

arXiv:2605.17135v1 Announce Type: new Abstract: Annotating large-scale LiDAR point clouds for 3D semantic segmentation is costly and time-consuming, which motivates the use of semi-supervised learning

COLSON: Controllable Learning-Based Social Navigation via Diffusion-Based Reinforcement Learning

SafetyDGX agent

arXiv:2503.13934v2 Announce Type: replace-cross Abstract: Mobile robot navigation in dynamic environments with pedestrian traffic is a key challenge in the development of autonomous mobile service rob

Confidence-Gated Robot Autonomy: When Does Uncertainty Actually Help?

SafetyDGX agent

arXiv:2605.18045v1 Announce Type: cross Abstract: Robotic systems often use predictive uncertainty to decide whether to act autonomously or defer to a fallback policy. In threshold-gated autonomy, unc

Consent Chain Degradation in Embodied Multi-Agent Systems: Bridging the Gap Between AI Agent Governance and Robot Ethics

SafetyDGX agent

arXiv:2605.16300v1 Announce Type: cross Abstract: Robotic systems are moving from isolated platforms to interconnected multi-agent ecosystems that operate in human environments. This shift raises a go

Contrastive Conceptor Activation Steering (COAST): Unlocking Vision-Language-Action Models through Hidden States

SafetyDGX agent

arXiv:2605.17144v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models leverage powerful perceptual priors from web-scale Vision-Language Model (VLM) pre-training, yet they remain surpr

Controlling Decision Drift in Multimodal Sentiment Analysis with Missing Modalities

SafetyDGX agent

arXiv:2605.16889v1 Announce Type: new Abstract: Multimodal sentiment analysis relies on textual, acoustic, and visual signals, yet real-world data often suffer from modality missing and quality imbala

Convex Dataset Valuation for Post-Training

SafetyDGX agent

arXiv:2605.16704v1 Announce Type: new Abstract: Improving LLM performance on downstream tasks sometimes requires leveraging auxiliary datasets during post-training. In practice, however, developers fa

COOPO: Cyclic Offline-Online Policy Optimization Algorithm

SafetyDGX agent

arXiv:2605.18675v1 Announce Type: cross Abstract: Offline reinforcement learning struggles with distributional shift and constrained performance due to static dataset limitations, while online RL dema

Cracks in the Foundation: A Civil Infrastructure Dataset to Challenge Vision Foundation Models

SafetyDGX agent

arXiv:2605.18413v1 Announce Type: new Abstract: Automated structural health monitoring is essential to prevent catastrophic infrastructure failures. Precise, pixel-level defect segmentation is needed

Cross-modal Affinity-aligned Multimodal Learning Analytics for Predicting Student Collaboration Satisfaction in Game-Based Learning

SafetyDGX agent

arXiv:2605.16806v1 Announce Type: cross Abstract: Collaborative game-based learning environments offer rich opportunities for small-group knowledge construction, yet automatically predicting student c

Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image Generation

SafetyDGX agent

arXiv:2605.17807v1 Announce Type: cross Abstract: Text-to-Image (T2I) generation has achieved remarkable progress in recent years. Meanwhile, reinforcement learning methods, particularly those based o

DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models

SafetyDGX agent

arXiv:2605.16342v1 Announce Type: cross Abstract: Diffusion large language models are a compelling alternative to autoregressive models, yet existing RL methods for diffusion treat all denoising steps

Dance Across Shifts: Forward-Facilitation Continual Test-Time Adaptation through Dynamic Style Bridging

SafetyDGX agent

arXiv:2605.18608v1 Announce Type: new Abstract: Continual Test-Time Adaptation (CTTA) aims to empower perception systems to handle dynamic distribution shifts encountered after deployment. Existing me

DARC: Disagreement-Aware Alignment via Risk-Constrained Decoding

SafetyDGX agent

arXiv:2603.08145v2 Announce Type: replace-cross Abstract: Preference-based alignment methods (e.g., RLHF, DPO) typically optimize a single scalar objective, implicitly averaging over heterogeneous hum

Decision-Aware Proximal Bridge Learning for Optimal Treatment Selection

SafetyDGX agent

arXiv:2605.16989v1 Announce Type: new Abstract: Individualized treatment selection with continuous actions requires accurate causal response estimation in decision-relevant regions, rather than unifor

Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation

SafetyDGX agent

arXiv:2605.16826v1 Announce Type: cross Abstract: Knowledge distillation is central to LLM post-training, yet its design space remains poorly understood, especially alongside reinforcement learning (R

Deep Reinforcement Learning Framework for Diversified Portfolio Management Across Global Equity Markets

SafetyDGX agent

arXiv:2605.17307v1 Announce Type: cross Abstract: This study develops and evaluates a deep reinforcement learning framework for dynamic portfolio allocation across global equity markets. The Soft Acto

Deep sequence models tend to memorize geometrically; it is unclear why

SafetyDGX agent

arXiv:2510.26745v3 Announce Type: replace-cross Abstract: Deep sequence models are said to store atomic facts predominantly in the form of associative memory: a brute-force lookup of co-occurring enti

DeTrack: A Benchmark and Altitude-Aware Dual World Model for Drone-embodied Tracking

Model ReleasesDGX agent

arXiv:2605.17451v1 Announce Type: new Abstract: Aerial object tracking has broad applications in public safety, emergency rescue, wildlife monitoring, and related fields. However, existing aerial trac

DEVIS-GRPO: Unleashing GRPO on Dynamic Extreme View Synthesis

SafetyDGX agent

arXiv:2605.16937v1 Announce Type: new Abstract: Trajectory-controlled video generation has become essential for controllable video generation. While current methods perform well under small-view camer

Dexora: Open-source VLA for High-DoF Bimanual Dexterity

SafetyDGX agent

arXiv:2605.18722v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently become a central direction in embodied AI, but current systems are restricted to either dual-gripper c

Differentiable Optimization Layers for Guaranteed Fairness in Deep Learning

SafetyDGX agent

arXiv:2605.17118v1 Announce Type: new Abstract: Differentiable optimization layers are traditionally integrated in predict-then-optimize frameworks where a neural model estimates parameters that subse

← Previous
1…158159160161162…242
Next →