AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,435 results
Safety

Toward an AI-Powered Computational Testbed for Workforce Policy

DGX agent

arXiv:2605.19064v1 Announce Type: cross Abstract: Workforce transformations are difficult to forecast and costly to mismanage. In particular, the integration of artificial intelligence into knowledge

safetyarxiv-cs-ai
20 May 2026
Safety
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Towards Distillation Guarantees under Algorithmic Alignment for Combinatorial Optimization

DGX agent

arXiv:2605.20074v1 Announce Type: new Abstract: Distillation transfers knowledge from a large model trained on broad data to a smaller, more efficient model suitable for deployment. In structured pred

safetyarxiv-cs-lg
20 May 2026
Safety

Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On

DGX agent

arXiv:2605.19035v1 Announce Type: new Abstract: The rapid advancement of Large Language Models has given rise to autonomous LLM-based agents capable of complex reasoning and execution. As these agents

safetyarxiv-cs-ai
20 May 2026
Safety

TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents

DGX agent

arXiv:2602.11767v3 Announce Type: replace Abstract: Advances in large language models (LLMs) are driving a shift toward using reinforcement learning (RL) to train agents from iterative, multi-turn int

safetyarxiv-cs-ai
20 May 2026
Safety

Universal Skeleton Understanding via Differentiable Rendering and MLLMs

DGX agent

arXiv:2603.18003v4 Announce Type: replace Abstract: Multimodal large language models (MLLMs) exhibit strong visual-language reasoning, yet cannot process structured, non-visual data such as human skel

safetyarxiv-cs-cv
20 May 2026
Safety

When Critics Disagree: Adaptive Reward Poisoning Attacks in RIS-Aided Wireless Control System

DGX agent

arXiv:2605.20037v1 Announce Type: cross Abstract: Reward-poisoning attacks present a significant risk to learning-based wireless control systems. Given this, we propose a Disagreement-Guided Reward Po

safetyarxiv-cs-ai
20 May 2026
Safety

When Preference Labels Fall Short: Aligning Diffusion Models from Real Data

DGX agent

arXiv:2605.19839v1 Announce Type: new Abstract: Preference alignment aims to guide generative models by learning from comparisons between preferred and non-preferred samples. In practice, most existin

safetyarxiv-cs-cv
20 May 2026
Safety

When Tabular Foundation Models Meet Strategic Tabular Data: A Prior Alignment Approach

DGX agent

arXiv:2605.19662v1 Announce Type: new Abstract: Tabular foundation models based on pretrained prior-data fitted networks~(PFNs) have shown strong generalization on diverse tabular tasks, but they are

safetyarxiv-cs-ai
20 May 2026
Safety

When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR

DGX agent

arXiv:2605.19425v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become the dominant paradigm for advanced reasoning in Large Language Models (LLMs), but rol

safetyarxiv-cs-ai
20 May 2026
Safety

Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making

DGX agent

arXiv:2602.07008v2 Announce Type: replace Abstract: Reliable models should not only predict correctly, but also justify decisions with acceptable evidence. Yet conventional supervised learning typical

safetyarxiv-cs-cv
20 May 2026
Safety

Worst-Group Equalized Odds Regularization for Multi-Attribute Fair Medical Image Classification

DGX agent

arXiv:2605.19214v1 Announce Type: cross Abstract: Diagnostic performance in medical AI varies systematically across demographic groups, yet subgroup AUC can mask clinically important disparities. At a

safetyarxiv-cs-cv
20 May 2026
Safety

A Fourier perspective on the learning dynamics of neural networks: from sample complexities to mechanistic insights

DGX agent

arXiv:2605.16913v1 Announce Type: cross Abstract: Neural networks trained with gradient-based methods exhibit a strong simplicity bias: they learn simpler statistical features of their data before mov

safetyarxiv-cs-lg
19 May 2026
Safety

A Simplex Witness Certificate for Constant Collapse in Variational Autoencoders

DGX agent

arXiv:2605.18224v1 Announce Type: cross Abstract: This note studies exact constant collapse in variational autoencoders, where the encoder mean becomes independent of the input. The goal is to make th

safetyarxiv-cs-ai
19 May 2026
Safety

A Visual Reinforcement Learning-Based Separate Primitive Policy for Peg-in-Hole Tasks

DGX agent

arXiv:2504.14820v2 Announce Type: replace Abstract: For peg-in-hole tasks, humans rely on binocular visual perception to locate the peg above the hole surface and then proceed with insertion. This pap

safetyarxiv-cs-ro
19 May 2026
Safety

Actionable World Representation

DGX agent

arXiv:2605.18743v1 Announce Type: new Abstract: Inspired by the emergent behaviors in large language models that generalized human intelligence, the research community is pursuing similar emergent cap

safetyarxiv-cs-ai
19 May 2026
Safety

Adaptive Control in Autonomous Driving via Real-Time Recurrent RL

DGX agent

arXiv:2602.02236v4 Announce Type: replace-cross Abstract: We study online fine-tuning of pretrained control policies for autonomous driving using Real-Time Recurrent Reinforcement Learning (RTRRL), a

safetyarxiv-cs-lg
19 May 2026
Safety

Adaptive Experimentation for Censored Survival Outcomes

DGX agent

arXiv:2605.18459v1 Announce Type: new Abstract: Adaptive experimentation enables efficient estimation of causal effects, but existing methods are not designed for survival data with censoring, where e

safetyarxiv-cs-lg
19 May 2026
Safety

Adaptive Generate-Rank-Verify: Inference-Time Search with Costly Verification

DGX agent

arXiv:2605.17609v1 Announce Type: new Abstract: Many inference-time language-model pipelines combine a cheap reward signal with an expensive verifier, such as exact answer checking in mathematical rea

safetyarxiv-cs-lg
19 May 2026
Safety

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment

DGX agent

arXiv:2605.17517v1 Announce Type: new Abstract: Recent advances in Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation. However, the visual represe

safetyarxiv-cs-ro
19 May 2026
Safety

Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces

DGX agent

arXiv:2605.17698v1 Announce Type: new Abstract: The deployment of Large Language Models (LLMs) as autonomous economic agents introduces systemic risks that extend beyond individual capability failures

safetyarxiv-cs-lg
19 May 2026
Safety

AI Agents May Always Fall for Prompt Injections

DGX agent

arXiv:2605.17634v1 Announce Type: cross Abstract: Prompt injection is the most critical vulnerability in deployed AI agents. Despite recent progress, we show that the prevailing defense paradigm (data

safetyarxiv-cs-cl
19 May 2026
Safety

AI of the People, by the People, for the People: A Social Choice Approach to Collective Control of Artificial Intelligence

DGX agent

arXiv:2605.16291v1 Announce Type: cross Abstract: With the growing adoption of AI systems, reasoning about how society can exert control over AI becomes an increasingly urgent problem. Existing work o

safetyarxiv-cs-ai
19 May 2026
Safety

AIM: Adversarial Information Masking for Faithfulness Evaluation of Saliency Maps

DGX agent

arXiv:2605.16905v1 Announce Type: cross Abstract: Post-hoc saliency methods are widely used to interpret deep neural networks, but their faithfulness is difficult to evaluate reliably. Existing evalua

safetyarxiv-cs-cv
19 May 2026
Safety

Algorithmic Cultivation: How Social Media Feeds Shape User Language

DGX agent

arXiv:2605.17010v1 Announce Type: cross Abstract: Algorithmic feeds have become primary environments for encountering information online, yet while they shape what people see, less is known about how

safetyarxiv-cs-ai
19 May 2026
Local Ai

ALIGN: A Vision-Language Framework for High-Accuracy Accident Location Inference through Geo-Spatial Neural Reasoning

DGX agent

arXiv:2511.06316v3 Announce Type: replace Abstract: In low- and middle-income countries, public safety and urban planning initiatives frequently face a critical shortage of accurate, location-specific

local-aiarxiv-cs-ai
19 May 2026
Safety

Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework

DGX agent

arXiv:2605.16516v1 Announce Type: cross Abstract: Long-term interaction with LLM-based systems may produce alignment drift: a gradual process in which system outputs become less constrained by the use

safetyarxiv-cs-ai
19 May 2026
Safety

AMATA: Adaptive Multi-Agent Trajectory Alignment for Knowledge-Intensive Question Answering

DGX agent

arXiv:2605.17352v1 Announce Type: new Abstract: Despite substantial advances in large language models (LLMs), generating factually consistent responses for knowledge-intensive question answering remai

safetyarxiv-cs-cl
19 May 2026
Safety

AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment

DGX agent

arXiv:2605.18529v1 Announce Type: new Abstract: The alignment of Large Language Models (LLMs) for complex reasoning heavily relies on Reinforcement Learning with Verifiable Rewards (RLVR). However, st

safetyarxiv-cs-ai
19 May 2026
Safety

An Assessment of Human vs. Model Uncertainty in Soft-Label Learning and Calibration

DGX agent

arXiv:2605.18648v1 Announce Type: cross Abstract: Central to human-aligned AI is understanding the benefits of human-elicited labels over synthetic alternatives. While human soft-labels improve calibr

safetyarxiv-cs-ai
19 May 2026
Safety

An Efficient Streaming Video Understanding Framework with Agentic Control

DGX agent

arXiv:2605.17921v1 Announce Type: new Abstract: Streaming video requires handling dynamic information density under strict latency budgets. Yet, existing methods typically employ static strategies, su

safetyarxiv-cs-cv
19 May 2026
Safety

An Empirical Study of Privacy Leakage Chains via Prompt Injection in Black-Box Chatbot Environments

DGX agent

arXiv:2605.18133v1 Announce Type: cross Abstract: LLM-based chatbot agents increasingly process user requests by combining natural-language reasoning with external tools such as web browsing. These ca

safetyarxiv-cs-ai
19 May 2026
Safety

AnchorDiff: Topology-Aware Masked Diffusion with Confidence-based Rewriting for Radiology Report Generation

DGX agent

arXiv:2605.17071v1 Announce Type: new Abstract: Radiology report generation (RRG) aims to automatically produce clinically accurate textual reports from medical images. Existing methods predominantly

safetyarxiv-cs-ai
19 May 2026
Safety

Anytime and Difficulty-Adaptive PAC-Bayes for Constrained Density-Ratio Network with Continual Learning Guarantees

DGX agent

arXiv:2605.17212v1 Announce Type: new Abstract: A unified framework for learning under covariate shift is presented, in which a constrained density-ratio network approximates the Radon-Nikodym derivat

safetyarxiv-cs-lg
19 May 2026
Safety

Are Multimodal LLMs Ready for Surveillance? A Reality Check on Zero-Shot Anomaly Detection in the Wild

DGX agent

arXiv:2603.04727v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have demonstrated impressive general competence in video understanding, yet their reliability for rea

safetyarxiv-cs-ai
19 May 2026
Safety

ARROW: Augmented Replay for RObust World models

DGX agent

arXiv:2603.11395v2 Announce Type: replace-cross Abstract: Continual reinforcement learning challenges agents to acquire new skills while retaining previously learned ones with the goal of improving pe

safetyarxiv-cs-ai
19 May 2026
Safety

Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews Large Language Model Decision-Making

DGX agent

arXiv:2605.17228v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in high-stakes domains such as clinical decision support and medical documentation. However, the

safetyarxiv-cs-cl
19 May 2026
Safety

AURORA: Contextual Orthogonalization for Geometric Representation Learning in Healthcare Foundation Models

DGX agent

arXiv:2605.17765v1 Announce Type: new Abstract: Recent healthcare foundation models have achieved strong predictive performance through large scale self supervised learning, yet their latent represent

safetyarxiv-cs-lg
19 May 2026
Safety

Automatic Generation of High-Performance RL Environments

DGX agent

arXiv:2603.12145v2 Announce Type: replace-cross Abstract: Translating complex reinforcement learning (RL) environments into high-performance implementations has traditionally required months of specia

safetyarxiv-cs-ai
19 May 2026
Safety

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment

DGX agent

arXiv:2605.17602v1 Announce Type: new Abstract: Aligning Text-to-Image (T2I) generation models with human preferences increasingly relies on image reward models that score or rank generated images acc

safetyarxiv-cs-ai
19 May 2026
Safety

Avoiding Structural Failure Modes in Tabular Fair SSL: Online Primal-Dual Allocation under Confidence Gating

DGX agent

arXiv:2605.16446v1 Announce Type: cross Abstract: Semi-supervised learning (SSL) enables prediction with limited labels, but high-stakes tabular applications (medical, credit, recidivism) require stat

safetyarxiv-cs-ai
19 May 2026
Safety

Benchmarking transferability of SSL pretraining to same and different modality segmentation tasks

DGX agent

arXiv:2605.18491v1 Announce Type: new Abstract: Methods: Nine SSL methods spanning four pretext-task families were pretrained from scratch using the same 10{,}412 3D CT scans (1.89~M 2D axial slices)

safetyarxiv-cs-cv
19 May 2026
Safety

Beyond Compliance: How AI Could Help Creative Writers by Refusing Them

DGX agent

arXiv:2605.16272v1 Announce Type: cross Abstract: Mainstream creativity support design prioritizes compliant AI for seamless writing interactions, but concerns over inappropriate AI reliance highlight

safetyarxiv-cs-ai
19 May 2026
Safety

Beyond Policy Optimization: A Data Curation Flywheel for Sparse-Reward Long-Horizon Planning

DGX agent

arXiv:2508.03018v2 Announce Type: replace Abstract: Large Language Reasoning Models have demonstrated remarkable success on static tasks, yet their application to multi-round agentic planning in inter

safetyarxiv-cs-ai
19 May 2026
Safety

Beyond RLHF: A Unified Theoretical Framework of Alignment

DGX agent

arXiv:2506.01523v2 Announce Type: replace Abstract: Alignment via reinforcement learning from human feedback (RLHF) has become the dominant paradigm for controlling the quality of outputs from large l

safetyarxiv-cs-lg
19 May 2026
Safety

Beyond Scaling: Agents Are Heading to the Edge

DGX agent

arXiv:2605.18535v1 Announce Type: new Abstract: The bottleneck of useful agentic intelligence has shifted from compressing world knowledge into a single model to executing a coordinated system. This p

safetyarxiv-cs-lg
19 May 2026
Safety

Beyond the Final Actor: Modeling the Dual Roles of Creator and Editor for Fine-Grained LLM-Generated Text Detection

DGX agent

arXiv:2604.04932v3 Announce Type: replace Abstract: The misuse of large language models (LLMs) requires precise detection of synthetic text. Existing works mainly follow binary or ternary classificati

safetyarxiv-cs-cl
19 May 2026
Safety

Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational Speech

DGX agent

arXiv:2605.17652v1 Announce Type: new Abstract: There are not enough established benchmarks for the task fo speech summarization. Creating new benchmarks demands human annotation, as LLMs could embed

safetyarxiv-cs-cl
19 May 2026
Safety

BIDO: A Biometric Identity Online Authentication Framework

DGX agent

arXiv:2605.16908v1 Announce Type: cross Abstract: Security systems demand continuous, cryptograph- ically robust identity verification without requiring subjects to carry physical tokens, smart cards,

safetyarxiv-cs-cv
19 May 2026
← Previous
1…171172173174175…260
Next →