AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,702 results
22 Apr 2026

RESFL: An Uncertainty-Aware Framework for Responsible Federated Learning by Balancing Privacy, Fairness and Utility

SafetyDGX agent

arXiv:2503.16251v2 Announce Type: replace-cross Abstract: Federated Learning (FL) has gained prominence in machine learning applications across critical domains by enabling collaborative model trainin

REVEAL: Multimodal Vision-Language Alignment of Retinal Morphometry and Clinical Risks for Incident AD and Dementia Prediction

SafetyDGX agent

arXiv:2604.18757v1 Announce Type: cross Abstract: The retina provides a unique, noninvasive window into Alzheimer's disease (AD) and dementia, capturing early structural changes through morphometric f

Rivian and Volkswagen Group's joint venture put Devin to work on testing and ticket triage across a software platform that will power up to …


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety
DGX agent

Rivian and Volkswagen Group's joint venture put Devin to work on testing and ticket triage across a software platform that will power up to 30M vehicles. Devin handles autonomous ticket triage in Slac

RL-ABC: Reinforcement Learning for Accelerator Beamline Control

SafetyDGX agent

arXiv:2604.19146v1 Announce Type: new Abstract: Particle accelerator beamline optimization is a high-dimensional control problem traditionally requiring significant expert intervention. We present RLA

Safety-Critical Contextual Control via Online Riemannian Optimization with World Models

SafetyDGX agent

arXiv:2604.19639v1 Announce Type: cross Abstract: Modern world models are becoming too complex to admit explicit dynamical descriptions. We study safety-critical contextual control, where a Planner mu

SEAT: Sparse Entity-Aware Tuning for Knowledge Adaptation while Preserving Epistemic Abstention

SafetyDGX agent

arXiv:2506.14387v3 Announce Type: replace Abstract: Adapting LLMs with new knowledge is increasingly important, but standard fine-tuning often erodes aligned epistemic abstention: the ability to ackno

Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring

SafetyDGX agent

arXiv:2604.18835v1 Announce Type: cross Abstract: We propose a scalable, multifactorial experimental framework that systematically probes LLM sensitivity to subtle semantic changes in pairwise documen

Sherpa.ai Privacy-Preserving Multi-Party Entity Alignment without Intersection Disclosure for Noisy Identifiers

SafetyDGX agent

arXiv:2604.19219v1 Announce Type: cross Abstract: Federated Learning (FL) enables collaborative model training among multiple parties without centralizing raw data. There are two main paradigms in FL:

Sources: Micron is pushing the US Congress to pass the 'MATCH Act', which would put new export restrictions on equipment its Chinese rivals use to make chips (Karen Freifeld/Reuters)

SafetyDGX agent

Karen Freifeld / Reuters: Sources: Micron is pushing the US Congress to pass the “MATCH Act”, which would put new export restrictions on equipment its Chinese rivals use to make chips — Micron Technol

SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model

SafetyDGX agent

arXiv:2604.19710v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models offer a promising autonomous driving paradigm for leveraging world knowledge and reasoning capabilities, especially

Stable-RAG: Mitigating Retrieval-Permutation-Induced Hallucinations in Retrieval-Augmented Generation

SafetyDGX agent

arXiv:2601.02993v4 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) has become a key paradigm for reducing factual hallucinations in Large Language Models (LLMs), yet little is kn

STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming

SafetyDGX agent

arXiv:2604.18976v1 Announce Type: new Abstract: While Large Language Models (LLMs) are widely used, they remain susceptible to jailbreak prompts that can elicit harmful or inappropriate responses. Thi

Symbolic Quantile Regression for the Interpretable Prediction of Conditional Quantiles

SafetyDGX agent

arXiv:2508.08080v2 Announce Type: replace Abstract: Symbolic Regression (SR) is a well-established framework for generating interpretable or white-box predictive models. Although SR has been successfu

Task-Adaptive Admittance Control for Human-Quadrotor Cooperative Load Transportation with Dynamic Cable-Length Regulation

SafetyDGX agent

arXiv:2604.18905v1 Announce Type: new Abstract: The collaboration between humans and robots is critical in many robotic applications, especially in those requiring physical human-robot interaction (pH

TEMPO: Scaling Test-time Training for Large Reasoning Models

SafetyDGX agent

arXiv:2604.19295v1 Announce Type: new Abstract: Test-time training (TTT) adapts model parameters on unlabeled test instances during inference time, which continuously extends capabilities beyond the r

Terrible

SafetyDGX agent

Terrible This is insane… The Virginia redistricting amendment on the ballot today is framed as a vote to 'restore fairness in the upcoming elections.' In reality, it turns a state that Kamala barely w

The Alignment Waltz: Jointly Training Agents to Collaborate for Safety

SafetyDGX agent

arXiv:2510.08240v2 Announce Type: replace Abstract: Harnessing the power of LLMs requires a delicate dance between being helpful and harmless. This creates a fundamental tension between two competing

The Data-Driven Censored Newsvendor Problem

SafetyDGX agent

arXiv:2412.01763v3 Announce Type: replace-cross Abstract: We study a censored variant of the data-driven newsvendor problem, where the decision-maker must select an ordering quantity that minimizes ex

The Essence of Balance for Self-Improving Agents in Vision-and-Language Navigation

SafetyDGX agent

arXiv:2604.19064v1 Announce Type: new Abstract: In vision-and-language navigation (VLN), self-improvement from policy-induced experience, using only standard VLN action supervision, critically depends

The PROPER Approach to Proactivity: Benchmarking and Advancing Knowledge Gap Navigation

SafetyDGX agent

arXiv:2601.09926v3 Announce Type: replace Abstract: Current approaches to proactive assistance move beyond the ask-and-respond paradigm by anticipating user needs. In practice, they either burden user

The signal is the ceiling: Measurement limits of LLM-predicted experience ratings from open-ended survey text

SafetyDGX agent

arXiv:2604.19645v1 Announce Type: new Abstract: An earlier paper (Hong, Potteiger, and Zapata 2026) established that an unoptimized GPT 4.1 prompt predicts fan-reported experience ratings within one p

The Triadic Loop: A Framework for Negotiating Alignment in AI Co-hosted Livestreaming

SafetyDGX agent

arXiv:2604.18850v1 Announce Type: cross Abstract: AI systems are increasingly embedded in multi-user social environments, yet most alignment frameworks conceptualize interaction as a dyadic relationsh

This actually happened. Smart second-graders know better. WATF. 🤯

SafetyDGX agent

Gary Marcus expresses skepticism or concern about an artificial intelligence claim or incident that he finds implausible, suggesting that even young children would recognize the flaw in the reasoning

Toward Clinically Acceptable Chest X-ray Report Generation: A Qualitative Retrospective Pilot Study of CXRMate-2

SafetyDGX agent

arXiv:2604.18967v1 Announce Type: new Abstract: Chest X-ray (CXR) radiology report generation (RRG) models have shown rapid progress, yet their clinical utility remains uncertain due to limited evalua

TRN-R1-Zero: Text-rich Network Reasoning via LLMs with Reinforcement Learning Only

SafetyDGX agent

arXiv:2604.19070v1 Announce Type: new Abstract: Zero-shot reasoning on text-rich networks (TRNs) remains a challenging frontier, as models must integrate textual semantics with relational structure wi

TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards

SafetyDGX agent

arXiv:2512.07761v3 Announce Type: replace Abstract: Large language models have seen widespread adoption, yet they remain vulnerable to multi-turn jailbreak attacks, threatening their safe deployment.

User Simulation in the Era of Generative AI: User Modeling, Synthetic Data Generation, and System Evaluation

SafetyDGX agent

arXiv:2501.04410v2 Announce Type: replace Abstract: User simulation is an emerging interdisciplinary topic with multiple critical applications in the era of Generative AI. It involves creating an inte

VIGIL: An Extensible System for Real-Time Detection and Mitigation of Cognitive Bias Triggers

SafetyDGX agent

arXiv:2604.03261v2 Announce Type: replace Abstract: The rise of generative AI is posing increasing risks to online information integrity and civic discourse. Most concretely, such risks can materialis

VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph

SafetyDGX agent

arXiv:2602.12735v2 Announce Type: replace-cross Abstract: Effectively retrieving, reasoning, and understanding multimodal information remains a critical challenge for agentic systems. Traditional Retr

Vision-Based Human Awareness Estimation for Enhanced Safety and Efficiency of AMRs in Industrial Warehouses

SafetyDGX agent

arXiv:2604.18627v1 Announce Type: new Abstract: Ensuring human safety is of paramount importance in warehouse environments that feature mixed traffic of human workers and autonomous mobile robots (AMR

Visual Adversarial Attack on Vision-Language Models for Autonomous Driving

SafetyDGX agent

arXiv:2411.18275v2 Announce Type: replace Abstract: Vision-language models (VLMs) have significantly advanced autonomous driving (AD) by enhancing reasoning capabilities. However, these models remain

VoteGCL: Enhancing Graph-based Recommendations with Majority-Voting LLM-Rerank Augmentation

SafetyDGX agent

arXiv:2507.21563v4 Announce Type: replace-cross Abstract: Recommendation systems often suffer from data sparsity caused by limited user-item interactions, which degrade their performance and amplify p

Weakly supervised framework for wildlife detection and counting in challenging Arctic environments: a case study on caribou (Rangifer tarandus)

SafetyDGX agent

arXiv:2601.18891v3 Announce Type: replace Abstract: Caribou across the Arctic has declined in recent decades, motivating scalable and accurate monitoring approaches to guide evidence-based conservatio

When Can We Trust Deep Neural Networks? Towards Reliable Industrial Deployment with an Interpretability Guide

SafetyDGX agent

arXiv:2604.19206v1 Announce Type: new Abstract: The deployment of AI systems in safety-critical domains, such as industrial defect inspection, autonomous driving, and medical diagnosis, is severely ha

21 Apr 2026

3 new ways Ads Advisor is making Google Ads safer and faster

SafetyDGX agent

Ads Advisor is an agentic conversational experience built with Gemini in Google Ads designed to help maximize performance based on business goals. The tool can help users get personalized answers, und

A Hamilton-Jacobi Reachability-Guided Search Framework for Efficient and Safe Indoor Planar Robot Navigation

SafetyDGX agent

arXiv:2604.17679v1 Announce Type: new Abstract: Autonomous navigation requires planning to reach a goal safely and efficiently in complex and potentially dynamic environments. Graph search-based algor

A High-Accuracy Optical Music Recognition Method Based on Bottleneck Residual Convolutions

SafetyDGX agent

arXiv:2604.16446v1 Announce Type: new Abstract: Optical Music Recognition (OMR) aims to convert printed or handwritten music score images into editable symbolic representations. This paper presents an

A Quasi-Experimental Developer Study of Security Training in LLM-Assisted Web Application Development

SafetyDGX agent

arXiv:2604.17763v1 Announce Type: cross Abstract: This paper presents a controlled quasi-experimental developer study examining whether a layer-based security training package is associated with impro

A Real-Time Bike-Pedestrian Safety System with Wide-Angle Perception and Evaluation Testbed for Urban Intersections

SafetyDGX agent

arXiv:2604.17046v1 Announce Type: new Abstract: Collisions between cyclists and pedestrians at urban intersections remain a persistent source of injuries, yet few systems attempt real-time warnings to

A Sensitivity Approach to Causal Inference Under Limited Overlap

SafetyDGX agent

arXiv:2511.22003v2 Announce Type: replace-cross Abstract: Limited overlap between treated and control groups is a key challenge in observational analysis. Standard approaches like trimming importance

A Text-To-Text Alignment Algorithm for Better Evaluation of Modern Speech Recognition Systems

SafetyDGX agent

arXiv:2509.24478v2 Announce Type: replace Abstract: Modern neural networks have greatly improved performance across speech recognition benchmarks. However, gains are often driven by frequent words wit

Adversarial Arena: Crowdsourcing Data Generation through Interactive Competition

SafetyDGX agent

arXiv:2604.17803v1 Announce Type: cross Abstract: Post-training Large Language Models requires diverse, high-quality data which is rare and costly to obtain, especially in low resource domains and for

Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations

SafetyDGX agent

arXiv:2510.16458v2 Announce Type: replace Abstract: Natural Language Inference (NLI) datasets often exhibit human label variation. To better understand these variations, explanation-based approaches a

Align Documents to Questions: Question-Oriented Document Rewriting for Retrieval-Augmented Generation

SafetyDGX agent

arXiv:2604.17325v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) enhances the factuality of Large Language Models (LLMs) by incorporating retrieved documents and/or generated conte

Aligning Backchannel and Dialogue Context Representations via Contrastive LLM Fine-Tuning

SafetyDGX agent

arXiv:2604.16622v1 Announce Type: new Abstract: Backchannels (e.g., `yeah', `mhm', and `right') are short, non-interruptive feedback signals whose lexical form and prosody jointly convey pragmatic mea

Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints

SafetyDGX agent

arXiv:2604.18489v1 Announce Type: cross Abstract: Large Language Models (LLMs) show promise in lyric-to-melody generation, but models trained with Supervised Fine-Tuning (SFT) often produce musically

Alignment Data Map for Efficient Preference Data Selection and Diagnosis

SafetyDGX agent

arXiv:2505.23114v3 Announce Type: replace Abstract: Human preference data is essential for aligning large language models (LLMs) with human values, but collecting such data is often costly and ineffic

An `Inverse' Experimental Framework to Estimate Market Efficiency

SafetyDGX agent

arXiv:2604.18130v1 Announce Type: new Abstract: Digital marketplaces processing billions of dollars annually represent critical infrastructure in sociotechnical ecosystems, yet their performance optim

Annotation-Assisted Learning of Treatment Policies From Multimodal Electronic Health Records

SafetyDGX agent

arXiv:2507.20993v3 Announce Type: replace Abstract: We study how to learn treatment policies from multimodal electronic health records (EHRs) that consist of tabular data and clinical text. These poli

APIs and limited releases for AI models are not a safety policy, they’re a business model (which is totally ok as long as you’re transparent…

SafetyDGX agent

APIs and limited releases for AI models are not a safety policy, they’re a business model (which is totally ok as long as you’re transparent about it). Especially on cyber-security, they give a false

Arch: An AI-Native Hardware Description Language for Register-Transfer Clocked Hardware Design

SafetyDGX agent

arXiv:2604.05983v2 Announce Type: replace-cross Abstract: We present Arch (AI-native Register-transfer Clocked Hardware), a hardware description language for micro-architecture specification and AI-as

ARCS: Autoregressive Circuit Synthesis with Topology-Aware Graph Attention and Spec Conditioning

SafetyDGX agent

arXiv:2603.29068v3 Announce Type: replace Abstract: This paper presents ARCS (Autoregressive Circuit Synthesis), a system for amortized analog circuit generation. ARCS produces complete, SPICE-simulat

Asset Harvester: Extracting 3D Assets from Autonomous Driving Logs for Simulation

SafetyDGX agent

arXiv:2604.18468v1 Announce Type: new Abstract: Closed-loop simulation is a core component of autonomous vehicle (AV) development, enabling scalable testing, training, and safety validation before rea

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs

SafetyDGX agent

arXiv:2511.02356v2 Announce Type: replace-cross Abstract: Despite extensive safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks. However, existing methods generally l

Attention-space Contrastive Guidance for Efficient Hallucination Mitigation in LVLMs

SafetyDGX agent

arXiv:2601.13707v2 Announce Type: replace Abstract: Hallucinations in large vision--language models (LVLMs) often arise when language priors dominate over visual evidence, leading to object misidentif

Audio-DeepThinker: Progressive Reasoning-Aware Reinforcement Learning for High-Quality Chain-of-Thought Emergence in Audio Language Models

SafetyDGX agent

arXiv:2604.18187v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) have made significant progress in audio understanding, yet they primarily operate as perception-and-answer systems

AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction

SafetyDGX agent

arXiv:2510.15339v3 Announce Type: replace Abstract: Building effective knowledge graphs (KGs) for Retrieval-Augmented Generation (RAG) is pivotal for advancing question answering (QA) systems. However

Autonomous Vehicle Collision Avoidance With Racing Parameterized Deep Reinforcement Learning

SafetyDGX agent

arXiv:2604.16702v1 Announce Type: new Abstract: Road traffic accidents are a leading cause of fatalities worldwide. In the US, human error causes 94% of crashes, resulting in excess of 7,000 pedestria

Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

SafetyDGX agent

arXiv:2604.18572v1 Announce Type: new Abstract: The Platonic Representation Hypothesis suggests that neural networks trained on different modalities (e.g., text and images) align and eventually conver

BASIS: Balanced Activation Sketching with Invariant Scalars for 'Ghost Backpropagation'

SafetyDGX agent

arXiv:2604.16324v1 Announce Type: new Abstract: The activation memory required for exact backpropagation scales linearly with network depth, context length, and feature dimensionality, forming an O(L

← Previous
1…186187188189190…212
Next →