AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,351 results
Safety

Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models

DGX agent

arXiv:2601.15220v2 Announce Type: replace Abstract: We identify a novel phenomenon in language models: benign fine-tuning of frontier models can lead to privacy collapse. We find that diverse, subtle

safetyarxiv-cs-cl
21 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring

DGX agent

arXiv:2512.12069v3 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) are vulnerable to a growing array of multimodal jailbreak attacks, necessitating defenses that are both g

safetyarxiv-cs-cl
21 Apr 2026
Safety

Test-Time Alignment via Hypothesis Reweighting

DGX agent

arXiv:2412.08812v2 Announce Type: replace Abstract: Reward models trained on aggregate preferences often fail to capture individual users' values, but existing adaptation methods such as fine-tuning o

safetyarxiv-cs-lg
21 Apr 2026
Safety

Language Models as Semantic Teachers: Post-Training Alignment for Medical Audio Understanding

DGX agent

arXiv:2512.04847v2 Announce Type: replace-cross Abstract: Pre-trained audio models excel at detecting acoustic patterns in auscultation sounds but often fail to grasp their clinical significance, limi

safetyarxiv-cs-ai
20 Apr 2026
Safety

Puppets or partners? Governing cyborg propaganda in the digital public square

DGX agent

arXiv:2602.13088v2 Announce Type: replace-cross Abstract: The distinction between genuine grassroots activism and automated influence operations is collapsing. While contemporary policy debates priori

safetyarxiv-cs-ai
20 Apr 2026
Safety

The AI industry insists they can manage the risks of superintelligence, but there are in fact zero widely agreed on or accepted solutions to…

DGX agent

The AI industry insists they can manage the risks of superintelligence, but there are in fact zero widely agreed on or accepted solutions to the problem of how one could even control something vastly

safetyconnor-leahy--x
20 Apr 2026
Safety

When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models

DGX agent

arXiv:2510.09689v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have been augmented with web search to overcome the limitations of the static knowledge boundary by accessing up-

safetyarxiv-cs-ai
20 Apr 2026
Safety

Blinded Multi-Rater Comparative Evaluation of a Large Language Model and Clinician-Authored Responses in CGM-Informed Diabetes Counseling

DGX agent

arXiv:2604.15124v1 Announce Type: new Abstract: Continuous glucose monitoring (CGM) is central to diabetes care, but explaining CGM patterns clearly and empathetically remains time-intensive. Evidence

safetyarxiv-cs-cl
17 Apr 2026
Safety

Inclusive Kitchen Design for Older Adults: Generative AI Visualizations to Support Mild Cognitive Impairment

DGX agent

arXiv:2604.13203v1 Announce Type: cross Abstract: Mild Cognitive Impairment (MCI) affects 15-20% of adults aged 65 and older, often making kitchen navigation and independent living difficult, particul

safetyarxiv-cs-ai
17 Apr 2026
Safety

Momentum-constrained Hybrid Heuristic Trajectory Optimization Framework with Residual-enhanced DRL for Visually Impaired Scenarios

DGX agent

arXiv:2604.14986v1 Announce Type: new Abstract: Safe and efficient assistive planning for visually impaired scenarios remains challenging, since existing methods struggle with multi-objective optimiza

safetyarxiv-cs-ro
17 Apr 2026
Safety

ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration

DGX agent

arXiv:2509.21823v2 Announce Type: replace Abstract: Reward is critical to the evaluation and training of large language models (LLMs). However, existing rule-based or model-based reward methods strugg

safetyarxiv-cs-ai
17 Apr 2026
Safety

Separation is Optimal for LQR under Intermittent Feedback

DGX agent

arXiv:2603.27833v2 Announce Type: cross Abstract: In this work, we first prove that the separation principle holds for communication-constrained LQR problems under i.i.d. zero-mean disturbances with a

safetyarxiv-cs-ro
17 Apr 2026
Safety

Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models

DGX agent

arXiv:2604.14629v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have shown remarkable capabilities in joint vision-language understanding, but their large scale poses significant challen

safetyarxiv-cs-cv
17 Apr 2026
Safety

TOPCELL: Topology Optimization of Standard Cell via LLMs

DGX agent

arXiv:2604.14237v1 Announce Type: new Abstract: Transistor topology optimization is a critical step in standard cell design, directly dictating diffusion sharing efficiency and downstream routability.

safetyarxiv-cs-lg
17 Apr 2026
Safety

Towards Scalable Lightweight GUI Agents via Multi-role Orchestration

DGX agent

arXiv:2604.13488v1 Announce Type: new Abstract: Autonomous Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) enable digital automation on end-user devices. Whil

safetyarxiv-cs-ai
17 Apr 2026
Safety

VoxSafeBench: Not Just What Is Said, but Who, How, and Where

DGX agent

arXiv:2604.14548v1 Announce Type: cross Abstract: As speech language models (SLMs) transition from personal devices into shared, multi-user environments, their responses must account for far more than

safetyarxiv-cs-lg
17 Apr 2026
Safety

What some leaders think of using universal income to mitigate AI-fueled layoffs: Musk calls it the 'best way', OpenAI's policy doc mentions a Public Wealth Fund (Siladitya Ray/Forbes)

DGX agent

Siladitya Ray / Forbes: What some leaders think of using universal income to mitigate AI-fueled layoffs: Musk calls it the “best way”, OpenAI's policy doc mentions a Public Wealth Fund — Topline — Elo

safetytechmeme
17 Apr 2026
Safety

Asymmetric-Loss-Guided Hybrid CNN-BiLSTM-Attention Model for Industrial RUL Prediction with Interpretable Failure Heatmaps

DGX agent

arXiv:2604.13459v1 Announce Type: new Abstract: Turbofan engine degradation under sustained operational stress necessitates robust prognostic systems capable of accurately estimating the Remaining Use

safetyarxiv-cs-lg
16 Apr 2026
Safety

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling

DGX agent

arXiv:2604.13561v1 Announce Type: new Abstract: Vision-language models trained with contrastive learning on paired medical images and reports show strong zero-shot diagnostic capabilities, yet the eff

safetyarxiv-cs-cv
16 Apr 2026
Safety

Context Sensitivity Improves Human-Machine Visual Alignment

DGX agent

arXiv:2604.13883v1 Announce Type: new Abstract: Modern machine learning models typically represent inputs as fixed points in a high-dimensional embedding space. While this approach has been proven pow

safetyarxiv-cs-cv
16 Apr 2026
Safety

RFK Jr. forces FDA to reconsider 12 unproven peptides after 2023 ban

DGX agent

In 2023, the FDA removed 19 peptides from the list of drugs that compounding pharmacies could produce, and in 2026 the FDA announced it will review whether to add back 7 of these peptides following pr

safetyars-technica
16 Apr 2026
Safety

Robust Verification of Controllers under State Uncertainty via Hamilton-Jacobi Reachability Analysis

DGX agent

arXiv:2511.14755v2 Announce Type: replace-cross Abstract: As perception-based controllers for autonomous systems become increasingly popular in the real world, it is important that we can formally ver

safetyarxiv-cs-lg
16 Apr 2026
Safety

VRAG-DFD: Verifiable Retrieval-Augmentation for MLLM-based Deepfake Detection

DGX agent

arXiv:2604.13660v1 Announce Type: new Abstract: In Deepfake Detection (DFD) tasks, researchers proposed two types of MLLM-based methods: complementary combination with small DFD detectors, or static f

safetyarxiv-cs-cv
16 Apr 2026
Safety

A Comparison of Reinforcement Learning and Optimal Control Methods for Path Planning

DGX agent

arXiv:2604.12628v1 Announce Type: cross Abstract: Path-planning for autonomous vehicles in threat-laden environments is a fundamental challenge. While traditional optimal control methods can find idea

safetyarxiv-cs-ro
15 Apr 2026
Safety

ARGen: Affect-Reinforced Generative Augmentation towards Vision-based Dynamic Emotion Perception

DGX agent

arXiv:2604.12255v1 Announce Type: cross Abstract: Dynamic facial expression recognition in the wild remains challenging due to data scarcity and long-tail distributions, which hinder models from effec

safetyarxiv-cs-ai
15 Apr 2026
Safety

LatentRefusal: Latent-Signal Refusal for Unanswerable Text-to-SQL Queries

DGX agent

arXiv:2601.10398v3 Announce Type: replace Abstract: In LLM-based text-to-SQL systems, unanswerable and underspecified user queries may generate not only incorrect text but also executable programs tha

safetyarxiv-cs-ai
15 Apr 2026
Safety

Risk-Calibrated Learning: Minimizing Fatal Errors in Medical AI

DGX agent

arXiv:2604.12693v1 Announce Type: new Abstract: Deep learning models often achieve expert-level accuracy in medical image classification but suffer from a critical flaw: semantic incoherence. These hi

safetyarxiv-cs-cv
15 Apr 2026
Safety

Robust Optimization for Mitigating Reward Hacking with Correlated Proxies

DGX agent

arXiv:2604.12086v1 Announce Type: new Abstract: Designing robust reinforcement learning (RL) agents in the presence of imperfect reward signals remains a core challenge. In practice, agents are often

safetyarxiv-cs-lg
15 Apr 2026
Safety

TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs

DGX agent

arXiv:2604.12232v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed across diverse domains, yet their vulnerability to jailbreak attacks, where adversarial inputs

safetyarxiv-cs-ai
15 Apr 2026
Safety

Understanding and Improving Continuous Adversarial Training for LLMs via In-context Learning Theory

DGX agent

arXiv:2604.12817v1 Announce Type: new Abstract: Adversarial training (AT) is an effective defense for large language models (LLMs) against jailbreak attacks, but performing AT on LLMs is costly. To im

safetyarxiv-cs-lg
15 Apr 2026
Safety

Anthropogenic Regional Adaptation in Multimodal Vision-Language Model

DGX agent

arXiv:2604.11490v1 Announce Type: new Abstract: While the field of vision-language (VL) has achieved remarkable success in integrating visual and textual information across multiple languages and doma

safetyarxiv-cs-ai
14 Apr 2026
Model Releases

Consistency of AI-Generated Exercise Prescriptions: A Repeated Generation Study Using a Large Language Model

DGX agent

arXiv:2604.11287v1 Announce Type: new Abstract: Background: Large language models (LLMs) have been explored as tools for generating personalized exercise prescriptions, yet the consistency of outputs

model-releasesarxiv-cs-ai
14 Apr 2026
Safety

Efficient KernelSHAP Explanations for Patch-based 3D Medical Image Segmentation

DGX agent

arXiv:2604.11775v1 Announce Type: cross Abstract: Perturbation-based explainability methods such as KernelSHAP provide model-agnostic attributions but are typically impractical for patch-based 3D medi

safetyarxiv-cs-ai
14 Apr 2026
Safety

Endogenous Information in Routing Games: Memory-Constrained Equilibria, Recall Braess Paradoxes, and Memory Design

DGX agent

arXiv:2604.11733v1 Announce Type: cross Abstract: We study routing games in which travelers optimize over routes that are remembered or surfaced, rather than over a fixed exogenous action set. The pap

safetyarxiv-cs-ai
14 Apr 2026
Safety

Enhancing Fine-Grained Spatial Grounding in 3D CT Report Generation via Discriminative Guidance

DGX agent

arXiv:2604.10437v1 Announce Type: new Abstract: Vision--language models (VLMs) for radiology report generation (RRG) can produce long-form chest CT reports from volumetric scans and show strong potent

safetyarxiv-cs-cv
14 Apr 2026
Safety

Examining EAP Students' AI Disclosure Intention: A Cognition-Affect-Conation Perspective

DGX agent

arXiv:2604.10991v1 Announce Type: cross Abstract: The growing use of generative artificial intelligence (AI) in academic writing has raised increasing concerns regarding transparency and academic inte

safetyarxiv-cs-ai
14 Apr 2026
Safety

FREE-Switch: Frequency-based Dynamic LoRA Switch for Style Transfer

DGX agent

arXiv:2604.10023v1 Announce Type: cross Abstract: With the growing availability of open-sourced adapters trained on the same diffusion backbone for diverse scenes and objects, combining these pretrain

safetyarxiv-cs-ai
14 Apr 2026
Safety

Learning to Focus: CSI-Free Hierarchical MARL for Reconfigurable Reflectors

DGX agent

arXiv:2604.05165v2 Announce Type: replace Abstract: Reconfigurable Intelligent Surfaces (RIS) has a potential to engineer smart radio environments for next-generation millimeter-wave (mmWave) networks

safetyarxiv-cs-ai
14 Apr 2026
Safety

Like a Hammer, It Can Build, It Can Break: Large Language Model Uses, Perceptions, and Adoption in Cybersecurity Operations on Reddit

DGX agent

arXiv:2604.09998v1 Announce Type: cross Abstract: Large language models (LLMs) have recently emerged as promising tools for augmenting Security Operations Center (SOC) workflows, with vendors increasi

safetyarxiv-cs-ai
14 Apr 2026
Safety

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering

DGX agent

arXiv:2604.09757v1 Announce Type: cross Abstract: Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains largely te

safetyarxiv-cs-ai
14 Apr 2026
Safety

MOSAIC: Multi-Domain Orthogonal Session Adaptive Intent Capture for Prescient Recommendations

DGX agent

arXiv:2604.10147v1 Announce Type: cross Abstract: Capturing user intent across heterogeneous behavioral domains stands as a fundamental challenge in session-based recommender systems. Yet, existing mu

safetyarxiv-cs-ai
14 Apr 2026
Safety

Multi-Model Synthetic Training for Mission-Critical Small Language Models

DGX agent

arXiv:2509.13047v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities across many domains, yet their application to specialized fields remain

safetyarxiv-cs-ai
14 Apr 2026
Safety

OOM-RL: Out-of-Money Reinforcement Learning Market-Driven Alignment for LLM-Based Multi-Agent Systems

DGX agent

arXiv:2604.11477v1 Announce Type: new Abstract: The alignment of Multi-Agent Systems (MAS) for autonomous software engineering is constrained by evaluator epistemic uncertainty. Current paradigms, suc

safetyarxiv-cs-ai
14 Apr 2026
Safety

PEMANT: Persona-Enriched Multi-Agent Negotiation for Travel

DGX agent

arXiv:2604.10475v1 Announce Type: new Abstract: Modeling household-level trip generation is fundamental to accurate demand forecasting, traffic flow estimation, and urban system planning. Existing stu

safetyarxiv-cs-ai
14 Apr 2026
Safety

Preference-Agile Multi-Objective Optimization for Real-time Vehicle Dispatching

DGX agent

arXiv:2604.10664v1 Announce Type: new Abstract: Multi-objective optimization (MOO) has been widely studied in literature because of its versatility in human-centered decision making in real-life appli

safetyarxiv-cs-ai
14 Apr 2026
Safety

SemaClaw: A Step Towards General-Purpose Personal AI Agents through Harness Engineering

DGX agent

arXiv:2604.11548v1 Announce Type: new Abstract: The rise of OpenClaw in early 2026 marks the moment when millions of users began deploying personal AI agents into their daily lives, delegating tasks r

safetyarxiv-cs-ai
14 Apr 2026
Safety

Sense Less, Infer More: Agentic Multimodal Transformers for Edge Medical Intelligence

DGX agent

arXiv:2604.10404v1 Announce Type: cross Abstract: Edge-based multimodal medical monitoring requires models that balance diagnostic accuracy with severe energy constraints. Continuous acquisition of EC

safetyarxiv-cs-lg
14 Apr 2026
Safety

Steered LLM Activations are Non-Surjective

DGX agent

arXiv:2604.09839v1 Announce Type: new Abstract: Activation steering is a popular white-box control technique that modifies model activations to elicit an abstract change in output behavior. It has als

safetyarxiv-cs-ai
14 Apr 2026
← Previous
1…3940414243…299
Next →