AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,351 results
21 Apr 2026

Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models

SafetyDGX agent

arXiv:2601.15220v2 Announce Type: replace Abstract: We identify a novel phenomenon in language models: benign fine-tuning of frontier models can lead to privacy collapse. We find that diverse, subtle

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring

SafetyDGX agent

arXiv:2512.12069v3 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) are vulnerable to a growing array of multimodal jailbreak attacks, necessitating defenses that are both g

Test-Time Alignment via Hypothesis Reweighting

SafetyDGX agent

arXiv:2412.08812v2 Announce Type: replace Abstract: Reward models trained on aggregate preferences often fail to capture individual users' values, but existing adaptation methods such as fine-tuning o

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
20 Apr 2026

Language Models as Semantic Teachers: Post-Training Alignment for Medical Audio Understanding

SafetyDGX agent

arXiv:2512.04847v2 Announce Type: replace-cross Abstract: Pre-trained audio models excel at detecting acoustic patterns in auscultation sounds but often fail to grasp their clinical significance, limi

Puppets or partners? Governing cyborg propaganda in the digital public square

SafetyDGX agent

arXiv:2602.13088v2 Announce Type: replace-cross Abstract: The distinction between genuine grassroots activism and automated influence operations is collapsing. While contemporary policy debates priori

The AI industry insists they can manage the risks of superintelligence, but there are in fact zero widely agreed on or accepted solutions to…

SafetyDGX agent

The AI industry insists they can manage the risks of superintelligence, but there are in fact zero widely agreed on or accepted solutions to the problem of how one could even control something vastly

When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models

SafetyDGX agent

arXiv:2510.09689v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have been augmented with web search to overcome the limitations of the static knowledge boundary by accessing up-

17 Apr 2026

Blinded Multi-Rater Comparative Evaluation of a Large Language Model and Clinician-Authored Responses in CGM-Informed Diabetes Counseling

SafetyDGX agent

arXiv:2604.15124v1 Announce Type: new Abstract: Continuous glucose monitoring (CGM) is central to diabetes care, but explaining CGM patterns clearly and empathetically remains time-intensive. Evidence

Inclusive Kitchen Design for Older Adults: Generative AI Visualizations to Support Mild Cognitive Impairment

SafetyDGX agent

arXiv:2604.13203v1 Announce Type: cross Abstract: Mild Cognitive Impairment (MCI) affects 15-20% of adults aged 65 and older, often making kitchen navigation and independent living difficult, particul

Momentum-constrained Hybrid Heuristic Trajectory Optimization Framework with Residual-enhanced DRL for Visually Impaired Scenarios

SafetyDGX agent

arXiv:2604.14986v1 Announce Type: new Abstract: Safe and efficient assistive planning for visually impaired scenarios remains challenging, since existing methods struggle with multi-objective optimiza

ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration

SafetyDGX agent

arXiv:2509.21823v2 Announce Type: replace Abstract: Reward is critical to the evaluation and training of large language models (LLMs). However, existing rule-based or model-based reward methods strugg

Separation is Optimal for LQR under Intermittent Feedback

SafetyDGX agent

arXiv:2603.27833v2 Announce Type: cross Abstract: In this work, we first prove that the separation principle holds for communication-constrained LQR problems under i.i.d. zero-mean disturbances with a

Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models

SafetyDGX agent

arXiv:2604.14629v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have shown remarkable capabilities in joint vision-language understanding, but their large scale poses significant challen

TOPCELL: Topology Optimization of Standard Cell via LLMs

SafetyDGX agent

arXiv:2604.14237v1 Announce Type: new Abstract: Transistor topology optimization is a critical step in standard cell design, directly dictating diffusion sharing efficiency and downstream routability.

Towards Scalable Lightweight GUI Agents via Multi-role Orchestration

SafetyDGX agent

arXiv:2604.13488v1 Announce Type: new Abstract: Autonomous Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) enable digital automation on end-user devices. Whil

VoxSafeBench: Not Just What Is Said, but Who, How, and Where

SafetyDGX agent

arXiv:2604.14548v1 Announce Type: cross Abstract: As speech language models (SLMs) transition from personal devices into shared, multi-user environments, their responses must account for far more than

What some leaders think of using universal income to mitigate AI-fueled layoffs: Musk calls it the 'best way', OpenAI's policy doc mentions a Public Wealth Fund (Siladitya Ray/Forbes)

SafetyDGX agent

Siladitya Ray / Forbes: What some leaders think of using universal income to mitigate AI-fueled layoffs: Musk calls it the “best way”, OpenAI's policy doc mentions a Public Wealth Fund — Topline — Elo

16 Apr 2026

Asymmetric-Loss-Guided Hybrid CNN-BiLSTM-Attention Model for Industrial RUL Prediction with Interpretable Failure Heatmaps

SafetyDGX agent

arXiv:2604.13459v1 Announce Type: new Abstract: Turbofan engine degradation under sustained operational stress necessitates robust prognostic systems capable of accurately estimating the Remaining Use

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling

SafetyDGX agent

arXiv:2604.13561v1 Announce Type: new Abstract: Vision-language models trained with contrastive learning on paired medical images and reports show strong zero-shot diagnostic capabilities, yet the eff

Context Sensitivity Improves Human-Machine Visual Alignment

SafetyDGX agent

arXiv:2604.13883v1 Announce Type: new Abstract: Modern machine learning models typically represent inputs as fixed points in a high-dimensional embedding space. While this approach has been proven pow

RFK Jr. forces FDA to reconsider 12 unproven peptides after 2023 ban

SafetyDGX agent

In 2023, the FDA removed 19 peptides from the list of drugs that compounding pharmacies could produce, and in 2026 the FDA announced it will review whether to add back 7 of these peptides following pr

Robust Verification of Controllers under State Uncertainty via Hamilton-Jacobi Reachability Analysis

SafetyDGX agent

arXiv:2511.14755v2 Announce Type: replace-cross Abstract: As perception-based controllers for autonomous systems become increasingly popular in the real world, it is important that we can formally ver

VRAG-DFD: Verifiable Retrieval-Augmentation for MLLM-based Deepfake Detection

SafetyDGX agent

arXiv:2604.13660v1 Announce Type: new Abstract: In Deepfake Detection (DFD) tasks, researchers proposed two types of MLLM-based methods: complementary combination with small DFD detectors, or static f

15 Apr 2026

A Comparison of Reinforcement Learning and Optimal Control Methods for Path Planning

SafetyDGX agent

arXiv:2604.12628v1 Announce Type: cross Abstract: Path-planning for autonomous vehicles in threat-laden environments is a fundamental challenge. While traditional optimal control methods can find idea

ARGen: Affect-Reinforced Generative Augmentation towards Vision-based Dynamic Emotion Perception

SafetyDGX agent

arXiv:2604.12255v1 Announce Type: cross Abstract: Dynamic facial expression recognition in the wild remains challenging due to data scarcity and long-tail distributions, which hinder models from effec

LatentRefusal: Latent-Signal Refusal for Unanswerable Text-to-SQL Queries

SafetyDGX agent

arXiv:2601.10398v3 Announce Type: replace Abstract: In LLM-based text-to-SQL systems, unanswerable and underspecified user queries may generate not only incorrect text but also executable programs tha

Risk-Calibrated Learning: Minimizing Fatal Errors in Medical AI

SafetyDGX agent

arXiv:2604.12693v1 Announce Type: new Abstract: Deep learning models often achieve expert-level accuracy in medical image classification but suffer from a critical flaw: semantic incoherence. These hi

Robust Optimization for Mitigating Reward Hacking with Correlated Proxies

SafetyDGX agent

arXiv:2604.12086v1 Announce Type: new Abstract: Designing robust reinforcement learning (RL) agents in the presence of imperfect reward signals remains a core challenge. In practice, agents are often

TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs

SafetyDGX agent

arXiv:2604.12232v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed across diverse domains, yet their vulnerability to jailbreak attacks, where adversarial inputs

Understanding and Improving Continuous Adversarial Training for LLMs via In-context Learning Theory

SafetyDGX agent

arXiv:2604.12817v1 Announce Type: new Abstract: Adversarial training (AT) is an effective defense for large language models (LLMs) against jailbreak attacks, but performing AT on LLMs is costly. To im

14 Apr 2026

Anthropogenic Regional Adaptation in Multimodal Vision-Language Model

SafetyDGX agent

arXiv:2604.11490v1 Announce Type: new Abstract: While the field of vision-language (VL) has achieved remarkable success in integrating visual and textual information across multiple languages and doma

Consistency of AI-Generated Exercise Prescriptions: A Repeated Generation Study Using a Large Language Model

Model ReleasesDGX agent

arXiv:2604.11287v1 Announce Type: new Abstract: Background: Large language models (LLMs) have been explored as tools for generating personalized exercise prescriptions, yet the consistency of outputs

Efficient KernelSHAP Explanations for Patch-based 3D Medical Image Segmentation

SafetyDGX agent

arXiv:2604.11775v1 Announce Type: cross Abstract: Perturbation-based explainability methods such as KernelSHAP provide model-agnostic attributions but are typically impractical for patch-based 3D medi

Endogenous Information in Routing Games: Memory-Constrained Equilibria, Recall Braess Paradoxes, and Memory Design

SafetyDGX agent

arXiv:2604.11733v1 Announce Type: cross Abstract: We study routing games in which travelers optimize over routes that are remembered or surfaced, rather than over a fixed exogenous action set. The pap

Enhancing Fine-Grained Spatial Grounding in 3D CT Report Generation via Discriminative Guidance

SafetyDGX agent

arXiv:2604.10437v1 Announce Type: new Abstract: Vision--language models (VLMs) for radiology report generation (RRG) can produce long-form chest CT reports from volumetric scans and show strong potent

Examining EAP Students' AI Disclosure Intention: A Cognition-Affect-Conation Perspective

SafetyDGX agent

arXiv:2604.10991v1 Announce Type: cross Abstract: The growing use of generative artificial intelligence (AI) in academic writing has raised increasing concerns regarding transparency and academic inte

FREE-Switch: Frequency-based Dynamic LoRA Switch for Style Transfer

SafetyDGX agent

arXiv:2604.10023v1 Announce Type: cross Abstract: With the growing availability of open-sourced adapters trained on the same diffusion backbone for diverse scenes and objects, combining these pretrain

Learning to Focus: CSI-Free Hierarchical MARL for Reconfigurable Reflectors

SafetyDGX agent

arXiv:2604.05165v2 Announce Type: replace Abstract: Reconfigurable Intelligent Surfaces (RIS) has a potential to engineer smart radio environments for next-generation millimeter-wave (mmWave) networks

Like a Hammer, It Can Build, It Can Break: Large Language Model Uses, Perceptions, and Adoption in Cybersecurity Operations on Reddit

SafetyDGX agent

arXiv:2604.09998v1 Announce Type: cross Abstract: Large language models (LLMs) have recently emerged as promising tools for augmenting Security Operations Center (SOC) workflows, with vendors increasi

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering

SafetyDGX agent

arXiv:2604.09757v1 Announce Type: cross Abstract: Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains largely te

MOSAIC: Multi-Domain Orthogonal Session Adaptive Intent Capture for Prescient Recommendations

SafetyDGX agent

arXiv:2604.10147v1 Announce Type: cross Abstract: Capturing user intent across heterogeneous behavioral domains stands as a fundamental challenge in session-based recommender systems. Yet, existing mu

Multi-Model Synthetic Training for Mission-Critical Small Language Models

SafetyDGX agent

arXiv:2509.13047v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities across many domains, yet their application to specialized fields remain

OOM-RL: Out-of-Money Reinforcement Learning Market-Driven Alignment for LLM-Based Multi-Agent Systems

SafetyDGX agent

arXiv:2604.11477v1 Announce Type: new Abstract: The alignment of Multi-Agent Systems (MAS) for autonomous software engineering is constrained by evaluator epistemic uncertainty. Current paradigms, suc

PEMANT: Persona-Enriched Multi-Agent Negotiation for Travel

SafetyDGX agent

arXiv:2604.10475v1 Announce Type: new Abstract: Modeling household-level trip generation is fundamental to accurate demand forecasting, traffic flow estimation, and urban system planning. Existing stu

Preference-Agile Multi-Objective Optimization for Real-time Vehicle Dispatching

SafetyDGX agent

arXiv:2604.10664v1 Announce Type: new Abstract: Multi-objective optimization (MOO) has been widely studied in literature because of its versatility in human-centered decision making in real-life appli

SemaClaw: A Step Towards General-Purpose Personal AI Agents through Harness Engineering

SafetyDGX agent

arXiv:2604.11548v1 Announce Type: new Abstract: The rise of OpenClaw in early 2026 marks the moment when millions of users began deploying personal AI agents into their daily lives, delegating tasks r

Sense Less, Infer More: Agentic Multimodal Transformers for Edge Medical Intelligence

SafetyDGX agent

arXiv:2604.10404v1 Announce Type: cross Abstract: Edge-based multimodal medical monitoring requires models that balance diagnostic accuracy with severe energy constraints. Continuous acquisition of EC

Steered LLM Activations are Non-Surjective

SafetyDGX agent

arXiv:2604.09839v1 Announce Type: new Abstract: Activation steering is a popular white-box control technique that modifies model activations to elicit an abstract change in output behavior. It has als

Towards Adaptive Open-Set Object Detection via Category-Level Collaboration Knowledge Mining

SafetyDGX agent

arXiv:2604.11195v1 Announce Type: cross Abstract: Existing object detectors often struggle to generalize across domains while adapting to emerging novel categories. Adaptive open-set object detection

Trajectory-based actuator identification via differentiable simulation

SafetyDGX agent

arXiv:2604.10351v1 Announce Type: new Abstract: Accurate actuation models are critical for bridging the gap between simulation and real robot behavior, yet obtaining high-fidelity actuator dynamics ty

Weird Generalization is Weirdly Brittle

SafetyDGX agent

arXiv:2604.10022v1 Announce Type: new Abstract: Weird generalization is a phenomenon in which models fine-tuned on data from a narrow domain (e.g. insecure code) develop surprising traits that manifes

13 Apr 2026

A Restore Britain Government would back British drivers. How? Raise the speed limit on motorways to 80mph. Remove all 20mph zones, other tha…

SafetyDGX agent

A Restore Britain Government would back British drivers. How? Raise the speed limit on motorways to 80mph. Remove all 20mph zones, other than those outside schools or areas with vulnerable individuals

Better Eyes, Better Thoughts: Why Vision Chain-of-Thought Fails in Medicine

SafetyDGX agent

arXiv:2603.06665v2 Announce Type: replace-cross Abstract: Large vision-language models (VLMs) often benefit from chain-of-thought (CoT) prompting in general domains, yet its efficacy in medical vision

CausalVAD: De-confounding End-to-End Autonomous Driving via Causal Intervention

SafetyDGX agent

arXiv:2603.18561v2 Announce Type: replace Abstract: Planning-oriented end-to-end driving models show great promise, yet they fundamentally learn statistical correlations instead of true causal relatio

ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion

SafetyDGX agent

arXiv:2604.09450v1 Announce Type: cross Abstract: Chest X-ray report generation (CXR-RG) has the potential to substantially alleviate radiologists' workload. However, conventional autoregressive visio

Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism

SafetyDGX agent

arXiv:2604.09544v1 Announce Type: cross Abstract: Large language models (LLMs) undergo alignment training to avoid harmful behaviors, yet the resulting safeguards remain brittle: jailbreaks routinely

On the Spectral Geometry of Cross-Modal Representations: A Functional Map Diagnostic for Multimodal Alignment

SafetyDGX agent

arXiv:2604.08579v1 Announce Type: cross Abstract: We study cross-modal alignment between independently pretrained vision (DINOv2) and language (all-MiniLM-L6-v2) encoders using the functional map fram

Semantic Intent Fragmentation: A Single-Shot Compositional Attack on Multi-Agent AI Pipelines

SafetyDGX agent

arXiv:2604.08608v1 Announce Type: cross Abstract: We introduce Semantic Intent Fragmentation (SIF), an attack class against LLM orchestration systems where a single, legitimately phrased request cause

12 Apr 2026

Is Elon Musk right about Sam Altman? just read the new Ronan Farrow and Andrew Marantz piece on Sam Altman in The New Yorker, and it is damn…

SafetyDGX agent

Is Elon Musk right about Sam Altman? just read the new Ronan Farrow and Andrew Marantz piece on Sam Altman in The New Yorker, and it is damning. And the answer is yes. It’s a deep dive based on never-

10 Apr 2026

Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection

SafetyDGX agent

arXiv:2604.07831v1 Announce Type: cross Abstract: Existing red-teaming studies on GUI agents have important limitations. Adversarial perturbations typically require white-box access, which is unavaila

← Previous
1…3132333435…240
Next →