AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,314 results
Safety

Navigating the Conceptual Multiverse

DGX agent

arXiv:2604.17815v1 Announce Type: cross Abstract: When language models answer open-ended problems, they implicitly make hidden decisions that shape their outputs, leaving users with uncontextualized a

safetyarxiv-cs-cl
21 Apr 2026
Safety

One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment

DGX agent
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2601.18731v2 Announce Type: replace Abstract: Alignment of Large Language Models (LLMs) aims to align outputs with human preferences, and personalized alignment further adapts models to individu

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

Please refuse to answer me! Mitigating Over-Refusal in Large Language Models via Adaptive Contrastive Decoding

DGX agent

arXiv:2604.17132v1 Announce Type: new Abstract: Safety-aligned large language models (LLMs) often generate refusal responses to harmless queries due to the over-refusal problem. However, existing meth

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

PowerCLIP: Powerset Alignment for Contrastive Pre-Training

DGX agent

arXiv:2511.23170v5 Announce Type: replace Abstract: Contrastive vision-language pre-training frameworks such as CLIP have demonstrated impressive zero-shot performance across a range of vision-languag

safetyarxiv-cs-cv
21 Apr 2026
Safety

Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models

DGX agent

arXiv:2601.15220v2 Announce Type: replace Abstract: We identify a novel phenomenon in language models: benign fine-tuning of frontier models can lead to privacy collapse. We find that diverse, subtle

safetyarxiv-cs-cl
21 Apr 2026
Safety

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring

DGX agent

arXiv:2512.12069v3 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) are vulnerable to a growing array of multimodal jailbreak attacks, necessitating defenses that are both g

safetyarxiv-cs-cl
21 Apr 2026
Safety

Test-Time Alignment via Hypothesis Reweighting

DGX agent

arXiv:2412.08812v2 Announce Type: replace Abstract: Reward models trained on aggregate preferences often fail to capture individual users' values, but existing adaptation methods such as fine-tuning o

safetyarxiv-cs-lg
21 Apr 2026
Safety

Language Models as Semantic Teachers: Post-Training Alignment for Medical Audio Understanding

DGX agent

arXiv:2512.04847v2 Announce Type: replace-cross Abstract: Pre-trained audio models excel at detecting acoustic patterns in auscultation sounds but often fail to grasp their clinical significance, limi

safetyarxiv-cs-ai
20 Apr 2026
Safety

Puppets or partners? Governing cyborg propaganda in the digital public square

DGX agent

arXiv:2602.13088v2 Announce Type: replace-cross Abstract: The distinction between genuine grassroots activism and automated influence operations is collapsing. While contemporary policy debates priori

safetyarxiv-cs-ai
20 Apr 2026
Safety

When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models

DGX agent

arXiv:2510.09689v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have been augmented with web search to overcome the limitations of the static knowledge boundary by accessing up-

safetyarxiv-cs-ai
20 Apr 2026
Safety

Blinded Multi-Rater Comparative Evaluation of a Large Language Model and Clinician-Authored Responses in CGM-Informed Diabetes Counseling

DGX agent

arXiv:2604.15124v1 Announce Type: new Abstract: Continuous glucose monitoring (CGM) is central to diabetes care, but explaining CGM patterns clearly and empathetically remains time-intensive. Evidence

safetyarxiv-cs-cl
17 Apr 2026
Safety

Inclusive Kitchen Design for Older Adults: Generative AI Visualizations to Support Mild Cognitive Impairment

DGX agent

arXiv:2604.13203v1 Announce Type: cross Abstract: Mild Cognitive Impairment (MCI) affects 15-20% of adults aged 65 and older, often making kitchen navigation and independent living difficult, particul

safetyarxiv-cs-ai
17 Apr 2026
Safety

Momentum-constrained Hybrid Heuristic Trajectory Optimization Framework with Residual-enhanced DRL for Visually Impaired Scenarios

DGX agent

arXiv:2604.14986v1 Announce Type: new Abstract: Safe and efficient assistive planning for visually impaired scenarios remains challenging, since existing methods struggle with multi-objective optimiza

safetyarxiv-cs-ro
17 Apr 2026
Safety

ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration

DGX agent

arXiv:2509.21823v2 Announce Type: replace Abstract: Reward is critical to the evaluation and training of large language models (LLMs). However, existing rule-based or model-based reward methods strugg

safetyarxiv-cs-ai
17 Apr 2026
Safety

Separation is Optimal for LQR under Intermittent Feedback

DGX agent

arXiv:2603.27833v2 Announce Type: cross Abstract: In this work, we first prove that the separation principle holds for communication-constrained LQR problems under i.i.d. zero-mean disturbances with a

safetyarxiv-cs-ro
17 Apr 2026
Safety

Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models

DGX agent

arXiv:2604.14629v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have shown remarkable capabilities in joint vision-language understanding, but their large scale poses significant challen

safetyarxiv-cs-cv
17 Apr 2026
Safety

TOPCELL: Topology Optimization of Standard Cell via LLMs

DGX agent

arXiv:2604.14237v1 Announce Type: new Abstract: Transistor topology optimization is a critical step in standard cell design, directly dictating diffusion sharing efficiency and downstream routability.

safetyarxiv-cs-lg
17 Apr 2026
Safety

Towards Scalable Lightweight GUI Agents via Multi-role Orchestration

DGX agent

arXiv:2604.13488v1 Announce Type: new Abstract: Autonomous Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) enable digital automation on end-user devices. Whil

safetyarxiv-cs-ai
17 Apr 2026
Safety

VoxSafeBench: Not Just What Is Said, but Who, How, and Where

DGX agent

arXiv:2604.14548v1 Announce Type: cross Abstract: As speech language models (SLMs) transition from personal devices into shared, multi-user environments, their responses must account for far more than

safetyarxiv-cs-lg
17 Apr 2026
Safety

Asymmetric-Loss-Guided Hybrid CNN-BiLSTM-Attention Model for Industrial RUL Prediction with Interpretable Failure Heatmaps

DGX agent

arXiv:2604.13459v1 Announce Type: new Abstract: Turbofan engine degradation under sustained operational stress necessitates robust prognostic systems capable of accurately estimating the Remaining Use

safetyarxiv-cs-lg
16 Apr 2026
Safety

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling

DGX agent

arXiv:2604.13561v1 Announce Type: new Abstract: Vision-language models trained with contrastive learning on paired medical images and reports show strong zero-shot diagnostic capabilities, yet the eff

safetyarxiv-cs-cv
16 Apr 2026
Safety

Context Sensitivity Improves Human-Machine Visual Alignment

DGX agent

arXiv:2604.13883v1 Announce Type: new Abstract: Modern machine learning models typically represent inputs as fixed points in a high-dimensional embedding space. While this approach has been proven pow

safetyarxiv-cs-cv
16 Apr 2026
Safety

Robust Verification of Controllers under State Uncertainty via Hamilton-Jacobi Reachability Analysis

DGX agent

arXiv:2511.14755v2 Announce Type: replace-cross Abstract: As perception-based controllers for autonomous systems become increasingly popular in the real world, it is important that we can formally ver

safetyarxiv-cs-lg
16 Apr 2026
Safety

VRAG-DFD: Verifiable Retrieval-Augmentation for MLLM-based Deepfake Detection

DGX agent

arXiv:2604.13660v1 Announce Type: new Abstract: In Deepfake Detection (DFD) tasks, researchers proposed two types of MLLM-based methods: complementary combination with small DFD detectors, or static f

safetyarxiv-cs-cv
16 Apr 2026
Safety

A Comparison of Reinforcement Learning and Optimal Control Methods for Path Planning

DGX agent

arXiv:2604.12628v1 Announce Type: cross Abstract: Path-planning for autonomous vehicles in threat-laden environments is a fundamental challenge. While traditional optimal control methods can find idea

safetyarxiv-cs-ro
15 Apr 2026
Safety

ARGen: Affect-Reinforced Generative Augmentation towards Vision-based Dynamic Emotion Perception

DGX agent

arXiv:2604.12255v1 Announce Type: cross Abstract: Dynamic facial expression recognition in the wild remains challenging due to data scarcity and long-tail distributions, which hinder models from effec

safetyarxiv-cs-ai
15 Apr 2026
Safety

LatentRefusal: Latent-Signal Refusal for Unanswerable Text-to-SQL Queries

DGX agent

arXiv:2601.10398v3 Announce Type: replace Abstract: In LLM-based text-to-SQL systems, unanswerable and underspecified user queries may generate not only incorrect text but also executable programs tha

safetyarxiv-cs-ai
15 Apr 2026
Safety

Risk-Calibrated Learning: Minimizing Fatal Errors in Medical AI

DGX agent

arXiv:2604.12693v1 Announce Type: new Abstract: Deep learning models often achieve expert-level accuracy in medical image classification but suffer from a critical flaw: semantic incoherence. These hi

safetyarxiv-cs-cv
15 Apr 2026
Safety

Robust Optimization for Mitigating Reward Hacking with Correlated Proxies

DGX agent

arXiv:2604.12086v1 Announce Type: new Abstract: Designing robust reinforcement learning (RL) agents in the presence of imperfect reward signals remains a core challenge. In practice, agents are often

safetyarxiv-cs-lg
15 Apr 2026
Safety

TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs

DGX agent

arXiv:2604.12232v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed across diverse domains, yet their vulnerability to jailbreak attacks, where adversarial inputs

safetyarxiv-cs-ai
15 Apr 2026
Safety

Understanding and Improving Continuous Adversarial Training for LLMs via In-context Learning Theory

DGX agent

arXiv:2604.12817v1 Announce Type: new Abstract: Adversarial training (AT) is an effective defense for large language models (LLMs) against jailbreak attacks, but performing AT on LLMs is costly. To im

safetyarxiv-cs-lg
15 Apr 2026
Safety

Anthropogenic Regional Adaptation in Multimodal Vision-Language Model

DGX agent

arXiv:2604.11490v1 Announce Type: new Abstract: While the field of vision-language (VL) has achieved remarkable success in integrating visual and textual information across multiple languages and doma

safetyarxiv-cs-ai
14 Apr 2026
Model Releases

Consistency of AI-Generated Exercise Prescriptions: A Repeated Generation Study Using a Large Language Model

DGX agent

arXiv:2604.11287v1 Announce Type: new Abstract: Background: Large language models (LLMs) have been explored as tools for generating personalized exercise prescriptions, yet the consistency of outputs

model-releasesarxiv-cs-ai
14 Apr 2026
Safety

Efficient KernelSHAP Explanations for Patch-based 3D Medical Image Segmentation

DGX agent

arXiv:2604.11775v1 Announce Type: cross Abstract: Perturbation-based explainability methods such as KernelSHAP provide model-agnostic attributions but are typically impractical for patch-based 3D medi

safetyarxiv-cs-ai
14 Apr 2026
Safety

Endogenous Information in Routing Games: Memory-Constrained Equilibria, Recall Braess Paradoxes, and Memory Design

DGX agent

arXiv:2604.11733v1 Announce Type: cross Abstract: We study routing games in which travelers optimize over routes that are remembered or surfaced, rather than over a fixed exogenous action set. The pap

safetyarxiv-cs-ai
14 Apr 2026
Safety

Enhancing Fine-Grained Spatial Grounding in 3D CT Report Generation via Discriminative Guidance

DGX agent

arXiv:2604.10437v1 Announce Type: new Abstract: Vision--language models (VLMs) for radiology report generation (RRG) can produce long-form chest CT reports from volumetric scans and show strong potent

safetyarxiv-cs-cv
14 Apr 2026
Safety

Examining EAP Students' AI Disclosure Intention: A Cognition-Affect-Conation Perspective

DGX agent

arXiv:2604.10991v1 Announce Type: cross Abstract: The growing use of generative artificial intelligence (AI) in academic writing has raised increasing concerns regarding transparency and academic inte

safetyarxiv-cs-ai
14 Apr 2026
Safety

FREE-Switch: Frequency-based Dynamic LoRA Switch for Style Transfer

DGX agent

arXiv:2604.10023v1 Announce Type: cross Abstract: With the growing availability of open-sourced adapters trained on the same diffusion backbone for diverse scenes and objects, combining these pretrain

safetyarxiv-cs-ai
14 Apr 2026
Safety

Learning to Focus: CSI-Free Hierarchical MARL for Reconfigurable Reflectors

DGX agent

arXiv:2604.05165v2 Announce Type: replace Abstract: Reconfigurable Intelligent Surfaces (RIS) has a potential to engineer smart radio environments for next-generation millimeter-wave (mmWave) networks

safetyarxiv-cs-ai
14 Apr 2026
Safety

Like a Hammer, It Can Build, It Can Break: Large Language Model Uses, Perceptions, and Adoption in Cybersecurity Operations on Reddit

DGX agent

arXiv:2604.09998v1 Announce Type: cross Abstract: Large language models (LLMs) have recently emerged as promising tools for augmenting Security Operations Center (SOC) workflows, with vendors increasi

safetyarxiv-cs-ai
14 Apr 2026
Safety

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering

DGX agent

arXiv:2604.09757v1 Announce Type: cross Abstract: Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains largely te

safetyarxiv-cs-ai
14 Apr 2026
Safety

MOSAIC: Multi-Domain Orthogonal Session Adaptive Intent Capture for Prescient Recommendations

DGX agent

arXiv:2604.10147v1 Announce Type: cross Abstract: Capturing user intent across heterogeneous behavioral domains stands as a fundamental challenge in session-based recommender systems. Yet, existing mu

safetyarxiv-cs-ai
14 Apr 2026
Safety

Multi-Model Synthetic Training for Mission-Critical Small Language Models

DGX agent

arXiv:2509.13047v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities across many domains, yet their application to specialized fields remain

safetyarxiv-cs-ai
14 Apr 2026
Safety

OOM-RL: Out-of-Money Reinforcement Learning Market-Driven Alignment for LLM-Based Multi-Agent Systems

DGX agent

arXiv:2604.11477v1 Announce Type: new Abstract: The alignment of Multi-Agent Systems (MAS) for autonomous software engineering is constrained by evaluator epistemic uncertainty. Current paradigms, suc

safetyarxiv-cs-ai
14 Apr 2026
Safety

PEMANT: Persona-Enriched Multi-Agent Negotiation for Travel

DGX agent

arXiv:2604.10475v1 Announce Type: new Abstract: Modeling household-level trip generation is fundamental to accurate demand forecasting, traffic flow estimation, and urban system planning. Existing stu

safetyarxiv-cs-ai
14 Apr 2026
Safety

Preference-Agile Multi-Objective Optimization for Real-time Vehicle Dispatching

DGX agent

arXiv:2604.10664v1 Announce Type: new Abstract: Multi-objective optimization (MOO) has been widely studied in literature because of its versatility in human-centered decision making in real-life appli

safetyarxiv-cs-ai
14 Apr 2026
Safety

SemaClaw: A Step Towards General-Purpose Personal AI Agents through Harness Engineering

DGX agent

arXiv:2604.11548v1 Announce Type: new Abstract: The rise of OpenClaw in early 2026 marks the moment when millions of users began deploying personal AI agents into their daily lives, delegating tasks r

safetyarxiv-cs-ai
14 Apr 2026
Safety

Sense Less, Infer More: Agentic Multimodal Transformers for Edge Medical Intelligence

DGX agent

arXiv:2604.10404v1 Announce Type: cross Abstract: Edge-based multimodal medical monitoring requires models that balance diagnostic accuracy with severe energy constraints. Continuous acquisition of EC

safetyarxiv-cs-lg
14 Apr 2026
← Previous
1…3536373839…257
Next →