AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,702 results
21 Apr 2026

MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models

SafetyDGX agent

arXiv:2604.17730v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly explored as scalable tools for mental health counseling, yet evaluating their safety remains challenging d

Mix and Match: Context Pairing for Scalable Topic-Controlled Educational Summarisation

SafetyDGX agent

arXiv:2604.18087v1 Announce Type: new Abstract: Topic-controlled summarisation enables users to generate summaries focused on specific aspects of source documents. This paper investigates a data augme

MoCo: A One-Stop Shop for Model Collaboration Research

SafetyDGX agent

arXiv:2601.21257v2 Announce Type: replace Abstract: Advancing beyond single monolithic language models (LMs), recent research increasingly recognizes the importance of model collaboration, where multi


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Modeling User Exploration Saturation: When Recommender Systems Should Stop Pushing Novelty

SafetyDGX agent

arXiv:2604.16419v1 Announce Type: cross Abstract: Fairness-aware recommender systems often mitigate bias by increasing exposure to under-represented or long-tail content, commonly through mechanisms t

Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition

SafetyDGX agent

arXiv:2604.17062v1 Announce Type: new Abstract: Zero-shot action recognition is challenging due to the semantic gap between seen and unseen classes. We present a novel framework that enhances CLIP wit

Multimodal Policy Internalization for Conversational Agents

SafetyDGX agent

arXiv:2510.09474v2 Announce Type: replace Abstract: Modern conversational agents like ChatGPT and Alexa+ rely on predefined policies specifying metadata, response styles, and tool-usage rules. As thes

Navigating Distribution Shifts in Medical Image Analysis: A Survey

SafetyDGX agent

arXiv:2411.05824v3 Announce Type: replace-cross Abstract: Medical Image Analysis (MedIA) has become indispensable in modern healthcare, enhancing clinical diagnostics and personalized treatment. Despi

Navigating the Conceptual Multiverse

SafetyDGX agent

arXiv:2604.17815v1 Announce Type: cross Abstract: When language models answer open-ended problems, they implicitly make hidden decisions that shape their outputs, leaving users with uncontextualized a

Negative Advantage Is a Double-Edged Sword: Calibrating Advantage in GRPO for Deep Search

SafetyDGX agent

arXiv:2604.18235v1 Announce Type: new Abstract: Deep search agents can autonomously initiate multi-turn interactions with search engines, thereby exhibiting strong question-answering capabilities. Suc

OmniVLA-RL: A Vision-Language-Action Model with Spatial Understanding and Online RL

SafetyDGX agent

arXiv:2604.17706v1 Announce Type: new Abstract: Visual-Language-Action (VLA) models represent a paradigm shift in embodied AI, yet existing frameworks often struggle with imprecise spatial perception,

On-Orbit Space AI: Federated, Multi-Agent, and Collaborative Algorithms for Satellite Constellations

SafetyDGX agent

arXiv:2604.16518v1 Announce Type: new Abstract: Satellite constellations are transforming space systems from isolated spacecraft into networked, software-defined platforms capable of on-orbit percepti

On Safety Risks in Experience-Driven Self-Evolving Agents

SafetyDGX agent

arXiv:2604.16968v1 Announce Type: new Abstract: Experience-driven self-evolution has emerged as a promising paradigm for improving the autonomy of large language model agents, yet its reliance on self

On the Convergence and Size Transferability of Continuous-depth Graph Neural Networks

SafetyDGX agent

arXiv:2510.03923v2 Announce Type: replace Abstract: Continuous-depth graph neural networks, also known as Graph Neural Differential Equations (GNDEs), combine the structural inductive bias of Graph Ne

On the Importance of Tactile Sensing for Imitation Learning: A Case Study on Robotic Match Lighting

SafetyDGX agent

arXiv:2504.13618v4 Announce Type: replace Abstract: The field of robotic manipulation has advanced significantly in recent years. At the sensing level, several novel tactile sensors have been develope

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization

SafetyDGX agent

arXiv:2509.23542v2 Announce Type: replace Abstract: The LLM-as-a-judge paradigm is widely used in both evaluating free-text model responses and reward modeling for model alignment and fine-tuning. Rec

One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment

SafetyDGX agent

arXiv:2601.18731v2 Announce Type: replace Abstract: Alignment of Large Language Models (LLMs) aims to align outputs with human preferences, and personalized alignment further adapts models to individu

Online Conformal Prediction with Adversarial Semi-bandit Feedback via Regret Minimization

SafetyDGX agent

arXiv:2604.17984v1 Announce Type: new Abstract: Uncertainty quantification is crucial in safety-critical systems, where decisions must be made under uncertainty. In particular, we consider the problem

Operationalizing Fairness in Text-to-Image Models: A Survey of Bias, Fairness Audits and Mitigation Strategies

SafetyDGX agent

arXiv:2604.16516v1 Announce Type: new Abstract: Text-to-Image (T2I) generation models have been widely adopted across various industries, yet are criticized for frequently exhibiting societal stereoty

OPSDL: On-Policy Self-Distillation for Long-Context Language Models

SafetyDGX agent

arXiv:2604.17535v1 Announce Type: new Abstract: Extending the effective context length of large language models (LLMs) remains a central challenge for real-world applications. While recent post-traini

OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection

SafetyDGX agent

arXiv:2511.21064v2 Announce Type: replace-cross Abstract: Open-Vocabulary Object Detection (OVOD) aims to enable detectors to generalize across categories by leveraging semantic information. Although

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling

SafetyDGX agent

arXiv:2510.24235v3 Announce Type: replace Abstract: Reward models (RMs) are central to reinforcement learning from human feedback (RLHF), providing the critical supervision signals that align large la

Peerispect: Claim Verification in Scientific Peer Reviews

SafetyDGX agent

arXiv:2604.17667v1 Announce Type: new Abstract: Peer review is central to scientific publishing, yet reviewers frequently include claims that are subjective, rhetorical, or misaligned with the submitt

PEPR: Privileged Event-based Predictive Regularization for Domain Generalization

SafetyDGX agent

arXiv:2602.04583v2 Announce Type: replace Abstract: Deep neural networks for visual perception are highly susceptible to domain shift, which poses a critical challenge for real-world deployment under

Plasticity Loss in Deep Reinforcement Learning: A Survey

SafetyDGX agent

arXiv:2411.04832v3 Announce Type: replace-cross Abstract: Plasticity refers to a network's ability to adapt to changing data distributions, which is crucial for the successful training of deep reinfor

Poetry overheard on http://B.sky:

SafetyDGX agent

Gary Marcus shared an observation or commentary about poetry encountered on a platform or service referenced as 'B.sky' (likely Bluesky, the decentralized social network). The post appears to document

Policy Testing in Markov Decision Processes

SafetyDGX agent

arXiv:2505.15342v2 Announce Type: replace-cross Abstract: We study the policy testing problem in discounted Markov decision processes (MDPs) in the fixed-confidence setting under a generative model wi

PoliLegalLM: A Technical Report on a Large Language Model for Political and Legal Affairs

SafetyDGX agent

arXiv:2604.17543v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable success in general-domain tasks, yet their direct application to the legal domain remains challeng

PowerCLIP: Powerset Alignment for Contrastive Pre-Training

SafetyDGX agent

arXiv:2511.23170v5 Announce Type: replace Abstract: Contrastive vision-language pre-training frameworks such as CLIP have demonstrated impressive zero-shot performance across a range of vision-languag

PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment

SafetyDGX agent

arXiv:2508.05132v2 Announce Type: replace Abstract: As medical LLMs transition to clinical deployment, assessing their ethical reasoning capability becomes critical. While achieving high accuracy on k

Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models

SafetyDGX agent

arXiv:2601.15220v2 Announce Type: replace Abstract: We identify a novel phenomenon in language models: benign fine-tuning of frontier models can lead to privacy collapse. We find that diverse, subtle

ProtoCLIP: Prototype-Aligned Latent Refinement for Robust Zero-Shot Chest X-Ray Classification

SafetyDGX agent

arXiv:2604.18444v1 Announce Type: cross Abstract: Zero-shot vision-language models (VLMs) have shown promise for chest radiograph classification, but their performance is often limited by confounding

Q-SINDy: Quantum-Kernel Sparse Identification of Nonlinear Dynamics with Provable Coefficient Debiasing

SafetyDGX agent

arXiv:2604.16779v1 Announce Type: cross Abstract: Quantum feature maps offer expressive embeddings for classical learning tasks, and augmenting sparse identification of nonlinear dynamics (SINDy) with

R3D2: Realistic 3D Asset Insertion via Diffusion for Autonomous Driving Simulation

SafetyDGX agent

arXiv:2506.07826v2 Announce Type: replace Abstract: Validating autonomous driving (AD) systems requires diverse and safety-critical testing, making photorealistic virtual environments essential. Tradi

RAYEN: Imposition of Hard Convex Constraints on Neural Networks

SafetyDGX agent

arXiv:2307.08336v2 Announce Type: replace Abstract: Despite the numerous applications of convex constraints in Robotics, enforcing them within learning-based frameworks remains an open challenge. Exis

Reasoning on the Manifold: Bidirectional Consistency for Self-Verification in Diffusion Language Models

SafetyDGX agent

arXiv:2604.16565v1 Announce Type: new Abstract: While Diffusion Large Language Models (dLLMs) offer structural advantages for global planning, efficiently verifying that they arrive at correct answers

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction

SafetyDGX agent

arXiv:2506.01770v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have achieved tremendous success in various tasks, yet concerns about their safety and security have emerged. In

RemoteShield: Enable Robust Multimodal Large Language Models for Earth Observation

SafetyDGX agent

arXiv:2604.17243v1 Announce Type: new Abstract: A robust Multimodal Large Language Model (MLLM) for Earth Observation should maintain consistent interpretation and reasoning under realistic input vari

Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning

SafetyDGX agent

arXiv:2601.14750v3 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting has achieved remarkable success in unlocking the reasoning capabilities of Large Language Models (LLMs). Although C

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring

SafetyDGX agent

arXiv:2512.12069v3 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) are vulnerable to a growing array of multimodal jailbreak attacks, necessitating defenses that are both g

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction

SafetyDGX agent

arXiv:2604.17328v1 Announce Type: new Abstract: This paper investigates the length problem in sequence-level relative reinforcement learning. We observe that, although existing methods partially allev

Retrieval-Augmented Multimodal Model for Fake News Detection

SafetyDGX agent

arXiv:2604.18112v1 Announce Type: new Abstract: In recent years, multimodal multidomain fake news detection has garnered increasing attention. Nevertheless, this direction presents two significant cha

Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF

SafetyDGX agent

arXiv:2604.17769v1 Announce Type: new Abstract: Ensuring the safety of large language models (LLMs) requires robust red teaming, yet the systematic synthesis of high-quality toxic data remains under-e

Reward Score Matching: Unifying Reward-based Fine-tuning for Flow and Diffusion Models

SafetyDGX agent

arXiv:2604.17415v1 Announce Type: cross Abstract: Reward-based fine-tuning aims to steer a pretrained diffusion or flow-based generative model toward higher-reward samples while remaining close to the

REZE: Representation Regularization for Domain-adaptive Text Embedding Pre-finetuning

SafetyDGX agent

arXiv:2604.17257v1 Announce Type: new Abstract: Recent text embedding models are often adapted to specialized domains via contrastive pre-finetuning (PFT) on a naive collection of scattered, heterogen

RISC-V Functional Safety for Autonomous Automotive Systems: An Analytical Framework and Research Roadmap for ML-Assisted Certification

SafetyDGX agent

arXiv:2604.17391v1 Announce Type: cross Abstract: RISC-V is emerging as a viable platform for automotive-grade embedded computing, with recent ISO 26262 ASIL-D certifications demonstrating readiness f

Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors

SafetyDGX agent

arXiv:2601.15625v2 Announce Type: replace Abstract: Large language models (LLMs) can call tools effectively, yet they remain brittle in multi-turn execution: after a tool-call error, smaller models of

S-GRPO: Unified Post-Training for Large Vision-Language Models

SafetyDGX agent

arXiv:2604.16557v1 Announce Type: cross Abstract: Current post-training methodologies for adapting Large Vision-Language Models (LVLMs) generally fall into two paradigms: Supervised Fine-Tuning (SFT)

SafeLM: Unified Privacy-Aware Optimization for Trustworthy Federated Large Language Models

SafetyDGX agent

arXiv:2604.16606v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in high-stakes domains, yet a unified treatment of their overlapping safety challenges remains

SaFeR-Steer: Evolving Multi-Turn MLLMs via Synthetic Bootstrapping and Feedback Dynamics

SafetyDGX agent

arXiv:2604.16358v1 Announce Type: cross Abstract: MLLMs are increasingly deployed in multi-turn settings, where attackers can escalate unsafe intent through the evolving visual-text history and exploi

Safer Trajectory Planning with CBF-guided Diffusion Model for Unmanned Aerial Vehicles

SafetyDGX agent

arXiv:2604.17527v1 Announce Type: new Abstract: Safe and agile trajectory planning is essential for autonomous systems, especially during complex aerobatic maneuvers. Motivated by the recent success o

Safety, Security, and Cognitive Risks in State-Space Models: A Systematic Threat Analysis with Spectral, Stateful, and Capacity Attacks

SafetyDGX agent

arXiv:2604.16424v1 Announce Type: cross Abstract: State-Space Models (SSMs) -- structured SSMs (S4, S4D, DSS, S5), selective SSMs (Mamba, Mamba-2), and hybrid architectures (Jamba) -- are deployed in

Scalable Neighborhood-Based Multi-Agent Actor-Critic

SafetyDGX agent

arXiv:2604.18190v1 Announce Type: new Abstract: We propose MADDPG-K, a scalable extension to Multi-Agent Deep Deterministic Policy Gradient (MADDPG) that addresses the computational limitations of cen

Scalable Physics-Informed Neural Differential Equations and Data-Driven Algorithms for HVAC Systems

SafetyDGX agent

arXiv:2604.18438v1 Announce Type: new Abstract: We present a scalable, data-driven simulation framework for large-scale heating, ventilation, and air conditioning (HVAC) systems that couples physics-i

See Through the Noise: Improving Domain Generalization in Gaze Estimation

SafetyDGX agent

arXiv:2604.16562v1 Announce Type: new Abstract: Generalizable gaze estimation methods have garnered increasing attention due to their critical importance in real-world applications and have achieved s

SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection

SafetyDGX agent

arXiv:2604.18476v1 Announce Type: new Abstract: Camera-only 3D object detection has emerged as a cost-effective and scalable alternative to LiDAR for autonomous driving, yet existing methods primarily

Sharpening Lightweight Models for Generalized Polyp Segmentation: A Boundary Guided Distillation from Foundation Models

SafetyDGX agent

arXiv:2604.17865v1 Announce Type: new Abstract: Automated polyp segmentation is critical for early colorectal cancer detection and its prevention, yet remains challenging due to weak boundaries, large

Shepherding UAV Swarm with Action Prediction Based on Movement Constraints

SafetyDGX agent

arXiv:2604.17189v1 Announce Type: new Abstract: In this study, we propose a new sheepdog-inspired control method for a swarm of small unmanned aerial vehicles (UAVs), which predicts the swarm behavior

Soft Label Pruning and Quantization for Large-Scale Dataset Distillation

SafetyDGX agent

arXiv:2604.18135v1 Announce Type: new Abstract: Large-scale dataset distillation requires storing auxiliary soft labels that can be 30-40x larger on ImageNet-1K and 200x larger on ImageNet-21K than th

Source-Free Domain Adaptation with Vision-Language Prior

SafetyDGX agent

arXiv:2604.17748v1 Announce Type: new Abstract: Source-Free Domain Adaptation (SFDA) seeks to adapt a source model, which is pre-trained on a supervised source domain, for a target domain, with only a

(Sparse) Attention to the Details: Preserving Spectral Fidelity in ML-based Weather Forecasting Models

SafetyDGX agent

arXiv:2604.16429v1 Announce Type: cross Abstract: We introduce Mosaic, a probabilistic weather forecasting model that addresses two principal sources of spectral degradation in ML-based weather predic

← Previous
1…189190191192193…212
Next →