AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,600 results
15 Apr 2026

Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training

SafetyDGX agent

arXiv:2509.25758v2 Announce Type: replace Abstract: The remarkable capabilities of modern large reasoning models are largely unlocked through post-training techniques such as supervised fine-tuning (S

This is what a black politician in South Africa said… “We will kill white women, we will kill white children, and we will even kill your pet…

SafetyDGX agent

This is what a black politician in South Africa said… “We will kill white women, we will kill white children, and we will even kill your pets' Violent language entering politics is a serious warning s

Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via Sequence-Level Likelihood

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.12736v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has significantly advanced the reasoning ability of large language models (LLMs), particularly in their mathem

Towards Generalized Certified Robustness with Multi-Norm Training

SafetyDGX agent

arXiv:2410.03000v3 Announce Type: replace Abstract: Existing certified training methods can only train models to be robust against a certain perturbation type (e.g. l_infty or l_2). However, an l_inft

Towards Platonic Representation for Table Reasoning: A Foundation for Permutation-Invariant Retrieval

SafetyDGX agent

arXiv:2604.12133v1 Announce Type: new Abstract: Historical approaches to Table Representation Learning (TRL) have largely adopted the sequential paradigms of Natural Language Processing (NLP). We argu

Uncertainty-Aware Image Classification In Biomedical Imaging Using Spectral-normalized Neural Gaussian Processes

SafetyDGX agent

arXiv:2602.02370v2 Announce Type: replace Abstract: Accurate histopathologic interpretation is key for clinical decision-making; however, current deep learning models for digital pathology are often o

Understanding and Improving Continuous Adversarial Training for LLMs via In-context Learning Theory

SafetyDGX agent

arXiv:2604.12817v1 Announce Type: new Abstract: Adversarial training (AT) is an effective defense for large language models (LLMs) against jailbreak attacks, but performing AT on LLMs is costly. To im

WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces

SafetyDGX agent

arXiv:2603.05295v3 Announce Type: replace Abstract: We introduce WebChain, the largest open-source dataset of human-annotated trajectories on real-world websites, designed to accelerate reproducible r

What happens when you systematically oversell the value of your product for years, while pretty much screwing society along the way? Eventua…

SafetyDGX agent

What happens when you systematically oversell the value of your product for years, while pretty much screwing society along the way? Eventually your customers figure it out. Surprisingly, people see A

Whole-Body Mobile Manipulation using Offline Reinforcement Learning on Sub-optimal Controllers

SafetyDGX agent

arXiv:2604.12509v1 Announce Type: cross Abstract: Mobile Manipulation (MoMa) of articulated objects, such as opening doors, drawers, and cupboards, demands simultaneous, whole-body coordination betwee

WiseOWL: A Methodology for Evaluating Ontological Descriptiveness and Semantic Correctness for Ontology Reuse and Ontology Recommendations

SafetyDGX agent

arXiv:2604.12025v1 Announce Type: new Abstract: The Semantic Web standardizes concept meaning for humans and machines, enabling machine-operable content and consistent interpretation that improves adv

XRZero-G0: Pushing the Frontier of Dexterous Robotic Manipulation with Interfaces, Quality and Ratios

SafetyDGX agent

arXiv:2604.13001v1 Announce Type: new Abstract: The acquisition of high-quality, action-aligned demonstration data remains a fundamental bottleneck in scaling foundation models for dexterous robot man

14 Apr 2026

3D Multi-View Stylization with Pose-Free Correspondences Matching for Robust 3D Geometry Preservation

SafetyDGX agent

arXiv:2604.09639v1 Announce Type: new Abstract: Artistic style transfer is well studied for images and videos, but extending it to multi-view 3D scenes remains difficult because stylization can disrup

A Comparative Theoretical Analysis of Entropy Control Methods in Reinforcement Learning

SafetyDGX agent

arXiv:2604.09676v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a key approach for enhancing reasoning in large language models (LLMs), yet scalable training is often hindered

A Dual-Positive Monotone Parameterization for Multi-Segment Bids and a Validity Assessment Framework for Reinforcement Learning Agent-based Simulation of Electricity Markets

SafetyDGX agent

arXiv:2604.10252v1 Announce Type: new Abstract: Reinforcement learning agent-based simulation (RL-ABS) has become an important tool for electricity market mechanism analysis and evaluation. In the mod

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning

SafetyDGX agent

arXiv:2601.16399v5 Announce Type: replace Abstract: We study a structured bi-level optimization problem where the upper-level objective is a smooth function and the lower-level problem is policy optim

A Mamba-Based Multimodal Network for Multiscale Blast-Induced Rapid Structural Damage Assessment

SafetyDGX agent

arXiv:2604.11709v1 Announce Type: new Abstract: Accurate and rapid structural damage assessment (SDA) is crucial for post-disaster management, helping responders prioritise resources, plan rescues, an

A mathematical theory of evolution for self-designing AIs

SafetyDGX agent

arXiv:2604.05142v2 Announce Type: replace Abstract: As artificial intelligence systems (AIs) become increasingly produced by recursive self-improvement, a form of evolution may emerge, with the traits

A Multilingual Dataset and Empirical Validation for the Mutual Reinforcement Effect in Information Extraction

SafetyDGX agent

arXiv:2407.10953v5 Announce Type: replace Abstract: The Mutual Reinforcement Effect (MRE) describes a phenomenon in information extraction where word-level and sentence-level tasks can mutually improv

A profile of BusPatrol, whose AI-powered cameras on 35K+ school buses in 24 US states record vehicles passing illegally, claiming they help reduce violations (Byard Duncan/Bloomberg)

SafetyDGX agent

Byard Duncan / Bloomberg: A profile of BusPatrol, whose AI-powered cameras on 35K+ school buses in 24 US states record vehicles passing illegally, claiming they help reduce violations — BusPatrol says

A Proposed Biomedical Data Policy Framework to Reduce Fragmentation, Improve Quality, and Incentivize Sharing in Indian Healthcare in the era of Artificial Intelligence and Digital Health

SafetyDGX agent

arXiv:2604.11125v1 Announce Type: new Abstract: India generates vast biomedical data through postgraduate research, government hospital services and audits, government schemes, private hospitals and t

A Queueing-Theoretic Framework for Dynamic Attack Surfaces: Data-Integrated Risk Analysis and Adaptive Defense

SafetyDGX agent

arXiv:2604.10427v1 Announce Type: cross Abstract: We develop a queueing-theoretic framework to model the temporal evolution of cyber-attack surfaces, where the number of active vulnerabilities is repr

Active Diffusion Matching: Score-based Iterative Alignment of Cross-Modal Retinal Images

SafetyDGX agent

arXiv:2604.10084v1 Announce Type: new Abstract: Objective: The study aims to address the challenge of aligning Standard Fundus Images (SFIs) and Ultra-Widefield Fundus Images (UWFIs), which is difficu

Adaptive Bidding Policies for First-Price Auctions with Budget Constraints under Non-stationarity

SafetyDGX agent

arXiv:2505.02796v2 Announce Type: replace-cross Abstract: We study how a budget-constrained bidder should learn to adaptively bid in repeated first-price auctions to maximize her cumulative payoff. Th

AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Afford Correspondence

SafetyDGX agent

arXiv:2604.10579v1 Announce Type: cross Abstract: Despite the recent success of modern imitation learning methods in robot manipulation, their performance is often constrained by geometric variations

Agentic Video Generation: From Text to Executable Event Graphs via Tool-Constrained LLM Planning

SafetyDGX agent

arXiv:2604.10383v1 Announce Type: new Abstract: Existing multi-agent video generation systems use LLM agents to orchestrate neural video generators, producing visually impressive but semantically unre

AI Integrity: A New Paradigm for Verifiable AI Governance

SafetyDGX agent

arXiv:2604.11065v1 Announce Type: new Abstract: AI systems increasingly shape high-stakes decisions in healthcare, law, defense, and education, yet existing governance paradigms -- AI Ethics, AI Safet

AI Organizations are More Effective but Less Aligned than Individual Agents

SafetyDGX agent

arXiv:2604.10290v1 Announce Type: new Abstract: AI is increasingly deployed in multi-agent systems; however, most research considers only the behavior of individual models. We experimentally show that

Anthropic believes that good transparency legislation needs to ensure public safety and accountability for the companies developing this pow…

SafetyDGX agent

Anthropic believes that good transparency legislation needs to ensure public safety and accountability for the companies developing this powerful technology, not provide a get-out-of-jail-free card ag

Anthropic details using AI agents to accelerate alignment research on 'weak-to-strong supervision', where a weak model supervises the training of a stronger one (Anthropic)

SafetyDGX agent

Anthropic: Anthropic details using AI agents to accelerate alignment research on “weak-to-strong supervision”, where a weak model supervises the training of a stronger one — Large language models' eve

Anthropogenic Regional Adaptation in Multimodal Vision-Language Model

SafetyDGX agent

arXiv:2604.11490v1 Announce Type: new Abstract: While the field of vision-language (VL) has achieved remarkable success in integrating visual and textual information across multiple languages and doma

// Artifacts as Memory Beyond the Agent Boundary // An agent doesn't always need a bigger memory buffer. Sometimes the environment itself re…

SafetyDGX agent

// Artifacts as Memory Beyond the Agent Boundary // An agent doesn't always need a bigger memory buffer. Sometimes the environment itself remembers on the agent's behalf. New research formalizes this

ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models

SafetyDGX agent

arXiv:2604.10065v1 Announce Type: cross Abstract: End-to-end full-duplex Speech Language Models (SLMs) require precise turn-taking for natural interaction. However, optimizing temporal dynamics via st

Assessing Model-Agnostic XAI Methods against EU AI Act Explainability Requirements

SafetyDGX agent

arXiv:2604.09628v1 Announce Type: cross Abstract: Explainable AI (XAI) has evolved in response to expectations and regulations, such as the EU AI Act, which introduces regulatory requirements on AI-po

Auto-regressive transformation for image alignment

SafetyDGX agent

arXiv:2505.04864v2 Announce Type: replace-cross Abstract: Existing methods for image alignment struggle in cases involving feature-sparse regions, extreme scale and field-of-view differences, and larg

Autonomous Diffractometry Enabled by Visual Reinforcement Learning

SafetyDGX agent

arXiv:2604.11773v1 Announce Type: cross Abstract: Automation underpins progress across scientific and industrial disciplines. Yet, automating tasks requiring interpretation of abstract visual informat

AWARE: Adaptive Whole-body Active Rotating Control for Enhanced LiDAR-Inertial Odometry under Human-in-the-Loop Interaction

SafetyDGX agent

arXiv:2604.10598v1 Announce Type: new Abstract: Human-in-the-loop (HITL) UAV operation is essential in complex and safety-critical aerial surveying environments, where human operators provide navigati

Awesome work by @jiaxinwen22, @liangqiu_1994, Joe Benton, and @janhkirchner! For more details, check out the blog post 👇 https://anthropic.…

SafetyDGX agent

Jan Leike praised collaborative work by researchers Jiaxin Wen, Liang Qiu, Joe Benton, and Jan Kirchner, directing followers to an Anthropic blog post for further details. The post appears to highligh

Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward

SafetyDGX agent

arXiv:2604.09748v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is an emerging paradigm that significantly boosts a Large Language Model's (LLM's) reasoning abi

Belief-Aware VLM Model for Human-like Reasoning

SafetyDGX agent

arXiv:2604.09686v1 Announce Type: new Abstract: Traditional neural network models for intent inference rely heavily on observable states and struggle to generalize across diverse tasks and dynamic env

Belief-State RWKV for Reinforcement Learning under Partial Observability

SafetyDGX agent

arXiv:2604.09671v1 Announce Type: new Abstract: We propose a stronger formulation of RL on top of RWKV-style recurrent sequence models, in which the fixed-size recurrent state is explicitly interprete

Below-ground Fungal Biodiversity Can be Monitored Using Self-Supervised Learning Satellite Features

SafetyDGX agent

arXiv:2604.09818v1 Announce Type: new Abstract: Mycorrhizal fungi are vital to terrestrial ecosystem functioning. Yet monitoring their biodiversity at landscape scales is often unfeasible due to time

Beyond Compliance: A Resistance-Informed Motivation Reasoning Framework for Challenging Psychological Client Simulation

SafetyDGX agent

arXiv:2604.10507v1 Announce Type: new Abstract: Psychological client simulators have emerged as a scalable solution for training and evaluating counselor trainees and psychological LLMs. Yet existing

Beyond Message Passing: A Semantic View of Agent Communication Protocols

SafetyDGX agent

arXiv:2604.02369v3 Announce Type: replace-cross Abstract: Agent communication protocols are becoming critical infrastructure for large language model (LLM) systems that must use tools, coordinate with

Beyond Monologue: Interactive Talking-Listening Avatar Generation with Conversational Audio Context-Aware Kernels

SafetyDGX agent

arXiv:2604.10367v1 Announce Type: new Abstract: Audio-driven human video generation has achieved remarkable success in monologue scenarios, largely driven by advancements in powerful video generation

Beyond Reconstruction: Reconstruction-to-Vector Diffusion for Hyperspectral Anomaly Detection

SafetyDGX agent

arXiv:2604.11390v1 Announce Type: new Abstract: While Hyperspectral Anomaly Detection (HAD) excels at identifying sparse targets in complex scenes, existing models remain trapped in a scalar 'reconstr

Bidirectional Learning of Facial Action Units and Expressions via Structured Semantic Mapping across Heterogeneous Datasets

SafetyDGX agent

arXiv:2604.10541v1 Announce Type: new Abstract: Facial action unit (AU) detection and facial expression (FE) recognition can be jointly viewed as affective facial behavior tasks, representing fine-gra

Binary Flow Matching: Prediction-Loss Space Alignment for Robust Learning

SafetyDGX agent

arXiv:2602.10420v2 Announce Type: replace Abstract: Flow matching has emerged as a powerful framework for generative modeling, with recent empirical successes highlighting the effectiveness of signal-

bioLeak: Leakage-Aware Modeling and Diagnostics for Machine Learning in R

SafetyDGX agent

arXiv:2604.10965v1 Announce Type: cross Abstract: Data leakage remains a recurrent source of optimistic bias in biomedical machine learning studies. Standard row-wise cross-validation and globally est

Brain-Grasp: Graph-based Saliency Priors for Improved fMRI-based Visual Brain Decoding

SafetyDGX agent

arXiv:2604.10617v1 Announce Type: cross Abstract: Recent progress in brain-guided image generation has improved the quality of fMRI-based reconstructions; however, fundamental challenges remain in pre

Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance

SafetyDGX agent

arXiv:2604.10590v1 Announce Type: cross Abstract: Multilingual Large Language Models (LLMs) struggle with cross-lingual tasks due to data imbalances between high-resource and low-resource languages, a

Bridging the RGB-IR Gap: Consensus and Discrepancy Modeling for Text-Guided Multispectral Detection

SafetyDGX agent

arXiv:2604.11234v1 Announce Type: new Abstract: Text-guided multispectral object detection uses text semantics to guide semantic-aware cross-modal interaction between RGB and IR for more robust percep

Budget-Aware Uncertainty for Radiotherapy Segmentation QA Using nnU-Net

SafetyDGX agent

arXiv:2604.11798v1 Announce Type: cross Abstract: Accurate delineation of the Clinical Target Volume (CTV) is essential for radiotherapy planning, yet remains time-consuming and difficult to assess, e

C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis

SafetyDGX agent

arXiv:2604.00013v2 Announce Type: replace-cross Abstract: Multimodal sentiment analysis aims to integrate textual, acoustic, and visual information for deep emotional understanding. Despite the progre

CAGenMol: Condition-Aware Diffusion Language Model for Goal-Directed Molecular Generation

SafetyDGX agent

arXiv:2604.11483v1 Announce Type: new Abstract: Goal-directed molecular generation requires satisfying heterogeneous constraints such as protein--ligand compatibility and multi-objective drug-like pro

Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs

SafetyDGX agent

arXiv:2604.10585v1 Announce Type: cross Abstract: Modern large language models (LLMs) are increasingly fine-tuned via reinforcement learning from human feedback (RLHF) or related reward optimisation s

China leaps ahead of the US in the race to control dangerously anthropomorphic AI.

SafetyDGX agent

China leaps ahead of the US in the race to control dangerously anthropomorphic AI. 🚨 BREAKING: China's new law on AI anthropomorphism has been officially enacted, and it is the world's STRICTEST law o

CID-TKG: Collaborative Historical Invariance and Evolutionary Dynamics Learning for Temporal Knowledge Graph Reasoning

SafetyDGX agent

arXiv:2604.09600v1 Announce Type: new Abstract: Temporal knowledge graph (TKG) reasoning aims to infer future facts at unseen timestamps from temporally evolving entities and relations. Despite recent

CityGuard: Graph-Aware Private Descriptors for Bias-Resilient Identity Search Across Urban Cameras

SafetyDGX agent

arXiv:2602.18047v3 Announce Type: replace Abstract: City-scale person re-identification across distributed cameras must handle severe appearance changes from viewpoint, occlusion, and domain shift whi

Claim2Vec: Embedding Fact-Check Claims for Multilingual Similarity and Clustering

SafetyDGX agent

arXiv:2604.09812v1 Announce Type: new Abstract: Recurrent claims present a major challenge for automated fact-checking systems designed to combat misinformation, especially in multilingual settings. W

← Previous
1…197198199200201…210
Next →