AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,351 results
15 Apr 2026

Scaffold-Conditioned Preference Triplets for Controllable Molecular Optimization with Large Language Models

SafetyDGX agent

arXiv:2604.12350v1 Announce Type: cross Abstract: Molecular property optimization is central to drug discovery, yet many deep learning methods rely on black-box scoring and offer limited control over

Scalable and General Whole-Body Control for Cross-Humanoid Locomotion

SafetyDGX agent

arXiv:2602.05791v2 Announce Type: replace Abstract: Learning-based whole-body controllers have become a key driver for humanoid robots, yet most existing approaches require robot-specific training. In

Schema-Adaptive Tabular Representation Learning with LLMs for Generalizable Multimodal Clinical Reasoning

SafetyDGX agent

arXiv:2604.11835v1 Announce Type: cross Abstract: Machine learning for tabular data remains constrained by poor schema generalization, a challenge rooted in the lack of semantic understanding of struc

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

SafetyDGX agent

arXiv:2604.12002v1 Announce Type: new Abstract: Current post-training methods in verifiable settings fall into two categories. Reinforcement learning (RLVR) relies on binary rewards, which are broadly

Simulation as Supervision: Mechanistic Pretraining for Scientific Discovery

SafetyDGX agent

arXiv:2507.08977v4 Announce Type: replace-cross Abstract: Scientific modeling faces a tradeoff between the interpretability of mechanistic theory and the predictive power of machine learning. While ex

So true. “Being right too soon is socially unacceptable”

SafetyDGX agent

So true. “Being right too soon is socially unacceptable” I remember reading this from Heinlein as a young man and it took experience for me to fully understand it. People and institutions will persist

SOAR: Self-Correction for Optimal Alignment and Refinement in Diffusion Models

SafetyDGX agent

arXiv:2604.12617v1 Announce Type: cross Abstract: The post-training pipeline for diffusion models currently has two stages: supervised fine-tuning (SFT) on curated data and reinforcement learning (RL)

StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback

SafetyDGX agent

arXiv:2510.20093v2 Announce Type: replace-cross Abstract: Although recent advancements in diffusion models have significantly enriched the quality of generated images, challenges remain in synthesizin

Task Alignment: A simple and effective proxy for model merging in computer vision

SafetyDGX agent

arXiv:2604.12935v1 Announce Type: new Abstract: Efficiently merging several models fine-tuned for different tasks, but stemming from the same pretrained base model, is of great practical interest. Des

Teaching LLMs Human-Like Editing of Inappropriate Argumentation via Reinforcement Learning

SafetyDGX agent

arXiv:2604.12770v1 Announce Type: new Abstract: Editing human-written text has become a standard use case of large language models (LLMs), for example, to make one's arguments more appropriate for a d

The front door to the internet hasn't changed -- but the person going through that door has changed from a human browsing 5 blue links, to a…

SafetyDGX agent

The front door to the internet hasn't changed -- but the person going through that door has changed from a human browsing 5 blue links, to an agent browsing a similar index in vastly different ways. A

The role of System 1 and System 2 semantic memory structure in human and LLM biases

SafetyDGX agent

arXiv:2604.12816v1 Announce Type: new Abstract: Implicit biases in both humans and large language models (LLMs) pose significant societal risks. Dual process theories propose that biases arise primari

The Stackelberg Speaker: Optimizing Persuasive Communication in Social Deduction Games

SafetyDGX agent

arXiv:2510.09087v2 Announce Type: replace Abstract: Large language model (LLM) agents have shown remarkable progress in social deduction games (SDGs). However, existing approaches primarily focus on i

Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training

SafetyDGX agent

arXiv:2509.25758v2 Announce Type: replace Abstract: The remarkable capabilities of modern large reasoning models are largely unlocked through post-training techniques such as supervised fine-tuning (S

Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via Sequence-Level Likelihood

SafetyDGX agent

arXiv:2604.12736v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has significantly advanced the reasoning ability of large language models (LLMs), particularly in their mathem

Towards Generalized Certified Robustness with Multi-Norm Training

SafetyDGX agent

arXiv:2410.03000v3 Announce Type: replace Abstract: Existing certified training methods can only train models to be robust against a certain perturbation type (e.g. l_infty or l_2). However, an l_inft

Towards Platonic Representation for Table Reasoning: A Foundation for Permutation-Invariant Retrieval

SafetyDGX agent

arXiv:2604.12133v1 Announce Type: new Abstract: Historical approaches to Table Representation Learning (TRL) have largely adopted the sequential paradigms of Natural Language Processing (NLP). We argu

WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces

SafetyDGX agent

arXiv:2603.05295v3 Announce Type: replace Abstract: We introduce WebChain, the largest open-source dataset of human-annotated trajectories on real-world websites, designed to accelerate reproducible r

What happens when you systematically oversell the value of your product for years, while pretty much screwing society along the way? Eventua…

SafetyDGX agent

What happens when you systematically oversell the value of your product for years, while pretty much screwing society along the way? Eventually your customers figure it out. Surprisingly, people see A

Whole-Body Mobile Manipulation using Offline Reinforcement Learning on Sub-optimal Controllers

SafetyDGX agent

arXiv:2604.12509v1 Announce Type: cross Abstract: Mobile Manipulation (MoMa) of articulated objects, such as opening doors, drawers, and cupboards, demands simultaneous, whole-body coordination betwee

WiseOWL: A Methodology for Evaluating Ontological Descriptiveness and Semantic Correctness for Ontology Reuse and Ontology Recommendations

SafetyDGX agent

arXiv:2604.12025v1 Announce Type: new Abstract: The Semantic Web standardizes concept meaning for humans and machines, enabling machine-operable content and consistent interpretation that improves adv

XRZero-G0: Pushing the Frontier of Dexterous Robotic Manipulation with Interfaces, Quality and Ratios

SafetyDGX agent

arXiv:2604.13001v1 Announce Type: new Abstract: The acquisition of high-quality, action-aligned demonstration data remains a fundamental bottleneck in scaling foundation models for dexterous robot man

14 Apr 2026

3D Multi-View Stylization with Pose-Free Correspondences Matching for Robust 3D Geometry Preservation

SafetyDGX agent

arXiv:2604.09639v1 Announce Type: new Abstract: Artistic style transfer is well studied for images and videos, but extending it to multi-view 3D scenes remains difficult because stylization can disrup

A Comparative Theoretical Analysis of Entropy Control Methods in Reinforcement Learning

SafetyDGX agent

arXiv:2604.09676v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a key approach for enhancing reasoning in large language models (LLMs), yet scalable training is often hindered

A Dual-Positive Monotone Parameterization for Multi-Segment Bids and a Validity Assessment Framework for Reinforcement Learning Agent-based Simulation of Electricity Markets

SafetyDGX agent

arXiv:2604.10252v1 Announce Type: new Abstract: Reinforcement learning agent-based simulation (RL-ABS) has become an important tool for electricity market mechanism analysis and evaluation. In the mod

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning

SafetyDGX agent

arXiv:2601.16399v5 Announce Type: replace Abstract: We study a structured bi-level optimization problem where the upper-level objective is a smooth function and the lower-level problem is policy optim

A mathematical theory of evolution for self-designing AIs

SafetyDGX agent

arXiv:2604.05142v2 Announce Type: replace Abstract: As artificial intelligence systems (AIs) become increasingly produced by recursive self-improvement, a form of evolution may emerge, with the traits

A Multilingual Dataset and Empirical Validation for the Mutual Reinforcement Effect in Information Extraction

SafetyDGX agent

arXiv:2407.10953v5 Announce Type: replace Abstract: The Mutual Reinforcement Effect (MRE) describes a phenomenon in information extraction where word-level and sentence-level tasks can mutually improv

A profile of BusPatrol, whose AI-powered cameras on 35K+ school buses in 24 US states record vehicles passing illegally, claiming they help reduce violations (Byard Duncan/Bloomberg)

SafetyDGX agent

Byard Duncan / Bloomberg: A profile of BusPatrol, whose AI-powered cameras on 35K+ school buses in 24 US states record vehicles passing illegally, claiming they help reduce violations — BusPatrol says

A Proposed Biomedical Data Policy Framework to Reduce Fragmentation, Improve Quality, and Incentivize Sharing in Indian Healthcare in the era of Artificial Intelligence and Digital Health

SafetyDGX agent

arXiv:2604.11125v1 Announce Type: new Abstract: India generates vast biomedical data through postgraduate research, government hospital services and audits, government schemes, private hospitals and t

A Queueing-Theoretic Framework for Dynamic Attack Surfaces: Data-Integrated Risk Analysis and Adaptive Defense

SafetyDGX agent

arXiv:2604.10427v1 Announce Type: cross Abstract: We develop a queueing-theoretic framework to model the temporal evolution of cyber-attack surfaces, where the number of active vulnerabilities is repr

Active Diffusion Matching: Score-based Iterative Alignment of Cross-Modal Retinal Images

SafetyDGX agent

arXiv:2604.10084v1 Announce Type: new Abstract: Objective: The study aims to address the challenge of aligning Standard Fundus Images (SFIs) and Ultra-Widefield Fundus Images (UWFIs), which is difficu

Adaptive Bidding Policies for First-Price Auctions with Budget Constraints under Non-stationarity

SafetyDGX agent

arXiv:2505.02796v2 Announce Type: replace-cross Abstract: We study how a budget-constrained bidder should learn to adaptively bid in repeated first-price auctions to maximize her cumulative payoff. Th

AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Afford Correspondence

SafetyDGX agent

arXiv:2604.10579v1 Announce Type: cross Abstract: Despite the recent success of modern imitation learning methods in robot manipulation, their performance is often constrained by geometric variations

Agentic Video Generation: From Text to Executable Event Graphs via Tool-Constrained LLM Planning

SafetyDGX agent

arXiv:2604.10383v1 Announce Type: new Abstract: Existing multi-agent video generation systems use LLM agents to orchestrate neural video generators, producing visually impressive but semantically unre

Anthropic details using AI agents to accelerate alignment research on 'weak-to-strong supervision', where a weak model supervises the training of a stronger one (Anthropic)

SafetyDGX agent

Anthropic: Anthropic details using AI agents to accelerate alignment research on “weak-to-strong supervision”, where a weak model supervises the training of a stronger one — Large language models' eve

// Artifacts as Memory Beyond the Agent Boundary // An agent doesn't always need a bigger memory buffer. Sometimes the environment itself re…

SafetyDGX agent

// Artifacts as Memory Beyond the Agent Boundary // An agent doesn't always need a bigger memory buffer. Sometimes the environment itself remembers on the agent's behalf. New research formalizes this

ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models

SafetyDGX agent

arXiv:2604.10065v1 Announce Type: cross Abstract: End-to-end full-duplex Speech Language Models (SLMs) require precise turn-taking for natural interaction. However, optimizing temporal dynamics via st

Assessing Model-Agnostic XAI Methods against EU AI Act Explainability Requirements

SafetyDGX agent

arXiv:2604.09628v1 Announce Type: cross Abstract: Explainable AI (XAI) has evolved in response to expectations and regulations, such as the EU AI Act, which introduces regulatory requirements on AI-po

Auto-regressive transformation for image alignment

SafetyDGX agent

arXiv:2505.04864v2 Announce Type: replace-cross Abstract: Existing methods for image alignment struggle in cases involving feature-sparse regions, extreme scale and field-of-view differences, and larg

Autonomous Diffractometry Enabled by Visual Reinforcement Learning

SafetyDGX agent

arXiv:2604.11773v1 Announce Type: cross Abstract: Automation underpins progress across scientific and industrial disciplines. Yet, automating tasks requiring interpretation of abstract visual informat

Belief-State RWKV for Reinforcement Learning under Partial Observability

SafetyDGX agent

arXiv:2604.09671v1 Announce Type: new Abstract: We propose a stronger formulation of RL on top of RWKV-style recurrent sequence models, in which the fixed-size recurrent state is explicitly interprete

Below-ground Fungal Biodiversity Can be Monitored Using Self-Supervised Learning Satellite Features

SafetyDGX agent

arXiv:2604.09818v1 Announce Type: new Abstract: Mycorrhizal fungi are vital to terrestrial ecosystem functioning. Yet monitoring their biodiversity at landscape scales is often unfeasible due to time

Beyond Compliance: A Resistance-Informed Motivation Reasoning Framework for Challenging Psychological Client Simulation

SafetyDGX agent

arXiv:2604.10507v1 Announce Type: new Abstract: Psychological client simulators have emerged as a scalable solution for training and evaluating counselor trainees and psychological LLMs. Yet existing

Beyond Message Passing: A Semantic View of Agent Communication Protocols

SafetyDGX agent

arXiv:2604.02369v3 Announce Type: replace-cross Abstract: Agent communication protocols are becoming critical infrastructure for large language model (LLM) systems that must use tools, coordinate with

Beyond Monologue: Interactive Talking-Listening Avatar Generation with Conversational Audio Context-Aware Kernels

SafetyDGX agent

arXiv:2604.10367v1 Announce Type: new Abstract: Audio-driven human video generation has achieved remarkable success in monologue scenarios, largely driven by advancements in powerful video generation

Beyond Reconstruction: Reconstruction-to-Vector Diffusion for Hyperspectral Anomaly Detection

SafetyDGX agent

arXiv:2604.11390v1 Announce Type: new Abstract: While Hyperspectral Anomaly Detection (HAD) excels at identifying sparse targets in complex scenes, existing models remain trapped in a scalar 'reconstr

Bidirectional Learning of Facial Action Units and Expressions via Structured Semantic Mapping across Heterogeneous Datasets

SafetyDGX agent

arXiv:2604.10541v1 Announce Type: new Abstract: Facial action unit (AU) detection and facial expression (FE) recognition can be jointly viewed as affective facial behavior tasks, representing fine-gra

Binary Flow Matching: Prediction-Loss Space Alignment for Robust Learning

SafetyDGX agent

arXiv:2602.10420v2 Announce Type: replace Abstract: Flow matching has emerged as a powerful framework for generative modeling, with recent empirical successes highlighting the effectiveness of signal-

bioLeak: Leakage-Aware Modeling and Diagnostics for Machine Learning in R

SafetyDGX agent

arXiv:2604.10965v1 Announce Type: cross Abstract: Data leakage remains a recurrent source of optimistic bias in biomedical machine learning studies. Standard row-wise cross-validation and globally est

Brain-Grasp: Graph-based Saliency Priors for Improved fMRI-based Visual Brain Decoding

SafetyDGX agent

arXiv:2604.10617v1 Announce Type: cross Abstract: Recent progress in brain-guided image generation has improved the quality of fMRI-based reconstructions; however, fundamental challenges remain in pre

Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance

SafetyDGX agent

arXiv:2604.10590v1 Announce Type: cross Abstract: Multilingual Large Language Models (LLMs) struggle with cross-lingual tasks due to data imbalances between high-resource and low-resource languages, a

Bridging the RGB-IR Gap: Consensus and Discrepancy Modeling for Text-Guided Multispectral Detection

SafetyDGX agent

arXiv:2604.11234v1 Announce Type: new Abstract: Text-guided multispectral object detection uses text semantics to guide semantic-aware cross-modal interaction between RGB and IR for more robust percep

Budget-Aware Uncertainty for Radiotherapy Segmentation QA Using nnU-Net

SafetyDGX agent

arXiv:2604.11798v1 Announce Type: cross Abstract: Accurate delineation of the Clinical Target Volume (CTV) is essential for radiotherapy planning, yet remains time-consuming and difficult to assess, e

C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis

SafetyDGX agent

arXiv:2604.00013v2 Announce Type: replace-cross Abstract: Multimodal sentiment analysis aims to integrate textual, acoustic, and visual information for deep emotional understanding. Despite the progre

Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs

SafetyDGX agent

arXiv:2604.10585v1 Announce Type: cross Abstract: Modern large language models (LLMs) are increasingly fine-tuned via reinforcement learning from human feedback (RLHF) or related reward optimisation s

China leaps ahead of the US in the race to control dangerously anthropomorphic AI.

SafetyDGX agent

China leaps ahead of the US in the race to control dangerously anthropomorphic AI. 🚨 BREAKING: China's new law on AI anthropomorphism has been officially enacted, and it is the world's STRICTEST law o

CID-TKG: Collaborative Historical Invariance and Evolutionary Dynamics Learning for Temporal Knowledge Graph Reasoning

SafetyDGX agent

arXiv:2604.09600v1 Announce Type: new Abstract: Temporal knowledge graph (TKG) reasoning aims to infer future facts at unseen timestamps from temporally evolving entities and relations. Despite recent

CityGuard: Graph-Aware Private Descriptors for Bias-Resilient Identity Search Across Urban Cameras

SafetyDGX agent

arXiv:2602.18047v3 Announce Type: replace Abstract: City-scale person re-identification across distributed cameras must handle severe appearance changes from viewpoint, occlusion, and domain shift whi

Claim2Vec: Embedding Fact-Check Claims for Multilingual Similarity and Clustering

SafetyDGX agent

arXiv:2604.09812v1 Announce Type: new Abstract: Recurrent claims present a major challenge for automated fact-checking systems designed to combat misinformation, especially in multilingual settings. W

← Previous
1…210211212213214…240
Next →