AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,812 results
9 Jun 2026

Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2606.09138v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) has become an important post-training paradigm for turning LLMs from static chatbots into interactive agents, giving

CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning

SafetyDGX agent

arXiv:2509.25004v2 Announce Type: replace Abstract: Online reinforcement learning with verifiable rewards (RLVR) has become an effective paradigm for improving the reasoning abilities of large languag

Code Is More Than Text: Uncertainty Estimation for Code Generation

SafetyDGX agent

arXiv:2606.09577v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as code generators, where silently wrong programs pose real safety and reliability risks. Relia


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Cold feet about coating the entire surface of the earth in data centers?

SafetyDGX agent

Gary Marcus likely discusses concerns about the environmental and practical implications of exponentially expanding data center infrastructure across the globe, questioning whether covering Earth's su

Comparative evaluation of training strategies using partially labelled datasets for segmentation of white matter hyperintensities and stroke lesions in FLAIR MRI

SafetyDGX agent

arXiv:2601.20503v2 Announce Type: replace-cross Abstract: White matter hyperintensities (WMH) and ischaemic stroke lesions (ISL) are key imaging biomarkers of cerebral small vessel disease (SVD) detec

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning

SafetyDGX agent

arXiv:2606.08088v1 Announce Type: new Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has recently become a key paradigm for improving the reasoning abilities of Large Language Models

Constrained Paraphrase Consistency for LLM Hallucination Detection

SafetyDGX agent

arXiv:2606.08158v1 Announce Type: cross Abstract: Large language models (LLMs) can generate factually inconsistent claims, motivating accurate and scalable hallucination detectors. Prior work largely

Constrained user-item allocation for e-commerce marketing campaigns

SafetyDGX agent

arXiv:2606.09623v1 Announce Type: new Abstract: When running marketing campaigns, retailers must decide which products to promote and which users to target. These decisions are inherently coupled: eff

Constraint-Aware Optimization for Robust Protein Stability Prediction

SafetyDGX agent

arXiv:2606.08100v1 Announce Type: new Abstract: Multimodal DeltaDelta G predictors integrating protein language models with inverse-folding representations achieve strong in-distribution accuracy on t

Context-Fractured Decomposition Attacks on Tool-Using LLM Agents: Exploiting Artifact Provenance Gaps

SafetyDGX agent

arXiv:2606.09084v1 Announce Type: cross Abstract: Tool-using LLM agents interact with the world through actions that persist state in artifacts (e.g., workspace files or logs). Consequently, jailbreak

Context Over Compute Human-in-the-Loop Outperforms Iterative Chain-of-Thought Prompting in Interview Answer Quality

SafetyDGX agent

arXiv:2603.09995v2 Announce Type: replace-cross Abstract: Behavioral interview evaluation using large language models presents unique challenges that require structured assessment, realistic interview

Contrast encodes inductive bias: separating slow noise from dynamics in predictive representation learning

SafetyDGX agent

arXiv:2606.07770v1 Announce Type: new Abstract: Self-supervised methods that learn representations and predict dynamics fully in the latent space, such as JEPA, have been shown to confuse slowly varyi

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers

SafetyDGX agent

arXiv:2606.07604v1 Announce Type: cross Abstract: Analyzing attention weights has become a standard approach for interpreting the information flow of Large Language Models (LLMs). However, this approa

Cooperative Long Rope Skipping via Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2606.08064v1 Announce Type: new Abstract: Humans exhibit remarkable motor agility, enabling a wide range of dynamic skills such as running and jumping, which highlights the great potential of hu

Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings

SafetyDGX agent

arXiv:2606.09409v1 Announce Type: new Abstract: Pairwise comparisons combined with aggregation methods like Elo have become central to evaluating generative models, yet concerns remain that they rewar

Cranio-Diff: Diffusion-based Cross-domain Craniofacial Reconstruction with 2D X-ray Skull Guidance and Structural Identity Constraints

SafetyDGX agent

arXiv:2606.09699v1 Announce Type: new Abstract: The state-of-the-art generative models, such as CycleGAN, Pix2Pix, and diffusion models have demonstrated remarkable performance in the face generation

Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing

SafetyDGX agent

arXiv:2606.07636v1 Announce Type: new Abstract: Editing a long-form video from heterogeneous footage requires more than selecting clips: an agent must preserve narrative intent across material prepara

Culturally-Adapted Red-Teaming Across East and Southeast Asian Contexts: A Methodological and Comparative Analysis

SafetyDGX agent

arXiv:2606.09178v1 Announce Type: cross Abstract: Multilingual safety evaluation of large language models (LLMs) has predominantly relied on direct translation (DT) of English benchmarks into target l

CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generation

SafetyDGX agent

arXiv:2601.15408v2 Announce Type: replace-cross Abstract: Medical vision-language models can automate the generation of radiology reports but struggle with accurate visual grounding and factual consis

Data Agents Under Attack: Vulnerabilities in LLM-Driven Analytical Systems

SafetyDGX agent

arXiv:2606.08661v1 Announce Type: cross Abstract: Data agents integrate LLM-driven reasoning with relational data access, executable analytical tools, and multi-step workflow orchestration, making the

Decentralized End-to-End Multi-AAV Pursuit Using Predictive Spatio-Temporal Observation via Deep Reinforcement Learning

SafetyDGX agent

arXiv:2603.24238v2 Announce Type: replace Abstract: Decentralized cooperative pursuit in cluttered environments is challenging for autonomous aerial swarms, especially under partial and noisy percepti

Decoupling Semantics and Logic: A Training-Free Coarse-to-Fine Pipeline for Video Retrieval-Augmented Generation

SafetyDGX agent

arXiv:2606.07924v1 Announce Type: cross Abstract: This paper presents our system description for the 2nd Workshop on Multimodal Augmented Generation via MultimodAl Retrieval (MAGMaR). Addressing the c

DexPIE: Stable Dexterous Policy Improvement from Real-World Experience

SafetyDGX agent

arXiv:2606.09615v1 Announce Type: cross Abstract: Dexterous manipulation presents substantial challenges for imitation learning due to its high-dimensional action space and complex contact-rich dynami

Diffuse AI Control on Fuzzy Tasks

SafetyDGX agent

arXiv:2606.08892v1 Announce Type: new Abstract: AI models deployed in critical domains, such as AI safety research, may subtly sabotage our efforts due to misalignment. Diffuse AI Control is a subfiel

Disentanglement with Holographic Reduced Representations

SafetyDGX agent

arXiv:2606.09725v1 Announce Type: new Abstract: Disentanglement, the separation of factors of variation in data using neural networks, remains a long-standing challenge in machine learning. Prior work

Distant Object Localisation from Noisy Image Segmentation Sequences

SafetyDGX agent

arXiv:2509.20906v3 Announce Type: replace Abstract: 3D object localisation based on a sequence of camera measurements is essential for safety-critical surveillance tasks, such as drone-based wildfire

Distilling LLM Reasoning into an Interpretable Policy Tree for Human-AI Collaboration

SafetyDGX agent

arXiv:2606.08596v1 Announce Type: new Abstract: Constructing efficient and reliable policies to assist humans is indispensable for human-AI collaboration. Existing methods mainly follow two lines of w

DIVERGE: Diversity-Enhanced RAG for Open-Ended Information Seeking

SafetyDGX agent

arXiv:2602.00238v2 Announce Type: replace-cross Abstract: Existing retrieval-augmented generation (RAG) systems often assume that each query has a single correct answer. This assumption overlooks open

Diverse Thinking Schemata Elicit Better Reasoning in Large Language Models

SafetyDGX agent

arXiv:2606.08974v1 Announce Type: new Abstract: Large reasoning models (LRMs) have attracted increasing attention for their ability to solve complex mathematical problems by generating extended reason

Do VLMs See What Sensors Feel? A Scalable Expert-Guided Design for Wheelchair Accessibility Assessment from Street View

SafetyDGX agent

arXiv:2606.07642v1 Announce Type: new Abstract: Assessing built-environment interaction, such as wheelchair accessibility, is difficult because real-world mobility is shaped by distributed, context-de

Does Persona Make LLMs K-pop Fans? A Pilot Study of LLM-Based Online Concert Audience Agents

SafetyDGX agent

arXiv:2606.07837v1 Announce Type: cross Abstract: A concert is a collective experience, but recorded performance videos are typically watched alone, stripping away the shared audience presence that ma

DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment

SafetyDGX agent

arXiv:2606.07678v1 Announce Type: cross Abstract: Safety alignment for large language models relies on preference data, but current pipelines often train on large, redundant datasets. Existing data se

@dpetrou @karpathy Yes. Locking in a permanent status quo power structure. Incredibly unsafe, and damaging for humanity's prospects.

SafetyDGX agent

Gary Marcus argues that establishing a permanent, locked-in power structure is fundamentally unsafe and harmful to humanity's long-term prospects. The statement appears to be part of a discussion with

Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition

SafetyDGX agent

arXiv:2603.12046v2 Announce Type: replace-cross Abstract: Audio-Visual Speech Recognition (AVSR) leverages both acoustic and visual information for robust recognition under noise. However, how models

Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation

SafetyDGX agent

arXiv:2606.08737v1 Announce Type: new Abstract: World action models inherit the predictive capability of world models, enabling action generation to be guided by anticipated future observations. Howev

DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning

SafetyDGX agent

arXiv:2606.08035v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a leading paradigm for enhancing visual reasoning in Multimodal Large Language Mode

EgoAERO: Learning Dexterous Manipulation from a Single Egocentric Video without Object Assets

SafetyDGX agent

arXiv:2606.08057v1 Announce Type: cross Abstract: Egocentric RGB-D videos offer a natural source of human dexterous manipulation demonstrations, but existing data is difficult to use for robot learnin

Emergent alignment and the projectability of ethical personas

SafetyDGX agent

arXiv:2606.09475v1 Announce Type: new Abstract: Work on `emergent misalignment' shows that finetuning LLMs on narrow tasks can induce broadly misaligned behavior. This supports the `persona selection'

Enhancing AI Interpretability and Safety through Localised Architectures

SafetyDGX agent

arXiv:2606.07998v1 Announce Type: cross Abstract: Recent advances in generative AI, especially powerful Large Language Models (LLMs) and Large Reasoning Models (LRMs), raise concerns over the interpre

Entropic Optimal Transport Eigenmaps for Nonlinear Alignment and Joint Embedding of High-Dimensional Datasets

SafetyDGX agent

arXiv:2407.01718v2 Announce Type: replace-cross Abstract: Embedding high-dimensional data into a low-dimensional space is an indispensable component of data analysis. In numerous applications, it is n

Escaping the KL Agreement Trap in On-Policy Distillation

SafetyDGX agent

arXiv:2606.09471v1 Announce Type: new Abstract: On-policy distillation (OPD) provides dense token-level supervision by asking a teacher to score student-generated rollouts. However, when the student d

Evaluating AI Investment Strategies

SafetyDGX agent

arXiv:2606.08791v1 Announce Type: cross Abstract: We study the problem of auditing a black-box algorithmic decision-maker from observable inputs and outputs alone. Our main result is an exact decompos

Evaluation of ML Resource Utilization Requires Model Life Cycle Assessment

SafetyDGX agent

arXiv:2606.07632v1 Announce Type: new Abstract: Proper accounting of the energy requirements and environmental impact of artificial intelligence (AI) systems is necessary for researchers, developers,

Exposing Hidden Biases in Text-to-Image Models via Automated Prompt Search

SafetyDGX agent

arXiv:2512.08724v3 Announce Type: replace Abstract: Text-to-image (TTI) diffusion models have achieved remarkable visual quality, yet they have been repeatedly shown to exhibit social biases across se

fable’s safety guardrails broken within an hour. raise your hand if you are surprised.

SafetyDGX agent

fable’s safety guardrails broken within an hour. raise your hand if you are surprised. We tested Anthropic’s new @claudeai Fable 5. It did not fail like an ordinary jailbreak. It failed more quietly.

FADRW: A Feature-Aware Modulated and Dynamically Reweighted Loss for Few-Shot Linguistic Steganalysis

SafetyDGX agent

arXiv:2606.07655v1 Announce Type: cross Abstract: The ubiquity of social media platforms facilitates malicious linguistic steganography, posing significant security risks. However, detection is severe

FADTI: Fourier and Attention Driven Diffusion for Multivariate Time Series Imputation

SafetyDGX agent

arXiv:2512.15116v2 Announce Type: replace-cross Abstract: Multivariate time series imputation is fundamental in applications such as healthcare, traffic forecasting, and biological modeling, where sen

FAME: Forecastability-Aware Mixture of Experts for Heterogeneous Time Series Forecasting

SafetyDGX agent

arXiv:2606.08896v1 Announce Type: new Abstract: Large-scale retail and industrial forecasting systems contain many heterogeneous time series whose lifecycle, sparsity, volatility, seasonality, spectra

Fast LLM-Based Semantic Filtering: From a Unified Framework to an Adaptive Two-Phase Method

SafetyDGX agent

arXiv:2606.08090v1 Announce Type: cross Abstract: Evaluating a natural-language yes/no predicate over a document corpus under an accuracy target - the semantic filter - is a cornerstone of LLM-based d

Few-step Cofolding with All-Atom Flow Maps

SafetyDGX agent

arXiv:2606.08375v1 Announce Type: new Abstract: All-atom generative modeling of 3D biomolecular complexes has emerged as the dominant paradigm for predicting the structure of proteins and protein-liga

FlowLet: Conditional 3D Brain MRI Synthesis using Wavelet Flow Matching

SafetyDGX agent

arXiv:2601.05212v2 Announce Type: replace Abstract: Brain Magnetic Resonance Imaging (MRI) plays a central role in studying neurological development, aging, and diseases. One key application is Brain

Frankenstein in the Pipeline: Computational Epistemicide in Facial Recognition

SafetyDGX agent

arXiv:2606.07628v1 Announce Type: cross Abstract: While the eugenic roots of computer vision are well-documented in critical technology studies, less attention has been paid to the operational mechani

From A to B to A: Palindromic Zero-Shot Voice Conversion with Non-Parallel Data

SafetyDGX agent

arXiv:2606.08843v1 Announce Type: cross Abstract: We present a voice conversion (VC) framework that utilizes K-Nearest Neighbors (KNN) retrieval over WavLM representations to align non-parallel source

@GaryMarcus @Pontifex Specifically that the Pope is a powerful, influential individual serving as head of one of humanity's most enduring in…

SafetyDGX agent

@GaryMarcus @Pontifex Specifically that the Pope is a powerful, influential individual serving as head of one of humanity's most enduring institutions who is asserting clearly & unequivocally the prim

Generalized Rank-based Evaluation for Knowledge Graph Completion: Perspectives, Framework, and Analyses

SafetyDGX agent

arXiv:2606.08921v1 Announce Type: new Abstract: Knowledge graph completion (KGC) aims to predict missing facts from an observed knowledge graph (KG), playing a crucial role in a wide range of real-wor

Generalizing Fair Top-k Selection: An Integrative Approach

SafetyDGX agent

arXiv:2603.04689v3 Announce Type: replace-cross Abstract: Fair top-k selection, which ensures appropriate proportional representation of members from minority or historically disadvantaged groups amon

Generative Reasoning Re-ranker

SafetyDGX agent

arXiv:2602.07774v5 Announce Type: replace-cross Abstract: Recent studies increasingly explore Large Language Models (LLMs) as a new paradigm for recommendation systems due to their scalability and wor

GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model

SafetyDGX agent

arXiv:2512.20978v2 Announce Type: replace-cross Abstract: Language Model (LM)-based generative modeling has emerged as a promising direction for TSE, offering potential for improved generalization and

Guided Discovery of New Behaviors using Diffusion Policies

SafetyDGX agent

arXiv:2606.08743v1 Announce Type: new Abstract: Diffusion models have become a powerful tool for generative modeling in robotics, with diffusion policies excelling at modeling multimodal action-trajec

GVC-Seg: Training-Free 3D Instance Segmentation via Geometric Visual Correspondence

SafetyDGX agent

arXiv:2606.08014v1 Announce Type: cross Abstract: Accurate 3D instance segmentation in point cloud data is critical for machine vision applications. Recent advancements leverage multiple pre-trained f

← Previous
1…7980818283…214
Next →