AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,708 results
12 May 2026

Sink vs. diagonal patterns as mechanisms for attention switch and oversmoothing prevention

SafetyDGX agent

arXiv:2605.08453v1 Announce Type: cross Abstract: This paper studies the role of sinks and diagonal patterns as attention switch and anti-oversmoothing mechanisms. We analyze geometric conditions unde

SKG-VLA: Scene Knowledge Graph Priors for Structured Scene Semantics and Multimodal Reasoning for Decision Making

SafetyDGX agent

arXiv:2605.09343v1 Announce Type: new Abstract: Decision making in large-scale complaint handling systems increasingly relies on heterogeneous evidence, including complaint narratives, screenshots, or

Skill-R1: Agent Skill Evolution via Reinforcement Learning

SafetyDGX agent

arXiv:2605.09359v1 Announce Type: cross Abstract: Agentic large language models often rely on skills, reusable natural language procedures that guide planning, action, and tool use. In practice, skill


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SLASH the Sink: Sharpening Structural Attention Inside LLMs

SafetyDGX agent

arXiv:2605.10503v1 Announce Type: new Abstract: Large Language Models (LLMs) show remarkable semantic understanding but often struggle with structural understanding when processing graph topologies in

SnareNet: Flexible Repair Layers for Neural Networks with Hard Constraints

SafetyDGX agent

arXiv:2602.09317v2 Announce Type: replace-cross Abstract: Neural networks are increasingly used as fast surrogate models across various domains, but unconstrained predictions can violate physical, ope

So basically Anthropic’s estimated valuation went up half a trillion dollars in a couple weeks (then back down a bit) on hype. Small wonder …

SafetyDGX agent

So basically Anthropic’s estimated valuation went up half a trillion dollars in a couple weeks (then back down a bit) on hype. Small wonder that companies like OpenAI, Anthropic and Tesla keep hyping

Sparsity Hurts: Simple Linear Adapter Can Boost Generalized Category Discovery

SafetyDGX agent

arXiv:2605.08183v1 Announce Type: new Abstract: Generalized Category Discovery (GCD) seeks to identify novel categories from unlabeled data while retaining the classification ability of seen categorie

speaks for itself

SafetyDGX agent

speaks for itself Musk's lawyer: 'Are you completely trustworthy?' Altman: 'I believe so.' Musk's lawyer: 'But, you know, you don't know whether you're completely trustworthy.' Altman: 'I'll just amen

Spectral Transformer Neural Processes

SafetyDGX agent

arXiv:2605.09498v1 Announce Type: cross Abstract: Time series, spatial data, and images are natural applications of Neural Processes. However, when such data exhibit strong periodicity and quasi-perio

Stable Long-Horizon PDE Forecasting via Latent Structured Spectral Propagators

SafetyDGX agent

arXiv:2605.10154v1 Announce Type: new Abstract: Long-horizon forecasting of time-dependent partial differential equations (PDEs) is critical for characterizing the sustained evolution of physical syst

StereoPolicy: Improving Robotic Manipulation Policies via Stereo Perception

SafetyDGX agent

arXiv:2605.09989v1 Announce Type: cross Abstract: Recent advances in robot imitation learning have yielded powerful visuomotor policies capable of manipulating a wide variety of objects directly from

STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering

SafetyDGX agent

arXiv:2604.01824v2 Announce Type: replace Abstract: We introduce STRIVE (SpatioTemporal Reinforcement with Importance-aware Variant Exploration), a structured reinforcement learning framework for vide

Structural Alignment Improves Graph Test-Time Adaptation

SafetyDGX agent

arXiv:2502.18334v5 Announce Type: replace Abstract: Graph-based learning excels at capturing interaction patterns in diverse domains like recommendation, fraud detection, and particle physics. However

Sub-JEPA: Subspace Gaussian Regularization for Stable End-to-End World Models

SafetyDGX agent

arXiv:2605.09241v1 Announce Type: cross Abstract: Joint-Embedding Predictive Architectures (JEPAs) provide a simpleframework for learning world models by predicting future latent representations.Howev

Summary: An International Agreement to Prevent the Premature Creation of Artificial Superintelligence

SafetyDGX agent

If anyone, anywhere builds a superhuman artificial intelligence using present methods, the most likely outcome is catastrophe. There have accordingly been widespread calls for an international agreeme

Supercharging Bayesian Inference with Reliable AI-Informed Priors

SafetyDGX agent

arXiv:2605.09834v1 Announce Type: cross Abstract: Modern predictive systems encode beliefs that can act as useful prior information for statistical inference in data-limited settings. Using them for p

Supervised Mixture-of-Experts for Surgical Grasping and Retraction

SafetyDGX agent

arXiv:2601.21971v2 Announce Type: replace-cross Abstract: Imitation learning has achieved remarkable success in robotic manipulation, yet its application to surgical robotics remains challenging due t

Survey-aware Machine Learning: A Guideline for Valid Population Health Inference based on Scoping Review

SafetyDGX agent

arXiv:2605.08963v1 Announce Type: cross Abstract: Machine Learning (ML) models trained on complex health surveys such as the National Health and Nutrition Examination Survey (NHANES) often ignore prim

Switching-Geometry Analysis of Deflated Q-Value Iteration

SafetyDGX agent

arXiv:2605.10811v1 Announce Type: cross Abstract: This paper develops a joint spectral radius (JSR) framework for analyzing rank-one deflated Q-value iteration (Q-VI) in discounted Markov decision pro

Synergistic Simplex: Cooperative Runtime Assurance for Safety-Critical Autonomous Systems

SafetyDGX agent

arXiv:2605.08190v1 Announce Type: new Abstract: Autonomous systems increasingly rely on machine-learning (ML) components for safety-critical tasks such as perception and control in autonomous vehicles

SynerMedGen: Synergizing Medical Multimodal Understanding with Generation via Task Alignment

SafetyDGX agent

arXiv:2605.08724v1 Announce Type: new Abstract: Unifying multimodal understanding and generation is a compelling frontier that is beginning to emerge in the medical field. However, the limited existin

Targeted Synthetic Control Method

SafetyDGX agent

arXiv:2602.04611v2 Announce Type: replace-cross Abstract: The synthetic control method (SCM) estimates causal effects in panel data with a single-treated unit by constructing a counterfactual outcome

TD3B: Transition-Directed Discrete Diffusion for Allosteric Binder Generation

SafetyDGX agent

arXiv:2605.09810v1 Announce Type: cross Abstract: Protein function is often controlled by ligands that bias the direction of state transitions, such as agonists and antagonists, rather than stabilizin

Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMs

SafetyDGX agent

arXiv:2605.09922v1 Announce Type: cross Abstract: While recent self-training approaches have reduced reliance on human-labeled data for aligning LLMs, they still face critical limitations: (i) sensiti

The Accountability Paradox: How Platform API Restrictions Undermine AI Transparency Mandates

SafetyDGX agent

arXiv:2505.11577v4 Announce Type: replace-cross Abstract: Recent application programming interface (API) restrictions on major social media platforms challenge compliance with the EU Digital Services

The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring

SafetyDGX agent

arXiv:2605.09225v1 Announce Type: cross Abstract: Jailbreak attacks -- adversarial prompts that bypass LLM alignment through purely linguistic manipulation -- pose a growing operational security threa

The Bystander Effect in Multi-Agent Reasoning: Quantifying Cognitive Loafing in Collaborative Interactions

SafetyDGX agent

arXiv:2605.10698v1 Announce Type: cross Abstract: Multi-agent systems (MAS) assume that collaborating inherently improves Large Language Model (LLM) reasoning. We challenge this by demonstrating that

The FT says that Amazon employees are doing random unnecessary task automations to consume tokens and to show their bosses that they're usin…

SafetyDGX agent

The FT says that Amazon employees are doing random unnecessary task automations to consume tokens and to show their bosses that they're using AI more https://www.ft.com/content/8ee0d3ef-9548-422d-8ff1

The Geometric Structure of Models Learning Sparse Data

SafetyDGX agent

arXiv:2605.08464v1 Announce Type: new Abstract: The manifold hypothesis (MH) is often used to explain how machine learning can overcome the curse of dimensionality. However, the MH is only applicable

The Grounding Gap: How LLMs Anchor the Meaning of Abstract Concepts Differently from Humans

SafetyDGX agent

arXiv:2605.08837v1 Announce Type: cross Abstract: Abstract concepts - justice, theory, availability - have no single perceivable referent; in the human brain, their meaning emerges from a web of exper

The Pokemon Theorem and other Fairness Impossibility Results

SafetyDGX agent

arXiv:2605.09221v1 Announce Type: cross Abstract: Fairness impossibility results often look like distinct scalar incompatibility statements. We show that several share one RKHS geometry: fairness crit

The Procrustean Bed of Time Series: The Optimization Bias in Point-wise Loss Functions

SafetyDGX agent

arXiv:2512.18610v3 Announce Type: replace Abstract: Intuitively, a more deterministic time series should be easier to forecast. However, point-wise loss functions (e.g., MSE and MAE), serving as diffe

The Safety-Aware Denoiser for Text Diffusion Models

SafetyDGX agent

arXiv:2605.08116v1 Announce Type: cross Abstract: Recent work on text diffusion models offers a promising alternative to autoregressive generation, but controlling their safety remains underexplored.

The Value of Mechanistic Priors in Sequential Decision Making

SafetyDGX agent

arXiv:2605.10018v1 Announce Type: new Abstract: Hybrid mechanistic models, physical priors with learned residuals, promise to reduce the data required for good decisions, but have no computable criter

The Wittgensteinian Representation Hypothesis: Is Language the Attractor of Multimodal Convergence?

SafetyDGX agent

arXiv:2605.09352v1 Announce Type: new Abstract: Understanding why independently trained neural networks from different modalities converge toward shared representations, and where this convergence lea

There will be no AI jobpocalypse. The story that AI will lead to massive unemployment is stoking unnecessary fear. AI — like any other techn…

SafetyDGX agent

There will be no AI jobpocalypse. The story that AI will lead to massive unemployment is stoking unnecessary fear. AI — like any other technology — does affect jobs, but telling overblown stories of l

Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection

SafetyDGX agent

arXiv:2605.10130v1 Announce Type: new Abstract: Existing open-vocabulary detectors focus on RGB images and fail to generalize to thermal imagery, where low texture and emissivity variations challenge

TIE: Time Interval Encoding for Video Generation over Events

SafetyDGX agent

arXiv:2605.10543v1 Announce Type: new Abstract: Director-style prompting, robotic action prediction, and interactive video agents demand temporal grounding over concurrent events -- a regime in which

tl;dr - let's not just remove negative behavior from models, but also add positive ones 👍

SafetyDGX agent

tl;dr - let's not just remove negative behavior from models, but also add positive ones 👍 If anyone builds it, everyone thrives. Over the past decade, a lot of important work on AI alignment has focus

TodyComm: Task-Oriented Dynamic Communication for Multi-Round LLM-based Multi-Agent System

SafetyDGX agent

arXiv:2602.03688v2 Announce Type: replace Abstract: Multi-round LLM-based multi-agent systems rely on effective communication structures to support collaboration across rounds. However, most existing

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning

SafetyDGX agent

arXiv:2508.20697v3 Announce Type: replace-cross Abstract: As large language models (LLMs) continue to grow in capability, so do the risks of harmful misuse through fine-tuning. While most prior studie

Total Generalized Variation regularization closes the gap between neural-eld and classical methods in seismic travel-time tomography

SafetyDGX agent

arXiv:2605.09960v1 Announce Type: cross Abstract: Travel-time tomography forces a trade-off between mesh resolution and stability in which the regularizer choice dominates what can be recovered. We in

Toward Reliable Sim-to-Real Predictability for MoE-based Robust Quadrupedal Locomotion

SafetyDGX agent

arXiv:2602.00678v4 Announce Type: replace Abstract: Reinforcement learning has shown strong promise for quadrupedal agile locomotion, even with proprioception-only sensing. In practice, however, sim-t

Towards Customized Multimodal Role-Play

SafetyDGX agent

arXiv:2605.08129v1 Announce Type: new Abstract: Unified multimodal understanding and generation models enable richer human-AI interaction. Yet jointly customizing a character's persona, dialogue style

TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment

SafetyDGX agent

arXiv:2605.10194v1 Announce Type: new Abstract: On-policy self-distillation (self-OPD) densifies reinforcement learning with verifiable rewards (RLVR) by letting a policy teach itself under privileged

Training-Free Cultural Alignment of Large Language Models via Persona Disagreement

SafetyDGX agent

arXiv:2605.10843v1 Announce Type: cross Abstract: Large language models increasingly mediate decisions that turn on moral judgement, yet a growing body of evidence shows that their implicit preference

Training with Harnesses: On-Policy Harness Self-Distillation for Complex Reasoning

SafetyDGX agent

arXiv:2605.08741v1 Announce Type: new Abstract: Inference-time harnesses substantially improve large language models on complex reasoning tasks. However, the intrinsic capabilities of the underlying m

Trajectory-Consistent Flow Matching for Robust Visuomotor Policy Learning

SafetyDGX agent

arXiv:2605.08511v1 Announce Type: new Abstract: Flow matching policies learn continuous velocity fields that transport noise to actions, enabling fast deterministic inference for robot manipulation. H

TrajTok: Learning Trajectory Tokens enables better Video Understanding

SafetyDGX agent

arXiv:2602.22779v2 Announce Type: replace Abstract: Tokenization in video models, typically through patchification, generates an excessive and redundant number of tokens. This severely limits video ef

TripleWin: Fixed-Point Equilibrium Pricing for Data-Model Coupled Markets

SafetyDGX agent

arXiv:2511.03368v2 Announce Type: replace Abstract: The rise of the machine learning (ML) model economy has intertwined markets for training datasets and pre-trained models. However, most pricing appr

Trustworthy AI: Ensuring Reliability and Accountability from Models to Agents

SafetyDGX agent

arXiv:2605.08964v1 Announce Type: new Abstract: In this thesis, we develop algorithms with theoretical guarantees for ensuring reliability and accountability of Machine Learning (ML) systems. As ML sy

UAV-Assisted Scan-to-Simulation for Landslides Using Physics-Informed Gaussian Splatting

SafetyDGX agent

arXiv:2605.10715v1 Announce Type: new Abstract: Landslide monitoring and simulation play an important role in urban safety assessment and disaster prevention. Existing landslide simulation pipelines t

Uncertainty-Aware and Decoder-Aligned Learning for Video Summarization

SafetyDGX agent

arXiv:2605.09507v1 Announce Type: new Abstract: Video summarization aims to produce a compact representation of a long video by selecting a subset of temporally important segments that best reflect hu

Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views

SafetyDGX agent

arXiv:2511.12878v4 Announce Type: replace Abstract: Forecasting how human hands move in egocentric views is critical for applications like augmented reality and human-robot policy transfer. Recently,

Unified Noise Steering for Efficient Human-Guided VLA Adaptation

SafetyDGX agent

arXiv:2605.10821v1 Announce Type: new Abstract: Diffusion-based vision-language-action (VLA) models have emerged as strong priors for robotic manipulation, yet adapting them to real-world distribution

Uniform Inductive Spatio-Temporal Kriging

SafetyDGX agent

arXiv:2603.05301v2 Announce Type: replace Abstract: Inductive spatio-temporal kriging infers signals at unobserved locations from observed sensors, but real-world observations are often incomplete and

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation

SafetyDGX agent

arXiv:2605.08729v1 Announce Type: new Abstract: Motion, speech, and sound effects are fundamental elements of human-centric videos, yet their heterogeneous temporal characteristics make joint generati

Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning

SafetyDGX agent

arXiv:2605.08765v1 Announce Type: cross Abstract: Unlearning in large language models (LLMs) aims to remove harmful training data while preserving overall utility. However, we find that existing metho

Unsupervised Process Reward Models

SafetyDGX agent

arXiv:2605.10158v1 Announce Type: new Abstract: Process Reward Models (PRMs) are a powerful mechanism for steering large language model reasoning by providing fine-grained, step-level supervision. How

Upholding Epistemic Agency: A Brouwerian Assertibility Constraint for Responsible AI

SafetyDGX agent

arXiv:2603.03971v2 Announce Type: replace-cross Abstract: Generative AI can convert uncertainty into hypersuasive, authoritative-seeming verdicts, displacing the justificatory work on which democratic

← Previous
1…152153154155156…212
Next →