AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,702 results
21 Apr 2026

Spectral bandits for smooth graph functions

SafetyDGX agent

arXiv:2604.18420v1 Announce Type: cross Abstract: Smooth functions on graphs have wide applications in manifold and semi-supervised learning. In this paper, we study a bandit problem where the payoffs

Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding

SafetyDGX agent

arXiv:2509.24328v2 Announce Type: replace Abstract: LLMs have low GPU efficiency and high latency due to autoregressive decoding. Speculative decoding (SD) mitigates this using a small draft model to

SpiralThinker: Latent Reasoning through an Iterative Process with Text-Latent Interleaving

SafetyDGX agent

arXiv:2511.08983v2 Announce Type: replace Abstract: Recent advances in large reasoning models have been driven by reinforcement learning and test-time scaling, accompanied by growing interest in laten


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models

SafetyDGX agent

arXiv:2604.16995v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a promising paradigm for training reasoning-oriented models by leveraging rule-based reward signals. However,

StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt Generation

SafetyDGX agent

arXiv:2601.04740v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly applied in specialized domains such as finance and healthcare, where they introduce unique safety risk

STL-Based Motion Planning and Uncertainty-Aware Risk Analysis for Human-Robot Collaboration with a Multi-Rotor Aerial Vehicle

SafetyDGX agent

arXiv:2509.10692v3 Announce Type: replace Abstract: This paper presents a motion planning and risk analysis framework for enhancing human-robot collaboration with a Multi-Rotor Aerial Vehicle. The pro

Structure-Aware Diversity Pursuit as an AI Safety Strategy against Homogenization

SafetyDGX agent

arXiv:2601.06116v2 Announce Type: replace-cross Abstract: Generative AI models reproduce the biases in the training data and can further amplify them through mode collapse. We refer to the resulting h

Sub-metre Lunar DEM Generation and Validation from Chandrayaan-2 OHRC Multi-View Imagery Using an Open-Source Pipeline

SafetyDGX agent

arXiv:2604.01032v3 Announce Type: replace Abstract: High-resolution digital elevation models (DEMs) of the lunar surface are essential for surface mobility planning, landing site characterization, and

Support Sufficiency as Consequence-Sensitive Compression in Belief Arbitration

SafetyDGX agent

arXiv:2604.16434v1 Announce Type: cross Abstract: When a system commits to a hypothesis, much of the evidential structure behind that commitment is lost to compression. Standard accounts assume that s

// Survey on Multi-Agent Systems // The paper traces the landscape from classical paradigms (consensus, distributed control, swarm intellige…

SafetyDGX agent

// Survey on Multi-Agent Systems // The paper traces the landscape from classical paradigms (consensus, distributed control, swarm intelligence, cooperative learning) to foundation-model-enabled MAS (

SynAgent: Generalizable Cooperative Humanoid Manipulation via Solo-to-Cooperative Agent Synergy

SafetyDGX agent

arXiv:2604.18557v1 Announce Type: new Abstract: Controllable cooperative humanoid manipulation is a fundamental yet challenging problem for embodied intelligence, due to severe data scarcity, complexi

SynopticBench: Evaluating Vision-Language Models on Generating Weather Forecast Discussions of the Future

SafetyDGX agent

arXiv:2604.16451v1 Announce Type: new Abstract: Recent advances in visual-language models (VLMs) have led to significant improvements in a plethora of complex multimodal tasks like image captioning, r

Synthia: Scalable Grounded Persona Generation from Social Media Data

SafetyDGX agent

arXiv:2507.14922v2 Announce Type: replace Abstract: Persona-driven simulations are increasingly used in computational social science, yet their validity critically depends on the fidelity of the under

Task Matters: Knowledge Requirements Shape LLM Responses to Context-Memory Conflict

SafetyDGX agent

arXiv:2506.06485v4 Announce Type: replace Abstract: Large language models (LLMs) draw on both contextual information and parametric memory, yet these sources can conflict. Prior studies have largely e

TeMuDance: Contrastive Alignment-Based Textual Control for Music-Driven Dance Generation

SafetyDGX agent

arXiv:2604.17005v1 Announce Type: new Abstract: Existing music-driven dance generation approaches have achieved strong realism and effective audio-motion alignment. However, they generally lack semant

Test-Time Alignment via Hypothesis Reweighting

SafetyDGX agent

arXiv:2412.08812v2 Announce Type: replace Abstract: Reward models trained on aggregate preferences often fail to capture individual users' values, but existing adaptation methods such as fine-tuning o

The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillation

SafetyDGX agent

arXiv:2604.16830v1 Announce Type: new Abstract: On-policy distillation (OPD) is an increasingly important paradigm for post-training language models. However, we identify a pervasive Scaling Law of Mi

The Impact of Off-Policy Training Data on Probe Generalisation

SafetyDGX agent

arXiv:2511.17408v4 Announce Type: replace-cross Abstract: Probing has emerged as a promising method for monitoring large language models (LLMs), enabling cheap inference-time detection of concerning b

The Less You Depend, The More You Learn: Synthesizing Novel Views from Sparse, Unposed Images with Minimal 3D Knowledge

SafetyDGX agent

arXiv:2506.09885v2 Announce Type: replace Abstract: Recent advances in feed-forward Novel View Synthesis (NVS) have led to a divergence between two design philosophies: bias-driven methods, which rely

The Provenance Gap in Clinical AI: Evidence-Traceable Temporal Knowledge Graphs for Rare Disease Reasoning

SafetyDGX agent

arXiv:2604.17114v1 Announce Type: new Abstract: Frontier large language models generate clinically accurate outputs, but their citations are often fabricated. We term this the Provenance Gap. We teste

The Umwelt Representation Hypothesis: Rethinking Universality

SafetyDGX agent

arXiv:2604.17960v1 Announce Type: cross Abstract: Recent studies reveal striking representational alignment between artificial neural networks (ANNs) and biological brains, leading to proposals that a

“There is no option to opt out” - Meta’s CTO “There was no option to opt out” - epigraph for life as we knew it?

SafetyDGX agent

“There is no option to opt out” - Meta’s CTO “There was no option to opt out” - epigraph for life as we knew it? Scoop! Meta staff are outraged over a new program that trains AI off their keystrokes a

They just indicted the Southern Poverty Law Center on 11 counts of fraud. Wow! The SPLC worked with Facebook, Old Twitter and Youtube on the…

SafetyDGX agent

They just indicted the Southern Poverty Law Center on 11 counts of fraud. Wow! The SPLC worked with Facebook, Old Twitter and Youtube on their safety teams, sometimes directly advising/flagging what c

Think before Go: Hierarchical Reasoning for Image-goal Navigation

SafetyDGX agent

arXiv:2604.17407v1 Announce Type: new Abstract: Image-goal navigation steers an agent to a target location specified by an image in unseen environments. Existing methods primarily handle this task by

Time Series Forecasting as Reasoning: A Slow-Thinking Approach with Reinforced LLMs

SafetyDGX agent

arXiv:2506.10630v2 Announce Type: replace Abstract: To advance time series forecasting (TSF), various methods have been proposed to improve prediction accuracy, evolving from statistical techniques to

Toward Consistent World Models with Multi-Token Prediction and Latent Semantic Enhancement

SafetyDGX agent

arXiv:2604.06155v2 Announce Type: replace-cross Abstract: Whether Large Language Models (LLMs) develop coherent internal world models remains a core debate. While conventional Next-Token Prediction (N

Toward Reusability of AI Models Using Dynamic Updates of AI Documentation

SafetyDGX agent

arXiv:2604.17626v1 Announce Type: cross Abstract: This work addresses the challenge of disseminating reusable artificial intelligence (AI) models accompanied by AI documentation (a.k.a., AI model card

Towards Robust Text-to-Image Person Retrieval: Multi-View Reformulation for Semantic Compensation

SafetyDGX agent

arXiv:2604.18376v1 Announce Type: new Abstract: In text-to-image person retrieval tasks, the diversity of natural language expressions and the implicitness of visual semantics often lead to the proble

Towards Trustworthy Depression Estimation via Disentangled Evidential Learning

SafetyDGX agent

arXiv:2604.16579v1 Announce Type: new Abstract: Automated depression estimation is highly vulnerable to signal corruption and ambient noise in real-world deployment. Prevailing deterministic methods p

Towards Universal Skeleton-Based Action Recognition

SafetyDGX agent

arXiv:2604.17013v1 Announce Type: new Abstract: With the development of robotics, skeleton-based action recognition has become increasingly important, as human-robot interaction requires understanding

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts

SafetyDGX agent

arXiv:2604.18473v1 Announce Type: new Abstract: Extending a fully post-trained language model with new domain capabilities is fundamentally limited by monolithic training paradigms: retraining from sc

Training-inference input alignment outweighs framework choice in longitudinal retinal image prediction

SafetyDGX agent

arXiv:2604.16955v1 Announce Type: new Abstract: Quantitative prediction of future retinal appearance from longitudinal imaging would support clinical decisions in progressive macular disease that curr

Training Language Models to Use Prolog as a Tool

SafetyDGX agent

arXiv:2512.07407v2 Announce Type: replace Abstract: Language models frequently produce plausible yet incorrect reasoning traces that are difficult to verify. We investigate fine-tuning models to use P

Transition-Matrix Regularization for Next Dialogue Act Prediction in Counselling Conversations

SafetyDGX agent

arXiv:2604.18539v1 Announce Type: new Abstract: This paper studies how empirical dialogue-flow statistics can be incorporated into Next Dialogue Act Prediction (NDAP). A KL regularization term is prop

Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment

SafetyDGX agent

arXiv:2405.13068v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have revolutionized various applications, making robust safety alignment essential to prevent harmful outputs. Cu

UniComp: A Unified Evaluation of Large Language Model Compression via Pruning, Quantization and Distillation

SafetyDGX agent

arXiv:2602.09130v3 Announce Type: replace Abstract: Model compression is increasingly essential for deploying large language models (LLMs), yet existing comparative studies largely focus on pruning an

UniCon: Unified Framework for Efficient Contrastive Alignment via Kernels

SafetyDGX agent

arXiv:2604.16678v1 Announce Type: new Abstract: Contrastive objectives power state-of-the-art multimodal models, but their training remains slow, relying on long stochastic optimization. We propose a

UniCSG: Unified High-Fidelity Content-Constrained Style-Driven Generation via Staged Semantic and Frequency Disentanglement

SafetyDGX agent

arXiv:2604.17850v1 Announce Type: new Abstract: Style transfer must match a target style while preserving content semantics. DiT-based diffusion models often suffer from content-style entanglement, le

Untrained CNNs Match Backpropagation at V1: A Systematic RSA Comparison of Four Learning Rules Against Human fMRI

SafetyDGX agent

arXiv:2604.16875v1 Announce Type: new Abstract: A central question in computational neuroscience is whether the learning rule used to train a neural network determines how well its internal representa

VA law REQUIRES that ballot language be 'a neutral explanation.' This language is CLEARLY illegal. A court ruled as much...and Democrats ign…

SafetyDGX agent

VA law REQUIRES that ballot language be 'a neutral explanation.' This language is CLEARLY illegal. A court ruled as much...and Democrats ignored the court and waited for a different court to punt on t

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech

SafetyDGX agent

arXiv:2604.17248v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases remain underexplored. Existing sp

VIDEOP2R: Video Understanding from Perception to Reasoning

SafetyDGX agent

arXiv:2511.11113v2 Announce Type: replace Abstract: Reinforcement fine-tuning (RFT), a two-stage framework consisting of supervised fine-tuning (SFT) and reinforcement learning (RL) has shown promisin

Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception

SafetyDGX agent

arXiv:2604.17475v1 Announce Type: cross Abstract: Small Vision-Language Models (SVLMs) are efficient task controllers but often suffer from visual brittleness and poor tool orchestration. They typical

We think ControlAI can turn $50M / year into a 10% chance of banning ASI. Most of the AI safety community has been far too coy about extinct…

SafetyDGX agent

We think ControlAI can turn $50M / year into a 10% chance of banning ASI. Most of the AI safety community has been far too coy about extinction risk. We're not. It's not that complicated: AI smarter t

Weakly-Supervised Referring Video Object Segmentation through Text Supervision

SafetyDGX agent

arXiv:2604.17797v1 Announce Type: new Abstract: Referring video object segmentation (RVOS) aims to segment the target instance in a video, referred by a text expression. Conventional approaches are mo

When Choices Become Risks: Safety Failures of Large Language Models under Multiple-Choice Constraints

SafetyDGX agent

arXiv:2604.16916v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is primarily evaluated under open-ended generation, where models can mitigate risk by refusing to respo

Where Do Self-Supervised Speech Models Become Unfair?

SafetyDGX agent

arXiv:2604.18249v1 Announce Type: new Abstract: Speech encoder models are known to model members of some speaker groups (SGs) better than others. However, there has been little work in establishing wh

Who Gets the Kidney? Human-AI Alignment, Indecision, and Moral Values

SafetyDGX agent

arXiv:2506.00079v2 Announce Type: replace-cross Abstract: The rapid integration of Large Language Models (LLMs) in high-stakes decision-making -- such as allocating scarce resources like donor organs

Why Agents Compromise Safety Under Pressure

SafetyDGX agent

arXiv:2603.14975v2 Announce Type: replace-cross Abstract: Large Language Model agents deployed in complex environments frequently encounter a conflict between maximizing goal achievement and adhering

Why AI Readiness Is an Organizational Learning Problem, Not a Technology Purchase

SafetyDGX agent

arXiv:2604.16369v1 Announce Type: cross Abstract: Global corporate AI investment reached $252.3 billion in 2024, yet only 6% of firms report significant earnings impact. This article argues that AI pr

Why do no 'AI safety' people ever say how LLM AIs failed to live up to the hype? Why do the safety people never share the failures of LLMs? …

SafetyDGX agent

Why do no 'AI safety' people ever say how LLM AIs failed to live up to the hype? Why do the safety people never share the failures of LLMs? This link is from Nov 2023 & shows how tech CEOs used 'safet

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention

SafetyDGX agent

arXiv:2510.04212v3 Announce Type: replace Abstract: The pursuit of computational efficiency has driven the adoption of low-precision formats for training transformer models. However, this progress is

ZoFia: Zero-Shot Fake News Detection with Entity-Guided Retrieval and Multi-LLM Interaction

SafetyDGX agent

arXiv:2511.01188v2 Announce Type: replace Abstract: The rapid spread of fake news threatens social stability and public trust, highlighting the urgent need for its effective detection. Although large

20 Apr 2026

A Systematic Study of Training-Free Methods for Trustworthy Large Language Models

SafetyDGX agent

arXiv:2604.15789v1 Announce Type: new Abstract: As Large Language Models (LLMs) receive increasing attention and are being deployed across various domains, their potential risks, including generating

A Wasserstein Geometric Framework for Hebbian Plasticity

SafetyDGX agent

arXiv:2604.16052v1 Announce Type: cross Abstract: We introduce the Tan-HWG framework (Hebbian-Wasserstein-Geometry), a geometric theory of Hebbian plasticity in which memory states are modeled as prob

Adaptive multi-fidelity optimization with fast learning rates

SafetyDGX agent

arXiv:2604.16239v1 Announce Type: cross Abstract: In multi-fidelity optimization, biased approximations of varying costs of the target function are available. This paper studies the problem of optimiz

Aerial Multi-Functional RIS in Fluid Antennas-Aided Full-Duplex Networks: A Self-Optimized Hybrid Deep Reinforcement Learning Approach

SafetyDGX agent

arXiv:2604.14309v2 Announce Type: replace-cross Abstract: To address high data traffic demands of sixth-generation (6G) networks, this paper proposes a novel architecture that integrates autonomous ae

💯 agree

SafetyDGX agent

💯 agree The AI bubble will burst says economist Ann Pettifor She is Chair of the Political Economy Research Centre's Advisory Board at Goldsmiths, a fellow of the New Economics Foundation, a director

AI Agents and Hard Choices

SafetyDGX agent

arXiv:2504.15304v2 Announce Type: replace Abstract: Can AI agents deal with hard choices -- cases where options are incommensurable because multiple objectives are pursued simultaneously? Adopting a t

“AI is at best a functional mimic, not a conscious experiencing subject. …. The real moral issue lies not in making AI conscious …. but in a…

SafetyDGX agent

“AI is at best a functional mimic, not a conscious experiencing subject. …. The real moral issue lies not in making AI conscious …. but in avoiding transforming humans into zombies” @GaryMarcus @OEIAC

← Previous
1…190191192193194…212
Next →