AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,704 results
29 Apr 2026

Sustained Gradient Alignment Mediates Subliminal Learning in a Multi-Step Setting: Evidence from MNIST Auxiliary Logit Distillation Experiment

SafetyDGX agent

arXiv:2604.25779v1 Announce Type: new Abstract: In the MNIST auxiliary logit distillation experiment, a student can acquire an unintended teacher trait despite distilling only on no-class logits throu

the ability of this man to sell things he doesn’t believe in is truly extraordinary

SafetyDGX agent

the ability of this man to sell things he doesn’t believe in is truly extraordinary SAM ALTMAN “OpenAI is structured as a nonprofit because we don’t ever want to be making decisions to benefit shareho

'The more you look at chatbots, the more you realize that they were rushed to market with very little consideration for the consequences,' h…

Safety

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

Gary Marcus critiques the rapid deployment of chatbots to market without adequate consideration of potential harms and consequences. The statement reflects concerns about insufficient safety testing a

The Musk-Altman trial is clearly a battle of egos—and it’s hard to root for either side—𝗯𝘂𝘁 𝗶𝘁’𝘀 𝗮𝗹𝘀𝗼 𝗮 𝘁𝗿𝗶𝗮𝗹 𝗮𝗯𝗼𝘂𝘁 𝘄…

SafetyDGX agent

The Musk-Altman trial is clearly a battle of egos—and it’s hard to root for either side—𝗯𝘂𝘁 𝗶𝘁’𝘀 𝗮𝗹𝘀𝗼 𝗮 𝘁𝗿𝗶𝗮𝗹 𝗮𝗯𝗼𝘂𝘁 𝘄𝗵𝗲𝘁𝗵𝗲𝗿 𝗢𝗽𝗲𝗻𝗔𝗜 𝘀𝗵𝗼𝘂𝗹𝗱 𝗯𝗲 𝗵𝗲𝗹𝗱 𝘁𝗼 𝗶𝘁𝘀 𝗽𝗿𝗼𝗺𝗶𝘀𝗲𝘀 𝘁𝗼 𝗯𝗲 𝗮 𝗻𝗼𝗻𝗽𝗿𝗼𝗳𝗶𝘁 𝘄𝗼𝗿𝗸𝗶𝗻𝗴 𝗳𝗼𝗿 𝘁𝗵𝗲 𝗯𝗲𝗻𝗲

The President of the United States trying to take $10 billion dollars from US taxpayers for his own personal gain—and he is probably going t…

SafetyDGX agent

This post appears to discuss allegations that a U.S. President sought to obtain $10 billion in taxpayer funds for personal use. The content is incomplete in the provided title, making it difficult to

The Rashomon Effect for Visualizing High-Dimensional Data

SafetyDGX agent

arXiv:2604.00485v2 Announce Type: replace Abstract: Dimension reduction (DR) is inherently non-unique: multiple embeddings can preserve the structure of high-dimensional data equally well while differ

The Russian Legislative Corpus

SafetyDGX agent

arXiv:2406.04855v3 Announce Type: replace Abstract: We present a comprehensive corpus of Russian primary and secondary legislation adopted between 1991 and 2025, comprising 304,382 texts (194,425,905

Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models

SafetyDGX agent

arXiv:2510.16340v2 Announce Type: replace Abstract: Recent advances in post-training techniques have endowed Large Language Models (LLMs) with enhanced capabilities for tackling complex, logic-intensi

this kind of armchair nonsense is especially hilarious on day when I defended Elon’s lawsuit on CNN and in my newsletter. 🙄

SafetyDGX agent

this kind of armchair nonsense is especially hilarious on day when I defended Elon’s lawsuit on CNN and in my newsletter. 🙄 @GaryMarcus @gnoble79 I'm not sure history is going to be kind to this take.

Three Models of RLHF Annotation: Extension, Evidence, and Authority

SafetyDGX agent

arXiv:2604.25895v1 Announce Type: cross Abstract: Preference-based alignment methods, most prominently Reinforcement Learning with Human Feedback (RLHF), use the judgments of human annotators to shape

To be a bit more clear for people who did not read the full article: A domestic data-center ban does not directly address the risks of extin…

SafetyDGX agent

To be a bit more clear for people who did not read the full article: A domestic data-center ban does not directly address the risks of extinction from ASI, and weakens the US. People may or may not al

TouchAI: Exploring human-AI perceptual alignment in touch through language model representations

SafetyDGX agent

arXiv:2406.06587v2 Announce Type: replace Abstract: Aligning large language models (LLMs) behaviour with human intent is critical for future AI. An important yet often overlooked aspect of this alignm

True or False: OpenAI will eventually become a massively profitable company, more than earning out all the money that went into it.

SafetyDGX agent

Gary Marcus poses a question about whether OpenAI will achieve sufficient profitability to justify the substantial capital investments made into the company. This reflects broader industry speculation

Unrequited Emotions: Investigating the Gaps in Motivation and Practice in Speech Emotion Recognition Research

SafetyDGX agent

arXiv:2604.25776v1 Announce Type: new Abstract: Critical analyses of emotion recognition technology have raised ethical concerns around task validity and potential downstream impacts, urging researche

VISION-SLS: Safe Perception-Based Control from Learned Visual Representations via System Level Synthesis

SafetyDGX agent

arXiv:2604.24894v1 Announce Type: cross Abstract: We propose VISION-SLS, a method for nonlinear output-feedback control from high-resolution RGB images which provides robust constraint satisfaction gu

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution

SafetyDGX agent

arXiv:2604.03472v2 Announce Type: replace Abstract: Co-evolutionary self-play, where one language model generates problems and another solves them, promises autonomous curriculum learning without huma

Voice, Bias, and Coreference: An Interpretability Study of Gender in Speech Translation

SafetyDGX agent

arXiv:2511.21517v2 Announce Type: replace Abstract: Unlike text, speech conveys information about the speaker, such as gender, through acoustic cues like pitch. This gives rise to modality-specific bi

We at @ControlAI are sometimes asked what we think of banning datacenter construction. At CAI, we focus on one issue: The risk of extinction…

SafetyDGX agent

We at @ControlAI are sometimes asked what we think of banning datacenter construction. At CAI, we focus on one issue: The risk of extinction from superintelligent AI. The only way to prevent this is t

“we might as well stop training radiologists” Geoff Hinton, 2016, vs the actual data, via Torsten Slok at Apollo

SafetyDGX agent

Gary Marcus shares Torsten Slok's analysis comparing Geoffrey Hinton's 2016 prediction that radiologist training should cease due to AI capabilities against actual empirical data on AI performance in

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient

SafetyDGX agent

arXiv:2604.25872v1 Announce Type: new Abstract: Training language models via reinforcement learning often relies on imperfect proxy rewards, since ground truth rewards that precisely define the intend

Zuckerberg bets big on biology? About as much he put into 5 employees in his AI startup for a year 🙄 He is *far* more interested AI for sel…

SafetyDGX agent

Gary Marcus critiques Mark Zuckerberg's claimed focus on biology research, suggesting the financial commitment is minimal compared to his AI investments and stating that Zuckerberg's actual priorities

28 Apr 2026

A Comparative analysis of Layer-wise Representational Capacity in AR and Diffusion LLMs

SafetyDGX agent

arXiv:2603.07475v2 Announce Type: replace Abstract: Autoregressive (AR) language models build representations incrementally via left-to-right prediction, while diffusion language models (dLLMs) are tr

A Decoupled Human-in-the-Loop System for Controlled Autonomy in Agentic Workflows

SafetyDGX agent

arXiv:2604.23049v1 Announce Type: new Abstract: AI agents are increasingly deployed to execute tasks and make decisions within agentic workflows, introducing new requirements for safe and controlled a

A Differentiable Framework for Global Circulation Model Precipitation Bias Correction

SafetyDGX agent

arXiv:2604.23045v1 Announce Type: new Abstract: Systematic biases in Global Circulation Model (GCM) outputs limit their direct applicability in regional planning, necessitating bias correction. Correc

A Lightweight Explainable Guardrail for Prompt Safety

SafetyDGX agent

arXiv:2602.15853v2 Announce Type: replace-cross Abstract: We propose a lightweight explainable guardrail (LEG) method to detect unsafe prompts. LEG uses a multi-task learning architecture to jointly l

A Multi-Dimensional Audit of Politically Aligned Large Language Models

SafetyDGX agent

arXiv:2604.24429v1 Announce Type: new Abstract: As the application of Large Language Models (LLMs) spreads across various industries, there are increasing concerns about the potential for their misuse

A Reward-Free Viewpoint on Multi-Objective Reinforcement Learning

SafetyDGX agent

arXiv:2604.24532v1 Announce Type: new Abstract: Many sequential decision-making tasks involve optimizing multiple conflicting objectives, requiring policies that adapt to different user preferences. I

A Self-Supervised Framework for Space Object Behaviour Characterisation

SafetyDGX agent

arXiv:2504.06176v3 Announce Type: replace-cross Abstract: Foundation Models, which leverage large neural networks pre-trained on unlabelled data before fine-tuning for specific tasks, are increasingly

A Taxonomy and Resolution Strategy for Client-Level Disagreements in Federated Learning

SafetyDGX agent

arXiv:2604.23386v1 Announce Type: cross Abstract: Federated Learning (FL) typically assumes unconditional collaboration, a premise that overlooks the complexities of real-world, multi-stakeholder envi

Adaptive Multi-Subspace Representation Steering for Attribute Alignment in Large Language Models

SafetyDGX agent

arXiv:2508.10599v4 Announce Type: replace Abstract: Activation steering offers a promising approach to controlling the behavior of Large Language Models by directly manipulating their internal activat

AdaRubric: Task-Adaptive Rubrics for LLM Agent Evaluation

SafetyDGX agent

arXiv:2603.21362v2 Announce Type: replace Abstract: LLM-as-Judge evaluation fails agent tasks because a fixed rubric cannot capture what matters for this task: code debugging demands Correctness and E

Additive Control Variates Dominate Self-Normalisation in Off-Policy Evaluation

SafetyDGX agent

arXiv:2602.14914v2 Announce Type: replace Abstract: Off-policy evaluation (OPE) is essential for assessing ranking and recommendation systems without costly online interventions. Self-Normalised Inver

Adversary-Free Counterfactual Prediction via Information-Regularized Representations

SafetyDGX agent

arXiv:2510.15479v2 Announce Type: replace Abstract: We study counterfactual prediction under assignment bias and propose a mathematically grounded, information-theoretic approach that removes treatmen

AI Safety Training Can be Clinically Harmful

SafetyDGX agent

arXiv:2604.23445v1 Announce Type: cross Abstract: Large language models are being deployed as mental health support agents at scale, yet only 16% of LLM-based chatbot interventions have undergone rigo

Algorithmic Administration and the EU AI Act: Legal Principles for Public Sector Use of AI

SafetyDGX agent

arXiv:2604.22765v1 Announce Type: cross Abstract: The increasing use of artificial intelligence (AI) by public authorities introduces both opportunities for innovation and significant challenges for t

Aligning with Your Own Voice: Self-Corrected Preference Learning for Hallucination Mitigation in LVLMs

SafetyDGX agent

arXiv:2604.24395v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) frequently suffer from hallucinations. Existing preference learning-based approaches largely rely on proprietary mo

AMAVA: Adaptive Motion-Aware Video-to-Audio Framework for Visually-Impaired Assistance

SafetyDGX agent

arXiv:2604.23909v1 Announce Type: new Abstract: Navigational aids for blind and low vision individuals struggle conveying dynamic real-world environments, leading to cognitive overload from continuous

An Automatic Ground Collision Avoidance System with Reinforcement Learning

SafetyDGX agent

arXiv:2604.24403v1 Announce Type: new Abstract: This article evaluates an artificial intelligence (AI)-based Automatic Ground Collision Avoidance System (AGCAS) designed for advanced jet trainers to e

An empirical evaluation of the risks of AI model updates using clinical data: stability, arbitrariness, and fairness

SafetyDGX agent

arXiv:2604.23954v1 Announce Type: new Abstract: Artificial Intelligence and Machine Learning (AI/ML) models used in clinical settings are increasingly deployed to support clinical decision-making. How

An Information-Geometric Framework for Stability Analysis of Large Language Models under Entropic Stress

SafetyDGX agent

arXiv:2604.24076v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed in high-stakes and operational settings, evaluation strategies based solely on aggregate accur

Analytica: Soft Propositional Reasoning for Robust and Scalable LLM-Driven Analysis

SafetyDGX agent

arXiv:2604.23072v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly tasked with complex real-world analysis (e.g., in financial forecasting, scientific discovery), yet t

ANCHOR: LLM-driven Subject Conditioning for Text-to-Image Synthesis

SafetyDGX agent

arXiv:2404.10141v2 Announce Type: replace-cross Abstract: Text-to-image (T2I) models have achieved remarkable progress in high-quality image synthesis, yet most benchmarks rely on simple, self-contain

AnemiaVision: Non-Invasive Anemia Detection via Smartphone Imagery Using EfficientNet-B3 with TrivialAugmentWide, Mixup Augmentation, and Persistent Patient History Management

SafetyDGX agent

arXiv:2604.22964v1 Announce Type: new Abstract: Anemia affects over one billion people globally and remains severely under-diagnosed in low-resource regions where laboratory blood tests are inaccessib

Animalbooth: multimodal feature enhancement for animal subject personalization

SafetyDGX agent

arXiv:2509.16702v2 Announce Type: replace Abstract: Personalized animal image generation is challenging due to rich appearance cues and large morphological variability. Existing approaches often exhib

ArgRE: Formal Argumentation for Conflict Resolution in Multi-Agent Requirements Negotiation

SafetyDGX agent

arXiv:2604.23124v1 Announce Type: cross Abstract: As software systems grow in complexity, they must satisfy an increasing number of competing quality attributes, making it essential to balance them in

As an avid cyclist, I was amused to see ChatGPT's “powerful new image engine' draw a bicycle with the 'brake' label pointing to empty space …

SafetyDGX agent

As an avid cyclist, I was amused to see ChatGPT's “powerful new image engine' draw a bicycle with the 'brake' label pointing to empty space where brakes are sometimes found on other bicycles. The poin

Autocorrelation Reintroduces Spectral Bias in KANs for Time Series Forecasting

SafetyDGX agent

arXiv:2604.23518v1 Announce Type: cross Abstract: Existing theory suggests that Kolmogorov-Arnold Networks (KANs) can overcome the spectral bias commonly observed in neural networks under the assumpti

Bellman Residual Minimization for Control: Geometry, Stationarity, and Convergence

SafetyDGX agent

arXiv:2601.18840v3 Announce Type: replace Abstract: Markov decision problems are most commonly solved via dynamic programming. Another approach is Bellman residual minimization, which directly minimiz

Betting for Sim-to-Real Performance Evaluation

SafetyDGX agent

arXiv:2604.24018v1 Announce Type: new Abstract: This paper studies the problem of robot performance evaluation, focusing on how to obtain accurate and efficient estimates of real-world behavior under

Beyond Binary Out-of-Distribution Detection: Characterizing Distributional Shifts with Multi-Statistic Diffusion Trajectories

SafetyDGX agent

arXiv:2510.17381v2 Announce Type: replace Abstract: Detecting out-of-distribution (OOD) data is critical for machine learning, be it for safety reasons or to enable open-ended learning. However, beyon

Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models

SafetyDGX agent

arXiv:2502.14888v4 Announce Type: replace-cross Abstract: The success of vision-language models is primarily attributed to effective alignment across modalities such as vision and language. However, m

Beyond Match Maximization and Fairness: Retention-Optimized Two-Sided Matching

SafetyDGX agent

arXiv:2602.15752v2 Announce Type: replace Abstract: On two-sided matching platforms such as online dating and recruiting, recommendation algorithms often aim to maximize the total number of matches. H

BMD-45: A Large-Scale CCTV Vehicle Detection Dataset for Urban Traffic in Developing Cities

SafetyDGX agent

arXiv:2604.24419v1 Announce Type: new Abstract: Robust vehicle detection from fixed CCTV cameras is critical for Intelligent Transportation Systems. Yet existing benchmarks predominantly feature relat

Breaking Lock-In: Preserving Steerability under Low-Data VLA Post-Training

SafetyDGX agent

arXiv:2604.23121v1 Announce Type: cross Abstract: Have you ever post-trained a generalist vision-language-action (VLA) policy on a small demonstration dataset, only to find that it stops responding to

Bridging Reasoning and Action: Hybrid LLM-RL Framework for Efficient Cross-Domain Task-Oriented Dialogue

SafetyDGX agent

arXiv:2604.23345v1 Announce Type: new Abstract: Cross-domain task-oriented dialogue requires reasoning over implicit and explicit feasibility constraints while planning long-horizon, multi-turn action

BVI-Mamba: Video Enhancement Using a Visual State-Space Model for Low-Light and Underwater Environments

SafetyDGX agent

arXiv:2604.23655v1 Announce Type: new Abstract: Videos captured in low-light and underwater conditions often suffer from distortions such as noise, low contrast, color imbalance, and blur. These issue

CA-IDD: Cross-Attention Guided Identity-Conditional Diffusion for Identity-Consistent Face Swapping

SafetyDGX agent

arXiv:2604.24493v1 Announce Type: new Abstract: Face swapping aims to optimize realistic facial image generation by leveraging the identity of a source face onto a target face while preserving pose, e

Can Compact Language Models Search Like Agents? Distillation-Guided Policy Optimization for Preserving Agentic RAG Capabilities

SafetyDGX agent

arXiv:2508.20324v4 Announce Type: replace Abstract: Reinforcement Learning has emerged as a dominant post-training approach to elicit agentic RAG behaviors such as search and planning from language mo

CAP-CoT: Cycle Adversarial Prompt for Improving Chain of Thoughts in LLM Reasoning

SafetyDGX agent

arXiv:2604.23270v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting has emerged as a simple and effective way to elicit step-by-step solutions from large language models (LLMs). However,

CAPSULE: Control-Theoretic Action Perturbations for Safe Uncertainty-Aware Reinforcement Learning

SafetyDGX agent

arXiv:2604.23576v1 Announce Type: cross Abstract: Ensuring safe exploration in high-dimensional systems with unknown dynamics remains a significant challenge. Existing safe reinforcement learning meth

← Previous
1…174175176177178…212
Next →