AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
Safety

Beyond Surface Judgments: Human-Grounded Risk Evaluation of LLM-Generated Disinformation

DGX agent

arXiv:2604.06820v1 Announce Type: new Abstract: Large language models (LLMs) can generate persuasive narratives at scale, raising concerns about their potential use in disinformation campaigns. Assess

safetyarxiv-cs-ai
10 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Bias Redistribution in Visual Machine Unlearning: Does Forgetting One Group Harm Another?

DGX agent

arXiv:2604.08111v1 Announce Type: cross Abstract: Machine unlearning enables models to selectively forget training data, driven by privacy regulations such as GDPR and CCPA. However, its fairness impl

safetyarxiv-cs-cv
10 Apr 2026
Model Releases

Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules

DGX agent

arXiv:2604.06233v1 Announce Type: new Abstract: Safety-trained language models routinely refuse requests for help circumventing rules. But not all rules deserve compliance. When users ask for help eva

model-releasesarxiv-cs-ai
10 Apr 2026
Safety

Brain3D: EEG-to-3D Decoding of Visual Representations via Multimodal Reasoning

DGX agent

arXiv:2604.08068v1 Announce Type: new Abstract: Decoding visual information from electroencephalography (EEG) has recently achieved promising results, primarily focusing on reconstructing two-dimensio

safetyarxiv-cs-cv
10 Apr 2026
Safety

CAFP: A Post-Processing Framework for Group Fairness via Counterfactual Model Averaging

DGX agent

arXiv:2604.07009v1 Announce Type: new Abstract: Ensuring fairness in machine learning predictions is a critical challenge, especially when models are deployed in sensitive domains such as credit scori

safetyarxiv-cs-ai
10 Apr 2026
Safety

Candles are more regulated than AI. And candle manufacturers aren’t AFAIK lobbying for absolute freedom from liability if they fuck up. Open…

DGX agent

Candles are more regulated than AI. And candle manufacturers aren’t AFAIK lobbying for absolute freedom from liability if they fuck up. OpenAI is truly appalling. @GaryMarcus @ESYudkowsky Seriously fu

safetygary-marcus--x
10 Apr 2026
Safety

CNN-based Surface Temperature Forecasts with Ensemble Numerical Weather Prediction

DGX agent

arXiv:2507.18937v3 Announce Type: replace-cross Abstract: Due to limited computational resources, medium-range temperature forecasts typically rely on low-resolution numerical weather prediction (NWP)

safetyarxiv-cs-ai
10 Apr 2026
Safety

Contextualising (Im)plausible Events Triggers Figurative Language

DGX agent

arXiv:2604.07885v1 Announce Type: new Abstract: This work explores the connection between (non-)literalness and plausibility at the example of subject-verb-object events in English. We design a system

safetyarxiv-cs-cl
10 Apr 2026
Safety

Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs

DGX agent

arXiv:2604.07518v1 Announce Type: new Abstract: Vision-Language Models often struggle with complex visual reasoning due to the visual information loss in textual CoT. Existing methods either add the c

safetyarxiv-cs-cl
10 Apr 2026
Safety

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

DGX agent

arXiv:2604.08527v1 Announce Type: new Abstract: On-policy distillation (OPD) trains student models under their own induced distribution while leveraging supervision from stronger teachers. We identify

safetyarxiv-cs-cl
10 Apr 2026
Safety

Discrete Flow Matching Policy Optimization

DGX agent

arXiv:2604.06491v1 Announce Type: cross Abstract: We introduce Discrete flow Matching policy Optimization (DoMinO), a unified framework for Reinforcement Learning (RL) fine-tuning Discrete Flow Matchi

safetyarxiv-cs-ai
10 Apr 2026
Safety

Drift-Based Policy Optimization: Native One-Step Policy Learning for Online Robot Control

DGX agent

arXiv:2604.03540v2 Announce Type: replace Abstract: Although multi-step generative policies achieve strong performance in robotic manipulation by modeling multimodal action distributions, they require

safetyarxiv-cs-ro
10 Apr 2026
Safety

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation

DGX agent

arXiv:2602.13669v4 Announce Type: replace Abstract: Recent multi-modal video generation models have achieved high visual quality, but their prohibitive latency and limited temporal stability hinder re

safetyarxiv-cs-cv
10 Apr 2026
Safety

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World

DGX agent

arXiv:2604.07607v1 Announce Type: cross Abstract: Robot learning increasingly depends on large and diverse data, yet robot data collection remains expensive and difficult to scale. Egocentric human da

safetyarxiv-cs-cv
10 Apr 2026
Safety

Epic corruption, endless propaganda, security force crackdowns and false flag operations, election interference by a hostile foreign dictato…

DGX agent

Epic corruption, endless propaganda, security force crackdowns and false flag operations, election interference by a hostile foreign dictatorship… That’s Orbán’s Hungary today and, unless America wake

safetygary-marcus--x
10 Apr 2026
Safety

Equivariant Multi-agent Reinforcement Learning for Multimodal Vehicle-to-Infrastructure Systems

DGX agent

arXiv:2604.06914v1 Announce Type: new Abstract: In this paper, we study a vehicle-to-infrastructure (V2I) system where distributed base stations (BSs) acting as road-side units (RSUs) collect multimod

safetyarxiv-cs-lg
10 Apr 2026
Safety

Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization

DGX agent

arXiv:2604.08476v1 Announce Type: new Abstract: Multimodal reasoning models (MRMs) trained with reinforcement learning with verifiable rewards (RLVR) show improved accuracy on visual reasoning benchma

safetyarxiv-cs-cv
10 Apr 2026
Safety

Force-Aware Residual DAgger via Trajectory Editing for Precision Insertion with Impedance Control

DGX agent

arXiv:2603.04038v2 Announce Type: replace Abstract: Imitation learning (IL) has shown strong potential for contact-rich precision insertion tasks. However, its practical deployment is often hindered b

safetyarxiv-cs-ro
10 Apr 2026
Safety

FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling

DGX agent

arXiv:2604.06916v1 Announce Type: cross Abstract: Reinforcement-Learning-based post-training has recently emerged as a promising paradigm for aligning text-to-image diffusion models with human prefere

safetyarxiv-cs-ai
10 Apr 2026
Safety

From Ground Truth to Measurement: A Statistical Framework for Human Labeling

DGX agent

arXiv:2604.07591v1 Announce Type: cross Abstract: Supervised machine learning assumes that labeled data provide accurate measurements of the concepts models are meant to learn. Yet in practice, human

safetyarxiv-cs-cl
10 Apr 2026
Safety

FVD: Inference-Time Alignment of Diffusion Models via Fleming-Viot Resampling

DGX agent

arXiv:2604.06779v1 Announce Type: new Abstract: We introduce Fleming-Viot Diffusion (FVD), an inference-time alignment method that resolves the diversity collapse commonly observed in Sequential Monte

safetyarxiv-cs-ai
10 Apr 2026
Safety

Garry @Kasparov63 retired from competitive chess over twenty years ago, and most or all of his tournament games are presumably publicly avai…

DGX agent

Garry @Kasparov63 retired from competitive chess over twenty years ago, and most or all of his tournament games are presumably publicly available - and yet I bet he still could crush any LLM that didn

safetygary-marcus--x
10 Apr 2026
Safety

Geometric Properties of the Voronoi Tessellation in Latent Semantic Manifolds of Large Language Models

DGX agent

arXiv:2604.06767v1 Announce Type: new Abstract: Language models operate on discrete tokens but compute in continuous vector spaces, inducing a Voronoi tessellation over the representation manifold. We

safetyarxiv-cs-lg
10 Apr 2026
Safety

GIFT: Group-Relative Implicit Fine-Tuning Integrates GRPO with DPO and UNA

DGX agent

arXiv:2510.23868v4 Announce Type: replace Abstract: This paper proposes extit{Group-relative Implicit Fine-Tuning (GIFT)}, a reinforcement learning framework for aligning large language models (LLMs

safetyarxiv-cs-lg
10 Apr 2026
Safety

Governed Capability Evolution for Embodied Agents: Safe Upgrade, Compatibility Checking, and Runtime Rollback for Embodied Capability Modules

DGX agent

arXiv:2604.08059v1 Announce Type: new Abstract: Embodied agents are increasingly expected to improve over time by updating their executable capabilities rather than rewriting the agent itself. Prior w

safetyarxiv-cs-ro
10 Apr 2026
Safety

Grasp as You Dream: Imitating Functional Grasping from Generated Human Demonstrations

DGX agent

arXiv:2604.07517v1 Announce Type: new Abstract: Building generalist robots capable of performing functional grasping in everyday, open-world environments remains a significant challenge due to the vas

safetyarxiv-cs-ro
10 Apr 2026
Safety

Guiding a Diffusion Model by Swapping Its Tokens

DGX agent

arXiv:2604.08048v1 Announce Type: new Abstract: Classifier-Free Guidance (CFG) is a widely used inference-time technique to boost the image quality of diffusion models. Yet, its reliance on text condi

safetyarxiv-cs-cv
10 Apr 2026
Safety

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles

DGX agent

arXiv:2604.07650v1 Announce Type: cross Abstract: The rapid growth of the large language model (LLM) ecosystem raises a critical question: are seemingly diverse models truly independent? Shared pretra

safetyarxiv-cs-cl
10 Apr 2026
Safety

How Psychological Learning Paradigms Shaped and Constrained Artificial Intelligence

DGX agent

arXiv:2603.18203v3 Announce Type: replace Abstract: Current artificial intelligence systems struggle with systematic compositional reasoning: the capacity to recombine known components in novel config

safetyarxiv-cs-cl
10 Apr 2026
Safety

How to Evaluate Speech Translation with Source-Aware Neural MT Metrics

DGX agent

arXiv:2511.03295v3 Announce Type: replace-cross Abstract: Automatic evaluation of ST systems is typically performed by comparing translation hypotheses with one or more reference translations. While e

safetyarxiv-cs-ai
10 Apr 2026
Safety

HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns

DGX agent

arXiv:2601.10198v3 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in reasoning and generation, serving as the foundation for advanced persona s

safetyarxiv-cs-cl
10 Apr 2026
Safety

I agree totally @Gary. I’ve been saying this since the ChatGPT moment in Nov 22. Thank you for saying it out loud @demishassabis. Too much t…

DGX agent

I agree totally @Gary. I’ve been saying this since the ChatGPT moment in Nov 22. Thank you for saying it out loud @demishassabis. Too much time and energy being spent on mitigating the unintended cons

safetygary-marcus--x
10 Apr 2026
Safety

I rest my case: Mythos isn’t AGI. It’s not even better at biology than the last model. It’s tuned to particular things, not a giant advance …

DGX agent

I rest my case: Mythos isn’t AGI. It’s not even better at biology than the last model. It’s tuned to particular things, not a giant advance towards general intelligence. Same as it ever was. @GaryMarc

safetygary-marcus--x
10 Apr 2026
Safety

Improving Semantic Uncertainty Quantification in Language Model Question-Answering via Token-Level Temperature Scaling

DGX agent

arXiv:2604.07172v1 Announce Type: new Abstract: Calibration is central to reliable semantic uncertainty quantification, yet prior work has largely focused on discrimination, neglecting calibration. As

safetyarxiv-cs-lg
10 Apr 2026
Safety

Incremental Residual Reinforcement Learning Toward Real-World Learning for Social Navigation

DGX agent

arXiv:2604.07945v1 Announce Type: new Abstract: As the demand for mobile robots continues to increase, social navigation has emerged as a critical task, driving active research into deep reinforcement

safetyarxiv-cs-ro
10 Apr 2026
Safety

Indeed, if the LLM crew would just stick to this narrative, I would have a *lot* less to say 🤷‍♂️

DGX agent

Indeed, if the LLM crew would just stick to this narrative, I would have a *lot* less to say 🤷‍♂️ @GaryMarcus All we want is truth. Instead of overhyping LLMs, they should keep narative: -LLMs are use

safetygary-marcus--x
10 Apr 2026
Safety

Inside-Out: Measuring Generalization in Vision Transformers Through Inner Workings

DGX agent

arXiv:2604.08192v1 Announce Type: cross Abstract: Reliable generalization metrics are fundamental to the evaluation of machine learning models. Especially in high-stakes applications where labeled tar

safetyarxiv-cs-cv
10 Apr 2026
Safety

Just launched at @aiDotEngineer : our official AGI Pills! prescribe one (1) if your colleague is saying we are hitting a wall and/or trying …

DGX agent

Just launched at @aiDotEngineer : our official AGI Pills! prescribe one (1) if your colleague is saying we are hitting a wall and/or trying to add inductive bias instead of Trusting The Model Media bt

safetyswyx--x
10 Apr 2026
Safety

Karma Mechanisms for Decentralised, Cooperative Multi Agent Path Finding

DGX agent

arXiv:2604.07970v1 Announce Type: cross Abstract: Multi-Agent Path Finding (MAPF) is a fundamental coordination problem in large-scale robotic and cyber-physical systems, where multiple agents must co

safetyarxiv-cs-ro
10 Apr 2026
Safety

KD-MARL: Resource-Aware Knowledge Distillation in Multi-Agent Reinforcement Learning

DGX agent

arXiv:2604.06691v1 Announce Type: new Abstract: Real world deployment of multi agent reinforcement learning MARL systems is fundamentally constrained by limited compute memory and inference time. Whil

safetyarxiv-cs-ai
10 Apr 2026
Safety

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents

DGX agent

arXiv:2512.17445v2 Announce Type: replace Abstract: LangDriveCTRL is a natural-language-controllable framework for editing real-world driving videos to synthesize diverse traffic scenarios. It represe

safetyarxiv-cs-cv
10 Apr 2026
Safety

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning

DGX agent

arXiv:2604.07941v1 Announce Type: new Abstract: Post-training has become central to turning pretrained large language models (LLMs) into aligned and deployable systems. Recent progress spans supervise

safetyarxiv-cs-cl
10 Apr 2026
Safety

Learning to Negotiate: Multi-Agent Deliberation for Collective Value Alignment in LLMs

DGX agent

arXiv:2603.10476v2 Announce Type: replace Abstract: LLM alignment has progressed in single-agent settings through paradigms such as RL with human feedback (RLHF), while recent work explores scalable a

safetyarxiv-cs-cl
10 Apr 2026
Safety

Limits of Difficulty Scaling: Hard Samples Yield Diminishing Returns in GRPO-Tuned SLMs

DGX agent

arXiv:2604.06298v1 Announce Type: new Abstract: Recent alignment work on Large Language Models (LLMs) suggests preference optimization can improve reasoning by shifting probability mass toward better

safetyarxiv-cs-lg
10 Apr 2026
Safety

MCLR: Improving Conditional Modeling via Inter-Class Likelihood-Ratio Maximization and Unifying Classifier-Free Guidance with Alignment Objectives

DGX agent

arXiv:2603.22364v2 Announce Type: replace-cross Abstract: Diffusion models have achieved state-of-the-art performance in generative modeling, but their success often relies heavily on classifier-free

safetyarxiv-cs-cv
10 Apr 2026
Safety

MDP modeling for multi-stage stochastic programs

DGX agent

arXiv:2509.22981v2 Announce Type: replace Abstract: We study a class of multi-stage stochastic programs, which incorporate modeling features from Markov decision processes (MDPs). This class includes

safetyarxiv-cs-lg
10 Apr 2026
Safety

MemReader: From Passive to Active Extraction for Long-Term Agent Memory

DGX agent

arXiv:2604.07877v1 Announce Type: new Abstract: Long-term memory is fundamental for personalized and autonomous agents, yet populating it remains a bottleneck. Existing systems treat memory extraction

safetyarxiv-cs-cl
10 Apr 2026
Safety

Mixture Proportion Estimation and Weakly-supervised Kernel Test for Conditional Independence

DGX agent

arXiv:2604.07191v1 Announce Type: cross Abstract: Mixture proportion estimation (MPE) aims to estimate class priors from unlabeled data. This task is a critical component in weakly supervised learning

safetyarxiv-cs-ai
10 Apr 2026
← Previous
1…271272273274275…299
Next →