AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,357 results
11 May 2026

MoCoTalk: Multi-Conditional Diffusion with Adaptive Router for Controllable Talking Head Generation

SafetyDGX agent

arXiv:2605.08050v1 Announce Type: new Abstract: Talking-head generation requires joint modeling of identity, head pose, facial expression, and mouth dynamics. Existing methods typically address only a

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models

SafetyDGX agent

arXiv:2602.07026v2 Announce Type: replace-cross Abstract: Despite the success of multimodal contrastive learning in aligning visual and linguistic representations, a persistent geometric anomaly, the

MPD^2-Router: Mask-aware Multi-expert Prior-regularized Dual-head Deferral Router in Glaucoma Screening and Diagnosis

SafetyDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.08024v1 Announce Type: new Abstract: Learning-to-defer (L2D) can make glaucoma screening safer by routing difficult/uncertain cases to humans, yet standard formulations overlook expert avai

Multi-environment Invariance Learning with Missing Data

SafetyDGX agent

arXiv:2601.07247v2 Announce Type: replace-cross Abstract: Learning models that can handle distribution shifts is a key challenge in domain generalization. Invariance learning, an approach that focuses

Multi-Environment POMDPs with Finite-Horizon Objectives

SafetyDGX agent

arXiv:2605.07537v1 Announce Type: new Abstract: Partially Observable Markov Decision Processes (POMDPs) are systems in which one agent interacts with a stochastic environment, and receives only partia

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation

SafetyDGX agent

arXiv:2603.16876v2 Announce Type: replace-cross Abstract: We propose MARL-Rad, a multi-modal multi-agent reinforcement learning framework for radiology report generation that trains the entire agentic

Multi-Objective Multi-Agent Bandits: From Learning Efficiency to Fairness Optimization

SafetyDGX agent

arXiv:2605.06864v1 Announce Type: new Abstract: We study multi-objective multi-agent multi-armed bandits (MO-MA-MAB) under stochastic rewards, where agents observe heterogeneous reward vectors and com

Mythos found a single vulnerability in cURL (along with three false positives, and one issue they classified as a bug). The founder/lead dev…

SafetyDGX agent

Mythos identified one genuine vulnerability in cURL while also reporting three false positives and one issue classified as a bug during their security analysis. The post references cURL's founder or l

@NameInteger @GaryMarcus @geoffreyhinton Hinton's argument was about encoding: LLMs don't encode text as text, but as weighted matrices. It …

SafetyDGX agent

@NameInteger @GaryMarcus @geoffreyhinton Hinton's argument was about encoding: LLMs don't encode text as text, but as weighted matrices. It makes no difference. Memorised data is memorised data no mat

No Forgetting Learning: Buffer-free Continual Learning Classification

SafetyDGX agent

arXiv:2503.04638v3 Announce Type: replace Abstract: Most Continual Learning (CL) methods maintain performance on earlier tasks by storing exemplars in a replay buffer, introducing memory overhead that

NoiseGate: Learning Per-Latent Timestep Schedules as Information Gating in World Action Models

SafetyDGX agent

arXiv:2605.07794v1 Announce Type: new Abstract: World Action Models (WAMs) are an emerging family of policies that tie robot action generation to future-observation modeling. In this work, we focus on

Not even surprised by horrific stories like these anymore. The mission of getting LLMs aligned with human values has largely been a failure.

SafetyDGX agent

Not even surprised by horrific stories like these anymore. The mission of getting LLMs aligned with human values has largely been a failure. NEW: ChatGPT advised the FSU shooter that a mass shooting w

OASES: Outcome-Aligned Search-Evaluation Co-Training for Agentic Search

SafetyDGX agent

arXiv:2604.03675v2 Announce Type: replace Abstract: Agentic search enables language models to solve knowledge-intensive tasks by adaptively acquiring external evidence over multiple steps. Reinforceme

Object Hallucination-Free Reinforcement Unlearning for Vision-Language Models

SafetyDGX agent

arXiv:2605.08031v1 Announce Type: new Abstract: Vision-language models (VLMs) raise growing concerns about privacy, copyright, and bias, motivating machine unlearning to remove sensitive knowledge. Ho

Offline Policy Optimization with Posterior Sampling

SafetyDGX agent

arXiv:2605.07393v1 Announce Type: new Abstract: A fundamental challenge in model-based offline reinforcement learning (RL) lies in the trade-off between generalization and robustness against exploitat

oh. my. god. 😱

SafetyDGX agent

oh. my. god. 😱 FT Exclusive: NHS England has granted external staff from companies including Palantir “unlimited access” to identifiable patient data while working on a part of its flagship data platf

On the Meta-Design of Allocation Problems

SafetyDGX agent

arXiv:2602.08786v4 Announce Type: replace-cross Abstract: There is an extensive literature that studies how to find optimal policies in resource allocation problems, taking the underlying design param

On Training in Imagination

SafetyDGX agent

arXiv:2605.06732v1 Announce Type: new Abstract: State-of-the-art model-based reinforcement learning methods train policies on imagined rollouts. These rollouts are trajectories generated by a learned

One estimate of how much annual revenue AI needs to “make sense”: 1.6 trillion. That’s four times what Google made in its best year. (total …

SafetyDGX agent

One estimate of how much annual revenue AI needs to “make sense”: 1.6 trillion. That’s four times what Google made in its best year. (total revenue so far is perhaps on order of 100 billion.) In 2024

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy

SafetyDGX agent

arXiv:2605.07931v1 Announce Type: cross Abstract: Vision-language-action (VLA) models increasingly rely on auxiliary world modules to plan over long horizons, yet how such modules should be parameteri

Online Allocation with Unknown Shared Supply

SafetyDGX agent

arXiv:2605.07080v1 Announce Type: new Abstract: Many real-world resource allocation systems, such as humanitarian logistics and vaccine distribution, must preposition limited supply across multiple lo

Openclaw token consumption fell by half in a month, per openrouter data What happened?

SafetyDGX agent

Openclaw token consumption dropped 50% within a month according to OpenRouter usage data, as reported by Gary Marcus. The post likely discusses potential causes for this significant decline, such as c

Optimal Recourse Summaries via Bi-Objective Decision Tree Learning

SafetyDGX agent

arXiv:2605.07598v1 Announce Type: new Abstract: Actionable Recourse provides individuals with actions they can take to change an unfavorable classifier outcome. While useful at the instance level, it

PACEvolve++: Improving Test-time Learning for Evolutionary Search Agents

SafetyDGX agent

arXiv:2605.07039v1 Announce Type: new Abstract: Large language models have become drivers of evolutionary search, but most systems rely on a fixed, prompt-elicited policy to sample next candidates. Th

Pan-FM: A Pan-Organ Foundation Model with Saliency-Guided Masking for Missing Robustness

SafetyDGX agent

arXiv:2605.07055v1 Announce Type: cross Abstract: Foundation models (FMs) have shown great promise in medical imaging, but most FMs are trained on unimodal data within isolated domains, such as brain

PaT: Planning-after-Trial for Efficient Test-Time Code Generation

SafetyDGX agent

arXiv:2605.07248v1 Announce Type: new Abstract: Beyond training-time optimization, scaling test-time computation has emerged as a key paradigm to extend the reasoning capabilities of Large Language Mo

Persistent-Transient Policy Evaluation for Markov Chains via Minimal Peripheral Quotients

SafetyDGX agent

arXiv:2602.00474v2 Announce Type: replace-cross Abstract: We study fixed-policy evaluation for finite Markov chains that may be reducible and periodic. Classical evaluation methods with gain and bias

Physical Simulators as Do-Operators: Causal Discovery under Latent Confounders for AI-for-Science

SafetyDGX agent

arXiv:2605.07467v1 Announce Type: cross Abstract: Existing interventional causal discovery methods -- IGSP, DCDI, ENCO -- assume causal sufficiency (no latent confounders) and rely on virtual interven

Physics-Based Benchmarking Metrics for Multimodal Synthetic Images

SafetyDGX agent

arXiv:2511.15204v3 Announce Type: replace-cross Abstract: Current state of the art measures like BLEU, CIDEr, VQA score, SigLIP-2 and CLIPScore are often unable to capture semantic or structural accur

PLOT: Progressive Localization via Optimal Transport in Neural Causal Abstraction

SafetyDGX agent

arXiv:2605.06979v1 Announce Type: cross Abstract: Causal abstraction offers a principled framework for mechanistic interpretability, aligning a high-level causal model with the low-level computation r

POETS: Uncertainty-Aware LLM Optimization via Compute-Efficient Policy Ensembles

SafetyDGX agent

arXiv:2605.07775v1 Announce Type: cross Abstract: Balancing exploration and exploitation is a core challenge in sequential decision-making and black-box optimization. We introduce POETS (extbf{Po}licy

Position: Mechanistic Interpretability Must Disclose Identification Assumptions for Causal Claims

SafetyDGX agent

arXiv:2605.08012v1 Announce Type: cross Abstract: Mechanistic interpretability papers increasingly use causal vocabulary: circuits, mediators, causal abstraction, monosemanticity. Such claims require

Post-training makes large language models less human-like

SafetyDGX agent

arXiv:2605.07632v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as surrogates for human participants, but it remains unclear which models best capture human behavi

ProtoSSL: Interpretable Prototype Learning from Unlabeled Time-Series Data

SafetyDGX agent

arXiv:2605.06943v1 Announce Type: new Abstract: In time-series domains where both predictive performance and interpretability are essential, deep neural networks achieve strong results but provide lim

Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignment

SafetyDGX agent

arXiv:2605.08064v1 Announce Type: new Abstract: Spatial intelligence in vision-language models (VLMs) attracts research interest with the practical demand to reason in the 3D world.Despite promising r

Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning

SafetyDGX agent

arXiv:2605.07804v1 Announce Type: cross Abstract: On-policy distillation (OPD) leverages dense teacher rewards to enhance reasoning models. However, scaling OPD to long-horizon tasks exposes a critica

Q-MMR: Off-Policy Evaluation via Recursive Reweighting and Moment Matching

SafetyDGX agent

arXiv:2605.06474v2 Announce Type: replace-cross Abstract: We present a novel theoretical framework, Q-MMR, for off-policy evaluation in finite-horizon MDPs. Q-MMR learns a set of scalar weights, one f

R-GTD: A Geometric Analysis of Gradient Temporal-Difference Learning in Singular Regimes

SafetyDGX agent

arXiv:2601.20599v2 Announce Type: replace-cross Abstract: Gradient temporal-difference (GTD) learning algorithms are widely used for off-policy policy evaluation with function approximation. However,

Radiologist-Guided Causal Concept Bottleneck Models for Chest X-Ray Interpretation

SafetyDGX agent

arXiv:2605.07785v1 Announce Type: new Abstract: Concept Bottleneck Models (CBMs) in medical imaging aim to improve model interpretability by predicting intermediate clinical concepts before final diag

Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners

SafetyDGX agent

arXiv:2605.08019v1 Announce Type: new Abstract: Humans rapidly learn abstract knowledge when encountering novel environments and flexibly deploy this knowledge to guide efficient and intelligent actio

ReasonEdit: Towards Interpretable Image Editing Evaluation via Reinforcement Learning

SafetyDGX agent

arXiv:2605.07477v1 Announce Type: new Abstract: Recent text-guided image editing (TIE) models have achieved remarkable progress, however, many edited results still suffer from artifacts, unintended mo

Receipts: https://open.substack.com/pub/garymarcus/p/deconstructing-geoffrey-hintons-weakest?r=8tdk6&utm_medium=ios

SafetyDGX agent

Gary Marcus analyzes and critiques Geoffrey Hinton's arguments regarding weaknesses in deep learning and artificial neural networks. The article likely examines specific technical or conceptual claims

ReCLIP++: Learn to Rectify the Bias of CLIP for Unsupervised Semantic Segmentation

SafetyDGX agent

arXiv:2408.06747v4 Announce Type: replace Abstract: Recent works utilize CLIP to perform the challenging unsupervised semantic segmentation task where only images without annotations are available. Ho

Reflections and New Directions for Human-Centered Large Language Models

SafetyDGX agent

arXiv:2605.06901v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly shaping the private and professional lives of users, with numerous applications in business, education, fi

Reinforcement Learning for Exponential Utility: Algorithms and Convergence in Discounted MDPs

SafetyDGX agent

arXiv:2605.08053v1 Announce Type: new Abstract: Reinforcement learning (RL) for exponential-utility optimization in discounted Markov decision processes (MDPs) lacks principled value-based algorithms.

RELO: Reinforcement Learning to Localize for Visual Object Tracking

SafetyDGX agent

arXiv:2605.07379v1 Announce Type: cross Abstract: Conventional visual object trackers localize targets using handcrafted spatial priors, often in the form of heatmaps. Such priors provide only surroga

Repeated Deceptive Path Planning against Learnable Observer

SafetyDGX agent

arXiv:2605.07174v1 Announce Type: new Abstract: We study the problem of deceptive path planning (DPP), where an agent aims to conceal its true destination from external observers. While existing work

reply to Hinton’s reply to me, for additional context:

SafetyDGX agent

reply to Hinton’s reply to me, for additional context: Dear @geoffreyhinton, I literally never said that AI systems “JUST regurgitate”; that’s plainly false. I don’t believe it, and I didn’t say it. (

Resource-Element Energy Difference for Noncoherent Over-the-Air Federated Learning

SafetyDGX agent

arXiv:2605.07263v1 Announce Type: cross Abstract: Over-the-air federated learning (OTA-FL) reduces uplink latency by exploiting waveform superposition, but conventional analog aggregation schemes typi

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding

SafetyDGX agent

arXiv:2605.07575v1 Announce Type: cross Abstract: Proactive streaming video understanding requires Video-LLMs to decide when to respond as a video unfolds, a task where existing methods often fall sho

Response Time Enhances Alignment with Heterogeneous Preferences

SafetyDGX agent

arXiv:2605.06987v1 Announce Type: new Abstract: Aligning large language models (LLMs) to human preferences typically relies on aggregating pooled feedback into a single reward model. However, this sta

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective

SafetyDGX agent

arXiv:2605.07331v1 Announce Type: cross Abstract: Reinforcement learning, including reinforcement learning with verifiable rewards (RLVR), has emerged as a powerful approach for LLM post-training. Cen

RIDER: 3D RNA Inverse Design with Reinforcement Learning-Guided Diffusion

SafetyDGX agent

arXiv:2602.16548v2 Announce Type: replace Abstract: The inverse design of RNA three-dimensional (3D) structures is crucial for engineering functional RNAs in synthetic biology and therapeutics. While

Risk-Consistent Multiclass Learning from Random Label-Subset Membership Queries

SafetyDGX agent

arXiv:2605.07413v1 Announce Type: new Abstract: Obtaining accurate class labels is often costly or unreliable, and may also be limited by privacy or other practical conditions. Compared with asking an

Robustness of Refugee-Matching Gains to Off-Policy Evaluation Choices

SafetyDGX agent

arXiv:2605.06686v1 Announce Type: new Abstract: Previous research has investigated the potential of refugee matching for boosting refugee outcomes, first considered by Bansak et al. (2018). This paper

Rollback-Free Stable Brick Structures Generation

SafetyDGX agent

arXiv:2605.06947v1 Announce Type: new Abstract: While autoregressive models have advanced 3D generation, creating physically stable brick structures remains a challenge due to the strict requirements

Rubric-based On-policy Distillation

SafetyDGX agent

arXiv:2605.07396v1 Announce Type: cross Abstract: On-policy distillation (OPD) is a powerful paradigm for model alignment, yet its reliance on teacher logits restricts its application to white-box sce

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence

SafetyDGX agent

arXiv:2605.06230v2 Announce Type: replace Abstract: As large models evolve from conversational assistants into autonomous agents, challenges increasingly arise from long-horizon decision making, tool

SAGE: Hierarchical LLM-Based Literary Evaluation through Ontology-Grounded Interpretive Dimensions

SafetyDGX agent

arXiv:2605.07102v1 Announce Type: new Abstract: Evaluating literary quality requires assessing interpretive dimensions such as cultural representation, emotional depth, and philosophical sophisticatio

Same Signal, Opposite Meaning: Direction-Informed Adaptive Learning for LLM Agents

SafetyDGX agent

arXiv:2605.06908v1 Announce Type: cross Abstract: Adaptive test-time compute for LLM agents aims to invoke extra computation only when it improves performance. Existing methods typically use confidenc

← Previous
1…178179180181182…240
Next →