AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,813 results
Safety

… and (if am not mistaken) the only one in the top 14 to have never turned a profit. Welcome to the world of vibe investing!

DGX agent

… and (if am not mistaken) the only one in the top 14 to have never turned a profit. Welcome to the world of vibe investing! OpenAI is likely to raise 60 billion in its IPO, more than double Saudi Ara

safetygary-marcus--x
21 May 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Anomaly-Informed Confidence Calibration for Vision-Based Safety Prediction

DGX agent

arXiv:2605.21109v1 Announce Type: new Abstract: Reliable confidence estimates are important for safely deploying vision-based controllers in autonomous racing, where safety predictions must be derived

safetyarxiv-cs-ro
21 May 2026
Safety

approximately equals “I am already rich and fuck you if you lose your job to AI”

DGX agent

approximately equals “I am already rich and fuck you if you lose your job to AI” Marc Andreessen to Joe Rogan: Why AI Workers Beat Human Workers⁣ ⁣ 'Never gets drunk. Never gets sick. Never gets depre

safetygary-marcus--x
21 May 2026
Safety

Automated Byzantine-Resilient Clustered Decentralized Federated Learning for Battery Intelligence in Connected EVs

DGX agent

arXiv:2605.21115v1 Announce Type: cross Abstract: Federated learning (FL) has emerged as a promising paradigm for managing electric vehicle (EV) battery data in intelligent transportation systems (ITS

safetyarxiv-cs-lg
21 May 2026
Safety

Automatically Learning Construction Injury Precursors from Text

DGX agent

arXiv:1907.11769v4 Announce Type: replace Abstract: In light of the increasing availability of digitally recorded safety reports in the construction industry, it is important to develop methods to exp

safetyarxiv-cs-cl
21 May 2026
Safety

AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged Signals

DGX agent

arXiv:2605.20643v1 Announce Type: cross Abstract: Self-distillation enables language models to learn on-policy from their own trajectories by using the same model as both student and teacher, with the

safetyarxiv-cs-cl
21 May 2026
Safety

Bayesian Preference Learning for Test-Time Steerable Reward Models

DGX agent

arXiv:2602.08819v2 Announce Type: replace-cross Abstract: Reward models are central to aligning language models with human preferences via reinforcement learning (RL). As RL is increasingly applied to

safetyarxiv-cs-cl
21 May 2026
Safety

Behavior-Consistent Deep Reinforcement Learning

DGX agent

arXiv:2605.21214v1 Announce Type: new Abstract: Reinforcement learning (RL) often exhibits high variance across training runs, leading to unreliable performance and posing a major challenge to deploym

safetyarxiv-cs-lg
21 May 2026
Safety

Beyond Text-to-SQL: An Agentic LLM System for Governed Enterprise Analytics APIs

DGX agent

arXiv:2605.21027v1 Announce Type: new Abstract: Enterprise analytics aims to make organizational data accessible for decision-making, yet non-technical users still face barriers when using traditional

safetyarxiv-cs-cl
21 May 2026
Safety

Beyond the Bellman Recursion: A Pontryagin-Guided Framework for Non-Exponential Discounting

DGX agent

arXiv:2605.20996v1 Announce Type: new Abstract: Most value-based and actor--critic reinforcement learning methods rely on Bellman-style recursions, yet these recursions collapse under non-exponential

safetyarxiv-cs-lg
21 May 2026
Safety

Biomedical AI may be headed for a replication crisis. (This work below is not about AI-generated reports; it’s about studies of biomedicine …

DGX agent

Biomedical AI may be headed for a replication crisis. (This work below is not about AI-generated reports; it’s about studies of biomedicine that use ML in their methods, and how they are evaluted.) In

safetygary-marcus--x
21 May 2026
Safety

Can Microcanonical Langevin Dynamics Leverage Mini-Batch Gradient Noise?

DGX agent

arXiv:2602.06500v2 Announce Type: replace Abstract: Scaling inference methods such as Markov chain Monte Carlo to high-dimensional models remains a central challenge in Bayesian deep learning. A promi

safetyarxiv-cs-lg
21 May 2026
Safety

Can Vision Models Truly Forget? Mirage: Representation-Level Certification of Visual Unlearning

DGX agent

arXiv:2605.20282v1 Announce Type: new Abstract: Machine unlearning in Vertical Federated Learning (VFL) has attracted growing interest, yet existing methods certify forgetting solely using output-leve

safetyarxiv-cs-cv
21 May 2026
Safety

can’t believe people assume that success on highly verifiable problems in math (where we don’t even know how many tests were performed and h…

DGX agent

can’t believe people assume that success on highly verifiable problems in math (where we don’t even know how many tests were performed and how many might have failed) iautomatically generalize to ever

safetygary-marcus--x
21 May 2026
Safety

Checking the math behind the latest headlines from OpenAI and Anthropic, link below:

DGX agent

Gary Marcus examines and fact-checks recent claims made by OpenAI and Anthropic in their public announcements and media coverage. The post links to detailed analysis questioning the mathematical valid

safetygary-marcus--x
21 May 2026
Safety

CHEM: Estimating and Understanding Hallucinations in Deep Learning for Image Processing

DGX agent

arXiv:2512.09806v2 Announce Type: replace Abstract: Deep learning-based methods have recently achieved significant success in image reconstruction problems. However, challenges have emerged, as these

safetyarxiv-cs-cv
21 May 2026
Safety

Choose Wisely and Privately: Proactive Client Selection for Fair and Efficient Federated Learning

DGX agent

arXiv:2605.20975v1 Announce Type: new Abstract: Federated Learning enables collaborative model training across decentralized data sources without data transfer. Averaging-based FL is limited by the pr

safetyarxiv-cs-lg
21 May 2026
Safety

Comparative Evaluation of Deep Learning Models for Fake Image Detection

DGX agent

arXiv:2605.20971v1 Announce Type: new Abstract: The growing sophistication of GAN-based image manipulation presents significant challenges for digital forensics. This study compares the performance of

safetyarxiv-cs-cv
21 May 2026
Safety

Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models

DGX agent

arXiv:2602.02304v2 Announce Type: replace-cross Abstract: Large-scale foundation models exhibit behavioral shifts when subjected to interventions such as scaling, fine-tuning, reinforcement learning w

safetyarxiv-cs-lg
21 May 2026
Safety

CompilerKV: Risk-Adaptive KV Compression via Offline Experience Compilation

DGX agent

arXiv:2602.08686v2 Announce Type: replace Abstract: Prefill-only KV compression freezes a token subset at the end of prefill and decodes from it without further eviction. The retention decision is the

safetyarxiv-cs-lg
21 May 2026
Safety

Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs

DGX agent

arXiv:2605.20555v1 Announce Type: new Abstract: We introduce a novel method that averages the logits of a frozen reference policy (e.g., SFT) and a trainable policy, and incorporate the method into Gr

safetyarxiv-cs-lg
21 May 2026
Safety

Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment

DGX agent

arXiv:2605.20834v1 Announce Type: cross Abstract: Direct Preference Optimization (DPO) has emerged as a popular alternative to Reinforcement Learning from Human Feedback (RLHF), offering theoretical e

safetyarxiv-cs-lg
21 May 2026
Safety

Conflict-Aware Active Perception and Control in 3D Gaussian Splatting Fields via Control Barrier Functions

DGX agent

arXiv:2605.20566v1 Announce Type: new Abstract: Active perception in uncertain environments requires robots to navigate safely while acquiring informative observations to reduce map uncertainty. These

safetyarxiv-cs-ro
21 May 2026
Safety

Consistently Informative Soft-Label Temperature for Knowledge Distillation

DGX agent

arXiv:2605.20357v1 Announce Type: new Abstract: Knowledge distillation (KD) transfers knowledge from a high-capacity teacher to a compact student by matching their predictive distributions, with tempe

safetyarxiv-cs-lg
21 May 2026
Safety

Correcting Stochastic Update Bias in Preconditioned Language Model Optimizers

DGX agent

arXiv:2605.20756v1 Announce Type: new Abstract: Preconditioned optimizers are central to language model training, but their stochastic update rules are usually treated as direct approximations to popu

safetyarxiv-cs-lg
21 May 2026
Safety

CRAFT: Conflict-Resolved Aggregation for Federated Training

DGX agent

arXiv:2605.21317v1 Announce Type: new Abstract: The aggregation of conflicting client updates remains a fundamental bottleneck in federated learning (FL) under heterogeneous data distributions. Naive

safetyarxiv-cs-lg
21 May 2026
Safety

Cross-lingual robustness of LLM-brain alignment and its computational roots

DGX agent

arXiv:2605.21049v1 Announce Type: new Abstract: Large language models (LLMs) reliably predict neural activity during language comprehension and transformer depth has been interpreted as mirroring hier

safetyarxiv-cs-cl
21 May 2026
Safety

Cumulative Meta-Learning from Active Learning Queries for Robustness to Spurious Correlations

DGX agent

arXiv:2605.20771v1 Announce Type: new Abstract: Spurious correlations in real-world datasets cause machine learning models to rely on irrelevant patterns, undermining reliability, generalization, and

safetyarxiv-cs-lg
21 May 2026
Safety

Data-Efficient Neural Operator Training via Physics-Based Active Learning

DGX agent

arXiv:2605.21348v1 Announce Type: new Abstract: Solving partial differential equations with neural operators significantly reduces computational costs but remains bottlenecked by high training data re

safetyarxiv-cs-lg
21 May 2026
Safety

Decision-Path Patterns as Tree Reliability Signals: Path-based Adaptive Weighting for Random Forest Classification

DGX agent

arXiv:2605.20716v1 Announce Type: new Abstract: Random forests aggregate tree votes by simple majority, treating all trees as equally informative. We observe that the topological pattern along each tr

safetyarxiv-cs-lg
21 May 2026
Safety

Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor

DGX agent

arXiv:2605.20402v1 Announce Type: new Abstract: MXFP4 arithmetic can dramatically accelerate reinforcement learning (RL) post-training of large language models (LLMs), yet the quantization error intro

safetyarxiv-cs-lg
21 May 2026
Safety

DeCoR: Design and Control Co-Optimization for Urban Streets Using Reinforcement Learning

DGX agent

arXiv:2605.21311v1 Announce Type: new Abstract: Modern vision systems can detect, track, and forecast urban actors at scale, yet translating perception outputs to urban design remains limited. We intr

safetyarxiv-cs-lg
21 May 2026
Safety

Decoupling Communication from Policy: Robust MARL under Bandwidth Constraints

DGX agent

arXiv:2605.21085v1 Announce Type: cross Abstract: Communication enables coordination in multi-agent reinforcement learning (MARL), but many real-world applications, e.g., search-and-rescue with drone

safetyarxiv-cs-lg
21 May 2026
Safety

Deep Attention Reweighting: Post-Hoc Attention-Based Feature Aggregation in CNNs for Disentangling Core and Spurious Features under Spurious Correlations

DGX agent

arXiv:2605.20732v1 Announce Type: new Abstract: Convolutional Neural Networks (CNNs) often exploit spurious correlations in datasets, learning superficially predictive yet causally irrelevant features

safetyarxiv-cs-cv
21 May 2026
Safety

DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards

DGX agent

arXiv:2605.21467v1 Announce Type: cross Abstract: Reinforcement learning from verifiable rewards (RLVR) has emerged as a central technique for improving the reasoning capabilities of large language mo

safetyarxiv-cs-cl
21 May 2026
Safety

Design for Manufacturing: A Manufacturability Knowledge-Integrated Reinforcement Learning Framework for Free-Form Pipe Routing in Aeroengines

DGX agent

arXiv:2605.20644v1 Announce Type: new Abstract: Design for manufacturing plays a critical role in advanced aeroengine development, where complex components necessitate careful consideration of manufac

safetyarxiv-cs-lg
21 May 2026
Safety

Disentangling Bias by Modeling Intra- and Inter-modal Causal Attention for Multimodal Sentiment Analysis

DGX agent

arXiv:2508.04999v2 Announce Type: replace Abstract: Multimodal sentiment analysis (MSA) aims to understand human emotions by integrating information from multiple modalities, such as text, audio, and

safetyarxiv-cs-lg
21 May 2026
Safety

Distill to Think, Foresee to Act: Cognitive-Physical Reinforcement Learning for Autonomous Driving

DGX agent

arXiv:2605.21139v1 Announce Type: new Abstract: Current end-to-end autonomous driving models are fundamentally constrained by the behavioral cloning ceiling of imitation learning. While reinforcement

safetyarxiv-cs-cv
21 May 2026
Safety

Distributed Direct Preference Optimization

DGX agent

arXiv:2605.20696v1 Announce Type: new Abstract: Preference-based reinforcement learning (RL) is a key paradigm for aligning policies with human judgments, yet its theoretical behavior in distributed s

safetyarxiv-cs-lg
21 May 2026
Safety

Distribution-Aware Reward: Reinforcement Learning over Predictive Distributions for LLM Regression

DGX agent

arXiv:2605.20740v1 Announce Type: cross Abstract: Large language models can predict real-valued quantities from heterogeneous inputs such as text, code, and molecular strings, but most training object

safetyarxiv-cs-cl
21 May 2026
Safety

Distributional Alignment as a Criterion for Designing Task Vectors in In-Context Learning

DGX agent

arXiv:2605.20730v1 Announce Type: new Abstract: In-context learning (ICL) allows large language models (LLMs) to adapt to new tasks through demonstrations, yet it suffers from escalating inference cos

safetyarxiv-cs-cl
21 May 2026
Safety

Divide-Prompt-Refine: a Training-Free, Structure-Aware Framework for Biomedical Abstract Generation

DGX agent

arXiv:2605.20628v1 Announce Type: new Abstract: Biomedical abstracts play a critical role in downstream NLP applications, such as information retrieval, biocuration, and biomedical knowledge discovery

safetyarxiv-cs-cl
21 May 2026
Safety

Do agents listen to you… or themselves? While evaling subagent behavior in deep agent systems, we noticed an interesting quirk in our agents…

DGX agent

Do agents listen to you… or themselves? While evaling subagent behavior in deep agent systems, we noticed an interesting quirk in our agents' alignment with hand-written system prompts vs. the instruc

safetyharrison-chase--x
21 May 2026
Safety

Do No Harm? Hallucination and Actor-Level Abuse in Web-Deployed Medical Large Language Models

DGX agent

arXiv:2605.20591v1 Announce Type: new Abstract: Medical large language models (LLMs), including custom medical GPTs (MedGPTs) and open-source models, are increasingly deployed on web platforms to prov

safetyarxiv-cs-cl
21 May 2026
Safety

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards

DGX agent

arXiv:2605.21180v1 Announce Type: new Abstract: Large language models show strong potential for automated code generation, but lack guarantees for correctness, quality, safety, and domain-specific con

safetyarxiv-cs-lg
21 May 2026
Safety

Enhancing Speech Large Language Models through Reinforced Behavior Alignment

DGX agent

arXiv:2509.03526v2 Announce Type: replace Abstract: The recent advancements of Large Language Models (LLMs) have spurred considerable research interest in extending their linguistic capabilities beyon

safetyarxiv-cs-cl
21 May 2026
Safety

EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models

DGX agent

arXiv:2510.05942v3 Announce Type: replace Abstract: We present EvalMORAAL, a transparent chain-of-thought (CoT) framework that uses two scoring methods (log-probabilities and direct ratings) plus a mo

safetyarxiv-cs-cl
21 May 2026
Safety

extit{Stochastic} MeanFlow Policies: One-Step Generative Control with Entropic Mirror Descent

DGX agent

arXiv:2605.21282v1 Announce Type: new Abstract: Online off-policy reinforcement learning (RL) is shaped by two coupled choices: the policy class and the update rule. Gaussian policies are fast and hav

safetyarxiv-cs-lg
21 May 2026
← Previous
1…157158159160161…267
Next →