AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,485 results
22 May 2026

Reducing Political Manipulation with Consistency Training

SafetyDGX agent

arXiv:2605.22771v1 Announce Type: new Abstract: Large language models (LLMs) exhibit systematic political bias across a variety of sensitive contexts. We find that LLMs handle counterpart topics from

Reinforcing VLAs in Task-Agnostic World Models

SafetyDGX agent

arXiv:2605.12334v2 Announce Type: replace Abstract: Post-training Vision-Language-Action (VLA) models via reinforcement learning (RL) in learned world models has emerged as an effective strategy to ad

Retaining Suboptimal Actions to Follow Shifting Optima in Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2602.17062v2 Announce Type: replace Abstract: Value decomposition is a core approach for cooperative multi-agent reinforcement learning (MARL). However, existing methods still rely on a single o

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented Reasoning

SafetyDGX agent

arXiv:2605.22511v1 Announce Type: cross Abstract: Post-training has become the dominant recipe for turning a language model into a competent search-augmented reasoning agent. A line of recent work pus

Self-Policy Distillation via Capability-Selective Subspace Projection

SafetyDGX agent

arXiv:2605.22675v1 Announce Type: new Abstract: Self-distillation bootstraps large language models (LLMs) by training on their own generations. However, existing methods either rely on external signal

SENIOR: Efficient Query Selection and Preference-Guided Exploration in Preference-based Reinforcement Learning

SafetyDGX agent

arXiv:2506.14648v2 Announce Type: replace Abstract: Preference-based Reinforcement Learning (PbRL) methods provide a solution to avoid reward engineering by learning reward models based on human prefe

Supervised Classification Heads as Semantic Prototypes: Unlocking Vision-Language Alignment via Weight Recycling

SafetyDGX agent

arXiv:2605.22484v1 Announce Type: new Abstract: Vision-Language Models (VLMs) excel at tasks like zero-shot classification and cross-modal retrieval by mapping images and text to a shared space, but t

TacO: Benchmarking Tactile Sensors for Object Manipulation

SafetyDGX agent

arXiv:2605.21976v1 Announce Type: new Abstract: Vision-based learning from demonstrations has achieved remarkable success in enabling robots to perform manipulation tasks and high-level semantic reaso

The Double Dilemma in Multi-Task Radiology Report Generation: A Gradient Dynamics Analysis and Solution

SafetyDGX agent

arXiv:2605.22635v1 Announce Type: cross Abstract: While multi-task learning based automatic radiology report generation (RRG) is widely adopted to ensure clinical consistency, most focus on architectu

The Erdős Proof and AI Capabilities

SafetyDGX agent

View the official memo here. An internal model at OpenAI has autonomously disproved a central conjecture in discrete geometry, a mathematical field with applications in cryptography, wireless device c

The new White House policy requiring green card applicants to apply from outside the US is a capricious attack on legal immigration. It will…

SafetyDGX agent

The new White House policy requiring green card applicants to apply from outside the US is a capricious attack on legal immigration. It will hurt families, leave us with fewer doctors, teachers and sc

The state of LLMs in one video #AI

SafetyDGX agent

Gary Marcus discusses the current state of large language models (LLMs), likely covering their capabilities, limitations, and practical applications in AI. The post probably addresses key challenges s

The White House keeps handing Dems opportunities on silver platters to point out how Big Tech CEOs being in bed with lawmakers means the Ame…

SafetyDGX agent

The White House keeps handing Dems opportunities on silver platters to point out how Big Tech CEOs being in bed with lawmakers means the American people lose. New: The AI exec. order was postponed bec

This is the most interesting paper I have read this week. The authors test a wide range of LLMs on a massive dataset of behavioural experime…

SafetyDGX agent

This is the most interesting paper I have read this week. The authors test a wide range of LLMs on a massive dataset of behavioural experiments, with more than 200,000 participants and nearly 26 milli

To folks who claim I am always wrong, I have a few questions: Was I wrong that these systems would continue to hallucinate and be untrustwor…

SafetyDGX agent

To folks who claim I am always wrong, I have a few questions: Was I wrong that these systems would continue to hallucinate and be untrustworthy? That Sam was a liar? That these companies would struggl

TriSweep: A Four-Drone Swarm Framework for Electromagnetic Side-Channel Analysis

SafetyDGX agent

arXiv:2605.22709v1 Announce Type: cross Abstract: Electromagnetic (EM) side-channel analysis traditionally assumes a stationary, close-proximity probe - a threat model that underestimates aerial adver

Unifying Masked Diffusion Models with Various Generation Orders and Beyond

SafetyDGX agent

arXiv:2602.02112v2 Announce Type: replace-cross Abstract: Masked diffusion models (MDMs) are a potential alternative to autoregressive models (ARMs) for language generation, but generation quality dep

UniSD: Towards a Unified Self-Distillation Framework for Large Language Models

SafetyDGX agent

arXiv:2605.06597v2 Announce Type: replace Abstract: Self-distillation (SD) offers a promising path for adapting large language models (LLMs) without relying on stronger external teachers. However, SD

Universal CT Representations from Anatomy to Disease Phenotype through Agglomerative Pretraining

SafetyDGX agent

arXiv:2605.21906v1 Announce Type: new Abstract: Computed tomography (CT) is a central to three-dimensional medical imaging, yet CT-based artificial intelligence remains fragmented across task-specific

update @emollick does note some related caveats further in his thread; my objections are to the top post but there was some nuance below it,…

SafetyDGX agent

Gary Marcus responds to Ethan Mollick's post about AI capabilities, noting that while he objects to claims made in the top-level post, Mollick does provide important caveats and nuance in subsequent r

Value-Gradient Hypothesis of RL for LLMs

SafetyDGX agent

arXiv:2605.21654v1 Announce Type: cross Abstract: Reinforcement learning substantially improves pretrained language models, but it remains understudied why critic-free methods such as PPO and GRPO wor

Vector Policy Optimization: Training for Diversity Improves Test-Time Search

SafetyDGX agent

arXiv:2605.22817v1 Announce Type: cross Abstract: Language models must now generalize out of the box to novel environments and work inside inference-scaling search procedures, such as AlphaEvolve, tha

What Does the Caption Really Say? Counterfactual Phrase Intervention for Compositional Data Selection in Vision-Language Pretraining

SafetyDGX agent

arXiv:2605.22651v1 Announce Type: new Abstract: CLIP-style contrastive pretraining typically curates web-scale image-text pairs using sample-level filtering signals, often based on pair-level alignmen

💯: Why OpenAI keeps taking childish shots at me, is exactly what @FrankRundatz says below: “you have the added audacity of being influentia…

SafetyDGX agent

💯: Why OpenAI keeps taking childish shots at me, is exactly what @FrankRundatz says below: “you have the added audacity of being influential enough to move the needle on the timing of OpenAI’s record-

Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization

SafetyDGX agent

arXiv:2605.21801v1 Announce Type: cross Abstract: Post-training has become central to improving reasoning and alignment in large language models, where critic-free models enable scalable learning from

'Would You Want an AI Tutor?' Understanding Stakeholder Perceptions of LLM-based Systems in the Classroom

SafetyDGX agent

arXiv:2503.02885v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have gained traction in educational settings, often framed as virtual tutors or teaching assistants. Following ea

X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction

SafetyDGX agent

arXiv:2605.05765v2 Announce Type: replace Abstract: Inspired by the development of OpenClaw, there is a growing demand for mobile-based personal agents capable of handling complex and intuitive intera

21 May 2026

3D Reconstruction and Knowledge Distillation to Improve Multi-View Image Models to Explore Spike Volume Estimation in Wheat

SafetyDGX agent

arXiv:2605.20940v1 Announce Type: new Abstract: Accurate estimation of wheat spike volume is important for yield component analysis and stress resilience assessment, yet field-based measurement remain

A 10,000-Year Global Stochastic Tropical Cyclone Catalog with Wind-Dependent Track Transitions (WHITS)

SafetyDGX agent

arXiv:2605.20494v1 Announce Type: new Abstract: Reliable assessment of tropical cyclone (TC) risk is limited by the brevity and spatial sparsity of the historical record, particularly for the rare, hi

A Systematic Comparison between Extractive Self-Explanations and Human Rationales in Text Classification

SafetyDGX agent

arXiv:2410.03296v4 Announce Type: replace Abstract: Instruction-tuned LLMs are able to provide extit{an} explanation about their output to users by generating self-explanations, without requiring the

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation

SafetyDGX agent

arXiv:2605.21125v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO), a prominent algorithm within the Reinforcement Learning from Verifiable Rewards (RLVR) framework, has achieve

AFD-INSTRUCTION: A Comprehensive Antibody Instruction Dataset with Functional Annotations for LLM-Based Understanding and Design

SafetyDGX agent

arXiv:2602.04916v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have significantly advanced protein representation learning. However, their capacity to interpret and design anti

all these dudes posting about anthropic’s profits without checking the fine print. (see my earlier tweet) there’s some fine print i believe …

SafetyDGX agent

all these dudes posting about anthropic’s profits without checking the fine print. (see my earlier tweet) there’s some fine print i believe that n the openai result too, and a lot that has not been di

Always read the fine print: Anthropic is projecting its first (slightly) profitable quarter ever, which is amazing—assuming it actually happ…

SafetyDGX agent

Always read the fine print: Anthropic is projecting its first (slightly) profitable quarter ever, which is amazing—assuming it actually happens —but if it does it will be in no small part because they

… and (if am not mistaken) the only one in the top 14 to have never turned a profit. Welcome to the world of vibe investing!

SafetyDGX agent

… and (if am not mistaken) the only one in the top 14 to have never turned a profit. Welcome to the world of vibe investing! OpenAI is likely to raise 60 billion in its IPO, more than double Saudi Ara

approximately equals “I am already rich and fuck you if you lose your job to AI”

SafetyDGX agent

approximately equals “I am already rich and fuck you if you lose your job to AI” Marc Andreessen to Joe Rogan: Why AI Workers Beat Human Workers⁣ ⁣ 'Never gets drunk. Never gets sick. Never gets depre

Automated Byzantine-Resilient Clustered Decentralized Federated Learning for Battery Intelligence in Connected EVs

SafetyDGX agent

arXiv:2605.21115v1 Announce Type: cross Abstract: Federated learning (FL) has emerged as a promising paradigm for managing electric vehicle (EV) battery data in intelligent transportation systems (ITS

AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged Signals

SafetyDGX agent

arXiv:2605.20643v1 Announce Type: cross Abstract: Self-distillation enables language models to learn on-policy from their own trajectories by using the same model as both student and teacher, with the

Bayesian Preference Learning for Test-Time Steerable Reward Models

SafetyDGX agent

arXiv:2602.08819v2 Announce Type: replace-cross Abstract: Reward models are central to aligning language models with human preferences via reinforcement learning (RL). As RL is increasingly applied to

Behavior-Consistent Deep Reinforcement Learning

SafetyDGX agent

arXiv:2605.21214v1 Announce Type: new Abstract: Reinforcement learning (RL) often exhibits high variance across training runs, leading to unreliable performance and posing a major challenge to deploym

Beyond Text-to-SQL: An Agentic LLM System for Governed Enterprise Analytics APIs

SafetyDGX agent

arXiv:2605.21027v1 Announce Type: new Abstract: Enterprise analytics aims to make organizational data accessible for decision-making, yet non-technical users still face barriers when using traditional

Beyond the Bellman Recursion: A Pontryagin-Guided Framework for Non-Exponential Discounting

SafetyDGX agent

arXiv:2605.20996v1 Announce Type: new Abstract: Most value-based and actor--critic reinforcement learning methods rely on Bellman-style recursions, yet these recursions collapse under non-exponential

Biomedical AI may be headed for a replication crisis. (This work below is not about AI-generated reports; it’s about studies of biomedicine …

SafetyDGX agent

Biomedical AI may be headed for a replication crisis. (This work below is not about AI-generated reports; it’s about studies of biomedicine that use ML in their methods, and how they are evaluted.) In

Can Microcanonical Langevin Dynamics Leverage Mini-Batch Gradient Noise?

SafetyDGX agent

arXiv:2602.06500v2 Announce Type: replace Abstract: Scaling inference methods such as Markov chain Monte Carlo to high-dimensional models remains a central challenge in Bayesian deep learning. A promi

Can Vision Models Truly Forget? Mirage: Representation-Level Certification of Visual Unlearning

SafetyDGX agent

arXiv:2605.20282v1 Announce Type: new Abstract: Machine unlearning in Vertical Federated Learning (VFL) has attracted growing interest, yet existing methods certify forgetting solely using output-leve

can’t believe people assume that success on highly verifiable problems in math (where we don’t even know how many tests were performed and h…

SafetyDGX agent

can’t believe people assume that success on highly verifiable problems in math (where we don’t even know how many tests were performed and how many might have failed) iautomatically generalize to ever

Checking the math behind the latest headlines from OpenAI and Anthropic, link below:

SafetyDGX agent

Gary Marcus examines and fact-checks recent claims made by OpenAI and Anthropic in their public announcements and media coverage. The post links to detailed analysis questioning the mathematical valid

Choose Wisely and Privately: Proactive Client Selection for Fair and Efficient Federated Learning

SafetyDGX agent

arXiv:2605.20975v1 Announce Type: new Abstract: Federated Learning enables collaborative model training across decentralized data sources without data transfer. Averaging-based FL is limited by the pr

Comparative Evaluation of Deep Learning Models for Fake Image Detection

SafetyDGX agent

arXiv:2605.20971v1 Announce Type: new Abstract: The growing sophistication of GAN-based image manipulation presents significant challenges for digital forensics. This study compares the performance of

Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models

SafetyDGX agent

arXiv:2602.02304v2 Announce Type: replace-cross Abstract: Large-scale foundation models exhibit behavioral shifts when subjected to interventions such as scaling, fine-tuning, reinforcement learning w

CompilerKV: Risk-Adaptive KV Compression via Offline Experience Compilation

SafetyDGX agent

arXiv:2602.08686v2 Announce Type: replace Abstract: Prefill-only KV compression freezes a token subset at the end of prefill and decodes from it without further eviction. The retention decision is the

Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs

SafetyDGX agent

arXiv:2605.20555v1 Announce Type: new Abstract: We introduce a novel method that averages the logits of a frozen reference policy (e.g., SFT) and a trainable policy, and incorporate the method into Gr

Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment

SafetyDGX agent

arXiv:2605.20834v1 Announce Type: cross Abstract: Direct Preference Optimization (DPO) has emerged as a popular alternative to Reinforcement Learning from Human Feedback (RLHF), offering theoretical e

Consistently Informative Soft-Label Temperature for Knowledge Distillation

SafetyDGX agent

arXiv:2605.20357v1 Announce Type: new Abstract: Knowledge distillation (KD) transfers knowledge from a high-capacity teacher to a compact student by matching their predictive distributions, with tempe

Correcting Stochastic Update Bias in Preconditioned Language Model Optimizers

SafetyDGX agent

arXiv:2605.20756v1 Announce Type: new Abstract: Preconditioned optimizers are central to language model training, but their stochastic update rules are usually treated as direct approximations to popu

CRAFT: Conflict-Resolved Aggregation for Federated Training

SafetyDGX agent

arXiv:2605.21317v1 Announce Type: new Abstract: The aggregation of conflicting client updates remains a fundamental bottleneck in federated learning (FL) under heterogeneous data distributions. Naive

Cross-lingual robustness of LLM-brain alignment and its computational roots

SafetyDGX agent

arXiv:2605.21049v1 Announce Type: new Abstract: Large language models (LLMs) reliably predict neural activity during language comprehension and transformer depth has been interpreted as mirroring hier

Cumulative Meta-Learning from Active Learning Queries for Robustness to Spurious Correlations

SafetyDGX agent

arXiv:2605.20771v1 Announce Type: new Abstract: Spurious correlations in real-world datasets cause machine learning models to rely on irrelevant patterns, undermining reliability, generalization, and

Data-Efficient Neural Operator Training via Physics-Based Active Learning

SafetyDGX agent

arXiv:2605.21348v1 Announce Type: new Abstract: Solving partial differential equations with neural operators significantly reduces computational costs but remains bottlenecked by high training data re

Decision-Path Patterns as Tree Reliability Signals: Path-based Adaptive Weighting for Random Forest Classification

SafetyDGX agent

arXiv:2605.20716v1 Announce Type: new Abstract: Random forests aggregate tree votes by simple majority, treating all trees as equally informative. We observe that the topological pattern along each tr

← Previous
1…153154155156157…242
Next →