AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,812 results
22 May 2026

Unifying Masked Diffusion Models with Various Generation Orders and Beyond

SafetyDGX agent

arXiv:2602.02112v2 Announce Type: replace-cross Abstract: Masked diffusion models (MDMs) are a potential alternative to autoregressive models (ARMs) for language generation, but generation quality dep

UniSD: Towards a Unified Self-Distillation Framework for Large Language Models

SafetyDGX agent

arXiv:2605.06597v2 Announce Type: replace Abstract: Self-distillation (SD) offers a promising path for adapting large language models (LLMs) without relying on stronger external teachers. However, SD

Universal CT Representations from Anatomy to Disease Phenotype through Agglomerative Pretraining

SafetyDGX agent

arXiv:2605.21906v1 Announce Type: new Abstract: Computed tomography (CT) is a central to three-dimensional medical imaging, yet CT-based artificial intelligence remains fragmented across task-specific


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

update @emollick does note some related caveats further in his thread; my objections are to the top post but there was some nuance below it,…

SafetyDGX agent

Gary Marcus responds to Ethan Mollick's post about AI capabilities, noting that while he objects to claims made in the top-level post, Mollick does provide important caveats and nuance in subsequent r

update @thestalwart found more recent models less vulnerable. would be good to do a broad study of this.

SafetyDGX agent

Gary Marcus notes that researcher @thestalwart has found more recent AI models to be less vulnerable to certain attacks or exploits, and suggests that a comprehensive study across multiple models woul

Value-Gradient Hypothesis of RL for LLMs

SafetyDGX agent

arXiv:2605.21654v1 Announce Type: cross Abstract: Reinforcement learning substantially improves pretrained language models, but it remains understudied why critic-free methods such as PPO and GRPO wor

Vector Policy Optimization: Training for Diversity Improves Test-Time Search

SafetyDGX agent

arXiv:2605.22817v1 Announce Type: cross Abstract: Language models must now generalize out of the box to novel environments and work inside inference-scaling search procedures, such as AlphaEvolve, tha

What Does the Caption Really Say? Counterfactual Phrase Intervention for Compositional Data Selection in Vision-Language Pretraining

SafetyDGX agent

arXiv:2605.22651v1 Announce Type: new Abstract: CLIP-style contrastive pretraining typically curates web-scale image-text pairs using sample-level filtering signals, often based on pair-level alignmen

💯: Why OpenAI keeps taking childish shots at me, is exactly what @FrankRundatz says below: “you have the added audacity of being influentia…

SafetyDGX agent

💯: Why OpenAI keeps taking childish shots at me, is exactly what @FrankRundatz says below: “you have the added audacity of being influential enough to move the needle on the timing of OpenAI’s record-

Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization

SafetyDGX agent

arXiv:2605.21801v1 Announce Type: cross Abstract: Post-training has become central to improving reasoning and alignment in large language models, where critic-free models enable scalable learning from

'Would You Want an AI Tutor?' Understanding Stakeholder Perceptions of LLM-based Systems in the Classroom

SafetyDGX agent

arXiv:2503.02885v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have gained traction in educational settings, often framed as virtual tutors or teaching assistants. Following ea

X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction

SafetyDGX agent

arXiv:2605.05765v2 Announce Type: replace Abstract: Inspired by the development of OpenClaw, there is a growing demand for mobile-based personal agents capable of handling complex and intuitive intera

Yeah, I don't think I've ever met someone who has directly worked for an extended period on safety or security at an AI company who thinks t…

SafetyDGX agent

Yeah, I don't think I've ever met someone who has directly worked for an extended period on safety or security at an AI company who thinks things are fine readiness wise or incentive wise etc. https:/

21 May 2026

3D Reconstruction and Knowledge Distillation to Improve Multi-View Image Models to Explore Spike Volume Estimation in Wheat

SafetyDGX agent

arXiv:2605.20940v1 Announce Type: new Abstract: Accurate estimation of wheat spike volume is important for yield component analysis and stress resilience assessment, yet field-based measurement remain

A 10,000-Year Global Stochastic Tropical Cyclone Catalog with Wind-Dependent Track Transitions (WHITS)

SafetyDGX agent

arXiv:2605.20494v1 Announce Type: new Abstract: Reliable assessment of tropical cyclone (TC) risk is limited by the brevity and spatial sparsity of the historical record, particularly for the rare, hi

A Systematic Comparison between Extractive Self-Explanations and Human Rationales in Text Classification

SafetyDGX agent

arXiv:2410.03296v4 Announce Type: replace Abstract: Instruction-tuned LLMs are able to provide extit{an} explanation about their output to users by generating self-explanations, without requiring the

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation

SafetyDGX agent

arXiv:2605.21125v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO), a prominent algorithm within the Reinforcement Learning from Verifiable Rewards (RLVR) framework, has achieve

AFD-INSTRUCTION: A Comprehensive Antibody Instruction Dataset with Functional Annotations for LLM-Based Understanding and Design

SafetyDGX agent

arXiv:2602.04916v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have significantly advanced protein representation learning. However, their capacity to interpret and design anti

AI-Assisted Competency Assessment from Egocentric Video in Simulation-Based Nursing Education

SafetyDGX agent

arXiv:2605.20233v1 Announce Type: new Abstract: Assessing learner competency in clinical simulation requires expert observation that is time-intensive, difficult to scale, and subject to inter-rater v

AI-based Prediction of Independent Construction Safety Outcomes from Universal Attributes

SafetyDGX agent

arXiv:1908.05972v3 Announce Type: replace Abstract: This paper significantly improves on, and finishes to validate, an approach proposed in previous research in which safety outcomes were predicted fr

AIMBio-Mat: An AI-Native FAIR Platform for Closed-Loop Materials Discovery and Biomedical Translation

SafetyDGX agent

arXiv:2605.21083v1 Announce Type: cross Abstract: Materials discovery and biomedical translation increasingly require models that can reason across composition, processing, structure, biological respo

all these dudes posting about anthropic’s profits without checking the fine print. (see my earlier tweet) there’s some fine print i believe …

SafetyDGX agent

all these dudes posting about anthropic’s profits without checking the fine print. (see my earlier tweet) there’s some fine print i believe that n the openai result too, and a lot that has not been di

Always read the fine print: Anthropic is projecting its first (slightly) profitable quarter ever, which is amazing—assuming it actually happ…

SafetyDGX agent

Always read the fine print: Anthropic is projecting its first (slightly) profitable quarter ever, which is amazing—assuming it actually happens —but if it does it will be in no small part because they

… and (if am not mistaken) the only one in the top 14 to have never turned a profit. Welcome to the world of vibe investing!

SafetyDGX agent

… and (if am not mistaken) the only one in the top 14 to have never turned a profit. Welcome to the world of vibe investing! OpenAI is likely to raise 60 billion in its IPO, more than double Saudi Ara

Anomaly-Informed Confidence Calibration for Vision-Based Safety Prediction

SafetyDGX agent

arXiv:2605.21109v1 Announce Type: new Abstract: Reliable confidence estimates are important for safely deploying vision-based controllers in autonomous racing, where safety predictions must be derived

approximately equals “I am already rich and fuck you if you lose your job to AI”

SafetyDGX agent

approximately equals “I am already rich and fuck you if you lose your job to AI” Marc Andreessen to Joe Rogan: Why AI Workers Beat Human Workers⁣ ⁣ 'Never gets drunk. Never gets sick. Never gets depre

Automated Byzantine-Resilient Clustered Decentralized Federated Learning for Battery Intelligence in Connected EVs

SafetyDGX agent

arXiv:2605.21115v1 Announce Type: cross Abstract: Federated learning (FL) has emerged as a promising paradigm for managing electric vehicle (EV) battery data in intelligent transportation systems (ITS

Automatically Learning Construction Injury Precursors from Text

SafetyDGX agent

arXiv:1907.11769v4 Announce Type: replace Abstract: In light of the increasing availability of digitally recorded safety reports in the construction industry, it is important to develop methods to exp

AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged Signals

SafetyDGX agent

arXiv:2605.20643v1 Announce Type: cross Abstract: Self-distillation enables language models to learn on-policy from their own trajectories by using the same model as both student and teacher, with the

Bayesian Preference Learning for Test-Time Steerable Reward Models

SafetyDGX agent

arXiv:2602.08819v2 Announce Type: replace-cross Abstract: Reward models are central to aligning language models with human preferences via reinforcement learning (RL). As RL is increasingly applied to

Behavior-Consistent Deep Reinforcement Learning

SafetyDGX agent

arXiv:2605.21214v1 Announce Type: new Abstract: Reinforcement learning (RL) often exhibits high variance across training runs, leading to unreliable performance and posing a major challenge to deploym

Beyond Text-to-SQL: An Agentic LLM System for Governed Enterprise Analytics APIs

SafetyDGX agent

arXiv:2605.21027v1 Announce Type: new Abstract: Enterprise analytics aims to make organizational data accessible for decision-making, yet non-technical users still face barriers when using traditional

Beyond the Bellman Recursion: A Pontryagin-Guided Framework for Non-Exponential Discounting

SafetyDGX agent

arXiv:2605.20996v1 Announce Type: new Abstract: Most value-based and actor--critic reinforcement learning methods rely on Bellman-style recursions, yet these recursions collapse under non-exponential

Biomedical AI may be headed for a replication crisis. (This work below is not about AI-generated reports; it’s about studies of biomedicine …

SafetyDGX agent

Biomedical AI may be headed for a replication crisis. (This work below is not about AI-generated reports; it’s about studies of biomedicine that use ML in their methods, and how they are evaluted.) In

Can Microcanonical Langevin Dynamics Leverage Mini-Batch Gradient Noise?

SafetyDGX agent

arXiv:2602.06500v2 Announce Type: replace Abstract: Scaling inference methods such as Markov chain Monte Carlo to high-dimensional models remains a central challenge in Bayesian deep learning. A promi

Can Vision Models Truly Forget? Mirage: Representation-Level Certification of Visual Unlearning

SafetyDGX agent

arXiv:2605.20282v1 Announce Type: new Abstract: Machine unlearning in Vertical Federated Learning (VFL) has attracted growing interest, yet existing methods certify forgetting solely using output-leve

can’t believe people assume that success on highly verifiable problems in math (where we don’t even know how many tests were performed and h…

SafetyDGX agent

can’t believe people assume that success on highly verifiable problems in math (where we don’t even know how many tests were performed and how many might have failed) iautomatically generalize to ever

Checking the math behind the latest headlines from OpenAI and Anthropic, link below:

SafetyDGX agent

Gary Marcus examines and fact-checks recent claims made by OpenAI and Anthropic in their public announcements and media coverage. The post links to detailed analysis questioning the mathematical valid

CHEM: Estimating and Understanding Hallucinations in Deep Learning for Image Processing

SafetyDGX agent

arXiv:2512.09806v2 Announce Type: replace Abstract: Deep learning-based methods have recently achieved significant success in image reconstruction problems. However, challenges have emerged, as these

Choose Wisely and Privately: Proactive Client Selection for Fair and Efficient Federated Learning

SafetyDGX agent

arXiv:2605.20975v1 Announce Type: new Abstract: Federated Learning enables collaborative model training across decentralized data sources without data transfer. Averaging-based FL is limited by the pr

Comparative Evaluation of Deep Learning Models for Fake Image Detection

SafetyDGX agent

arXiv:2605.20971v1 Announce Type: new Abstract: The growing sophistication of GAN-based image manipulation presents significant challenges for digital forensics. This study compares the performance of

Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models

SafetyDGX agent

arXiv:2602.02304v2 Announce Type: replace-cross Abstract: Large-scale foundation models exhibit behavioral shifts when subjected to interventions such as scaling, fine-tuning, reinforcement learning w

CompilerKV: Risk-Adaptive KV Compression via Offline Experience Compilation

SafetyDGX agent

arXiv:2602.08686v2 Announce Type: replace Abstract: Prefill-only KV compression freezes a token subset at the end of prefill and decodes from it without further eviction. The retention decision is the

Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs

SafetyDGX agent

arXiv:2605.20555v1 Announce Type: new Abstract: We introduce a novel method that averages the logits of a frozen reference policy (e.g., SFT) and a trainable policy, and incorporate the method into Gr

Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment

SafetyDGX agent

arXiv:2605.20834v1 Announce Type: cross Abstract: Direct Preference Optimization (DPO) has emerged as a popular alternative to Reinforcement Learning from Human Feedback (RLHF), offering theoretical e

Conflict-Aware Active Perception and Control in 3D Gaussian Splatting Fields via Control Barrier Functions

SafetyDGX agent

arXiv:2605.20566v1 Announce Type: new Abstract: Active perception in uncertain environments requires robots to navigate safely while acquiring informative observations to reduce map uncertainty. These

Consistently Informative Soft-Label Temperature for Knowledge Distillation

SafetyDGX agent

arXiv:2605.20357v1 Announce Type: new Abstract: Knowledge distillation (KD) transfers knowledge from a high-capacity teacher to a compact student by matching their predictive distributions, with tempe

Correcting Stochastic Update Bias in Preconditioned Language Model Optimizers

SafetyDGX agent

arXiv:2605.20756v1 Announce Type: new Abstract: Preconditioned optimizers are central to language model training, but their stochastic update rules are usually treated as direct approximations to popu

CRAFT: Conflict-Resolved Aggregation for Federated Training

SafetyDGX agent

arXiv:2605.21317v1 Announce Type: new Abstract: The aggregation of conflicting client updates remains a fundamental bottleneck in federated learning (FL) under heterogeneous data distributions. Naive

Cross-lingual robustness of LLM-brain alignment and its computational roots

SafetyDGX agent

arXiv:2605.21049v1 Announce Type: new Abstract: Large language models (LLMs) reliably predict neural activity during language comprehension and transformer depth has been interpreted as mirroring hier

Cumulative Meta-Learning from Active Learning Queries for Robustness to Spurious Correlations

SafetyDGX agent

arXiv:2605.20771v1 Announce Type: new Abstract: Spurious correlations in real-world datasets cause machine learning models to rely on irrelevant patterns, undermining reliability, generalization, and

Data-Efficient Neural Operator Training via Physics-Based Active Learning

SafetyDGX agent

arXiv:2605.21348v1 Announce Type: new Abstract: Solving partial differential equations with neural operators significantly reduces computational costs but remains bottlenecked by high training data re

Decision-Path Patterns as Tree Reliability Signals: Path-based Adaptive Weighting for Random Forest Classification

SafetyDGX agent

arXiv:2605.20716v1 Announce Type: new Abstract: Random forests aggregate tree votes by simple majority, treating all trees as equally informative. We observe that the topological pattern along each tr

Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor

SafetyDGX agent

arXiv:2605.20402v1 Announce Type: new Abstract: MXFP4 arithmetic can dramatically accelerate reinforcement learning (RL) post-training of large language models (LLMs), yet the quantization error intro

DeCoR: Design and Control Co-Optimization for Urban Streets Using Reinforcement Learning

SafetyDGX agent

arXiv:2605.21311v1 Announce Type: new Abstract: Modern vision systems can detect, track, and forecast urban actors at scale, yet translating perception outputs to urban design remains limited. We intr

Decoupling Communication from Policy: Robust MARL under Bandwidth Constraints

SafetyDGX agent

arXiv:2605.21085v1 Announce Type: cross Abstract: Communication enables coordination in multi-agent reinforcement learning (MARL), but many real-world applications, e.g., search-and-rescue with drone

Deep Attention Reweighting: Post-Hoc Attention-Based Feature Aggregation in CNNs for Disentangling Core and Spurious Features under Spurious Correlations

SafetyDGX agent

arXiv:2605.20732v1 Announce Type: new Abstract: Convolutional Neural Networks (CNNs) often exploit spurious correlations in datasets, learning superficially predictive yet causally irrelevant features

DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards

SafetyDGX agent

arXiv:2605.21467v1 Announce Type: cross Abstract: Reinforcement learning from verifiable rewards (RLVR) has emerged as a central technique for improving the reasoning capabilities of large language mo

Design for Manufacturing: A Manufacturability Knowledge-Integrated Reinforcement Learning Framework for Free-Form Pipe Routing in Aeroengines

SafetyDGX agent

arXiv:2605.20644v1 Announce Type: new Abstract: Design for manufacturing plays a critical role in advanced aeroengine development, where complex components necessitate careful consideration of manufac

Disentangling Bias by Modeling Intra- and Inter-modal Causal Attention for Multimodal Sentiment Analysis

SafetyDGX agent

arXiv:2508.04999v2 Announce Type: replace Abstract: Multimodal sentiment analysis (MSA) aims to understand human emotions by integrating information from multiple modalities, such as text, audio, and

← Previous
1…125126127128129…214
Next →