AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,813 results
4 Jun 2026

Achieving Rotation-Invariant Convolution via Non-Learnable Orientation Alignment Operators

SafetyDGX agent

arXiv:2404.11309v2 Announce Type: replace Abstract: Achieving rotational invariance in deep neural networks without data augmentation is a research hotspot. Intrinsic invariance enables features to ca

Activation Steering of Video Generation Models via Reduced-Order Linear Optimal Control

SafetyDGX agent

arXiv:2606.04775v1 Announce Type: cross Abstract: Text-to-video (T2V) models trained on large-scale web data can generate undesired content, motivating interventions that reduce harmful outputs withou

Adaptive Calibration for Fair and Performant Facial Recognition

SafetyDGX agent

arXiv:2606.04469v1 Announce Type: cross Abstract: We introduce Adaptive Calibration (AC), a novel calibration strategy for facial recognition that maps cosine similarity between normalized embeddings


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Adaptive Information Control for Search-Augmented LLM Reasoning

SafetyDGX agent

arXiv:2602.01672v2 Announce Type: replace Abstract: Search-augmented reasoning agents interleave multi-step reasoning with external retrieval, but uncontrolled retrieval can introduce redundant eviden

AI enthusiasts are in a race against time, AI skeptics are in a race against entropy

SafetyDGX agent

AI enthusiasts are in a race against time, AI skeptics are in a race against entropy Charity Majors neatly captures the dynamic between AI enthusiasts and AI skeptics, both of whom are trying to build

Anycast Performance in Context

SafetyDGX agent

arXiv:2606.04298v1 Announce Type: cross Abstract: IP anycast lets a service advertise one address from many physical sites, leaving BGP to map each client to a site. It is central to the DNS root serv

Be Fair! Can Machine Learning Engineering Agents Adhere to Fairness Constraints?

SafetyDGX agent

arXiv:2606.04971v1 Announce Type: new Abstract: Machine learning engineering (MLE) agents promise to automate end-to-end ML pipeline development from raw data and natural language instructions, potent

Beyond Static Priors: Dynamic Neural Guidance for Large-Scale Ant Colony Optimization

SafetyDGX agent

arXiv:2606.04039v1 Announce Type: cross Abstract: Neural-guided Ant Colony Optimization (ACO) suffers from a fundamental training-inference misalignment: policies are typically trained to generate sta

Beyond Symmetric Alignment: Spectral Diagnostics of Modality Imbalance in Vision-Language Models in the Medical Domain

SafetyDGX agent

arXiv:2606.04613v1 Announce Type: new Abstract: Vision-Language Models (VLMs) struggle when applied to medical image-text data, yet the tools available to diagnose this failure remain limited. Existin

BiasGRPO: Stabilizing Bias Mitigation in High-Variance Reward Landscapes via Group-Relative Policy Optimization

SafetyDGX agent

arXiv:2606.04807v1 Announce Type: new Abstract: Mitigating social bias in Large Language Models (LLMs) presents a distinct alignment challenge: unlike verifiable tasks, bias lacks a single ground trut

Blessing from Human-AI Interaction: Super Reinforcement Learning in Confounded Environments

SafetyDGX agent

arXiv:2209.15448v3 Announce Type: replace Abstract: As AI becomes more prevalent throughout society, effective methods of integrating humans and AI systems that leverage their respective strengths and

Causal Multi-fidelity Surrogate Forward and Inverse Models for ICF Implosions

SafetyDGX agent

arXiv:2509.05510v3 Announce Type: replace-cross Abstract: Continued progress in inertial confinement fusion (ICF) requires solving inverse problems relating experimental observations to simulation inp

Certified Neural Approximations of Nonlinear Dynamics

SafetyDGX agent

arXiv:2505.15497v3 Announce Type: replace Abstract: Neural networks hold great potential to act as approximate models of nonlinear dynamical systems, with the resulting neural approximations enabling

Channel-Oriented Design for EEG-to-Music Reconstruction

SafetyDGX agent

arXiv:2606.04040v1 Announce Type: cross Abstract: Brain-computer interfaces aim to decode naturalistic stimuli from neural signals, yet most progress to date has focused on vision and language. In thi

Confidence Before Answering: A Paradigm Shift for Efficient LLM Uncertainty Estimation

SafetyDGX agent

arXiv:2603.05881v2 Announce Type: replace Abstract: Reliable deployment of large language models (LLMs) requires accurate uncertainty estimation. Existing methods are predominantly answer-first, produ

COP-Q: Safety-First Reinforcement Learning for Robot Control via Cholesky-Ordered Projection

SafetyDGX agent

arXiv:2606.04749v1 Announce Type: cross Abstract: Safe robot control requires maximizing return while satisfying safety constraints. In off-policy safe reinforcement learning, reward and safety Q-valu

CoRe-MoE: Contrastive Reweighted Mixture of Experts for Multi-Terrain Humanoid Locomotion with Gait Adaptation

SafetyDGX agent

arXiv:2606.04718v1 Announce Type: cross Abstract: Humans primarily rely on walking and running to traverse complex terrains, without resorting to unnecessarily complex motion patterns. Similarly, huma

Covert Influence Between Language Models

SafetyDGX agent

arXiv:2606.04071v1 Announce Type: cross Abstract: As language models increasingly consume one another's outputs, covert influence -- a phenomenon where a sender's payload (the behavioral disposition i

Crafting Your Evolving Dreams: Concept-Incremental Versatile Customization

SafetyDGX agent

arXiv:2606.04797v1 Announce Type: new Abstract: Custom diffusion models (CDMs) have garnered significant interest owing to their remarkable capacity for generating personalized concepts. However, the

Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks

SafetyDGX agent

arXiv:2601.22396v2 Announce Type: replace-cross Abstract: Despite the growing utility of Large Language Models (LLMs) for simulating human behavior, the extent to which these synthetic personas accura

Customizing the Inductive Biases of Softmax Attention using Structured Matrices

SafetyDGX agent

arXiv:2509.07963v2 Announce Type: replace Abstract: The core component of attention is the scoring function, which transforms the inputs into low-dimensional queries and keys and takes the dot product

death of tokenmaxxing = potentially a very serious issue for all three big IPOs.

SafetyDGX agent

death of tokenmaxxing = potentially a very serious issue for all three big IPOs. 🚨 Sam Altman warns OpenAi and Anthropic are experiencing severe pullback on Ai spending as companies put significant re

DiffAero: A GPU-Accelerated Differentiable Simulation Framework for Efficient Quadrotor Policy Learning

SafetyDGX agent

arXiv:2509.10247v1 Announce Type: cross Abstract: This letter introduces DiffAero, a lightweight, GPU-accelerated, and fully differentiable simulation framework designed for efficient quadrotor contro

Distribution-Free Risk-Aware Planning and Control Under Uncertainty Using Conformal Spectral Risk Control

SafetyDGX agent

arXiv:2606.04185v1 Announce Type: new Abstract: Safe navigation in dynamic and uncertain environments often relies on accurate estimation of, or assumptions about, the true underlying uncertainty. How

DPM++: Dynamic Masked Metric Learning for Occluded Person Re-identification

SafetyDGX agent

arXiv:2605.06637v2 Announce Type: replace Abstract: Although person re-identification has made impressive progress, occlusion caused by obstacles remains an unsettled issue in real applications. The d

DuDi: Dual-Signal Distillation with Cross-Lingual Verbalizer

SafetyDGX agent

arXiv:2606.04694v1 Announce Type: new Abstract: Small language models (SLMs) are efficient and scalable, but their multilingual capabilities degrade severely at sub-billion scales, especially for Sout

DVGT: Driving Visual Geometry Transformer

SafetyDGX agent

arXiv:2512.16919v2 Announce Type: replace-cross Abstract: Perceiving and reconstructing 3D scene geometry from visual inputs is crucial for autonomous driving. However, there still lacks a driving-tar

Dynamic Multi-Pair Trading Strategy in Cryptocurrency Markets with Deep Reinforcement Learning

SafetyDGX agent

arXiv:2606.04574v1 Announce Type: new Abstract: This study aims to determine whether the application of Deep Reinforcement Learning (DRL) as a specialized execution overlay can enhance pair trading in

Dynamic Policy Learning for Legged Robot with Simplified Model Pretraining and Model-Homotopy-Inspired Transfer

SafetyDGX agent

arXiv:2512.24698v2 Announce Type: replace Abstract: Generating dynamic motions for legged robots remains a challenging problem. While reinforcement learning has achieved notable success in various leg

Edge of Stability Selectively Shapes Learning Across the Data Distribution

SafetyDGX agent

arXiv:2606.04212v1 Announce Type: new Abstract: Existing analyses of the edge of stability (EoS) treat it as a global property of optimization. We show that it is also selective: the stability constra

Efficient Adversarial Attacks on High-dimensional Offline Bandits

SafetyDGX agent

arXiv:2602.01658v2 Announce Type: replace-cross Abstract: Bandit algorithms have recently emerged as a powerful tool for evaluating machine learning models, including generative image models and large

Elon Musk petitioned the FTC in May to end its 2022 order restricting Twitter's data use, claiming Twitter no longer exists as X merged with xAI and then SpaceX (Ashley Belanger/Ars Technica)

SafetyDGX agent

Ashley Belanger / Ars Technica: Elon Musk petitioned the FTC in May to end its 2022 order restricting Twitter's data use, claiming Twitter no longer exists as X merged with xAI and then SpaceX — Criti

Enhancing the MADDPG Algorithm for Multi-Agent Learning via Action Inference and Importance Sampling

SafetyDGX agent

arXiv:2606.05021v1 Announce Type: new Abstract: We investigate multi-agent deep reinforcement learning and propose two enhancements to the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algor

Expert-Aware Refusal Steering

SafetyDGX agent

arXiv:2606.04160v1 Announce Type: new Abstract: Safety alignment in instruction-tuned large language models (LLMs) depends on a model's ability to reliably refuse to respond to harmful or disallowed r

Explainably Safe Reinforcement Learning

SafetyDGX agent

arXiv:2606.04634v1 Announce Type: new Abstract: Trust in a decision-making system requires both safety guarantees and the ability to interpret and understand its behavior. This is particularly importa

Extending Fair Null-Space Projections for Continuous Attributes to Kernel Methods

SafetyDGX agent

arXiv:2511.03304v2 Announce Type: replace-cross Abstract: With the on-going integration of machine learning systems into the everyday social life of millions the notion of fairness becomes an ever inc

Feels like a good time to resurface this one Mine and @jaswu_'s basic point: cheaper AI complicates the narrative for OpenAI and Anthropic w…

SafetyDGX agent

Feels like a good time to resurface this one Mine and @jaswu_'s basic point: cheaper AI complicates the narrative for OpenAI and Anthropic when they eventually try to go public. Could also ripple acro

Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning

SafetyDGX agent

arXiv:2603.07445v2 Announce Type: replace Abstract: Large language models (LLMs) often require fine-tuning (FT) to perform well on downstream tasks, but FT can induce safety-alignment drift even when

FLAGG: Flexible Autoregressive Graph Generation

SafetyDGX agent

arXiv:2606.05067v1 Announce Type: new Abstract: The Deep Graph Generation's panorama spans two extremes: one-shot and sequential models. The former generates nodes and edges jointly, while the latter

Fog of Love: Engineering Virtuous Agent Behavior with Affinity-based Reinforcement Learning in a Game Environment

SafetyDGX agent

arXiv:2606.04750v1 Announce Type: new Abstract: Instilling virtuous behavior in artificial intelligence has seen increasing interest. One of the techniques proposed is known as affinity-based reinforc

Formal Semantics for Agentic Tool Protocols: A Process Calculus Approach

SafetyDGX agent

arXiv:2603.24747v2 Announce Type: replace Abstract: The emergence of large language model agents capable of invoking external tools has created urgent need for formal verification of agent protocols.

From Agent Traces to Trust: Evidence Tracing and Execution Provenance in LLM Agents

SafetyDGX agent

arXiv:2606.04990v1 Announce Type: cross Abstract: Large language model (LLM)-based agents increasingly solve complex tasks by interacting with external tools, retrieval systems, memory modules, enviro

GARL: Game-Theoretic Reinforcement Learning for Multi-Agent Strategic Prioritisation

SafetyDGX agent

arXiv:2606.05002v1 Announce Type: new Abstract: LLM-based multi-agent systems are increasingly used for strategic decision-making tasks. In such settings, performance depends not only on individual mo

Generalizable Multi-Task Learning for Wireless Networks Using Prompt Decision Transformers

SafetyDGX agent

arXiv:2606.04328v1 Announce Type: cross Abstract: Future wireless networks demand rapid adaptation to highly heterogeneous environments and dynamic task configurations, necessitating a shift from conv

Generalization of World Models under Environmental Variability for Vision-based Quadrotor Navigation

SafetyDGX agent

arXiv:2606.05015v1 Announce Type: new Abstract: World models, learned generative models that predict how an environment evolves, have become a promising tool for sample-efficient robot learning. Yet h

Geometry-Aware Distillation for Prompt Tuning Biomedical Vision-Language Models

SafetyDGX agent

arXiv:2606.04922v1 Announce Type: cross Abstract: Current prompt-based and adapter-based tuning of vision-language models (VLMs) is attractive for medical imaging, where clinical data sensitivity favo

Geospatial Foundation Models to Enable Progress on Sustainable Development Goals

SafetyDGX agent

arXiv:2505.24528v3 Announce Type: replace Abstract: Foundation Models (FMs) are large-scale, pre-trained artificial intelligence (AI) systems that have revolutionized natural language processing and c

Global Sketch-Based Watermarking for Diffusion Language Models

SafetyDGX agent

arXiv:2606.04486v1 Announce Type: cross Abstract: Watermarking methods for language models have been studied extensively in the autoregressive setting, where tokens are generated sequentially. These w

Good Reasoning Makes Good Demonstrations: Implicit Reasoning Quality Supervision via In-Context Reinforcement Learning

SafetyDGX agent

arXiv:2603.09803v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) improves reasoning in large language models but treats all correct solutions equally, potentia

GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewards

SafetyDGX agent

arXiv:2606.04889v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (e.g. GRPO) is now a common way to improve mathematical reasoning in Large Language Models (LLMs). Howeve

HapTile: A Haptic-Informed Vision-Tactile-Language-Action Dataset for Contact-Rich Imitation Learning

SafetyDGX agent

arXiv:2606.04825v1 Announce Type: new Abstract: Despite the importance of tactile sensing for reliable manipulation, most existing Vision-Language-Action (VLA) datasets remain vision-only, and those t

How does Elon get off on just lying straight up about the mainstream media? Can someone like BBC sue him for defamation? He claimed “Legacy …

SafetyDGX agent

How does Elon get off on just lying straight up about the mainstream media? Can someone like BBC sue him for defamation? He claimed “Legacy mainstream media, same ones who wrote about George Floyd mil

Hybrid Adversarial Defence for Natural Language Understanding Tasks

SafetyDGX agent

arXiv:2606.04612v1 Announce Type: new Abstract: Large Language Models (LLMs) are vulnerable both to hallucination and adversarial manipulation. Although these problems are closely related, existing de

I think @Levie is overstating the positive case for employment in the (near term) AI era but that most people have overstated the negative c…

SafetyDGX agent

I think @Levie is overstating the positive case for employment in the (near term) AI era but that most people have overstated the negative case, and that the truth is somewhere in between. Which is to

If we can’t trust Goldman on which IPO to buy… who can we trust? WeWork. Goldman Sachs marketed the real estate firm at an inflated 96 billi…

SafetyDGX agent

If we can’t trust Goldman on which IPO to buy… who can we trust? WeWork. Goldman Sachs marketed the real estate firm at an inflated 96 billion dollar tech valuation. The IPO was pulled after the prosp

If you could buy into exactly one of the three mega IPOs, which would it be?

SafetyDGX agent

Gary Marcus poses a hypothetical investment question asking followers to choose among three major IPOs, likely seeking comparative analysis of high-profile public offerings. The post invites discussio

if you had 100k to invest in OpenAI and/or Anthropic IPOs which would you go for?

SafetyDGX agent

Gary Marcus discusses investment strategy between potential OpenAI and Anthropic IPOs, likely weighing factors such as the companies' technological capabilities, market positioning, business models, a

If you – or your retirement funds - get taken for a ride on SpaceX blame hype guys like this, who don’t even mention that Goldman is the lea…

SafetyDGX agent

If you – or your retirement funds - get taken for a ride on SpaceX blame hype guys like this, who don’t even mention that Goldman is the lead left on the deal. GOLDMAN SEES SPACEX AI REVENUE EXPLODING

Imbuing Large Language Models with Bidirectional Logic for Robust Chain Repair

SafetyDGX agent

arXiv:2606.05030v1 Announce Type: new Abstract: Autoregressive chain-of-thought (CoT) reasoning in large language models (LLMs) is fundamentally forward-directed: each step conditions only on prior to

In-Context Graphical Inference

SafetyDGX agent

arXiv:2606.05042v1 Announce Type: cross Abstract: Marginal inference in discrete graphical models forces a choice between exactness and scalability: exact algorithms are intractable for high-treewidth

← Previous
1…8990919293…214
Next →