AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,702 results
22 Apr 2026

Allo{SR}^2: Rectifying One-Step Super-Resolution to Stay Real via Allomorphic Generative Flows

SafetyDGX agent

arXiv:2604.19238v1 Announce Type: new Abstract: Real-world image super-resolution (Real-SR) has been revolutionized by leveraging the powerful generative priors of large-scale diffusion and flow-based

Anthropic’s own internal security blows.

SafetyDGX agent

Anthropic’s own internal security blows. Anthropic said Mythos was too dangerous to release. Then four random guys in a Discord gained access on day one by guessing the URL... This is pretty insane: →

ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System

SafetyDGX agent

arXiv:2604.18789v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is central to aligning Large Language Models (LLMs), yet it introduces a critical vulnerability: an im


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

ARM: Advantage Reward Modeling for Long-Horizon Manipulation

SafetyDGX agent

arXiv:2604.03037v2 Announce Type: replace-cross Abstract: Long-horizon robotic manipulation remains challenging for reinforcement learning (RL) because sparse rewards provide limited guidance for cred

Assessing VLM-Driven Semantic-Affordance Inference for Non-Humanoid Robot Morphologies

SafetyDGX agent

arXiv:2604.19509v1 Announce Type: new Abstract: Vision-language models (VLMs) have demonstrated remarkable capabilities in understanding human-object interactions, but their application to robotic sys

ASVSim (AirSim for Surface Vehicles): A High-Fidelity Simulation Framework for Autonomous Surface Vehicle Research

SafetyDGX agent

arXiv:2506.22174v2 Announce Type: replace-cross Abstract: The transport industry has recently shown significant interest in unmanned surface vehicles (USVs), specifically for port and inland waterway

Attention-based Multi-modal Deep Learning Model of Spatio-temporal Crop Yield Prediction with Satellite, Soil and Climate Data

SafetyDGX agent

arXiv:2604.19217v1 Announce Type: cross Abstract: Crop yield prediction is one of the most important challenge, which is crucial to world food security and policy-making decisions. The conventional fo

Auditing LLMs for Algorithmic Fairness in Casenote-Augmented Tabular Prediction

SafetyDGX agent

arXiv:2604.19204v1 Announce Type: cross Abstract: LLMs are increasingly being considered for prediction tasks in high-stakes social service settings, but their algorithmic fairness properties in this

Australia's eSafety Commissioner issues transparency notices to Roblox, Microsoft's Minecraft, and other online gaming platforms to detail child safety measures (Renju Jose/Reuters)

SafetyDGX agent

Renju Jose / Reuters: Australia's eSafety Commissioner issues transparency notices to Roblox, Microsoft's Minecraft, and other online gaming platforms to detail child safety measures — Australia's int

AutoAWG: Adverse Weather Generation with Adaptive Multi-Controls for Automotive Videos

SafetyDGX agent

arXiv:2604.18993v1 Announce Type: cross Abstract: Perception robustness under adverse weather remains a critical challenge for autonomous driving, with the core bottleneck being the scarcity of real-w

BAPO: Boundary-Aware Policy Optimization for Reliable Agentic Search

SafetyDGX agent

arXiv:2601.11037v2 Announce Type: replace Abstract: RL-based agentic search enables LLMs to solve complex questions via dynamic planning and external search. While this approach significantly enhances

Benchmarking Misuse Mitigation Against Covert Adversaries

SafetyDGX agent

arXiv:2506.06414v2 Announce Type: replace-cross Abstract: Existing language model safety evaluations focus on overt attacks and low-stakes tasks. In reality, an attacker can easily subvert existing sa

Beyond Bellman: High-Order Generator Regression for Continuous-Time Policy Evaluation

SafetyDGX agent

arXiv:2604.18972v1 Announce Type: cross Abstract: We study finite-horizon continuous-time policy evaluation from discrete closed-loop trajectories under time-inhomogeneous dynamics. The target value s

Beyond Linear Probes: Dynamic Safety Monitoring for Language Models

SafetyDGX agent

arXiv:2509.26238v4 Announce Type: replace Abstract: Monitoring large language models' (LLMs) activations is an effective way to detect harmful requests before they lead to unsafe outputs. However, tra

Beyond Marginal Distributions: A Framework to Evaluate the Representativeness of Demographic-Aligned LLMs

SafetyDGX agent

arXiv:2601.15755v3 Announce Type: replace Abstract: Large language models are increasingly used to represent human opinions, values, or beliefs, and their steerability towards these ideals is an activ

Beyond Semantic Similarity: A Component-Wise Evaluation Framework for Medical Question Answering Systems with Health Equity Implications

SafetyDGX agent

arXiv:2604.19281v1 Announce Type: cross Abstract: The use of Large Language Models (LLMs) to support patients in addressing medical questions is becoming increasingly prevalent. However, most of the m

Beyond the Bellman Fixed Point: Geometry and Fast Policy Identification in Value Iteration

SafetyDGX agent

arXiv:2604.17457v2 Announce Type: replace-cross Abstract: Dynamic programming is one of the most fundamental methodologies for solving Markov decision problems. Among its many variants, Q-value iterat

Breaking the Illusion: Consensus-Based Generative Mitigation of Adversarial Illusions in Multi-Modal Embeddings

SafetyDGX agent

arXiv:2511.21893v2 Announce Type: replace Abstract: Multi-modal foundation models align images, text, and other modalities in a shared embedding space but remain vulnerable to adversarial illusions [3

Bridging Semantics and Geometry: A Decoupled LVLM-SAM Framework for Reasoning Segmentation in Optical Remote Sensing

SafetyDGX agent

arXiv:2512.19302v2 Announce Type: replace Abstract: Large Vision--Language Models (LVLMs) hold great promise for advancing optical remote sensing (RS) analysis, yet existing reasoning segmentation fra

CAHAL: Clinically Applicable resolution enHAncement for Low-resolution MRI scans

SafetyDGX agent

arXiv:2604.18781v1 Announce Type: new Abstract: Large-scale automated morphometric analysis of brain MRI is limited by the thick-slice, anisotropic acquisitions prevalent in routine clinical practice.

Capturing Classic Authorial Style in Long-Form Story Generation with GRPO Fine-Tuning

SafetyDGX agent

arXiv:2512.05747v3 Announce Type: replace Abstract: Evaluating and optimising authorial style in long-form story generation remains challenging because style is often assessed with ad hoc prompting an

CentaurTA Studio: A Self-Improving Human-Agent Collaboration System for Thematic Analysis

SafetyDGX agent

arXiv:2604.18589v1 Announce Type: cross Abstract: Thematic analysis is difficult to scale: manual workflows are labor-intensive, while fully automated pipelines often lack controllability and transpar

Chain-of-Thought as a Lens: Evaluating Structured Reasoning Alignment between Human Preferences and Large Language Models

SafetyDGX agent

arXiv:2511.06168v3 Announce Type: replace Abstract: This paper primarily demonstrates a method to quantitatively assess the alignment between multi-step, structured reasoning in large language models

ChatGPT doesn’t know its whisk from its elbow

SafetyDGX agent

Gary Marcus critiques ChatGPT's lack of embodied understanding and spatial reasoning, arguing that the language model struggles with physical concepts that humans intuitively grasp through bodily expe

Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models

SafetyDGX agent

arXiv:2510.26782v3 Announce Type: replace-cross Abstract: A world model is an internal model that simulates how the world evolves. Given past observations and actions, it predicts the future physical

Counting Worlds Branching Time Semantics for post-hoc Bias Mitigation in generative AI

SafetyDGX agent

arXiv:2604.19431v1 Announce Type: cross Abstract: Generative AI systems are known to amplify biases present in their training data. While several inference-time mitigation strategies have been propose

CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers

SafetyDGX agent

arXiv:2604.19632v1 Announce Type: new Abstract: Graphic design images consist of multiple editable layers, such as text, background, and decorative elements, while most generative models produce raste

Curvature-Aware PCA with Geodesic Tangent Space Aggregation for Semi-Supervised Learning

SafetyDGX agent

arXiv:2604.18816v1 Announce Type: cross Abstract: Principal Component Analysis (PCA) is a fundamental tool for representation learning, but its global linear formulation fails to capture the structure

Debiased neural operators for estimating functionals

SafetyDGX agent

arXiv:2604.19296v1 Announce Type: new Abstract: Neural operators are widely used to approximate solution maps of complex physical systems. In many applications, however, the goal is not to recover the

Decomposed Trust: Privacy, Adversarial Robustness, Ethics, and Fairness in Low-Rank LLMs

SafetyDGX agent

arXiv:2511.22099v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have driven major advances across domains, yet their massive size hinders deployment in resource-constrained sett

Denoising, Fast and Slow: Difficulty-Aware Adaptive Sampling for Image Generation

SafetyDGX agent

arXiv:2604.19141v1 Announce Type: new Abstract: Diffusion- and flow-based models usually allocate compute uniformly across space, updating all patches with the same timestep and number of function eva

Developing a Robotic Surgery Training System for Wide Accessibility and Research

SafetyDGX agent

arXiv:2505.20562v2 Announce Type: replace Abstract: Robotic surgery represents a major breakthrough in medical interventions, which has revolutionized surgical procedures. However, the high cost and l

Diagnosable ColBERT: Debugging Late-Interaction Retrieval Models Using a Learned Latent Space as Reference

SafetyDGX agent

arXiv:2604.19566v1 Announce Type: cross Abstract: Reliable biomedical and clinical retrieval requires more than strong ranking performance: it requires a practical way to find systematic model failure

Diamond Maps: Efficient Reward Alignment via Stochastic Flow Maps

SafetyDGX agent

arXiv:2602.05993v2 Announce Type: replace-cross Abstract: Flow and diffusion models produce high-quality samples, but adapting them to user preferences or constraints post-training remains costly and

Diff-SBSR: Learning Multimodal Feature-Enhanced Diffusion Models for Zero-Shot Sketch-Based 3D Shape Retrieval

SafetyDGX agent

arXiv:2604.19135v1 Announce Type: new Abstract: This paper presents the first exploration of text-to-image diffusion models for zero-shot sketch-based 3D shape retrieval (ZS-SBSR). Existing sketch-bas

DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval

SafetyDGX agent

arXiv:2604.19432v1 Announce Type: new Abstract: Vision foundation models have shown great promise for open-set 3D object retrieval (3DOR) through efficient adaptation to multi-view images. Leveraging

Discovering a Shared Logical Subspace: Steering LLM Logical Reasoning via Alignment of Natural-Language and Symbolic Views

SafetyDGX agent

arXiv:2604.19716v1 Announce Type: new Abstract: Large Language Models (LLMs) still struggle with multi-step logical reasoning. Existing approaches either purely refine the reasoning chain in natural l

Do Emotions Influence Moral Judgment in Large Language Models?

SafetyDGX agent

arXiv:2604.19125v1 Announce Type: new Abstract: Large language models have been extensively studied for emotion recognition and moral reasoning as distinct capabilities, yet the extent to which emotio

Documents and sources: insurers including QBE and Beazley are moving to cap cyber policy payouts for losses and regulatory fines tied to AI use and 'LLMjacking' (Lee Harris/Financial Times)

SafetyDGX agent

Lee Harris / Financial Times: Documents and sources: insurers including QBE and Beazley are moving to cap cyber policy payouts for losses and regulatory fines tied to AI use and “LLMjacking” — Beazley

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling

SafetyDGX agent

arXiv:2604.19544v1 Announce Type: new Abstract: Multimodal reward models (MRMs) play a crucial role in aligning Multimodal Large Language Models (MLLMs) with human preferences. Training a good MRM req

Dual Triangle Attention: Effective Bidirectional Attention Without Positional Embeddings

SafetyDGX agent

arXiv:2604.18603v1 Announce Type: cross Abstract: Bidirectional transformers are the foundation of many sequence modeling tasks across natural, biological, and chemical language domains, but they are

Efforts to revive chip manufacturing in Pennsylvania have been left in limbo by President Trump's sudden upending of US semiconductor policy over the past year (Michael Acton/Financial Times)

SafetyDGX agent

Michael Acton / Financial Times: Efforts to revive chip manufacturing in Pennsylvania have been left in limbo by President Trump's sudden upending of US semiconductor policy over the past year — High-

Enhancing Construction Worker Safety in Extreme Heat: A Machine Learning Approach Utilizing Wearable Technology for Predictive Health Analytics

SafetyDGX agent

arXiv:2604.19559v1 Announce Type: new Abstract: Construction workers are highly vulnerable to heat stress, yet tools that translate real-time physiological data into actionable safety intelligence rem

Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient Transformers

SafetyDGX agent

arXiv:2510.18358v2 Announce Type: replace-cross Abstract: Uncertainty quantification (UQ) is essential for deploying deep neural networks in safety-critical settings. Although methods like Deep Ensemb

Evaluating LLM-Driven Summarisation of Parliamentary Debates with Computational Argumentation

SafetyDGX agent

arXiv:2604.19331v1 Announce Type: new Abstract: Understanding how policy is debated and justified in parliament is a fundamental aspect of the democratic process. However, the volume and complexity of

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training

SafetyDGX agent

arXiv:2604.19485v1 Announce Type: cross Abstract: Reinforcement learning (RL) for LLM post-training faces a fundamental design choice: whether to use a learned critic as a baseline for policy optimiza

ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors

SafetyDGX agent

arXiv:2603.15956v2 Announce Type: replace-cross Abstract: Learning generalizable and robust behavior cloning policies requires large volumes of high-quality robotics data. While human demonstrations (

Failure Modes in Multi-Hop QA: The Weakest Link Effect and the Recognition Bottleneck

SafetyDGX agent

arXiv:2601.12499v2 Announce Type: replace Abstract: Despite scaling to massive context windows, Large Language Models (LLMs) struggle with multi-hop reasoning due to inherent position bias, which caus

Fairness Audits of Institutional Risk Models in Deployed ML Pipelines

SafetyDGX agent

arXiv:2604.19468v1 Announce Type: cross Abstract: Fairness audits of institutional risk models are critical for understanding how deployed machine learning pipelines allocate resources. Drawing on mul

FairTree: Subgroup Fairness Auditing of Machine Learning Models with Bias-Variance Decomposition

SafetyDGX agent

arXiv:2604.19357v1 Announce Type: new Abstract: The evaluation of machine learning models typically relies mainly on performance metrics based on loss functions, which risk to overlook changes in perf

Fascinating how AI is getting better at diagrams like these (at least for ones that you could easily find on web search) but still making so…

SafetyDGX agent

Fascinating how AI is getting better at diagrams like these (at least for ones that you could easily find on web search) but still making some pretty wacky errors — like confusing where the rear brake

FASE : A Fairness-Aware Spatiotemporal Event Graph Framework for Predictive Policing

SafetyDGX agent

arXiv:2604.18644v1 Announce Type: cross Abstract: Predictive policing systems that allocate patrol resources based solely on predicted crime risk can unintentionally amplify racial disparities through

FASTER: Value-Guided Sampling for Fast RL

SafetyDGX agent

arXiv:2604.19730v1 Announce Type: cross Abstract: Some of the most performant reinforcement learning algorithms today can be prohibitively expensive as they use test-time scaling methods such as sampl

FB-NLL: A Feature-Based Approach to Tackle Noisy Labels in Personalized Federated Learning

SafetyDGX agent

arXiv:2604.19729v1 Announce Type: new Abstract: Personalized Federated Learning (PFL) aims to learn multiple task-specific models rather than a single global model across heterogeneous data distributi

Filing: Tron founder Justin Sun sues the Trump family's World Liberty Financial, alleging it unfairly locked up his WLFI holdings and threatened and defamed him (CoinDesk)

SafetyDGX agent

CoinDesk: Filing: Tron founder Justin Sun sues the Trump family's World Liberty Financial, alleging it unfairly locked up his WLFI holdings and threatened and defamed him — World Liberty unfairly froz

Fitted Q Evaluation Without Bellman Completeness via Stationary Weighting

SafetyDGX agent

arXiv:2512.23805v2 Announce Type: replace-cross Abstract: Fitted Q-evaluation (FQE) is a foundational method for off-policy evaluation in reinforcement learning, but existing theory typically relies o

Framelet-Based Blind Image Restoration with Minimax Concave Regularization

SafetyDGX agent

arXiv:2604.19314v1 Announce Type: new Abstract: Recovering corrupted images is one of the most challenging problems in image processing. Among various restoration tasks, blind image deblurring has bee

From Particles to Perils: SVGD-Based Hazardous Scenario Generation for Autonomous Driving Systems Testing

SafetyDGX agent

arXiv:2604.18918v1 Announce Type: cross Abstract: Simulation-based testing of autonomous driving systems (ADS) must uncover realistic and diverse failures in dense, heterogeneous traffic. However, exi

Gives new meaning to “Rear Brake Lever”!

SafetyDGX agent

This post likely references a humorous or unexpected use case involving a rear brake lever, possibly demonstrating an unintended design flaw, unconventional application, or double meaning related to b

God these people are annoying. Obnoxious comment and the guy can’t be bothered to notice the front tire that is labeled as a fork 🙄 Or to n…

SafetyDGX agent

God these people are annoying. Obnoxious comment and the guy can’t be bothered to notice the front tire that is labeled as a fork 🙄 Or to notice the front brake that’s lost its cable and is hovering u

← Previous
1…184185186187188…212
Next →