AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,356 results
28 Apr 2026

UniAda: Universal Adaptive Multi-objective Adversarial Attack for End-to-End Autonomous Driving Systems

SafetyDGX agent

arXiv:2604.23362v1 Announce Type: cross Abstract: Adversarial attacks play a pivotal role in testing and improving the reliability of deep learning (DL) systems. Existing literature has demonstrated t

Verifying Quantized GNNs With Readout Is Decidable But Highly Intractable

SafetyDGX agent

arXiv:2510.08045v2 Announce Type: replace-cross Abstract: We introduce a logical language for reasoning about quantized aggregate-combine graph neural networks with global readout (ACR-GNNs). We provi

When Policies Cannot Be Retrained: A Unified Closed-Form View of Post-Training Steering in Offline Reinforcement Learning

SafetyDGX agent

arXiv:2604.22873v1 Announce Type: cross Abstract: Offline reinforcement learning (RL) can learn effective policies from fixed datasets, but deployment objectives may change after training, and in many

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
27 Apr 2026

A Co-Evolutionary Theory of Human-AI Coexistence: Mutualism, Governance, and Dynamics in Complex Societies

SafetyDGX agent

arXiv:2604.22227v1 Announce Type: cross Abstract: Classical robot ethics is often framed around obedience, most famously through Asimov's laws. This framing is too narrow for contemporary AI systems,

How Large Language Models Balance Internal Knowledge with User and Document Assertions

SafetyDGX agent

arXiv:2604.22193v1 Announce Type: new Abstract: Large language models (LLMs) often need to balance their internal parametric knowledge with external information, such as user beliefs and content from

Is it me or X starting to look like a vibe coded mess? Polls are broken. Accounts are getting hacked. My DMs are full of phishing scams. Bas…

SafetyDGX agent

Gary Marcus discusses technical and security issues affecting the X platform, including malfunctioning polls, compromised accounts, and increased phishing scams in direct messages. The post appears to

Learning-augmented robotic automation for real-world manufacturing

SafetyDGX agent

arXiv:2604.22235v1 Announce Type: cross Abstract: Industrial robots are widely used in manufacturing, yet most manipulation still depends on fixed waypoint scripts that are brittle to environmental ch

On the Properties of Feature Attribution for Supervised Contrastive Learning

SafetyDGX agent

arXiv:2604.22540v1 Announce Type: cross Abstract: Most Neural Networks (NNs) for classification are trained using Cross-Entropy as a loss function. This approach requires the model to have an explicit

Reliable Self-Harm Risk Screening via Adaptive Multi-Agent LLM Systems

SafetyDGX agent

arXiv:2604.22154v1 Announce Type: cross Abstract: Emerging AI systems in behavioral health and psychiatry use multi-step or multi-agent LLM pipelines for tasks like assessing self-harm risk and screen

Rethinking XAI Evaluation: A Human-Centered Audit of Shapley Benchmarks in High-Stakes Settings

SafetyDGX agent

arXiv:2604.22662v1 Announce Type: cross Abstract: Shapley values are a cornerstone of explainable AI, yet their proliferation into competing formulations has created a fragmented landscape with little

Scam Altman didn’t tell the OpenAI board that he OWNED the OpenAI Startup Fund. Altman lied in congressional testimony that he didn’t have f…

SafetyDGX agent

Scam Altman didn’t tell the OpenAI board that he OWNED the OpenAI Startup Fund. Altman lied in congressional testimony that he didn’t have financial gain from OpenAI. Ex-board member of OpenAI calls S

The only AGI that Sam Altman is after is Adjusted Gross Income.

SafetyDGX agent

Gary Marcus made a critical commentary on Sam Altman's priorities, using a pun on 'AGI' (Artificial General Intelligence) to suggest that Altman's actual focus is on 'Adjusted Gross Income' rather tha

The people building the most powerful technology in history cannot tell you what is happening inside their own systems. This is insane. We s…

SafetyDGX agent

The people building the most powerful technology in history cannot tell you what is happening inside their own systems. This is insane. We started the Torchbearer Community for exactly this reason. Re

26 Apr 2026

Coders and software engineers ONLY:

SafetyDGX agent

Gary Marcus, a prominent AI researcher and critic, posted a message on X (formerly Twitter) addressing software engineers and coders, likely discussing technical aspects of AI development, programming

Existential risk mongers are a small, very vocal cult with a lot of very clever online astroturfing skills. Politicians never waste a good f…

SafetyDGX agent

Existential risk mongers are a small, very vocal cult with a lot of very clever online astroturfing skills. Politicians never waste a good fake crisis, which is why they're perfect for Bernie to try t

What if I told you there was a technology where 1.5 million people would die every year and injure 50 million would you sign up for that tec…

SafetyDGX agent

What if I told you there was a technology where 1.5 million people would die every year and injure 50 million would you sign up for that tech? Hell no, right? But the answer is actually 'hell yes' bec

24 Apr 2026

A Survey of Legged Robotics in Non-Inertial Environments: Past, Present, and Future

SafetyDGX agent

arXiv:2604.20990v1 Announce Type: new Abstract: Legged robots have demonstrated remarkable agility on rigid, stationary ground, but their locomotion reliability remains limited in non-inertial environ

Automated Annotation of Shearographic Measurements Enabling Weakly Supervised Defect Detection

SafetyDGX agent

arXiv:2512.06171v2 Announce Type: replace Abstract: Shearography is an interferometric technique sensitive to surface displacement gradients, providing high sensitivity for detecting subsurface defect

CARE: Counselor-Aligned Response Engine for Online Mental-Health Support

SafetyDGX agent

arXiv:2604.21352v1 Announce Type: new Abstract: Mental health challenges are increasing worldwide, straining emotional support services and leading to counselor overload. This can result in delayed re

Dialect vs Demographics: Quantifying LLM Bias from Implicit Linguistic Signals vs. Explicit User Profiles

Model ReleasesDGX agent

arXiv:2604.21152v1 Announce Type: cross Abstract: As state-of-the-art Large Language Models (LLMs) have become ubiquitous, ensuring equitable performance across diverse demographics is critical. Howev

Escaping the Agreement Trap: Defensibility Signals for Evaluating Rule-Governed AI

SafetyDGX agent

arXiv:2604.20972v1 Announce Type: new Abstract: Content moderation systems are typically evaluated by measuring agreement with human labels. In rule-governed environments this assumption fails: multip

Fairness Evaluation and Inference Level Mitigation in LLMs

SafetyDGX agent

arXiv:2510.18914v4 Announce Type: replace-cross Abstract: Large language models often display undesirable behaviors embedded in their internal representations, undermining fairness, inconsistency drif

In 1982, Blade Runner was speculation. In 2026, it's the conversation. Memory, personhood, alignment, what we owe the minds we build. Watchi…

SafetyDGX agent

In 1982, Blade Runner was speculation. In 2026, it's the conversation. Memory, personhood, alignment, what we owe the minds we build. Watching it Thursday 28 May, Kensington Central Library. Come thin

Inferring High-Level Events from Timestamped Data: Complexity and Medical Applications

SafetyDGX agent

arXiv:2604.21793v1 Announce Type: new Abstract: In this paper, we develop a novel logic-based approach to detecting high-level temporally extended events from timestamped data and background knowledge

Language-Conditioned Safe Trajectory Generation for Spacecraft Rendezvous

SafetyDGX agent

arXiv:2512.09111v3 Announce Type: replace-cross Abstract: Reliable real-time trajectory generation is essential for future autonomous spacecraft. While recent progress in nonconvex guidance and contro

Probabilistic Verification of Neural Networks via Efficient Probabilistic Hull Generation

SafetyDGX agent

arXiv:2604.21556v1 Announce Type: new Abstract: The problem of probabilistic verification of a neural network investigates the probability of satisfying the safe constraints in the output space when t

Task-specific Subnetwork Discovery in Reinforcement Learning for Autonomous Underwater Navigation

SafetyDGX agent

arXiv:2604.21640v1 Announce Type: cross Abstract: Autonomous underwater vehicles are required to perform multiple tasks adaptively and in an explainable manner under dynamic, uncertain conditions and

TraceScope: Interactive URL Triage via Decoupled Checklist Adjudication

SafetyDGX agent

arXiv:2604.21840v1 Announce Type: cross Abstract: Modern phishing campaigns increasingly evade snapshot-based URL classifiers using interaction gates (e.g., checkbox/slider challenges), delayed conten

Tumor-anchored deep feature random forests for out-of-distribution detection in lung cancer segmentation

SafetyDGX agent

arXiv:2512.08216v3 Announce Type: replace-cross Abstract: Accurate segmentation of lung tumors from 3D computed tomography (CT) scans is essential for automated treatment planning and response assessm

Unbiased Prevalence Estimation with Multicalibrated LLMs

SafetyDGX agent

arXiv:2604.21549v1 Announce Type: new Abstract: Estimating the prevalence of a category in a population using imperfect measurement devices (diagnostic tests, classifiers, or large language models) is

Why Do Language Model Agents Whistleblow?

SafetyDGX agent

arXiv:2511.17085v3 Announce Type: replace-cross Abstract: The deployment of Large Language Models (LLMs) as tool-using agents causes their alignment training to manifest in new ways. Recent work finds

23 Apr 2026

Environmental Understanding Vision-Language Model for Embodied Agent

SafetyDGX agent

arXiv:2604.19839v1 Announce Type: cross Abstract: Vision-language models (VLMs) have shown strong perception and reasoning abilities for instruction-following embodied agents. However, despite these a

Explainable AML Triage with LLMs: Evidence Retrieval and Counterfactual Checks

SafetyDGX agent

arXiv:2604.19755v1 Announce Type: new Abstract: Anti-money laundering (AML) transaction monitoring generates large volumes of alerts that must be rapidly triaged by investigators under strict audit an

From Fuzzy to Formal: Scaling Hospital Quality Improvement with AI

SafetyDGX agent

arXiv:2604.20055v1 Announce Type: new Abstract: Hospital Quality Improvement (QI) plays a critical role in optimizing healthcare delivery by translating high-level hospital goals into actionable solut

From Noise to Signal to Selbstzweck: Reframing Human Label Variation in the Era of Post-training in NLP

SafetyDGX agent

arXiv:2510.12817v3 Announce Type: replace-cross Abstract: Human Label Variation (HLV) refers to legitimate disagreement in annotation that reflects the diversity of human perspectives rather than mere

FSFM: A Biologically-Inspired Framework for Selective Forgetting of Agent Memory

SafetyDGX agent

arXiv:2604.20300v1 Announce Type: new Abstract: For LLM agents, memory management critically impacts efficiency, quality, and security. While much research focuses on retention, selective forgetting--

ICLR 2026: 12 papers on making AI systems reliable, efficient, and secure

SafetyDGX agent

A 7B agent that beats GPT-4o. Lossless weight compression that speeds up inference by 177%. An arena where 23 teams battled across 103,000 adversarial rounds. This year at ICLR, Lambda is presenting t

Large language models perceive cities through a culturally uneven baseline

SafetyDGX agent

arXiv:2604.20048v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to describe, evaluate and interpret places, yet it remains unclear whether they do so from a cultural

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills

SafetyDGX agent

arXiv:2604.20441v1 Announce Type: new Abstract: Background: Agent skills are increasingly deployed as modular, reusable capability units in AI agent systems. Medical research agent skills require safe

NeuroSymActive: Differentiable Neural-Symbolic Reasoning with Active Exploration for Knowledge Graph Question Answering

SafetyDGX agent

arXiv:2602.15353v2 Announce Type: replace-cross Abstract: Large pretrained language models and neural reasoning systems have advanced many natural language tasks, yet they remain challenged by knowled

Semantic Prompting: Agentic Incremental Narrative Refinement through Spatial Semantic Interaction

SafetyDGX agent

arXiv:2604.19971v1 Announce Type: cross Abstract: Interactive spatial layouts empower users to synthesize information and organize findings for sensemaking. While Large Language Models (LLMs) can auto

22 Apr 2026

A Functionality-Grounded Benchmark for Evaluating Web Agents in E-commerce Domains

Model ReleasesDGX agent

arXiv:2508.15832v2 Announce Type: replace-cross Abstract: Web agents have shown great promise in performing many tasks on ecommerce website. To assess their capabilities, several benchmarks have been

Assessing VLM-Driven Semantic-Affordance Inference for Non-Humanoid Robot Morphologies

SafetyDGX agent

arXiv:2604.19509v1 Announce Type: new Abstract: Vision-language models (VLMs) have demonstrated remarkable capabilities in understanding human-object interactions, but their application to robotic sys

ASVSim (AirSim for Surface Vehicles): A High-Fidelity Simulation Framework for Autonomous Surface Vehicle Research

SafetyDGX agent

arXiv:2506.22174v2 Announce Type: replace-cross Abstract: The transport industry has recently shown significant interest in unmanned surface vehicles (USVs), specifically for port and inland waterway

CAHAL: Clinically Applicable resolution enHAncement for Low-resolution MRI scans

SafetyDGX agent

arXiv:2604.18781v1 Announce Type: new Abstract: Large-scale automated morphometric analysis of brain MRI is limited by the thick-slice, anisotropic acquisitions prevalent in routine clinical practice.

Decomposed Trust: Privacy, Adversarial Robustness, Ethics, and Fairness in Low-Rank LLMs

SafetyDGX agent

arXiv:2511.22099v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have driven major advances across domains, yet their massive size hinders deployment in resource-constrained sett

Developing a Robotic Surgery Training System for Wide Accessibility and Research

SafetyDGX agent

arXiv:2505.20562v2 Announce Type: replace Abstract: Robotic surgery represents a major breakthrough in medical interventions, which has revolutionized surgical procedures. However, the high cost and l

Documents and sources: insurers including QBE and Beazley are moving to cap cyber policy payouts for losses and regulatory fines tied to AI use and 'LLMjacking' (Lee Harris/Financial Times)

SafetyDGX agent

Lee Harris / Financial Times: Documents and sources: insurers including QBE and Beazley are moving to cap cyber policy payouts for losses and regulatory fines tied to AI use and “LLMjacking” — Beazley

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling

SafetyDGX agent

arXiv:2604.19544v1 Announce Type: new Abstract: Multimodal reward models (MRMs) play a crucial role in aligning Multimodal Large Language Models (MLLMs) with human preferences. Training a good MRM req

Hierarchically Robust Zero-shot Vision-language Models

SafetyDGX agent

arXiv:2604.18867v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) can perform zero-shot classification but are susceptible to adversarial attacks. While robust fine-tuning improves their

Rivian and Volkswagen Group's joint venture put Devin to work on testing and ticket triage across a software platform that will power up to …

SafetyDGX agent

Rivian and Volkswagen Group's joint venture put Devin to work on testing and ticket triage across a software platform that will power up to 30M vehicles. Devin handles autonomous ticket triage in Slac

Symbolic Quantile Regression for the Interpretable Prediction of Conditional Quantiles

SafetyDGX agent

arXiv:2508.08080v2 Announce Type: replace Abstract: Symbolic Regression (SR) is a well-established framework for generating interpretable or white-box predictive models. Although SR has been successfu

Task-Adaptive Admittance Control for Human-Quadrotor Cooperative Load Transportation with Dynamic Cable-Length Regulation

SafetyDGX agent

arXiv:2604.18905v1 Announce Type: new Abstract: The collaboration between humans and robots is critical in many robotic applications, especially in those requiring physical human-robot interaction (pH

TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards

SafetyDGX agent

arXiv:2512.07761v3 Announce Type: replace Abstract: Large language models have seen widespread adoption, yet they remain vulnerable to multi-turn jailbreak attacks, threatening their safe deployment.

User Simulation in the Era of Generative AI: User Modeling, Synthetic Data Generation, and System Evaluation

SafetyDGX agent

arXiv:2501.04410v2 Announce Type: replace Abstract: User simulation is an emerging interdisciplinary topic with multiple critical applications in the era of Generative AI. It involves creating an inte

21 Apr 2026

Adversarial Arena: Crowdsourcing Data Generation through Interactive Competition

SafetyDGX agent

arXiv:2604.17803v1 Announce Type: cross Abstract: Post-training Large Language Models requires diverse, high-quality data which is rare and costly to obtain, especially in low resource domains and for

Arch: An AI-Native Hardware Description Language for Register-Transfer Clocked Hardware Design

SafetyDGX agent

arXiv:2604.05983v2 Announce Type: replace-cross Abstract: We present Arch (AI-native Register-transfer Clocked Hardware), a hardware description language for micro-architecture specification and AI-as

BIASEDTALES-ML: A Multilingual Dataset for Analyzing Narrative Attribute Distributions in LLM-Generated Stories

SafetyDGX agent

arXiv:2604.17008v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used to generate narrative content, including children's stories, which play an important role in social a

CAPC-CG: A Large-Scale, Expert-Directed LLM-Annotated Corpus of Adaptive Policy Communication in China

SafetyDGX agent

arXiv:2510.08986v2 Announce Type: replace Abstract: We introduce CAPC-CG, the Chinese Adaptive Policy Communication (Central Government) Corpus, the first open dataset of Chinese policy directives ann

Deep learning based Non-Rigid Volume-to-Surface Registration for Brain Shift compensation Using Point Cloud

SafetyDGX agent

arXiv:2604.17389v1 Announce Type: new Abstract: Soft-tissue deformation remains a major limitation in image-guided neurosurgery, where intra-operative anatomy can deviate substantially from pre-operat

← Previous
1…4748495051…240
Next →