AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,707 results
5 May 2026

IPS: In-Prompt Process Supervision for Short Video Content Moderation

SafetyDGX agent

arXiv:2412.15251v3 Announce Type: replace Abstract: Multimodal large language models (MLLMs) are effective at capturing the semantics of short video content; however, they often fail to attend to the

🆕 @katiemiller has started following @GaryMarcus

SafetyDGX agent

Katie Miller began following Gary Marcus on X (formerly Twitter). Gary Marcus is a cognitive scientist and AI researcher known for his public commentary on artificial intelligence and technology polic

Knowledge-Based Design Requirements for Generative Social Robots in Higher Education

SafetyDGX agent

arXiv:2602.12873v4 Announce Type: replace-cross Abstract: Generative social robots (GSRs) powered by large language models enable adaptive, conversational tutoring but also introduce risks such as mis


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Lateral String Stability for Vehicle Platoons: Formulation, Definition, and Analysis

SafetyDGX agent

arXiv:2605.01731v1 Announce Type: new Abstract: Platooning of connected and automated vehicles provides significant benefits in terms of energy efficiency, traffic throughput, and, most critically, sa

Learning to Act Through Contact: A Unified View of Multi-Task Robot Learning

SafetyDGX agent

arXiv:2510.03599v2 Announce Type: replace Abstract: We present a unified framework for multi-task locomotion and manipulation policy learning grounded in a contact-explicit representation. Instead of

Less is More: Geometric Unlearning for LLMs with Minimal Data Disclosure

SafetyDGX agent

arXiv:2605.01735v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed in real-world systems, they must support post-hoc removal of specific content to meet privacy

Less Precise Can Be More Reliable: A Systematic Evaluation of Quantization's Impact on VLMs Beyond Accuracy

SafetyDGX agent

arXiv:2509.21173v5 Announce Type: replace Abstract: Vision-Language Models (VLMs) such as CLIP have revolutionized zero-shot classification and safety-critical tasks, including Out-of-Distribution (OO

Linking spatial biology and clinical histology via Haiku

SafetyDGX agent

arXiv:2605.00925v1 Announce Type: cross Abstract: Integrating molecular, morphological, and clinical data is essential for basic and translational biomedical research, yet systematic frameworks for jo

LLM-Augmented Semantic Steering of Text Embedding Projection Spaces

SafetyDGX agent

arXiv:2605.01957v1 Announce Type: cross Abstract: Low-dimensional projections of text embeddings support visual analysis of document collections, but their spatial organization may not reflect the rel

LLM-Based Agentic Negotiation for 6G: Addressing Uncertainty Neglect and Tail-Event Risk

SafetyDGX agent

arXiv:2511.19175v2 Announce Type: replace-cross Abstract: A critical barrier to the trustworthiness of sixth-generation (6G) agentic autonomous networks is the uncertainty neglect bias; a cognitive te

LLM-VA: Resolving the Jailbreak-Overrefusal Trade-off via Vector Alignment

SafetyDGX agent

arXiv:2601.19487v2 Announce Type: replace Abstract: Safety-aligned LLMs suffer from two failure modes: jailbreak (answering harmful inputs) and over-refusal (declining benign queries). Existing vector

Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness

SafetyDGX agent

arXiv:2506.24056v2 Announce Type: replace-cross Abstract: RLHF-style alignment trains language models to refuse unsafe requests, but how much operational margin does this refusal rest on? We introduce

Low-Latency Video Anonymization for Crowd Anomaly Detection: Privacy Versus Performance

SafetyDGX agent

arXiv:2410.18717v2 Announce Type: replace Abstract: Recent advancements in artificial intelligence hold ample potential for monitoring applications using surveillance cameras. However, concerns about

LVLM-Aided Alignment of Task-Specific Vision Models

SafetyDGX agent

arXiv:2512.21985v2 Announce Type: replace Abstract: In high-stakes domains, small task-specific vision models are crucial due to their low computational requirements and the availability of numerous m

Machine Learning Enhanced Laser Spectroscopy for Multi-Species Gas Detection in Complex and Harsh Environments

SafetyDGX agent

arXiv:2605.01306v1 Announce Type: cross Abstract: Laser absorption spectroscopy (LAS) is a well-established technique for non-intrusive measurement of gas species in combustion and atmospheric environ

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate

SafetyDGX agent

arXiv:2605.01347v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own trajectories under token-level teacher supervision, but existing methods are capped by a single

Major new class action lawsuit accuses Meta of copyright infringement around AI training. It says they trained on pirated books, and that AI…

SafetyDGX agent

Major new class action lawsuit accuses Meta of copyright infringement around AI training. It says they trained on pirated books, and that AI books flooding the market demonstrates market harm. These l

Manifold-Constrained Adversarial Training for Long-Tailed Robustness via Geometric Alignment

SafetyDGX agent

arXiv:2605.02183v1 Announce Type: new Abstract: Adversarial training is effective on balanced datasets, but its robustness degrades under longtailed class distributions, where tail classes suffer high

Mean Testing under Truncation beyond Gaussian

SafetyDGX agent

arXiv:2605.01335v1 Announce Type: cross Abstract: We characterize the fundamental limits of high-dimensional mean testing under arbitrary truncation, where samples are drawn from the conditional distr

Meta says it will expand Instagram teen account safeguards to 27 EU countries, and plans to roll them out on Facebook in the US, ahead of the UK and EU in June (Foo Yun Chee/Reuters)

SafetyDGX agent

Foo Yun Chee / Reuters: Meta says it will expand Instagram teen account safeguards to 27 EU countries, and plans to roll them out on Facebook in the US, ahead of the UK and EU in June — Meta Platforms

Minimizing Collateral Damage in Activation Steering

SafetyDGX agent

arXiv:2605.01167v1 Announce Type: new Abstract: Activation steering is a method for controlling Large Language Model (LLM) behavior by intervening in its internal representations to increase the align

MIRA: A Score for Conditional Distribution Accuracy and Model Comparison

SafetyDGX agent

arXiv:2605.02014v1 Announce Type: cross Abstract: We introduce Mira, a sample-based score for assessing the accuracy of a candidate conditional distribution using only joint samples from the true data

Mitigating Misalignment Contagion by Steering with Implicit Traits

SafetyDGX agent

arXiv:2605.02751v1 Announce Type: cross Abstract: Language models (LMs) are increasingly used in high-stakes, multi-agent settings, where following instructions and maintaining value alignment are cri

MOC-3D: Manifold-Order Consistency for Text-to-3D Generation

SafetyDGX agent

arXiv:2605.01743v1 Announce Type: new Abstract: With the burgeoning development of fields such as the Metaverse, Virtual Reality (VR), and Digital Twins, text-to-3D generation has emerged as a researc

Model-Based Proactive Cost Generation for Learning Safe Policies Offline with Limited Violation Data

SafetyDGX agent

arXiv:2605.01356v1 Announce Type: new Abstract: Learning constraint-satisfying policies from offline data without risky online interaction is crucial for safety-critical decision making. Conventional

Momentum-Anchored Multi-Scale Fusion Model for Long-Tailed Chest X-Ray Classification

SafetyDGX agent

arXiv:2605.02292v1 Announce Type: new Abstract: Chest X-ray classification suffers from severe class imbalance where gradient updates bias toward majority classes, causing feature drift and poor perfo

MTA: Multi-Granular Trajectory Alignment for Large Language Model Distillation

SafetyDGX agent

arXiv:2605.01374v1 Announce Type: new Abstract: Knowledge distillation is a key technique for compressing large language models (LLMs), but most existing methods align representations at fixed layers

Multi-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and Reasoning

SafetyDGX agent

arXiv:2605.01736v1 Announce Type: new Abstract: Understanding the geometric and semantic structure of environments is essential for embodied navigation and reasoning. Existing semantic mapping methods

Multi-User Dueling Bandits: A Fair Approach using Nash Social Welfare

SafetyDGX agent

arXiv:2605.01961v1 Announce Type: new Abstract: Learning from human preference data is becoming a useful tool, from fine-tuning large language models to training reinforcement learning agents. However

Multi-View Hierarchical Representation Learning of Fetal Hemodynamics for Maternal Hypertension Detection at the Edge

SafetyDGX agent

arXiv:2605.00872v1 Announce Type: cross Abstract: Hypertensive disorders of pregnancy remain a leading cause of maternal and fetal morbidity worldwide, yet diagnosis relies on intermittent cuff-based

Multimodal Data Curation Through Ranked Retrieval

SafetyDGX agent

arXiv:2605.01163v1 Announce Type: cross Abstract: Shared embedding spaces are widely used for multimodal search and data curation. In practice, two problems often limit how well this works. First, emb

NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation Tasks

SafetyDGX agent

arXiv:2508.02046v4 Announce Type: replace-cross Abstract: Recent advances in Graphical User Interface (GUI) and embodied navigation have driven progress, yet these domains have largely evolved in isol

Now, now. Be nice. Greg has a been great witness!. For Elon.

SafetyDGX agent

This post by AI researcher Gary Marcus appears to reference a witness or testimony related to Elon Musk, using a lighthearted tone to encourage civil discourse. Without access to the specific tweet co

Online Safety Filter for Deformable Object Manipulation with Horizon Agnostic Neural Operators

SafetyDGX agent

arXiv:2605.01069v1 Announce Type: new Abstract: Safety critical control of robotic manipulation tasks involving deformable media such as fluids, cloth, and soft objects remains challenging because exi

OpenAI: The Movie Act I: ChatGPT sets records! Act II: Sam gets fired, rehired Act III: Greg’s diary blows it all up, in court.

SafetyDGX agent

This post by AI researcher Gary Marcus summarizes major events in OpenAI's recent history across three acts: ChatGPT's record-breaking success, the dramatic firing and rehiring of CEO Sam Altman, and

Parking Assistance for Trailer-Truck Transport Vehicles Using Sensor Fusion and Motion Planning

SafetyDGX agent

arXiv:2605.02716v1 Announce Type: new Abstract: Autonomous driving technology has rapidly evolved over the past decade, offering significant improvements in transportation efficiency, safety, and cost

Patient-Specific Optimization for Mandibular Reconstruction Planning with Enhanced Bone Union

SafetyDGX agent

arXiv:2605.01084v1 Announce Type: new Abstract: Mandibular reconstruction with vascularized bone grafts is complicated by donor-host nonunion, and current virtual surgical planning produces a geometri

Perceptual Flow Network for Visually Grounded Reasoning

SafetyDGX agent

arXiv:2605.02730v1 Announce Type: new Abstract: Despite the success of Large-Vision Language Models (LVLMs), general optimization objectives (e.g., standard MLE) fail to constrain visual trajectories,

PRCD-MAP: Learning How Much to Trust Imperfect Priors in Causal Discovery

SafetyDGX agent

arXiv:2605.01669v1 Announce Type: cross Abstract: External priors of unknown reliability create a brittle trade-off in causal discovery: blind trust amplifies errors, blind rejection wastes signal. Re

ProPACT: A Proactive AI-Driven Adaptive Collaborative Tutor for Pair Programming

SafetyDGX agent

arXiv:2605.02703v1 Announce Type: cross Abstract: Effective pair programming depends on coordination of attention, cognitive effort, and joint regulation over time, yet most adaptive learning systems

Protein-Conditioned Multi-Objective Reinforcement Learning for Full-Length mRNA Design

SafetyDGX agent

arXiv:2605.01513v1 Announce Type: new Abstract: Designing therapeutic messenger RNA (mRNA) requires creating full-length transcripts that carefully balance stability, translation efficiency, and immun

ProtoFair: Fair Self-Supervised Contrastive Learning via Pseudo-Counterfactual Pairs

SafetyDGX agent

arXiv:2605.01971v1 Announce Type: new Abstract: Self-supervised learning methods learn high-quality visual representations, yet recent studies show that these representations often capture demographic

RA-CMF: Region-Adaptive Conditional MeanFlow for CT Image Reconstruction

SafetyDGX agent

arXiv:2605.00901v1 Announce Type: new Abstract: The use of CT imaging is important for screening, diagnosis, therapy planning, and prognosis of lung cancers. Unfortunately, due to differences in imagi

Rationality Measurement and Theory for Reinforcement Learning Agents

SafetyDGX agent

arXiv:2602.04737v2 Announce Type: replace Abstract: This paper proposes a suite of rationality measures and associated theory for reinforcement learning agents, a property increasingly critical yet ra

RECAP: Transparent Inference-Time Emotion Alignment for Medical Dialogue Systems

SafetyDGX agent

arXiv:2509.10746v3 Announce Type: replace Abstract: Large language models in healthcare often produce emotionally flat or opaque responses, failing to provide the transparent reasoning required for cl

Registration-Free Learnable Multi-View Capture of Faces in Dense Semantic Correspondence

SafetyDGX agent

arXiv:2605.01450v1 Announce Type: new Abstract: Recent frameworks like ToFu and TEMPEH provide an automated alternative to classical registration pipelines by predicting 3D meshes in dense semantic co

Reinforcement Learning from Compiler and Language Server Feedback

SafetyDGX agent

arXiv:2510.22907v2 Announce Type: replace Abstract: Coding agents fail when text-level guesses outrun program facts: they hallucinate APIs, drift to the wrong symbol, and apply edits without evidence

Reliability-Oriented Multilingual Orthopedic Diagnosis: A Domain-Adaptive Modeling and a Conceptual Validation Framework

SafetyDGX agent

arXiv:2605.02266v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly proposed for clinical decision support including multilingual diagnosis in low-resource settings. However,

Remote Action Generation: Remote Control with Minimal Communication

SafetyDGX agent

arXiv:2605.01833v1 Announce Type: cross Abstract: We address the challenge of remote control where one or more actors, lacking direct reward access, are steered by a controller over a communication-co

Representation learning from OCT images

SafetyDGX agent

arXiv:2605.02589v1 Announce Type: new Abstract: Optical Coherence Tomography (OCT) has become one of the most used imaging modality in ophthalmology. It provides high-resolution, non-invasive visualiz

Research on Vision-Language Question Answering Models for Industrial Robots

SafetyDGX agent

arXiv:2605.01483v1 Announce Type: new Abstract: A hierarchical cross-modal fusion model is proposed for vision-language question answering (VLQA) in industrial robotics, targeting the challenges of se

Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance

SafetyDGX agent

arXiv:2605.01325v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have enhanced traditional LLMs with visual capabilities through the integration of vision encoders. While recent works hav

Retrieval-Guided Generation for Safer Histopathology Image Captioning

SafetyDGX agent

arXiv:2605.00893v1 Announce Type: new Abstract: Generative vision-language models can produce fluent medical image captions but remain prone to hallucination, over-specific diagnostic claims, and fact

Rhythm: Learning Interactive Whole-Body Control for Dual Humanoids

SafetyDGX agent

arXiv:2603.02856v2 Announce Type: replace Abstract: Realizing interactive whole-body control for multi-humanoid systems is critical for unlocking complex collaborative capabilities in shared environme

Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction

SafetyDGX agent

arXiv:2602.10101v2 Announce Type: replace Abstract: 3D spatial perception is fundamental to generalizable robotic manipulation, yet obtaining reliable, high-quality 3D geometry remains challenging. De

Safe Planning in Interactive Environments via Iterative Policy Updates and Adversarially Robust Conformal Prediction

SafetyDGX agent

arXiv:2511.10586v2 Announce Type: replace-cross Abstract: Safe planning of an autonomous agent in interactive environments -- such as the control of a self-driving vehicle among pedestrians -- poses a

SAGA: A Robust Self-Attention and Goal-Aware Anchor-based Planner for Safe UAV Autonomous Navigation

SafetyDGX agent

arXiv:2605.02301v1 Announce Type: new Abstract: Agile unmanned aerial vehicle (UAV) navigation in cluttered environments demands a planning architecture that is both computationally efficient and stru

Semantic-Contact Fields for Category-Level Generalizable Tactile Tool Manipulation

SafetyDGX agent

arXiv:2602.13833v2 Announce Type: replace Abstract: Generalizing tool manipulation requires both semantic planning and precise physical control. Modern generalist robot policies, such as Vision-Langua

Semantic Risk-Aware Heuristic Planning for Robotic Navigation in Dynamic Environments: An LLM-Inspired Approach

SafetyDGX agent

arXiv:2605.02862v1 Announce Type: new Abstract: The integration of Large Language Model (LLM) reasoning principles into classical robot path planning represents a rapidly emerging research direction.

Sentra-Guard: A Real-Time Multilingual Defense Against Adversarial LLM Prompts

SafetyDGX agent

arXiv:2510.22628v2 Announce Type: replace-cross Abstract: This paper presents a real-time modular defense system named Sentra-Guard. The system detects and mitigates jailbreak and prompt injection att

← Previous
1…165166167168169…212
Next →