AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
Human
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,596 results
12 Aug 2026

A Convolutional Layer Activation Dimensionality Reduction for Out-of-Distribution and Adversarial Attack Detection Methods

SafetyDGX agent

arXiv:2608.10203v1 Announce Type: new Abstract: Despite the success of convolutional neural networks in image classification tasks and their general application in multi-modal models, their susceptibi

A Joint-Distribution Route to Fair Representations with Continuous Sensitive Attributes

SafetyDGX agent

arXiv:2608.10470v1 Announce Type: new Abstract: Fair representation learning with a continuous sensitive attribute S requires a representation Z that is statistically independent of S. Existing criter

A Neural Network Based Teleoperation for Remote Controlled Vehicles

SafetyDGX agent
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.10367v1 Announce Type: new Abstract: Direct teleoperation of vehicles faces critical technical bottlenecks: communication latency and the operator's inability to physically perceive unmodel

Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents

SafetyDGX agent

arXiv:2608.11110v1 Announce Type: new Abstract: When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compa

AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss

SafetyDGX agent

arXiv:2608.11205v1 Announce Type: new Abstract: Frechet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-le

APCReg: Anatomical-Prior-Guided Coarse-to-Fine CBCT--IOS Registration via Multi-View Projection and Reliability-Controlled Residual Correction

SafetyDGX agent

arXiv:2608.09993v1 Announce Type: cross Abstract: Registration between cone-beam computed tomography (CBCT) and intraoral scans (IOS) is essential for patient-specific surgical planning. However, disp

Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety

SafetyDGX agent

arXiv:2512.06227v3 Announce Type: replace Abstract: Real-world indicators play an important role in many Natural Language Processing (NLP) applications, such as life events for mental health analysis

Beyond Forecasting: Recasting Volatility Control as a Routing Problem

SafetyDGX agent

arXiv:2608.10375v1 Announce Type: cross Abstract: Volatility control converts risk estimates into portfolio exposure, yet existing approaches often rely on a fixed volatility estimator or a pre-define

BooST: Bridging Semantics and Motions for Efficient Skill Transfer

SafetyDGX agent

arXiv:2608.10600v1 Announce Type: cross Abstract: Skill abstraction---the process of learning reusable and temporally extended behaviors---has emerged as a key paradigm for improving sample efficiency

CARE: Confidence-Aware Reasoning for Reliable Medical VQA

SafetyDGX agent

arXiv:2608.10964v1 Announce Type: cross Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual qu

Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory

SafetyDGX agent

arXiv:2608.09937v1 Announce Type: new Abstract: Recent work in NLP has probed large language models for their understanding of cultural norms across countries. However, this work typically considers d

ConfTriage: A Calibration-Aware LLM Triage Framework for Pulmonary Nodule Malignancy with Selective Specialist Deferral

SafetyDGX agent

arXiv:2608.10885v1 Announce Type: new Abstract: Pulmonary nodule malignancy prediction typically depends on image-trained specialist deep learning (DL) models that require substantial annotated imagin

ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering

SafetyDGX agent

arXiv:2608.10996v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has been especially effective in mathematics and coding, where answers can be checked automatically. Many

Convergence of Sign-based Random Reshuffling Algorithms for Nonconvex Optimization

SafetyDGX agent

arXiv:2310.15976v4 Announce Type: replace Abstract: signSGD is attractive in nonconvex optimization because it communicates sign-valued rather than full-precision gradients. Several standard analyses

CRHT: A Continuous Regression Hybrid Transformer for Vessel Trajectory Prediction with Online Cluster Sampling

SafetyDGX agent

arXiv:2608.10256v1 Announce Type: new Abstract: Accurate vessel trajectory prediction is critical for maritime safety and anomaly detection, yet existing models often struggle with geographic bias and

Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning

SafetyDGX agent

arXiv:2608.10473v1 Announce Type: cross Abstract: Offline-to-online (O2O) reinforcement learning aims to leverage policies pretrained on static datasets while improving them through online interaction

Data Attribution of Emergent Misalignment with Persona Features

SafetyDGX agent

arXiv:2608.11025v1 Announce Type: new Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leadi

Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition Agents

SafetyDGX agent

arXiv:2608.10441v1 Announce Type: cross Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expen

DIMOS: Disentangling Instance-level Moving Object Segmentation

SafetyDGX agent

arXiv:2606.12826v2 Announce Type: replace-cross Abstract: Moving instance segmentation (MIS) attracts increasing attention due to its broad applications in traffic surveillance, autonomous driving, an

Do AI weather models miss extremes?

SafetyDGX agent

arXiv:2608.09972v1 Announce Type: cross Abstract: First-generation AI weather models are often reported to underperform at extremes, mostly in reanalysis-based evaluations of deterministic regression

Do Time-Series Forecasters Use the Right History: Recoverability, Recovery, and Functional Use of Temporal Delays

SafetyDGX agent

arXiv:2608.10433v1 Announce Type: new Abstract: Forecast accuracy does not tell us which past inputs produced a prediction. We separate three questions for time-series models with known delay structur

Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving

SafetyDGX agent

arXiv:2608.10386v1 Announce Type: new Abstract: Sample-efficient reinforcement learning for autonomous driving is often limited by the trade-off between data efficiency and model bias. While world mod

Dual-Loop Self-Evolution via Verifiable Emotion Feedback for Multi-Turn Empathetic Dialogue

SafetyDGX agent

arXiv:2608.10626v1 Announce Type: new Abstract: Large language models have demonstrated conversational capabilities, yet empathetic competence remains challenging. Empathetic support is inherently mul

Dual Space Preconditioning for Gradient Descent in the Overparameterized Regime

SafetyDGX agent

arXiv:2603.10485v3 Announce Type: replace-cross Abstract: In this work, we study the convergence properties of the Dual Space Preconditioned Gradient Descent, encompassing optimizers such as Normalize

Dual Stress: Runtime Safety Monitoring for Safety-Constrained MPC Navigation

SafetyDGX agent

arXiv:2608.10791v1 Announce Type: new Abstract: Runtime hazard monitors for autonomous naviga- tion are conventionally built from geometric quantities: predicted clearance, time to collision, and requ

Efficient Hypergradient Descent for Inverse Reinforcement Learning

SafetyDGX agent

arXiv:2608.11052v1 Announce Type: new Abstract: Inverse reinforcement learning (IRL) aims to recover a reward function under which the resulting policy reproduces the behavior observed in expert demon

ELMER: Evolutionary Language Model that Explores and Refines

SafetyDGX agent

arXiv:2608.10196v1 Announce Type: cross Abstract: Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space. Syntactic edit size is a

Enhancing Automated Essay Scoring With Three Techniques: Two-Stage Fine-Tuning, Score Alignment, and Self-Training

SafetyDGX agent

arXiv:2602.01747v2 Announce Type: replace Abstract: Automated Essay Scoring (AES) plays a crucial role in education by providing scalable and efficient assessment tools. However, in real-world setting

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes

SafetyDGX agent

arXiv:2608.10209v1 Announce Type: new Abstract: Feedback signals used to train Large Language Models (LLMs) are the primary driver of their behavior and our main lever for instilling alignment with hu

Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases

SafetyDGX agent

arXiv:2608.10503v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. The NL

Expert-Guided g-computation with Large Language Models for Estimating Causal Effects on Timings: Applications to Hospital Quality Improvement

SafetyDGX agent

arXiv:2608.10339v1 Announce Type: cross Abstract: Hospital quality improvement (QI) programs routinely face multiple candidate interventions to optimize hospital flow, but existing methods struggle to

FedCGR: Federated Cross-Domain Generative Recommendation

SafetyDGX agent

arXiv:2608.10929v1 Announce Type: new Abstract: Cross-domain recommendation (CDR) transfers preference knowledge across related domains, but federated deployment makes cross-domain alignment difficult

FoR-SALE: Frame of Reference-guided Spatial Adjustment in LLM-based Diffusion Editing

SafetyDGX agent

arXiv:2509.23452v2 Announce Type: replace-cross Abstract: Current text-to-image generation models, even state-of-the-art models, exhibit a significant performance gap when spatial expressions are desc

From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop

SafetyDGX agent

arXiv:2608.11171v1 Announce Type: cross Abstract: The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings pap

From Prediction to Incrementality: Causal Optimization for Large-Scale Targeting and Recommendation

SafetyDGX agent

arXiv:2608.10182v1 Announce Type: cross Abstract: Large-scale targeting and recommendation systems are typically built around predictive scores fed into heuristic or local allocation. When the busines

Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Models

SafetyDGX agent

arXiv:2601.17387v3 Announce Type: replace Abstract: Multilingual speech-text models rely on cross-modal language alignment to transfer knowledge between speech and text, but it remains unclear whether

Hidden in Plain Sight: Diffusion-Based Unrestricted Robotic Attacks on Vision-Language-Action Models

SafetyDGX agent

arXiv:2608.10393v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown strong capabilities in controlling robots across diverse manipulation tasks. However, their adversarial r

Hierarchical Empirical-Bayes Naive Bayes: Minimax Smoothing and Calibration with AODE Extension

SafetyDGX agent

arXiv:2608.11162v1 Announce Type: new Abstract: The Naive Bayes (NB) classifier remains a standard choice for categorical data, yet its widely used smoothing rules, such as Laplace, Lidstone, Krichevs

Hip Energized Monopedal Hopping

SafetyDGX agent

arXiv:2608.10387v1 Announce Type: new Abstract: We present a novel stepping strategy for pitch unlocked planar monopeds where the reaction torques from stabilizing pitch with a conventional PD + feedf

How to Verify Consistency of Probabilistic Claims

SafetyDGX agent

arXiv:2608.11181v1 Announce Type: cross Abstract: When a probabilistic predictor answers many conditional-probability queries, are its answers self-consistent, and can this be verified in polynomial t

IADD-TR: Intervention-Aware Dynamics Decoupling with Targeted Regularization for Model-Based Reinforcement Learning

SafetyDGX agent

arXiv:2608.10634v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL), which learns environment dynamics to generate synthetic experience, is a promising approach to sample-efficie

Injecting Hallucinations in Autonomous Vehicles: A Component-Agnostic Safety Evaluation Framework

SafetyDGX agent

arXiv:2510.07749v2 Announce Type: replace Abstract: Perception failures in autonomous vehicles (AV) remain a major safety concern because they are the basis for many accidents. To study how these fail

INSIDE the Student's Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators

SafetyDGX agent

arXiv:2608.10492v1 Announce Type: new Abstract: Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them. In education, w

IO Factory: Simulating AI-Enabled Influence Campaigns at Scale

SafetyDGX agent

arXiv:2608.10920v1 Announce Type: new Abstract: We introduce IO Factory, an AI-driven framework for simulating information and influence campaigns as fully integrated, traceable processes. The threat

Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen)

SafetyDGX agent

arXiv:2509.06191v2 Announce Type: replace-cross Abstract: Recent 3D generative models, which are capable of generating full object shapes from just a few images, now open up new opportunities in robot

Leveraging Large Language Models for Causal Discovery: a Constraint-based, Argumentation-driven Approach

SafetyDGX agent

arXiv:2602.16481v2 Announce Type: replace Abstract: Causal discovery seeks to uncover causal relations from data, typically represented as causal graphs, and is essential for predicting the effects of

LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations

SafetyDGX agent

arXiv:2602.09924v4 Announce Type: replace-cross Abstract: Running LLMs with extended reasoning on every problem is expensive, but determining which inputs actually require additional compute remains c

LoRCA: LoRA Cycle Adaptation for Histology to HiP-CT Translation with DINOv3

SafetyDGX agent

arXiv:2608.10002v1 Announce Type: cross Abstract: Hierarchical Phase-Contrast Tomography (HiP-CT) is a synchrotron based X-ray imaging technique that enables non-destructive, volumetric imaging of int

MARCO: Click-Intent Decomposition for Calibrated Ads Conversion Prediction

SafetyDGX agent

arXiv:2608.10562v1 Announce Type: new Abstract: Not all clicks are equal. Industrial ads ranking decouples conversion probability into click-through rate (CTR) and post-click conversion rate (CVR), ye

MarkNull: Model-Agnostic Watermark Removal in AI-Generated Images via On-Manifold Latent Manipulation

SafetyDGX agent

arXiv:2608.10166v1 Announce Type: cross Abstract: Digital watermarking has emerged as a critical technique for provenance and copyright attribution in AI-generated imagery, yet its robustness against

MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models

SafetyDGX agent

arXiv:2608.10635v1 Announce Type: cross Abstract: Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain chall

MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale

SafetyDGX agent

arXiv:2608.10333v1 Announce Type: new Abstract: LLM agents execute heterogeneous sequences of model calls within a single task: some invocations require careful reasoning, while others are structured

MIDAS: Mutual Information Disentanglement with Uncertainty-Aware Fusion for Incomplete Multimodal Sentiment Analysis

SafetyDGX agent

arXiv:2608.09986v1 Announce Type: new Abstract: Most existing multimodal sentiment analysis approaches assume access to complete multimodal inputs. However, real-world applications frequently encounte

Mitigating Bus Bunching with Reinforcement Learning Enhanced by Semantic Stop Embedding

SafetyDGX agent

arXiv:2608.10207v1 Announce Type: new Abstract: Bus bunching degrades service regularity and increases passenger waiting in high-frequency transit. Existing reinforcement-learning-based holding contro

Most biomedical publications show signs of LLM-assisted writing

SafetyDGX agent

arXiv:2608.10715v1 Announce Type: cross Abstract: Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing can be valua

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models

SafetyDGX agent

arXiv:2608.10864v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong image and video understanding, yet their visual-spatial representations remain geometrically fragile,

Multiplayer Nash Preference Optimization

SafetyDGX agent

arXiv:2509.23102v4 Announce Type: replace Abstract: Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences. Ho

Navigating the Proximity-Safety Balance: Constraint Decomposition for Human Following in Pedestrian Crowds

SafetyDGX agent

arXiv:2608.10056v1 Announce Type: cross Abstract: Following a target human in crowded environments involves an inherent conflict between staying close to the target and navigating safely among surroun

Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models

SafetyDGX agent

arXiv:2608.10405v1 Announce Type: cross Abstract: Many studies have shown that specially crafted inputs can induce large language models (LLMs) to generate excessively long outputs, resulting in signi

Observational Policy Ranking for SMB Financial Guidance from Multi-Action Accounting Logs

SafetyDGX agent

arXiv:2608.10050v1 Announce Type: new Abstract: Small and medium-sized businesses need timely financial guidance, yet historical accounting logs record self-selected and often co-occurring business ch

← Previous
123…210
Next →