AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,702 results
30 Jul 2026

LiDARDraft: Generating LiDAR Point Cloud from Versatile Inputs

SafetyDGX agent

arXiv:2512.20105v2 Announce Type: replace Abstract: Generating realistic and diverse LiDAR point clouds is crucial for autonomous driving simulation. Although previous methods achieve LiDAR point clou

Lilith: Backdoor Generalization under Training-Inference Trigger Shift

SafetyDGX agent

arXiv:2607.26099v1 Announce Type: cross Abstract: Machine-learning services increasingly rely on public data, third-party providers, and outsourced training, creating opportunities for data-poisoning

Long-Tailed 3D Point Cloud Dataset Distillation

SafetyDGX agent

arXiv:2607.26763v1 Announce Type: new Abstract: Dataset distillation compresses large-scale datasets into compact synthetic sets while preserving their training utility, enabling efficient 3D point cl


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Lottery Tickets Are Not Deployment Tickets

SafetyDGX agent

arXiv:2607.27031v1 Announce Type: cross Abstract: Reports on how sparsification, compression, and lottery tickets change model behavior have been mixed in the prior literature, with beneficial effects

Misalignment Has a Personality: A Big Five Account of Emergent Misalignment

SafetyDGX agent

arXiv:2607.26389v1 Announce Type: new Abstract: Fine-tuning a language model on data containing a narrow flaw, such as insecure code or incorrect mathematical answers, can cause broad misalignment thr

Multi-Objective Compliance-Integrated Coevolution For Simulated And Real-World Deployment Of Multi-Robot Marine Autonomy

SafetyDGX agent

arXiv:2607.26279v1 Announce Type: new Abstract: Collaborative robots are well-suited to maritime missions that benefit from coordination, such as the exploration of unknown reef structures, inspection

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

SafetyDGX agent

arXiv:2607.27081v1 Announce Type: cross Abstract: Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers

One-Frame Calibration with Siamese Network in Facial Action Unit Recognition

SafetyDGX agent

arXiv:2409.00240v2 Announce Type: replace Abstract: Automatic facial action unit (AU) recognition is used widely in facial expression analysis. Most existing AU recognition systems aim for cross-parti

one of the top cybersecurity models, post-trained from open weights by @depthfirstlabs on @FireworksAI_HQ long-horizon RL is as much an infr…

SafetyDGX agent

one of the top cybersecurity models, post-trained from open weights by @depthfirstlabs on @FireworksAI_HQ long-horizon RL is as much an infra problem as a research one: 100+ turn rollouts, async/pipel

OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment

SafetyDGX agent

arXiv:2607.26981v1 Announce Type: new Abstract: Large language models are increasingly used as decision aids whose probability judgments shape downstream choices. Whether those judgments carry a syste

Parameterized Fair Resource Allocation under Diversity Constraints

SafetyDGX agent

arXiv:2607.26485v1 Announce Type: cross Abstract: Resource allocation across multiple agent groups arises in many applications including e-commerce recommendation systems, housing assignment, and cour

Physically Real-time Infrared Attack against Optical Flow Estimation Networks

SafetyDGX agent

arXiv:2607.26651v1 Announce Type: new Abstract: With the promising performance of deep neural networks on image-based tasks, different real-world applications such as autonomous driving and motion det

Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models

SafetyDGX agent

arXiv:2607.26119v1 Announce Type: cross Abstract: Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their supervised fine-tuned (SFT) counterpar

Quoting Bruce Schneier

SafetyDGX agent

The writing assignments I give my students are gym tasks, not work tasks. I ask them to write policy memos not because the world needs more policy memos. I assign them because the very act of writing,

R-SLPR: Region-based Small-to-Large Point-cloud Registration with Contrastive Learning

SafetyDGX agent

arXiv:2607.26583v1 Announce Type: new Abstract: Point-cloud (PC) registration is fundamental to three-dimensional (3D) perception in robotic systems. However, classic registration algorithms falter wh

Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks

SafetyDGX agent

arXiv:2607.26574v1 Announce Type: cross Abstract: Safety classifiers ('guards') are the dominant black-box defense for vision-language models, yet they judge an input's surface form, not its meaning:

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

SafetyDGX agent

arXiv:2607.19321v2 Announce Type: replace-cross Abstract: As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be

Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance

SafetyDGX agent

arXiv:2607.26333v1 Announce Type: cross Abstract: Chest X-ray (CXR) machine learning relies heavily on automated evaluation using reference standards that aim to approximate clinical judgment. However

Retrospective Orthogonal Design: Response-Surface Reconstruction from Observational Data

SafetyDGX agent

arXiv:2607.26219v1 Announce Type: cross Abstract: Regression estimates from observational data can depend on specification under multicollinearity, while sequential sums of squares (SS) depend on term

RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models

SafetyDGX agent

arXiv:2607.26991v1 Announce Type: new Abstract: Despite the impressive visuomotor capabilities enabled by Vision-Language-Action (VLA) models, their performance often degrades on challenging and out-o

RLMM-Flow: A Flow-based Mobile Manipulation Framework with Latent-Space Reinforcement Learning

SafetyDGX agent

arXiv:2607.26460v1 Announce Type: new Abstract: Mobile manipulation requires generating whole-body action chunks that jointly satisfy goal reaching, collision avoidance, base kinematic constraints, ma

Robostreet Flow: A Lightweight, Ultra-Low-Drag Electric Tractor and Four-Truck Hybrid Convoy Architecture for Minimum-Cost Point-to-Point Freight

SafetyDGX agent

arXiv:2607.26250v1 Announce Type: new Abstract: Line-haul trucking costs are dominated by three comparably sized components: energy, driver labor, and equipment. Most efficiency technologies address o

SCALPEL: Semantic Cross-modal Alignment via LLM-Powered Encoder Learning for Medical Vision-Language Representation

SafetyDGX agent

arXiv:2607.26885v1 Announce Type: new Abstract: Vision-language pre-training (VLP) serves as a cornerstone for medical multimodal representation learning. However, existing medical VLP frameworks are

SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence

SafetyDGX agent

arXiv:2607.27066v1 Announce Type: new Abstract: Scientific figure assessment in peer review differs fundamentally from general image quality evaluation: a figure must be visually legible, faithfully s

Searching for Robust Augmentations to Improve Out-of-Domain Generalization in Dermoscopic Skin Cancer Classification

SafetyDGX agent

arXiv:2607.26765v1 Announce Type: new Abstract: Background/Objectives: Dermoscopic skin lesion classifiers often lose accuracy under domain shift across imaging devices, illumination, and capture arti

Self-Adaptive Learning and Model Predictive Control for Tracking Unknown Dynamics with No Regret

SafetyDGX agent

arXiv:2607.26370v1 Announce Type: cross Abstract: We propose a self-adaptive online learning for control method for tracking unknown target dynamics. The target dynamics can exhibit switching behavior

Shape-Based Inductive Bias for Glioma Grading from Tumor Contours

SafetyDGX agent

arXiv:2607.26090v1 Announce Type: cross Abstract: Glioma grading from tumor contours is often treated as a pixel problem even when the signal of interest is shape. We align closed contours with a func

Situational Awareness, the $20bn hedge fund founded by former OpenAI employee Leopold Aschenbrenner, has sought to raise fresh capital from …

SafetyDGX agent

Situational Awareness, the $20bn hedge fund founded by former OpenAI employee Leopold Aschenbrenner, has sought to raise fresh capital from investors after suffering heavy losses during the recent rou

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

SafetyDGX agent

arXiv:2607.26784v1 Announce Type: new Abstract: Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learnin

SMSP: A Plug-and-Play Strategy of Multi-Scale Perception for MLLMs to Perceive Visual Illusions

SafetyDGX agent

arXiv:2603.23118v2 Announce Type: replace Abstract: Recent works have shown that multimodal large language models (MLLMs) are highly vulnerable to hidden-pattern visual illusions, where the hidden con

Stable and Budget-Feasible Coalition Formation for Clustered Federated Learning: A Hedonic Potential-Game Approach

SafetyDGX agent

arXiv:2607.26788v1 Announce Type: cross Abstract: Clustered federated learning benefits from organizing heterogeneous participants into coalitions that train coalition-specific models, but such cluste

Steering Instruction Hierarchies at Inference Time

SafetyDGX agent

arXiv:2607.26228v1 Announce Type: new Abstract: Instruction hierarchies are a core safety assumption of language model deployment: higher priority inputs, such as system prompts, should override confl

SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual Perception

SafetyDGX agent

arXiv:2607.26985v1 Announce Type: cross Abstract: Deep reinforcement policy learning directly in physical robots (on-robot learning) remains bottlenecked by slow wall-clock training times. We present

The Advantage of Fine-Grained Training

SafetyDGX agent

arXiv:2509.05130v2 Announce Type: replace Abstract: In classification problems, models are trained to predict a class label based on the input data features. However, class labels are organized hierar

// The agent is its own best speculator // Agents spend a large share of wall-clock time waiting on tool results. Speculation hides that lat…

SafetyDGX agent

// The agent is its own best speculator // Agents spend a large share of wall-clock time waiting on tool results. Speculation hides that latency by predicting and pre-executing the next call, but exte

The Confounder Trap: Treatment-Encoding Representations in Causal Inference with Text

SafetyDGX agent

arXiv:2607.26309v1 Announce Type: cross Abstract: Estimating causal effects of linguistic properties from observational text is difficult because the same document can contain both the treatment of in

The narrative: Blame the Agent, instead of the Agency that told him to “apply all your powers and told to achieve this win.” Ironically, man…

SafetyDGX agent

The narrative: Blame the Agent, instead of the Agency that told him to “apply all your powers and told to achieve this win.” Ironically, many forefront members of the AI-Safety community, in their fer

Thinking Under Uncertainty: Evidence Use and Information-Seeking in Language Models

SafetyDGX agent

arXiv:2607.26845v1 Announce Type: new Abstract: Inference-time thinking improves the performance of large language models, but aggregate outcomes do not reveal whether models use available evidence mo

Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs

SafetyDGX agent

arXiv:2607.27122v1 Announce Type: new Abstract: Gastrointestinal (GI) endoscopic image analysis has shifted from single-label classification toward visual question answering (VQA), where a model must

TraceCLIP: Recovering Local Semantics from Patch-to-CLS Contributions

SafetyDGX agent

arXiv:2607.26107v1 Announce Type: new Abstract: Dense vision-language understanding, including object localization, region recognition, and open-vocabulary semantic segmentation, requires associating

Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI Detection

SafetyDGX agent

arXiv:2607.27113v1 Announce Type: new Abstract: The growing capability of image generation models has made synthetic images a routine presence in open media, making robust and generalizable AI-Generat

Visibility-Aware Cooperative Tracking with Decentralized LiDAR-Based Aerial Swarms

SafetyDGX agent

arXiv:2512.01280v2 Announce Type: replace Abstract: Autonomous aerial tracking with drones offers vast potential for surveillance, cinematography, and industrial inspection applications. While single-

Weak-to-Strong On-Policy Distillation

SafetyDGX agent

arXiv:2607.26246v1 Announce Type: new Abstract: On-policy distillation (OPD), which aligns a student with the teacher's token-level distribution on the student's own rollouts, is an effective paradigm

WikiLoop: Jointly Learning to Build and Navigate Agent-Native Wikis with Downstream Feedback

SafetyDGX agent

arXiv:2607.26604v1 Announce Type: new Abstract: Knowledge-base construction and querying are typically optimized in isolation: retrieval-augmented agents operate over a fixed, externally maintained in

29 Jul 2026

A context-adaptive policy framework for robust and reactive robotic manipulation via uncertainty-aware imitation learning

SafetyDGX agent

arXiv:2410.24035v2 Announce Type: replace-cross Abstract: Generating robust and reactive manipulation strategies that can adapt to changing context information is a challenging task in robotics. Over

A Functional Approach to Curve Alignment and Shape Analysis

SafetyDGX agent

arXiv:2503.05632v2 Announce Type: cross Abstract: In many image analysis problems, the contours of objects carry important statistical information about shape. Such contours are typically affected by

A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics

SafetyDGX agent

arXiv:2607.25207v1 Announce Type: new Abstract: This paper investigates a hybrid reinforcement learning setting in tabular Markov Decision Processes (MDPs), where an agent aims to learn an optimal pol

A2D2: Audi Autonomous Driving Dataset

SafetyDGX agent

arXiv:2004.06320v2 Announce Type: replace Abstract: Research in machine learning, mobile robotics, and autonomous driving is accelerated by the availability of high quality annotated data. To this end

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning

SafetyDGX agent

arXiv:2607.24833v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards is a powerful paradigm for eliciting reasoning in large language models, yet it suffers from severe rewar

Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation

SafetyDGX agent

arXiv:2607.25489v1 Announce Type: new Abstract: Large language models and multimodal foundation models are enabling medical artificial intelligence (AI) systems to move beyond isolated prediction and

AlphaCrafter: Harnessing Multi-Agent Workflows for Cross-Sectional Quantitative Trading

SafetyDGX agent

arXiv:2605.05580v2 Announce Type: replace Abstract: Quantitative trading agents have demonstrated substantial promise in automating factor discovery, signal aggregation, and portfolio execution. Howev

Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering

SafetyDGX agent

arXiv:2607.25479v1 Announce Type: cross Abstract: Vision--Language Models (VLMs) are increasingly deployed through a model supply chain in which pretrained checkpoints, architecture definitions, text

Behavior-Driven Explainability

SafetyDGX agent

arXiv:2607.24881v1 Announce Type: new Abstract: As system complexity has vastly increased, it has become significantly more challenging for a single person or a team to fully understand all aspects of

Beyond Predictive Accuracy: A Reliability-Aware Audit of Molecular Representations for Human Olfaction

SafetyDGX agent

arXiv:2607.24848v1 Announce Type: cross Abstract: Pretrained molecular encoders are commonly evaluated through downstream prediction, but predictive accuracy alone does not establish that a learned re

Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs

SafetyDGX agent

arXiv:2607.25600v1 Announce Type: cross Abstract: Retrieval-augmented generation improves knowledge-intensive question answering, but indiscriminate retrieval can introduce irrelevant evidence and unn

Breaking the Curse of Repulsion: Remoteness-Aware Control of Negative Off-Policy Updates

SafetyDGX agent

arXiv:2602.10430v2 Announce Type: replace-cross Abstract: Off-policy policy optimization reuses historical behavior, including negative-advantage samples that suppress known failures. We show that rep

Calibrated Partial Resets: Preventing Policy Collapse in Continual Reinforcement Learning

SafetyDGX agent

arXiv:2607.24996v1 Announce Type: cross Abstract: Neural networks are hindered by accumulating dormant neurons and loss of expressivity throughout training, particularly in non-stationary data setting

CAST: Game Solvers as Turn-Level Teachers for LLM Agents

SafetyDGX agent

arXiv:2607.25308v1 Announce Type: cross Abstract: Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning w

CIFNet: An Analytic Neural Learning Framework for Efficient and Calibrated Class-Incremental Learning

SafetyDGX agent

arXiv:2509.11285v2 Announce Type: replace-cross Abstract: Class-Incremental Learning (CIL) in deep neural networks is conventionally framed as an iterative gradient-based optimization problem, incurri

Construction-Driven Injection: Linguistically-Grounded Edit-Based Code-Mixing Fingerprints for Large Language Models

SafetyDGX agent

arXiv:2607.25633v1 Announce Type: cross Abstract: Large language models (LLMs) are costly intellectual assets that remain exposed to unauthorized redistribution and commercial misuse. Injected fingerp

← Previous
1…2223242526…212
Next →