AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,435 results
Safety

Intelligent Offloading in Vehicular Edge Computing: A Comprehensive Review of Deep Reinforcement Learning Approaches and Architectures

DGX agent

arXiv:2502.06963v3 Announce Type: replace-cross Abstract: The increasing complexity of Intelligent Transportation Systems (ITS) has led to significant interest in computational offloading to external

safetyarxiv-cs-ai
27 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Intuitions of Machine Learning Researchers about Transfer Learning for Medical Image Classification

DGX agent

arXiv:2510.00902v2 Announce Type: replace Abstract: Transfer learning is crucial for medical imaging, yet the selection of source datasets often relies on researchers' intuition rather than systematic

safetyarxiv-cs-cv
27 May 2026
Safety

It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty

DGX agent

arXiv:2605.27288v1 Announce Type: cross Abstract: Large language models (LLMs) are known to abandon their initial stance to conform to user pushback. While prior research largely attributes this behav

safetyarxiv-cs-ai
27 May 2026
Safety

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models

DGX agent

arXiv:2605.26409v1 Announce Type: cross Abstract: Evaluating and mitigating a generative system's susceptibility to jailbreak attacks is critical to its safe deployment. Given the number of deployable

safetyarxiv-cs-ai
27 May 2026
Safety

KARMA: Karma-Aligned Reward Model Adaptation

DGX agent

arXiv:2605.26738v1 Announce Type: new Abstract: Human communication depends on implicit social signals where effectiveness is shaped by tone, context, and conversational norms rather than semantic con

safetyarxiv-cs-cl
27 May 2026
Safety

LearnedCache: An eBPF-Integrated Perceptron-Based Eviction Policy for the Linux Page Cache

DGX agent

arXiv:2605.26168v1 Announce Type: cross Abstract: Linux is the foundation of the digital age, accounting for the majority of the cloud and mobile OS markets. Any device that runs Linux uses the Linux

safetyarxiv-cs-lg
27 May 2026
Safety

Learning Dynamic Graph Representations through Timespan View Contrasts

DGX agent

arXiv:2605.27063v1 Announce Type: new Abstract: The rich information underlying graphs has inspired further investigation of unsupervised graph representation. Existing studies mainly depend on node f

safetyarxiv-cs-lg
27 May 2026
Safety

Learning to Orchestrate Agents under Uncertainty

DGX agent

arXiv:2605.27073v1 Announce Type: new Abstract: Adaptive orchestration of heterogeneous agents requires making sequential delegation decisions under uncertain and evolving agent behaviour, e.g., coord

safetyarxiv-cs-lg
27 May 2026
Safety

Learning to Reason Efficiently with Discounted Reinforcement Learning

DGX agent

arXiv:2510.23486v2 Announce Type: replace Abstract: Large reasoning models (LRMs) often consume excessive tokens, inflating computational cost and latency. More broadly, in goal reaching sequential de

safetyarxiv-cs-lg
27 May 2026
Safety

Less is More: Early Stopping Rollout for On-Policy Distillation

DGX agent

arXiv:2605.27028v1 Announce Type: cross Abstract: On-policy distillation has recently emerged as a promising alternative to standard sequence-level imitation, training a student by scoring its own rol

safetyarxiv-cs-ai
27 May 2026
Safety

Linear and Neural Dueling Bandits with Delayed Feedback

DGX agent

arXiv:2605.26554v1 Announce Type: cross Abstract: Contextual dueling bandits form a cornerstone of preference-based decision-making, with critical applications in recommender systems and large languag

safetyarxiv-cs-ai
27 May 2026
Safety

MAIGO: Mitigating Lost-in-Conversation with History-Cleaned On-Policy Self-Distillation

DGX agent

arXiv:2605.27186v1 Announce Type: new Abstract: Large language models often solve tasks from a fully specified prompt but degrade when the same requirements unfold over multiple turns, known as the lo

safetyarxiv-cs-cl
27 May 2026
Safety

MATCHA: Matching Text via Contrastive Semantic Alignment

DGX agent

arXiv:2605.27345v1 Announce Type: new Abstract: Reliable evaluation is essential for understanding large language model (LLM) performance, yet today's go-to metrics, namely token-overlap scores (e.g.,

safetyarxiv-cs-cl
27 May 2026
Safety

MechRL: Reinforcement Learning Agents Perform Circuit Discovery for Mechanistic Interpretability

DGX agent

arXiv:2605.26343v1 Announce Type: new Abstract: Mechanistic interpretability has identified small sets of attention heads that implement specific behaviours in transformer language models, but recover

safetyarxiv-cs-lg
27 May 2026
Safety

MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning

DGX agent

arXiv:2605.26154v1 Announce Type: cross Abstract: LLM-driven agents are capable of selecting external tools to complete users' tasks. However, attackers could compromise such process, steering agents

safetyarxiv-cs-ai
27 May 2026
Safety

Mildly Overparameterized ReLU Networks on Orthogonal Data: Incremental Learning and Implicit Bias

DGX agent

arXiv:2605.27097v1 Announce Type: new Abstract: The successful training of neural networks hinges on the use of first order optimization methods, yet the theoretical characterization of these methods

safetyarxiv-cs-lg
27 May 2026
Safety

Monte Carlo Permutation Search

DGX agent

arXiv:2510.06381v2 Announce Type: replace-cross Abstract: We propose Monte Carlo Permutation Search (MCPS), a general-purpose Monte Carlo Tree Search (MCTS) algorithm that improves upon the GRAVE algo

safetyarxiv-cs-ai
27 May 2026
Safety

Multi-Stakeholder LLM Alignment: Decomposing Estimation from Aggregation

DGX agent

arXiv:2605.26878v1 Announce Type: new Abstract: Multi-stakeholder tasks require one output to satisfy users with conflicting preferences. Holistic LLM judges conflate utility estimation and utility ag

safetyarxiv-cs-ai
27 May 2026
Safety

MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation

DGX agent

arXiv:2602.09878v2 Announce Type: replace Abstract: World-model-based imagine-then-act becomes a promising paradigm for robotic manipulation, yet existing approaches typically support either purely im

safetyarxiv-cs-cv
27 May 2026
Safety

Olaf-World: Orienting Latent Actions for Video World Modeling

DGX agent

arXiv:2602.10104v2 Announce Type: replace-cross Abstract: Scaling action-controllable world models is limited by the scarcity of action labels. While latent action learning promises to extract control

safetyarxiv-cs-ai
27 May 2026
Safety

On the Push-Based Asynchronous Federated Learning: A Bias-Correction Aggregation Approach

DGX agent

arXiv:2605.26162v1 Announce Type: cross Abstract: Asynchronous decentralized federated learning (ADFL) eliminates central coordination and global synchronization, making it attractive for large-scale

safetyarxiv-cs-ai
27 May 2026
Safety

On the Role of Inductive Bias in Time-Series Pretraining: A Case Study in Learning Generalizable Representations for Clinical Time Series

DGX agent

arXiv:2605.26194v1 Announce Type: new Abstract: Clinical time-series learning is routinely constrained by small, heterogeneous cohorts and protocol drift, while its downstream use spans both classific

safetyarxiv-cs-lg
27 May 2026
Safety

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks

DGX agent

arXiv:2605.26526v1 Announce Type: new Abstract: Recent defenses for safeguarding open-weight large language models (LLMs) are intended to prevent adversarial usage. Underlying these defenses is an ass

safetyarxiv-cs-lg
27 May 2026
Safety

Over-Alignment vs Over-Fitting: The Role of Feature Learning Strength in Generalization

DGX agent

arXiv:2602.00827v2 Announce Type: replace Abstract: Feature learning strength (FLS), i.e., the inverse of the effective output scaling of a model, plays a critical role in shaping the optimization dyn

safetyarxiv-cs-lg
27 May 2026
Safety

Pair-In, Pair-Out: Latent Multi-Token Prediction for Efficient LLMs

DGX agent

arXiv:2605.27255v1 Announce Type: cross Abstract: Long chain-of-thought reasoning has made autoregressive decoding the dominant inference cost of modern large language models. Existing methods target

safetyarxiv-cs-ai
27 May 2026
Safety

PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization

DGX agent

arXiv:2507.16679v3 Announce Type: replace-cross Abstract: In-Context Learning has shown great potential for aligning Large Language Models (LLMs) with human values, helping reduce harmful outputs and

safetyarxiv-cs-ai
27 May 2026
Safety

Position: Machine Learning for Heart Transplant Allocation Policy Optimization Should Account for Incentives

DGX agent

arXiv:2602.04990v3 Announce Type: replace Abstract: The allocation of scarce donor organs constitutes one of the most consequential algorithmic challenges in healthcare. While the field is rapidly tra

safetyarxiv-cs-lg
27 May 2026
Safety

PyCAT4: A Hierarchical Vision Transformer-based Framework for 3D Human Pose Estimation

DGX agent

arXiv:2508.02806v3 Announce Type: replace Abstract: Recently, a significant improvement in the accuracy of 3D human pose estimation has been achieved by combining convolutional neural networks (CNNs)

safetyarxiv-cs-cv
27 May 2026
Safety

Quantized Keys Steal Attention: Bias Correction for KV-Cache Compression in Video Diffusion

DGX agent

arXiv:2605.26266v1 Announce Type: cross Abstract: Chunk-wise autoregressive video diffusion models rely on a KV cache of previously generated chunks to avoid redundant computation, but this cache quic

safetyarxiv-cs-ai
27 May 2026
Safety

Real Images, Worse Judgments: Evaluating Vision-Language Models on Concreteness and Imagery

DGX agent

arXiv:2605.27315v1 Announce Type: new Abstract: Visual inputs are often assumed to improve language understanding in multimodal models. We examine this assumption by asking whether vision-language mod

safetyarxiv-cs-cl
27 May 2026
Safety

Rethinking the Trust Region in LLM Reinforcement Learning

DGX agent

arXiv:2602.04879v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a cornerstone for fine-tuning Large Language Models (LLMs), with Proximal Policy Optimization (PPO) ser

safetyarxiv-cs-ai
27 May 2026
Safety

Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective

DGX agent

arXiv:2605.26441v1 Announce Type: cross Abstract: This paper addresses the challenging task of weakly-supervised video temporal grounding. Existing approaches are generally based on the moment proposa

safetyarxiv-cs-ai
27 May 2026
Safety

RICE-PO: Turning Retrieval Interactions into Credit Signals for Reasoning Agents

DGX agent

arXiv:2605.26352v1 Announce Type: new Abstract: Retrieval is increasingly moving from one-shot matching toward interactive reasoning, where language agents iteratively inspect evidence, reformulate qu

safetyarxiv-cs-cl
27 May 2026
Safety

Sample Complexity of Policy Gradient for Log-Growth Control

DGX agent

arXiv:2605.26640v1 Announce Type: cross Abstract: We study the sample complexity of policy gradient for log-growth control -- the problem of learning, from observed state transitions, a feedback gain

safetyarxiv-cs-lg
27 May 2026
Safety

Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization

DGX agent

arXiv:2605.26282v1 Announce Type: new Abstract: Model-based reinforcement learning (RL) can be effectively supported at scale through the use of world models. However, in practice, scaling such approa

safetyarxiv-cs-lg
27 May 2026
Safety

SCENT: Aligning Mass Spectra with Molecular Structure for Olfactory Perception

DGX agent

arXiv:2605.27009v1 Announce Type: new Abstract: Predicting human olfactory perception from molecular structure has seen remarkable progress, yet these approaches require explicit chemical structure at

safetyarxiv-cs-lg
27 May 2026
Safety

SCKAN: Structural Consensus-based KAN Prototype Learning for Semi-Supervised Pancreas Segmentation

DGX agent

arXiv:2605.27032v1 Announce Type: new Abstract: Accurate pancreas segmentation is critical for early cancer diagnosis, where annotation scarcity necessitates Semi-Supervised Learning (SSL). However, d

safetyarxiv-cs-cv
27 May 2026
Safety

Securing Multi-Agent Systems Against Corruptions via Node Contribution Backpropagation

DGX agent

arXiv:2510.19420v2 Announce Type: replace-cross Abstract: Multi-Agent Systems (MAS) have become a prevalent paradigm for Large Language Model (LLM) applications. However, the complex multi-agent desig

safetyarxiv-cs-ai
27 May 2026
Safety

Self-Improvement Imitation with Biologically Guided Search for Protein Design Under Oracle Budgets

DGX agent

arXiv:2605.26690v1 Announce Type: cross Abstract: Protein sequence optimization under tight oracle budgets requires methods that explore vast combinatorial spaces while making each evaluation informat

safetyarxiv-cs-ai
27 May 2026
Safety

Signal-to-Noise Ratio and Sample Size Govern Representational Alignment in Neural Networks

DGX agent

arXiv:2605.26973v1 Announce Type: cross Abstract: Neural networks are known to develop latent representations that are aligned, namely structurally similar across networks trained with different archi

safetyarxiv-cs-lg
27 May 2026
Safety

SketchAssist: A Practical Assistant for Semantic Edits and Precise Local Redrawing

DGX agent

arXiv:2512.14140v2 Announce Type: replace Abstract: Sketch editing requires jointly handling high-level semantic changes and precise local redrawing, a combination that is particularly challenging for

safetyarxiv-cs-cv
27 May 2026
Safety

SL-BiLEM: Structured Learnable Behavior-in-the-Loop Epidemic Modeling for Forecasting and Policy Evaluation

DGX agent

arXiv:2605.26704v1 Announce Type: cross Abstract: Epidemic forecasting faces a fundamental challenge: human behavior dynamically responds to disease spread, creating feedback loops that induce distrib

safetyarxiv-cs-ai
27 May 2026
Safety

Spectral Principal Paths: A Spectral Perspective on Linear Representation Formation in LLMs

DGX agent

arXiv:2506.08543v3 Announce Type: replace Abstract: High-level representations have become a central focus in enhancing AI transparency and control, shifting attention from individual neurons or circu

safetyarxiv-cs-cv
27 May 2026
Safety

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training

DGX agent

arXiv:2605.26606v1 Announce Type: cross Abstract: Reinforcement learning (RL) is the dominant paradigm for post-training large language models. However, in the online, on-policy setting, rollout gener

safetyarxiv-cs-ai
27 May 2026
Safety

SQARL: A Size-Agnostic Reinforcement Learning approach for Circuit Allocation in Distributed Quantum Architectures

DGX agent

arXiv:2605.27027v1 Announce Type: new Abstract: The scaling of quantum processors is currently limited by technical challenges such as decoherence and cross-talk. As the number of qubits grows, interf

safetyarxiv-cs-lg
27 May 2026
Safety

StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

DGX agent

arXiv:2605.27140v1 Announce Type: new Abstract: Reinforcement learning for multi-turn agents suffers from a credit-assignment mismatch: rewards are sparse and trajectory-level, while success often hin

safetyarxiv-cs-ai
27 May 2026
Safety

TAGRPO: Boosting GRPO on Image-to-Video Generation with Direct Trajectory Alignment

DGX agent

arXiv:2601.05729v2 Announce Type: replace Abstract: Recent studies have demonstrated the efficacy of integrating Group Relative Policy Optimization (GRPO) into flow matching models, particularly for t

safetyarxiv-cs-cv
27 May 2026
Safety

The Coverage Illusion: From Pre-retrieval Routing Failure to Post-retrieval Cascades in a Production RAG System

DGX agent

arXiv:2605.27220v1 Announce Type: new Abstract: In modern RAG pipelines, query augmentation methods such as HyDE and query expansion are applied to every query, resulting in substantial LLM inference

safetyarxiv-cs-cl
27 May 2026
← Previous
1…157158159160161…260
Next →