AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,435 results
Safety

Rewarding Structural Conformance of Reasoning using Process Mining

DGX agent

arXiv:2510.25065v3 Announce Type: replace Abstract: Recent advances in sparse reward policy gradient methods have enabled effective reinforcement learning (RL)-based language model post-training. Howe

safetyarxiv-cs-ai
26 May 2026
Safety
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Right-Sizing Communication and Recommendation Set Size in AI-Assisted Search

DGX agent

arXiv:2605.23944v1 Announce Type: new Abstract: We model the interaction between a user and an AI driven recommendation system. The user initiates the process by conveying preference information throu

safetyarxiv-cs-ai
26 May 2026
Safety

RiskBridge: Turning CVEs into Business-Aligned Patch Priorities

DGX agent

arXiv:2601.06201v2 Announce Type: replace-cross Abstract: Enterprises are confronted with an unprecedented escalation in cybersecurity vulnerabilities, with thousands of new CVEs disclosed each month.

safetyarxiv-cs-ai
26 May 2026
Safety

SEAL: Synergistic Co-Evolution of Agents and Learning Environments

DGX agent

arXiv:2605.24426v1 Announce Type: new Abstract: Large Language Model (LLM) agents are increasingly improved through interaction, yet most self-evolution methods adapt either the policy or the learning

safetyarxiv-cs-cl
26 May 2026
Safety

Selective Latent Thinking: Adaptive Compression of LLM Reasoning Chains

DGX agent

arXiv:2605.25745v1 Announce Type: new Abstract: Explicit chain-of-thought (CoT) reasoning substantially improves the reasoning ability of large language models (LLMs), but incurs high inference cost d

safetyarxiv-cs-cl
26 May 2026
Safety

SliceWorld: A Predictive and Controllable World-State Model for CT Report Generation

DGX agent

arXiv:2605.24371v1 Announce Type: cross Abstract: CT report generation (CTRG) requires models to summarize three-dimensional anatomical context and pathological findings from hundreds of axial slices.

safetyarxiv-cs-cl
26 May 2026
Safety

Smoother Action Chunking Flow Policy via Prior-Corrected Orthogonal Trust-Region Guidance

DGX agent

arXiv:2605.24433v1 Announce Type: cross Abstract: Flow-matching robot policies commonly use action-chunking inference for efficient closed-loop control, but chunk boundaries can introduce discontinuou

safetyarxiv-cs-lg
26 May 2026
Model Releases

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models

DGX agent

arXiv:2506.18543v2 Announce Type: replace-cross Abstract: The rapid proliferation of Large Language Models (LLMs) has heightened concerns regarding their exposure to jailbreak attacks, which craft adv

model-releasesarxiv-cs-ai
26 May 2026
Safety

SpecAlign: A Semantic Alignment Framework for SystemVerilog Assertion Generation

DGX agent

arXiv:2605.25181v1 Announce Type: new Abstract: Existing Large Language Model (LLM) approaches to SystemVerilog Assertion (SVA) generation primarily focus on syntactic validity and formal verification

safetyarxiv-cs-ai
26 May 2026
Safety

StakeBench: Evaluating Language Understanding Grounded in Market Commitment

DGX agent

arXiv:2605.26074v1 Announce Type: cross Abstract: Existing financial NLP benchmarks often rely on labels supplied by outside observers, measuring how language is perceived rather than what speakers ha

safetyarxiv-cs-ai
26 May 2026
Safety

STaT: Resolving Shape Distortion in Non-Stationary Time Series via Tri-Modal Synergy

DGX agent

arXiv:2605.25943v1 Announce Type: new Abstract: Recent research in time series forecasting frequently investigates the integration of textual and visual modalities with numerical models to better navi

safetyarxiv-cs-lg
26 May 2026
Safety

Stop Comparing LLM Agents Without Disclosing the Harness

DGX agent

arXiv:2605.23950v1 Announce Type: new Abstract: This position paper argues that, for long-horizon tasks evaluated across models with comparable frontier capability, the agent execution harness, namely

safetyarxiv-cs-ai
26 May 2026
Safety

Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent Games

DGX agent

arXiv:2605.04906v2 Announce Type: replace Abstract: While Large Language Models (LLMs) excel in certain reasoning tasks, they struggle in multi-agent games where the final outcome depends on the joint

safetyarxiv-cs-ai
26 May 2026
Safety

Subspace-Guided Semantic and Topological Invariant Registration for Annotation-Free Ultrasound Plane Quality Control

DGX agent

arXiv:2605.25396v1 Announce Type: cross Abstract: Reliable quality control (QC) of ultrasound images is essential for both real-time acquisition guidance and retrospective clinical audit, yet existing

safetyarxiv-cs-ai
26 May 2026
Safety

Summoning the Oracle to Slay It: Mitigating Look-Ahead Bias in Financial Backtesting with Large Language Models

DGX agent

arXiv:2605.24564v1 Announce Type: new Abstract: Backtesting large language models (LLMs) on historical financial data is unreliable because pre-training cuts off after the events happened. An LLM trai

safetyarxiv-cs-ai
26 May 2026
Safety

TapSampling: Inference-Time Sampling with a Task-Progress-Understanding Verifier for Robotic Manipulation

DGX agent

arXiv:2605.25547v1 Announce Type: new Abstract: Existing embodied control research demonstrates remarkable performance improvements by scaling training data and model size. We instead explore inferenc

safetyarxiv-cs-ro
26 May 2026
Safety

Task-Aligned Self-Supervised Learning for Medical Image Analysis: A Systematic Review and Practical Design Guidelines

DGX agent

arXiv:2605.23995v1 Announce Type: cross Abstract: Self-supervised learning (SSL) has emerged as a promising paradigm for addressing the annotation bottleneck in medical imaging by learning representat

safetyarxiv-cs-ai
26 May 2026
Safety

The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible

DGX agent

arXiv:2605.25739v1 Announce Type: new Abstract: We prove that no reinforcement learning policy with confidence-gated autonomy can simultaneously achieve maximum helpfulness, optimal calibration, and f

safetyarxiv-cs-lg
26 May 2026
Safety

The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth

DGX agent

arXiv:2605.24856v1 Announce Type: cross Abstract: Concept formation in transformer language models is depth-extended, not a single-layer event: concepts emerge gradually across a contiguous region of

safetyarxiv-cs-ai
26 May 2026
Safety

The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks

DGX agent

arXiv:2602.16340v3 Announce Type: replace Abstract: We study the implicit bias of momentum-based optimizers on smooth homogeneous models. We show that extit{momentum steepest descent} algorithms like

safetyarxiv-cs-lg
26 May 2026
Safety

The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes

DGX agent

arXiv:2605.11182v2 Announce Type: replace Abstract: On-policy distillation (OPD) and on-policy self-distillation (OPSD) have emerged as promising post-training methods for large language models, offer

safetyarxiv-cs-ai
26 May 2026
Safety

The Path Matters: Learning a Token-Commitment Policy for Diffusion Language Models

DGX agent

arXiv:2605.24697v1 Announce Type: cross Abstract: Diffusion large language models promise faster generation by refining many token positions in parallel, but this parallelism introduces a hidden contr

safetyarxiv-cs-ai
26 May 2026
Safety

TopoAlign: Topology-Aware Visual Representation Alignment

DGX agent

arXiv:2605.25541v1 Announce Type: cross Abstract: Neural networks encode inputs as high-dimensional vectors, known as representations, that capture how models process data by encoding task-relevant st

safetyarxiv-cs-ai
26 May 2026
Safety

Toward Reliable Design of LLM-Enabled Agentic Workflows: Optimizing Latency-Reliability-Cost Tradeoffs

DGX agent

arXiv:2605.23929v1 Announce Type: new Abstract: Modern AI systems increasingly rely on workflows composed of multiple interacting agents, some powered by large language models (LLMs) and others by con

safetyarxiv-cs-ai
26 May 2026
Safety

Towards Cognitively-Faithful Decision-Making Models to Improve AI Alignment

DGX agent

arXiv:2509.04445v2 Announce Type: replace Abstract: Recent AI trends seek to align AI models to learned human-centric objectives, such as personal preferences, utility, or societal values. Using stand

safetyarxiv-cs-lg
26 May 2026
Safety

Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content

DGX agent

arXiv:2509.12672v2 Announce Type: replace Abstract: The volume of machine-generated content online has grown dramatically due to the widespread use of Large Language Models (LLMs), leading to new chal

safetyarxiv-cs-cl
26 May 2026
Safety

Towards Low-Gravity Planetary Exploration using Reinforcement Learning for Walking, Jumping, and In-flight Attitude Control

DGX agent

arXiv:2605.24643v1 Announce Type: new Abstract: This paper presents reinforcement learning (RL) policies for dynamic quadrupedal locomotion in planetary exploration scenarios. Building on a taskoptimi

safetyarxiv-cs-ro
26 May 2026
Safety

Towards the Connection between Activation Sparsity and Flat Minima

DGX agent

arXiv:2605.25612v1 Announce Type: cross Abstract: The observation that activation sparsity emerges in MLP blocks of standardly trained Transformers offers an opportunity to drastically reduce computat

safetyarxiv-cs-ai
26 May 2026
Safety

Towards Understanding Adam Convergence on Highly Degenerate Polynomials

DGX agent

arXiv:2603.09581v2 Announce Type: replace Abstract: Adam is a widely used optimization algorithm in deep learning, yet the specific class of objective functions where it exhibits inherent advantages r

safetyarxiv-cs-lg
26 May 2026
Safety

Trait-Aware Policy Optimization for Autoregressive Multi-Trait Essay Scoring

DGX agent

arXiv:2605.25731v1 Announce Type: new Abstract: Multi-trait essay scoring aims to provide fine-grained evaluation of writing quality across multiple dimensions. However, how to effectively post-train

safetyarxiv-cs-cl
26 May 2026
Safety

Trust-Aware Joint Feature-Prediction Discrepancy for Robust Domain Adaptation

DGX agent

arXiv:2605.25119v1 Announce Type: cross Abstract: Domain adaptation aims to mitigate performance degradation caused by distribution shifts between a labeled source domain and an unlabeled or sparsely

safetyarxiv-cs-ai
26 May 2026
Safety

Uncertainty-DTW for Sequences and Visual Tokens

DGX agent

arXiv:2605.25110v1 Announce Type: cross Abstract: Aligning structured data is a fundamental problem in computer vision and machine learning, underlying tasks such as time series analysis, human action

safetyarxiv-cs-ai
26 May 2026
Safety

Unifying Value Alignment and Assignment in Cross-Domain Offline Reinforcement Learning with Heterogeneous Datasets

DGX agent

arXiv:2605.24862v1 Announce Type: new Abstract: Cross-domain offline reinforcement learning (RL) aims to learn a policy in the target domain with a limited target domain dataset and a source domain da

safetyarxiv-cs-lg
26 May 2026
Safety

UniRank: End-to-End Domain-Specific Reranking of Hybrid Text-Image Candidates

DGX agent

arXiv:2603.29897v2 Announce Type: replace-cross Abstract: Reranking is a critical component in many information retrieval pipelines. Despite remarkable progress in text-only settings, multimodal reran

safetyarxiv-cs-ai
26 May 2026
Safety

Universal Boosts, Specific Suppressors: Sparse Autoencoder Steering of Medical Vision-Language Models

DGX agent

arXiv:2605.24977v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) often hallucinate findings when generating chest X-ray reports: they fabricate findings that are not present in

safetyarxiv-cs-cl
26 May 2026
Safety

VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding

DGX agent

arXiv:2605.25952v1 Announce Type: cross Abstract: Despite the remarkable progress achieved by recent efficient methods in accelerating multimodal understanding, they still suffer from noticeable perfo

safetyarxiv-cs-ai
26 May 2026
Safety

Vision-Guided Outdoor Flight and Obstacle Evasion via Reinforcement Learning

DGX agent

arXiv:2605.24449v1 Announce Type: cross Abstract: Although quadcopters boast impressive traversal capabilities enabled by their omnidirectional maneuverability, the need for continuous pilot control i

safetyarxiv-cs-lg
26 May 2026
Safety

What Gets Cited: Competitive GEO in AI Answer Engines

DGX agent

arXiv:2605.25517v1 Announce Type: new Abstract: AI answer engines generate answers from retrieved pages but cite only a few sources. This makes visibility depend not just on ranking, but on being cite

safetyarxiv-cs-ai
26 May 2026
Safety

When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs

DGX agent

arXiv:2605.24202v1 Announce Type: new Abstract: Multi-agent LLM workflows route inference through specialized roles to lift end-task accuracy, but jointly training those roles with reinforcement learn

safetyarxiv-cs-ai
26 May 2026
Safety

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards

DGX agent

arXiv:2605.25864v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable advancements in reasoning capabilities empowered by Reinforcement Learning with Verifiable Rewar

safetyarxiv-cs-cl
26 May 2026
Safety

X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models

DGX agent

arXiv:2605.25044v1 Announce Type: new Abstract: Learning universal policies from cross-embodied data remains a fundamental challenge in robotics. Although Vision-Language-Action (VLA) models are pre-t

safetyarxiv-cs-ro
26 May 2026
Safety

XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation

DGX agent

arXiv:2510.06672v3 Announce Type: replace Abstract: Reinforcement learning algorithms such as GRPO have driven recent advances in large language model (LLM) reasoning. While scaling the number of roll

safetyarxiv-cs-lg
26 May 2026
Safety

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding

DGX agent

arXiv:2605.05997v2 Announce Type: replace Abstract: Dynamic spatial reasoning from monocular video is essential for bridging visual intelligence and the physical world, yet remains challenging for vis

safetyarxiv-cs-cv
25 May 2026
Safety

ALIVE: Awakening LLM Reasoning via Adversarial Learning and Instructive Verbal Evaluation

DGX agent

arXiv:2602.05472v2 Announce Type: replace Abstract: The quest for expert-level reasoning in Large Language Models (LLMs) has been hampered by a persistent extit{reward bottleneck}: traditional reinfor

safetyarxiv-cs-ai
25 May 2026
Safety

ARMS: Automatic Reward Shaping for Sparse-Reward Multi-Agent Reinforcement Learning

DGX agent

arXiv:2605.23562v1 Announce Type: cross Abstract: Sparse rewards are a major bottleneck in multi-agent reinforcement learning (MARL), where simultaneous learning induces non-stationarity and makes rew

safetyarxiv-cs-ai
25 May 2026
Safety

Assessing Predictive Models for Fairness Based on Movement Patterns

DGX agent

arXiv:2605.23234v1 Announce Type: new Abstract: Assessing the spatial fairness of predictive models involves establishing whether they are statistically penalizing (favoring) individuals associated wi

safetyarxiv-cs-lg
25 May 2026
Safety

B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation

DGX agent

arXiv:2605.23500v1 Announce Type: new Abstract: Segmentation is a fundamental task in computer vision, underpinning pixel-level scene understanding and serving as a cornerstone for applications rangin

safetyarxiv-cs-cv
25 May 2026
Safety

Beyond Binary Edits Robust Multimodal Knowledge Editing with Adversarial Subspace Alignment

DGX agent

arXiv:2605.23780v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) need efficient mechanisms to update knowledge without degrading existing capabilities. While intrinsic multimod

safetyarxiv-cs-ai
25 May 2026
← Previous
1…162163164165166…260
Next →