AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
All
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,600 results
Safety

Preference-Agile Multi-Objective Optimization for Real-time Vehicle Dispatching

DGX agent

arXiv:2604.10664v1 Announce Type: new Abstract: Multi-objective optimization (MOO) has been widely studied in literature because of its versatility in human-centered decision making in real-life appli

safetyarxiv-cs-ai
14 Apr 2026
Safety
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Premier: Personalized Preference Modulation with Learnable User Embedding in Text-to-Image Generation

DGX agent

arXiv:2603.20725v2 Announce Type: replace Abstract: Text-to-image generation has advanced rapidly, yet it still struggles to capture the nuanced user preferences. Existing approaches typically rely on

safetyarxiv-cs-cv
14 Apr 2026
Safety

Principles Do Not Apply Themselves: A Hermeneutic Perspective on AI Alignment

DGX agent

arXiv:2604.10673v1 Announce Type: new Abstract: AI alignment is often framed as the task of ensuring that an AI system follows a set of stated principles or human preferences, but general principles r

safetyarxiv-cs-ai
14 Apr 2026
Safety

PRISM Risk Signal Framework: Hierarchy-Based Red Lines for AI Behavioral Risk

DGX agent

arXiv:2604.11070v1 Announce Type: new Abstract: Current approaches to AI safety define red lines at the case level: specific prompts, specific outputs, specific harms. This paper argues that red lines

safetyarxiv-cs-ai
14 Apr 2026
Safety

Prompt Injection as Role Confusion

DGX agent

arXiv:2603.12277v3 Announce Type: replace-cross Abstract: Language models remain vulnerable to prompt injection attacks despite extensive safety training. We trace this failure to role confusion: mode

safetyarxiv-cs-ai
14 Apr 2026
Safety

Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation

DGX agent

arXiv:2604.10030v1 Announce Type: new Abstract: Video diffusion models have achieved remarkable progress in generating high-quality videos. However, these models struggle to represent the temporal suc

safetyarxiv-cs-cv
14 Apr 2026
Safety

ProUIE: A Macro-to-Micro Progressive Learning Method for LLM-based Universal Information Extraction

DGX agent

arXiv:2604.10633v1 Announce Type: new Abstract: LLM-based universal information extraction (UIE) methods often rely on additional information beyond the original training data, which increases trainin

safetyarxiv-cs-cl
14 Apr 2026
Safety

Proximal Supervised Fine-Tuning

DGX agent

arXiv:2508.17784v2 Announce Type: replace-cross Abstract: Supervised fine-tuning (SFT) of foundation models often leads to poor generalization, where prior capabilities deteriorate after tuning on new

safetyarxiv-cs-ai
14 Apr 2026
Safety

QFS-Composer: Query-focused summarization pipeline for less resourced languages

DGX agent

arXiv:2604.10687v1 Announce Type: new Abstract: Large language models (LLMs) demonstrate strong performance in text summarization, yet their effectiveness drops significantly across languages with res

safetyarxiv-cs-cl
14 Apr 2026
Safety

QShield: Securing Neural Networks Against Adversarial Attacks using Quantum Circuits

DGX agent

arXiv:2604.10933v1 Announce Type: cross Abstract: Deep neural networks remain highly vulnerable to adversarial perturbations, limiting their reliability in security- and safety-critical applications.

safetyarxiv-cs-ai
14 Apr 2026
Safety

Quantifying the Climate Risk of Generative AI: Region-Aware Carbon Accounting with G-TRACE and the AI Sustainability Pyramid

DGX agent

arXiv:2511.04776v2 Announce Type: replace-cross Abstract: Generative Artificial Intelligence (GenAI) represents a rapidly expanding digital infrastructure whose energy demand and associated CO2 emissi

safetyarxiv-cs-cl
14 Apr 2026
Safety

RAG-KT: Cross-platform Explainable Knowledge Tracing with Multi-view Fusion Retrieval Generation

DGX agent

arXiv:2604.10960v1 Announce Type: new Abstract: Knowledge Tracing (KT) infers a student's knowledge state from past interactions to predict future performance. Conventional Deep Learning (DL)-based KT

safetyarxiv-cs-ai
14 Apr 2026
Safety

RealSR-R1: Reinforcement Learning for Real-World Image Super-Resolution with Vision-Language Chain-of-Thought

DGX agent

arXiv:2506.16796v4 Announce Type: replace Abstract: Real-World Image Super-Resolution is one of the most challenging task in image restoration. However, existing methods struggle with an accurate unde

safetyarxiv-cs-cv
14 Apr 2026
Safety

Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging

DGX agent

arXiv:2604.11399v1 Announce Type: cross Abstract: Multimodal adaptation equips large language models (LLMs) with perceptual capabilities, but often weakens the reasoning ability inherited from languag

safetyarxiv-cs-cl
14 Apr 2026
Safety

Rebooting Microreboot: Architectural Support for Safe, Parallel Recovery in Microservice Systems

DGX agent

arXiv:2604.09963v1 Announce Type: cross Abstract: Microreboot enables fast recovery by restarting only the failing component, but in modern microservices naive restarts are unsafe: dense dependencies

safetyarxiv-cs-ai
14 Apr 2026
Safety

Regularized Entropy Information Adaptation with Temporal-Awareness Networks for Simultaneous Speech Translation

DGX agent

arXiv:2604.09916v1 Announce Type: new Abstract: Simultaneous Speech Translation (SimulST) requires balancing high translation quality with low latency. Recent work introduced REINA, a method that trai

safetyarxiv-cs-lg
14 Apr 2026
Safety

Reinforcement Learning for Intensity Control: An Application to Choice-Based Network Revenue Management

DGX agent

arXiv:2406.05358v3 Announce Type: replace Abstract: Intensity control is a class of continuous-time dynamic optimization problems with many important applications in Operations Research including queu

safetyarxiv-cs-lg
14 Apr 2026
Safety

Relative Entropy Pathwise Policy Optimization

DGX agent

arXiv:2507.11019v4 Announce Type: replace Abstract: Score-function based methods for policy learning, such as REINFORCE and PPO, have delivered strong results in game-playing and robotics, yet their h

safetyarxiv-cs-lg
14 Apr 2026
Safety

Reliable and Real-Time Highway Trajectory Planning via Hybrid Learning-Optimization Frameworks

DGX agent

arXiv:2508.04436v2 Announce Type: replace Abstract: Autonomous highway driving involves high-speed safety risks due to limited reaction time, where rare but dangerous events may lead to severe consequ

safetyarxiv-cs-ro
14 Apr 2026
Safety

Reliable Evaluation Protocol for Low-Precision Retrieval

DGX agent

arXiv:2508.03306v4 Announce Type: replace-cross Abstract: Lowering the numerical precision of model parameters and computations is widely adopted to improve the efficiency of retrieval systems. Howeve

safetyarxiv-cs-ai
14 Apr 2026
Safety

Resilient Write: A Six-Layer Durable Write Surface for LLM Coding Agents

DGX agent

arXiv:2604.10842v1 Announce Type: cross Abstract: LLM-powered coding agents increasingly rely on tool-use protocols such as the Model Context Protocol~(MCP) to read and write files on a developer's wo

safetyarxiv-cs-ai
14 Apr 2026
Safety

Rethinking LLM Watermark Detection in Black-Box Settings: A Non-Intrusive Third-Party Framework

DGX agent

arXiv:2603.14968v2 Announce Type: replace-cross Abstract: While watermarking serves as a critical mechanism for LLM provenance, existing secret-key schemes tightly couple detection with injection, req

safetyarxiv-cs-cl
14 Apr 2026
Safety

Rethinking Token-Level Credit Assignment in RLVR: A Polarity-Entropy Analysis

DGX agent

arXiv:2604.11056v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has substantially improved the reasoning ability of Large Language Models (LLMs). However, its s

safetyarxiv-cs-ai
14 Apr 2026
Safety

Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference

DGX agent

arXiv:2604.11496v1 Announce Type: cross Abstract: Dual-encoder Vision-Language Models (VLMs) such as CLIP are often characterized as bag-of-words systems due to their poor performance on compositional

safetyarxiv-cs-cl
14 Apr 2026
Safety

Revisiting Epistemic Markers in Confidence Estimation: Can Markers Accurately Reflect Large Language Models' Uncertainty?

DGX agent

arXiv:2505.24778v3 Announce Type: replace Abstract: As large language models (LLMs) are increasingly used in high-stakes domains, accurately assessing their confidence is crucial. Humans typically exp

safetyarxiv-cs-cl
14 Apr 2026
Safety

Risk Awareness Injection: Calibrating Vision-Language Models for Safety without Compromising Utility

DGX agent

arXiv:2602.03402v3 Announce Type: replace Abstract: Vision language models (VLMs) extend the reasoning capabilities of large language models (LLMs) to cross-modal settings, yet remain highly vulnerabl

safetyarxiv-cs-ai
14 Apr 2026
Safety

RoboStereo: Dual-Tower 4D Embodied World Models for Unified Policy Optimization

DGX agent

arXiv:2603.12639v2 Announce Type: replace Abstract: Scalable Embodied AI faces fundamental constraints due to prohibitive costs and safety risks of real-world interaction. While Embodied World Models

safetyarxiv-cs-cv
14 Apr 2026
Safety

Robust Real-Time Coordination of CAVs: A Distributed Optimization Framework under Uncertainty

DGX agent

arXiv:2508.21322v2 Announce Type: replace Abstract: Achieving both safety guarantees and real-time performance in cooperative vehicle coordination remains a fundamental challenge, particularly in dyna

safetyarxiv-cs-ro
14 Apr 2026
Safety

Safe Human-to-Humanoid Motion Imitation Using Control Barrier Functions

DGX agent

arXiv:2604.11447v1 Announce Type: new Abstract: Ensuring operational safety is critical for human-to-humanoid motion imitation. This paper presents a vision-based framework that enables a humanoid rob

safetyarxiv-cs-ro
14 Apr 2026
Safety

SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering

DGX agent

arXiv:2508.11290v3 Announce Type: replace Abstract: LLMs increasingly exhibit over-refusal behavior, where safety mechanisms cause models to reject benign instructions that seemingly resemble harmful

safetyarxiv-cs-cl
14 Apr 2026
Safety

Safety Guarantees in Zero-Shot Reinforcement Learning for Cascade Dynamical Systems

DGX agent

arXiv:2604.10429v1 Announce Type: new Abstract: This paper considers the problem of zero-shot safety guarantees for cascade dynamical systems. These are systems where a subset of the states (the inner

safetyarxiv-cs-ai
14 Apr 2026
Safety

SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting

DGX agent

arXiv:2604.10688v1 Announce Type: cross Abstract: On-policy reinforcement learning has become the dominant paradigm for reasoning alignment in large language models, yet its sparse, outcome-level rewa

safetyarxiv-cs-ai
14 Apr 2026
Safety

SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents

DGX agent

arXiv:2604.07791v2 Announce Type: replace Abstract: Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have demonstrated significant potential in single-turn reasoning tasks. Wit

safetyarxiv-cs-ai
14 Apr 2026
Safety

See Fair, Speak Truth: Equitable Attention Improves Grounding and Reduces Hallucination in Vision-Language Alignment

DGX agent

arXiv:2604.09749v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) frequently hallucinate objects that are absent from the visual input, often because attention during decoding i

safetyarxiv-cs-cv
14 Apr 2026
Safety

Self-Organizing Dual-Buffer Adaptive Clustering Experience Replay (SODACER) for Safe Reinforcement Learning in Optimal Control

DGX agent

arXiv:2601.06540v2 Announce Type: replace-cross Abstract: This paper proposes a novel reinforcement learning framework, named Self-Organizing Dual-buffer Adaptive Clustering Experience Replay (SODACER

safetyarxiv-cs-ai
14 Apr 2026
Safety

SemaClaw: A Step Towards General-Purpose Personal AI Agents through Harness Engineering

DGX agent

arXiv:2604.11548v1 Announce Type: new Abstract: The rise of OpenClaw in early 2026 marks the moment when millions of users began deploying personal AI agents into their daily lives, delegating tasks r

safetyarxiv-cs-ai
14 Apr 2026
Safety

Sense Less, Infer More: Agentic Multimodal Transformers for Edge Medical Intelligence

DGX agent

arXiv:2604.10404v1 Announce Type: cross Abstract: Edge-based multimodal medical monitoring requires models that balance diagnostic accuracy with severe energy constraints. Continuous acquisition of EC

safetyarxiv-cs-lg
14 Apr 2026
Safety

SHE: Stepwise Hybrid Examination Reinforcement Learning Framework for E-commerce Search Relevance

DGX agent

arXiv:2510.07972v3 Announce Type: replace Abstract: Query-product relevance prediction is vital for AI-driven e-commerce, yet current LLM-based approaches face a dilemma: SFT and DPO struggle with lon

safetyarxiv-cs-ai
14 Apr 2026
Safety

Shuffling the Data, Stretching the Step-size: Sharper Bias in constant step-size SGD

DGX agent

arXiv:2604.10373v1 Announce Type: cross Abstract: From adversarial robustness to multi-agent learning, many machine learning tasks can be cast as finite-sum min-max optimization or, more generally, as

safetyarxiv-cs-lg
14 Apr 2026
Safety

SIMPLER: H&E-Informed Representation Learning for Structured Illumination Microscopy

DGX agent

arXiv:2604.10334v1 Announce Type: new Abstract: Structured Illumination Microscopy (SIM) enables rapid, high-contrast optical sectioning of fresh tissue without staining or physical sectioning, making

safetyarxiv-cs-cv
14 Apr 2026
Safety

Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

DGX agent

arXiv:2604.10674v1 Announce Type: cross Abstract: Reinforcement learning (RL) has been widely used to train LLM agents for multi-turn interactive tasks, but its sample efficiency is severely limited b

safetyarxiv-cs-ai
14 Apr 2026
Safety

SLALOM: Simulation Lifecycle Analysis via Longitudinal Observation Metrics for Social Simulation

DGX agent

arXiv:2604.11466v1 Announce Type: cross Abstract: Large Language Model (LLM) agents offer a potentially-transformative path forward for generative social science but face a critical crisis of validity

safetyarxiv-cs-ai
14 Apr 2026
Safety

Some opportunities you just can't turn down, like when @GaryMarcus wants to come talk about AI at BugBash. Fewer than 20 tickets left, come …

DGX agent

Gary Marcus, a prominent AI researcher and critic known for his skepticism of current deep learning approaches, was invited to speak at BugBash, an event hosted by Antithesis. The post, shared by Anti

safetygary-marcus--x
14 Apr 2026
Safety

Spatiotemporal-Aware Bit-Flip Injection on DNN-based Advanced Driver Assistance Systems (extended version)

DGX agent

arXiv:2604.03753v2 Announce Type: replace-cross Abstract: Modern advanced driver assistance systems (ADAS) rely on deep neural networks (DNNs) for perception and planning. Since DNNs' parameters resid

safetyarxiv-cs-lg
14 Apr 2026
Safety

Speaking to No One: Ontological Dissonance and the Double Bind of Conversational AI

DGX agent

arXiv:2604.10833v1 Announce Type: cross Abstract: Recent reports indicate that sustained interaction with conversational artificial intelligence (AI) systems can, in a small subset of users, contribut

safetyarxiv-cs-ai
14 Apr 2026
Safety

StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation

DGX agent

arXiv:2510.05057v2 Announce Type: replace-cross Abstract: A fundamental challenge in embodied intelligence is developing expressive and compact state representations for efficient world modeling and d

safetyarxiv-cs-cv
14 Apr 2026
Safety

Steered LLM Activations are Non-Surjective

DGX agent

arXiv:2604.09839v1 Announce Type: new Abstract: Activation steering is a popular white-box control technique that modifies model activations to elicit an abstract change in output behavior. It has als

safetyarxiv-cs-ai
14 Apr 2026
Safety

Stop Fixating on Prompts: Reasoning Hijacking and Constraint Tightening for Red-Teaming LLM Agents

DGX agent

arXiv:2604.05549v2 Announce Type: replace Abstract: With the widespread application of LLM-based agents across various domains, their complexity has introduced new security threats. Existing red-team

safetyarxiv-cs-cl
14 Apr 2026
← Previous
1…251252253254255…263
Next →