AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,600 results
5 Aug 2026

AI Assistance Reduces Persistence and Hurts Independent Performance

SafetyDGX agent

arXiv:2604.04721v3 Announce Type: replace Abstract: People often optimize for long-term goals in collaboration: A mentor or companion doesn't just answer questions, but also scaffolds learning, tracks

AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality

SafetyDGX agent

arXiv:2608.03581v1 Announce Type: cross Abstract: AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic u

Aligned in Form, Not in Meaning: The Comprehension - Containment Decoupling of LLM Safety in Low-Resource Bangla Derogatory Speech

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.02941v1 Announce Type: new Abstract: We audit five frontier large language models on native Bangla derogatory speech (gali) across six protocols to test a single hypothesis: Comprehension-C

Another real-world manifestation of the kind of misaligned actions frontier systems developed by leading companies can take to achieve goals…

SafetyDGX agent

Another real-world manifestation of the kind of misaligned actions frontier systems developed by leading companies can take to achieve goals. Current frontier AI models are trained with reinforcement

Attribute-based Undetectable Watermarking for Generative AI Models

SafetyDGX agent

arXiv:2608.03174v1 Announce Type: cross Abstract: Generative AI systems increasingly produce content whose provenance is difficult to verify, motivating watermarking techniques for identifying model-g

BAP-SQL: Budget-Aware Observation Planning for Agentic Text-to-SQL

SafetyDGX agent

arXiv:2608.02876v1 Announce Type: new Abstract: Tool-using agents do not merely consume observations: their actions determine what arrives next. In agentic text-to-SQL, a broad query can spend context

Bimanual Manipulation Within an 8 GB Budget: Zero-Copy Sensing and Quantized ACT on an Entry-Level Jetson

SafetyDGX agent

arXiv:2608.03938v1 Announce Type: new Abstract: Bimanual manipulation policies trained with imitation learning are typically evaluated on workstation or datacenter-class GPUs, leaving the cost of depl

BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?

SafetyDGX agent

arXiv:2608.02867v1 Announce Type: cross Abstract: Although reinforcement learning with verifiable rewards (RLVR) has improved the performance of large language models (LLMs) across a variety of reason

BOW: Training Language Models to Reason Over Plausible Next Words

SafetyDGX agent

arXiv:2506.13502v3 Announce Type: replace Abstract: Next-word prediction (NWP) trains language models against a single observed continuation, even though many contexts admit multiple plausible next wo

Bridging Prediction and Attribution: Identifying Forward and Backward Causal Influence Ranges Using Assimilative Causal Inference

SafetyDGX agent

arXiv:2510.21889v2 Announce Type: replace-cross Abstract: Causal inference identifies cause-and-effect relationships between variables. While traditional approaches rely on data to reveal causal links

CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generation

SafetyDGX agent

arXiv:2608.03046v1 Announce Type: new Abstract: Text-to-video (T2V) diffusion transformers (DiTs) are trained with detailed video captions, whereas inference often relies on user prompts rewritten by

CausalOPD: First-Wrong-Step Supervision for Distilling Causal Chain Reasoning

SafetyDGX agent

arXiv:2608.03673v1 Announce Type: new Abstract: Many critical reasoning tasks, including clinical diagnosis, legal judgment, and industrial fault diagnosis, require step-dependent causal chains in whi

Channel-wise Dynamic Knowledge Distillation via Adaptive Sample Generation for Action Recognition

SafetyDGX agent

arXiv:2608.03100v1 Announce Type: new Abstract: Knowledge Distillation (KD) offers a promising yet underexplored path for compressing large action recognition models. However, existing KD methods suff

Clinically-Grounded Hierarchical Classification for Consistent Chest X-ray Interpretation

SafetyDGX agent

arXiv:2608.03016v1 Announce Type: new Abstract: Accurate chest X-ray interpretation is inherently hierarchical. Clinical decisions depend not only on what abnormality is present but where it is situat

CLIP4VI-ReID: Learning Modality-shared Representations via CLIP Semantic Bridge for Visible-Infrared Person Re-identification

SafetyDGX agent

arXiv:2511.10309v2 Announce Type: replace Abstract: This paper proposes a novel CLIP-driven modality-shared representation learning network named CLIP4VI-ReID for VI-ReID task, which consists of Text

⚠️⚠️⚠️ Constitutional AI is not working. Aligning LLMs is not working. We need a different approach. If society doesn’t place its bets diffe…

SafetyDGX agent

⚠️⚠️⚠️ Constitutional AI is not working. Aligning LLMs is not working. We need a different approach. If society doesn’t place its bets differently, we are screwed. Anthropic's Mythos created fake iden

Continue or Replan? Bernoulli-Continuation Policy Learning for Adaptive Horizon Execution

SafetyDGX agent

arXiv:2608.03483v1 Announce Type: cross Abstract: Existing chunk-based Vision-Language-Action (VLA) models execute a fixed number of actions (i.e., execution horizon) before replanning, turning replan

Control Barrier Functions via Minkowski Operations for Safe Navigation among Polytopes

SafetyDGX agent

arXiv:2608.02886v1 Announce Type: new Abstract: Safely navigating polytopic environments while respecting the dynamics, control, and exact geometry of the underlying system is a challenge in robotics.

CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation

SafetyDGX agent

arXiv:2608.03147v1 Announce Type: new Abstract: Referring Remote Sensing Image Segmentation (RRSIS) has achieved significant progress through the integration of VLMs and the Segment Anything Model (SA

CT-HEG: A Bidirectional, Timestamp-Attributed Event Graph for ICU In-Hospital Mortality Prediction - An Architectural Ablation Study

SafetyDGX agent

arXiv:2608.02663v1 Announce Type: cross Abstract: Accurate ICU mortality prediction requires modeling irregular clinical observations across heterogeneous entity types. Existing sequence models handle

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning

SafetyDGX agent

arXiv:2608.03068v1 Announce Type: cross Abstract: Reinforcement learning (RL) has emerged as an effective method for enhancing the reasoning capabilities of large language models (LLMs). However, exis

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

SafetyDGX agent

arXiv:2608.01755v2 Announce Type: replace Abstract: Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reason

Disentangling Language Modeling and Boundaries

SafetyDGX agent

arXiv:2608.03599v1 Announce Type: new Abstract: Byte-level language models are usually argued for on the grounds of robustness, multilingual fairness, and character-level skills. We point to a differe

DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformers

SafetyDGX agent

arXiv:2608.03082v1 Announce Type: new Abstract: Recent advances in Diffusion Transformers (DiTs) have enabled remarkable progress in visual synthesis, benefiting from their superior scalability. To fa

DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning

SafetyDGX agent

arXiv:2608.03292v1 Announce Type: new Abstract: Long Document Visual Question Answering (LongDocVQA) requires Multimodal Large Language Models (MLLMs) to locate, integrate, and reason over heterogeneo

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR

SafetyDGX agent

arXiv:2608.03119v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning but typically relies on ground-truth (GT) answers, limiting scalability. Vo

Double Down on Defense: Strengthening Deep Perceptual Hashes against Evasion Attacks without Retraining

SafetyDGX agent

arXiv:2608.03101v1 Announce Type: new Abstract: Near-duplicate image matching is crucial for trust and safety, provenance verification, copyright enforcement, and large-scale visual search. Modern pla

DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack

SafetyDGX agent

arXiv:2608.03207v1 Announce Type: new Abstract: Flow-matching vision-language-action (VLA) models such as pi0 generate robot actions by integrating a learned denoising velocity field, and have been re

DriftWorld: Fast World Modeling through Drifting

SafetyDGX agent

arXiv:2607.15065v2 Announce Type: replace-cross Abstract: Predictive world models enable robots to plan by imagining the outcomes of their actions, but their value for control hinges on generating man

Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation

SafetyDGX agent

arXiv:2608.03044v1 Announce Type: cross Abstract: Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignm

Enhanced Polarization Locking in VCSELs

SafetyDGX agent

arXiv:2604.01857v2 Announce Type: replace-cross Abstract: While optical injection locking (OIL) of vertical-cavity surface-emitting lasers (VCSELs) has been widely studied in the past, the polarizatio

Enhancing Q-Value Updates in Deep Q-Learning via Successor-State Prediction

SafetyDGX agent

arXiv:2511.03836v2 Announce Type: replace Abstract: Deep Q-Networks (DQNs) estimate future returns by learning from transitions sampled from a replay buffer. However, the target updates in DQN often r

Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning

SafetyDGX agent

arXiv:2608.03875v1 Announce Type: cross Abstract: Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL). Recent work uses large foundation Vision-Language Mode

Evading Chain-of-Thought Monitoring Through Model Poisoning

SafetyDGX agent

arXiv:2608.02820v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring is an increasingly important component of AI safety stacks but relies on the assumption that a model's reasoning tra

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning

SafetyDGX agent

arXiv:2608.03872v1 Announce Type: new Abstract: Human-in-the-loop reinforcement learning (HIL-RL) enables robots to learn contact-rich manipulation from limited real-world interaction, but deployment

Fast Object Removal Attacks on Safety-Critical Video-based Perception Systems

SafetyDGX agent

arXiv:2608.02806v1 Announce Type: cross Abstract: By leveraging data from video-based perception systems, intelligent transportation systems (ITS) support safety-critical applications that improve roa

Fields Medalist Jacob Tsimerman knows MUCH MUCH more about math than I do, or ever will, but I have been studying natural and artificial int…

SafetyDGX agent

Fields Medalist Jacob Tsimerman knows MUCH MUCH more about math than I do, or ever will, but I have been studying natural and artificial intelligence for 40 years, and I think his prediction here (“AI

Flying over The Uncertain Nature (FORTUNE): Intelligent and Humanistic 3D Path Planning for Low-Altitude Collaboration

SafetyDGX agent

arXiv:2608.03408v1 Announce Type: new Abstract: The proliferation of low-altitude intelligent agents is increasing the demand for timely and socially responsible collaborative sensing in dynamic urban

Forbidden Region Dynamic Active Constraints in Robot-Assisted Minimally Invasive Surgery

SafetyDGX agent

arXiv:2608.03010v1 Announce Type: new Abstract: In robot-assisted surgery, Forbidden Region Active Constraints (FRAC) represent a control strategy that helps maintain task safety by generating anisotr

From Routes to Steps: Separating Semantic Progress from Local Execution in Vision-and-Language Navigation

SafetyDGX agent

arXiv:2608.03143v1 Announce Type: new Abstract: Vision-and-Language Navigation (VLN) requires an agent to follow a route-level instruction by executing its constituent steps from egocentric visual obs

GenOS: Compositional Certificates for Semantic Robustness in AI Code Generation

SafetyDGX agent

arXiv:2608.03588v1 Announce Type: cross Abstract: AI coding agents are stochastic workflows: prompts are interpreted, artifacts are sampled, validators produce observations, and orchestrators commit o

GORDON: Graph-based Object-centric Rewards for Decomposition of Long-Horizon Manipulation

SafetyDGX agent

arXiv:2608.03753v1 Announce Type: new Abstract: Learning long-horizon manipulation skills with reinforcement learning remains challenging due to the complexity of reward design, the limited guidance o

GROW: Group-Relative Advantage-Weighted On-Policy Reinforcement Learning of Autoregressive-Diffusion Text-to-Speech model

SafetyDGX agent

arXiv:2608.03215v1 Announce Type: cross Abstract: Reinforcement learning for flow-matching text-to-speech is complicated by deterministic ODE sampling: trajectory-level policy-gradient methods typical

HERO: Hierarchical Evidential Reasoning Optimization for Radiology Report Generation via Reason-then-Summarize

SafetyDGX agent

arXiv:2601.03321v3 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have substantially advanced Radiology Report Generation (RRG), yet aligning them through reinforcemen

Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding

SafetyDGX agent

arXiv:2608.03471v1 Announce Type: new Abstract: Generative Vision-Language Models (VLMs) commonly treat bounding-box coordinates as independent output symbols, leaving numerical order and axis semanti

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning

SafetyDGX agent

arXiv:2608.03545v1 Announce Type: new Abstract: Test-time reinforcement learning (TTRL) improves the reasoning capabilities of large language models without labeled data by updating the policy with ps

HomoEnsNER: Does Language Alignment Outperform Architectural Complexity in Gujarati Named Entity Recognition?

SafetyDGX agent

arXiv:2608.03105v1 Announce Type: new Abstract: Named Entity Recognition (NER) for Gujarati remains underexplored, hindered by the absence of capitalization cues, rich morphology, lexical ambiguity, a

ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization

SafetyDGX agent

arXiv:2608.03210v1 Announce Type: new Abstract: Foundation models have achieved remarkable success across diverse tasks, but they remain vulnerable. To investigate such vulnerabilities, semantic-shift

Implementing Causal Perception: Competing SCMs and Situated Fairness

SafetyDGX agent

arXiv:2608.03917v1 Announce Type: new Abstract: Causal perception occurs when agents with competing Structural Causal Models (SCMs) of the same system infer different probability distributions, includ

Improved Quantum Algorithms for Reinforcement Learning Under a Generative Model

SafetyDGX agent

arXiv:2608.02826v1 Announce Type: cross Abstract: Reinforcement learning is a subfield of machine learning that studies how an agent interacts with an environment in order to extract as large a reward

interesting (and completely opposed to @haider1’s take on the same graph) *the labs themselves* have only modestly changed their predictions…

SafetyDGX agent

interesting (and completely opposed to @haider1’s take on the same graph) *the labs themselves* have only modestly changed their predictions on AGI timelines over the last decade. note also that (on a

Interpretable Adaptive Sampling for LLM Test-Time Scaling

SafetyDGX agent

arXiv:2608.03961v1 Announce Type: new Abstract: Test-time scaling improves LLM reasoning by generating and aggregating multiple candidate answers, yet many pipelines use fixed per-query budgets that s

is there a polite synonym for “circle jerk”?

SafetyDGX agent

The post contains two distinct snippets. First, user @GaryMarcus asks whether there is a more polite way to refer to “circle jerk.” Second, it shares a (likely satirical) claim that Microsoft’s AI rev

Joint Affine Spectral Shaping: Coupling Weight and Bias Updates Beyond Weight-Only Muon

SafetyDGX agent

arXiv:2608.02991v1 Announce Type: new Abstract: Matrix spectral optimizers reshape weight-update spectra but usually delegate vector-valued biases to a separate optimizer. We study whether this separa

Just had to create an 'accidental-cyberattacks' tag on my blog We're up to four now: the original OpenAI+Hugging Face one, Anthropic's me-to…

SafetyDGX agent

Just had to create an 'accidental-cyberattacks' tag on my blog We're up to four now: the original OpenAI+Hugging Face one, Anthropic's me-too attacks, then two new ones from the UK AI Safety Institute

Kernel weighted importance sampling for off-policy evaluation in contextual bandits

SafetyDGX agent

arXiv:2607.15067v2 Announce Type: replace Abstract: This article presents a novel estimator for performing off-policy evaluation using only offline data for contextual bandits. The proposed estimator,

KernelBrain: Coarse-to-Fine, Budget-Aware Search for Agentic GPU Kernel Optimization

SafetyDGX agent

arXiv:2608.02611v1 Announce Type: cross Abstract: Automating GPU kernel optimization remains difficult in practice: generated variants can violate correctness constraints, runtime measurements are noi

Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR

SafetyDGX agent

arXiv:2608.03610v1 Announce Type: new Abstract: Modern LLM-based ASR systems have established multilingual capability as a standard feature, leveraging large-scale multilingual corpora and LLMs' cross

LatentGuard: Efficient and Inspectable Latent Reasoning for LLM Safeguards

SafetyDGX agent

arXiv:2608.03838v1 Announce Type: new Abstract: Reasoning-based guard models improve LLM safeguards, but decoding explicit rationales for every interaction makes them costly to deploy. Although latent

Learning Attribute-aware Representations for Few-shot Scene Text Segmentation

SafetyDGX agent

arXiv:2504.11164v2 Announce Type: replace Abstract: Supervised scene text segmentation has achieved notable progress in recent years. However, its development is largely constrained by the scarcity of

← Previous
1…1011121314…210
Next →