AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,816 results
Safety

what a time to be alive

DGX agent

what a time to be alive wow, Opus 4.8 is very... argument-happy? it picked a fight with me about my usage of the word 'ontology', and when we eventually got back on the same page philosophically, told

safetygary-marcus--x
30 May 2026
Safety

Where is the power/value of money physically located? It's clearly not in the actual physical bills. Why are we more scared of an elderly ma…

Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

Where is the power/value of money physically located? It's clearly not in the actual physical bills. Why are we more scared of an elderly mafia boss than their much more physically dangerous underling

safetyconnor-leahy--x
30 May 2026
Safety

Why LLMs rarely payoff—and what I have been saying literally for 7 years—confirmed yet again: LLMs can’t handle the truth. (Nor apparently c…

DGX agent

Why LLMs rarely payoff—and what I have been saying literally for 7 years—confirmed yet again: LLMs can’t handle the truth. (Nor apparently can my critics, who keep saying I am “always wrong”, when I h

safetygary-marcus--x
30 May 2026
Safety

A 25B fund just refused SpaceX at any price. It says the company can't be worth more than 1T, half the $1.8T IPO target, and Musk's 85% co…

DGX agent

A 25B fund just refused SpaceX at any price. It says the company can't be worth more than 1T, half the $1.8T IPO target, and Musk's 85% control makes it impossible to fix from inside. https://thenextw

safetygary-marcus--x
29 May 2026
Safety

A Fully Convolutional Approach to Denoising Structural Dynamics Data from X-Ray Photon Correlation Spectroscopy

DGX agent

arXiv:2605.29975v1 Announce Type: new Abstract: We present a fully convolutional denoising autoencoder (FC-DAE) for denoising two-time intensity-intensity correlation functions (C_2) in X-ray photon c

safetyarxiv-cs-lg
29 May 2026
Safety

A Geometric View of SRC: Learning Representations for Stable Residual Inference

DGX agent

arXiv:2605.29673v1 Announce Type: cross Abstract: Reconstruction-based inference assigns a class by comparing class-wise reconstruction residuals; Sparse Representation Classification (SRC) is a canon

safetyarxiv-cs-cv
29 May 2026
Safety

A Heterogeneous Architecture for Robot RL Beyond GPU-Dominant Paradigms

DGX agent

arXiv:2605.30313v1 Announce Type: new Abstract: Simulation-based RL for contemporary robot control is increasingly organized around GPU-resident simulation: physics, rollout collection, and learning a

safetyarxiv-cs-ro
29 May 2026
Safety

A Modular Architecture for Typologically Controlled Lexicon Generation

DGX agent

arXiv:2605.28824v1 Announce Type: new Abstract: Constructing artificial lexicons that are pronounceable, typologically plausible, and semantically structured remains an open challenge in computational

safetyarxiv-cs-cl
29 May 2026
Safety

A Predictive Law for On-Policy Self-Distillation From World Feedback

DGX agent

arXiv:2605.30070v1 Announce Type: cross Abstract: Moving beyond simple scalar rewards toward richer world feedback is a natural path to more scalable RL post-training. On-policy self-distillation (OPS

safetyarxiv-cs-ai
29 May 2026
Safety

A Review of Learning-Based Motion Planning: Toward a Data-Driven Optimal Control Approach

DGX agent

arXiv:2512.11944v2 Announce Type: replace-cross Abstract: Motion planning for autonomous driving (AD) faces a critical trade-off. While traditional rule-based pipelines offer verifiable safety and int

safetyarxiv-cs-ai
29 May 2026
Safety

A Study on Question-Answer Dataset for LLM Safety Evaluation with a Focus on Illegal Activities

DGX agent

arXiv:2605.29340v1 Announce Type: new Abstract: In this paper, we discuss question-answer dataset for LLM safety evaluation, with a focus on illegal activities. Specifically, on the basis of manual an

safetyarxiv-cs-cl
29 May 2026
Safety

ActTraitBench: Quantifying the Knowledge-Decision Gap in Large Language Models via Human-Grounded Behavioral Validation

DGX agent

arXiv:2605.29791v1 Announce Type: new Abstract: While Large Language Models (LLMs) can convincingly simulate personas in explicit self-reports, they often deviate in implicit behavioral decisions, rev

safetyarxiv-cs-cl
29 May 2026
Safety

Adaptive Interviewing for Persona Simulation in LLMs: Evidence-Grounded Reasoning Improves Decision Alignment

DGX agent

arXiv:2605.29458v1 Announce Type: cross Abstract: Accurately simulating the decisions of a specific individual remains challenging for large language models (LLMs), partly because persona information

safetyarxiv-cs-ai
29 May 2026
Safety

AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching

DGX agent

arXiv:2603.01006v2 Announce Type: replace-cross Abstract: REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teacher

safetyarxiv-cs-ai
29 May 2026
Safety

Agora: Toward Autonomous Bug Detection in Production-Level Consensus Protocols with LLM Agents

DGX agent

arXiv:2605.29910v1 Announce Type: cross Abstract: Consensus protocols form the backbone of distributed systems and blockchains, where implementation bugs can cause data corruption and financial losses

safetyarxiv-cs-ai
29 May 2026
Safety

AIRGuard: Guarding Agent Actions with Runtime Authority Control

DGX agent

arXiv:2605.28914v1 Announce Type: cross Abstract: Tool-using language agents turn model decisions into external side effects: they read files, run scripts, call APIs, send messages, and invoke Model C

safetyarxiv-cs-ai
29 May 2026
Safety

AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing

DGX agent

arXiv:2605.29434v1 Announce Type: cross Abstract: Existing sentence-level watermarking methods enhance robustness to paraphrasing by anchoring watermarks in sentence semantics. However, their prefix-b

safetyarxiv-cs-ai
29 May 2026
Safety

Anytime-Valid Federated Conformal RAG for LLM Swarms

DGX agent

arXiv:2605.29139v1 Announce Type: cross Abstract: Federated Conformal RAG (FC-RAG) provides distribution-free coverage for a bandwidth-limited swarm of weak language models, but only at a fixed horizo

safetyarxiv-cs-lg
29 May 2026
Safety

Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and Cost-Aware Evaluation

DGX agent

arXiv:2605.30031v1 Announce Type: cross Abstract: Large Audio Language Models (LALMs) expand jailbreak risks from token-level prompting to the full speech perception-to-reasoning pipeline, where unsaf

safetyarxiv-cs-ai
29 May 2026
Safety

Auditing Training Data in Generative Music Models via Black-Box Membership Inference

DGX agent

arXiv:2605.29202v1 Announce Type: new Abstract: Recent advances in text-to-music generation enable high-fidelity synthesis of structured musical audio, raising growing concerns about data provenance,

safetyarxiv-cs-lg
29 May 2026
Safety

Automating Low-Risk Code Review at Meta: RADAR, Risk Calibration, and Review Efficiency

DGX agent

arXiv:2605.30208v1 Announce Type: cross Abstract: AI-assisted coding tools have altered software production. At Meta, significant lines of code per human-landed diff grew by 105.9% year over year and

safetyarxiv-cs-ai
29 May 2026
Safety

BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation

DGX agent

arXiv:2605.28994v1 Announce Type: new Abstract: AI tools to support real world decision making must be able to build simulation models that inform their recommendations and render them interpretable.

safetyarxiv-cs-ai
29 May 2026
Safety

Behavior-Induced Mirror-Prox Temporal-Difference Learning for Faster Off-Policy Prediction

DGX agent

arXiv:2605.28849v1 Announce Type: new Abstract: Gradient temporal-difference methods provide stable off-policy prediction with linear function approximation, but their practical performance is strongl

safetyarxiv-cs-ai
29 May 2026
Safety

Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures

DGX agent

arXiv:2605.29629v1 Announce Type: new Abstract: Attack Success Rate (ASR) evaluates each jailbreak with a single yes/no label at the end of generation, telling us whether a failure happened but not ho

safetyarxiv-cs-ai
29 May 2026
Safety

Beyond Bilingual Transfer: Multilingual Code-Switching in Instruction Tuning

DGX agent

arXiv:2605.29414v1 Announce Type: cross Abstract: Recent studies have shown that code-switching data (CSD), in which multiple languages are mixed within the same context, can improve cross-lingual tra

safetyarxiv-cs-ai
29 May 2026
Safety

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling

DGX agent

arXiv:2605.29697v1 Announce Type: new Abstract: In Agentic Search, trajectory-level outcome rewards fail to quantify the behavioral contributions of individual steps, while existing step-level reward

safetyarxiv-cs-ai
29 May 2026
Safety

bold set of counterpredictions, from @scaling01:

DGX agent

bold set of counterpredictions, from @scaling01: Cold take on what comes next: - OpenAI will flourish - Anthropic will continue to be profitable - Google will not catch up to Anthropic or OpenAI - no

safetygary-marcus--x
29 May 2026
Safety

BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models

DGX agent

arXiv:2605.30226v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for grounding visual-language understanding into real-world robotic manipulat

safetyarxiv-cs-ai
29 May 2026
Safety

Bridging the Sim-to-Real Gap in Reinforcement Learning-Based Industrial Dispatching through Execution Semantics

DGX agent

arXiv:2605.29078v1 Announce Type: new Abstract: Event-driven scheduling policies are increasingly deployed in industrial environments, where decisions are made under asynchronous and partially observe

safetyarxiv-cs-ai
29 May 2026
Safety

Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations

DGX agent

arXiv:2601.08064v2 Announce Type: replace Abstract: Confidence estimation (CE) indicates how reliable the answers of large language models are and impacts user trust and decision-making. Existing eval

safetyarxiv-cs-cl
29 May 2026
Safety

Causal Interventions on Continuous Variables: A Case Study on Verb Bias in Steering Vectors for In-Context Learning

DGX agent

arXiv:2605.29971v1 Announce Type: new Abstract: Causal interventions in language model representations have largely targeted discrete features, like grammatical number. However, language models must a

safetyarxiv-cs-cl
29 May 2026
Safety

Causal-JEPA: Learning World Models through Object-Level Latent Masking

DGX agent

arXiv:2602.11389v2 Announce Type: replace Abstract: World models require robust relational understanding to support prediction, reasoning, and control. While object-centric representations provide a u

safetyarxiv-cs-ai
29 May 2026
Safety

CB-SLICE: Concept-Based Interpretable Error Slice Discovery

DGX agent

arXiv:2605.29836v1 Announce Type: cross Abstract: Despite strong average-case performance, deep learning models often exhibit systematic errors on specific population groups, known as error slices. Id

safetyarxiv-cs-ai
29 May 2026
Safety

Certified Policy Optimisation for Nested Causal Bandits via PAC-Bayes Risk

DGX agent

arXiv:2605.29788v1 Announce Type: new Abstract: Critical sequential decisions are rarely single-timescale: a strategic decision causally shapes the context in which every subsequent tactical choice is

safetyarxiv-cs-ai
29 May 2026
Safety

Colored Noise Diffusion Sampling

DGX agent

arXiv:2605.30332v1 Announce Type: new Abstract: Diffusion models achieve state-of-the-art image synthesis, with their generative trajectories fundamentally exhibiting a spectral bias, resolving low-fr

safetyarxiv-cs-cv
29 May 2026
Safety

Comparative evaluation of photogrammetric reconstruction methods and 3D Gaussian Splatting for road surface roughness analysis

DGX agent

arXiv:2605.29452v1 Announce Type: new Abstract: Image-based 3D reconstruction offers a low-cost alternative to traditional sensor-based techniques for road surface assessment. This study compares four

safetyarxiv-cs-cv
29 May 2026
Safety

Crafting Desirable Climate Trajectories with RL Explored Socio-Environmental Simulations

DGX agent

arXiv:2410.07287v2 Announce Type: replace-cross Abstract: Climate change poses an existential threat, necessitating effective climate policies to enact impactful change. Decisions in this domain are i

safetyarxiv-cs-ai
29 May 2026
Safety

CRITIC-R1: Learning Structured Critics for Retrieval-Augmented Generation

DGX agent

arXiv:2605.29886v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) improves knowledge-intensive question answering by incorporating external evidence. However, existing RAG methods

safetyarxiv-cs-ai
29 May 2026
Safety

Cycle Consistency in Video Object-Centric Learning

DGX agent

arXiv:2605.30211v1 Announce Type: new Abstract: Self-supervised video Object-Centric Learning (OCL) aims to discover distinct objects and associate them across time, whereas self-supervised Multi-Obje

safetyarxiv-cs-cv
29 May 2026
Safety

DAMEL: Dual-Axis Multi-Expert Learning for Class-Imbalanced Learning

DGX agent

arXiv:2605.30135v1 Announce Type: cross Abstract: Various algorithms have been proposed to address the challenges posed by class-imbalanced learning from real-world data with long-tailed distributions

safetyarxiv-cs-ai
29 May 2026
Safety

DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Generation

DGX agent

arXiv:2605.29522v1 Announce Type: new Abstract: As scientific literature grows rapidly, automated survey generation has become a key capability for AI scientists and human researchers. However, existi

safetyarxiv-cs-ai
29 May 2026
Safety

DefSynUS: Real-time Patient-specific Intrahepatic Vessel Identification via Deformation-Aware CT-US Domain Adaptation

DGX agent

arXiv:2605.29570v1 Announce Type: new Abstract: Purpose: Laparoscopic ultrasound (LUS) enhances the safety of liver surgery by visualizing intrahepatic vessels in real-time. Still, vessel identificati

safetyarxiv-cs-cv
29 May 2026
Safety

Deja View: Looping Transformers for Multi-View 3D Reconstruction

DGX agent

arXiv:2605.30215v1 Announce Type: new Abstract: Recent feed-forward 3D reconstruction transformers have scaled to over a billion parameters, following the broader trend of increasing model capacity in

safetyarxiv-cs-cv
29 May 2026
Safety

Discovering Cooperative Pipelines: Autoresearch for Sequential Social Dilemmas

DGX agent

arXiv:2605.30003v1 Announce Type: cross Abstract: We study two-level autoresearch for cooperation: an outer-loop AI agent autonomously redesigns the inner-loop pipeline of an LLM policy-synthesis syst

safetyarxiv-cs-ai
29 May 2026
Safety

DLM-SWAI: Steering Diffusion Language Models Before They Unmask

DGX agent

arXiv:2605.29626v1 Announce Type: cross Abstract: Steering language model generation toward desired textual properties is essential for practical deployment, and inference-time methods are particularl

safetyarxiv-cs-ai
29 May 2026
Safety

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias

DGX agent

arXiv:2605.29152v1 Announce Type: new Abstract: Randomly initialized neural networks induce a prior over functions, but the predictor used in practice is produced only after training. We ask how much

safetyarxiv-cs-lg
29 May 2026
Safety

Draft-OPD: On-Policy Distillation for Speculative Draft Models

DGX agent

arXiv:2605.29343v1 Announce Type: new Abstract: Speculative decoding accelerates large language model inference by pairing a target model with a lightweight draft model whose proposed tokens are verif

safetyarxiv-cs-cl
29 May 2026
Safety

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

DGX agent

arXiv:2510.27607v3 Announce Type: replace Abstract: Augmenting vision-language-action models (VLAs) with world models is promising for robotic policy learning but faces challenges in jointly predictin

safetyarxiv-cs-cv
29 May 2026
← Previous
1…131132133134135…267
Next →