AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,814 results
Safety

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training

DGX agent

arXiv:2606.04272v1 Announce Type: new Abstract: The standard LLM training pipeline applies reinforcement learning (RL) only after pre-training and supervised fine-tuning (SFT). We question this status

safetyarxiv-cs-lg
4 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models

DGX agent

arXiv:2502.01576v2 Announce Type: replace Abstract: Multi-modal Large Language Models (MLLMs) excel in vision-language tasks but remain vulnerable to visual adversarial perturbations that can induce h

safetyarxiv-cs-cv
4 Jun 2026
Safety

RSC: Decentralized Rigid Formation Flocking for Large-Scale Swarms via Hybrid Predictive Control and Online Reconfiguration

DGX agent

arXiv:2606.04248v1 Announce Type: new Abstract: Decentralized rigid formation flocking requires a swarm of autonomous agents to maintain a predetermined geometric configuration while moving, relying s

safetyarxiv-cs-ro
4 Jun 2026
Safety

RUBAS: Rubric-Based Reinforcement Learning for Agent Safety

DGX agent

arXiv:2606.04051v1 Announce Type: cross Abstract: The evolution of LLMs into tool-enabled agents creates a new class of safety challenges associated with real-world execution rather than simple text g

safetyarxiv-cs-ai
4 Jun 2026
Safety

Scaling Self-Evolving Agents via Parametric Memory

DGX agent

arXiv:2606.04536v1 Announce Type: new Abstract: Existing memory-augmented LLM agents store past experience exclusively in prompt space, as textual summaries or retrieved passages, while keeping model

safetyarxiv-cs-ai
4 Jun 2026
Safety

Scenario Generation for Risk-Aware Reinforcement Learning with Probably Approximately Safe Guarantees

DGX agent

arXiv:2606.04812v1 Announce Type: cross Abstract: Guaranteeing safety is critical to the deployment of reinforcement learning (RL) agents in the real-world, especially as policies learned using deep R

safetyarxiv-cs-ai
4 Jun 2026
Safety

See also: https://x.com/emollick/status/1799046218917822682?s=20

DGX agent

See also: https://x.com/emollick/status/1799046218917822682?s=20 Alignment, of a sort: this paper conducts what they call a “moral Turing Test,” asking people to compare GPT-4o to humans on ethical qu

safetyethan-mollick--x
4 Jun 2026
Safety

Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to Signing

DGX agent

arXiv:2512.08094v2 Announce Type: replace Abstract: The goal of this work is to develop a universal approach for aligning subtitles (i.e., spoken language text with corresponding timestamps) to contin

safetyarxiv-cs-cl
4 Jun 2026
Safety

Selective Coupling of Decoupled Informative Regions: Masked Attention Alignment for Data-Free Quantization of Vision Transformers

DGX agent

arXiv:2606.04373v1 Announce Type: cross Abstract: Data-Free Quantization (DFQ) addresses data security concerns by synthesizing samples, without accessing real data. It has garnered increasing attenti

safetyarxiv-cs-ai
4 Jun 2026
Safety

Self-Distilled Policy Gradient

DGX agent

arXiv:2606.04036v1 Announce Type: new Abstract: On-policy self-distillation, where a language model conditions on privileged context to supervise its own generations, is a promising source of dense su

safetyarxiv-cs-lg
4 Jun 2026
Safety

Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model

DGX agent

arXiv:2512.21917v3 Announce Type: replace-cross Abstract: Policy alignment to preference data typically assumes a known link function between observed preferences and latent rewards (e.g., Bradley-Ter

safetyarxiv-cs-ai
4 Jun 2026
Safety

shit is getting real

DGX agent

shit is getting real Sam Altman, Dario Amodei, Demis Hassabis and many others have signed a letter urging Congress to increase security on orders of synthetic nucleic acids - and the equipment needed

safetygary-marcus--x
4 Jun 2026
Safety

Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents

DGX agent

arXiv:2510.13704v2 Announce Type: replace-cross Abstract: Recent works have proposed accelerating the wall-clock training time of actor-critic methods via the use of large-scale environment paralleliz

safetyarxiv-cs-ai
4 Jun 2026
Safety

Smart Transportation Without Neurons -- Fair Metro Network Expansion with Tabular Reinforcement Learning

DGX agent

arXiv:2606.04167v1 Announce Type: cross Abstract: We tackle the Metro Network Expansion Problem (MNEP), a subset of the Transport Network Design Problem (TNDP), which focuses on expanding metro system

safetyarxiv-cs-ai
4 Jun 2026
Safety

SocialCoach: Personalized Social Skill Learning with RL-based Agentic Tutoring and Practice

DGX agent

arXiv:2606.04155v1 Announce Type: cross Abstract: Social skills such as negotiation and leadership are crucial for personal and professional success in today's interconnected world. However, scalable

safetyarxiv-cs-cl
4 Jun 2026
Safety

SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference Optimization

DGX agent

arXiv:2505.11166v3 Announce Type: replace-cross Abstract: Despite advances in pretraining with extended context sizes, large language models (LLMs) still face challenges in effectively utilizing real-

safetyarxiv-cs-ai
4 Jun 2026
Safety

SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space

DGX agent

arXiv:2511.20102v3 Announce Type: replace Abstract: Sparse attention reduces the quadratic complexity of full self-attention but faces two challenges: (1) an attention gap, where applying sparse atten

safetyarxiv-cs-cl
4 Jun 2026
Safety

Starmer is right about this; would be ironic if the net of it was to keep Starmer in power for longer.

DGX agent

Starmer is right about this; would be ironic if the net of it was to keep Starmer in power for longer. British Prime Minister Keir Starmer has accused Elon Musk of 'whipping up division' and interferi

safetygary-marcus--x
4 Jun 2026
Safety

Stumbling Into AI Emotional Dependence: How Routine AI Interactions Reshape Human Connection

DGX agent

arXiv:2606.04150v1 Announce Type: new Abstract: Public discourse and emerging policy typically assume that AI emotional support is a deliberate act: a lonely user consciously seeking comfort from a de

safetyarxiv-cs-ai
4 Jun 2026
Safety

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success

DGX agent

arXiv:2601.18175v2 Announce Type: replace Abstract: A widely used technique for improving policies is success conditioning, in which one collects trajectories, identifies those that achieve a desired

safetyarxiv-cs-ai
4 Jun 2026
Safety

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering

DGX agent

arXiv:2602.06911v2 Announce Type: replace-cross Abstract: As increasingly capable open-weight large language models (LLMs) are deployed, improving their tamper resistance against unsafe modifications,

safetyarxiv-cs-ai
4 Jun 2026
Safety

Test-time reward-guided alignment of language models by importance sampling on pre-logit space

DGX agent

arXiv:2510.26219v3 Announce Type: replace-cross Abstract: Test-time alignment of large language models (LLMs) attracts attention because fine-tuning of LLMs requires high computational costs. In this

safetyarxiv-cs-ai
4 Jun 2026
Safety

Testing Neural Networks via Bayesian-Guided Exploration of Decision Landscapes

DGX agent

arXiv:2606.04314v1 Announce Type: new Abstract: As neural networks are increasingly deployed in safety-critical domains, testing is essential to evaluate and improve their reliability. Existing testin

safetyarxiv-cs-lg
4 Jun 2026
Safety

The Accountability Horizon: An Impossibility Theorem for Governing Human-Agent Collectives

DGX agent

arXiv:2604.07778v2 Announce Type: replace Abstract: Existing accountability frameworks for AI systems, legal, ethical, and regulatory, rest on a shared assumption: for any consequential outcome, at le

safetyarxiv-cs-ai
4 Jun 2026
Safety

The Digital Apprentice: A Framework for Human-Directed Agentic AI Development

DGX agent

arXiv:2606.04321v1 Announce Type: new Abstract: Agentic AI deployments face a recurring design tension: heavy human oversight limits scale, while broad autonomy outruns accountability. Neither posture

safetyarxiv-cs-ai
4 Jun 2026
Safety

The Invisible Lottery: How Subtle Cues Steer Algorithm Choice in LLM Code Generation

DGX agent

arXiv:2606.04057v1 Announce Type: cross Abstract: Large language models (LLMs) now generate substantial production code, often for tasks with multiple valid algorithmic solutions. Incidental prompt cu

safetyarxiv-cs-ai
4 Jun 2026
Safety

The Loss Is Not Enough: Sampling Conditions and Inductive Bias in Contrastive Representation Learning

DGX agent

arXiv:2606.04280v1 Announce Type: cross Abstract: Contrastive learning has become a leading paradigm for self-supervised representation learning, yet the conditions under which it recovers meaningful

safetyarxiv-cs-ai
4 Jun 2026
Safety

The Right Measure for Physics-Constrained Generation: A Co-Area Correction for Posterior-Consistent PDE Inverse Problems

DGX agent

arXiv:2606.04804v1 Announce Type: new Abstract: Generative models -- diffusion and flow matching -- are increasingly used to solve partial differential equation (PDE) inverse problems, enforcing the g

safetyarxiv-cs-lg
4 Jun 2026
Safety

Think Fast and Far: Long-Horizon Online POMDP Planning via Rapid State Sampling

DGX agent

arXiv:2606.04355v1 Announce Type: new Abstract: Partially Observable Markov Decision Processes (POMDPs) are a general and principled framework for motion planning under uncertainty. Despite tremendous

safetyarxiv-cs-ro
4 Jun 2026
Safety

THIS. I fully agree with @andrewyang, especially since GenAi leverages IP from a huge range of humans that were not adequately compensated.

DGX agent

THIS. I fully agree with @andrewyang, especially since GenAi leverages IP from a huge range of humans that were not adequately compensated. We should tax AI. https://www.cnbc.com/video/2026/06/04/andr

safetygary-marcus--x
4 Jun 2026
Safety

Three Predictions: 1. Some form of AI, probably neurosymbolic in nature, will come that is far more economical and data- and energy-efficien…

DGX agent

Three Predictions: 1. Some form of AI, probably neurosymbolic in nature, will come that is far more economical and data- and energy-efficient than LLMs, and it will make an absolute fortune. 2. LLMs,

safetygary-marcus--x
4 Jun 2026
Safety

Toward Multi-Domain and Long-Tailed Quantization via Feature Alignment and Scaling

DGX agent

arXiv:2606.04920v1 Announce Type: cross Abstract: Quantizing deep neural networks is essential for efficient inference on resource-constrained devices. However, most existing methods are designed for

safetyarxiv-cs-cv
4 Jun 2026
Safety

Towards Pretraining Text Encoders for TabPFN

DGX agent

arXiv:2606.04876v1 Announce Type: new Abstract: Tabular foundation models, such as TabPFN, achieve strong performance on tabular datasets with numerical and categorical data, but do not natively handl

safetyarxiv-cs-lg
4 Jun 2026
Safety

Trace-Mediated Peak Bias: Bridging Temporal Credit Assignment and Cognitive Heuristics in Deep Reinforcement Learning

DGX agent

arXiv:2606.04735v1 Announce Type: cross Abstract: Temporal credit assignment is central to both biological and artificial intelligence, yet its interaction with non-linear function approximation is po

safetyarxiv-cs-ai
4 Jun 2026
Safety

Transferable Multi-Bit Watermarking Across Frozen Diffusion Models via Latent Consistency Bridges

DGX agent

arXiv:2603.20304v2 Announce Type: replace Abstract: As generative AI advances, global governance frameworks increasingly mandate verifiable content provenance. However, existing watermarking technique

safetyarxiv-cs-cv
4 Jun 2026
Safety

TransTac: Visuo-Tactile Modality Transition via Ultraviolet-Encoded Transparent Elastomers

DGX agent

arXiv:2606.04477v1 Announce Type: new Abstract: Vision-based tactile sensors (VBTS) recover high-resolution contact geometry but typically rely on opaque elastomer layers that prevent visual transpare

safetyarxiv-cs-ro
4 Jun 2026
Safety

Trump’s budget director Russ Vought is the most dangerous person you’ve never heard of. And he just proposed turning every federal grant int…

DGX agent

Trump’s budget director Russ Vought is the most dangerous person you’ve never heard of. And he just proposed turning every federal grant into a loyalty test. His plan would make funding for cancer res

safetyyann-lecun--x
4 Jun 2026
Safety

Two House lawmakers unveil bipartisan AI legislation that would override some state AI laws and require top AI developers to implement risk-management plans (Politico)

DGX agent

Politico: Two House lawmakers unveil bipartisan AI legislation that would override some state AI laws and require top AI developers to implement risk-management plans — But it's the proposal to preemp

safetytechmeme
4 Jun 2026
Safety

U-Net-Accelerated Quality-Diversity Optimization for Climate-Adaptive Urban Layouts

DGX agent

arXiv:2606.04658v1 Announce Type: cross Abstract: Optimizing urban layouts for climate adaptation requires balancing building density with cold-air ventilation. Because physics-based climate simulatio

safetyarxiv-cs-lg
4 Jun 2026
Safety

UniFair: A unified fair clustering approach based on separation and compactness

DGX agent

arXiv:2606.04777v1 Announce Type: new Abstract: Clustering is increasingly used to support high-impact decisions, yet standard objectives such as k-means can produce clusterings that treat demographic

safetyarxiv-cs-lg
4 Jun 2026
Safety

Unlocking Proactivity in Task-Oriented Dialogue

DGX agent

arXiv:2605.22240v2 Announce Type: replace Abstract: Proactive task-oriented dialogue (TOD), such as outbound sales, demands a persuasive agent that actively probes the user's concerns and steers the c

safetyarxiv-cs-ai
4 Jun 2026
Safety

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models

DGX agent

arXiv:2602.19101v2 Announce Type: replace-cross Abstract: Value alignment of Large Language Models (LLMs) requires us to empirically measure these models' actual, acquired representation of value. Amo

safetyarxiv-cs-ai
4 Jun 2026
Safety

VentAgent: When LLMs Learn to Breathe -- Multi-Objective Arbitration for ARDS Ventilation

DGX agent

arXiv:2606.04632v1 Announce Type: cross Abstract: Mechanical ventilation for Acute Respiratory Distress Syndrome (ARDS) requires balancing competing physiological goals, including oxygenation, lung pr

safetyarxiv-cs-cl
4 Jun 2026
Safety

veriFIRE: an Industrial Case Study in Verifying Consistency Properties for a DNN-Based Wildfire Detection System

DGX agent

arXiv:2606.04121v1 Announce Type: cross Abstract: We present our ongoing work on the veriFIRE project: a collaboration between industry and academia, aimed at applying verification to increase the rel

safetyarxiv-cs-lg
4 Jun 2026
Safety

vibe coding meme of the day, via @emanuelmaiberg @404mediaco

DGX agent

This post likely references a humorous or critical meme about 'vibe coding'—a colloquial term for writing code based on intuition rather than rigorous testing or established best practices. The post a

safetygary-marcus--x
4 Jun 2026
Safety

VT-3DAD: Cross-Category 3D Anomaly Detection via Visual-Text Normal Space Alignment

DGX agent

arXiv:2606.04369v1 Announce Type: new Abstract: Few-shot cross-category 3D anomaly detection aims to determine whether an unknown point cloud belongs to a target normal category using only a few norma

safetyarxiv-cs-cv
4 Jun 2026
Safety

WAM-Nav: Asymmetric Latent World-Action Modeling for Unified Visual Navigation

DGX agent

arXiv:2606.04907v1 Announce Type: new Abstract: Visual navigation requires generating smooth and collision-free trajectories under complex geometric and physical constraints. Existing reactive policie

safetyarxiv-cs-ro
4 Jun 2026
Safety

Watch the goal post shift unfold in real time: AGI used to be doing anything a person, including an expert, could do, and by the end of the …

DGX agent

Watch the goal post shift unfold in real time: AGI used to be doing anything a person, including an expert, could do, and by the end of the interview it’s not. I call this the AI bait and switch. As I

safetygary-marcus--x
4 Jun 2026
← Previous
1…114115116117118…267
Next →