AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,494 results
Safety

Potential-Guided Flow Matching for Vision-Language-Action Policy Improvement

DGX agent

arXiv:2606.04968v1 Announce Type: new Abstract: Large vision-language-action (VLA) policies are increasingly trained as conditional generative models over action chunks. Yet deployment produces mixed-

safetyarxiv-cs-ro
4 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Probing Outcome-Level Resemblance and Mechanism-Level Alignment in LLM Risk Decisions: Evidence from the St. Petersburg Game

DGX agent

arXiv:2606.04978v1 Announce Type: new Abstract: LLMs can appear cautious in risk decision-making tasks, yet cautious-looking outputs do not necessarily indicate alignment with human decision-making me

safetyarxiv-cs-cl
4 Jun 2026
Safety

ps no fair just mimicing walkthroughs.

DGX agent

This post likely discusses concerns about AI systems that merely imitate or reproduce existing walkthroughs and instructional content rather than demonstrating genuine understanding or original proble

safetygary-marcus--x
4 Jun 2026
Safety

Reducing the Filtering Effect in Public School Admissions: A Bias-aware Analysis for Targeted Interventions

DGX agent

arXiv:2004.10846v5 Announce Type: replace-cross Abstract: Problem definition: Traditionally, New York City's top 8 public schools have selected candidates solely based on their scores in the Specializ

safetyarxiv-cs-lg
4 Jun 2026
Safety

REGAIN: REconciliation GAIN-driven Auxiliary Direction Learning

DGX agent

arXiv:2606.04380v1 Announce Type: cross Abstract: Forecast reconciliation usually starts from a fixed measurement system and asks how forecasts should be projected onto a coherent space. We ask a diff

safetyarxiv-cs-lg
4 Jun 2026
Safety

Reinforcement Learning from Rich Feedback with Distributional DAgger

DGX agent

arXiv:2606.05152v1 Announce Type: cross Abstract: Reasoning models have advanced rapidly, but the dominant reinforcement learning from verifiable rewards (RLVR) recipe remains surprisingly narrow: sam

safetyarxiv-cs-ai
4 Jun 2026
Safety

Remember when Dario called pausing AI 'most extreme'? '...we should just pause ... that extreme position doesn't make much sense to me eithe…

DGX agent

Remember when Dario called pausing AI 'most extreme'? '...we should just pause ... that extreme position doesn't make much sense to me either.' Anthropic now: 'We believe it would be good for the worl

safetyconnor-leahy--x
4 Jun 2026
Safety

RePercENT: Scaling Disentangled Representation Learning Beyond Two Modalities

DGX agent

arXiv:2606.05109v1 Announce Type: new Abstract: To leverage the full potential of multimodal data, we need representations that go beyond the state-of-the-art alignment and fusion approaches and explo

safetyarxiv-cs-lg
4 Jun 2026
Safety

Representation Matters in Randomized Smoothing for Audio Classification

DGX agent

arXiv:2606.04210v1 Announce Type: cross Abstract: Randomized smoothing (RS) certifies robustness in the vector space where Gaussian noise is added. In audio classification, this space is often not uni

safetyarxiv-cs-lg
4 Jun 2026
Safety

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning

DGX agent

arXiv:2606.04923v1 Announce Type: cross Abstract: Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. However, policy models

safetyarxiv-cs-ai
4 Jun 2026
Safety

Rethinking Continual Experience Internalization for Self-Evolving LLM Agents

DGX agent

arXiv:2606.04703v1 Announce Type: new Abstract: Experience internalization converts contextual experience from past interactions into reusable parametric capability, offering a promising path toward c

safetyarxiv-cs-cl
4 Jun 2026
Safety

Rethinking Sales Lead Scoring with LLM-based Hierarchical Preference Ranking

DGX agent

arXiv:2606.04387v1 Announce Type: cross Abstract: Sales lead conversion in high-stakes domains (e.g., automotive, real estate) differs fundamentally from e-commerce recommendation due to prolonged dec

safetyarxiv-cs-ai
4 Jun 2026
Safety

Reusing Trajectories in Policy Gradients Enables Fast Convergence

DGX agent

arXiv:2506.06178v3 Announce Type: replace Abstract: Policy gradient (PG) methods are a class of effective reinforcement learning algorithms, particularly when dealing with continuous control problems.

safetyarxiv-cs-lg
4 Jun 2026
Safety

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training

DGX agent

arXiv:2606.04272v1 Announce Type: new Abstract: The standard LLM training pipeline applies reinforcement learning (RL) only after pre-training and supervised fine-tuning (SFT). We question this status

safetyarxiv-cs-lg
4 Jun 2026
Safety

Scaling Self-Evolving Agents via Parametric Memory

DGX agent

arXiv:2606.04536v1 Announce Type: new Abstract: Existing memory-augmented LLM agents store past experience exclusively in prompt space, as textual summaries or retrieved passages, while keeping model

safetyarxiv-cs-ai
4 Jun 2026
Safety

See also: https://x.com/emollick/status/1799046218917822682?s=20

DGX agent

See also: https://x.com/emollick/status/1799046218917822682?s=20 Alignment, of a sort: this paper conducts what they call a “moral Turing Test,” asking people to compare GPT-4o to humans on ethical qu

safetyethan-mollick--x
4 Jun 2026
Safety

Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to Signing

DGX agent

arXiv:2512.08094v2 Announce Type: replace Abstract: The goal of this work is to develop a universal approach for aligning subtitles (i.e., spoken language text with corresponding timestamps) to contin

safetyarxiv-cs-cl
4 Jun 2026
Safety

Selective Coupling of Decoupled Informative Regions: Masked Attention Alignment for Data-Free Quantization of Vision Transformers

DGX agent

arXiv:2606.04373v1 Announce Type: cross Abstract: Data-Free Quantization (DFQ) addresses data security concerns by synthesizing samples, without accessing real data. It has garnered increasing attenti

safetyarxiv-cs-ai
4 Jun 2026
Safety

Self-Distilled Policy Gradient

DGX agent

arXiv:2606.04036v1 Announce Type: new Abstract: On-policy self-distillation, where a language model conditions on privileged context to supervise its own generations, is a promising source of dense su

safetyarxiv-cs-lg
4 Jun 2026
Safety

Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model

DGX agent

arXiv:2512.21917v3 Announce Type: replace-cross Abstract: Policy alignment to preference data typically assumes a known link function between observed preferences and latent rewards (e.g., Bradley-Ter

safetyarxiv-cs-ai
4 Jun 2026
Safety

shit is getting real

DGX agent

shit is getting real Sam Altman, Dario Amodei, Demis Hassabis and many others have signed a letter urging Congress to increase security on orders of synthetic nucleic acids - and the equipment needed

safetygary-marcus--x
4 Jun 2026
Safety

Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents

DGX agent

arXiv:2510.13704v2 Announce Type: replace-cross Abstract: Recent works have proposed accelerating the wall-clock training time of actor-critic methods via the use of large-scale environment paralleliz

safetyarxiv-cs-ai
4 Jun 2026
Safety

Smart Transportation Without Neurons -- Fair Metro Network Expansion with Tabular Reinforcement Learning

DGX agent

arXiv:2606.04167v1 Announce Type: cross Abstract: We tackle the Metro Network Expansion Problem (MNEP), a subset of the Transport Network Design Problem (TNDP), which focuses on expanding metro system

safetyarxiv-cs-ai
4 Jun 2026
Safety

SocialCoach: Personalized Social Skill Learning with RL-based Agentic Tutoring and Practice

DGX agent

arXiv:2606.04155v1 Announce Type: cross Abstract: Social skills such as negotiation and leadership are crucial for personal and professional success in today's interconnected world. However, scalable

safetyarxiv-cs-cl
4 Jun 2026
Safety

SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference Optimization

DGX agent

arXiv:2505.11166v3 Announce Type: replace-cross Abstract: Despite advances in pretraining with extended context sizes, large language models (LLMs) still face challenges in effectively utilizing real-

safetyarxiv-cs-ai
4 Jun 2026
Safety

SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space

DGX agent

arXiv:2511.20102v3 Announce Type: replace Abstract: Sparse attention reduces the quadratic complexity of full self-attention but faces two challenges: (1) an attention gap, where applying sparse atten

safetyarxiv-cs-cl
4 Jun 2026
Safety

Starmer is right about this; would be ironic if the net of it was to keep Starmer in power for longer.

DGX agent

Starmer is right about this; would be ironic if the net of it was to keep Starmer in power for longer. British Prime Minister Keir Starmer has accused Elon Musk of 'whipping up division' and interferi

safetygary-marcus--x
4 Jun 2026
Safety

Stumbling Into AI Emotional Dependence: How Routine AI Interactions Reshape Human Connection

DGX agent

arXiv:2606.04150v1 Announce Type: new Abstract: Public discourse and emerging policy typically assume that AI emotional support is a deliberate act: a lonely user consciously seeking comfort from a de

safetyarxiv-cs-ai
4 Jun 2026
Safety

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success

DGX agent

arXiv:2601.18175v2 Announce Type: replace Abstract: A widely used technique for improving policies is success conditioning, in which one collects trajectories, identifies those that achieve a desired

safetyarxiv-cs-ai
4 Jun 2026
Safety

Test-time reward-guided alignment of language models by importance sampling on pre-logit space

DGX agent

arXiv:2510.26219v3 Announce Type: replace-cross Abstract: Test-time alignment of large language models (LLMs) attracts attention because fine-tuning of LLMs requires high computational costs. In this

safetyarxiv-cs-ai
4 Jun 2026
Safety

The Accountability Horizon: An Impossibility Theorem for Governing Human-Agent Collectives

DGX agent

arXiv:2604.07778v2 Announce Type: replace Abstract: Existing accountability frameworks for AI systems, legal, ethical, and regulatory, rest on a shared assumption: for any consequential outcome, at le

safetyarxiv-cs-ai
4 Jun 2026
Safety

The Digital Apprentice: A Framework for Human-Directed Agentic AI Development

DGX agent

arXiv:2606.04321v1 Announce Type: new Abstract: Agentic AI deployments face a recurring design tension: heavy human oversight limits scale, while broad autonomy outruns accountability. Neither posture

safetyarxiv-cs-ai
4 Jun 2026
Safety

The Invisible Lottery: How Subtle Cues Steer Algorithm Choice in LLM Code Generation

DGX agent

arXiv:2606.04057v1 Announce Type: cross Abstract: Large language models (LLMs) now generate substantial production code, often for tasks with multiple valid algorithmic solutions. Incidental prompt cu

safetyarxiv-cs-ai
4 Jun 2026
Safety

The Loss Is Not Enough: Sampling Conditions and Inductive Bias in Contrastive Representation Learning

DGX agent

arXiv:2606.04280v1 Announce Type: cross Abstract: Contrastive learning has become a leading paradigm for self-supervised representation learning, yet the conditions under which it recovers meaningful

safetyarxiv-cs-ai
4 Jun 2026
Safety

The Right Measure for Physics-Constrained Generation: A Co-Area Correction for Posterior-Consistent PDE Inverse Problems

DGX agent

arXiv:2606.04804v1 Announce Type: new Abstract: Generative models -- diffusion and flow matching -- are increasingly used to solve partial differential equation (PDE) inverse problems, enforcing the g

safetyarxiv-cs-lg
4 Jun 2026
Safety

Think Fast and Far: Long-Horizon Online POMDP Planning via Rapid State Sampling

DGX agent

arXiv:2606.04355v1 Announce Type: new Abstract: Partially Observable Markov Decision Processes (POMDPs) are a general and principled framework for motion planning under uncertainty. Despite tremendous

safetyarxiv-cs-ro
4 Jun 2026
Safety

THIS. I fully agree with @andrewyang, especially since GenAi leverages IP from a huge range of humans that were not adequately compensated.

DGX agent

THIS. I fully agree with @andrewyang, especially since GenAi leverages IP from a huge range of humans that were not adequately compensated. We should tax AI. https://www.cnbc.com/video/2026/06/04/andr

safetygary-marcus--x
4 Jun 2026
Safety

Three Predictions: 1. Some form of AI, probably neurosymbolic in nature, will come that is far more economical and data- and energy-efficien…

DGX agent

Three Predictions: 1. Some form of AI, probably neurosymbolic in nature, will come that is far more economical and data- and energy-efficient than LLMs, and it will make an absolute fortune. 2. LLMs,

safetygary-marcus--x
4 Jun 2026
Safety

Toward Multi-Domain and Long-Tailed Quantization via Feature Alignment and Scaling

DGX agent

arXiv:2606.04920v1 Announce Type: cross Abstract: Quantizing deep neural networks is essential for efficient inference on resource-constrained devices. However, most existing methods are designed for

safetyarxiv-cs-cv
4 Jun 2026
Safety

Towards Pretraining Text Encoders for TabPFN

DGX agent

arXiv:2606.04876v1 Announce Type: new Abstract: Tabular foundation models, such as TabPFN, achieve strong performance on tabular datasets with numerical and categorical data, but do not natively handl

safetyarxiv-cs-lg
4 Jun 2026
Safety

Trace-Mediated Peak Bias: Bridging Temporal Credit Assignment and Cognitive Heuristics in Deep Reinforcement Learning

DGX agent

arXiv:2606.04735v1 Announce Type: cross Abstract: Temporal credit assignment is central to both biological and artificial intelligence, yet its interaction with non-linear function approximation is po

safetyarxiv-cs-ai
4 Jun 2026
Safety

Transferable Multi-Bit Watermarking Across Frozen Diffusion Models via Latent Consistency Bridges

DGX agent

arXiv:2603.20304v2 Announce Type: replace Abstract: As generative AI advances, global governance frameworks increasingly mandate verifiable content provenance. However, existing watermarking technique

safetyarxiv-cs-cv
4 Jun 2026
Safety

TransTac: Visuo-Tactile Modality Transition via Ultraviolet-Encoded Transparent Elastomers

DGX agent

arXiv:2606.04477v1 Announce Type: new Abstract: Vision-based tactile sensors (VBTS) recover high-resolution contact geometry but typically rely on opaque elastomer layers that prevent visual transpare

safetyarxiv-cs-ro
4 Jun 2026
Safety

Trump’s budget director Russ Vought is the most dangerous person you’ve never heard of. And he just proposed turning every federal grant int…

DGX agent

Trump’s budget director Russ Vought is the most dangerous person you’ve never heard of. And he just proposed turning every federal grant into a loyalty test. His plan would make funding for cancer res

safetyyann-lecun--x
4 Jun 2026
Safety

U-Net-Accelerated Quality-Diversity Optimization for Climate-Adaptive Urban Layouts

DGX agent

arXiv:2606.04658v1 Announce Type: cross Abstract: Optimizing urban layouts for climate adaptation requires balancing building density with cold-air ventilation. Because physics-based climate simulatio

safetyarxiv-cs-lg
4 Jun 2026
Safety

UniFair: A unified fair clustering approach based on separation and compactness

DGX agent

arXiv:2606.04777v1 Announce Type: new Abstract: Clustering is increasingly used to support high-impact decisions, yet standard objectives such as k-means can produce clusterings that treat demographic

safetyarxiv-cs-lg
4 Jun 2026
Safety

Unlocking Proactivity in Task-Oriented Dialogue

DGX agent

arXiv:2605.22240v2 Announce Type: replace Abstract: Proactive task-oriented dialogue (TOD), such as outbound sales, demands a persuasive agent that actively probes the user's concerns and steers the c

safetyarxiv-cs-ai
4 Jun 2026
Safety

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models

DGX agent

arXiv:2602.19101v2 Announce Type: replace-cross Abstract: Value alignment of Large Language Models (LLMs) requires us to empirically measure these models' actual, acquired representation of value. Amo

safetyarxiv-cs-ai
4 Jun 2026
← Previous
1…156157158159160…302
Next →