AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
All
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,812 results
Safety

SocraticPO: Policy Optimization via Interactive Guidance

DGX agent

arXiv:2606.09887v1 Announce Type: cross Abstract: Reinforcement learning (RL) for large language models usually supervises reasoning with scalar outcome rewards, such as binary correctness. Such rewar

safetyarxiv-cs-ai
10 Jun 2026
Safety

SoK: Colluding Adversaries in Machine Learning Pipelines

Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

arXiv:2606.10091v1 Announce Type: cross Abstract: Machine learning (ML) models are susceptible to various security, privacy, and fairness risks. Adversaries with different characteristics (i.e., objec

safetyarxiv-cs-lg
10 Jun 2026
Safety

Sources: Trump administration officials have told CAISI to halt publication of its model assessments while an EO President Trump signed last week is implemented (Amrith Ramkumar/Wall Street Journal)

DGX agent

Amrith Ramkumar / Wall Street Journal: Sources: Trump administration officials have told CAISI to halt publication of its model assessments while an EO President Trump signed last week is implemented

safetytechmeme
10 Jun 2026
Safety

Speaker Group Encoding in Self-supervised Speech Recognition Models

DGX agent

arXiv:2606.10654v1 Announce Type: new Abstract: We investigate what self-supervised speech recognition models (S3Ms) learn about speaker groups (SGs). We examine several states of S3Ms: pretrained, fi

safetyarxiv-cs-cl
10 Jun 2026
Safety

SpeechJBB: Probing Safety Alignment and Comprehension in Large Audio Language Models under Code-Switched Speech

DGX agent

arXiv:2606.06037v2 Announce Type: cross Abstract: Large audio language models (LALMs) are increasingly deployed in real-world applications, yet their safety alignment is still primarily evaluated on m

safetyarxiv-cs-cl
10 Jun 2026
Safety

Standard Language Ideology in AI-Generated Language

DGX agent

arXiv:2406.08726v3 Announce Type: replace Abstract: Large language models (LLMs) generate text that reinforces standard language ideology: a bias towards certain language varieties that are granted mo

safetyarxiv-cs-cl
10 Jun 2026
Safety

STEDiff: Strengthening Text Embedding for Text-to-Image Alignment in Diffusion Model

DGX agent

arXiv:2606.10653v1 Announce Type: new Abstract: Although pretrained text-to-image (T2I) generation models can produce high-quality images, they often fail to faithfully reflect the semantic intent of

safetyarxiv-cs-cv
10 Jun 2026
Safety

Stop Early, Spend Less: Hidden-State Probes as a Practical Recipe for Streaming Moderation of LLM Outputs

DGX agent

arXiv:2606.10487v1 Announce Type: cross Abstract: Deploying large language models in user-facing systems requires efficient output safety filtering. Existing approaches typically rely on a separate mo

safetyarxiv-cs-ai
10 Jun 2026
Safety

Structure-Preserving Learning Improves Geometry Generalization in Neural PDEs

DGX agent

arXiv:2602.02788v2 Announce Type: replace-cross Abstract: We aim to develop physics foundation models for science and engineering that provide real-time solutions to Partial Differential Equations (PD

safetyarxiv-cs-ai
10 Jun 2026
Safety

subtle shift: some folks seemed to have shifted from expecting truly exponential progress to being happy they can find measurable progress a…

DGX agent

subtle shift: some folks seemed to have shifted from expecting truly exponential progress to being happy they can find measurable progress at all. and another, even larger group has grown concerned ab

safetygary-marcus--x
10 Jun 2026
Safety

Support sufficiency as action-sufficient compression: a single-cycle rate-regret formulation

DGX agent

arXiv:2606.09858v1 Announce Type: cross Abstract: Robust decision-making requires compression. A system that forms a rich support state cannot usually preserve its full structure at the point of actio

safetyarxiv-cs-ai
10 Jun 2026
Safety

Synthesizable Molecular Generation via Soft-constrained GFlowNets with Rich Chemical Priors

DGX agent

arXiv:2602.04119v2 Announce Type: replace Abstract: The application of generative models for experimental drug discovery campaigns is severely limited by the difficulty of designing molecules de novo

safetyarxiv-cs-lg
10 Jun 2026
Safety

Task Robustness via Re-Labelling Vision-Action Robot Data

DGX agent

arXiv:2606.10918v1 Announce Type: cross Abstract: The recent trend in scaling models for robot learning has resulted in impressive policies that can perform various manipulation tasks and generalize t

safetyarxiv-cs-lg
10 Jun 2026
Safety

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition

DGX agent

arXiv:2606.09883v1 Announce Type: cross Abstract: Large language models (LLMs) have made remarkable progress in reasoning tasks, largely driven by post-training paradigms, especially reinforcement lea

safetyarxiv-cs-ai
10 Jun 2026
Safety

Test-time Adversarial Takeover: A Real-time Hijacking Interface against Robotic Diffusion Policies

DGX agent

arXiv:2606.10371v1 Announce Type: cross Abstract: Diffusion-based action generation has become a foundational component of embodied AI, but its reliance on visual conditioning leaves deployed visuomot

safetyarxiv-cs-ai
10 Jun 2026
Safety

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning

DGX agent

arXiv:2606.11087v1 Announce Type: cross Abstract: Expressive continuous control policies, such as diffusion and flow models, form the backbone of recent advances in scaling imitation learning for simu

safetyarxiv-cs-ai
10 Jun 2026
Safety

The essay also covers what AI’s steep trajectory means for jobs and the economy, scientific progress, civil liberties, and geopolitics.

DGX agent

Dario Amodei's essay discusses the broad societal implications of AI's rapid advancement, examining its potential impacts across multiple domains including employment, economic disruption, accelerated

safetydario-amodei--x
10 Jun 2026
Safety

The Role of Feedback Alignment in Self-Distillation

DGX agent

arXiv:2606.11173v1 Announce Type: new Abstract: Conditioning a language model on additional context, such as feedback on a previous attempt, typically improves its response. Self-distillation trains t

safetyarxiv-cs-ai
10 Jun 2026
Safety

The Whale That Outswam Evolution: Swarm Intelligence Maximises Memory in Connectome Reservoirs

DGX agent

arXiv:2606.09902v1 Announce Type: cross Abstract: Reservoir computing exploits the fixed dynamics of a recurrent network for temporal processing, requiring only a trained linear readout. Biological ne

safetyarxiv-cs-ai
10 Jun 2026
Safety

there is an epidemic of this scam, with stolen picture and random user names and handles. note my reply lol and do not get taken.

DGX agent

Gary Marcus warns about a widespread scam epidemic involving fake profiles that use stolen pictures and randomly generated usernames/handles to deceive people. He shares an example of his response to

safetygary-marcus--x
10 Jun 2026
Safety

This is genuinely big news, major unintended consequences.

DGX agent

This is genuinely big news, major unintended consequences. 🚨Breaking news that could be huge, and enormously bad for GenAI, if other countries make similar decisions. https://the-decoder.com/landmark-

safetygary-marcus--x
10 Jun 2026
Safety

this is important. and scary.

DGX agent

this is important. and scary. CAISI has reportedly been directed to stop publishing public model assessments as the new AI EO gets implemented. Natsec engagement on AI is essential. But pulling CAISI'

safetygary-marcus--x
10 Jun 2026
Safety

Today I'm publishing a new essay, Policy on the AI Exponential. AI is progressing extremely fast—much faster than the policy process was bui…

DGX agent

Today I'm publishing a new essay, Policy on the AI Exponential. AI is progressing extremely fast—much faster than the policy process was built to handle. The essay lays out where I think the technolog

safetyboris-cherny--x
10 Jun 2026
Safety

Toward Calibrated, Fair, and accurate Deepfake Detection

DGX agent

arXiv:2606.09881v1 Announce Type: cross Abstract: Deepfake detectors show large performance gaps across demographic groups. Existing fairness approaches require demographic labels, retraining, or sacr

safetyarxiv-cs-cv
10 Jun 2026
Safety

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning

DGX agent

arXiv:2606.11119v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is a promising approach for enhancing reasoning and agentic behavior in large language models. H

safetyarxiv-cs-ai
10 Jun 2026
Safety

Trading Utility for Dynamic Fairness in Multiple Resource Division with Sequential Demand

DGX agent

arXiv:2606.10472v1 Announce Type: cross Abstract: Dynamic multi-resource allocation is a central problem in shared computing environments, where users' demands arrive sequentially and resources must b

safetyarxiv-cs-lg
10 Jun 2026
Safety

Two to Tango: Coupled Task-Reference Selection for Safe LLM Fine-tuning

DGX agent

arXiv:2606.09866v1 Announce Type: cross Abstract: Fine-tuning safety aligned large language models (LLMs) on downstream data improves adaptation but may erode learned safety behavior. Existing methods

safetyarxiv-cs-ai
10 Jun 2026
Safety

Uncovering Vulnerability of Vision-Language-Action Models under Joint-Level Physical Faults

DGX agent

arXiv:2606.10501v1 Announce Type: new Abstract: Deploying Vision-Language-Action (VLA) models in real robotic systems requires robustness not only to semantic and perceptual variations, but also to em

safetyarxiv-cs-ro
10 Jun 2026
Safety

UniPET: a universal network for high-quality PET image denoising across varied dose reduction factors

DGX agent

arXiv:2606.11131v1 Announce Type: new Abstract: Most existing deep learning-based PET image denoising methods assume a fixed and known dose reduction factor (DRF) for low-dose PET images. However, the

safetyarxiv-cs-cv
10 Jun 2026
Safety

Using Probabilistic Programs to Train Inductive Reasoning in Large Language Models

DGX agent

arXiv:2606.09856v1 Announce Type: cross Abstract: Post-training Large Language Models (LLMs) for reasoning typically focuses on deductive tasks such as mathematics and coding where correctness is veri

safetyarxiv-cs-ai
10 Jun 2026
Safety

Visual-TCAV: Concept-based Attribution and Saliency Maps for Post-hoc Explainability in Image Classification

DGX agent

arXiv:2411.05698v3 Announce Type: replace-cross Abstract: Convolutional Neural Networks (CNNs) have shown remarkable performance in image classification. However, interpreting their predictions is cha

safetyarxiv-cs-ai
10 Jun 2026
Safety

Warren to SEC, lightly paraphrased: “Do your f’ing job, and don’t let retail investors get screwed”

DGX agent

Senator Elizabeth Warren criticized the SEC for insufficient enforcement and investor protection, urging the agency to strengthen oversight and prevent harm to retail investors. The post, shared by AI

safetygary-marcus--x
10 Jun 2026
Safety

What if, Germany locked itself out of the LLM race and • Had students who actually learned things in high school, instead of turning in prom…

DGX agent

What if, Germany locked itself out of the LLM race and • Had students who actually learned things in high school, instead of turning in prompt outputs they barely read • Emerged from the sea of slop •

safetygary-marcus--x
10 Jun 2026
Safety

What Should a Skill Remember? Quality--Cost Trade-offs in Cost-Aware Skill Rewriting for Language Model Agents

DGX agent

arXiv:2606.09421v2 Announce Type: replace Abstract: Large language model agents increasingly rely on skills: reusable procedural documents encoding workflows, tool use, implementation patterns, valida

safetyarxiv-cs-cl
10 Jun 2026
Safety

When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models

DGX agent

arXiv:2512.06343v3 Announce Type: replace-cross Abstract: Reward models are central to Large Language Model (LLM) alignment within the framework of RLHF. The standard objective used in reward modeling

safetyarxiv-cs-ai
10 Jun 2026
Safety

When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models

DGX agent

arXiv:2606.10740v1 Announce Type: new Abstract: Failures in multi-turn reasoning models are largely invisible to terminal-score evaluation. A model can lock onto an unsafe stance early in a long dialo

safetyarxiv-cs-ai
10 Jun 2026
Safety

When to Align, When to Predict: A Phase Diagram for Multimodal Learning

DGX agent

arXiv:2606.11190v1 Announce Type: new Abstract: Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic

safetyarxiv-cs-lg
10 Jun 2026
Safety

When you hear AI 'safety' you should hear 'censorship' and 'control' instead. All of us surveilled and spied by safeguards of loving grace. …

DGX agent

When you hear AI 'safety' you should hear 'censorship' and 'control' instead. All of us surveilled and spied by safeguards of loving grace. Today it's intelligent Terms of Service control. You can't d

safetyyann-lecun--x
10 Jun 2026
Safety

Why are AI research restrictions treated differently from every other safeguard? @theemozilla and @karan4d, co-founders of Nous Research: 'O…

DGX agent

Why are AI research restrictions treated differently from every other safeguard? @theemozilla and @karan4d, co-founders of Nous Research: 'On the bio stuff... just saying no to the user and being hone

safetynous-research--x
10 Jun 2026
Safety

Wow. This is worth watching, might have huge impact. @SenWarren makes some valid points, and the SEC owes her public answers

DGX agent

Wow. This is worth watching, might have huge impact. @SenWarren makes some valid points, and the SEC owes her public answers Sen. Warren calls on SEC to delay SpaceX IPO @CNBC https://www.cnbc.com/202

safetygary-marcus--x
10 Jun 2026
Safety

YUBI: Yielding Universal Bidigital Interface for Bimanual Dexterous Manipulation at Scale

DGX agent

arXiv:2606.10244v1 Announce Type: cross Abstract: We introduce Yielding Universal Bidigital Interface (YUBI), a finger-aligned gripper designed to enable intuitive, ergonomic, and scalable data collec

safetyarxiv-cs-ai
10 Jun 2026
Safety

6G Empowering Future Robotics: A Vision for Next-Generation Autonomous Systems

DGX agent

arXiv:2602.12246v2 Announce Type: replace-cross Abstract: The convergence of robotics and next-generation communication is a critical driver of technological advancement. As the world transitions from

safetyarxiv-cs-ro
9 Jun 2026
Safety

A Finetuned SpeechLLM for Joint Multi-Granular L2 Assessment and Natural-Language Rationales

DGX agent

arXiv:2606.09470v1 Announce Type: cross Abstract: Automated L2 speech assessment can assign proficiency labels, but often lacks interpretability. We propose a rubric-guided SpeechLLM for multi-aspect,

safetyarxiv-cs-ai
9 Jun 2026
Safety

A Geometric Unification of Concept Learning with Concept Cones

DGX agent

arXiv:2512.07355v2 Announce Type: replace Abstract: Two traditions of interpretability have evolved side by side but seldom spoken to each other: Concept Bottleneck Models (CBMs), which prescribe what

safetyarxiv-cs-ai
9 Jun 2026
Safety

A Joint Finite-Sample Certificate for Adaptive Selective Conformal Risk Control

DGX agent

arXiv:2606.08517v1 Announce Type: new Abstract: Selective predictors answer on confident inputs and abstain elsewhere; deploying one safely needs a single finite-sample certificate that simultaneously

safetyarxiv-cs-lg
9 Jun 2026
Safety

A Mixed Diet Makes DINO An Omnivorous Vision Encoder

DGX agent

arXiv:2602.24181v2 Announce Type: replace-cross Abstract: Pre-trained vision encoders like DINOv2 have demonstrated exceptional performance on unimodal tasks. However, we observe that their features a

safetyarxiv-cs-ai
9 Jun 2026
Safety

A practical probabilistic framework for deformable image registration uncertainty in radiotherapy dose propagation

DGX agent

arXiv:2606.09253v1 Announce Type: new Abstract: Deformable image registration (DIR) is widely used in radiotherapy for dose propagation and accumulation, but uncertainty in the underlying deformation

safetyarxiv-cs-cv
9 Jun 2026
Safety

A Unifying Lens on Reward Uncertainty in RLHF

DGX agent

arXiv:2606.09073v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) is bottlenecked by reward hacking, where the policy exploits errors in a proxy reward model (RM) and

safetyarxiv-cs-ai
9 Jun 2026
← Previous
1…979899100101…267
Next →