AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,809 results
10 Jun 2026

Pareto-Guided Teacher Alignment for Fair Personalized Text Generation

SafetyDGX agent

arXiv:2606.10126v1 Announce Type: cross Abstract: Personalized persuasive text generation can improve relevance and engagement, but demographic conditioning may also introduce unequal framing across g

Prediction (which I have made before and which increasingly seems likely) OpenAI is gonna fall, and take a lot down with it.

SafetyDGX agent

Prediction (which I have made before and which increasingly seems likely) OpenAI is gonna fall, and take a lot down with it. BREAKING: SoftBank shares fall 9% after its attempt secure a $6 billion mar

Question: tech market is down about 10% over the last week. Will that affect SpaceX’s IPO? Or will everything proceed as planned, since it i…

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

A question posed on X regarding whether a recent 10% decline in the tech market will impact SpaceX's planned IPO timing and execution. The post, from AI researcher and entrepreneur Gary Marcus, raises

really good point. who else?

SafetyDGX agent

really good point. who else? BREAKING; Bill Gates just told Congress that Jeffrey Epstein had sought to use Gates’ affairs in an effort to blackmail him. So the question now is, if he did this to Bill

Reasoning or Memorization? Direction-Aware Diversity Exploration in LLM Reinforcement Learning

SafetyDGX agent

arXiv:2606.10346v1 Announce Type: new Abstract: Reinforcement learning has become a key paradigm for eliciting reasoning abilities in large language models, where exploration is crucial for discoverin

Reasoning over Semantic IDs Enhances Generative Recommendation

SafetyDGX agent

arXiv:2603.23183v2 Announce Type: replace-cross Abstract: Recent advances in generative recommendation have leveraged pretrained LLMs by formulating sequential recommendation as autoregressive generat

Recovering the Zipfian Distribution in Unsupervised Term Discovery

SafetyDGX agent

arXiv:2606.10781v1 Announce Type: cross Abstract: Unsupervised term discovery involves segmenting unlabelled speech into word- or syllable-like units and clustering these into a lexicon of candidate t

Representation Curriculum: Stagewise Training for Robust Ranking and Allocation

SafetyDGX agent

arXiv:2606.09891v1 Announce Type: cross Abstract: Ranking in digital marketplaces is a dynamic exposure-allocation mechanism: displayed items shape discovery trajectories and success events logged by

Rethinking Embodied Navigation via Relational Inductive Bias

SafetyDGX agent

arXiv:2606.10348v1 Announce Type: new Abstract: Object navigation requires an agent to locate a target in an unknown environment through visual observations. Existing methods typically rely on open-vo

RoboNaldo: Accurate, Stable and Powerful Humanoid Soccer Shooting via Motion-Guided Curriculum Reinforcement Learning

SafetyDGX agent

arXiv:2606.11092v1 Announce Type: cross Abstract: Elite humanoid soccer shooting requires whole-body stability, high-impulse whole-body interactions, and accuracy to targets. Motion tracking-driven re

Robust Regression of General ReLUs with Queries

SafetyDGX agent

arXiv:2606.11130v1 Announce Type: new Abstract: We study the task of agnostically learning general (as opposed to homogeneous) ReLUs under the Gaussian distribution with respect to the squared loss. I

Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution

SafetyDGX agent

arXiv:2606.10917v1 Announce Type: new Abstract: Although Large Language Model (LLM) agents have demonstrated strong performance on complex tasks, their learning is often limited by inefficient interac

SD-GRPO: Verifiable Segment Decomposition for Long-Form Vision-Language Generation

SafetyDGX agent

arXiv:2606.09871v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) and its variants, originally developed for Large Language Models (LLMs), have recently been applied to Multi

Selection, Not Salience: The Shape and Limits of Personalization in Social Highlighting

SafetyDGX agent

arXiv:2606.10398v1 Announce Type: cross Abstract: Does personalizing what a reader sees pay off, and where does it stop? Using a social web highlighter and a co-readership identity control (the same d

Selective Disk Bispectrum: A Complete and Rotation Invariant Image Descriptor

SafetyDGX agent

arXiv:2511.19706v2 Announce Type: replace-cross Abstract: Rotation invariance is a fundamental requirement across many computer vision tasks. Historically, this inductive bias has been encoded through

Self-Distillation Policy Optimization via Visual Feedback: Bridging Code and Visual Artifacts

SafetyDGX agent

arXiv:2606.10334v1 Announce Type: new Abstract: Code-generating large language models (LLMs) increasingly produce visual artifacts such as charts, web pages, and slides by writing programs that are ex

Self-EmoQ: Plutchik-Guided Value-based Planning to Drive Streaming Emotional TTS

SafetyDGX agent

arXiv:2606.09837v1 Announce Type: cross Abstract: Emotional interaction is increasingly crucial for conversational AI, yet current systems lack a self-emotion determination mechanism to drive the stre

Self-Supervised Relevance Modelling in Autonomous Driving via Counterfactual Analysis

SafetyDGX agent

arXiv:2606.10688v1 Announce Type: new Abstract: Autonomous driving relies on computationally intensive perception pipelines to continuously detect and track objects in the surrounding environment. Whi

SinkRec: Mitigating Semantic State Sink in Long Sequence Recommendation with Memory-Conditioned Gated Delta Networks

SafetyDGX agent

arXiv:2606.09888v1 Announce Type: new Abstract: Linear attention provides an efficient backbone for long-sequence recommendation by avoiding the quadratic cost of standard Transformers, but its compre

Sketch-to-Layout: A Human-Centric Computational Agent for Constraint-Aware Synthesis of Modular Photobioreactors

SafetyDGX agent

arXiv:2606.09849v1 Announce Type: cross Abstract: Building-integrated photobioreactors (PBRs) offer a pathway for carbon-neutral architecture, yet deployment is hindered by configuration complexity an

So far today • Banks to SoftBank: No thanks; OpenAI stock ain’t worth what you say. • Germany to Google: LLMs can be held liable • Senator W…

SafetyDGX agent

So far today • Banks to SoftBank: No thanks; OpenAI stock ain’t worth what you say. • Germany to Google: LLMs can be held liable • Senator Warren to SEC and SpaceX: Fix this mess; it’’s foul play to s

SocraticPO: Policy Optimization via Interactive Guidance

SafetyDGX agent

arXiv:2606.09887v1 Announce Type: cross Abstract: Reinforcement learning (RL) for large language models usually supervises reasoning with scalar outcome rewards, such as binary correctness. Such rewar

SoK: Colluding Adversaries in Machine Learning Pipelines

SafetyDGX agent

arXiv:2606.10091v1 Announce Type: cross Abstract: Machine learning (ML) models are susceptible to various security, privacy, and fairness risks. Adversaries with different characteristics (i.e., objec

Sources: Trump administration officials have told CAISI to halt publication of its model assessments while an EO President Trump signed last week is implemented (Amrith Ramkumar/Wall Street Journal)

SafetyDGX agent

Amrith Ramkumar / Wall Street Journal: Sources: Trump administration officials have told CAISI to halt publication of its model assessments while an EO President Trump signed last week is implemented

Speaker Group Encoding in Self-supervised Speech Recognition Models

SafetyDGX agent

arXiv:2606.10654v1 Announce Type: new Abstract: We investigate what self-supervised speech recognition models (S3Ms) learn about speaker groups (SGs). We examine several states of S3Ms: pretrained, fi

SpeechJBB: Probing Safety Alignment and Comprehension in Large Audio Language Models under Code-Switched Speech

SafetyDGX agent

arXiv:2606.06037v2 Announce Type: cross Abstract: Large audio language models (LALMs) are increasingly deployed in real-world applications, yet their safety alignment is still primarily evaluated on m

Standard Language Ideology in AI-Generated Language

SafetyDGX agent

arXiv:2406.08726v3 Announce Type: replace Abstract: Large language models (LLMs) generate text that reinforces standard language ideology: a bias towards certain language varieties that are granted mo

STEDiff: Strengthening Text Embedding for Text-to-Image Alignment in Diffusion Model

SafetyDGX agent

arXiv:2606.10653v1 Announce Type: new Abstract: Although pretrained text-to-image (T2I) generation models can produce high-quality images, they often fail to faithfully reflect the semantic intent of

Stop Early, Spend Less: Hidden-State Probes as a Practical Recipe for Streaming Moderation of LLM Outputs

SafetyDGX agent

arXiv:2606.10487v1 Announce Type: cross Abstract: Deploying large language models in user-facing systems requires efficient output safety filtering. Existing approaches typically rely on a separate mo

Structure-Preserving Learning Improves Geometry Generalization in Neural PDEs

SafetyDGX agent

arXiv:2602.02788v2 Announce Type: replace-cross Abstract: We aim to develop physics foundation models for science and engineering that provide real-time solutions to Partial Differential Equations (PD

subtle shift: some folks seemed to have shifted from expecting truly exponential progress to being happy they can find measurable progress a…

SafetyDGX agent

subtle shift: some folks seemed to have shifted from expecting truly exponential progress to being happy they can find measurable progress at all. and another, even larger group has grown concerned ab

Support sufficiency as action-sufficient compression: a single-cycle rate-regret formulation

SafetyDGX agent

arXiv:2606.09858v1 Announce Type: cross Abstract: Robust decision-making requires compression. A system that forms a rich support state cannot usually preserve its full structure at the point of actio

Synthesizable Molecular Generation via Soft-constrained GFlowNets with Rich Chemical Priors

SafetyDGX agent

arXiv:2602.04119v2 Announce Type: replace Abstract: The application of generative models for experimental drug discovery campaigns is severely limited by the difficulty of designing molecules de novo

Task Robustness via Re-Labelling Vision-Action Robot Data

SafetyDGX agent

arXiv:2606.10918v1 Announce Type: cross Abstract: The recent trend in scaling models for robot learning has resulted in impressive policies that can perform various manipulation tasks and generalize t

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition

SafetyDGX agent

arXiv:2606.09883v1 Announce Type: cross Abstract: Large language models (LLMs) have made remarkable progress in reasoning tasks, largely driven by post-training paradigms, especially reinforcement lea

Test-time Adversarial Takeover: A Real-time Hijacking Interface against Robotic Diffusion Policies

SafetyDGX agent

arXiv:2606.10371v1 Announce Type: cross Abstract: Diffusion-based action generation has become a foundational component of embodied AI, but its reliance on visual conditioning leaves deployed visuomot

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning

SafetyDGX agent

arXiv:2606.11087v1 Announce Type: cross Abstract: Expressive continuous control policies, such as diffusion and flow models, form the backbone of recent advances in scaling imitation learning for simu

The essay also covers what AI’s steep trajectory means for jobs and the economy, scientific progress, civil liberties, and geopolitics.

SafetyDGX agent

Dario Amodei's essay discusses the broad societal implications of AI's rapid advancement, examining its potential impacts across multiple domains including employment, economic disruption, accelerated

The Role of Feedback Alignment in Self-Distillation

SafetyDGX agent

arXiv:2606.11173v1 Announce Type: new Abstract: Conditioning a language model on additional context, such as feedback on a previous attempt, typically improves its response. Self-distillation trains t

The Whale That Outswam Evolution: Swarm Intelligence Maximises Memory in Connectome Reservoirs

SafetyDGX agent

arXiv:2606.09902v1 Announce Type: cross Abstract: Reservoir computing exploits the fixed dynamics of a recurrent network for temporal processing, requiring only a trained linear readout. Biological ne

there is an epidemic of this scam, with stolen picture and random user names and handles. note my reply lol and do not get taken.

SafetyDGX agent

Gary Marcus warns about a widespread scam epidemic involving fake profiles that use stolen pictures and randomly generated usernames/handles to deceive people. He shares an example of his response to

This is genuinely big news, major unintended consequences.

SafetyDGX agent

This is genuinely big news, major unintended consequences. 🚨Breaking news that could be huge, and enormously bad for GenAI, if other countries make similar decisions. https://the-decoder.com/landmark-

this is important. and scary.

SafetyDGX agent

this is important. and scary. CAISI has reportedly been directed to stop publishing public model assessments as the new AI EO gets implemented. Natsec engagement on AI is essential. But pulling CAISI'

Today I'm publishing a new essay, Policy on the AI Exponential. AI is progressing extremely fast—much faster than the policy process was bui…

SafetyDGX agent

Today I'm publishing a new essay, Policy on the AI Exponential. AI is progressing extremely fast—much faster than the policy process was built to handle. The essay lays out where I think the technolog

Toward Calibrated, Fair, and accurate Deepfake Detection

SafetyDGX agent

arXiv:2606.09881v1 Announce Type: cross Abstract: Deepfake detectors show large performance gaps across demographic groups. Existing fairness approaches require demographic labels, retraining, or sacr

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2606.11119v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is a promising approach for enhancing reasoning and agentic behavior in large language models. H

Trading Utility for Dynamic Fairness in Multiple Resource Division with Sequential Demand

SafetyDGX agent

arXiv:2606.10472v1 Announce Type: cross Abstract: Dynamic multi-resource allocation is a central problem in shared computing environments, where users' demands arrive sequentially and resources must b

Two to Tango: Coupled Task-Reference Selection for Safe LLM Fine-tuning

SafetyDGX agent

arXiv:2606.09866v1 Announce Type: cross Abstract: Fine-tuning safety aligned large language models (LLMs) on downstream data improves adaptation but may erode learned safety behavior. Existing methods

Uncovering Vulnerability of Vision-Language-Action Models under Joint-Level Physical Faults

SafetyDGX agent

arXiv:2606.10501v1 Announce Type: new Abstract: Deploying Vision-Language-Action (VLA) models in real robotic systems requires robustness not only to semantic and perceptual variations, but also to em

UniPET: a universal network for high-quality PET image denoising across varied dose reduction factors

SafetyDGX agent

arXiv:2606.11131v1 Announce Type: new Abstract: Most existing deep learning-based PET image denoising methods assume a fixed and known dose reduction factor (DRF) for low-dose PET images. However, the

Using Probabilistic Programs to Train Inductive Reasoning in Large Language Models

SafetyDGX agent

arXiv:2606.09856v1 Announce Type: cross Abstract: Post-training Large Language Models (LLMs) for reasoning typically focuses on deductive tasks such as mathematics and coding where correctness is veri

Visual-TCAV: Concept-based Attribution and Saliency Maps for Post-hoc Explainability in Image Classification

SafetyDGX agent

arXiv:2411.05698v3 Announce Type: replace-cross Abstract: Convolutional Neural Networks (CNNs) have shown remarkable performance in image classification. However, interpreting their predictions is cha

Warren to SEC, lightly paraphrased: “Do your f’ing job, and don’t let retail investors get screwed”

SafetyDGX agent

Senator Elizabeth Warren criticized the SEC for insufficient enforcement and investor protection, urging the agency to strengthen oversight and prevent harm to retail investors. The post, shared by AI

What if, Germany locked itself out of the LLM race and • Had students who actually learned things in high school, instead of turning in prom…

SafetyDGX agent

What if, Germany locked itself out of the LLM race and • Had students who actually learned things in high school, instead of turning in prompt outputs they barely read • Emerged from the sea of slop •

What Should a Skill Remember? Quality--Cost Trade-offs in Cost-Aware Skill Rewriting for Language Model Agents

SafetyDGX agent

arXiv:2606.09421v2 Announce Type: replace Abstract: Large language model agents increasingly rely on skills: reusable procedural documents encoding workflows, tool use, implementation patterns, valida

When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models

SafetyDGX agent

arXiv:2512.06343v3 Announce Type: replace-cross Abstract: Reward models are central to Large Language Model (LLM) alignment within the framework of RLHF. The standard objective used in reward modeling

When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models

SafetyDGX agent

arXiv:2606.10740v1 Announce Type: new Abstract: Failures in multi-turn reasoning models are largely invisible to terminal-score evaluation. A model can lock onto an unsafe stance early in a long dialo

When to Align, When to Predict: A Phase Diagram for Multimodal Learning

SafetyDGX agent

arXiv:2606.11190v1 Announce Type: new Abstract: Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic

When you hear AI 'safety' you should hear 'censorship' and 'control' instead. All of us surveilled and spied by safeguards of loving grace. …

SafetyDGX agent

When you hear AI 'safety' you should hear 'censorship' and 'control' instead. All of us surveilled and spied by safeguards of loving grace. Today it's intelligent Terms of Service control. You can't d

Why are AI research restrictions treated differently from every other safeguard? @theemozilla and @karan4d, co-founders of Nous Research: 'O…

SafetyDGX agent

Why are AI research restrictions treated differently from every other safeguard? @theemozilla and @karan4d, co-founders of Nous Research: 'On the bio stuff... just saying no to the user and being hone

← Previous
1…7778798081…214
Next →