AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,813 results
4 Jun 2026

Incredible. Goldman Sachs is the lead left on the SpaceX IPO, and somehow @ft fails to mention this in the headline below 🤦‍♂️

SafetyDGX agent

Incredible. Goldman Sachs is the lead left on the SpaceX IPO, and somehow @ft fails to mention this in the headline below 🤦‍♂️ Goldman Sachs expects SpaceX’s AI revenue to surge 100 times by 2030 http

Inference-Time Vulnerability Beyond Shallow Safety: Alignment Along Generation Trajectories

SafetyDGX agent

arXiv:2606.04778v1 Announce Type: new Abstract: Safety-aligned Large Language Models (LLMs) remain vulnerable to interventions during inference that redirect generation toward harmful outputs. Recent

Instance-Level Post Hoc Uncertainty Quantification in Object Detection

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.04656v1 Announce Type: cross Abstract: Object detection is a safety-critical component of autonomous driving. It is essential to quantify the uncertainty in bounding-box predictions for saf

Instant-Fold: In-Context Imitation Learning for Deformable Object Manipulation

SafetyDGX agent

arXiv:2606.04269v1 Announce Type: cross Abstract: Deformable object manipulation (DOM) is challenging due to high-dimensional, partially observable states that evolve through long-horizon, topology-ch

Inverse Critical Experiment Design via Gradient Optimization and a Multigroup Attention-Based Neural Network Architecture

SafetyDGX agent

arXiv:2606.04033v1 Announce Type: new Abstract: The validation of advanced nuclear reactor designs and fuel concepts requires critical experiments with high neutronic similarity to the target technolo

It seems like @GaryMarcus was right: the AI revenue models are imploding

SafetyDGX agent

It seems like @GaryMarcus was right: the AI revenue models are imploding Sam Altman said AI budgeting has recently become a 'huge issue' for some companies, something that 'never came up' earlier this

it takes balls to offer an IPO on a trillion dollar valuation when you are just months from bankruptcy. but maybe that’s what is happening. …

SafetyDGX agent

Gary Marcus critiques a company's decision to pursue an IPO at a trillion-dollar valuation despite allegedly being on the brink of financial collapse. The post suggests this represents either audaciou

Kinda crazy (but typical, sadly) that some moron just claimed that I was “doomer” who said that all jobs would be replaced when I actually p…

SafetyDGX agent

Kinda crazy (but typical, sadly) that some moron just claimed that I was “doomer” who said that all jobs would be replaced when I actually publicly predicted the opposite! Receipt, from my essay “25 p

KODA: Contrastive Representation Comparison and Alignment for Vision-Language Foundation Models

SafetyDGX agent

arXiv:2606.04180v1 Announce Type: new Abstract: Vision-language foundation models such as CLIP and SigLIP provide widely used representations for multimodal learning systems. While these models are ty

La stratégie nationale en matière d’IA dévoilée aujourd’hui prône le développement d’une technologie sécuritaire, éthique, digne de confianc…

SafetyDGX agent

La stratégie nationale en matière d’IA dévoilée aujourd’hui prône le développement d’une technologie sécuritaire, éthique, digne de confiance, et bénéfique pour l’ensemble de la société — ce sont les

Large Language Models in K-12 Education: Alignment with State Curriculum Standards and Student Personas

SafetyDGX agent

arXiv:2606.04846v1 Announce Type: new Abstract: As Large Language Models (LLMs) become increasingly popular in educational settings, they raise important questions about the ethical implications of th

LaVIDE: Language-Prompted Satellite Change Detection via Map-Image Alignment

SafetyDGX agent

arXiv:2411.19758v2 Announce Type: replace-cross Abstract: Remote sensing change detection based on a map reference and an up-to-date image boosts timely observation of the Earth's surface when earlier

Learning Empirically Admissible Neural Heuristics for Combinatorial Search

SafetyDGX agent

arXiv:2606.04860v1 Announce Type: cross Abstract: Finding optimal solution paths for combinatorial puzzles like the Rubik's Cube, sliding tile puzzles, and Lights Out remains a classical challenge in

Learning While Acting: A Skill-Enhanced Test-Time Co-Evolution Framework for Online Lifelong Learning Agents

SafetyDGX agent

arXiv:2606.04815v1 Announce Type: cross Abstract: Lifelong learning is essential for Large Language Model (LLM) agents operating in dynamic, interactive environments. However, existing lifelong learni

Listening to the Workforce: Measuring Construction Worker Safety Attitudes from Social Media Discourse Using LLMs

SafetyDGX agent

arXiv:2606.04450v1 Announce Type: new Abstract: Worker safety attitudes are key determinants of whether protective practices are applied or bypassed on construction sites. Yet measuring them at scale

M3imic: Learning a Versatile Whole-Body Controller for Multimodal Motion Mimicking

SafetyDGX agent

arXiv:2606.04829v1 Announce Type: new Abstract: Building a general-purpose whole-body controller is essential for enabling diverse motion capabilities in humanoid robots across a wide range of downstr

MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models

SafetyDGX agent

arXiv:2606.04027v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) generate text by iteratively denoising partially masked sequences under bidirectional context, exposing a safe

MATCH: Multi-faceted Adaptive Topo-Consistency for Semi-Supervised Histopathology Segmentation

SafetyDGX agent

arXiv:2510.01532v2 Announce Type: replace Abstract: In semi-supervised segmentation, capturing meaningful semantic structures from unlabeled data is essential. This is particularly challenging in hist

Measuring Model Robustness via Fisher Information: Spectral Bounds, Theoretical Guarantees, and Practical Algorithms

SafetyDGX agent

arXiv:2606.04767v1 Announce Type: cross Abstract: The robustness of deep neural networks is crucial for safety-critical deployments, yet existing evaluation methods are often attack-dependent and lack

MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs

SafetyDGX agent

arXiv:2511.07107v3 Announce Type: replace Abstract: Ensuring the safety of Large Language Models (LLMs) is critical for real-world deployment. However, current safety measures often fail to address im

Meta's Oversight Board says Meta's account deactivations lack due process, violations are flagged without clarity, and there's little support for appeals (Sarah Perez/TechCrunch)

SafetyDGX agent

Sarah Perez / TechCrunch: Meta's Oversight Board says Meta's account deactivations lack due process, violations are flagged without clarity, and there's little support for appeals — Meta's Oversight B

(Mis)generalization of Helpful-only Fine-tuning

SafetyDGX agent

arXiv:2606.04413v1 Announce Type: new Abstract: Helpful-only models, that is, models that are trained to always follow user intent, are valuable for dangerous capability evaluations and other areas of

MM-BizRAG: Rethinking Multimodal Retrieval-Augmented Generation for General Purpose Enterprise Q&A

SafetyDGX agent

arXiv:2606.04231v1 Announce Type: cross Abstract: Recent advances in multimodal retrieval-augmented generation (MM-RAG) have shifted toward minimal parsing, relying on page-level images for producing

Model Evaluations: Prove Your Routing Policy Actually Works

SafetyDGX agent

This article discusses methods and tools for evaluating routing policies in machine learning models, likely covering techniques to validate that model routing decisions are effective and functioning a

MorphoQuant: Modality-Aware Quantization for Omni-modal Large Language Models

SafetyDGX agent

arXiv:2606.04349v1 Announce Type: cross Abstract: Conventional Post-Training Quantization (PTQ) methods struggle with 4-bit Omni-modal Large Language Models (OLLMs) due to the extreme distribution het

Multi-Agent Next-Best-View Optimization for Risk-Averse Planning

SafetyDGX agent

arXiv:2606.04158v1 Announce Type: new Abstract: Multi-agent Next-Best-View (NBV) selection for safe path planning in uncertain and unknown environments requires informative, safety-aware, and efficien

MusaCoder: Native GPU Kernel Generation with Full-Stack Training on Moore Threads GPU

SafetyDGX agent

arXiv:2606.04847v1 Announce Type: cross Abstract: Native GPU kernel generation turns high-level tensor programs into executable, efficient low-level code. Existing Large Language Models (LLMs) struggl

Neetyabhas: A Framework for Uncertainty-Aware Public Policy Optimization in Rational Agent-Based Models

SafetyDGX agent

arXiv:2606.04562v1 Announce Type: new Abstract: Purpose The WHO's COVID-19 non-pharmaceutical interventions (e.g., lockdowns, vaccinations) effectively curb transmission but impose heavy economic stra

OA-CutMix: Correcting the Label Bias of CutMix

SafetyDGX agent

arXiv:2606.04820v1 Announce Type: cross Abstract: CutMix has become the de facto standard mixing augmentation, yet its label assignment rests on a flawed assumption: The area of the pasted patch faith

Off-Distribution Voices: Fanfiction Subgenres as Universal Vernacular Jailbreaks for Aligned LLMs

SafetyDGX agent

arXiv:2606.04483v1 Announce Type: new Abstract: Existing jailbreaks against aligned LLMs are discrete artifacts whose surface forms are easy to fingerprint and patch. We argue that the real failure mo

Offering a “damn, I’m truly impressed” award, for the first person or team to build a domain-general AI system that can play @baldursgate3, …

SafetyDGX agent

Offering a “damn, I’m truly impressed” award, for the first person or team to build a domain-general AI system that can play @baldursgate3, start to finish. Will make the prize especially sweet if you

Offering @kevinroose $100,000 bet that no domain-general AI system can do this before his book comes out. (mimicking walkthroughs doesn’t co…

SafetyDGX agent

Offering @kevinroose $100,000 bet that no domain-general AI system can do this before his book comes out. (mimicking walkthroughs doesn’t count.) Offering a “damn, I’m truly impressed” award, for the

On-the-fly Repulsion in the Contextual Space for Rich Diversity in Diffusion Transformers

SafetyDGX agent

arXiv:2603.28762v2 Announce Type: replace-cross Abstract: Modern Text-to-Image (T2I) diffusion models have achieved remarkable semantic alignment, yet they often suffer from a significant lack of vari

Optimal Transport Flow Matching by Design

SafetyDGX agent

arXiv:2606.04092v1 Announce Type: new Abstract: Flow matching models learn to transport samples from a simple prior distribution to a complex data distribution. When prior-data pairs are coupled via o

Optimal Transport under Group Fairness Constraints

SafetyDGX agent

arXiv:2601.07144v3 Announce Type: replace-cross Abstract: Ensuring fairness in matching algorithms is a key challenge in allocating scarce resources and positions. Focusing on Optimal Transport (OT),

OSCAR: Omni-Embodiment Skeleton-Conditioned World Action Model for Robotics

SafetyDGX agent

arXiv:2606.04463v1 Announce Type: new Abstract: We present OSCAR, a precise action-conditioned video world model that generalizes across different robot embodiments and enables robot policy evaluation

Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data

SafetyDGX agent

arXiv:2601.15158v4 Announce Type: replace-cross Abstract: Transformers trained via Reinforcement Learning (RL) with outcome-based supervision can spontaneously develop the ability to generate intermed

Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning

SafetyDGX agent

arXiv:2601.07408v2 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) has emerged as a promising critic-free reinforcement learning paradigm for reasoning tasks. However, stand

Path-conditioned training: a principled way to rescale ReLU neural networks

SafetyDGX agent

arXiv:2602.19799v2 Announce Type: replace-cross Abstract: Despite recent algorithmic advances, we still lack principled ways to leverage the well-documented rescaling symmetries in ReLU neural network

PerceptTwin: Semantic Scene Reconstruction for Iterative LLM Planning and Verification

SafetyDGX agent

arXiv:2606.04226v1 Announce Type: cross Abstract: Simulation environments are useful for both robot policy learning and planning verification and validation. Traditionally, the process of creating a s

PersonaTree: Structured Lifecycle Memory for Person Understanding in LLM Agents

SafetyDGX agent

arXiv:2606.04780v1 Announce Type: new Abstract: Persistent LLM agents require memory representations that make the formation of person understanding explicit across long term interaction. Existing age

Physics-Informed Machine Learning for Short-Term Flood Prediction

SafetyDGX agent

arXiv:2606.04143v1 Announce Type: cross Abstract: Accurate flood forecasting is essential for mitigating disaster risks and protecting communities. However, purely data-driven machine learning models

Plan First, Judge Later, Run Better: A DMAIC-Inspired Agentic System for Industrial Anomaly Detection

SafetyDGX agent

arXiv:2606.04599v1 Announce Type: new Abstract: Large language model (LLM) agents have shown promise in automating complex data-analysis workflows, but their reliable deployment remains challenging in

Plug-and-Play Diffusion Meets ADMM: Dual-Variable Coupling for Robust Medical Image Reconstruction

SafetyDGX agent

arXiv:2602.23214v2 Announce Type: replace Abstract: Plug-and-Play diffusion prior (PnPDP) frameworks have emerged as a powerful paradigm for solving imaging inverse problems by treating pretrained gen

POLARIS: Guiding Small Models to Write Long Stories

SafetyDGX agent

arXiv:2606.04095v1 Announce Type: cross Abstract: Small open-weight models struggle at long-form creative writing: their generated stories either fall far short of the requested length, or their quali

Policy Gradient for Continuous-Time Robust Markov Decision Processes

SafetyDGX agent

arXiv:2606.04335v1 Announce Type: new Abstract: The framework of robust Markov decision processes (RMDPs) allows the design of reinforcement learning agents that satisfy performance guarantees under w

Potential-Guided Flow Matching for Vision-Language-Action Policy Improvement

SafetyDGX agent

arXiv:2606.04968v1 Announce Type: new Abstract: Large vision-language-action (VLA) policies are increasingly trained as conditional generative models over action chunks. Yet deployment produces mixed-

Probing Outcome-Level Resemblance and Mechanism-Level Alignment in LLM Risk Decisions: Evidence from the St. Petersburg Game

SafetyDGX agent

arXiv:2606.04978v1 Announce Type: new Abstract: LLMs can appear cautious in risk decision-making tasks, yet cautious-looking outputs do not necessarily indicate alignment with human decision-making me

ps no fair just mimicing walkthroughs.

SafetyDGX agent

This post likely discusses concerns about AI systems that merely imitate or reproduce existing walkthroughs and instructional content rather than demonstrating genuine understanding or original proble

Reducing the Filtering Effect in Public School Admissions: A Bias-aware Analysis for Targeted Interventions

SafetyDGX agent

arXiv:2004.10846v5 Announce Type: replace-cross Abstract: Problem definition: Traditionally, New York City's top 8 public schools have selected candidates solely based on their scores in the Specializ

REGAIN: REconciliation GAIN-driven Auxiliary Direction Learning

SafetyDGX agent

arXiv:2606.04380v1 Announce Type: cross Abstract: Forecast reconciliation usually starts from a fixed measurement system and asks how forecasts should be projected onto a coherent space. We ask a diff

Reinforcement Learning from Rich Feedback with Distributional DAgger

SafetyDGX agent

arXiv:2606.05152v1 Announce Type: cross Abstract: Reasoning models have advanced rapidly, but the dominant reinforcement learning from verifiable rewards (RLVR) recipe remains surprisingly narrow: sam

Remember when Dario called pausing AI 'most extreme'? '...we should just pause ... that extreme position doesn't make much sense to me eithe…

SafetyDGX agent

Remember when Dario called pausing AI 'most extreme'? '...we should just pause ... that extreme position doesn't make much sense to me either.' Anthropic now: 'We believe it would be good for the worl

RePercENT: Scaling Disentangled Representation Learning Beyond Two Modalities

SafetyDGX agent

arXiv:2606.05109v1 Announce Type: new Abstract: To leverage the full potential of multimodal data, we need representations that go beyond the state-of-the-art alignment and fusion approaches and explo

Representation Matters in Randomized Smoothing for Audio Classification

SafetyDGX agent

arXiv:2606.04210v1 Announce Type: cross Abstract: Randomized smoothing (RS) certifies robustness in the vector space where Gaussian noise is added. In audio classification, this space is often not uni

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning

SafetyDGX agent

arXiv:2606.04923v1 Announce Type: cross Abstract: Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. However, policy models

Rethinking Continual Experience Internalization for Self-Evolving LLM Agents

SafetyDGX agent

arXiv:2606.04703v1 Announce Type: new Abstract: Experience internalization converts contextual experience from past interactions into reusable parametric capability, offering a promising path toward c

Rethinking Sales Lead Scoring with LLM-based Hierarchical Preference Ranking

SafetyDGX agent

arXiv:2606.04387v1 Announce Type: cross Abstract: Sales lead conversion in high-stakes domains (e.g., automotive, real estate) differs fundamentally from e-commerce recommendation due to prolonged dec

Reusing Trajectories in Policy Gradients Enables Fast Convergence

SafetyDGX agent

arXiv:2506.06178v3 Announce Type: replace Abstract: Policy gradient (PG) methods are a class of effective reinforcement learning algorithms, particularly when dealing with continuous control problems.

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training

SafetyDGX agent

arXiv:2606.04272v1 Announce Type: new Abstract: The standard LLM training pipeline applies reinforcement learning (RL) only after pre-training and supervised fine-tuning (SFT). We question this status

← Previous
1…9091929394…214
Next →