AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,707 results
4 May 2026

Adaptive Equilibrium: Dynamic Weighting Framework for Generalized Interruption of DeepFake Models

SafetyDGX agent

arXiv:2605.00443v1 Announce Type: cross Abstract: The advancement of generalized deepfake disruption is constrained by the interruption imbalance, a fundamental bottleneck inherent to the generation o

Agent Capsules: Quality-Gated Granularity Control for Multi-Agent LLM Pipelines

SafetyDGX agent

arXiv:2605.00410v1 Announce Type: new Abstract: A multi-agent pipeline with N agents typically issues N LLM calls per run. Merging agents into fewer calls (compound execution) promises token savings,

Almost every warning that I have issued over the last several years has come true. You better hope to God or Darwin or whoever you believe i…

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Almost every warning that I have issued over the last several years has come true. You better hope to God or Darwin or whoever you believe in that this one is wrong. Accidental nuclear war is a potent

🌐At the @UN, @Yoshua_Bengio briefed @antonioguterres on AI governance, followed by a reception with @CanadaUN & @DavidLametti opened by @Eg…

SafetyDGX agent

🌐At the @UN, @Yoshua_Bengio briefed @antonioguterres on AI governance, followed by a reception with @CanadaUN & @DavidLametti opened by @EgriseldaL As AI risks grow, global coordination matters. 🔗AI D

Augmented Lagrangian Multiplier Network for State-wise Safety in Reinforcement Learning

SafetyDGX agent

arXiv:2605.00667v1 Announce Type: new Abstract: Safety is a primary challenge in real-world reinforcement learning (RL). Formulating safety requirements as state-wise constraints has become a prominen

Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign Prompts

SafetyDGX agent

arXiv:2508.06361v4 Announce Type: replace Abstract: Large Language Models (LLMs) are widely deployed in reasoning, planning, and decision-making tasks, making their trustworthiness critical. A signifi

Beyond Suffixes: Token Position in GCG Adversarial Attacks on Large Language Models

SafetyDGX agent

arXiv:2602.03265v2 Announce Type: replace Abstract: Large Language Models (LLMs) have seen widespread adoption across multiple domains, creating an urgent need for robust safety alignment mechanisms.

Bias in Large Language Models: Origin, Evaluation, and Mitigation

SafetyDGX agent

arXiv:2411.10915v2 Announce Type: replace Abstract: Large Language Models (LLMs) have revolutionized natural language processing, but their susceptibility to biases poses significant challenges. This

BlenderRAG: High-Fidelity 3D Object Generation via Retrieval-Augmented Code Synthesis

SafetyDGX agent

arXiv:2605.00632v1 Announce Type: new Abstract: Automatic generation of executable Blender code from natural language remains challenging, with state-of-the-art LLMs producing frequent syntactic error

BOLT: Online Lightweight Adaptation for Preparation-Free Heterogeneous Cooperative Perception

SafetyDGX agent

arXiv:2605.00405v1 Announce Type: new Abstract: Most existing heterogeneous cooperative perception methods depend on prior preparation like offline joint training or tailored collaborator-model adapta

Can Small Language Models Handle Context-Summarized Multi-Turn Customer-Service QA? A Synthetic Data-Driven Comparative Evaluation

SafetyDGX agent

arXiv:2602.00665v3 Announce Type: replace Abstract: Customer-service question answering (QA) systems increasingly rely on conversational language understanding. While Large Language Models (LLMs) achi

Confirmed 2.5 years later, by Greg Brockman.

SafetyDGX agent

Confirmed 2.5 years later, by Greg Brockman. Remember how Sam Altman told the US Senate he had no “direct” investment in OpenAI? and how they gushed over his apparent selflessness? 👉He didn’t mention

Conformalized Quantum DeepONet Ensembles for Scalable Operator Learning with Distribution-Free Uncertainty

SafetyDGX agent

arXiv:2605.00330v1 Announce Type: new Abstract: Operator learning enables fast surrogate modeling of high-dimensional dynamical systems, but existing approaches face two fundamental limitations: quadr

Data Deletion Can Help in Adaptive RL

SafetyDGX agent

arXiv:2605.00298v1 Announce Type: new Abstract: Deploying reinforcement learning policies in the real world requires adapting to time-varying environments. We study this problem in the contextual Mark

Debate-Enhanced Pseudo Labeling and Frequency-Aware Progressive Debiasing for Weakly-Supervised Camouflaged Object Detection with Scribble Annotations

SafetyDGX agent

arXiv:2512.20260v5 Announce Type: replace Abstract: Weakly-Supervised Camouflaged Object Detection (WSCOD) aims to locate and segment objects that are visually concealed within their surrounding scene

Decentralized Proximal Stochastic Gradient Langevin Dynamics

SafetyDGX agent

arXiv:2605.00723v1 Announce Type: cross Abstract: We propose Decentralized Proximal Stochastic Gradient Langevin Dynamics (DE-PSGLD), a decentralized Markov chain Monte Carlo (MCMC) algorithm for samp

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment

SafetyDGX agent

arXiv:2506.00166v2 Announce Type: replace-cross Abstract: Existing paradigms for ensuring AI safety, such as guardrail models and alignment training, often compromise either inference efficiency or de

@dromanocpm @GaryMarcus Two stood up. I know that one told the truth. Gary Marcus and Sam Altman, Senate Hearing on Oversight for Artificial…

SafetyDGX agent

Gary Marcus and Sam Altman testified before the Senate regarding AI oversight, with Marcus noting that two individuals stood up during the hearing and asserting that one of them told the truth. This r

Dynamic-TD3: A Novel Algorithm for UAV Path Planning with Dynamic Obstacle Trajectory Prediction

SafetyDGX agent

arXiv:2605.00059v1 Announce Type: new Abstract: Deep reinforcement learning (DRL) finds extensive application in autonomous drone navigation within complex, high-risk environments. However, its practi

Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory

SafetyDGX agent

arXiv:2605.00238v1 Announce Type: new Abstract: Automated short answer grading (ASAG) with large language models (LLMs) is commonly evaluated with aggregate metrics such as macro-F1 and Cohen's kappa.

Exploring LLM biases to manipulate AI search overview

SafetyDGX agent

arXiv:2605.00012v1 Announce Type: cross Abstract: Modern large language models (LLMs) are used in many business applications in general, and specifically in web search systems and applications that ge

Fair Dataset Distillation via Cross-Group Barycenter Alignment

SafetyDGX agent

arXiv:2605.00185v1 Announce Type: new Abstract: Dataset Distillation aims to compress a large dataset into a small synthetic one while maintaining predictive performance. We show that as different dem

Fairness of Classifiers in the Presence of Constraints between Features

SafetyDGX agent

arXiv:2605.00592v1 Announce Type: new Abstract: In Machine Learning, an accepted definition of fairness of a decision taken by a classifier is that it should not depend on protected features, such as

Foundation AI Models for Aerosol Optical Depth Estimation from PACE Satellite Data

SafetyDGX agent

arXiv:2605.00678v1 Announce Type: new Abstract: Aerosol Optical Depth (AOD) retrieval is essential for Earth observation, supporting applications from air quality monitoring to climate studies. Conven

FreeRet: MLLMs as Training-Free Retrievers

SafetyDGX agent

arXiv:2509.24621v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) are emerging as versatile foundations for mixed-modality retrieval. Yet, they often require heavy post-hoc

@GaryMarcus to yann - yank overhyped AI claims after they've been marcussed https://julian-goldstahl.blogspot.com/2025/09/dictionairy.html

SafetyDGX agent

Gary Marcus criticizes overhyped AI claims, with the term 'marcussed' referring to his practice of debunking exaggerated AI assertions. The post appears to reference a discussion with Yann LeCun about

🚨 GREG BROCKMAN JUST CONFESSED UNDER OATH Q: You have an ownership interest in this cap profit company. Brockman: That is accurate. Q: And …

SafetyDGX agent

🚨 GREG BROCKMAN JUST CONFESSED UNDER OATH Q: You have an ownership interest in this cap profit company. Brockman: That is accurate. Q: And you invested 0 in order to acquire that interest. Correct? Br

Greg Brockman’s diary is well on its way to being the most famous, impactful diary of the 21st century.

SafetyDGX agent

Greg Brockman’s diary is well on its way to being the most famous, impactful diary of the 21st century. takeaway from Greg Brockman testimony at Elon vs. OpenAI trial today is that no grown man should

Hey filthy rich AI warlords who have alienated most of the population, don’t say I didn’t try to warn you.🤷‍♂️

SafetyDGX agent

Hey filthy rich AI warlords who have alienated most of the population, don’t say I didn’t try to warn you.🤷‍♂️ “Bubble or not, the AI backlash is validating what one researcher and critic has been say

How 'emotion AI', the use of facial and sentiment analysis tools to track workers' moods, is seeping into white-collar jobs amid concerns over privacy and bias (Ellen Cushing/The Atlantic)

SafetyDGX agent

Ellen Cushing / The Atlantic: How “emotion AI”, the use of facial and sentiment analysis tools to track workers' moods, is seeping into white-collar jobs amid concerns over privacy and bias — The good

How Language Models Process Out-of-Distribution Inputs: A Two-Pathway Framework

SafetyDGX agent

arXiv:2605.00269v1 Announce Type: new Abstract: Recent white-box OOD detection methods for LLMs -- including CED, RAUQ, and WildGuard confidence scores -- appear effective, but we show they are struct

HyCOP: Hybrid Composition Operators for Interpretable Learning of PDEs

SafetyDGX agent

arXiv:2605.00820v1 Announce Type: cross Abstract: We introduce HyCOP, a modular framework that learns parametric PDE solution operators by composing simple modules (advection, diffusion, learned closu

i love the recent explosion of interactive data visualizations. more plz

SafetyDGX agent

i love the recent explosion of interactive data visualizations. more plz Who actually shapes AI policy in the U.S.? We mapped 1,812 entities: 745 people, 918 organizations, 2,925 relationships. Fronti

If you need proof of why AI output can't be trusted by default watch that page for more than 1 minute for any given LLM that produces output…

SafetyDGX agent

If you need proof of why AI output can't be trusted by default watch that page for more than 1 minute for any given LLM that produces output in there. An amazing time capsule of the unevenness of curr

Impact of Task Phrasing on Presumptions in Large Language Models

SafetyDGX agent

arXiv:2605.00436v1 Announce Type: new Abstract: Concerns with the safety and reliability of applying large-language models (LLMs) in unpredictable real-world applications motivate this study, which ex

Import AI 455: AI systems are about to start building themselves.

SafetyDGX agent

This newsletter entry discusses advancements in automating AI research and development processes, exploring how artificial intelligence systems are becoming capable of autonomously designing and impro

InpaintSLat: Inpainting Structured 3D Latents via Initial Noise Optimization

SafetyDGX agent

arXiv:2605.00664v1 Announce Type: new Abstract: We present a training-free approach for controllable 3D inpainting based on initial noise optimization. In the structured 3D latent diffusion framework,

Intelligent Elastic Feature Fading: Enabling Model Retrain-Free Feature Efficiency Rollouts at Scale

SafetyDGX agent

arXiv:2605.00324v1 Announce Type: cross Abstract: Large-scale ranking systems depend on thousands of features derived from user behavior across multiple time horizons. Typically requires model retrain

Jensen Huang said Nvidia's market share of AI accelerators in China has 'now dropped to zero' and that US export policy 'has already largely backfired' (Anton Shilov/Tom's Hardware)

SafetyDGX agent

Anton Shilov / Tom's Hardware: Jensen Huang said Nvidia's market share of AI accelerators in China has “now dropped to zero” and that US export policy “has already largely backfired” — US export restr

Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs

SafetyDGX agent

arXiv:2408.11513v2 Announce Type: replace Abstract: This paper focuses on learning a Constrained Markov Decision Process (CMDP) via general parameterized policies. We propose a Primal-Dual based Regul

Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding

SafetyDGX agent

arXiv:2605.00642v1 Announce Type: cross Abstract: Graphical User Interface (GUI) grounding maps natural language instructions to the visual coordinates of target elements and serves as a core capabili

Learning Coarse-to-Fine Osteoarthritis Representations under Noisy Hierarchical Labels

SafetyDGX agent

arXiv:2605.00718v1 Announce Type: new Abstract: Knee osteoarthritis (OA) assessment involves a natural but often underused label hierarchy: a coarse binary OA decision and a fine-grained Kellgren--Law

Learning How and What to Memorize: Cognition-Inspired Two-Stage Optimization for Evolving Memory

SafetyDGX agent

arXiv:2605.00702v1 Announce Type: new Abstract: Large language model (LLM) agents require long-term user memory for consistent personalization, but limited context windows hinder tracking evolving pre

Learning physically grounded traffic accident reconstruction from public accident reports

SafetyDGX agent

arXiv:2605.00050v1 Announce Type: cross Abstract: Traffic accidents are routinely documented in textual reports, yet physically grounded accident reconstruction remains difficult because detailed scen

Learning while Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies

SafetyDGX agent

arXiv:2605.00416v1 Announce Type: new Abstract: Generalist robot policies increasingly benefit from large-scale pretraining, but offline data alone is insufficient for robust real-world deployment. De

Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving

SafetyDGX agent

arXiv:2605.00556v1 Announce Type: cross Abstract: Partial driving automation creates a tension: drivers remain legally responsible for vehicle behaviour, yet their active control is significantly redu

MemRouter: Memory-as-Embedding Routing for Long-Term Conversational Agents

SafetyDGX agent

arXiv:2605.00356v1 Announce Type: new Abstract: Long-term conversational agents must decide which turns to store in external memory, yet recent systems rely on autoregressive LLM generation at every t

Meritocratic Fairness in Budgeted Combinatorial Multi-armed Bandits via Shapley Values

SafetyDGX agent

arXiv:2605.00762v1 Announce Type: new Abstract: We propose a new framework for meritocratic fairness in budgeted combinatorial multi-armed bandits with full-bandit feedback (BCMAB-FBF). Unlike semi-ba

Mesh Field Theory: Port-Hamiltonian Formulation of Mesh-Based Physics

SafetyDGX agent

arXiv:2605.00394v1 Announce Type: new Abstract: We present Mesh Field Theory (MeshFT) and its neural realization, MeshFT-Net: a structure-preserving framework for mesh-based continuum physics that cle

Model-Based Reinforcement Learning with Double Oracle Efficiency in Policy Optimization and Offline Estimation

SafetyDGX agent

arXiv:2605.00393v1 Announce Type: new Abstract: Reinforcement learning (RL) in large environments often suffers from severe computational bottlenecks, as conventional regret minimization algorithms re

New Mexico child safety trial: New Mexico asks a judge to declare Meta a public nuisance and to order it to pay $3.7B and overhaul its apps to protect children (Diana Novak Jones/Reuters)

SafetyDGX agent

Diana Novak Jones / Reuters: New Mexico child safety trial: New Mexico asks a judge to declare Meta a public nuisance and to order it to pay $3.7B and overhaul its apps to protect children — The U.S.

NEW paper from Sakana AI (ICLR 2026). A 7B Conductor model just hit SOTA on GPQA-Diamond and LiveCodeBench by orchestrating other LLMs inste…

SafetyDGX agent

NEW paper from Sakana AI (ICLR 2026). A 7B Conductor model just hit SOTA on GPQA-Diamond and LiveCodeBench by orchestrating other LLMs instead of solving problems itself. (great paper! bookmark it!) T

Online Self-Calibration Against Hallucination in Vision-Language Models

SafetyDGX agent

arXiv:2605.00323v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) often suffer from hallucinations, generating descriptions that include visual details absent from the input image.

Optimal Spatio-Temporal Decoupling for Bayesian Conformal Prediction

SafetyDGX agent

arXiv:2605.00432v1 Announce Type: new Abstract: Online Conformal Prediction (CP) struggles to balance temporal adaptability and structural stability. Feedback-driven methods (e.g., Adaptive Conformal

Optimizing Resource-Constrained Non-Pharmaceutical Interventions for Multi-Cluster Outbreak Control Using Hierarchical Reinforcement Learning

SafetyDGX agent

arXiv:2603.19397v2 Announce Type: replace Abstract: Non-pharmaceutical interventions (NPIs), such as diagnostic testing and quarantine, are crucial for controlling infectious disease outbreaks but are

Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations

SafetyDGX agent

arXiv:2605.00227v1 Announce Type: new Abstract: There are growing concerns about the risks posed by AI companion applications designed for emotional engagement. Existing safety evaluations often rely

PORTool: Importance-Aware Policy Optimization with Rewarded Tree for Multi-Tool-Integrated Reasoning

SafetyDGX agent

arXiv:2510.26020v2 Announce Type: replace Abstract: Multi-tool-integrated reasoning enables LLM-empowered tool-use agents to solve complex tasks by interleaving natural-language reasoning with calls t

Pose-Aware Diffusion for 3D Generation

SafetyDGX agent

arXiv:2605.00345v1 Announce Type: new Abstract: Generating pose-aligned 3D objects is challenging due to the spatial mismatches and transformation ambiguities inherent in decoupled canonical-then-rota

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

SafetyDGX agent

arXiv:2411.02327v4 Announce Type: replace Abstract: In the past year, video-based large language models (Video LLMs) have achieved impressive progress, particularly in their ability to process long vi

PrefMoE: Robust Preference Modeling with Mixture-of-Experts Reward Learning

SafetyDGX agent

arXiv:2605.00384v1 Announce Type: new Abstract: Preference-based reinforcement learning offers a scalable alternative to manual reward engineering by learning reward structures from comparative feedba

← Previous
1…167168169170171…212
Next →