AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,814 results
1 Jun 2026

Reading Between the Citations: A Typed Claim Network for Scientific Literature

SafetyDGX agent

arXiv:2605.30966v1 Announce Type: cross Abstract: Knowledge graphs over corpora of inter-referencing documents - scholarly papers, legal opinions, policy briefs - encode the topology of reference but

REAL: Regression-Aware Reinforcement Learning for LLM-as-a-Judge

SafetyDGX agent

arXiv:2603.17145v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed as automated evaluators that assign numeric scores to model outputs, a paradigm known a

Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses

SafetyDGX agent

arXiv:2504.11972v3 Announce Type: replace Abstract: Extractive QA tasks are commonly evaluated using Exact Match (EM) and F1-score, but these metrics often fail to reflect true model performance. Rece


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Reinforced sequential Monte Carlo for amortised sampling

SafetyDGX agent

arXiv:2510.11711v2 Announce Type: replace Abstract: This paper proposes a synergy of amortised and particle-based methods for sampling from distributions defined by unnormalised density functions. We

Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards

SafetyDGX agent

arXiv:2605.31328v1 Announce Type: new Abstract: Emergent misalignment (EM) is the surprising tendency of language models to become broadly misaligned after fine-tuning on narrowly misaligned examples.

Reinterpreting Safety Thresholds as Neuron Spiking Thresholds

SafetyDGX agent

arXiv:2605.30368v1 Announce Type: cross Abstract: Surrogate Safety Measures (SSMs) are extensively utilised in the evaluation of traffic risk in automated driving contexts. However, the majority of SS

Representation Collapse in Sequential Post-Training of Large Language Models

SafetyDGX agent

arXiv:2605.30524v1 Announce Type: new Abstract: Large language models are now adapted through chains of post-training stages rather than through a single instruction-tuning pass. This paper studies wh

Rethinking Multimodal Few-Shot 3D Point Cloud Segmentation: From Fused Refinement to Decoupled Arbitration

SafetyDGX agent

arXiv:2601.01456v2 Announce Type: replace-cross Abstract: In this paper, we revisit multimodal few-shot 3D point cloud semantic segmentation (FS-PCS), identifying a conflict in 'Fuse-then-Refine' para

// Reusable Context Engineering // Context bloat quietly kills long-horizon runs, but you can fix it from the outside without fine-tuning th…

SafetyDGX agent

// Reusable Context Engineering // Context bloat quietly kills long-horizon runs, but you can fix it from the outside without fine-tuning the underlying agent. (bookmark this) Context management is us

Revisiting Zeroth-Order Hessian Approximation: A Single-Step Policy Optimization Lens

SafetyDGX agent

arXiv:2605.30960v1 Announce Type: new Abstract: Accurate Zeroth-Order (ZO) Hessian estimation is a cornerstone of derivative-free methods, essential for tasks such as bilevel optimization, Bayesian in

Routing on the Stiefel Manifold: When Does Adaptive Subspace Selection Help for Cross-Domain EEG Decoding?

SafetyDGX agent

arXiv:2605.31043v1 Announce Type: cross Abstract: Cross-domain EEG decoding remains challenging despite advances in Riemannian deep learning: covariance matrices from different subjects occupy systema

Safeguarding Text-to-Image Generation via Inference-Time Prompt-Noise Optimization

SafetyDGX agent

arXiv:2412.03876v2 Announce Type: replace Abstract: Text-to-Image (T2I) diffusion models are widely recognized for their ability to generate high-quality and diverse images based on text prompts. Howe

Scalable Constrained Multi-Agent Reinforcement Learning via State Augmentation and Consensus for Separable Dynamics

SafetyDGX agent

arXiv:2605.30461v1 Announce Type: cross Abstract: We present a distributed approach for constrained Multi-Agent Reinforcement Learning (MARL) that combines state-augmented policy learning with distrib

Scaling Multi-Agent Environment Co-Design with Diffusion Models

SafetyDGX agent

arXiv:2511.03100v2 Announce Type: replace-cross Abstract: The agent-environment co-design paradigm jointly optimises agent policies and environment configurations in search of improved system performa

SCOPE: Selective Conformal Optimized Pairwise LLM Judging

SafetyDGX agent

arXiv:2602.13110v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used as scalable judges in pairwise evaluation, but they remain prone to miscalibration and bias

SDM-Q: Cost-Aware Staged Decision-Making for Multi-Omics Classification with Deep Q-Learning

SafetyDGX agent

arXiv:2605.31014v1 Announce Type: new Abstract: Multi-omics data provide complementary molecular characterizations of disease phenotypes and play an important role in disease diagnosis and subtype cla

Secure AI agents with Policy and Lambda interceptors in Amazon Bedrock AgentCore gateway

SafetyDGX agent

In this post, we use a lakehouse data agent to demonstrate how you can use Policy for deterministic access control and Lambda interceptors for dynamic validation. We then show how to combine Lambda in

Seeing Before Agreeing: Aligning Multi-Agent Consensus with Visual Evidence

SafetyDGX agent

arXiv:2605.30698v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved strong performance on visual question answering (VQA). To mitigate individual hallucinations and blind spo

Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures

SafetyDGX agent

arXiv:2605.30608v1 Announce Type: new Abstract: Learning a shared representation between spoken text and gesture is central to co-speech gesture retrieval, synthesis, and understanding, but remains ch

SemStruct: Contextualizing Semantic Embeddings with Structural Information for Schema Matching

SafetyDGX agent

arXiv:2605.30729v1 Announce Type: new Abstract: Schema matching is a fundamental step in integrating heterogeneous data sources. While Pre-trained Language Models (PLMs) have revolutionized this task

Simulation of collision avoidance behavior in crowd movement by data-driven approach

SafetyDGX agent

arXiv:2605.31210v1 Announce Type: cross Abstract: Crowd movement simulation is essential for pedestrian safety management and facility layout optimization. Data-driven models enhance trajectory predic

Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents

SafetyDGX agent

arXiv:2605.30723v1 Announce Type: new Abstract: LLM agents increasingly retrieve externally curated skills-procedural instructions retrieved at decision time-to improve performance on long-horizon int

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO

SafetyDGX agent

arXiv:2605.30789v1 Announce Type: cross Abstract: We identify a new dimension for enhancing rollout diversity in Group Relative Policy Optimization (GRPO) for LLMs. While GRPO relies on diverse rollou

Softly Constrained Denoisers for Diffusion Models Applied to Partial Differential Equations

SafetyDGX agent

arXiv:2512.14980v4 Announce Type: replace Abstract: Diffusion models have become a powerful generative prior for solutions of partial differential equations (PDEs). Existing approaches enforce physica

SpongeBob: Sync-Aware Harmonious Audio-Visual Generative Editing

SafetyDGX agent

arXiv:2605.25193v2 Announce Type: replace Abstract: Visual and acoustic events in the physical world are inherently coupled, yet existing video editing methods typically adopt decoupled pipelines, lac

Stateful Online Monitoring Catches Distributed Agent Attacks

SafetyDGX agent

arXiv:2605.31593v1 Announce Type: cross Abstract: Language models can find thousands of severe software vulnerabilities, and agents are increasingly being misused for cyberattacks. To avoid detection,

Structural Bias Beyond Homophily: A Study of Fairness in Link Prediction

SafetyDGX agent

arXiv:2602.11802v2 Announce Type: replace Abstract: Graph link prediction (LP) plays a critical role in socially impactful applications such as job recommendation and friendship formation, making fair

Structure-Induced Information for Rerooting Levin Tree Search

SafetyDGX agent

arXiv:2605.30664v1 Announce Type: new Abstract: Subgoal-based policy tree search, which uses a policy to guide search, is effective for complex single-agent deterministic problems but often relies on

Supervised Learning as Lossy Compression: Characterizing Generalization and Sample Complexity via Finite Blocklength Analysis

SafetyDGX agent

arXiv:2602.04107v2 Announce Type: replace Abstract: This paper presents a novel information-theoretic perspective on generalization in machine learning by framing the learning problem within the conte

Supervised Training Rapidly Degrades Early Visual Cortex Alignment Across Biologically Plausible Learning Rules

SafetyDGX agent

arXiv:2605.30556v1 Announce Type: new Abstract: Random, untrained neural networks consistently match or exceed trained networks in representational similarity to early visual cortex. This puzzling fin

Surface Constraint Policy for Learning Surface-Constrained and Dynamically Feasible Robot Skills

SafetyDGX agent

arXiv:2605.31321v1 Announce Type: new Abstract: Diffusion-based imitation learning methods have driven rapid progress in robot dexterous manipulation tasks. However, they have limitations when applied

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation

SafetyDGX agent

arXiv:2511.11440v3 Announce Type: replace-cross Abstract: Performance gains of Vision Language Models (VLMs) obtained by fine-tuning are generally based on ad hoc data collection and annotation of rea

TALON: Token-Aligned Lightweight Adapters for 6-DoF Spacecraft Pose Estimation

SafetyDGX agent

arXiv:2605.31217v1 Announce Type: new Abstract: Monocular 6-DoF spacecraft pose estimation methods predominantly process individual frames, discarding the temporal information present in an image sequ

TARIC: Memory-Augmented Traversability-Aware Outdoor VLN under Interrupted Semantic Cues

SafetyDGX agent

arXiv:2605.31121v1 Announce Type: cross Abstract: Outdoor vision-language navigation (VLN) in long-range, open-world environments is frequently disrupted by semantic-cue interruptions, where informati

Task-Focused Memorization for Multimodal Agents

SafetyDGX agent

arXiv:2605.31075v1 Announce Type: new Abstract: Long-term memory is essential for multimodal agents to build coherent experience, accumulate world knowledge, and achieve continual learning. However, c

THE DEFINITIVE AI CONVERSATION OF THE YEAR. MUST LISTEN. @JG_Nuke @GaryMarcus @MacrostrategyP https://open.substack.com/pub/georgenoble/p/ai…

SafetyDGX agent

THE DEFINITIVE AI CONVERSATION OF THE YEAR. MUST LISTEN. @JG_Nuke @GaryMarcus @MacrostrategyP https://open.substack.com/pub/georgenoble/p/ai-the-biggest-capital-misallocation-d18?r=35saq&utm_medium=io

The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement

SafetyDGX agent

arXiv:2605.30888v1 Announce Type: new Abstract: Building strong reward models (RMs) for language model alignment is bottlenecked by the cost and difficulty of acquiring diverse and reliable preference

The Global Landscape of Environmental AI Regulation: From the Cost of Reasoning to a Right to Green AI

SafetyDGX agent

arXiv:2603.00068v2 Announce Type: replace-cross Abstract: Artificial intelligence (AI) systems impose substantial and growing environmental costs, yet transparency about these impacts has declined eve

The OpenAI Foundation is doing a lot of wonderful things. Helping society become resilient to AI is going to be incredibly important. Much m…

SafetyDGX agent

The OpenAI Foundation is doing a lot of wonderful things. Helping society become resilient to AI is going to be incredibly important. Much more to come here! AI is advancing quickly. Society’s ability

The Refutability Gap: Challenges in Validating Reasoning by Large Language Models

SafetyDGX agent

arXiv:2601.02380v4 Announce Type: replace-cross Abstract: Recent reports claim that Large Language Models (LLMs) have achieved the ability to derive new science and exhibit human-level general intelli

The Sword, Shield, and Achilles' Heel: Characterizing the Linguistic Inductive Bias of Large Language Models for Spatial Reasoning in Navigation Planning

SafetyDGX agent

arXiv:2605.31404v1 Announce Type: cross Abstract: Large Language Model (LLM)-based navigation systems commonly construct explicit spatial representations (e.g., topological graphs, semantic raster map

“The technology worked. The value didn’t arrive,” Bain concluded in the report. https://www.bloomberg.com/news/newsletters/2026-06-01/bain-s…

SafetyDGX agent

“The technology worked. The value didn’t arrive,” Bain concluded in the report. https://www.bloomberg.com/news/newsletters/2026-06-01/bain-survey-ai-delivers-less-cost-reduction-than-many-firms-predic

This was right five years ago, and still is: “Large scale pretrained models are certainly likely to figure prominently in artificial intelli…

SafetyDGX agent

This was right five years ago, and still is: “Large scale pretrained models are certainly likely to figure prominently in artificial intelligence for the near future, and play an important role in com

Traceable by Design: An LLM Pipeline and Dashboard for EU Regulatory Consultation Analysis

SafetyDGX agent

arXiv:2605.30995v1 Announce Type: cross Abstract: Public consultations generate large volumes of data in the form of stakeholder submissions that are practically unfeasible to analyse manually. We pre

Trust-Region Behavior Blending for On-Policy Distillation

SafetyDGX agent

arXiv:2605.31159v1 Announce Type: cross Abstract: On-policy distillation (OPD) trains a student on prefixes sampled from its own policy while matching a stronger teacher. This addresses the prefix mis

TunerDiT: Training-free Progressive Steering of Diffusion Transformer for Multi-Event Video Generation

SafetyDGX agent

arXiv:2605.31590v1 Announce Type: cross Abstract: Text-to-video (T2V) generation faces challenging questions when generating videos with long horizons containing multiple events. Inspired by the intri

TUX: Measuring Human--AI Tacit Understanding

SafetyDGX agent

arXiv:2605.30930v1 Announce Type: cross Abstract: As large language models (LLMs) increasingly act as collaborative partners, human--AI alignment is often evaluated through explicit task success, accu

Ubiquity of Emergent Hebbian Dynamics in Regularized Learning

SafetyDGX agent

arXiv:2505.18069v3 Announce Type: replace Abstract: Hebbian and anti-Hebbian plasticity are widely observed in the brain and are classically modeled as mechanistic, local homosynaptic rules stabilized

Uncertainty-Aware and Temporally Regulated Expert Advice in Reinforcement Learning for Autonomous Driving

SafetyDGX agent

arXiv:2605.30576v1 Announce Type: new Abstract: Exploration in reinforcement learning for autonomous driving is inherently unsafe: agents must experience novel behaviors to learn, yet exploration can

Unfolding Generative Flows with Koopman Operators: Trajectory-Preserving Linearization

SafetyDGX agent

arXiv:2506.22304v3 Announce Type: replace-cross Abstract: Continuous Normalizing Flows (CNFs) enable elegant generative modeling but remain bottlenecked by their iterative nature requiring costly samp

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception

SafetyDGX agent

arXiv:2605.31521v1 Announce Type: new Abstract: Semantic speech tokenizers have become a widely used interface for Audio-LLMs, owing to their compact single-codebook design and strong linguistic align

UniRTL: Unifying Code and Graph for Robust RTL Representation Learning

SafetyDGX agent

arXiv:2605.31040v1 Announce Type: new Abstract: Developing effective representations for register transfer level (RTL) designs is crucial for accelerating the hardware design workflow. Existing approa

Unsupervised Defect Detection for Surgical Instruments

SafetyDGX agent

arXiv:2509.21561v2 Announce Type: replace Abstract: Ensuring the safety of surgical instruments requires reliable detection of visual defects. However, manual inspection is prone to error, and existin

Used Car Salesbots? Honesty and Credulity of LLMs as Bargaining Agents under Partial Information

SafetyDGX agent

arXiv:2605.31445v1 Announce Type: cross Abstract: In this work we study agents in simulated bargaining scenarios, where a buyer and a seller communicate through a text channel and attempt to negotiate

UXR PoV for Neuroinclusive Emotion Regulation

SafetyDGX agent

arXiv:2605.31131v1 Announce Type: cross Abstract: Attention-deficit/hyperactivity disorder (ADHD) is a psychiatric disorder which presents itself in individuals through patterns of developmentally ina

Value Functions as Supermartingale Certificates

SafetyDGX agent

arXiv:2605.31524v1 Announce Type: new Abstract: Certification methods for stochastic systems provide sufficient proof rules, based on real-valued supermartingale certificates, to determine the almost-

VeriGate: Verifier-Gated Step-Level Supervision for GRPO

SafetyDGX agent

arXiv:2605.30451v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) is an effective recipe for training reasoning models with verifier-based outcome rewards, but its supervision

Vision-Language Models Suppress Female Representations Under Ambiguous Input

SafetyDGX agent

arXiv:2605.31556v1 Announce Type: cross Abstract: Alignment teaches vision-language models (VLMs) to avoid expressing demographic biases, and when gender is clearly visible they largely succeed. Far l

Wall-OSS-0.5 Technical Report

SafetyDGX agent

arXiv:2605.30877v1 Announce Type: new Abstract: Large-scale Vision-Language-Action (VLA) pretraining is increasingly adopted as the foundation for robot policies, yet the evidence for pretrained VLAs

@webisticsdawg @GaryMarcus Personally I think everyone w a 401(k) should be absolutely livid. I'm disgusted by this, it's so brazen. How dar…

SafetyDGX agent

@webisticsdawg @GaryMarcus Personally I think everyone w a 401(k) should be absolutely livid. I'm disgusted by this, it's so brazen. How dare the richest man in the world pick working Americans' pocke

← Previous
1…103104105106107…214
Next →