AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,600 results
6 Aug 2026

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation

SafetyDGX agent

arXiv:2608.04788v1 Announce Type: cross Abstract: Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on

Among proponents of neurosymbolic architectures, there had been some debate over the years about whether the outer level would be symbolic (…

SafetyDGX agent

Among proponents of neurosymbolic architectures, there had been some debate over the years about whether the outer level would be symbolic (i.e. a harness that calls neural models) or whether the oute

An Inline Control Architecture for Language Models in Intelligent Transportation Systems

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.04065v1 Announce Type: cross Abstract: Vehicle-to-everything (V2X) systems increasingly incorporate large language models (LLMs) for semantic tasks such as message summarization, operator a

Approximate Multi-Objective Search Under Rulebooks

SafetyDGX agent

arXiv:2608.04398v1 Announce Type: cross Abstract: Robotic planning often involves multiple objectives with complex priority relationships, such as safety, efficiency, and regulatory compliance. Rulebo

Arnold: A multi-task, multi-embodiment muscle transformer policy

SafetyDGX agent

arXiv:2508.18066v2 Announce Type: replace-cross Abstract: Controlling high-dimensional and nonlinear musculoskeletal models of the human body is a foundational scientific challenge. Recent machine lea

ATLAS: Adaptive Topological Learning with Abstract Successors for Continual Learning

SafetyDGX agent

arXiv:2608.04334v1 Announce Type: cross Abstract: Contemporary model-free reinforcement learning algorithms can achieve very high performance, but have low sample efficiency and are not robust to chan

Attention Fusion for Bridge Deck Delamination Detection

SafetyDGX agent

arXiv:2512.20113v4 Announce Type: replace Abstract: Subsurface delaminations in reinforced concrete bridge decks escape conventional visual inspection, and the two principal sensing techniques used to

Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification

SafetyDGX agent

arXiv:2607.16868v2 Announce Type: replace Abstract: Large Language Models often produce confidently stated yet unreliable outputs, posing critical challenges for deployment in safety-sensitive applica

Beyond the Dirac Delta: Mitigating Diversity Collapse in Reinforcement Fine-Tuning for Versatile Image Generation

SafetyDGX agent

arXiv:2601.12401v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as a powerful paradigm for fine-tuning large-scale generative models, such as diffusion and flow model

Breadcrumbing Search Agents

SafetyDGX agent

arXiv:2608.04565v1 Announce Type: cross Abstract: LLM-based search agents are widely used for information-seeking tasks, but their reliance on external tool returns introduces a critical security risk

BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

SafetyDGX agent

arXiv:2608.05042v1 Announce Type: new Abstract: Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot m

Calibrating Artificial Guilt: Neurally Grounded Reward Shaping for Prosocial Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2608.04663v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning often adds social terms to individual rewards, yet the scale of those terms is usually chosen by hand. We

Calibrating Transformer Attention via Task-Space Sensitivity Feedback

SafetyDGX agent

arXiv:2512.20661v2 Announce Type: replace Abstract: Transformer-based pre-trained language models (PLMs) excel in text classification but suffer from attention dilution and attention sink effects, for

CARGO-VL: Counterfactual Arbitration with Risk-Constrained Group Optimization for Vision-Language Models

SafetyDGX agent

arXiv:2608.04509v1 Announce Type: new Abstract: Vision-language systems combine images with retrieved text, but these sources can disagree or jointly fail to support an answer. Reliable models must id

CheckOne: Lightweight Fault Detection and Mitigation for Vision Transformers

SafetyDGX agent

arXiv:2608.04035v1 Announce Type: cross Abstract: The wide adoption of Vision Transformers (ViTs) in safety-critical applications raises reliability concerns related to hardware faults. Algorithm-Base

CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models

SafetyDGX agent

arXiv:2608.04302v1 Announce Type: new Abstract: Benchmarking video-language models has largely focused on short clips and single-sentence metrics, leaving open whether current systems can generate acc

C’mon, it’s not game over for Google Seven reasons why not, excerpted from my newsletter:

SafetyDGX agent

C’mon, it’s not game over for Google Seven reasons why not, excerpted from my newsletter: Jeff Dean and Demis Hassabis are the two most important AI executives at Google. Jeff is leaving and Demis is

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention

SafetyDGX agent

arXiv:2608.04396v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have driven significant progress in robotic manipulation, yet they fundamentally struggle with the vision-override p

Compass: Continuously Aligning Social Media Feeds via In-Situ Reflections

SafetyDGX agent

arXiv:2608.04274v1 Announce Type: cross Abstract: Social media recommendation feeds often optimize for users' immediate impulses rather than preferences they would hold after deeper reflection. Some s

Contrastive Diffusion Alignment: Learning Structured Latents for Controllable Generation

SafetyDGX agent

arXiv:2510.14190v3 Announce Type: replace Abstract: Diffusion models excel at generation, but their latent spaces are high dimensional and not explicitly organized for interpretation or control. We in

Control agent behaviors and cost beyond a single action: new capabilities in Amazon Bedrock AgentCore

SafetyDGX agent

Learn about new capabilities in Amazon Bedrock AgentCore: temporal policies powered by Dogwood, a new open source policy language for AI agents, and rate limiting on the gateway. These features give y

Corrigibility Transformation: Constructing Goals That Accept Updates

SafetyDGX agent

arXiv:2510.15395v2 Announce Type: replace Abstract: An AI agent will learn a desired goal more effectively if it does not resist the training process, but many partially learned goals incentivize an A

COSMO: Consensus-Driven Shift Modulation for Source-Free Domain Adaptation

SafetyDGX agent

arXiv:2608.04604v1 Announce Type: new Abstract: Source-free domain adaptation (SFDA) adapts a source-trained model to an unlabeled target domain without source data, a practical setting under privacy

critical history and context that a lot of people have conveniently forgotten

SafetyDGX agent

critical history and context that a lot of people have conveniently forgotten For a very long time most high-performing AI models were end-to-end neural models; vector input -> vector output, with onl

Curiosity-Diffuser: Curiosity Guide Diffusion Models for Reliability

SafetyDGX agent

arXiv:2503.14833v2 Announce Type: replace-cross Abstract: One of the bottlenecks in robotic intelligence is the instability of neural network models. This leads to risks when applying intelligence in

DAC-Pose: Dual-Agent Collaborative Framework for Pose-Guided Human Generation

SafetyDGX agent

arXiv:2608.04622v1 Announce Type: new Abstract: AI agents have emerged as a powerful new paradigm in generative image synthesis, enabling systems to perform complex semantic reasoning rather than pass

Data-Aware and Scalable Sensitivity Analysis for Decision Tree Ensembles

SafetyDGX agent

arXiv:2602.07453v2 Announce Type: replace Abstract: Decision tree ensembles are widely used in critical domains, making robustness and sensitivity analysis essential to their trustworthiness. We study

DataRx: Missingness-Aware Sampling for Safer Large Language Model Task-Specific Fine-Tuning

SafetyDGX agent

arXiv:2608.04322v1 Announce Type: new Abstract: Task-specific fine-tuning can improve the performance of large language models (LLMs) on downstream tasks. However, our study reveals that task-specific

Differentiating Through Dual Prices: End-to-End Policy Learning Under Capacity Constraints

SafetyDGX agent

arXiv:2608.04669v1 Announce Type: new Abstract: Many social services assign scarce resources, such as housing assistance or hospital interventions, to people who arrive one at a time: each arrival mus

DXC partners with Primary on zero-trust security for enterprise AI

SafetyDGX agent

DXC Technology Co. today announced a partnership with security startup Primary that makes the information technology services company the exclusive managed services partner for Primary’s zero-trust pl

Enabling Urgency-aware Robot Swarm Intralogistics using Smart IoT Tags

SafetyDGX agent

arXiv:2608.04721v1 Announce Type: new Abstract: Warehouse items differ in how urgently they must be moved: perishable goods, pharmaceutical shipments, and just-in-time production materials must be del

EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment

SafetyDGX agent

arXiv:2608.04472v1 Announce Type: cross Abstract: The development of foundation models (FMs) is crucial for advancing endoscopic image analysis. However, existing endoscopy FMs mainly rely on self-sup

Evaluating the Diagnostic Robustness of Vision-Language Models Under Visual and Textual Perturbations

SafetyDGX agent

arXiv:2608.04885v1 Announce Type: cross Abstract: Standard accuracy metrics for VLMs often mask significant reliability failures in sensitive domains. In this work, we utilize a histopathology-validat

Exact Model-Free Policy Iteration for Co-safe LTL Planning

SafetyDGX agent

arXiv:2608.05047v1 Announce Type: cross Abstract: This work studies model-free reinforcement learning for co-safe linear temporal logic (sc-LTL) objectives in finite Markov decision processes, which c

Feedback Loops and Code Perturbations in LLM-based Software Engineering: A Case Study on a C-to-Rust Translation System

SafetyDGX agent

arXiv:2512.02567v2 Announce Type: replace-cross Abstract: The advent of strong generative AI has a considerable impact on various software engineering tasks such as code repair, test generation, or la

Fewer Tokens, Smaller Cache: Reward-Coordinated Efficient Reasoning

SafetyDGX agent

arXiv:2608.04771v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) excel on complex tasks through long chain-of-thought (CoT) reasoning, but their lengthy intermediate steps cause severe ov

Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation

SafetyDGX agent

arXiv:2602.19161v2 Announce Type: replace Abstract: Latent diffusion models have enabled high-quality video synthesis, yet their inference remains costly and time-consuming. As diffusion transformers

FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory

SafetyDGX agent

arXiv:2608.04530v1 Announce Type: new Abstract: GUI agents must remember both useful experience from earlier tasks and unfinished progress in the current interaction. Latent memory offers a compact so

From Brute Force to Semantic Insight: Performance-Guided Data Transformation Design with LLMs

SafetyDGX agent

arXiv:2601.03808v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved notable performance in code synthesis; however, data-aware augmentation remains a limiting factor, handle

Gary Marcus won! @GaryMarcus

SafetyDGX agent

Gary Marcus won! @GaryMarcus I would have assumed it was fairly obvious, but in case it's not: a million-line codebase (also known as a 'harness'), running at inference time, orchestrating thousands o

Generative Optimization for Incentivized Advertising with Global Level Constraints

SafetyDGX agent

arXiv:2608.04421v1 Announce Type: cross Abstract: Incentivized advertising allocates monetary or virtual rewards to drive user engagement, where a key challenge is optimizing continuous incentive magn

GeoReward: Mitigating Contextual Variable Overestimation in Vision-Language Models for Cross-Market Preference Prediction

SafetyDGX agent

arXiv:2608.04504v1 Announce Type: cross Abstract: Vision-language models excel in many multimodal tasks but remain prone to a subtle yet impactful failure mode: they tend to overestimate dominant visu

GFlowNet Training by Policy Gradients

SafetyDGX agent

arXiv:2408.05885v3 Announce Type: replace Abstract: Generative Flow Networks (GFlowNets) have been shown effective to generate combinatorial objects with desired properties. We here propose a new GFlo

Governing Execution Risk in Agentic AI Systems: A Trajectory-Guided Framework for Red Teaming

SafetyDGX agent

arXiv:2608.04018v1 Announce Type: cross Abstract: AI agents are increasingly embedded in organizational workflows, where they interact with external information sources and invoke digital tools to per

Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent

SafetyDGX agent

arXiv:2608.04772v1 Announce Type: cross Abstract: Scaling supervision for multi-turn medical agents is difficult because expert dialogue annotation is costly and clinical conversations are privacy-res

HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents

SafetyDGX agent

arXiv:2608.02009v2 Announce Type: replace Abstract: Retrieval-augmented search agents answer multi-hop questions by repeatedly issuing search queries and accumulating evidence. This creates a stopping

HCRide: Harmonizing Passenger Fairness and Driver Preference for Human-Centered Ride-Hailing

SafetyDGX agent

arXiv:2508.04811v2 Announce Type: replace Abstract: Order dispatch systems play a vital role in ride-hailing services, which directly influence operator revenue, driver profit, and passenger experienc

Integrated Noise and Safety Management in UAM via A Unified Reinforcement Learning Framework

SafetyDGX agent

arXiv:2508.16440v2 Announce Type: replace-cross Abstract: Urban Air Mobility (UAM) envisions the widespread use of small aerial vehicles to transform transportation in dense urban environments. Howeve

Interpreting GFlowNets for Drug Discovery: What probes can and cannot show

SafetyDGX agent

arXiv:2511.19264v2 Announce Type: replace-cross Abstract: Generative Flow Networks (GFlowNets) construct molecules through sequential decisions, but their internal policies remain opaque, limiting ado

It was the verification problem all along, while the masses were distracted by the alignment problem. Recursive self improvement? How does t…

SafetyDGX agent

It was the verification problem all along, while the masses were distracted by the alignment problem. Recursive self improvement? How does the observer observe itself and know that it changed for the

Joint UAV Flight and Opportunistic Routing under Reinforcement Learning for Delay-Tolerant Networks

SafetyDGX agent

arXiv:2608.04590v1 Announce Type: new Abstract: The growing deployment of delay-tolerant networks (DTNs) has made store-carry-forward (SCF) communication indispensable under sparse connectivity. Howev

Language Models Generalize to Human-like Word Order Preferences

SafetyDGX agent

arXiv:2608.05028v1 Announce Type: new Abstract: A central question in language acquisition is whether linguistic biases can emerge from general learning mechanisms operating over underdetermined input

Learning to Resolve Neutron Resonances with Fully Convolutional Neural Networks

SafetyDGX agent

arXiv:2608.04027v1 Announce Type: new Abstract: This work investigates the feasibility of augmenting traditional R-Matrix codes with a robust machine learning framework for automatically detecting neu

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control

SafetyDGX agent

arXiv:2608.05084v1 Announce Type: new Abstract: Diffusion policies are a powerful policy class for continuous control, but their iterative denoising process creates a substantial computational bottlen

Lesion Detection in CT with Frozen Self-Distilled Features: SALT, a Spatially Adaptive Label-Guided Temperature

SafetyDGX agent

arXiv:2608.05100v1 Announce Type: new Abstract: Self-supervised pretraining objectives are spatially uniform: the teacher temperature and the per-patch loss weight are identical everywhere in the imag

LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment

SafetyDGX agent

arXiv:2606.18709v2 Announce Type: replace Abstract: Existing work on LLM-based educational assessment has focused largely on item difficulty, but difficulty alone does not indicate whether an item mea

Local Violation Certification for Linear Predict-Then-Optimize Pipelines

SafetyDGX agent

arXiv:2608.04474v1 Announce Type: new Abstract: Data-driven decision pipelines combining predictive machine learning models with downstream optimization software are increasingly used to make high-sta

Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions

SafetyDGX agent

arXiv:2608.02491v2 Announce Type: replace Abstract: Language models have taken on the role of a very new type of technology, by virtue of their 'human-ness' and rapid integration into users' daily liv

Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning

SafetyDGX agent

arXiv:2608.04347v1 Announce Type: new Abstract: Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competen

Manipulation-Proof Oblivious Audits against Deceptive Model Providers

SafetyDGX agent

arXiv:2608.04365v1 Announce Type: new Abstract: Audits have emerged as a critical instrument for algorithmic governance, providing a mechanism for external scrutiny and governance of machine learning

← Previous
1…89101112…210
Next →