AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,357 results
13 Aug 2026

Variable Selection in the Context of AI Fairness

SafetyDGX agent

arXiv:2608.11251v1 Announce Type: cross Abstract: Fairness in AI systems has become more important with recent regulatory demands, such as the EU AI Act. Traditional approaches often do not take into

Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study

SafetyDGX agent

arXiv:2608.11649v1 Announce Type: new Abstract: As users increasingly turn to Large Language Models (LLMs) for information and advice on political matters, particularly during election periods, the po

Why AI Detection Fails for Academic Integrity

SafetyDGX agent

arXiv:2608.11256v1 Announce Type: new Abstract: Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
12 Aug 2026

A Joint-Distribution Route to Fair Representations with Continuous Sensitive Attributes

SafetyDGX agent

arXiv:2608.10470v1 Announce Type: new Abstract: Fair representation learning with a continuous sensitive attribute S requires a representation Z that is statistically independent of S. Existing criter

Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents

SafetyDGX agent

arXiv:2608.11110v1 Announce Type: new Abstract: When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compa

// Actions Speak Louder Than Words // Multilingual agent evaluation compares final answers and throws the trajectory away. The trajectory fi…

SafetyDGX agent

// Actions Speak Louder Than Words // Multilingual agent evaluation compares final answers and throws the trajectory away. The trajectory fixes cost, latency, failure mode, and auditability. New resea

AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss

SafetyDGX agent

arXiv:2608.11205v1 Announce Type: new Abstract: Frechet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-le

APCReg: Anatomical-Prior-Guided Coarse-to-Fine CBCT--IOS Registration via Multi-View Projection and Reliability-Controlled Residual Correction

SafetyDGX agent

arXiv:2608.09993v1 Announce Type: cross Abstract: Registration between cone-beam computed tomography (CBCT) and intraoral scans (IOS) is essential for patient-specific surgical planning. However, disp

Beyond Forecasting: Recasting Volatility Control as a Routing Problem

SafetyDGX agent

arXiv:2608.10375v1 Announce Type: cross Abstract: Volatility control converts risk estimates into portfolio exposure, yet existing approaches often rely on a fixed volatility estimator or a pre-define

Big Tech's reliance on OpenAI and Anthropic for growth is systemically pervasive. Both companies are massively unprofitable — and much of wh…

SafetyDGX agent

Big Tech's reliance on OpenAI and Anthropic for growth is systemically pervasive. Both companies are massively unprofitable — and much of what they spend is Big Tech's own money, counted right back as

BooST: Bridging Semantics and Motions for Efficient Skill Transfer

SafetyDGX agent

arXiv:2608.10600v1 Announce Type: cross Abstract: Skill abstraction---the process of learning reusable and temporally extended behaviors---has emerged as a key paradigm for improving sample efficiency

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning

Model ReleasesDGX agent

arXiv:2608.10204v1 Announce Type: new Abstract: Safe reinforcement learning maximizes reward subject to safety constraints. For Constrained Markov Decision Processes, the linear-programming view over

CARE: Confidence-Aware Reasoning for Reliable Medical VQA

SafetyDGX agent

arXiv:2608.10964v1 Announce Type: cross Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual qu

Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory

SafetyDGX agent

arXiv:2608.09937v1 Announce Type: new Abstract: Recent work in NLP has probed large language models for their understanding of cultural norms across countries. However, this work typically considers d

ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering

SafetyDGX agent

arXiv:2608.10996v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has been especially effective in mathematics and coding, where answers can be checked automatically. Many

Convergence of Sign-based Random Reshuffling Algorithms for Nonconvex Optimization

SafetyDGX agent

arXiv:2310.15976v4 Announce Type: replace Abstract: signSGD is attractive in nonconvex optimization because it communicates sign-valued rather than full-precision gradients. Several standard analyses

Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning

SafetyDGX agent

arXiv:2608.10473v1 Announce Type: cross Abstract: Offline-to-online (O2O) reinforcement learning aims to leverage policies pretrained on static datasets while improving them through online interaction

Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition Agents

SafetyDGX agent

arXiv:2608.10441v1 Announce Type: cross Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expen

DIMOS: Disentangling Instance-level Moving Object Segmentation

SafetyDGX agent

arXiv:2606.12826v2 Announce Type: replace-cross Abstract: Moving instance segmentation (MIS) attracts increasing attention due to its broad applications in traffic surveillance, autonomous driving, an

Do AI weather models miss extremes?

SafetyDGX agent

arXiv:2608.09972v1 Announce Type: cross Abstract: First-generation AI weather models are often reported to underperform at extremes, mostly in reanalysis-based evaluations of deterministic regression

Do Time-Series Forecasters Use the Right History: Recoverability, Recovery, and Functional Use of Temporal Delays

SafetyDGX agent

arXiv:2608.10433v1 Announce Type: new Abstract: Forecast accuracy does not tell us which past inputs produced a prediction. We separate three questions for time-series models with known delay structur

Dual-Loop Self-Evolution via Verifiable Emotion Feedback for Multi-Turn Empathetic Dialogue

SafetyDGX agent

arXiv:2608.10626v1 Announce Type: new Abstract: Large language models have demonstrated conversational capabilities, yet empathetic competence remains challenging. Empathetic support is inherently mul

Dual Space Preconditioning for Gradient Descent in the Overparameterized Regime

SafetyDGX agent

arXiv:2603.10485v3 Announce Type: replace-cross Abstract: In this work, we study the convergence properties of the Dual Space Preconditioned Gradient Descent, encompassing optimizers such as Normalize

Efficient Hypergradient Descent for Inverse Reinforcement Learning

SafetyDGX agent

arXiv:2608.11052v1 Announce Type: new Abstract: Inverse reinforcement learning (IRL) aims to recover a reward function under which the resulting policy reproduces the behavior observed in expert demon

Enhancing Automated Essay Scoring With Three Techniques: Two-Stage Fine-Tuning, Score Alignment, and Self-Training

SafetyDGX agent

arXiv:2602.01747v2 Announce Type: replace Abstract: Automated Essay Scoring (AES) plays a crucial role in education by providing scalable and efficient assessment tools. However, in real-world setting

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes

SafetyDGX agent

arXiv:2608.10209v1 Announce Type: new Abstract: Feedback signals used to train Large Language Models (LLMs) are the primary driver of their behavior and our main lever for instilling alignment with hu

Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases

SafetyDGX agent

arXiv:2608.10503v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. The NL

FedCGR: Federated Cross-Domain Generative Recommendation

SafetyDGX agent

arXiv:2608.10929v1 Announce Type: new Abstract: Cross-domain recommendation (CDR) transfers preference knowledge across related domains, but federated deployment makes cross-domain alignment difficult

FoR-SALE: Frame of Reference-guided Spatial Adjustment in LLM-based Diffusion Editing

SafetyDGX agent

arXiv:2509.23452v2 Announce Type: replace-cross Abstract: Current text-to-image generation models, even state-of-the-art models, exhibit a significant performance gap when spatial expressions are desc

From Prediction to Incrementality: Causal Optimization for Large-Scale Targeting and Recommendation

SafetyDGX agent

arXiv:2608.10182v1 Announce Type: cross Abstract: Large-scale targeting and recommendation systems are typically built around predictive scores fed into heuristic or local allocation. When the busines

Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Models

SafetyDGX agent

arXiv:2601.17387v3 Announce Type: replace Abstract: Multilingual speech-text models rely on cross-modal language alignment to transfer knowledge between speech and text, but it remains unclear whether

Hierarchical Empirical-Bayes Naive Bayes: Minimax Smoothing and Calibration with AODE Extension

SafetyDGX agent

arXiv:2608.11162v1 Announce Type: new Abstract: The Naive Bayes (NB) classifier remains a standard choice for categorical data, yet its widely used smoothing rules, such as Laplace, Lidstone, Krichevs

Hip Energized Monopedal Hopping

SafetyDGX agent

arXiv:2608.10387v1 Announce Type: new Abstract: We present a novel stepping strategy for pitch unlocked planar monopeds where the reaction torques from stabilizing pitch with a conventional PD + feedf

IADD-TR: Intervention-Aware Dynamics Decoupling with Targeted Regularization for Model-Based Reinforcement Learning

SafetyDGX agent

arXiv:2608.10634v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL), which learns environment dynamics to generate synthetic experience, is a promising approach to sample-efficie

INSIDE the Student's Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators

SafetyDGX agent

arXiv:2608.10492v1 Announce Type: new Abstract: Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them. In education, w

IO Factory: Simulating AI-Enabled Influence Campaigns at Scale

SafetyDGX agent

arXiv:2608.10920v1 Announce Type: new Abstract: We introduce IO Factory, an AI-driven framework for simulating information and influence campaigns as fully integrated, traceable processes. The threat

Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen)

SafetyDGX agent

arXiv:2509.06191v2 Announce Type: replace-cross Abstract: Recent 3D generative models, which are capable of generating full object shapes from just a few images, now open up new opportunities in robot

Leveraging Large Language Models for Causal Discovery: a Constraint-based, Argumentation-driven Approach

SafetyDGX agent

arXiv:2602.16481v2 Announce Type: replace Abstract: Causal discovery seeks to uncover causal relations from data, typically represented as causal graphs, and is essential for predicting the effects of

LoRCA: LoRA Cycle Adaptation for Histology to HiP-CT Translation with DINOv3

SafetyDGX agent

arXiv:2608.10002v1 Announce Type: cross Abstract: Hierarchical Phase-Contrast Tomography (HiP-CT) is a synchrotron based X-ray imaging technique that enables non-destructive, volumetric imaging of int

MARCO: Click-Intent Decomposition for Calibrated Ads Conversion Prediction

SafetyDGX agent

arXiv:2608.10562v1 Announce Type: new Abstract: Not all clicks are equal. Industrial ads ranking decouples conversion probability into click-through rate (CTR) and post-click conversion rate (CVR), ye

MarkNull: Model-Agnostic Watermark Removal in AI-Generated Images via On-Manifold Latent Manipulation

SafetyDGX agent

arXiv:2608.10166v1 Announce Type: cross Abstract: Digital watermarking has emerged as a critical technique for provenance and copyright attribution in AI-generated imagery, yet its robustness against

MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale

SafetyDGX agent

arXiv:2608.10333v1 Announce Type: new Abstract: LLM agents execute heterogeneous sequences of model calls within a single task: some invocations require careful reasoning, while others are structured

MIDAS: Mutual Information Disentanglement with Uncertainty-Aware Fusion for Incomplete Multimodal Sentiment Analysis

SafetyDGX agent

arXiv:2608.09986v1 Announce Type: new Abstract: Most existing multimodal sentiment analysis approaches assume access to complete multimodal inputs. However, real-world applications frequently encounte

Mitigating Bus Bunching with Reinforcement Learning Enhanced by Semantic Stop Embedding

SafetyDGX agent

arXiv:2608.10207v1 Announce Type: new Abstract: Bus bunching degrades service regularity and increases passenger waiting in high-frequency transit. Existing reinforcement-learning-based holding contro

Most biomedical publications show signs of LLM-assisted writing

SafetyDGX agent

arXiv:2608.10715v1 Announce Type: cross Abstract: Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing can be valua

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models

SafetyDGX agent

arXiv:2608.10864v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong image and video understanding, yet their visual-spatial representations remain geometrically fragile,

Multiplayer Nash Preference Optimization

SafetyDGX agent

arXiv:2509.23102v4 Announce Type: replace Abstract: Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences. Ho

Observational Policy Ranking for SMB Financial Guidance from Multi-Action Accounting Logs

SafetyDGX agent

arXiv:2608.10050v1 Announce Type: new Abstract: Small and medium-sized businesses need timely financial guidance, yet historical accounting logs record self-selected and often co-occurring business ch

OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation

SafetyDGX agent

arXiv:2603.19201v3 Announce Type: replace Abstract: Contact-rich manipulation tasks, such as wiping and assembly, require accurate perception of contact forces, friction changes, and state transitions

On the Importance of Geometric Nonlinearity and Temperature-Dependent Properties in Multi-Material Thermo-Mechanical Topology Optimization

SafetyDGX agent

arXiv:2608.10344v1 Announce Type: cross Abstract: Thermo-mechanical compliant devices are commonly designed with small-strain linear elasticity and temperature-independent material properties, even th

On The Statistical Limits of Self-Improving Agents

SafetyDGX agent

arXiv:2510.04399v3 Announce Type: replace Abstract: We develop a learning-theoretic framework for analyzing self-improving agents by decomposing self-modification into five axes. Within this framework

OpenAI VP of Global Policy Ann O'Leary says AI policy in the US is anchored in the states and informed by California's AI transparency law passed in 2025 (Chase DiFeliciantonio/Politico)

SafetyDGX agent

Chase DiFeliciantonio / Politico: OpenAI VP of Global Policy Ann O'Leary says AI policy in the US is anchored in the states and informed by California's AI transparency law passed in 2025 — SACRAMENTO

Operationalising Relative Causal Knowledge: Backbone Identifiability from Private Reports on a Shared Outcome

SafetyDGX agent

arXiv:2608.10664v1 Announce Type: new Abstract: The Relativity of Causal Knowledge (RCK) explains how a network of agents with different structural causal models can exchange causal knowledge through

Pair-Centric Graph Rewiring for Over-Squashing via Optimal Transport-Guided Communication Alignment

SafetyDGX agent

arXiv:2608.10619v1 Announce Type: new Abstract: Message-passing neural networks (MPNNs) often struggle when task-relevant information is distributed across distant regions of a graph, since local prop

Partially Observable Learning for Multi-Platform Dispatch Optimization

SafetyDGX agent

arXiv:2608.10897v1 Announce Type: new Abstract: Instant delivery platforms have become a critical component of urban logistics, increasingly relying on crowdsourced couriers to fulfill highly dynamic

Physics-Informed Machine Learning in Prognostics and Health Management: A Systematic Literature Review

SafetyDGX agent

arXiv:2608.10047v1 Announce Type: cross Abstract: In modern industry, keeping complex systems reliable, safe, and efficient hinges on Prognostics and Health Management (PHM). Machine Learning (ML) has

Policy Convergence and Divergence Across National and Within Regional AI Strategies: A Policy Design Element Analysis

SafetyDGX agent

arXiv:2608.11006v1 Announce Type: cross Abstract: Governments worldwide have responded to the rapid expansion of AI by publishing national and regional AI strategies. Comparing national and regional A

Position Encoding in Transformers: From Absolute and Relative Methods to Rotary Position Embeddings and Long-Context Scaling

SafetyDGX agent

arXiv:2608.10021v1 Announce Type: new Abstract: Self-attention models content-dependent interactions between tokens but does not by itself encode token order. Position encoding addresses this limitati

Predicting Space Groups of Double Perovskites by LLM with Dynamic Few-Shot Learning

SafetyDGX agent

arXiv:2608.10483v1 Announce Type: new Abstract: Double perovskites (DPs) offer broad compositional tunability, but predicting the space groups (SGs) of stable structures remains difficult because avai

Procedural Fairness Failures in RLHF from Preference Averaging

SafetyDGX agent

arXiv:2608.10126v1 Announce Type: cross Abstract: Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous preferences into a single reward model, assuming preference homogeneity. Wh

← Previous
1…5253545556…240
Next →