AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,600 results
11 Aug 2026

Matching Supervision to the Student's Learning Capacity: A Unified Framework for On-Policy Self-Distillation

SafetyDGX agent

arXiv:2608.08176v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) improves the reasoning abilities of LLMs by internalizing privileged context into model parameters through self-disti

MedCalc-R1: Knowledge-Guided Reward Framework for Medical Mathematical Reasoning

SafetyDGX agent

arXiv:2608.08623v1 Announce Type: new Abstract: In Reinforcement Learning with Verifiable Rewards (RLVR) frameworks for mathematical reasoning tasks, floating-point results are typically evaluated usi

Metanormative Theory for RL-Based Moral Agents

SafetyDGX agent

arXiv:2608.08220v1 Announce Type: new Abstract: The overlapping disciplines of machine ethics and value alignment are concerned with designing artificial agents that are aligned with human values and


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

MGMCL: Multi-Granularity Manifold Contrastive Learning With Neural ODEs for Cross-Subject EEG Emotion Recognition

SafetyDGX agent

arXiv:2608.08440v1 Announce Type: new Abstract: Cross-subject electroencephalogram (EEG)-based emotion recognition remains challenging due to substantial inter-individual variability and discrete form

Mismatch Matters: On-Policy Distillation Beyond Token Agreement

SafetyDGX agent

arXiv:2608.09836v1 Announce Type: new Abstract: On-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerate agreement,

Mitigating Gender Bias in English to Romanian Machine Translation

SafetyDGX agent

arXiv:2608.08606v1 Announce Type: cross Abstract: Machine translation (MT) systems often fail to correctly translate gender, especially when converting from a gender-neutral language like English to a

ML-Based Hierarchical Prediction for Practical Energy Scheduling in Dynamic NTN-WPT Systems

SafetyDGX agent

arXiv:2608.08804v1 Announce Type: cross Abstract: With advancements in long-distance wireless power transfer (WPT) and space-based energy technologies, integrating WPT into non-terrestrial networks (N

Model-Based Systems Engineering Framework for SysML-Driven Design of Autonomous UAVs

SafetyDGX agent

arXiv:2608.09547v1 Announce Type: new Abstract: Autonomous Unmanned Aerial Vehicles (UAVs) are complex cyber-physical systems that require the coordinated integration of flight control, navigation, pe

Model the Edit, Not the Image: Visual Autoregressive Editing from a Source-Centric Perspective

SafetyDGX agent

arXiv:2608.09057v1 Announce Type: new Abstract: Next-scale visual autoregressive models (VARs) have emerged as a powerful generative paradigm, producing high-quality images through efficient coarse-to

Motif 3: Technical Report

SafetyDGX agent

arXiv:2608.09119v1 Announce Type: new Abstract: We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. Each spar

MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling

SafetyDGX agent

arXiv:2608.08553v1 Announce Type: new Abstract: Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from

Multi-Agent AI Safety as an Institutional Design Problem

SafetyDGX agent

arXiv:2608.09828v1 Announce Type: cross Abstract: AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent wo

Multi-Branch Policy Optimization for Multimodal Large Language Models

SafetyDGX agent

arXiv:2608.07581v1 Announce Type: cross Abstract: Group-based reinforcement learning methods for multimodal large language models typically rely on trajectory-level credit assignment that applies a si

Multimodal Model Diffing for Feature Discovery and Control

SafetyDGX agent

arXiv:2608.09928v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to

MultiShadow: Multi-Object Shadow Generation for Image Compositing via Diffusion Model

SafetyDGX agent

arXiv:2603.02743v4 Announce Type: replace Abstract: Realistic shadow generation is crucial for achieving seamless image compositing, yet existing methods primarily focus on single-object insertion and

Neural Message Passing on Structural Interaction Graphs for Fully-Inductive Graph Neural Networks

SafetyDGX agent

arXiv:2608.08567v1 Announce Type: new Abstract: A central obstacle in building graph foundation models is the input heterogeneity in terms of feature space dimensionality, semantics, and structure. Su

NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models

SafetyDGX agent

arXiv:2509.03985v2 Announce Type: replace-cross Abstract: In deployment and application, large language models (LLMs) typically undergo safety alignment to prevent illegal and unethical outputs. Howev

North Africa's Missing Framework: NLP-Driven Mental Healthcare in Algeria and Implications for Low-resource Settings

SafetyDGX agent

arXiv:2608.08607v1 Announce Type: new Abstract: Mental health disorders are a leading cause of disability worldwide, yet Natural Language Processing (NLP) research for mental healthcare has remained c

OD-Gear: Online Decomposition and Group Sampling for Expert-Guided Adversarial Routing in Scalable Capacitated Vehicle Routing

SafetyDGX agent

arXiv:2602.00488v3 Announce Type: replace Abstract: Solving large-scale capacitated vehicle routing problems (CVRP) is hindered by the high complexity of classical heuristics and the limited generaliz

On the use of foundation models in cognitive science

SafetyDGX agent

arXiv:2608.07812v1 Announce Type: new Abstract: A host of recent studies have evaluated the cognitive and developmental alignment of Foundation Models (FMs). These investigations include evaluations o

OnEvoMemory: Evolving Memory through Online Robot Rollouts for Pretrained Robot Policies

SafetyDGX agent

arXiv:2608.08749v1 Announce Type: new Abstract: Long-horizon robot manipulation requires policies to track completed subtasks and critical interaction events. However, existing memory mechanisms heavi

OWN YOUR INTELLIGENCE Last year, building on open-weight models was primarily a cost rationalization exercise. Slightly worse performance fo…

SafetyDGX agent

OWN YOUR INTELLIGENCE Last year, building on open-weight models was primarily a cost rationalization exercise. Slightly worse performance for a much cheaper price. Now, it is increasingly an existenti

PAM: Training Policy-Aligned Moderation Filters at Scale

SafetyDGX agent

arXiv:2505.19766v4 Announce Type: replace Abstract: Large language models (LLMs) remain vulnerable to misalignment and jailbreaks, making external safeguards like moderation filters essential, yet exi

PAST: Privileged Adaptation from Complete Student Trajectories for On-Policy Self-Distillation

SafetyDGX agent

arXiv:2608.08726v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) uses a privileged teacher to supervise a reasoning model on prefixes sampled from its own rollouts. Yet each rollou

Penalizing Length: Uncovering Systematic Bias in Quality Estimation Metrics

SafetyDGX agent

arXiv:2510.22028v4 Announce Type: replace Abstract: Quality Estimation (QE) metrics are vital in machine translation for reference-free evaluation and increasingly serve as selection criteria in data

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots

SafetyDGX agent

arXiv:2608.09931v1 Announce Type: new Abstract: Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Dist

Persuasive and Compliant Tendencies Predict Group Decision-Making in Humans and Language Models

SafetyDGX agent

arXiv:2608.08199v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly involved in group decision-making with other LLMs and humans. Yet it remains unclear whether their influen

PhysAttNet: Enhancing Predictive Performance in Industrial and Astrophysical Time Series via Physics-Informed Attention

SafetyDGX agent

arXiv:2608.07681v1 Announce Type: new Abstract: Accurate and robust time series forecasting is essential in many applications involving physical processes, such as manufacturing monitoring and astroph

Physics-Informed Policy Iteration for High-Dimensional Hamilton--Jacobi--Bellman Equations: Interior Error Bounds without Boundary Data

SafetyDGX agent

arXiv:2508.01718v2 Announce Type: replace Abstract: We develop a physics-informed policy-iteration method for stationary second-order Hamilton--Jacobi--Bellman equations arising in continuous-time sto

PIVOT: Preference-based Intervention Vectors for Pedagogical Tutor Steering

SafetyDGX agent

arXiv:2608.07509v1 Announce Type: cross Abstract: LLMs are increasingly used for conversational tutoring, but effective tutoring requires more than correct answers. Tutors must choose when to scaffold

Planning/RL for a stochastic single-player merge puzzle: afterstates, previewed chance events, and long-horizon throughput [D]

SafetyDGX agent

I am working on an AI for a small single-player merge puzzle and would appreciate pointers to related algorithms, papers, or existing implementations. It resembles 2048 in its action -> afterstate ->

Position Bias in Ordinal Classification: A Systematic Evaluation

SafetyDGX agent

arXiv:2608.08869v1 Announce Type: new Abstract: Large language models are increasingly used for ordinal classification, yet semantically equivalent changes to prompt organization can alter their predi

PQC in Plaintext: Google Cloud’s post-quantum cryptography roadmap

SafetyDGX agent

Securing infrastructure and services against a future cryptographically-relevant quantum computer has been a goal for Google for a decade, and we’ve dedicated ourselves to help developers by advancing

Pragmatic Attack Surface: Vulnerabilities of Implicit Context in Large Language Models

SafetyDGX agent

arXiv:2608.09551v1 Announce Type: new Abstract: In the era of large language models (LLMs), attackers often manipulate natural language to elicit unsafe or harmful outputs, creating a new natural lang

Privacy-Preserving Data Drift Detection and Recovery for Large-Scale LLM Applications via Proxy Representations

SafetyDGX agent

arXiv:2608.08245v1 Announce Type: cross Abstract: LLM applications deployed at scale face a fundamental challenge: privacy constraints prevent direct inspection of user interactions, making it difficu

Private Anytime Selective-Risk Certification for Federated Retrieval-Augmented Generation: Guarantees and Empirical Limits

SafetyDGX agent

arXiv:2608.07913v1 Announce Type: cross Abstract: Selective-risk certificates promise that accepted outputs meet a declared error target. We develop Fed-SRC, a score-agnostic certificate for federated

Privileged Likelihood Is Not Automatically Value: Three Checks for Token Credit in On-Policy Self-Distillation

SafetyDGX agent

arXiv:2608.09263v1 Announce Type: new Abstract: Outcome verifiers score completed reasoning traces but do not assign credit to intermediate tokens. Privileged self-distillation attempts to fill this g

Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation

SafetyDGX agent

arXiv:2608.09228v1 Announce Type: cross Abstract: On-Policy Self-Distillation (OPSD) is commonly interpreted as the transfer of privileged information: a teacher observes the verified solution to the

Proxy OPD: On-Policy Distillation with Transferable Relative Proxy Update

SafetyDGX agent

arXiv:2607.11505v2 Announce Type: replace-cross Abstract: Post-training for large language models typically couples policy exploration with model optimization, hindering the reuse of high-reward behav

Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression

SafetyDGX agent

arXiv:2608.08960v1 Announce Type: new Abstract: Multi-step language-model agents repeatedly process growing interaction histories, leading to substantial context costs. Vision--text compression reduce

Real-Time Nonlinear MPC via Sequential Quadratic Programming with Structure-Exploiting ADMM and Interior-Point Methods for Underactuated Double-Pendulum Swing-Up

SafetyDGX agent

arXiv:2608.09272v1 Announce Type: cross Abstract: The 4th 'AI Olympics with RealAIGym' competition, to be held at IJCAI-ECAI 2026 in Bremen, challenges participants to develop a global control policy

Reconfigurable Structural Robotic Assembly: Interlocking 3D Aggregations with Self-Aligning Compound Nested Lattice Modules

SafetyDGX agent

arXiv:2608.07576v1 Announce Type: new Abstract: Robotic construction systems often treat the material system and the robot as separate design problems, locating intelligence primarily in hardware, sen

Reflex First, Reflect Later: Latency-Aware Embodied LLM Agents for Dynamic Response

SafetyDGX agent

arXiv:2506.07223v2 Announce Type: replace Abstract: Large language models (LLMs) have substantially improved the planning capabilities of embodied agents, enabling their deployment in dynamic and safe

Regret of exploratory policy improvement and q-learning

SafetyDGX agent

arXiv:2411.01302v2 Announce Type: replace Abstract: We study the convergence of q-learning and related algorithms introduced by Jia and Zhou (J. Mach. Learn. Res., 24 (2023), 161) for controlled diffu

REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment

SafetyDGX agent

arXiv:2608.07931v1 Announce Type: new Abstract: Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment. Hallucinations in

Removing Infrastructure Barriers in Human-Robot Collaboration Through Wireless Reconfigurable Cells

SafetyDGX agent

arXiv:2608.09658v1 Announce Type: cross Abstract: Human-Robot Collaboration (HRC) plays a vital role in dynamic, high mix, low volume industrial scenarios such as remanufacturing, which frequently fac

Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models

SafetyDGX agent

arXiv:2508.16406v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) remain vulnerable to jailbreak attacks, which attempt to elicit harmful responses from LLMs. The evolving nature

Retrieval-Augmented Generation-Based Color Restoration for Low-Light Image Enhancement

SafetyDGX agent

arXiv:2608.08211v1 Announce Type: cross Abstract: Recent low-light image enhancement (LLIE) methods have driven brightness and structural fidelity close to that of normally-exposed images, yet their o

RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation

SafetyDGX agent

arXiv:2608.09226v1 Announce Type: cross Abstract: Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these procedures ar

RobustDefect-LLM: Explainable and Robustness-Aware Industrial Surface Defect Classification with Decision Support and AI-Assisted Reporting

SafetyDGX agent

arXiv:2608.08589v1 Announce Type: new Abstract: This paper presents RobustDefect-LLM, an industrial surface-defect inspection framework integrating deep-learning classification, operator-facing visual

RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance

SafetyDGX agent

arXiv:2608.09853v1 Announce Type: cross Abstract: General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from

SAFE-CHEM: Uncertainty-Aware Policy Switching for Robust Robotic Chemistry

SafetyDGX agent

arXiv:2608.09303v1 Announce Type: cross Abstract: The deployment of autonomous robotic systems in chemistry laboratories is accelerating experimental workflows and providing the foundational data for

Safety Cost of Steering Vectors Is Separable and Reducible

SafetyDGX agent

arXiv:2608.08383v1 Announce Type: new Abstract: Steering vectors are a lightweight tool for controlling LLM behavior. However, emerging evidence shows that steering vectors can unintentionally comprom

Satellite Trajectory Optimization via Proximal Policy Optimization for Space Debris Avoidance

SafetyDGX agent

arXiv:2608.09628v1 Announce Type: new Abstract: Collision avoidance systems are commonly used to avoid fragmentation events occurring in Low-Earth Orbit (LEO) and Geosynchronous Equatorial Orbit (GEO)

SC-Diff: Semantically Calibrated Diffusion for Visible-to-Infrared Image Translation

SafetyDGX agent

arXiv:2608.08555v1 Announce Type: new Abstract: Visible-to-infrared image translation provides a practical way to expand infrared training data using abundant visible images. Diffusion models are prom

Scalable extensions to given-data Sobol' index estimators

SafetyDGX agent

arXiv:2509.09078v3 Announce Type: replace-cross Abstract: Given-data methods for variance-based sensitivity analysis have significantly advanced the feasibility of Sobol' index computation for computa

ScaleSense: Cost-Intelligent Scaling Framework via Learned Resource Estimation in Alibaba AnalyticDB

SafetyDGX agent

arXiv:2608.07945v1 Announce Type: cross Abstract: Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage from compute, yet determining the optimal resource alloc

SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning

SafetyDGX agent

arXiv:2608.07959v1 Announce Type: new Abstract: Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multi

Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards

SafetyDGX agent

arXiv:2608.07531v1 Announce Type: cross Abstract: Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing ext

Self Supervised Learning from Automatically Generated Demonstrations for Visual Robotic Manipulation

SafetyDGX agent

arXiv:2608.07553v1 Announce Type: new Abstract: Robotic manipulation often requires object specific programming, manual data annotation, or calibrated perception pipelines, which limits rapid deployme

← Previous
1…34567…210
Next →