AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlog
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,478 results
Safety

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

DGX agent

arXiv:2608.12253v1 Announce Type: cross Abstract: Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that

safetyarxiv-cs-ai
13 Aug 2026
Safety
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR

DGX agent

arXiv:2608.11368v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) spends most of its compute generating groups of long reasoning trajectories. Recent allocators red

safetyarxiv-cs-lg
13 Aug 2026
Safety

Policy-as-logic for robust reasoning over rules

DGX agent

arXiv:2608.11905v1 Announce Type: new Abstract: In many practical applications of generative AI systems, from tax rules to airline baggage allowance, responses to natural language queries must respect

safetyarxiv-cs-ai
13 Aug 2026
Safety

Policy-Induced Hand Priors in Humanoid Dual-Arm Manipulation: Diagnosing and Mitigating Initial-Pose Dependence

DGX agent

arXiv:2608.11769v1 Announce Type: new Abstract: Vision-language-action (VLA) policies are expected to operate robustly across variations in the robot's initial configuration, yet aggregate task succes

safetyarxiv-cs-ro
13 Aug 2026
Safety

Post-Training with Policy Gradients: Optimality and the Base Model Barrier

DGX agent

arXiv:2603.06957v2 Announce Type: replace-cross Abstract: We study post-training linear autoregressive models with outcome and process rewards. Given a context oldsymbol{x}, the model must predict the

safetyarxiv-cs-ai
13 Aug 2026
Safety

Prompt-Driven Exploration

DGX agent

arXiv:2607.08837v2 Announce Type: replace-cross Abstract: Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers. Standard methods inject

safetyarxiv-cs-ai
13 Aug 2026
Safety

Proportional Committee Elections with Positive and Negative Votes

DGX agent

arXiv:2503.01985v2 Announce Type: replace-cross Abstract: In the classic committee election setting each voter approves a subset of candidates and the goal is to select k winners based on these prefer

safetyarxiv-cs-ai
13 Aug 2026
Safety

RA-ClipScore: Making Generative Model Evaluation More Interpretable

DGX agent

arXiv:2608.12088v1 Announce Type: new Abstract: Generative models can produce images nearly indistinguishable from real data, yet rigorous and interpretable evaluation remains challenging. Conventiona

safetyarxiv-cs-cv
13 Aug 2026
Safety

Redistribution-based Cost Inference Improves Sparse Safe Offline RL

DGX agent

arXiv:2608.12306v1 Announce Type: cross Abstract: Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level stop-feedback:

safetyarxiv-cs-ai
13 Aug 2026
Safety

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation

DGX agent

arXiv:2608.11698v1 Announce Type: cross Abstract: On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolation methods

safetyarxiv-cs-ai
13 Aug 2026
Safety

Reoptimization Algorithms for Contextual Bandits with Knapsack Constraints

DGX agent

arXiv:2608.11383v1 Announce Type: new Abstract: We study new algorithms for Contextual Bandits with Knapsack. In these problems, there are finitely many types of customers, products, and resources. Ea

safetyarxiv-cs-lg
13 Aug 2026
Safety

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning

DGX agent

arXiv:2608.11587v1 Announce Type: cross Abstract: Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalist

safetyarxiv-cs-cl
13 Aug 2026
Safety

ScaleVid: Geometry-Aware Video Object Scaling with Mesh-Free Inference

DGX agent

arXiv:2608.12232v1 Announce Type: new Abstract: Geometry-aware video object scaling aims to anisotropically resize the object along object-centric axes while preserving geometric plausibility, tempora

safetyarxiv-cs-cv
13 Aug 2026
Safety

Small Data Explainer -- The impact of small data methods in everyday life

DGX agent

arXiv:2507.11773v2 Announce Type: replace-cross Abstract: The emergence of breakthrough artificial intelligence (AI) techniques has led to a renewed focus on how small data settings, i.e., settings wi

safetyarxiv-cs-ai
13 Aug 2026
Safety

Stolen LLM Reasoning: How come OpenAI, Anthrophic, Google have the same vulnerabilities?

DGX agent

If you haven't checked the paper: https://arxiv.org/abs/2608.09867 TLDR: the authors show that you can swap out the 'encrypted' reasoning of the biggest model, like Opus, Sol, and put them into weaker

safetyr-localllama
13 Aug 2026
Safety

Through Van Gogh's Eyes: Global Style Transfer with Diffusion Mod

DGX agent

arXiv:2608.11546v1 Announce Type: new Abstract: Artistic image synthesis aims to recreate the expressive visual identity of a target artist, yet existing methods often fail to capture an artist's glob

safetyarxiv-cs-cv
13 Aug 2026
Safety

ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

DGX agent

arXiv:2608.11878v1 Announce Type: cross Abstract: Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. Howeve

safetyarxiv-cs-cl
13 Aug 2026
Safety

Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion

DGX agent

arXiv:2608.11794v1 Announce Type: cross Abstract: The growing role of AI-generated content and AI-enabled systems in public communication has led regulators to demand clear disclosure of content prove

safetyarxiv-cs-ai
13 Aug 2026
Safety

Towards a Formal Definition of Agent Memory: Basis, Span, Optimality, and the Sequential Memory Problem

DGX agent

arXiv:2608.11654v1 Announce Type: new Abstract: Despite the wide deployment of memory in large-model agents, there is no unified formal account of what a memory is or when it is optimal. This paper ta

safetyarxiv-cs-lg
13 Aug 2026
Safety

Towards Human Motion World Models via Executable Behaviour Representations

DGX agent

arXiv:2604.18064v2 Announce Type: replace Abstract: Human motion world models should capture motion's intentionality by being executable: adaptable to different actions and capable of assessing motion

safetyarxiv-cs-ai
13 Aug 2026
Safety

Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling

DGX agent

arXiv:2608.11829v1 Announce Type: cross Abstract: On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. It is commonly believed to enable the stu

safetyarxiv-cs-cl
13 Aug 2026
Safety

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

DGX agent

arXiv:2608.11752v1 Announce Type: new Abstract: Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content,

safetyarxiv-cs-cv
13 Aug 2026
Safety

Variable Selection in the Context of AI Fairness

DGX agent

arXiv:2608.11251v1 Announce Type: cross Abstract: Fairness in AI systems has become more important with recent regulatory demands, such as the EU AI Act. Traditional approaches often do not take into

safetyarxiv-cs-ai
13 Aug 2026
Safety

Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study

DGX agent

arXiv:2608.11649v1 Announce Type: new Abstract: As users increasingly turn to Large Language Models (LLMs) for information and advice on political matters, particularly during election periods, the po

safetyarxiv-cs-cl
13 Aug 2026
Safety

Why AI Detection Fails for Academic Integrity

DGX agent

arXiv:2608.11256v1 Announce Type: new Abstract: Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as

safetyarxiv-cs-lg
13 Aug 2026
Safety

A Joint-Distribution Route to Fair Representations with Continuous Sensitive Attributes

DGX agent

arXiv:2608.10470v1 Announce Type: new Abstract: Fair representation learning with a continuous sensitive attribute S requires a representation Z that is statistically independent of S. Existing criter

safetyarxiv-cs-lg
12 Aug 2026
Safety

Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents

DGX agent

arXiv:2608.11110v1 Announce Type: new Abstract: When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compa

safetyarxiv-cs-cl
12 Aug 2026
Safety

// Actions Speak Louder Than Words // Multilingual agent evaluation compares final answers and throws the trajectory away. The trajectory fi…

DGX agent

// Actions Speak Louder Than Words // Multilingual agent evaluation compares final answers and throws the trajectory away. The trajectory fixes cost, latency, failure mode, and auditability. New resea

safetydair-ai--x
12 Aug 2026
Safety

AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss

DGX agent

arXiv:2608.11205v1 Announce Type: new Abstract: Frechet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-le

safetyarxiv-cs-cv
12 Aug 2026
Safety

APCReg: Anatomical-Prior-Guided Coarse-to-Fine CBCT--IOS Registration via Multi-View Projection and Reliability-Controlled Residual Correction

DGX agent

arXiv:2608.09993v1 Announce Type: cross Abstract: Registration between cone-beam computed tomography (CBCT) and intraoral scans (IOS) is essential for patient-specific surgical planning. However, disp

safetyarxiv-cs-cv
12 Aug 2026
Safety

Beyond Forecasting: Recasting Volatility Control as a Routing Problem

DGX agent

arXiv:2608.10375v1 Announce Type: cross Abstract: Volatility control converts risk estimates into portfolio exposure, yet existing approaches often rely on a fixed volatility estimator or a pre-define

safetyarxiv-cs-ai
12 Aug 2026
Safety

Big Tech's reliance on OpenAI and Anthropic for growth is systemically pervasive. Both companies are massively unprofitable — and much of wh…

DGX agent

Big Tech's reliance on OpenAI and Anthropic for growth is systemically pervasive. Both companies are massively unprofitable — and much of what they spend is Big Tech's own money, counted right back as

safetygary-marcus--x
12 Aug 2026
Safety

BooST: Bridging Semantics and Motions for Efficient Skill Transfer

DGX agent

arXiv:2608.10600v1 Announce Type: cross Abstract: Skill abstraction---the process of learning reusable and temporally extended behaviors---has emerged as a key paradigm for improving sample efficiency

safetyarxiv-cs-cv
12 Aug 2026
Model Releases

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning

DGX agent

arXiv:2608.10204v1 Announce Type: new Abstract: Safe reinforcement learning maximizes reward subject to safety constraints. For Constrained Markov Decision Processes, the linear-programming view over

model-releasesarxiv-cs-lg
12 Aug 2026
Safety

CARE: Confidence-Aware Reasoning for Reliable Medical VQA

DGX agent

arXiv:2608.10964v1 Announce Type: cross Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual qu

safetyarxiv-cs-ai
12 Aug 2026
Safety

Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory

DGX agent

arXiv:2608.09937v1 Announce Type: new Abstract: Recent work in NLP has probed large language models for their understanding of cultural norms across countries. However, this work typically considers d

safetyarxiv-cs-cl
12 Aug 2026
Safety

ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering

DGX agent

arXiv:2608.10996v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has been especially effective in mathematics and coding, where answers can be checked automatically. Many

safetyarxiv-cs-cl
12 Aug 2026
Safety

Convergence of Sign-based Random Reshuffling Algorithms for Nonconvex Optimization

DGX agent

arXiv:2310.15976v4 Announce Type: replace Abstract: signSGD is attractive in nonconvex optimization because it communicates sign-valued rather than full-precision gradients. Several standard analyses

safetyarxiv-cs-lg
12 Aug 2026
Safety

Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning

DGX agent

arXiv:2608.10473v1 Announce Type: cross Abstract: Offline-to-online (O2O) reinforcement learning aims to leverage policies pretrained on static datasets while improving them through online interaction

safetyarxiv-cs-ai
12 Aug 2026
Safety

Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition Agents

DGX agent

arXiv:2608.10441v1 Announce Type: cross Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expen

safetyarxiv-cs-cl
12 Aug 2026
Safety

DIMOS: Disentangling Instance-level Moving Object Segmentation

DGX agent

arXiv:2606.12826v2 Announce Type: replace-cross Abstract: Moving instance segmentation (MIS) attracts increasing attention due to its broad applications in traffic surveillance, autonomous driving, an

safetyarxiv-cs-ai
12 Aug 2026
Safety

Do AI weather models miss extremes?

DGX agent

arXiv:2608.09972v1 Announce Type: cross Abstract: First-generation AI weather models are often reported to underperform at extremes, mostly in reanalysis-based evaluations of deterministic regression

safetyarxiv-cs-ai
12 Aug 2026
Safety

Do Time-Series Forecasters Use the Right History: Recoverability, Recovery, and Functional Use of Temporal Delays

DGX agent

arXiv:2608.10433v1 Announce Type: new Abstract: Forecast accuracy does not tell us which past inputs produced a prediction. We separate three questions for time-series models with known delay structur

safetyarxiv-cs-lg
12 Aug 2026
Safety

Dual-Loop Self-Evolution via Verifiable Emotion Feedback for Multi-Turn Empathetic Dialogue

DGX agent

arXiv:2608.10626v1 Announce Type: new Abstract: Large language models have demonstrated conversational capabilities, yet empathetic competence remains challenging. Empathetic support is inherently mul

safetyarxiv-cs-cl
12 Aug 2026
Safety

Dual Space Preconditioning for Gradient Descent in the Overparameterized Regime

DGX agent

arXiv:2603.10485v3 Announce Type: replace-cross Abstract: In this work, we study the convergence properties of the Dual Space Preconditioned Gradient Descent, encompassing optimizers such as Normalize

safetyarxiv-cs-lg
12 Aug 2026
Safety

Efficient Hypergradient Descent for Inverse Reinforcement Learning

DGX agent

arXiv:2608.11052v1 Announce Type: new Abstract: Inverse reinforcement learning (IRL) aims to recover a reward function under which the resulting policy reproduces the behavior observed in expert demon

safetyarxiv-cs-lg
12 Aug 2026
Safety

Enhancing Automated Essay Scoring With Three Techniques: Two-Stage Fine-Tuning, Score Alignment, and Self-Training

DGX agent

arXiv:2602.01747v2 Announce Type: replace Abstract: Automated Essay Scoring (AES) plays a crucial role in education by providing scalable and efficient assessment tools. However, in real-world setting

safetyarxiv-cs-cl
12 Aug 2026
Safety

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes

DGX agent

arXiv:2608.10209v1 Announce Type: new Abstract: Feedback signals used to train Large Language Models (LLMs) are the primary driver of their behavior and our main lever for instilling alignment with hu

safetyarxiv-cs-ai
12 Aug 2026
← Previous
1…6768697071…302
Next →