AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,153 results
Model Releases

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

DGX agent

arXiv:2506.07468v4 Announce Type: replace-cross Abstract: Conventional large language model (LLM) safety alignment relies on a reactive, disjoint loop: attackers exploit a static model, then defenders

model-releasesarxiv-cs-cl
7 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Conversational Human Audio-visual Talking Dialogue Generation

DGX agent

arXiv:2607.02799v1 Announce Type: new Abstract: Large-scale dyadic interactive audio-visual dialogue (DIAD) datasets provide fundamental data resources for developing humanoid interactive virtual agen

researcharxiv-cs-cv
7 Jul 2026
Safety

CRRL: A Causality-Based Reinforcement Learning Framework for Autonomous System Recovery

DGX agent

arXiv:2607.03177v1 Announce Type: cross Abstract: Traditional reinforcement learning (RL) for recovery in autonomous systems lacks causal understanding and generalizes poorly to novel failure scenario

safetyarxiv-cs-ai
7 Jul 2026
Model Releases

Enhancing Video Physical Consistency via Role-aware Joint Training and Modality-decoupled Denoising

DGX agent

arXiv:2607.04653v1 Announce Type: new Abstract: While modern video diffusion models excel in visual fidelity, maintaining long-range physical consistency remains a formidable challenge. Conventional p

model-releasesarxiv-cs-cv
7 Jul 2026
Safety

Governing Generative AI Across Financial Institutions: An SR 26-2-Compatible Framework for Generative AI Risk Control

DGX agent

arXiv:2607.04103v1 Announce Type: cross Abstract: The release of SR 26-2 marks a significant modernization of U.S. model risk management by replacing SR 11-7 with a more risk-based and materiality-sen

safetyarxiv-cs-lg
7 Jul 2026
Safety

High-Precision Formation Control for Heterogeneous Multi-Robot Systems via Hierarchical Hybrid Physics-Informed Deep Reinforcement Learning

DGX agent

arXiv:2607.03512v1 Announce Type: new Abstract: Existing classical control methods commonly require precise models and struggle to cope with model uncertainties and external disturbances, while end-to

safetyarxiv-cs-ro
7 Jul 2026
Safety

How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Krugel, and Uhl (2025)

DGX agent

arXiv:2603.22730v2 Announce Type: replace Abstract: Pfeffer, Krugel, and Uhl (2025) report that OpenAI's reasoning model o1-mini produces more utilitarian responses to the trolley problem and footbrid

safetyarxiv-cs-cl
7 Jul 2026
Safety

Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation

DGX agent

arXiv:2510.08807v3 Announce Type: replace-cross Abstract: From loco-motion to dextrous manipulation, humanoid robots have made remarkable strides in demonstrating complex full-body capabilities. Howev

safetyarxiv-cs-lg
7 Jul 2026
Model Releases

ICR-RL: Deep Reinforcement Learning via In-Context Regression

DGX agent

arXiv:2509.11259v2 Announce Type: replace-cross Abstract: Recent advancements in machine learning have largely been driven by foundation models (FMs) trained on large, diverse datasets, enabling them

model-releasesarxiv-cs-ai
7 Jul 2026
Safety

Improving LLMs via Validator-to-Generator Alignment

DGX agent

arXiv:2607.02668v1 Announce Type: new Abstract: Large language models are inconsistent: varying prompts or including unrelated information can lead to unexpected changes in model outputs. The generato

safetyarxiv-cs-cl
7 Jul 2026
Applications

Knowledge-Centric Information Systems

DGX agent

arXiv:2607.02609v1 Announce Type: cross Abstract: For decades, data engineering has developed mature architectural principles for integrating, governing, validating, cataloging, and serving organizati

applicationsarxiv-cs-ai
7 Jul 2026
Safety

Mitigating Covariate Shift in Imitation Learning for Autonomous Vehicles Using Latent Space Generative World Models

DGX agent

arXiv:2409.16663v5 Announce Type: replace-cross Abstract: We propose the use of latent space generative world models to address the covariate shift problem in autonomous driving. A world model is a ne

safetyarxiv-cs-cv
7 Jul 2026
Applications

Multi-Robot Open Adaptive Teaming Across Unseen Environments, Partners, and Scales

DGX agent

arXiv:2607.04972v1 Announce Type: cross Abstract: Deploying robot teams in the real world requires simultaneous adaptation to unseen environments, unknown partners, and varying team sizes, yet existin

applicationsarxiv-cs-ai
7 Jul 2026
Model Releases

Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs

DGX agent

arXiv:2607.04371v1 Announce Type: new Abstract: We present Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super optimized for interactive deployment. We designed the model to maxim

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Octax: Accelerated CHIP-8 Arcade Environments for Reinforcement Learning in JAX

DGX agent

arXiv:2510.01764v3 Announce Type: replace Abstract: Reinforcement learning (RL) research requires diverse, challenging environments that are both tractable and scalable. While modern video games may o

model-releasesarxiv-cs-lg
7 Jul 2026
Model Releases

OmniLayout: A Schematic-Coupled Multimodal Benchmark for Constraint-Aware Geometric Reasoning in PCB Layout

DGX agent

arXiv:2607.03261v1 Announce Type: new Abstract: Recent large language models (LLMs) have demonstrated remarkable progress in 3D spatial reasoning, spatial grounding, and fine-grained geometric underst

model-releasesarxiv-cs-cv
7 Jul 2026
Safety

Overloading Large Vision-Language Models for Jailbreaking

DGX agent

arXiv:2607.02961v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) exhibit remarkable vision-language capabilities and are increasingly deployed in real-world applications such as pe

safetyarxiv-cs-cv
7 Jul 2026
Model Releases

Predicting Biased Human Decision-Making with Large Language Models in Conversational Settings

DGX agent

arXiv:2601.11049v2 Announce Type: replace-cross Abstract: We examine whether large language models (LLMs) can predict biased decision-making in conversational settings, and whether their predictions c

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog

DGX agent

arXiv:2607.04438v1 Announce Type: cross Abstract: Research dissemination, turning a paper into a poster, a talk video, and a blog post, is still a manual last mile. Prior automation treats each artifa

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Responsibility Distribution Estimation in Ego-View Accident Videos with Multimodal Large Language Models

DGX agent

arXiv:2607.03591v1 Announce Type: cross Abstract: Recent studies on multimodal traffic accident understanding have mainly relied on infrastructure-camera footage, satellite imagery, or structured cras

model-releasesarxiv-cs-ai
7 Jul 2026
Safety

RL-Ballast: Ship Ballast Water Path Planning and Clog Prediction via Reinforcement Learning

DGX agent

arXiv:2607.04906v1 Announce Type: new Abstract: Under the Shipping 4.0 paradigm, autonomous and reduced-crew vessels require intelligent internal systems to maintain operational safety and structural

safetyarxiv-cs-lg
7 Jul 2026
Model Releases

S-EMBER: A Large-Scale Benchmark for Streaming Egocentric Memory Retrieval

DGX agent

arXiv:2607.02689v1 Announce Type: cross Abstract: As wearable devices enable continuous first-person recording, AI assistants must reason across long time horizons to recall past experiences-a capabil

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Seduced by the Narrative: Assessing Rule Adherence in Semi-Open Textual Sandboxes

DGX agent

arXiv:2607.02802v1 Announce Type: cross Abstract: As LLMs are increasingly deployed as autonomous adjudicators in semi-open textual game environments, robust rule adherence becomes critical when user

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Spectral Rewiring for Exploration, Purification, and Model Merging

DGX agent

arXiv:2607.03065v1 Announce Type: cross Abstract: Reinforcement learning has become a standard post-training recipe for large language models, but dense full-parameter updates create two deployment-re

model-releasesarxiv-cs-ai
7 Jul 2026
Safety

STRATOS: Bridging the Symbolic-to-Numeric Gap in Spatio-Temporal Text-to-SQL for Meteorological Data

DGX agent

arXiv:2607.03501v1 Announce Type: cross Abstract: Copernicus, the European Union's Earth observation program, produces petabytes of Earth observation and climate data, offering immense potential for r

safetyarxiv-cs-ai
7 Jul 2026
Model Releases

TREK: Distill to Explore, Reinforce to Refine

DGX agent

arXiv:2607.05339v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) is effective when the current policy already samples useful reasoning trajectories, but it stalls on hard pr

model-releasesarxiv-cs-ai
7 Jul 2026
Safety

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment

DGX agent

arXiv:2607.04728v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training for large language models (LLMs) follows a efficient paradigm of 'rollout then update', which inevitably res

safetyarxiv-cs-ai
7 Jul 2026
Research

Understanding electricity consumption behaviour through Inverse Reinforcement Learning

DGX agent

arXiv:2607.03176v1 Announce Type: new Abstract: Understanding how households consume electricity in response to socioeconomic and climatic drivers is important for decision-makers designing energy pol

researcharxiv-cs-lg
7 Jul 2026
Model Releases

Unified Audio Intelligence Without Regressing on Text Intelligence

DGX agent

arXiv:2607.05196v1 Announce Type: cross Abstract: Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A

model-releasesarxiv-cs-ai
7 Jul 2026
Hardware

Wan-Streamer v0.2: Higher Resolution, Same Latency

DGX agent

arXiv:2607.04443v1 Announce Type: cross Abstract: We present Wan-Streamer v0.2, a latency-preserving upgrade of the native-streaming, end-to-end audio-visual interaction model. v0.2 keeps the v0.1 mod

hardwarearxiv-cs-ai
7 Jul 2026
Model Releases

Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence

DGX agent

arXiv:2604.08169v2 Announce Type: replace Abstract: Alignment in LLMs is more brittle than commonly assumed: misalignment can be induced by adversarial prompts, benign fine-tuning, emergent misalignme

model-releasesarxiv-cs-ai
3 Jul 2026
Research

AI Virtue: What is 'Good' Knowledge in the Age of Artificial Intelligence?

DGX agent

arXiv:2607.01776v1 Announce Type: cross Abstract: In the age of AI, what will be good knowledge? This article, which is accepted and forthcoming in a special issue of Modern Fiction Studies on 'Cultur

researcharxiv-cs-ai
3 Jul 2026
Model Releases

HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety

DGX agent

arXiv:2607.02079v1 Announce Type: new Abstract: We present HaloGuard 1.0, an open-weights implementation of the constitutional-classifier paradigm for input safety. It achieves state-of-the-art perfor

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models

DGX agent

arXiv:2607.01813v1 Announce Type: cross Abstract: Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vulnerable to

model-releasesarxiv-cs-ai
3 Jul 2026
Safety

Overthink-Triggered Slowdown Attacks on LVLM-Based Robotic Systems

DGX agent

arXiv:2607.01518v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have been increasingly integrated into robotic systems. However, these models may exhibit overthinking behaviors,

safetyarxiv-cs-ro
3 Jul 2026
Safety

Playing 20 Question Game with Policy-Based Reinforcement Learning

DGX agent

arXiv:1808.07645v5 Announce Type: replace-cross Abstract: The 20 Questions (Q20) game is a well known game which encourages deductive reasoning and creativity. In the game, the answerer first thinks o

safetyarxiv-cs-ai
3 Jul 2026
Model Releases

PreScience: A Dataset and Benchmark for Scientific Forecasting

DGX agent

arXiv:2602.20459v2 Announce Type: replace Abstract: Can AI systems trained on the existing scientific record forecast the advances that will follow? We introduce PreScience, a dataset and benchmark fo

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics

DGX agent

arXiv:2510.09517v2 Announce Type: replace Abstract: Despite rapid advances in large language models (LLMs), statistical reasoning remains underrepresented in existing LLM benchmarks, which often do no

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

AGC-Bench: Measuring Artificial General Creativity

DGX agent

arXiv:2607.01152v1 Announce Type: new Abstract: Creativity research has debated whether creativity is domain-specific (e.g., visual, writing, science), and if it is psychometrically separable from gen

model-releasesarxiv-cs-cl
2 Jul 2026
Model Releases

Beyond Document Grounding: Span-Level Hallucination Detection over Code, Tool Output, and Documents

DGX agent

arXiv:2607.00895v1 Announce Type: new Abstract: Hallucination detection for retrieval-augmented generation (RAG) is usually evaluated on natural-language document evidence. However, grounded generatio

model-releasesarxiv-cs-cl
2 Jul 2026
Safety

Bounded Morality: Defining the Space of Moral Computation

DGX agent

arXiv:2607.00002v1 Announce Type: new Abstract: Moral cognition has traditionally been modeled as adherence to fixed ethical theories--deontology, consequentialism, virtue ethics--implemented as stati

safetyarxiv-cs-ai
2 Jul 2026
Safety

Distributed Multi Robot Lunar Cargo Transportation via Phase Decomposed Reinforcement Learning

DGX agent

arXiv:2607.00160v1 Announce Type: new Abstract: Modular reconfigurable robotic systems provide a scalable solution for cooperative surface operations in future lunar missions. However, cooperative car

safetyarxiv-cs-ro
2 Jul 2026
Safety

ECoSim: Data Efficient Fine-Tuning for Controllable Traffic Simulation

DGX agent

arXiv:2607.00545v1 Announce Type: new Abstract: Controllable traffic simulation is critical for testing autonomous driving systems, yet existing approaches often require retraining large generative mo

safetyarxiv-cs-cv
2 Jul 2026
Model Releases

FLYNN: Robust Neural Network for Robot Navigation using Fly Brain Topology

DGX agent

arXiv:2607.00025v1 Announce Type: cross Abstract: While deep learning models achieve state-of-the-art performance in complex tasks, they remain brittle when faced with new environments or sensory depr

model-releasesarxiv-cs-ai
2 Jul 2026
Local Ai

Interact3D: Compositional 3D Generation of Interactive Objects

DGX agent

arXiv:2603.16085v2 Announce Type: replace-cross Abstract: Recent breakthroughs in 3D generation have enabled the synthesis of high-fidelity individual assets. However, generating 3D compositional obje

local-aiarxiv-cs-ai
2 Jul 2026
Model Releases

Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training

DGX agent

arXiv:2607.01232v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a central component of post-training large language models (LLMs), yet little is understood about how RL adapta

model-releasesarxiv-cs-cl
2 Jul 2026
Safety

Learning Gait-Aware Quadruped Locomotion with Temporal Logic Specifications

DGX agent

arXiv:2607.00442v1 Announce Type: cross Abstract: Reinforcement learning (RL) for quadruped locomotion commonly depends on fixed, hand-crafted, and Markovian reward functions that limit both interpret

safetyarxiv-cs-ai
2 Jul 2026
Safety

Learning Generalizable Skill Policy with Data-Efficient Unsupervised RL

DGX agent

arXiv:2607.00392v1 Announce Type: cross Abstract: Unsupervised Reinforcement Learning (URL) aims to pre-train scalable, skill-conditioned policies without extrinsic rewards, serving as a foundation fo

safetyarxiv-cs-ai
2 Jul 2026
← Previous
1…209210211212213…233
Next →