AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries90,316
  • Agents7,714
  • Applications5,508
  • Concepts5
  • Hardware1,905
  • Industry6,191
  • Local Ai5,052
  • Model Releases24,539
  • Research20,616
  • Safety13,635
  • Syntheses17
  • Tools1,678
  • Tutorials3,456

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries90,316
  • Agents7,714
  • Applications5,508
  • Concepts5
  • Hardware1,905
  • Industry6,191
  • Local Ai5,052
  • Model Releases24,539
  • Research20,616
  • Safety13,635
  • Syntheses17
  • Tools1,678
  • Tutorials3,456

Source
Human
90,316Total entries
1Added by human
90,315Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
90,315 results
29 May 2026

ReasonBreak: Probing Vulnerabilities in Reasoning-Enabled Vision-Language-Action Models for Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.29114v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models with integrated reasoning have been proposed for end-to-end autonomous driving, assuming a tight coupling between

Reasoning about Reasoning: BAPO Bounds on Chain-of-Thought Token Complexity in LLMs

ResearchDGX agent

arXiv:2602.02909v2 Announce Type: replace Abstract: Inference-time scaling via chain-of-thought (CoT) reasoning is a major driver of state-of-the-art LLM performance, but it comes with substantial lat

Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning

Model ReleasesDGX agent
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2602.00994v2 Announce Type: replace Abstract: Agentic Reinforcement Learning (ARL) trains large language models to interleave reasoning with external tool execution to solve complex tasks. Most

Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization

Model ReleasesDGX agent

arXiv:2605.29327v1 Announce Type: new Abstract: Efficient Distillation (EDistill) compresses large language models (LLMs) by structured pruning parameters and tuning lightweight modules with high trai

Reasoning that Travels: Dissecting How Chain-of-Thought Transfers Across Models

TutorialsDGX agent

arXiv:2605.28913v1 Announce Type: new Abstract: Large reasoning models (LRMs) often generate extensive chain-of-thought (CoT) traces before producing a final answer. As explicit textual artifacts, the

Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought

Model ReleasesDGX agent

arXiv:2603.05488v4 Announce Type: replace-cross Abstract: We provide evidence of performative chain-of-thought (CoT) in reasoning models, where a model becomes strongly confident in its final answer,

Reasoning While Asking: Transforming Reasoning Large Language Models from Passive Solvers to Proactive Inquirers

SafetyDGX agent

arXiv:2601.22139v2 Announce Type: replace-cross Abstract: Reasoning-oriented Large Language Models (LLMs) have achieved remarkable progress with Chain-of-Thought (CoT) prompting, yet they remain funda

Reasoning with Sampling: Cutting at Decision Points

Local AiDGX agent

arXiv:2605.30327v1 Announce Type: cross Abstract: Frontier reasoning models are produced by posttraining base language models with reinforcement learning. Recent work has challenged this by showing th

ReasonLight: A Multimodal Foundation Model-Enhanced Reinforcement Learning Framework for Zero-Shot Traffic Signal Control

SafetyDGX agent

arXiv:2605.29425v1 Announce Type: new Abstract: Reinforcement learning (RL) has shown promise in traffic signal control (TSC). However, its reliance on predefined states limits responsiveness to obser

ReasonOps: Operator Segmentation for LLM Reasoning Traces

Model ReleasesDGX agent

arXiv:2605.29192v1 Announce Type: new Abstract: Chain-of-thought traces from large reasoning models can span tens of thousands of tokens, yet we lack a vocabulary for describing their internal structu

Reconstructing software engineering around AI is going to take work (even as the ability of AI to code increases at a rapid rate). Organizat…

ApplicationsDGX agent

Reconstructing software engineering around AI is going to take work (even as the ability of AI to code increases at a rapid rate). Organizations are ideally spending tokens for two things: 1) building

Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMs

Model ReleasesDGX agent

arXiv:2605.30021v1 Announce Type: new Abstract: Many open-ended instructions have multiple valid answers that users can benefit from seeing, but post-training often narrows an LLM's output space towar

Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents

SafetyDGX agent

arXiv:2605.29447v1 Announce Type: cross Abstract: While GUI agents have advanced rapidly, they often lack the robustness to recover from their own errors, hindering real-world deployment. To bridge th

Rectified LpJEPA: Joint-Embedding Predictive Architectures with Sparse and Maximum-Entropy Representations

TutorialsDGX agent

arXiv:2602.01456v2 Announce Type: replace-cross Abstract: Joint-Embedding Predictive Architectures (JEPA) learn view-invariant representations and admit projection-based distribution matching for coll

Recurrent Structural Policy Gradient for Partially Observable Mean Field Games

SafetyDGX agent

arXiv:2602.20141v2 Announce Type: replace Abstract: Mean Field Games (MFGs) provide a principled framework for modelling interactions in large population systems. However, algorithmic progress has bee

Reducing Experimental Testing in Space Propulsion Film Cooling Analyses by Pixelwise Generative Image Interpolation

ResearchDGX agent

arXiv:2605.29911v1 Announce Type: cross Abstract: We propose a machine learning approach for image regression from sparse experimental measurements. We show the application of the proposed method on f

Redundant or Necessary? A Benchmark for Detecting Redundant Steps in Agent Trajectories

Model ReleasesDGX agent

arXiv:2605.29893v1 Announce Type: new Abstract: LLM-based agents have demonstrated strong capabilities in solving complex tasks through multi-step reasoning and tool use. However, existing evaluation

Reinforcement Learning with Robust Rubric Rewards

ResearchDGX agent

arXiv:2605.30244v1 Announce Type: cross Abstract: While Reinforcement Learning with Verifiable Rewards (RLVR) is effective for deterministically checkable tasks, many vision-language tasks are partial

Reinforcing Few-step Generators via Reward-Tilted Distribution Matching

SafetyDGX agent

arXiv:2605.26108v2 Announce Type: replace Abstract: Recent advances in few-step diffusion distillation have enabled efficient image generation, yet aligning these models with human preferences remains

Rel-MOSS: Towards Imbalanced Relational Deep Learning on Relational Databases

ResearchDGX agent

arXiv:2603.07916v2 Announce Type: replace Abstract: In recent advances, to enable a fully data-driven learning paradigm on relational databases (RDB), relational deep learning (RDL) is proposed to str

Relational Rank Geometry in Transformers: Detecting and Steering Hidden-State Relation Frames

Model ReleasesDGX agent

arXiv:2605.29634v1 Announce Type: new Abstract: Transformer hidden states are often interpreted through local or low-order objects: neurons, sparse features, attention heads, residual-stream direction

Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents

Model ReleasesDGX agent

arXiv:2605.29224v1 Announce Type: cross Abstract: AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incorporating

Reliability shouldn't require reserving GPUs. Serverless 2.0 is live on Fireworks: one API, 3 serving paths. → Standard: elastic default → P…

ToolsDGX agent

Reliability shouldn't require reserving GPUs. Serverless 2.0 is live on Fireworks: one API, 3 serving paths. → Standard: elastic default → Priority: sheds last under congestion, pricing ~1.5x standard

Reliable Reasoning with Large Language Models via Preference-Based Maximum Satisfiability

ResearchDGX agent

arXiv:2605.29687v1 Announce Type: new Abstract: Large Language Models (LLMs) excel at understanding natural language but struggle with optimisation tasks involving multiple constraints and user-define

Replicable Simulation-Based Robot Validation through Provenance

SafetyDGX agent

arXiv:2605.29973v1 Announce Type: new Abstract: Robot behavior is often validated through simulation-based testing, yet the replicability of such campaigns depends critically on transparent documentat

REPOT: Recoverable Program-of-Thought via Checkpoint Repair

Model ReleasesDGX agent

arXiv:2605.30052v1 Announce Type: cross Abstract: One-shot Program-of-Thought (PoT) emits a Python program that prints a primitive-action plan; a single invalid action silently invalidates the traject

Representation Alignment Rests on Linear Structure

SafetyDGX agent

arXiv:2605.28870v1 Announce Type: cross Abstract: We investigate the Platonic Representation Hypothesis (PRH) through a tripartite statistical framework of representations: signal, bias, and noise. {1

Representation Signatures and Risk-Feedback Alignment in LLM Trading Agents

SafetyDGX agent

arXiv:2605.28850v1 Announce Type: new Abstract: We study behavioral alignment and representation dynamics of large language model (LLM) agents in financial decision environments. Using TradeArena, an

Representation Unlearning: Forgetting through Information Compression

Model ReleasesDGX agent

arXiv:2601.21564v2 Announce Type: replace Abstract: Machine unlearning seeks to remove the influence of specific training data from a model, a need driven by privacy regulations and robustness concern

Resolution as a Direction: Vector-Panning Feature Alignment for Cross-Resolution Re-Identification

SafetyDGX agent

arXiv:2510.00936v2 Announce Type: replace Abstract: Cross-resolution person re-identification (CR-ReID) remains challenging in practical surveillance, where camera quality and capture distance lead to

Resolution Diagnostics for Paired LLM Evaluation

ResearchDGX agent

arXiv:2605.30315v1 Announce Type: new Abstract: Across two public LLM leaderboards, many displayed pairwise rankings do not meet a conventional paired-test resolution target under the actual paired ev

Resolving Endpoint Underfitting in Diffusion Bridges via Noise Alignment

SafetyDGX agent

arXiv:2605.28962v1 Announce Type: new Abstract: Diffusion bridge models offer a powerful framework for connecting two data distributions, such as in image restoration and translation. Many existing me

REST3D: Reconstructing Physically Stable 3D Scenes from a Single Image

SafetyDGX agent

arXiv:2605.30338v1 Announce Type: new Abstract: Reconstructing physically stable 3D scenes from a single RGB image enables casual images to be converted into simulation-ready digital assets for applic

Rethinking FID Through the Geometry of the Reference Dataset

ResearchDGX agent

arXiv:2605.29335v1 Announce Type: cross Abstract: Frechet Inception Distance (FID) is widely used to evaluate image generators, yet lower FID does not always correspond to better sample quality. We sh

Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth

Model ReleasesDGX agent

arXiv:2605.29234v1 Announce Type: new Abstract: We study large-scale literature search from two complementary angles: improving the retrieval pipeline, and stress-testing the human reference list as a

Rethinking Post-Training Recipes for Multimodal Time-Series Forecasting

Model ReleasesDGX agent

arXiv:2605.29401v1 Announce Type: new Abstract: Time-Series Foundation Models (TSFMs) excel at zero-shot unimodal forecasting using numerical data, but unlike LLMs they cannot consume multimodal, non-

Rethinking Stepwise Model Routing: A Cost-Efficient Table Reasoning Perspective

ResearchDGX agent

arXiv:2605.29319v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong performance on table reasoning tasks but incur substantial inference cost due to long reasoning traces. Ste

Return-to-Go Is More Than a Number: Q-Guided Alignment for Return-Conditioned Supervised Learning

Model ReleasesDGX agent

arXiv:2605.29028v1 Announce Type: cross Abstract: Conditioned Sequence Models (CSMs) learn policies by treating return-to-go (RTG) as a control signal. However, existing CSMs often treat the RTGs as s

Review Arcade: On the Human Alignment and Gameability of LLM Reviews

SafetyDGX agent

arXiv:2605.28897v1 Announce Type: new Abstract: LLM-generated reviews for scientific papers are gaining considerable traction and are even being officially piloted by major conferences. We have to ass

Revisiting Observation Reduction for Web Agents: Comprehensive Evaluation with a Lightweight Framework

AgentsDGX agent

arXiv:2605.29397v1 Announce Type: new Abstract: HTML observations in LLM-based web agents are extremely long, and while many reduction methods have been proposed, it remains unclear which methods redu

RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with Large Language Models

AgentsDGX agent

arXiv:2603.18859v2 Announce Type: replace Abstract: Reinforcement learning (RL) shows promise for enhancing LLM agentic reasoning, yet sparse terminal rewards hinder fine-grained optimization. Process

RHO: Robust Holistic OSM-Based Metric Cross-View Geo-Localization

Model ReleasesDGX agent

arXiv:2603.27758v2 Announce Type: replace Abstract: Metric Cross-View Geo-Localization (MCVGL) aims to estimate the 3-DoF camera pose (position and heading) by matching ground and satellite images. In

Ridge Regression from Poisson Resetting: A Renewal Perspective on Spectral Regularization

ResearchDGX agent

arXiv:2605.30059v1 Announce Type: new Abstract: We connect stochastic resetting from non-equilibrium statistical physics with ridge regularization in statistical learning. For linear gradient flow, re

Riemannian AmbientFlow: Towards Simultaneous Manifold Learning and Generative Modeling from Corrupted Data

ResearchDGX agent

arXiv:2601.18728v2 Announce Type: replace Abstract: Modern generative modeling methods have demonstrated strong performance in learning complex data distributions from clean samples. In many scientifi

RightNow-Arabic-0.5B-Turbo: An Open Sub-1B Arabic Language Model via Vocabulary Injection and Edge-First Deployment

Model ReleasesDGX agent

arXiv:2605.28827v1 Announce Type: new Abstract: Open Arabic large language models split into two classes: sub-1B multilingual models that treat Arabic as an afterthought (Qwen2.5-0.5B, Falcon-H1-0.5B)

Risk-averse Fair Multi-class Classification

Model ReleasesDGX agent

arXiv:2509.05771v2 Announce Type: replace-cross Abstract: We develop a new classification framework based on the theory of coherent risk measures and systemic risk. The proposed approach is suitable f

RL2ML: Finite-Rollout Surrogate Objectives from Reinforcement Learning to Maximum Likelihood

SafetyDGX agent

arXiv:2605.30154v1 Announce Type: new Abstract: Correctness-based Reinforcement Learning with Verifiable Rewards (RLVR) trains language models from binary feedback on sampled outputs, but the objectiv

RoboWits: Unexpected Challenges for Robotic Creative Problem Solving

Model ReleasesDGX agent

arXiv:2605.30326v1 Announce Type: cross Abstract: The ability to reason, adapt, and creatively solve problems under unexpected challenges is essential for robots operating in real-world environments.

Robust and Efficient Guardrails with Latent Reasoning

Model ReleasesDGX agent

arXiv:2605.29068v1 Announce Type: new Abstract: Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in real-world applications. Existing safety guardrai

Robust and Efficient Writer-Independent IMU-Based Handwriting Recognition

ApplicationsDGX agent

arXiv:2502.20954v3 Announce Type: replace Abstract: Handwriting recognition (HWR) using inertial measurement unit (IMU) data remains challenging due to variations in writing styles and the limited ava

Robust and Generalizable Safety Steering for Text-to-Image Diffusion Transformers

SafetyDGX agent

arXiv:2605.30049v1 Announce Type: new Abstract: Diffusion Transformers have become a powerful backbone for text-to-image generation, but their layered and cross-modal generation process makes safety c

Robust Cross-Domain Generalization Using Unlabeled Target Data with Source-Domain Supervision

TutorialsDGX agent

arXiv:2605.29122v1 Announce Type: new Abstract: It is often desirable to generalize medical imaging AI models trained with dense annotations to data acquired from different ultrasound scanners or clin

Robust Frequency-Calibrated Virtual EEG Channel Generation from Four Frontal Electrodes for Wearable EEG Augmentation

ResearchDGX agent

arXiv:2605.29263v1 Announce Type: new Abstract: Low-channel wearable electroencephalography (EEG) is attractive for long-term monitoring, but four frontal electrodes provide only a sparse and spatiall

Rocket Report: A dark day for Blue Origin; Pentagon eyes new launch site

IndustryDGX agent

Blue Origin's New Glenn rocket exploded on the launch pad at Cape Canaveral on May 28 during a static fire engine test, destroying the 321-foot rocket . The explosion froze all 24 of Amazon's contract

Rooted Absorbed Prefix Trajectory Balance with Submodular Replay for GFlowNet Training

SafetyDGX agent

arXiv:2603.00454v2 Announce Type: replace-cross Abstract: Generative Flow Networks (GFlowNets) enable fine-tuning large language models to approximate reward-proportional posteriors, but they remain p

Routing by Reaching: Composition of Pre-trained GFlowNets for Multi-Objective Generation

TutorialsDGX agent

arXiv:2602.21565v2 Announce Type: replace Abstract: Generative Flow Networks (GFlowNets) learn to sample diverse candidates in proportion to a reward function, making them well-suited for scientific d

RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains

SafetyDGX agent

arXiv:2605.29156v1 Announce Type: cross Abstract: Pointwise reward modeling offers critical signals for LLM post-training, yet struggles with absolute scoring in subjective, non-verifiable settings. R

Rubric-Guided Process Reward for Stepwise Model Routing

SafetyDGX agent

arXiv:2605.29310v1 Announce Type: new Abstract: Stepwise model routing improves the efficiency of Large Reasoning Models (LRMs) by assigning each reasoning step to a suitable model. Recent methods for

RUD (rapid unscheduled disassembly) events are not unusual in the rocket world

IndustryDGX agent

RUD (rapid unscheduled disassembly) is a term used in the aerospace industry to describe unexpected vehicle failures or explosions during rocket testing and launches. Elon Musk's statement acknowledge

Run Docker containers inside Vercel Sandbox

ToolsDGX agent

Vercel announced the ability to run Docker containers directly within Vercel Sandbox, enabling developers to execute containerized applications and services as part of their development and testing wo

← Previous
1…790791792793794…1506
Next →