AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,188
  • Agents7,322
  • Applications5,231
  • Concepts5
  • Hardware1,770
  • Industry6,109
  • Local Ai4,762
  • Model Releases22,797
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,188
  • Agents7,322
  • Applications5,231
  • Concepts5
  • Hardware1,770
  • Industry6,109
  • Local Ai4,762
  • Model Releases22,797
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlog
85,188Total entries
1Added by human
85,187Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,688 results
Local Ai

Reasoning or Fluency? Dissecting Probabilistic Confidence in Best-of-N Selection

DGX agent

arXiv:2601.13735v2 Announce Type: replace Abstract: Probabilistic confidence metrics are increasingly adopted as proxies for reasoning quality in Best-of-N selection, under the assumption that higher

local-aiarxiv-cs-ai
4 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Local Ai

Recover-LoRA for Aggressive Quantization: Reclaiming Accuracy in 2-Bit Language Models via Low-Rank Adaptation with Knowledge Distillation on Synthetic Data

DGX agent

arXiv:2606.04238v1 Announce Type: cross Abstract: Aggressive weight quantization to 2-bit precision offers substantial throughput and memory gains for large language model (LLM) inference, but typical

local-aiarxiv-cs-ai
4 Jun 2026
Safety

Reinforcement Learning from Rich Feedback with Distributional DAgger

DGX agent

arXiv:2606.05152v1 Announce Type: cross Abstract: Reasoning models have advanced rapidly, but the dominant reinforcement learning from verifiable rewards (RLVR) recipe remains surprisingly narrow: sam

safetyarxiv-cs-ai
4 Jun 2026
Safety

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning

DGX agent

arXiv:2606.04923v1 Announce Type: cross Abstract: Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. However, policy models

safetyarxiv-cs-ai
4 Jun 2026
Safety

Rethinking Sales Lead Scoring with LLM-based Hierarchical Preference Ranking

DGX agent

arXiv:2606.04387v1 Announce Type: cross Abstract: Sales lead conversion in high-stakes domains (e.g., automotive, real estate) differs fundamentally from e-commerce recommendation due to prolonged dec

safetyarxiv-cs-ai
4 Jun 2026
Research

Revisiting Model Stitching In the Foundation Model Era

DGX agent

arXiv:2603.12433v3 Announce Type: replace-cross Abstract: Model stitching, connecting early layers of one model (source) to later layers of another (target) via a light stitch layer, has served as a p

researcharxiv-cs-ai
4 Jun 2026
Model Releases

Revisiting Vul-RAG: Reproducibility and Replicability of RAG-based Vulnerability Detection with Open-Weight Models

DGX agent

arXiv:2606.04739v1 Announce Type: cross Abstract: Large language models (LLMs) have shown strong potential for automated software vulnerability detection, particularly in retrieval-augmented generatio

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

Rollout-Level Advantage-Prioritized Experience Replay for GRPO

DGX agent

arXiv:2606.04560v1 Announce Type: cross Abstract: Reinforcement learning from verifiable rewards with GRPO is a standard approach for post-training reasoning LLMs. It remains sample inefficient. Each

model-releasesarxiv-cs-ai
4 Jun 2026
Local Ai

RowNet: A Memory Transformer for Tabular Regression

DGX agent

arXiv:2606.04445v1 Announce Type: cross Abstract: Real estate valuation is a structured regression problem in which prices are governed by heterogeneous feature types, sparse regional effects, nonline

local-aiarxiv-cs-ai
4 Jun 2026
Safety

RUBAS: Rubric-Based Reinforcement Learning for Agent Safety

DGX agent

arXiv:2606.04051v1 Announce Type: cross Abstract: The evolution of LLMs into tool-enabled agents creates a new class of safety challenges associated with real-world execution rather than simple text g

safetyarxiv-cs-ai
4 Jun 2026
Model Releases

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety

DGX agent

arXiv:2603.10044v2 Announce Type: replace-cross Abstract: A safety score earned on a benchmark need not predict how the same model behaves once it is wrapped in an agentic scaffold the benchmark never

model-releasesarxiv-cs-ai
4 Jun 2026
Research

SaliMory: Orchestrating Cognitive Memory for Conversational Agents

DGX agent

arXiv:2606.04120v1 Announce Type: cross Abstract: Conversational agents that serve as lifelong companions must maintain persistent memory across all interactions. However, simply expanding context win

researcharxiv-cs-ai
4 Jun 2026
Model Releases

SAM 3D: 3Dfy Anything in Images

DGX agent

arXiv:2511.16624v2 Announce Type: replace-cross Abstract: We present SAM 3D, a generative model for visually grounded 3D object reconstruction, predicting geometry, texture, and layout from a single i

model-releasesarxiv-cs-ai
4 Jun 2026
Hardware

Scaling Novel Graph Generation via Lightweight Structure-Guided Autoregressive Models

DGX agent

arXiv:2606.04287v1 Announce Type: cross Abstract: Generating realistic and diverse graphs is a key problem in machine learning, with applications in molecular discovery, circuit design, cybersecurity,

hardwarearxiv-cs-ai
4 Jun 2026
Safety

Scaling Self-Evolving Agents via Parametric Memory

DGX agent

arXiv:2606.04536v1 Announce Type: new Abstract: Existing memory-augmented LLM agents store past experience exclusively in prompt space, as textual summaries or retrieved passages, while keeping model

safetyarxiv-cs-ai
4 Jun 2026
Safety

Scenario Generation for Risk-Aware Reinforcement Learning with Probably Approximately Safe Guarantees

DGX agent

arXiv:2606.04812v1 Announce Type: cross Abstract: Guaranteeing safety is critical to the deployment of reinforcement learning (RL) agents in the real-world, especially as policies learned using deep R

safetyarxiv-cs-ai
4 Jun 2026
Research

SCI-PRM: A Tool Aware Process Reward Model for Scientific Reasoning Verification

DGX agent

arXiv:2606.04579v1 Announce Type: new Abstract: While Process Reward Models (PRMs) have achieved remarkable success in mathematical reasoning, their application in complex scientific domains-such as b

researcharxiv-cs-ai
4 Jun 2026
Safety

Selective Coupling of Decoupled Informative Regions: Masked Attention Alignment for Data-Free Quantization of Vision Transformers

DGX agent

arXiv:2606.04373v1 Announce Type: cross Abstract: Data-Free Quantization (DFQ) addresses data security concerns by synthesizing samples, without accessing real data. It has garnered increasing attenti

safetyarxiv-cs-ai
4 Jun 2026
Model Releases

Self-Evolving Deep Research via Joint Generation and Evaluation

DGX agent

arXiv:2606.04507v1 Announce Type: cross Abstract: Large Language Models (LLMs) have become increasingly adopted in daily applications, with deep research standing out as a particularly important capab

model-releasesarxiv-cs-ai
4 Jun 2026
Agents

Self-Reflective APIs: Structure Beats Verbosity for AI Agent Recovery

DGX agent

arXiv:2606.05037v1 Announce Type: cross Abstract: When an AI agent calls an API and hits a validation error, it needs more than what went wrong -- it needs what to do next. A self-reflective API retur

agentsarxiv-cs-ai
4 Jun 2026
Agents

Semantic Constraint Synthesis for Adaptive Trajectory Optimization via Large Language Models

DGX agent

arXiv:2606.04123v1 Announce Type: cross Abstract: Trajectory optimization is a critical component for enabling safe and reliable autonomous operations in space exploration. As space missions increase

agentsarxiv-cs-ai
4 Jun 2026
Safety

Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model

DGX agent

arXiv:2512.21917v3 Announce Type: replace-cross Abstract: Policy alignment to preference data typically assumes a known link function between observed preferences and latent rewards (e.g., Bradley-Ter

safetyarxiv-cs-ai
4 Jun 2026
Agents

SePO: Self-Evolving Prompt Agent for System Prompt Optimization

DGX agent

arXiv:2606.04465v1 Announce Type: cross Abstract: System prompt optimization improves agent behavior without modifying the underlying model, yielding human-readable, model-agnostic instructions. Exist

agentsarxiv-cs-ai
4 Jun 2026
Research

SFMambaNet: Spectral-Frequency Enhanced Selective State Space Model for Correspondence Pruning

DGX agent

arXiv:2606.04493v1 Announce Type: cross Abstract: Correspondence pruning aims to identify inliers from an initial set of correspondences. Most existing Graph Neural Network (GNN)-based methods rely on

researcharxiv-cs-ai
4 Jun 2026
Research

SharedRequest: Privacy-Preserving Model-Agnostic Inference for Large Language Models

DGX agent

arXiv:2606.05004v1 Announce Type: cross Abstract: With the widespread deployment of public large language models (LLMs) such as ChatGPT, protecting user prompt privacy has become an increasingly criti

researcharxiv-cs-ai
4 Jun 2026
Agents

ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling

DGX agent

arXiv:2603.02697v2 Announce Type: replace-cross Abstract: This paper presents ShareVerse, a video generation framework enabling multi-agent shared world modeling, addressing the gap in existing works

agentsarxiv-cs-ai
4 Jun 2026
Model Releases

Signed Dual Attention: Capturing Signed Dependencies in Time Series Forecasting

DGX agent

arXiv:2606.04833v1 Announce Type: cross Abstract: Initially developed for natural language processing, Transformer architectures and attention mechanisms are now central to a wide range of deep learni

model-releasesarxiv-cs-ai
4 Jun 2026
Safety

Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents

DGX agent

arXiv:2510.13704v2 Announce Type: replace-cross Abstract: Recent works have proposed accelerating the wall-clock training time of actor-critic methods via the use of large-scale environment paralleliz

safetyarxiv-cs-ai
4 Jun 2026
Research

Simulate, Reason, Decide: Scientific Reasoning with LLMs for Simulation-Driven Decision Making

DGX agent

arXiv:2606.04505v1 Announce Type: new Abstract: Scientific simulators are increasingly being integrated into LLM-driven systems for high-stakes simulation-driven decision-making. However, existing fra

researcharxiv-cs-ai
4 Jun 2026
Model Releases

SMAC-Talk: A Natural Language Extension of the StarCraft Multi-Agent Challenge for Large Language Models

DGX agent

arXiv:2606.04202v1 Announce Type: new Abstract: As LLMs become more widely deployed, they are increasingly expected to work alongside other AI agents rather than operating in isolation. Effective coor

model-releasesarxiv-cs-ai
4 Jun 2026
Research

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots

DGX agent

arXiv:2606.04503v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has greatly advanced large reasoning models (LRMs), but it requires timely training on a huge fu

researcharxiv-cs-ai
4 Jun 2026
Safety

Smart Transportation Without Neurons -- Fair Metro Network Expansion with Tabular Reinforcement Learning

DGX agent

arXiv:2606.04167v1 Announce Type: cross Abstract: We tackle the Metro Network Expansion Problem (MNEP), a subset of the Transport Network Design Problem (TNDP), which focuses on expanding metro system

safetyarxiv-cs-ai
4 Jun 2026
Safety

SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference Optimization

DGX agent

arXiv:2505.11166v3 Announce Type: replace-cross Abstract: Despite advances in pretraining with extended context sizes, large language models (LLMs) still face challenges in effectively utilizing real-

safetyarxiv-cs-ai
4 Jun 2026
Tutorials

Sparse Mixture-of-Experts Reward Models Learn Interpretable and Specialized Experts for Personalized Preference Modeling

DGX agent

arXiv:2606.04284v1 Announce Type: cross Abstract: Preference modeling plays a central role in reinforcement learning from human feedback (RLHF), enabling large language models (LLMs) to align with hum

tutorialsarxiv-cs-ai
4 Jun 2026
Research

Spatial Transcriptomics as Images for Large-Scale Pretraining

DGX agent

arXiv:2603.13432v4 Announce Type: replace-cross Abstract: Spatial Transcriptomics (ST) profiles thousands of gene expression values at discrete spots with precise coordinates on tissue sections, prese

researcharxiv-cs-ai
4 Jun 2026
Research

Spectral Scaling Laws of Muon

DGX agent

arXiv:2606.04058v1 Announce Type: cross Abstract: Orthonormalized update rules have rapidly become a leading choice of optimizer for training large language models, with recent open-source state-of-th

researcharxiv-cs-ai
4 Jun 2026
Model Releases

Speculative Thinking: Enhancing Small-Model Reasoning with Large Model Guidance at Inference Time

DGX agent

arXiv:2504.12329v2 Announce Type: replace-cross Abstract: Recent advances leverage post-training to enhance model reasoning performance, which typically requires costly training pipelines and still su

model-releasesarxiv-cs-ai
4 Jun 2026
Applications

SSSD: Simply-Scalable Speculative Decoding

DGX agent

arXiv:2411.05894v3 Announce Type: replace-cross Abstract: Speculative Decoding has emerged as a popular technique for accelerating inference in Large Language Models. However, most existing approaches

applicationsarxiv-cs-ai
4 Jun 2026
Model Releases

StandardE2E: A Unified Framework for End-to-End Autonomous Driving Datasets

DGX agent

arXiv:2606.04271v1 Announce Type: cross Abstract: Autonomous driving has shifted from modular perception-prediction-planning stacks toward end-to-end (E2E) models that map sensor inputs directly to ve

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

StepPRM-RTL: Stepwise Process-Reward Guided LLM Fine-Tuning for Enhanced RTL Synthesis

DGX agent

arXiv:2606.04246v1 Announce Type: new Abstract: Automatic generation of RTL code for digital hardware designs remains challenging due to long-horizon reasoning, multi-step dependencies, and strict cor

model-releasesarxiv-cs-ai
4 Jun 2026
Agents

Strabo: Declarative Specification and Implementation of Agentic Interaction Protocols

DGX agent

arXiv:2606.05043v1 Announce Type: new Abstract: The last few years have witnessed major advances in the modeling and implementation of multiagent systems based on declarative interaction protocols. Ou

agentsarxiv-cs-ai
4 Jun 2026
Model Releases

Streaming Communication in Multi-Agent Reasoning

DGX agent

arXiv:2606.05158v1 Announce Type: cross Abstract: Multi-agent reasoning systems adopt a 'generate-then-transfer' paradigm that forces end-to-end latency to scale linearly with pipeline depth. We intro

model-releasesarxiv-cs-ai
4 Jun 2026
Safety

Stumbling Into AI Emotional Dependence: How Routine AI Interactions Reshape Human Connection

DGX agent

arXiv:2606.04150v1 Announce Type: new Abstract: Public discourse and emerging policy typically assume that AI emotional support is a deliberate act: a lonely user consciously seeking comfort from a de

safetyarxiv-cs-ai
4 Jun 2026
Safety

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success

DGX agent

arXiv:2601.18175v2 Announce Type: replace Abstract: A widely used technique for improving policies is success conditioning, in which one collects trajectories, identifies those that achieve a desired

safetyarxiv-cs-ai
4 Jun 2026
Research

Supportive Token Revealing for Fast Diffusion Language Model Decoding

DGX agent

arXiv:2606.04236v1 Announce Type: cross Abstract: Discrete diffusion language models can generate text efficiently by updating multiple masked positions in parallel, but this parallelism introduces a

researcharxiv-cs-ai
4 Jun 2026
Agents

SUSD: Structured Unsupervised Skill Discovery through State Factorization

DGX agent

arXiv:2602.01619v2 Announce Type: replace-cross Abstract: Unsupervised Skill Discovery (USD) aims to autonomously learn a diverse set of skills without relying on extrinsic rewards. One of the most co

agentsarxiv-cs-ai
4 Jun 2026
Model Releases

SymTRELLIS: Symmetry-Enforced Voxel Latents for 3D Generation

DGX agent

arXiv:2606.04108v1 Announce Type: cross Abstract: Single-view 3D generative models have achieved impressive visual quality, yet they are not designed to satisfy structural or functional requirements,

model-releasesarxiv-cs-ai
4 Jun 2026
Research

Synthetic Personalities: How Well Can LLMs Mimic Individual Respondents Using Socio-Economic Microdata?

DGX agent

arXiv:2606.04592v1 Announce Type: cross Abstract: LLM-based digital twins promise to scale and accelerate market research, but most published twins are either coarse persona bots conditioned on a few

researcharxiv-cs-ai
4 Jun 2026
← Previous
1…206207208209210…452
Next →