AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
11 May 2026

PPI-Net connects molecular protein interactions to functional processes in disease

ResearchDGX agent

arXiv:2605.07838v1 Announce Type: cross Abstract: Understanding how molecular alterations propagate across biological systems to drive disease remains a central challenge. Although high-throughput pro

Predictive but Not Plannable: RC-aux for Latent World Models

Local AiDGX agent

arXiv:2605.07278v1 Announce Type: cross Abstract: A latent world model may achieve accurate short-horizon prediction while still inducing a latent space that is poorly aligned with planning. A key iss

Pretraining a Foundation Model for Small-Molecule Natural Products

ResearchDGX agent

arXiv:2503.17656v4 Announce Type: replace-cross Abstract: Natural products, as metabolites from microorganisms, animals, or plants, exhibit diverse biological activities, making them crucial for drug


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Proactive Instance Navigation with Comparative Judgment for Ambiguous User Queries

AgentsDGX agent

arXiv:2605.06223v2 Announce Type: replace Abstract: Natural-language instance navigation becomes challenging when the initial user request does not uniquely specify the target instance. A practical ag

ProactiveMobile: A Comprehensive Benchmark for Boosting Proactive Intelligence on Mobile Devices

Model ReleasesDGX agent

arXiv:2602.21858v4 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have made significant progress in mobile agent development, yet their capabilities are predominantly confin

Prompt Engineering Strategies for LLM-based Qualitative Coding of Psychological Safety in Software Engineering Communities: A Controlled Empirical Study

Model ReleasesDGX agent

arXiv:2605.07422v1 Announce Type: cross Abstract: Qualitative analysis plays a pivotal role in understanding the human and social aspects of software engineering. However, it remains a demanding proce

ProteinJEPA: Latent prediction complements protein language models

ResearchDGX agent

arXiv:2605.07554v1 Announce Type: cross Abstract: Protein language models are trained primarily with masked language modeling (MLM), which predicts amino-acid identities at masked positions. We ask wh

Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning

SafetyDGX agent

arXiv:2605.07804v1 Announce Type: cross Abstract: On-policy distillation (OPD) leverages dense teacher rewards to enhance reasoning models. However, scaling OPD to long-horizon tasks exposes a critica

PSK@EEUCA 2026: Fine-Tuning Large Language Models with Synthetic Data Augmentation for Multi-Class Toxicity Detection in Gaming Chat

Model ReleasesDGX agent

arXiv:2605.07201v1 Announce Type: cross Abstract: This paper describes our system for the EEUCA 2026 Shared Task on Understanding Toxic Behavior in Gaming Communities. The task involves classifying Wo

Q-MMR: Off-Policy Evaluation via Recursive Reweighting and Moment Matching

SafetyDGX agent

arXiv:2605.06474v2 Announce Type: replace-cross Abstract: We present a novel theoretical framework, Q-MMR, for off-policy evaluation in finite-horizon MDPs. Q-MMR learns a set of scalar weights, one f

Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation

Model ReleasesDGX agent

arXiv:2605.07647v1 Announce Type: cross Abstract: Automated short answer scoring (ASAS) is shifting from discriminative, fine-tuned models to large language models (LLMs) used in few-shot settings. Th

Query-efficient model evaluation using cached responses

Model ReleasesDGX agent

arXiv:2605.07096v1 Announce Type: cross Abstract: Evaluating a new model on an existing benchmark is often necessary to understand its behavior before deployment. For modern evaluation frameworks, gen

Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding

Model ReleasesDGX agent

arXiv:2605.07141v1 Announce Type: cross Abstract: Open-world referring segmentation requires grounding unconstrained language expressions to precise pixel-level regions. Existing multimodal large lang

R-GTD: A Geometric Analysis of Gradient Temporal-Difference Learning in Singular Regimes

SafetyDGX agent

arXiv:2601.20599v2 Announce Type: replace-cross Abstract: Gradient temporal-difference (GTD) learning algorithms are widely used for off-policy policy evaluation with function approximation. However,

R^3L: Reasoning 3D Layouts from Relative Spatial Relations

ResearchDGX agent

arXiv:2605.06758v1 Announce Type: cross Abstract: Relative spatial relations provide a compact representation of spatial structure and are fundamental to relative spatial reasoning in 3D layout genera

Randomness is sometimes necessary for coordination

Model ReleasesDGX agent

arXiv:2605.06825v1 Announce Type: new Abstract: Full parameter sharing is standard in cooperative multi-agent reinforcement learning (MARL) for homogeneous agents. Under permutation-symmetric observat

Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners

SafetyDGX agent

arXiv:2605.08019v1 Announce Type: new Abstract: Humans rapidly learn abstract knowledge when encountering novel environments and flexibly deploy this knowledge to guide efficient and intelligent actio

ReasonSTL: Bridging Natural Language and Signal Temporal Logic via Tool-Augmented Process-Rewarded Learning

Model ReleasesDGX agent

arXiv:2605.06483v2 Announce Type: replace Abstract: Signal Temporal Logic (STL) is an expressive formal language for specifying spatio-temporal requirements over real-valued, real-time signals. It has

Reformulating KV Cache Eviction Problem for Long-Context LLM Inference

Local AiDGX agent

arXiv:2605.07234v1 Announce Type: cross Abstract: Large language models (LLMs) support long-context inference but suffer from substantial memory and runtime overhead due to Key-Value (KV) Cache growth

Region4Web: Rethinking Observation Space Granularity for Web Agents

Model ReleasesDGX agent

arXiv:2605.07134v1 Announce Type: cross Abstract: Web agents perceive web pages through an observation space, yet its granularity has remained an underexamined design choice. Existing work treats obse

Regulating Branch Parallelism in LLM Serving

ResearchDGX agent

arXiv:2605.06914v1 Announce Type: cross Abstract: Recent methods expose intra-request parallelism in LLM outputs, allowing independent branches to decode concurrently. Existing serving systems execute

RELO: Reinforcement Learning to Localize for Visual Object Tracking

SafetyDGX agent

arXiv:2605.07379v1 Announce Type: cross Abstract: Conventional visual object trackers localize targets using handcrafted spatial priors, often in the form of heatmaps. Such priors provide only surroga

Render, Don't Decode: Weight-Space World Models with Latent Structural Disentanglement

HardwareDGX agent

arXiv:2605.06298v2 Announce Type: replace-cross Abstract: Training world models on vast quantities of unlabelled videos is a critical step toward fully autonomous intelligence. However, the prevailing

Rep2Text: Decoding Full Text from a Single LLM Token Representation

Model ReleasesDGX agent

arXiv:2511.06571v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have achieved remarkable progress across diverse tasks, yet their internal mechanisms remain largely opaque. In t

Repeated Deceptive Path Planning against Learnable Observer

SafetyDGX agent

arXiv:2605.07174v1 Announce Type: new Abstract: We study the problem of deceptive path planning (DPP), where an agent aims to conceal its true destination from external observers. While existing work

Replicating Human Motivated Reasoning Studies with LLMs

ResearchDGX agent

arXiv:2601.16130v2 Announce Type: replace-cross Abstract: Motivated reasoning - the idea that individuals processing information may be motivated to either arrive at accurate beliefs or arrive at desi

Resource-Element Energy Difference for Noncoherent Over-the-Air Federated Learning

SafetyDGX agent

arXiv:2605.07263v1 Announce Type: cross Abstract: Over-the-air federated learning (OTA-FL) reduces uplink latency by exploiting waveform superposition, but conventional analog aggregation schemes typi

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding

SafetyDGX agent

arXiv:2605.07575v1 Announce Type: cross Abstract: Proactive streaming video understanding requires Video-LLMs to decide when to respond as a video unfolds, a task where existing methods often fall sho

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective

SafetyDGX agent

arXiv:2605.07331v1 Announce Type: cross Abstract: Reinforcement learning, including reinforcement learning with verifiable rewards (RLVR), has emerged as a powerful approach for LLM post-training. Cen

Retina-RAG: Retrieval-Augmented Vision-Language Modeling for Joint Retinal Diagnosis and Clinical Report Generation

Model ReleasesDGX agent

arXiv:2605.06173v2 Announce Type: replace-cross Abstract: Diabetic Retinopathy (DR) is a leading cause of preventable blindness among working-age adults worldwide, yet most automated screening systems

Revisiting Adam for Streaming Reinforcement Learning

ResearchDGX agent

arXiv:2605.06764v1 Announce Type: cross Abstract: Learning from a sequence of interactions, as soon as observations are perceived and acted upon, without explicitly storing them, holds the promise of

Revisiting Transformer Layer Parameterization Through Causal Energy Minimization

Model ReleasesDGX agent

arXiv:2605.07588v1 Announce Type: cross Abstract: Transformer blocks typically combine multi-head attention (MHA) for token mixing with gated MLPs for token-wise feature transformation, yet many choic

RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation

Model ReleasesDGX agent

arXiv:2605.07129v1 Announce Type: cross Abstract: Large Language Models (LLMs) have emerged as a promising paradigm for next-generation recommender systems, offering strong semantic understanding and

Rubric-based On-policy Distillation

SafetyDGX agent

arXiv:2605.07396v1 Announce Type: cross Abstract: On-policy distillation (OPD) is a powerful paradigm for model alignment, yet its reliance on teacher logits restricts its application to white-box sce

Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning

Model ReleasesDGX agent

arXiv:2605.08061v1 Announce Type: new Abstract: We argue that decomposing reward into weighted, verifiable criteria and using an LLM judge to score them provides a partial-credit optimization signal:

RuleSafe-VL: Evaluating Rule-Conditioned Decision Reasoning in Vision-Language Content Moderation

Model ReleasesDGX agent

arXiv:2605.07760v1 Announce Type: new Abstract: Platform content moderation applies explicit policy rules and context-dependent conditions to decide whether user content is allowed, restricted, or rem

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence

SafetyDGX agent

arXiv:2605.06230v2 Announce Type: replace Abstract: As large models evolve from conversational assistants into autonomous agents, challenges increasingly arise from long-horizon decision making, tool

Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents

Model ReleasesDGX agent

arXiv:2605.07630v1 Announce Type: cross Abstract: When a phone-use agent avoids harm, does that show safety, or simply inability to act? Existing evaluations often cannot tell. A harmful outcome may b

Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks

Model ReleasesDGX agent

arXiv:2605.05995v2 Announce Type: replace-cross Abstract: The safety alignment of Large Language Models (LLMs) remains vulnerable to Harmful Fine-tuning (HFT). While existing defenses impose constrain

Saliency-Aware Regularized Quantization Calibration for Large Language Models

ResearchDGX agent

arXiv:2605.05693v2 Announce Type: replace Abstract: Post-training quantization (PTQ) is an effective approach for deploying large language models (LLMs) under memory and latency constraints. Most exis

SAM 3D Animal: Promptable Animal 3D Reconstruction from Images in the Wild

ResearchDGX agent

arXiv:2605.07604v1 Announce Type: cross Abstract: 3D animal reconstruction in the wild remains challenging due to large species variation, frequent occlusions, and the prevalence of multi-animal scene

Same Brain, Different Prediction: How Preprocessing Choices Undermine EEG Decoding Reliability

TutorialsDGX agent

arXiv:2605.07212v1 Announce Type: cross Abstract: Electroencephalography (EEG) is a cornerstone of brain-computer interfaces and clinical neuroscience, yet deep learning models are typically trained a

Same Signal, Opposite Meaning: Direction-Informed Adaptive Learning for LLM Agents

SafetyDGX agent

arXiv:2605.06908v1 Announce Type: cross Abstract: Adaptive test-time compute for LLM agents aims to invoke extra computation only when it improves performance. Existing methods typically use confidenc

SB-TRPO: Towards Safe Reinforcement Learning with Hard Constraints

SafetyDGX agent

arXiv:2512.23770v3 Announce Type: replace-cross Abstract: In safety-critical domains, reinforcement learning (RL) agents must often satisfy strict, zero-cost safety constraints while accomplishing tas

Scalable Option Learning in High-Throughput Environments

ResearchDGX agent

arXiv:2509.00338v3 Announce Type: replace-cross Abstract: Hierarchical reinforcement learning (RL) has the potential to enable effective decision-making over long timescales. Existing approaches, whil

SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation

Model ReleasesDGX agent

arXiv:2605.08043v1 Announce Type: cross Abstract: While text-to-image models have made strong progress in visual fidelity, faithfully realizing complex visual intents remains challenging because many

ScrapeGraphAI-100k: Dataset for Schema-Constrained LLM Generation

Model ReleasesDGX agent

arXiv:2602.15189v2 Announce Type: replace-cross Abstract: Producing output that conforms to a specified JSON schema underlies tool use, structured extraction, and knowledge base construction in modern

Script Sensitivity: Benchmarking Language Models on Unicode, Romanized and Mixed-Script Sinhala

ApplicationsDGX agent

arXiv:2601.14958v3 Announce Type: replace-cross Abstract: The performance of Language Models (LMs) on low-resource, morphologically rich languages like Sinhala remains largely unexplored, particularly

Searching for Privacy Risks in LLM Agents via Simulation

AgentsDGX agent

arXiv:2508.10880v3 Announce Type: replace-cross Abstract: The widespread deployment of LLM-based agents is likely to introduce a critical privacy threat: malicious agents that proactively engage other

Self-Programmed Execution for Language-Model Agents

SafetyDGX agent

arXiv:2605.06898v1 Announce Type: new Abstract: At the heart of existing language model agents is a fixed orchestrator program responsible for the state transition between consecutive turns. This pape

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding

HardwareDGX agent

arXiv:2605.07897v1 Announce Type: cross Abstract: Online streaming video understanding requires models to process continuous visual inputs and respond to user queries in real time, where the unbounded

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression

Model ReleasesDGX agent

arXiv:2502.01941v3 Announce Type: replace-cross Abstract: While Key-Value (KV) cache compression is essential for efficient LLM inference, current evaluations disproportionately focus on sparse retrie

Shaping the Future of Mathematics in the Age of AI

ResearchDGX agent

arXiv:2603.24914v2 Announce Type: replace-cross Abstract: Artificial intelligence is transforming mathematics at a speed and scale that demand active engagement from the mathematical community. We exa

SHRED: Retain-Set-Free Unlearning via Self-Distillation with Logit Demotion

ResearchDGX agent

arXiv:2605.07482v1 Announce Type: cross Abstract: Machine unlearning for large language models (LLMs) aims to selectively remove memorized content such as private data, copyrighted text, or hazardous

Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair

Local AiDGX agent

arXiv:2605.07276v1 Announce Type: new Abstract: Code-agent RL often receives weak feedback: rollout-time signals are reliable and executable, but capture only necessary or surface conditions for task

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning

SafetyDGX agent

arXiv:2605.06130v2 Announce Type: replace Abstract: A persistent skill library allows language model agents to reuse successful strategies across tasks. Maintaining such a library requires three coupl

Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models

ResearchDGX agent

arXiv:2509.25584v2 Announce Type: replace Abstract: Vision-language models achieve incredible performance across a wide range of tasks, but their large size makes inference costly. Recent work has sho

SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks

Model ReleasesDGX agent

arXiv:2603.24755v2 Announce Type: replace-cross Abstract: Software development is iterative, yet agentic coding benchmarks hide design issues through their single-shot setup. Recent iterative benchmar

SOD: Step-wise On-policy Distillation for Small Language Model Agents

SafetyDGX agent

arXiv:2605.07725v1 Announce Type: cross Abstract: Tool-integrated reasoning (TIR) is difficult to scale to small language models due to instability in long-horizon tool interactions and limited model

SOM: Structured Opponent Modeling for LLM-based Agents via Structural Causal Model

AgentsDGX agent

arXiv:2605.07301v1 Announce Type: new Abstract: Accurately predicting opponents' behavior from interactions is a fundamental capability for large language model (LLM)-based agents in multi-agent and g

← Previous
1…282283284285286…358
Next →