AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
16 Apr 2026

Memory Transfer Learning: How Memories are Transferred Across Domains in Coding Agents

ApplicationsDGX agent

arXiv:2604.14004v1 Announce Type: cross Abstract: Memory-based self-evolution has emerged as a promising paradigm for coding agents. However, existing approaches typically restrict memory utilization

Memp: Exploring Agent Procedural Memory

AgentsDGX agent

arXiv:2508.06433v4 Announce Type: replace Abstract: Large Language Models (LLMs) based agents excel at diverse tasks, yet they suffer from brittle procedural memory that is manually engineered or enta

MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments

Model ReleasesDGX agent

arXiv:2604.13418v1 Announce Type: new Abstract: Motivated by the underspecified, multi-hop nature of search queries and the multimodal, heterogeneous, and often conflicting nature of real-world web re


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates

Model ReleasesDGX agent

arXiv:2512.04844v2 Announce Type: replace Abstract: Expanding the linguistic diversity of instruct large language models (LLMs) is crucial for global accessibility but is often hindered by the relianc

MM-Doc-R1: Training Agents for Long Document Visual Question Answering through Multi-turn Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.13579v1 Announce Type: new Abstract: Conventional Retrieval-Augmented Generation (RAG) systems often struggle with complex multi-hop queries over long documents due to their single-pass ret

MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models

Model ReleasesDGX agent

arXiv:2505.07591v2 Announce Type: replace Abstract: Instruction following refers to the ability of large language models (LLMs) to generate outputs that satisfy all specified constraints. Existing res

MUSE: Multi-Domain Chinese User Simulation via Self-Evolving Profiles and Rubric-Guided Alignment

Local AiDGX agent

arXiv:2604.13828v1 Announce Type: new Abstract: User simulators are essential for the scalable training and evaluation of interactive AI systems. However, existing approaches often rely on shallow use

Native Hybrid Attention for Efficient Sequence Modeling

ResearchDGX agent

arXiv:2510.07019v3 Announce Type: replace Abstract: Transformers excel at sequence modeling but face quadratic complexity, while linear attention offers improved efficiency but often compromises recal

Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Models

ResearchDGX agent

arXiv:2601.11340v2 Announce Type: replace Abstract: Chain-of-Thought reasoning has significantly enhanced the problem-solving capabilities of Large Language Models. Unfortunately, current models gener

Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning

SafetyDGX agent

arXiv:2506.08125v3 Announce Type: replace-cross Abstract: Large language models (LLMs) show strong reasoning abilities but often produce unnecessarily long explanations that reduce efficiency. Althoug

OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs

ResearchDGX agent

arXiv:2604.13073v1 Announce Type: new Abstract: Modern multimodal large language models (MLLMs) generate fluent responses from interleaved text, image, audio, and video inputs. However, identifying wh

Parameter Importance is Not Static: Evolving Parameter Isolation for Supervised Fine-Tuning

Model ReleasesDGX agent

arXiv:2604.14010v1 Announce Type: cross Abstract: Supervised Fine-Tuning (SFT) of large language models often suffers from task interference and catastrophic forgetting. Recent approaches alleviate th

ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian

ResearchDGX agent

arXiv:2511.01619v2 Announce Type: replace Abstract: ParlaSpeech is a collection of spoken parliamentary corpora currently spanning four Slavic languages - Croatian, Czech, Polish and Serbian - all tog

Peer-Predictive Self-Training for Language Model Reasoning

Model ReleasesDGX agent

arXiv:2604.13356v1 Announce Type: new Abstract: Mechanisms for continued self-improvement of language models without external supervision remain an open challenge. We propose Peer-Predictive Self-Trai

PersonaVLM: Long-Term Personalized Multimodal LLMs

Model ReleasesDGX agent

arXiv:2604.13074v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) serve as daily assistants for millions. However, their ability to generate responses aligned with individual pr

pi-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data

AgentsDGX agent

arXiv:2604.14054v1 Announce Type: cross Abstract: Deep search agents have emerged as a promising paradigm for addressing complex information-seeking tasks, but their training remains challenging due t

QuantileMark: A Message-Symmetric Multi-bit Watermark for LLMs

ResearchDGX agent

arXiv:2604.13786v1 Announce Type: new Abstract: As large language models become standard backends for content generation, practical provenance increasingly requires multi-bit watermarking. In provider

RadAgents: Multimodal Agentic Reasoning for Chest X-ray Interpretation with Radiologist-like Workflows

AgentsDGX agent

arXiv:2509.20490v4 Announce Type: replace-cross Abstract: Agentic systems offer a potential path to solve complex clinical tasks through collaboration among specialized agents, augmented by tool use a

RAG or Learning? Understanding the Limits of LLM Adaptation under Continuous Knowledge Drift in the Real World

Model ReleasesDGX agent

arXiv:2604.05096v2 Announce Type: replace Abstract: Large language models (LLMs) acquire most of their knowledge during pretraining, which ties them to a fixed snapshot of the world and makes adaptati

Red Skills or Blue Skills? A Dive Into Skills Published on ClawHub

Model ReleasesDGX agent

arXiv:2604.13064v1 Announce Type: new Abstract: Skill ecosystems have emerged as an increasingly important layer in Large Language Model (LLM) agent systems, enabling reusable task packaging, public d

Reducing Hallucinations in LLMs via Factuality-Aware Preference Learning

SafetyDGX agent

arXiv:2601.03027v3 Announce Type: replace Abstract: Preference alignment methods such as RLHF and Direct Preference Optimization (DPO) improve instruction following, but they can also reinforce halluc

Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution

AgentsDGX agent

arXiv:2512.10696v2 Announce Type: replace-cross Abstract: Procedural memory enables large language model (LLM) agents to internalize 'how-to' knowledge, theoretically reducing redundant trial-and-erro

Reward Design for Physical Reasoning in Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.13993v1 Announce Type: cross Abstract: Physical reasoning over visual inputs demands tight integration of visual perception, domain knowledge, and multi-step symbolic inference. Yet even st

Rhetorical Questions in LLM Representations: A Linear Probing Study

Local AiDGX agent

arXiv:2604.14128v1 Announce Type: new Abstract: Rhetorical questions are asked not to seek information but to persuade or signal stance. How large language models internally represent them remains unc

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization

SafetyDGX agent

arXiv:2508.00222v5 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Reward (RLVR) has significantly advanced the complex reasoning abilities of Large Language Models (LLMs

Robust Reward Modeling for Large Language Models via Causal Decomposition

Model ReleasesDGX agent

arXiv:2604.13833v1 Announce Type: new Abstract: Reward models are central to aligning large language models, yet they often overfit to spurious cues such as response length and overly agreeable tone.

Saber: An Efficient Sampling with Adaptive Acceleration and Backtracking Enhanced Remasking for Diffusion Language Model

ResearchDGX agent

arXiv:2510.18165v2 Announce Type: replace-cross Abstract: Diffusion language models (DLMs) are emerging as a powerful and promising alternative to the dominant autoregressive paradigm, offering inhere

Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight Models

Model ReleasesDGX agent

arXiv:2510.14232v2 Announce Type: replace-cross Abstract: Competitive programming has become a rigorous benchmark for evaluating the reasoning and problem-solving capabilities of large language models

Social media polarization during conflict: Insights from an ideological stance dataset on Israel-Palestine Reddit comments

ResearchDGX agent

arXiv:2502.00414v2 Announce Type: replace Abstract: In politically sensitive scenarios like wars, social media serves as a platform for polarized discourse and expressions of strong ideological stance

Sparse or Dense? A Mechanistic Estimation of Computation Density in Transformer-based LLMs

ResearchDGX agent

arXiv:2601.22795v2 Announce Type: replace Abstract: Transformer-based large language models (LLMs) are comprised of billions of parameters arranged in deep and wide computational graphs. Several studi

SpatialEvo: Self-Evolving Spatial Intelligence via Deterministic Geometric Environments

Model ReleasesDGX agent

arXiv:2604.14144v1 Announce Type: cross Abstract: Spatial reasoning over three-dimensional scenes is a core capability for embodied intelligence, yet continuous model improvement remains bottlenecked

SPG: Sandwiched Policy Gradient for Masked Diffusion Language Models

SafetyDGX agent

arXiv:2510.09541v3 Announce Type: replace Abstract: Diffusion large language models (dLLMs) are emerging as an efficient alternative to autoregressive models due to their ability to decode multiple to

Syn-TurnTurk: A Synthetic Dataset for Turn-Taking Prediction in Turkish Dialogues

Model ReleasesDGX agent

arXiv:2604.13620v1 Announce Type: new Abstract: Managing natural dialogue timing is a significant challenge for voice-based chatbots. Most current systems usually rely on simple silence detection, whi

Synthesizing Instruction-Tuning Datasets with Contrastive Decoding

Model ReleasesDGX agent

arXiv:2604.13538v1 Announce Type: new Abstract: Using responses generated by high-performing large language models (LLMs) for instruction tuning has become a widely adopted approach. However, the exis

Text-as-Signal: Quantitative Semantic Scoring with Embeddings, Logprobs, and Noise Reduction

Model ReleasesDGX agent

arXiv:2604.13056v1 Announce Type: new Abstract: This paper presents a practical pipeline for turning text corpora into quantitative semantic signals. Each news item is represented as a full-document e

The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious

Model ReleasesDGX agent

arXiv:2604.13051v1 Announce Type: new Abstract: There is debate about whether LLMs can be conscious. We investigate a distinct question: if a model claims to be conscious, how does this affect its dow

TLoRA+: A Low-Rank Parameter-Efficient Fine-Tuning Method for Large Language Models

Model ReleasesDGX agent

arXiv:2604.13368v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) aims to adapt pre-trained models to specific tasks using relatively small and domain-specific datasets. Among P

ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded Execution

Model ReleasesDGX agent

arXiv:2604.13787v1 Announce Type: new Abstract: Large Language Models (LLMs) enhance their problem-solving capability by utilizing external tools. However, in open-world scenarios with massive and evo

ToolSpec: Accelerating Tool Calling via Schema-Aware and Retrieval-Augmented Speculative Decoding

AgentsDGX agent

arXiv:2604.13519v1 Announce Type: new Abstract: Tool calling has greatly expanded the practical utility of large language models (LLMs) by enabling them to interact with external applications. As LLM

Training-Free Test-Time Contrastive Learning for Large Language Models

AgentsDGX agent

arXiv:2604.13552v1 Announce Type: new Abstract: Large language models (LLMs) demonstrate strong reasoning capabilities, but their performance often degrades under distribution shift. Existing test-tim

TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration

Model ReleasesDGX agent

arXiv:2604.14116v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have empowered AI research agents to perform isolated scientific tasks, automating complex, real-world workflows, s

TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning Tasks

SafetyDGX agent

arXiv:2601.10245v2 Announce Type: replace-cross Abstract: Multi-step reasoning tasks like mathematical problem solving are vulnerable to cascading failures, where a single incorrect step leads to comp

Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations

ResearchDGX agent

arXiv:2601.07422v2 Announce Type: replace Abstract: Despite their impressive capabilities, large language models (LLMs) frequently generate hallucinations. Previous work shows that their internal stat

Two-Stage Regularization-Based Structured Pruning for LLMs

Model ReleasesDGX agent

arXiv:2505.18232v3 Announce Type: replace-cross Abstract: The deployment of large language models (LLMs) is largely hindered by their large number of parameters. Structural pruning has emerged as a pr

UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding

Local AiDGX agent

arXiv:2604.14113v1 Announce Type: cross Abstract: GUI grounding, which localizes interface elements from screenshots given natural language queries, remains challenging for small icons and dense layou

Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization

Local AiDGX agent

arXiv:2604.13197v1 Announce Type: new Abstract: Process reward models (PRMs) provide fine-grained reward signals along the reasoning process, but training reliable PRMs often requires step annotations

Using reasoning LLMs to extract SDOH events from clinical notes

ResearchDGX agent

arXiv:2604.13502v1 Announce Type: new Abstract: Social Determinants of Health (SDOH) refer to environmental, behavioral, and social conditions that influence how individuals live, work, and age. SDOH

ValueGround: Evaluating Culture-Conditioned Visual Value Grounding in MLLMs

Model ReleasesDGX agent

arXiv:2604.06484v2 Announce Type: replace Abstract: Cultural values are expressed not only through language but also through visual scenes and everyday social practices. Yet existing evaluations of cu

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors

ResearchDGX agent

arXiv:2604.02486v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have achieved impressive performance across a wide range of multimodal tasks. However, they often fail on tasks

WebXSkill: Skill Learning for Autonomous Web Agents

AgentsDGX agent

arXiv:2604.13318v1 Announce Type: cross Abstract: Autonomous web agents powered by large language models (LLMs) have shown promise in completing complex browser tasks, yet they still struggle with lon

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning?

Model ReleasesDGX agent

arXiv:2503.23137v2 Announce Type: replace-cross Abstract: Understanding humor-particularly when it involves complex, contradictory narratives that require comparative reasoning-remains a significant c

Who Gets Flagged? The Pluralistic Evaluation Gap in AI Content Watermarking

SafetyDGX agent

arXiv:2604.13776v1 Announce Type: cross Abstract: Watermarking is becoming the default mechanism for AI content authentication, with governance policies and frameworks referencing it as infrastructure

Working Notes on Late Interaction Dynamics: Analyzing Targeted Behaviors of Late Interaction Models

Model ReleasesDGX agent

arXiv:2603.26259v2 Announce Type: replace-cross Abstract: While Late Interaction models exhibit strong retrieval performance, many of their underlying dynamics remain understudied, potentially hiding

WorkRB: A Community-Driven Evaluation Framework for AI in the Work Domain

Model ReleasesDGX agent

arXiv:2604.13055v1 Announce Type: new Abstract: Today's evolving labor markets rely increasingly on recommender systems for hiring, talent management, and workforce analytics, with natural language pr

YOCO++: Enhancing YOCO with KV Residual Connections for Efficient LLM Inference

ResearchDGX agent

arXiv:2604.13556v1 Announce Type: new Abstract: Cross-layer key-value (KV) compression has been found to be effective in efficient inference of large language models (LLMs). Although they reduce the m

15 Apr 2026

AAPO: Enhancing the Reasoning Capabilities of LLMs with Advantage Margin

SafetyDGX agent

arXiv:2505.14264v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as an effective approach for enhancing the reasoning capabilities of large language models (LLMs), esp

ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning

Model ReleasesDGX agent

arXiv:2602.11236v2 Announce Type: replace-cross Abstract: Building general-purpose embodied agents across diverse hardware remains a central challenge in robotics, often framed as the ''one-brain, man

Accelerating Speculative Decoding with Block Diffusion Draft Trees

ResearchDGX agent

arXiv:2604.12989v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive language models by using a lightweight drafter to propose multiple future tokens, which the target model

Adaptive Test-Time Scaling for Zero-Shot Respiratory Audio Classification

ResearchDGX agent

arXiv:2604.12647v1 Announce Type: cross Abstract: Automated respiratory audio analysis promises scalable, non-invasive disease screening, yet progress is limited by scarce labeled data and costly expe

Advancing Multi-Agent RAG Systems with Minimalist Reinforcement Learning

SafetyDGX agent

arXiv:2505.17086v4 Announce Type: replace Abstract: Large Language Models (LLMs) equipped with modern Retrieval-Augmented Generation (RAG) systems often employ multi-turn interaction pipelines to inte

← Previous
1…117118119120121…128
Next →