AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
Agents

Practical Online KV Cache Compaction for LLM Agents: An Empirical Study

DGX agent

arXiv:2608.00902v1 Announce Type: new Abstract: LLM agents accumulate long trajectories of reasoning steps, tool calls, and environment feedback, making the KV cache a major inference bottleneck. KV c

agentsarxiv-cs-cl
4 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise

DGX agent

arXiv:2608.02372v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in task-oriented dialogue systems that support multi-step decision-making in high-stakes domains

model-releasesarxiv-cs-cl
4 Aug 2026
Research

Predicting Startup Exit from Textual Descriptors - A Computational Linguistics Framework

DGX agent

arXiv:2608.00045v1 Announce Type: new Abstract: This study shows that textual descriptors alone can predict early-stage startup success, defined as Exit, without relying on contextual, financial, or h

researcharxiv-cs-cl
4 Aug 2026
Agents

Progressive Agent Skill Generation via Reinforcement Learning

DGX agent

arXiv:2608.01678v1 Announce Type: cross Abstract: Existing skill generation methods largely rely on heuristics or pipeline-style consolidation, which must be specially designed for different evidence

agentsarxiv-cs-cl
4 Aug 2026
Model Releases

Prompt-Induced Waste in Large Reasoning Models: A Preregistered Two-Harness Benchmark of Coding Agents

DGX agent

arXiv:2608.01347v1 Announce Type: new Abstract: Large reasoning models used as coding agents incur costs from deliberation, tool calls, and repeated agent turns, yet the causal effect of prompt wordin

model-releasesarxiv-cs-cl
4 Aug 2026
Research

Pruned BPE: Post-training Visibility Pruning and Token Reallocation for Byte Pair Encoding

DGX agent

arXiv:2608.00837v1 Announce Type: new Abstract: Byte Pair Encoding (BPE) is widely used for subword tokenization, but standard BPE exposes every learned merge token to the downstream model, including

researcharxiv-cs-cl
4 Aug 2026
Research

QR-Erase: Efficient Subspace-Based Machine Unlearning with Layer Localization

DGX agent

arXiv:2608.01422v1 Announce Type: new Abstract: Machine unlearning seeks to remove targeted information from trained models without requiring costly retraining. Existing optimization-based methods oft

researcharxiv-cs-cl
4 Aug 2026
Research

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics

DGX agent

arXiv:2608.01522v1 Announce Type: cross Abstract: Teaching a language model a skill it has not mastered is obstructed by three recurring difficulties: training data is scarce, ground-truth reasoning t

researcharxiv-cs-cl
4 Aug 2026
Safety

Quick on the Uptake: Eliciting Implicit Intents from Human Demonstrations for Personalized Mobile-Use Agents

DGX agent

arXiv:2508.08645v3 Announce Type: replace Abstract: As multimodal large language models advance rapidly, the automation of mobile tasks has become increasingly feasible through the use of mobile-use a

safetyarxiv-cs-cl
4 Aug 2026
Model Releases

Qwen-CUA: Native Computer Use for (almost) Everything

DGX agent

arXiv:2608.02352v1 Announce Type: cross Abstract: Native computer use offers a general interface for agents to operate almost any software available to people, but requires long-horizon state tracking

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

RADAR: Rubric-Aware Dependency and Redundancy Analysis for LLM-as-Judge Evaluation

DGX agent

arXiv:2608.01810v1 Announce Type: new Abstract: Rubric-based LLM-as-judge pipelines often assume that evaluation criteria provide independent signals. In practice, however, criteria can be behaviorall

model-releasesarxiv-cs-cl
4 Aug 2026
Applications

RAG Strategies for Natural Language-Based SQL Query and REST API Call Generation

DGX agent

arXiv:2602.07086v2 Announce Type: replace-cross Abstract: Enterprise software systems commonly expose business functionality through both relational databases and REST APIs. Accessing these interfaces

applicationsarxiv-cs-cl
4 Aug 2026
Research

RAGOCR: Optical Compression of Retrieval-Augmented Text via Visual Representation

DGX agent

arXiv:2608.00765v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has become essential for knowledge-intensive question answering, yet scaling RAG pipelines remains challenging due

researcharxiv-cs-cl
4 Aug 2026
Model Releases

REFLEX: Rethinking MoE Inference as Refinement-Aware Compute Allocation in Diffusion Language Models

DGX agent

arXiv:2608.01784v1 Announce Type: cross Abstract: Mixture-of-experts (MoE) models increase parameter capacity by activating only a small subset of experts for each token. This conditional-computation

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction

DGX agent

arXiv:2608.01247v1 Announce Type: new Abstract: Query-agnostic KV cache eviction compresses a context once and reuses the resulting cache for arbitrary future queries, but performance can collapse und

model-releasesarxiv-cs-cl
4 Aug 2026
Research

Rethinking and formalising the state across languages: a unified computational learning theory account

DGX agent

arXiv:2608.00523v1 Announce Type: new Abstract: The linguistic notion of state has traditionally been restricted to the construct (annexation) state of Afroasiatic languages and treated as a language-

researcharxiv-cs-cl
4 Aug 2026
Research

Retrieval Augmented Biomedical Question Answering with Weak Question Recovery and Neural Reranking for BioASQ Task 14b

DGX agent

arXiv:2608.01468v1 Announce Type: new Abstract: This work presents DS@GT ARC BioASQ team's work for a biomedical question answering pipeline, integrating multi-source query expansion, neural reranking

researcharxiv-cs-cl
4 Aug 2026
Research

Revisiting Generalization Across Difficulty Levels: It's Not So Easy

DGX agent

arXiv:2511.21692v2 Announce Type: replace Abstract: We investigate how well large language models (LLMs) generalize across different task difficulties, a key question for effective data curation and e

researcharxiv-cs-cl
4 Aug 2026
Local Ai

RH-RAG: Trustworthy Long-Form Generation for Privacy-Constrained Settings

DGX agent

arXiv:2608.01311v1 Announce Type: new Abstract: Generating long-form content from extensive internal reports remains challenging for organizations operating under strict privacy and security constrain

local-aiarxiv-cs-cl
4 Aug 2026
Research

Rhythm of the Deep: Two-Tier Combinatorial Structure in Sperm Whale Codas Revealed by Acoustic Unit Induction

DGX agent

arXiv:2606.16084v2 Announce Type: replace-cross Abstract: Sperm-whale codas are conventionally described as recurring click-count and timing patterns. We show instead that their waveforms contain a tw

researcharxiv-cs-cl
4 Aug 2026
Model Releases

Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks

DGX agent

arXiv:2608.02442v1 Announce Type: cross Abstract: Scientific reasoning benchmarks typically evaluate large language models (LLMs) using final-answer accuracy. However, a correct answer does not necess

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

RING: Retrieval-Internalized Generation for Continual Large-Scale Knowledge Injection

DGX agent

arXiv:2608.01630v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) improves factuality but adds latency and engineering overhead at serving time. We propose RING (Retrieval-Internali

model-releasesarxiv-cs-cl
4 Aug 2026
Safety

RMSWeb: Reflection, Failure-Mode Mining, and Salvage-DS for Web Agent Reinforcement Learning

DGX agent

arXiv:2608.00335v1 Announce Type: cross Abstract: Compact web agents can reduce deployment cost, but training them poses challenges in both data collection and post-SFT reinforcement learning (RL). Su

safetyarxiv-cs-cl
4 Aug 2026
Model Releases

Role Steering of Language Models for Social Simulations

DGX agent

arXiv:2608.00023v1 Announce Type: new Abstract: Social simulations built from language-model agents need role-conditioned behavior that can be checked before agents are placed into a simulated populat

model-releasesarxiv-cs-cl
4 Aug 2026
Safety

Romanized Arabic Across Dialects: Views, Usage Patterns, and Linguistic Variation

DGX agent

arXiv:2608.02555v1 Announce Type: new Abstract: Arabizi refers to Arabic written in Latin script. Although previous studies have shown that the prevalence and usage of Arabizi vary by factors such as

safetyarxiv-cs-cl
4 Aug 2026
Agents

RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States

DGX agent

arXiv:2608.02508v1 Announce Type: cross Abstract: Learning-based memory systems for self-evolving LLM agents face two tightly coupled challenges. First, trajectory-indexed utilities grow with the inte

agentsarxiv-cs-cl
4 Aug 2026
Agents

RubricReviewer: From Direct Critique to Objective and Comprehensive Rubric-Driven Peer Review

DGX agent

arXiv:2608.00005v1 Announce Type: new Abstract: Peer review at major venues is under unprecedented submission pressure, motivating the use of large language models (LLMs) as review assistants. Existin

agentsarxiv-cs-cl
4 Aug 2026
Model Releases

S^4R: Selective Sampling, Subspaces, and Sparse Reconstruction for Compressed Long-Context KV Caching

DGX agent

arXiv:2608.00528v1 Announce Type: new Abstract: The growth of context window lengths in Large Language Models (LLMs) significantly enhances their long-context capabilities but incurs prohibitive memor

model-releasesarxiv-cs-cl
4 Aug 2026
Research

Same violence, different answer: how AI responds to coercive control against women across languages

DGX agent

arXiv:2608.01436v1 Announce Type: cross Abstract: Women experiencing coercive control, a form of intimate partner violence increasingly conducted through digital devices, are turning to conversational

researcharxiv-cs-cl
4 Aug 2026
Model Releases

SCHEDBench: A Benchmark for Evaluating LLM Constraint Faithfulness in Natural-Language Combinatorial Scheduling

DGX agent

arXiv:2608.00991v1 Announce Type: cross Abstract: This paper introduces SCHEDBench, a natural-language benchmark for evaluating combinatorial scheduling constraint faithfulness under surface-form vari

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step

DGX agent

arXiv:2608.02358v1 Announce Type: new Abstract: To operate robustly in open-world environments, autonomous agents should be able to infer the behavior of unfamiliar systems through interaction alone,

model-releasesarxiv-cs-cl
4 Aug 2026
Research

SeDeM: Selective Decompression of Hidden-State Memories for Long-Context Question Answering

DGX agent

arXiv:2608.00311v1 Announce Type: new Abstract: Long-context inference with large language models (LLMs) is costly: self-attention during prefill scales quadratically with sequence length, and the key

researcharxiv-cs-cl
4 Aug 2026
Research

Select-And-Extract: A Lightweight Plugin for Retrieval-Augmented Generation

DGX agent

arXiv:2608.00658v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) for language model (LM) systems fundamentally has two failure modes: retrieval failure and reading failure. The for

researcharxiv-cs-cl
4 Aug 2026
Safety

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

DGX agent

arXiv:2507.02778v3 Announce Type: replace Abstract: Although large language models (LLMs) have transformed AI, they still make errors and follow unproductive reasoning paths. Self-correction is vital

safetyarxiv-cs-cl
4 Aug 2026
Safety

Self-Improving Large Language Models via Progressive Experience Evolution

DGX agent

arXiv:2608.02139v1 Announce Type: new Abstract: Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism for transformin

safetyarxiv-cs-cl
4 Aug 2026
Safety

Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer

DGX agent

arXiv:2608.01585v1 Announce Type: new Abstract: Language model benchmarking is a difficult task. Outcome reasoning alone does not test the model's conceptualization of language and popular open-source

safetyarxiv-cs-cl
4 Aug 2026
Model Releases

SERL-SQL: Selective Hindsight Distillation for Text-to-SQL Reinforcement Agentic Learning

DGX agent

arXiv:2608.00485v1 Announce Type: new Abstract: Recent Text-to-SQL systems increasingly rely on multi-turn interaction, execution feedback, and reinforcement learning. However, most existing methods u

model-releasesarxiv-cs-cl
4 Aug 2026
Agents

ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors

DGX agent

arXiv:2608.01204v1 Announce Type: new Abstract: Conversational investment advisors influence not only what users know, but also how they make subsequent decisions as market conditions evolve. Existing

agentsarxiv-cs-cl
4 Aug 2026
Agents

SIEVE: Selective Integrity Verification and Escalation for Defending LLM Agents against Indirect Prompt Injection

DGX agent

arXiv:2512.06716v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used as the core of agentic systems due to their strong reasoning, planning, and tool-use capabi

agentsarxiv-cs-cl
4 Aug 2026
Research

Sixteen models, fewer than two voices: measuring ensemble dispersion where no answer is uniquely correct

DGX agent

arXiv:2608.00285v1 Announce Type: new Abstract: Sixteen language models drawn from ten families produced, on average, the semantic diversity of 1.69 distinct formulations of a psychotherapeutic case,

researcharxiv-cs-cl
4 Aug 2026
Model Releases

SLMs as Multi-Agent Routers: A Progressive SFT and Reinforcement Learning Approach

DGX agent

arXiv:2608.00030v1 Announce Type: new Abstract: Specialised retrieval agents typically surface higher quality results than general-purpose search, but selecting the optimal agent for a given query rem

model-releasesarxiv-cs-cl
4 Aug 2026
Safety

Slot2Text: Object-Centric Visual Tokenization for Efficient and Spatially Traceable Surgical MLLMs

DGX agent

arXiv:2608.01473v1 Announce Type: cross Abstract: Multimodal large language models (MLLM) for surgical scene understanding typically inject hundreds of dense visual tokens into a language model, leadi

safetyarxiv-cs-cl
4 Aug 2026
Model Releases

SoM-1K: A Thousand-Problem Benchmark Dataset for Strength of Materials

DGX agent

arXiv:2509.21079v2 Announce Type: replace Abstract: Foundation models have shown remarkable capabilities in various domains, but their performance on complex, multimodal engineering problems remains l

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models

DGX agent

arXiv:2608.01899v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduc

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation

DGX agent

arXiv:2608.01666v1 Announce Type: new Abstract: However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open q

model-releasesarxiv-cs-cl
4 Aug 2026
Agents

SWE-Touch: Benchmarking Coding Agents When Users Touch the Code

DGX agent

arXiv:2608.02499v1 Announce Type: cross Abstract: Real-world software development requires coding agents to operate in shared workspaces where users may inspect and modify code during an ongoing task,

agentsarxiv-cs-cl
4 Aug 2026
Model Releases

TAB-PO: Preference Optimization with a Token-Level Adaptive Barrier for Token-Critical Structured Generation

DGX agent

arXiv:2603.00025v3 Announce Type: replace Abstract: Direct Preference Optimization (DPO) is effective for offline alignment but poorly matched to ontology-driven structured prediction, where preferred

model-releasesarxiv-cs-cl
4 Aug 2026
Local Ai

TELLER: Non-intrusive Cross-Layer Root-Cause Analysis for LLM Inference

DGX agent

arXiv:2608.01975v1 Announce Type: cross Abstract: Large language model (LLM) inference has evolved from an offline workload into a continuously operated software service, yet root-cause analysis remai

local-aiarxiv-cs-cl
4 Aug 2026
← Previous
1…1213141516…160
Next →