AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Research

LaTER: Efficient Test-Time Reasoning via Latent Exploration and Explicit Verification

DGX agent

arXiv:2605.07315v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning improves large language models (LLMs) on difficult tasks, but it also makes inference expensive because every intermedi

researcharxiv-cs-cl
11 May 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Learning Agent Routing From Early Experience

DGX agent

arXiv:2605.07180v1 Announce Type: new Abstract: LLM agents achieve strong performance on complex reasoning tasks but incur high latency and compute cost. In practice, many queries fall within the capa

model-releasesarxiv-cs-cl
11 May 2026
Tutorials

LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction

DGX agent

arXiv:2605.06676v1 Announce Type: cross Abstract: Long-context inference in Large Language Models (LLMs) is bottlenecked by the linear growth of Key-Value (KV) cache memory. Existing KV cache compress

tutorialsarxiv-cs-cl
11 May 2026
Agents

LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling

DGX agent

arXiv:2605.08083v1 Announce Type: new Abstract: Test-time scaling (TTS) has become an effective approach for improving large language model performance by allocating additional computation during infe

agentsarxiv-cs-cl
11 May 2026
Model Releases

Mean-Pooled Cosine Similarity is Not Length-Invariant: Theory and Cross-Domain Evidence for a Length-Invariant Alternative

DGX agent

arXiv:2605.07345v1 Announce Type: new Abstract: Mean-pooled cosine similarity is the default metric for comparing neural representations across languages, modalities, and tasks. We establish that this

model-releasesarxiv-cs-cl
11 May 2026
Research

Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors

DGX agent

arXiv:2605.07847v1 Announce Type: new Abstract: As user simulators are increasingly used for interactive training and evaluation of AI assistants, it is essential that they represent the diverse behav

researcharxiv-cs-cl
11 May 2026
Research

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM

DGX agent

arXiv:2605.05927v2 Announce Type: replace Abstract: Speech large language models (SLMs) are typically built from text large language model (TLM) checkpoints, yet they still suffer from a substantial m

researcharxiv-cs-cl
11 May 2026
Model Releases

MIPIAD: Multilingual Indirect Prompt Injection Attack Defense with Qwen -- TF-IDF Hybrid and Meta-Ensemble Learning

DGX agent

arXiv:2605.07269v1 Announce Type: new Abstract: Indirect prompt injection remains a persistent weakness in retrieval-augmented and tool-using LLM systems, and the problem becomes harder to characteris

model-releasesarxiv-cs-cl
11 May 2026
Research

Multi-Dimensional Evaluation of LLMs for Grammatical Error Correction

DGX agent

arXiv:2605.07635v1 Announce Type: new Abstract: Automated assistants for Grammatical Error Correction are now embedded in educational platforms serving millions of learners, yet three critical gaps re

researcharxiv-cs-cl
11 May 2026
Model Releases

MultiSoc-4D: A Benchmark for Diagnosing Instruction-Induced Label Collapse in Closed-Set LLM Annotation of Bengali Social Media

DGX agent

arXiv:2605.06940v1 Announce Type: new Abstract: Annotation automation via Large Language Models (LLMs) is the core approach for scaling NLP datasets; however, LLM behavior with respect to closed-set i

model-releasesarxiv-cs-cl
11 May 2026
Model Releases

NCL-UoR at SemEval-2026 Task 5: Embedding-Based Methods, Fine-Tuning, and LLMs for Word Sense Plausibility Rating

DGX agent

arXiv:2603.08256v2 Announce Type: replace Abstract: Word sense plausibility rating requires predicting the human-perceived plausibility of a given word sense on a 1-5 scale in the context of short nar

model-releasesarxiv-cs-cl
11 May 2026
Model Releases

Neural Neural Scaling Laws

DGX agent

arXiv:2601.19831v2 Announce Type: replace-cross Abstract: Neural scaling laws predict how language model performance improves with increased training inputs. While aggregate metrics like validation lo

model-releasesarxiv-cs-cl
11 May 2026
Tutorials

Not All Tokens Learn Alike: Attention Entropy Reveals Heterogeneous Signals in RL Reasoning

DGX agent

arXiv:2605.07660v1 Announce Type: new Abstract: Reinforcement-learning-based post-training has become a key approach for improving the reasoning ability of large language models, but its token-level l

tutorialsarxiv-cs-cl
11 May 2026
Model Releases

NSMQ Riddles: A Benchmark of Scientific and Mathematical Riddles for Quizzing Large Language Models

DGX agent

arXiv:2605.07051v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown good performance on various science educational benchmarks, demonstrating their potential for use in science and

model-releasesarxiv-cs-cl
11 May 2026
Research

On the Complexity of the Matching Problem of Regular Expressions with Backreferences

DGX agent

arXiv:2605.07289v1 Announce Type: cross Abstract: ReDoS is a well-known type of algorithmic complexity attack, where an adversary supplies maliciously crafted strings to a regular expression matching

researcharxiv-cs-cl
11 May 2026
Model Releases

OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling

DGX agent

arXiv:2605.07815v1 Announce Type: cross Abstract: Muon improves neural-network training by orthogonalizing matrix-valued updates, but it leaves each layer's update magnitude controlled mostly by a glo

model-releasesarxiv-cs-cl
11 May 2026
Research

Overview of the TREC 2025 RAGTIME Track

DGX agent

arXiv:2602.10024v2 Announce Type: replace-cross Abstract: The principal goal of the RAG TREC Instrument for Multilingual Evaluation (RAGTIME) track at TREC is to study report generation from multiling

researcharxiv-cs-cl
11 May 2026
Safety

PaT: Planning-after-Trial for Efficient Test-Time Code Generation

DGX agent

arXiv:2605.07248v1 Announce Type: new Abstract: Beyond training-time optimization, scaling test-time computation has emerged as a key paradigm to extend the reasoning capabilities of Large Language Mo

safetyarxiv-cs-cl
11 May 2026
Research

PolySQL: Scaling Text-to-SQL Evaluation Across SQL Dialects via Automated Backend Isomorphism

DGX agent

arXiv:2605.07796v1 Announce Type: new Abstract: SQL dialects vary in syntax, types, and functions across database engines. Text-to-SQL benchmarks, however, predominantly support only SQLite. This crea

researcharxiv-cs-cl
11 May 2026
Research

ProtSent: Protein Sentence Transformers

DGX agent

arXiv:2605.06830v1 Announce Type: cross Abstract: Protein language models (pLMs) produce per-residue representations that capture evolutionary and structural information, yet their mean-pooled sequenc

researcharxiv-cs-cl
11 May 2026
Hardware

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory

DGX agent

arXiv:2605.06675v1 Announce Type: cross Abstract: Large language models cache all previously computed key-value (KV) pairs during generation, and this KV cache grows linearly with sequence length, mak

hardwarearxiv-cs-cl
11 May 2026
Safety

Reflections and New Directions for Human-Centered Large Language Models

DGX agent

arXiv:2605.06901v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly shaping the private and professional lives of users, with numerous applications in business, education, fi

safetyarxiv-cs-cl
11 May 2026
Research

Reliable Chain-of-Thought via Prefix Consistency

DGX agent

arXiv:2605.07654v1 Announce Type: cross Abstract: Large Language Models often improve accuracy on reasoning tasks by sampling multiple Chain-of-Thought (CoT) traces and aggregating them with majority

researcharxiv-cs-cl
11 May 2026
Model Releases

ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards

DGX agent

arXiv:2510.00568v3 Announce Type: replace Abstract: Search agents powered by Large Language Models (LLMs) have demonstrated significant potential in tackling knowledge-intensive tasks. Reinforcement l

model-releasesarxiv-cs-cl
11 May 2026
Research

Rethinking Dense Sequential Chains: Reasoning Language Models Can Extract Answers from Sparse, Order-Shuffling Chain-of-Thoughts

DGX agent

arXiv:2605.07307v1 Announce Type: new Abstract: Modern reasoning language models generate dense, sequential chain-of-thought traces implicitly assuming that every token contributes and that steps must

researcharxiv-cs-cl
11 May 2026
Research

Rethinking Experience Utilization in Self-Evolving Language Model Agents

DGX agent

arXiv:2605.07164v1 Announce Type: new Abstract: Self-evolving agents improve by accumulating and reusing experience from past interactions. Existing work has largely focused on how experience is const

researcharxiv-cs-cl
11 May 2026
Tutorials

Rethinking State Tracking in Recurrent Models Through Error Control Dynamics

DGX agent

arXiv:2605.07755v1 Announce Type: cross Abstract: The theory of state tracking in recurrent architectures has predominantly focused on expressive capacity: whether a fixed architecture can theoretical

tutorialsarxiv-cs-cl
11 May 2026
Model Releases

Rethinking Weight Tying: Pseudo-Inverse Tying for LM Stable Training and Updates

DGX agent

arXiv:2602.04556v2 Announce Type: replace Abstract: Weight tying is widely used in compact language models to reduce parameters by sharing the token table between the input embedding and the output pr

model-releasesarxiv-cs-cl
11 May 2026
Research

Retrieval Heads are Dynamic

DGX agent

arXiv:2602.11162v2 Announce Type: replace Abstract: Recent studies have identified 'retrieval heads' in Large Language Models (LLMs) responsible for extracting information from input contexts. However

researcharxiv-cs-cl
11 May 2026
Research

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning

DGX agent

arXiv:2605.07106v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have made remarkable progress on vision-language reasoning, yet most methods still compress visual evidence int

researcharxiv-cs-cl
11 May 2026
Model Releases

S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models

DGX agent

arXiv:2503.05085v2 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have fundamentally reshaped speech-to-speech (S2S) systems, enabling increasingly natural spoken int

model-releasesarxiv-cs-cl
11 May 2026
Safety

SAGE: Hierarchical LLM-Based Literary Evaluation through Ontology-Grounded Interpretive Dimensions

DGX agent

arXiv:2605.07102v1 Announce Type: new Abstract: Evaluating literary quality requires assessing interpretive dimensions such as cultural representation, emotional depth, and philosophical sophisticatio

safetyarxiv-cs-cl
11 May 2026
Model Releases

SCENE: Recognizing Social Norms and Sanctioning in Group Chats

DGX agent

arXiv:2605.07823v1 Announce Type: new Abstract: Online group chats are social spaces with implicit behavior patterns that, when broken, are often met with social sanctioning from the group. The abilit

model-releasesarxiv-cs-cl
11 May 2026
Agents

Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability

DGX agent

arXiv:2605.07110v1 Announce Type: new Abstract: Computer-use agents(CUAs)are moving frombounded benchmarks toward real software environments, wherethey operate browsers, desktops, mobile applications,

agentsarxiv-cs-cl
11 May 2026
Research

SEIF: Self-Evolving Reinforcement Learning for Instruction Following

DGX agent

arXiv:2605.07465v1 Announce Type: new Abstract: Instruction following is a fundamental capability of large language models (LLMs), yet continuously improving this capability remains challenging. Exist

researcharxiv-cs-cl
11 May 2026
Research

Self-Consolidating Language Models: Continual Knowledge Incorporation from Context

DGX agent

arXiv:2605.07076v1 Announce Type: new Abstract: Large language models (LLMs) increasingly receive information as streams of passages, conversations, and long-context workflows. While longer context wi

researcharxiv-cs-cl
11 May 2026
Model Releases

SEQUOR: A Multi-Turn Benchmark for Realistic Constraint Following

DGX agent

arXiv:2605.06353v2 Announce Type: replace Abstract: In a conversation, a helpful assistant must reliably follow user directives, even as they refine, modify, or contradict earlier requests. Yet most i

model-releasesarxiv-cs-cl
11 May 2026
Research

Sign-Based Optimizers Are Effective Under Heavy-Tailed Noise

DGX agent

arXiv:2602.07425v2 Announce Type: replace-cross Abstract: While adaptive gradient methods are the workhorse of modern machine learning, sign-based optimization algorithms such as Lion and Muon have re

researcharxiv-cs-cl
11 May 2026
Safety

SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation

DGX agent

arXiv:2605.07711v1 Announce Type: new Abstract: On-policy distillation (OPD) is a standard tool for transferring teacher behavior to a smaller student, but it implicitly assumes that teacher and stude

safetyarxiv-cs-cl
11 May 2026
Model Releases

SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair

DGX agent

arXiv:2605.07001v1 Announce Type: cross Abstract: Architectural code smells erode software maintainability and are costly to repair manually, yet unlike localized bugs, they require cross-module reaso

model-releasesarxiv-cs-cl
11 May 2026
Hardware

Sparser, Faster, Lighter Transformer Language Models

DGX agent

arXiv:2603.23198v2 Announce Type: replace-cross Abstract: Scaling autoregressive large language models (LLMs) has driven unprecedented progress but comes with vast computational costs. In this work, w

hardwarearxiv-cs-cl
11 May 2026
Research

SpecBlock: Block-Iterative Speculative Decoding with Dynamic Tree Drafting

DGX agent

arXiv:2605.07243v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by drafting a tree of candidate continuations and verifying it in one target forward. Existing drafters f

researcharxiv-cs-cl
11 May 2026
Research

SSP-based construction of evaluation-annotated data for fine-grained aspect-based sentiment analysis

DGX agent

arXiv:2605.07446v1 Announce Type: new Abstract: We report the construction of a Korean evaluation-annotated corpus, hereafter called 'Evaluation Annotated Dataset (EVAD)', and its use in Aspect-Based

researcharxiv-cs-cl
11 May 2026
Research

Statistical Patterns in the Equations of Physics and the Emergence of a Meta-Law of Nature

DGX agent

arXiv:2408.11065v2 Announce Type: replace-cross Abstract: Physics seeks to uncover the laws of Nature and express them through mathematical equations. Despite the vast diversity of natural phenomena,

researcharxiv-cs-cl
11 May 2026
Model Releases

TajPersLexon: A Tajik-Persian Lexical Resource and Hybrid Model for Cross-Script Low-Resource NLP

DGX agent

arXiv:2605.06886v1 Announce Type: new Abstract: This work introduces TajPersLexon, a curated Tajik--Persian parallel lexical resource of 40,112 word and short-phrase pairs for cross-script lexical ret

model-releasesarxiv-cs-cl
11 May 2026
Local Ai

TCMIIES: A Browser-Based LLM-Powered Intelligent Information Extraction System for Academic Literature

DGX agent

arXiv:2605.07507v1 Announce Type: new Abstract: The exponential growth of academic publications has created an urgent need for automated tools capable of extracting structured knowledge from unstructu

local-aiarxiv-cs-cl
11 May 2026
Model Releases

Teaching Language Models to Think in Code

DGX agent

arXiv:2605.07237v1 Announce Type: new Abstract: Tool-integrated reasoning (TIR) has emerged as a dominant paradigm for mathematical problem solving in language models, combining natural language (NL)

model-releasesarxiv-cs-cl
11 May 2026
Safety

TextLDM: Language Modeling with Continuous Latent Diffusion

DGX agent

arXiv:2605.07748v1 Announce Type: new Abstract: Diffusion Transformers (DiT) trained with flow matching in a VAE latent space have unified visual generation across images and videos. A natural next st

safetyarxiv-cs-cl
11 May 2026
← Previous
1…102103104105106…161
Next →