AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
4 Aug 2026

SERL-SQL: Selective Hindsight Distillation for Text-to-SQL Reinforcement Agentic Learning

Model ReleasesDGX agent

arXiv:2608.00485v1 Announce Type: new Abstract: Recent Text-to-SQL systems increasingly rely on multi-turn interaction, execution feedback, and reinforcement learning. However, most existing methods u

ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors

AgentsDGX agent

arXiv:2608.01204v1 Announce Type: new Abstract: Conversational investment advisors influence not only what users know, but also how they make subsequent decisions as market conditions evolve. Existing

SIEVE: Selective Integrity Verification and Escalation for Defending LLM Agents against Indirect Prompt Injection

AgentsDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2512.06716v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used as the core of agentic systems due to their strong reasoning, planning, and tool-use capabi

Sixteen models, fewer than two voices: measuring ensemble dispersion where no answer is uniquely correct

ResearchDGX agent

arXiv:2608.00285v1 Announce Type: new Abstract: Sixteen language models drawn from ten families produced, on average, the semantic diversity of 1.69 distinct formulations of a psychotherapeutic case,

SLMs as Multi-Agent Routers: A Progressive SFT and Reinforcement Learning Approach

Model ReleasesDGX agent

arXiv:2608.00030v1 Announce Type: new Abstract: Specialised retrieval agents typically surface higher quality results than general-purpose search, but selecting the optimal agent for a given query rem

Slot2Text: Object-Centric Visual Tokenization for Efficient and Spatially Traceable Surgical MLLMs

SafetyDGX agent

arXiv:2608.01473v1 Announce Type: cross Abstract: Multimodal large language models (MLLM) for surgical scene understanding typically inject hundreds of dense visual tokens into a language model, leadi

SoM-1K: A Thousand-Problem Benchmark Dataset for Strength of Materials

Model ReleasesDGX agent

arXiv:2509.21079v2 Announce Type: replace Abstract: Foundation models have shown remarkable capabilities in various domains, but their performance on complex, multimodal engineering problems remains l

SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models

Model ReleasesDGX agent

arXiv:2608.01899v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduc

Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation

Model ReleasesDGX agent

arXiv:2608.01666v1 Announce Type: new Abstract: However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open q

SWE-Touch: Benchmarking Coding Agents When Users Touch the Code

AgentsDGX agent

arXiv:2608.02499v1 Announce Type: cross Abstract: Real-world software development requires coding agents to operate in shared workspaces where users may inspect and modify code during an ongoing task,

TAB-PO: Preference Optimization with a Token-Level Adaptive Barrier for Token-Critical Structured Generation

Model ReleasesDGX agent

arXiv:2603.00025v3 Announce Type: replace Abstract: Direct Preference Optimization (DPO) is effective for offline alignment but poorly matched to ontology-driven structured prediction, where preferred

TELLER: Non-intrusive Cross-Layer Root-Cause Analysis for LLM Inference

Local AiDGX agent

arXiv:2608.01975v1 Announce Type: cross Abstract: Large language model (LLM) inference has evolved from an offline workload into a continuously operated software service, yet root-cause analysis remai

TextNCA: Neural Cellular Automata for Language Modeling via Hierarchical Local Attention

Model ReleasesDGX agent

arXiv:2608.02050v1 Announce Type: new Abstract: Can a strictly local, iterated, weight-shared computation primitive support language modelling, and which of those three properties actually drives the

The Holistic Storage of Verb+Up Phrases in Text-based and Audio-based Language Models

ResearchDGX agent

arXiv:2606.13993v2 Announce Type: replace Abstract: A crucial aspect of linguistic capability is the ability to trade off between stored representations and abstract knowledge: one must retrieve learn

The Learning Objective Governs Perceptual Narrowing: A Cross-Lingual, Layer-Wise, Ten-Seed Study of Self-Supervised Speech Encoders

Model ReleasesDGX agent

arXiv:2608.00507v1 Announce Type: new Abstract: Perceptual narrowing---the developmental loss of non-native phoneme discrimination in the first year of life itep{werker1984}---is a canonical developme

The methodology of Constructing the Large-Scale Dataset for Detecting Presuicidal and Anti-Suicidal Signals in Social Media Texts in Russian

ResearchDGX agent

arXiv:2608.00497v1 Announce Type: new Abstract: The suicide is a terrifying act of a person who is misled by his own mental state. This problem arises across many countries. Sadly, Russia also has qui

The Role of Disfluencies in Speech Translation

Model ReleasesDGX agent

arXiv:2608.02138v1 Announce Type: new Abstract: Current speech translation systems, including SpeechLLMs, are trained on cleaned text and tend to strip disfluencies like filled pauses and false starts

Through the LENS: Local Geometric Decomposition of Vision-Language Model Representations

Local AiDGX agent

arXiv:2608.00561v1 Announce Type: cross Abstract: Vision-language models (VLMs) process image patches and text tokens in a shared residual stream, but the local geometry through which the two modaliti

TIDES: A Longitudinal Bilingual Dataset for Modeling Multi-Party Social Dynamics

ApplicationsDGX agent

arXiv:2608.01724v1 Announce Type: new Abstract: Group conversations are fundamental to human collaboration, yet standard large language models (LLMs) still struggle with the complexities of multi-part

Token-Native Storage: Read and Write in your Agent's Language

AgentsDGX agent

arXiv:2608.02376v1 Announce Type: cross Abstract: Search and database engines still store text as UTF-8, a format built for humans. But the systems that increasingly read and write that text (embedder

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning

SafetyDGX agent

arXiv:2608.01743v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a central paradigm for large language model (LLM) post-training, but optimization toward new objectives can deg

Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization

Model ReleasesDGX agent

arXiv:2603.08091v2 Announce Type: replace Abstract: Large language model (LLM)-based judges are widely adopted for automated evaluation and reward modeling, yet their judgments are often affected by j

Towards a theory of morphology-driven marking in the lexicon: The case of the state

ResearchDGX agent

arXiv:2604.03422v2 Announce Type: replace Abstract: All languages have a noun category, but its realisation varies considerably. Depending on the language, semantic and/or morphosyntactic differences

TRACE-TS: Attribution-Grounded and Traceable Sensor-Language Reasoning for Human Activity Understanding

Local AiDGX agent

arXiv:2608.00200v1 Announce Type: cross Abstract: Wearable sensors capture fine-grained motion patterns that support rich behavioral understanding, yet most existing methods reduce these signals to ac

Training-Free versus Training-Based Intent Classification in LLMs: Accuracy, Robustness, and Failure Modes

ResearchDGX agent

arXiv:2608.02415v1 Announce Type: new Abstract: Intent classification in Large Language Models (LLMs) involves categorizing user prompts into predefined classes. For instance, given a user prompt, the

TRAM: Enhancing Multimodal Reasoning with Trajectory-Derived Auxiliary Memory

ResearchDGX agent

arXiv:2608.01922v1 Announce Type: new Abstract: Multimodal Large Reasoning Models (MLRMs) have achieved strong performance on tasks requiring visual understanding and multi-step inference. However, as

Transformers perform adaptive partial pooling

ResearchDGX agent

arXiv:2602.03980v2 Announce Type: replace Abstract: Any language model must decide what to say in novel contexts based on information from similar contexts. But what about contexts that are not novel

TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs

Model ReleasesDGX agent

arXiv:2608.00640v1 Announce Type: new Abstract: Large language models are increasingly viewed as a potential means of mitigating global health inequities, yet their outputs often reflect dominant high

TrimMoE A communication aware and adaptive depth framework for distributed edge inference

Model ReleasesDGX agent

arXiv:2608.00573v1 Announce Type: cross Abstract: Serving Mixture-of-Experts (MoE) large language models across distributed edge servers is bottlenecked by the cross-server expert transmission. The ex

Trustworthiness Costs of Domain Adaptation in Small Language Models:A Cross-Architecture Empirical Study

Model ReleasesDGX agent

arXiv:2608.00042v1 Announce Type: new Abstract: Domain adaptation of small language models (SLMs) has emerged as a practical strategy for deploying capable NLP systems in resource-constrained, high-st

TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning

SafetyDGX agent

arXiv:2510.03519v2 Announce Type: replace Abstract: Time series reasoning is crucial to decision-making in diverse domains, including finance, energy, and scientific discovery. While existing time ser

Two-Stage Bengali Sentiment Classification: Domain Adaptation Through Continual Learning and Parameter-Efficient Fine-Tuning

Model ReleasesDGX agent

arXiv:2608.01471v1 Announce Type: new Abstract: Understanding sentiment in low-resource languages remains a key challenge for Natural Language Processing (NLP), particularly when domain-specific data

UEmbed: Unified Sparse and Dense Multimodal Embeddings

AgentsDGX agent

arXiv:2608.02583v1 Announce Type: cross Abstract: Sparse retrieval underpins modern search systems, from web search to retrieval-augmented generation. Existing work has introduced Learned Sparse Retri

Understanding Sparse Attention Selectivity in Long-Context Foundation Models via Counterfactual Evaluation

ResearchDGX agent

arXiv:2608.01676v1 Announce Type: new Abstract: Sparse attention is widely deployed in long-context serving stacks, yet no framework audits how discarding blocks changes the influence of specific cont

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering

ResearchDGX agent

arXiv:2608.01147v1 Announce Type: cross Abstract: Knowledge-Based Visual Question Answering (KB-VQA) requires retrieving relevant entity knowledge from external sources to answer visually grounded que

Unleashing the Potential of Large Language Models: A Blueprint for Real-Time, Enterprise-Ready Deployments

SafetyDGX agent

arXiv:2608.00419v1 Announce Type: cross Abstract: Large language models deployed in real-time, regulated settings face knowledge staleness, catastrophic forgetting, hallucination, and weak feedback lo

Unpacking Hateful Memes: Presupposed Context and False Claims

ResearchDGX agent

arXiv:2510.09935v2 Announce Type: replace Abstract: While memes are often humorous, they are frequently used to disseminate hate, causing serious harm to individuals and society. Current approaches to

Unsupervised Multidomain Approaches to Named Entity Recognition with Small Datasets

ResearchDGX agent

arXiv:2608.00984v1 Announce Type: new Abstract: This paper explores the challenges and the methodologies associated with learning quality representations in scenarios with unlabelled small or limited

V-Mem: Modality-Routed Retrieval for Long-Term Multimodal Agentic Memory

AgentsDGX agent

arXiv:2608.01543v1 Announce Type: cross Abstract: Interaction between users and LLM agents is increasingly multimodal: conversations interleave text with images, and a later question may target either

Verification Without Sufficiency: Per-Chunk Filtering Fails on Multi-Hop RAG, and Decomposition Repairs It

ResearchDGX agent

arXiv:2608.00585v1 Announce Type: new Abstract: Verification for retrieval-augmented generation usually scores each retrieved chunk and drops the ones that fail. We show this cannot work for multi-hop

Verifier-Induced Support Reshaping in On-Policy Optimization

SafetyDGX agent

arXiv:2608.00220v1 Announce Type: cross Abstract: We show that on-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while making successful behaviors for l

Visualising Information Flow in Word Embeddings with Diffusion Tensor Imaging

ResearchDGX agent

arXiv:2601.05713v2 Announce Type: replace Abstract: Understanding how large language models (LLMs) represent natural language is a central challenge in natural language processing (NLP) research. Many

What Makes Position Zero Special? A Mechanistic Study of Position Zero Attention Sinks in LLMs

Model ReleasesDGX agent

arXiv:2603.06591v2 Announce Type: replace-cross Abstract: Transformers frequently allocate disproportionate attention to specific tokens, a phenomenon known as attention sinks. Causal large language m

What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs

Model ReleasesDGX agent

arXiv:2608.00013v1 Announce Type: new Abstract: Choosing the right large language model (LLM) backbone is the most consequential decision when building a vision-language model (VLM), yet it remains fu

When LLM Essays Outscore Student Essays: What a Korean Writing Rubric Rewards and Where Readers Disagree

ResearchDGX agent

arXiv:2601.19913v4 Announce Type: replace Abstract: LLMs now help students plan, draft, and revise essays. Educational assessment therefore faces a basic question: how should student and LLM writing b

When Only the Final Text Survives: Implicit Execution Tracing for Multi-Agent Auditing

AgentsDGX agent

arXiv:2603.17445v5 Announce Type: replace-cross Abstract: When a multi-agent system produces an incorrect or harmful answer, who is accountable if execution logs and agent identifiers are unavailable?

When Retrieval Helps and Distracts: Evaluating Evidence-Generating LLMs for Biomedical Claim Verification

Model ReleasesDGX agent

arXiv:2608.01409v1 Announce Type: new Abstract: Biomedical fact-checking systems must do more than predict whether a claim is supported, contradicted, or unaddressed: they should also produce evidence

When Words Divide: Diachronic Ideological Polarization in Political Discourse on Social Media

ResearchDGX agent

arXiv:2608.01176v1 Announce Type: new Abstract: Political polarization has become a defining feature of online discourse, yet its long-term evolution remains poorly understood. We present a longitudin

Where did the ambiguity go? Examining how multimodal models interpret polysemous words

ResearchDGX agent

arXiv:2608.00410v1 Announce Type: cross Abstract: Human language is highly polysemous. Many common words (e.g., 'bank' or 'palm') carry several distinct meanings that shape what humans communicate and

Who Should Be Generated? Justifying Demographic Targets in Open-Ended Generation

SafetyDGX agent

arXiv:2608.02551v1 Announce Type: cross Abstract: Fairness evaluation concerns not only what a model produces, but also what its outputs ought to be compared against. When a model generates 'a CEO in

Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy

ResearchDGX agent

arXiv:2608.01017v1 Announce Type: new Abstract: A language model that abandons a correct medical answer under user pushback is more dangerous than one that was simply wrong, because it lends the credi

Writing-System-Level Tokenizer Adaptation for Byte-Level BPE

Model ReleasesDGX agent

arXiv:2608.00582v1 Announce Type: new Abstract: Pretrained byte-level BPE tokenizers can segment underrepresented languages inefficiently. Replacing a tokenizer changes the meaning of nearly every tok

XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding

Model ReleasesDGX agent

arXiv:2608.00036v1 Announce Type: new Abstract: Real-world document tasks often ask professionals to answer questions from annual reports, regulations, clinical guidelines, and technical manuals that

3 Aug 2026

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements

Model ReleasesDGX agent

arXiv:2607.28661v1 Announce Type: new Abstract: Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding

Authorship Verification of Transcribed German-Language Videos

ResearchDGX agent

arXiv:2607.29168v1 Announce Type: new Abstract: Authorship Verification (AV) represents an important subfield of digital text forensics and addresses the fundamental question of whether two texts were

Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications

Model ReleasesDGX agent

arXiv:2607.28840v1 Announce Type: new Abstract: Large language models are increasingly deployed in financial applications that combine retrieval, proprietary data, tool use, orchestration logic, monit

BLADE: Boundary-Expanded and Layer-Adaptive Dynamic Exit for Efficient LLM Reasoning

ResearchDGX agent

arXiv:2607.28966v1 Announce Type: new Abstract: Large language models often improve task performance by generating long reasoning traces, but the resulting computation is frequently wasted on redundan

Bridging the Question-Answer Gap in Retrieval-Augmented Generation: Hypothetical Prompt Embeddings

SafetyDGX agent

arXiv:2607.29402v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems synergize retrieval mechanisms with generative language models to enhance the accuracy and relevance of r

Can LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMs

Model ReleasesDGX agent

arXiv:2607.28634v1 Announce Type: new Abstract: The estimation of item difficulty plays a key role in both formative assessment and large-scale high-stakes summative assessments. This study explores h

Can Zero-Shot LLMs Predict Child Malnutrition? A Fairness and Temporal Robustness Study

SafetyDGX agent

arXiv:2607.29082v1 Announce Type: new Abstract: Child malnutrition remains a major public health challenge in low- and middle-income countries, particularly in South Asia, where early identification o

← Previous
1…1011121314…128
Next →