AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
12 May 2026

Source or It Didn't Happen: A Multi-Agent Framework for Citation Hallucination Detection

Model ReleasesDGX agent

arXiv:2605.08583v1 Announce Type: new Abstract: Large language models are increasingly used in scientific writing, yet they can fabricate citation-shaped references that appear plausible but fail bibl

Sparse Layers are Critical to Scaling Looped Language Models

ResearchDGX agent

arXiv:2605.09165v1 Announce Type: cross Abstract: Looped language models repeat a set of transformer layers through depth, reducing memory costs and providing natural early-exit points at loop boundar

Sparse Reward Subsystem in Large Language Models

TutorialsDGX agent

arXiv:2602.00986v2 Announce Type: replace Abstract: Recent studies show that LLM hidden states encode reward-related information, such as answer correctness and model confidence. However, existing app


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?

Model ReleasesDGX agent

arXiv:2602.03916v3 Announce Type: replace-cross Abstract: Spatial reasoning is a fundamental aspect of human cognition, yet it remains a major challenge for contemporary vision-language models (VLMs).

Spherical Flows for Sampling Categorical Data

ResearchDGX agent

arXiv:2605.05629v2 Announce Type: replace-cross Abstract: We study the problem of learning generative models for discrete sequences in a continuous embedding space. Whereas prior approaches typically

SSA: Improving Performance With a Better Scoring Function

ResearchDGX agent

arXiv:2508.14685v4 Announce Type: replace Abstract: While transformer models exhibit strong in-context learning (ICL) abilities, they often fail to generalize under simple distribution shifts. We anal

Statistical Scouting Finds Debate-Safe but Not Debate-Useful Cases: A Matched-Ceiling Study of Open-Weight LLM Reasoning Protocols

Model ReleasesDGX agent

arXiv:2605.09618v1 Announce Type: new Abstract: When should a language model answer directly, sample and vote, or engage in multi-agent debate? Recent work shows voting often explains much of the gain

Structured Recurrent Mixers for Massively Parallelized Sequence Generation

ResearchDGX agent

arXiv:2605.08696v1 Announce Type: new Abstract: Over the last two decades, language modeling has experienced a shift from predominantly recurrent architectures that process tokens sequentially during

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data

Model ReleasesDGX agent

arXiv:2605.10129v1 Announce Type: new Abstract: Large language models (LLMs) rely on web-scale corpora for pre-training. The noise inherent in these datasets tends to obscure meaningful patterns and u

TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2605.09539v1 Announce Type: new Abstract: Multi-agent systems (MAS) have emerged as a promising paradigm for solving complex tasks. Recent work has explored self-evolving MAS that automatically

Talk to Your Slides: High-Efficiency Slide Editing via Language-Driven Structured Data Manipulation

Model ReleasesDGX agent

arXiv:2505.11604v5 Announce Type: replace Abstract: Editing presentation slides is a frequent yet tedious task, ranging from creative layout design to repetitive text maintenance. While recent GUI-bas

Task-Aware Calibration: Provably Optimal Decoding in LLMs

ResearchDGX agent

arXiv:2605.10202v1 Announce Type: cross Abstract: LLM decoding often relies on the model's predictive distribution to generate an output. Consequently, misalignment with respect to the true generating

TELL-TALE: Task Efficient LLMs with Task Aware Layer Elimination

ResearchDGX agent

arXiv:2510.22767v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) typically come with a fixed architecture, despite growing evidence that not all layers contribute equally to ever

Temporal Tokenization Strategies for Event Sequence Modeling with Large Language Models

ApplicationsDGX agent

arXiv:2512.13618v3 Announce Type: replace Abstract: Representing continuous time is a critical and under-explored challenge in modeling temporal event sequences with large language models (LLMs). Vari

Test-Time Speculation

Model ReleasesDGX agent

arXiv:2605.09329v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by using a fast draft model to generate tokens and a more accurate target model to verify them. Its perfo

The Association of Transformer-based Sentiment Analysis with Symptom Distress and Deterioration in Routine Psychotherapy Care

ResearchDGX agent

arXiv:2605.09838v1 Announce Type: new Abstract: Sentiment analysis has been of long-standing interest in psychotherapy research. Recently, the Transformer deep learning architecture has produced text-

The Astonishing Ability of Large Language Models to Parse Jabberwockified Language

ResearchDGX agent

arXiv:2602.23928v2 Announce Type: replace Abstract: We show that large language models (LLMs) have an astonishing ability to recover meaning from severely degraded English texts. Texts in which conten

The Extrapolation Cliff in On-Policy Distillation of Near-Deterministic Structured Outputs

Model ReleasesDGX agent

arXiv:2605.08737v1 Announce Type: cross Abstract: On-policy distillation (OPD) is widely used for LLM post-training. When pushed with a reward-extrapolation coefficient lambda > 1, the student can lif

The Impact of Editorial Intervention on Detecting Native Language Traces

ResearchDGX agent

arXiv:2605.10216v1 Announce Type: new Abstract: Native Language Identification (NLI) is the task of determining an author's native language (L1) from their non-native writings. With the advent of huma

The Realignment Problem: When Right becomes Wrong in LLMs

Model ReleasesDGX agent

arXiv:2511.02623v2 Announce Type: replace Abstract: Post-training alignment of large language models (LLMs) relies on large-scale human annotations guided by policy specifications that change over tim

The Truth Lies Somewhere in the Middle (of the Generated Tokens)

Local AiDGX agent

arXiv:2605.09969v1 Announce Type: cross Abstract: How should hidden states generated autoregressively be collapsed into a representation that reflects a language model's internal state? Despite tokens

Through the Lens of Character: Resolving Modality-Role Interference in Multimodal Role-Playing Agent

AgentsDGX agent

arXiv:2605.09443v1 Announce Type: cross Abstract: The advancement of Multimodal Large Language Models (MLLMs) has expanded Role-Playing Agents (RPAs) into visually grounded environments. However, huma

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning

SafetyDGX agent

arXiv:2508.20697v3 Announce Type: replace-cross Abstract: As large language models (LLMs) continue to grow in capability, so do the risks of harmful misuse through fine-tuning. While most prior studie

ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering

AgentsDGX agent

arXiv:2510.20036v2 Announce Type: replace Abstract: Large language model (LLM) agents rely on external tools to solve complex tasks, but real-world toolsets often contain redundant tools with overlapp

Topological Data Analysis Applications in Natural Language Processing: A Survey

ApplicationsDGX agent

arXiv:2411.10298v5 Announce Type: replace Abstract: The surge of data available on the Internet has driven the adoption of a wide range of computational methods for analyzing and extracting insights f

Toward Multi-Database Query Reasoning for Text2Cypher

TutorialsDGX agent

arXiv:2605.10373v1 Announce Type: cross Abstract: Large language models have significantly improved natural language interfaces to databases by translating user questions into executable queries. In p

Towards Compact Sign Language Translation: Frame Rate and Model Size Trade-offs

Model ReleasesDGX agent

arXiv:2605.09554v1 Announce Type: new Abstract: Sign Language Translation (SLT) converts sign language videos into spoken-language text, bridging communication between Deaf and hearing communities. Cu

Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents

Model ReleasesDGX agent

arXiv:2605.10832v1 Announce Type: new Abstract: Multimodal deep search requires an agent to solve open-world problems by chaining search, tool use, and visual reasoning over evolving textual and visua

TRACER: Verifiable Generative Provenance for Multimodal Tool-Using Agents

Model ReleasesDGX agent

arXiv:2605.09934v1 Announce Type: new Abstract: Multimodal large language models increasingly solve vision-centric tasks by calling external tools for visual inspection, OCR, retrieval, calculation, a

Training with Harnesses: On-Policy Harness Self-Distillation for Complex Reasoning

SafetyDGX agent

arXiv:2605.08741v1 Announce Type: new Abstract: Inference-time harnesses substantially improve large language models on complex reasoning tasks. However, the intrinsic capabilities of the underlying m

Two Ways to De-Bias an LLM-as-a-Judge: A Continuous-Score Comparison of Hierarchical Bayesian Calibration and Neural-ODE Score Transport

Model ReleasesDGX agent

arXiv:2605.09227v1 Announce Type: new Abstract: [Abridged] Using a Large Language Model (LLM) as an automatic rater (LLM-as-a-judge) is cheap but potentially biased: some judges run lenient, others st

UPA: Unsupervised Prompt Agent via Tree-Based Search and Selection

Local AiDGX agent

arXiv:2601.23273v2 Announce Type: replace Abstract: Prompt agents have recently emerged as a promising paradigm for automated prompt optimization, framing prompt discovery as a sequential decision-mak

UserGPT Technical Report

Model ReleasesDGX agent

arXiv:2605.08766v1 Announce Type: cross Abstract: Personalized user understanding from large-scale digital traces remains a fundamental challenge. Traditional user profiling methods rely on discrimina

V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning

SafetyDGX agent

arXiv:2605.10172v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have achieved remarkable success in general perception, yet complex multi-step visual reasoning remains a per

Verbalized Algorithms: Classical Algorithms are All You Need (Mostly)

ResearchDGX agent

arXiv:2509.08150v5 Announce Type: replace Abstract: Reasoning is a fundamentally algorithmic task. Yet current work on LLM-based reasoning relies on free-form generation whose theoretical guarantees (

VISTA: A Generative Egocentric Video Framework for Daily Assistance

SafetyDGX agent

arXiv:2605.10579v1 Announce Type: new Abstract: Training AI agents to proactively assist humans in daily activities, from routine household tasks to urgent safety situations, requires large-scale visu

When Efficient Communication Explains Convexity

ResearchDGX agent

arXiv:2602.02821v2 Announce Type: replace Abstract: Much recent work has argued that the variation in the languages of the world can be explained from the perspective of efficient communication; in pa

Where do aspectual variants of light verb constructions belong?

ResearchDGX agent

arXiv:2605.10605v1 Announce Type: new Abstract: Expressions with an aspectual variant of a light verb, e.g. 'take on debt' vs. 'have debt', are frequent in texts but often difficult to classify betwee

Where Does Long-Context Supervision Actually Go? Effective-Context Exposure Balancing

Model ReleasesDGX agent

arXiv:2605.10544v1 Announce Type: new Abstract: Long-context adaptation is often viewed as window scaling, but this misses a token-level supervision mismatch: in packed training with document masking,

Why is prompting hard? Understanding prompts on binary sequence predictors

ResearchDGX agent

arXiv:2502.10760v2 Announce Type: replace Abstract: Frontier models can be prompted or conditioned to do many tasks, but finding good prompts is not always easy, nor is understanding some performant p

WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation

Model ReleasesDGX agent

arXiv:2605.10912v1 Announce Type: new Abstract: Large language and vision-language models increasingly power agents that act on a user's behalf through command-line interface (CLI) harnesses. However,

XPERT: Expert Knowledge Transfer for Effective Training of Language Models

ResearchDGX agent

arXiv:2605.08842v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) language models organize knowledge into explicitly routed expert modules, making expert-level representations traceable and ana

YEZE at SemEval-2026 Task 9: Detecting Multilingual, Multicultural and Multievent Online Polarization via Heterogeneous Ensembling

ResearchDGX agent

arXiv:2605.06231v2 Announce Type: replace Abstract: This paper presents our system for SemEval-2026 Task 9: Detecting Multilingual, Multicultural and Multievent Online Polarization, which identifies p

11 May 2026

A Comparative Analysis of Classical Machine Learning and Deep Learning Approaches for Sentiment Classification on IMDb Movie Reviews

ResearchDGX agent

arXiv:2605.07811v1 Announce Type: new Abstract: This paper presents a comparative study of classical machine learning and deep learning methods for sentiment classification on the IMDb movie reviews d

A Reproducible Multi-Architecture Baseline for Token-Level Chinese Metaphor Identification under the MIPVU Framework

Model ReleasesDGX agent

arXiv:2605.07170v1 Announce Type: new Abstract: Metaphor is pervasive in everyday language, yet token-level computational identification of metaphor-related words in Chinese under the MIPVU framework

Accurate and Efficient Statistical Testing for Word Semantic Breadth

SafetyDGX agent

arXiv:2605.08048v1 Announce Type: new Abstract: Measuring the breadth of a word's meaning, or its spread across contexts, has become feasible with contextualized token embeddings. A word type can be r

Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning

Model ReleasesDGX agent

arXiv:2602.19612v3 Announce Type: replace Abstract: Machine Unlearning (MU) enables Large Language Models (LLMs) to remove unsafe or outdated information. However, existing work assumes that all facts

Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?

Model ReleasesDGX agent

arXiv:2605.07937v1 Announce Type: new Abstract: Long-horizon AI agents execute complex workflows spanning hundreds of sequential actions, yet a single wrong assumption early on can cascade into irreve

Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded Reasoning

Model ReleasesDGX agent

arXiv:2502.07143v3 Announce Type: replace Abstract: The severe shortage of medical doctors limits access to timely and reliable healthcare, leaving millions underserved. Large language models (LLMs) o

Bayesian Attention Mechanism: A Probabilistic Framework for Positional Encoding and Context Length Extrapolation

ResearchDGX agent

arXiv:2505.22842v4 Announce Type: replace Abstract: Transformer-based language models rely on positional encoding (PE) to handle token order and support context length extrapolation. However, existing

Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility

Model ReleasesDGX agent

arXiv:2605.06856v1 Announce Type: cross Abstract: Generative AI systems achieve impressive performance on standard benchmarks yet fail to deliver real-world utility, a disconnect we identify across 28

Beyond Factual Accuracy: Evaluating Global Reasoning Integrity in RAG Systems with LogicScore

Model ReleasesDGX agent

arXiv:2601.15050v4 Announce Type: replace Abstract: Current evaluation methods for Retrieval Augmented Generation (RAG) suffer from extit{factual myopia}: they relentlessly emphasize factual accuracy

Beyond 'I cannot fulfill this request': Alleviating Rigid Rejection in LLMs via Label Enhancement

SafetyDGX agent

arXiv:2605.07883v1 Announce Type: new Abstract: Large Language Models (LLMs) rely on safety alignment to obey safe requests while refusing harmful ones. However, traditional refusal mechanisms often l

Beyond Reasoning: Reinforcement Learning Unlocks Parametric Knowledge in LLMs

ResearchDGX agent

arXiv:2605.07153v1 Announce Type: new Abstract: Reinforcement learning (RL) has achieved remarkable success in LLM reasoning, but whether it can also improve direct recall of parametric knowledge rema

Beyond Single Ground Truth: Reference Monism as Epistemic Injustice in ASR Evaluation

ApplicationsDGX agent

arXiv:2605.07084v1 Announce Type: new Abstract: Automatic speech recognition (ASR) evaluation compares system output to ground truth transcripts, with Word Error Rate (WER) quantifying the distance be

Bridging Textual Profiles and Latent User Embeddings for Personalization

ResearchDGX agent

arXiv:2605.06981v1 Announce Type: cross Abstract: Personalized systems rely on user representations to connect behavioral history with downstream recommendation applications. Existing methods typicall

Can David Beat Goliath? On Multi-Hop Reasoning with Resource-Constrained Agents

SafetyDGX agent

arXiv:2601.21699v2 Announce Type: replace Abstract: Multi-turn reasoning agents solve complex questions by decomposing them into intermediate retrieval or tool-use steps, for accumulating supporting e

Can LLMs Take Retrieved Information with a Grain of Salt?

ApplicationsDGX agent

arXiv:2605.06919v1 Announce Type: new Abstract: Large language models have demonstrated impressive retrieval-augmented capabilities. However, a crucial area remains underexplored: their ability to app

Chain-based Distillation for Effective Initialization of Variable-Sized Small Language Models

Model ReleasesDGX agent

arXiv:2605.07783v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong performance but remain costly to deploy in resource-constrained settings. Training small language models (SL

ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring

Model ReleasesDGX agent

arXiv:2605.07415v1 Announce Type: cross Abstract: Referring expression grounding is a core problem in visual grounding and is widely used as a diagnostic of spatial grounding and reasoning in vision a

← Previous
1…8081828384…129
Next →