AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
11 May 2026

CktFormalizer: Autoformalization of Natural Language into Circuit Representations

HardwareDGX agent

arXiv:2605.07782v1 Announce Type: new Abstract: LLMs can generate hardware descriptions from natural language specifications, but the resulting Verilog often contains width mismatches, combinational l

CLIPer: Tailoring Diverse User Preference via Classifier-Guided Inference-Time Personalization

ResearchDGX agent

arXiv:2605.07162v1 Announce Type: new Abstract: Personalized LLMs can significantly enhance user experiences by tailoring responses to preferences such as helpfulness, conciseness, and humor. However,

Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2510.07926v2 Announce Type: replace Abstract: Despite demonstrating remarkable performance across a wide range of tasks, large language models (LLMs) have also been found to frequently produce o

Conformal Path Reasoning: Trustworthy Knowledge Graph Question Answering via Path-Level Calibration

TutorialsDGX agent

arXiv:2605.08077v1 Announce Type: new Abstract: Knowledge Graph Question Answering (KGQA) has shown promise for grounded and interpretable reasoning, yet existing approaches often fail to provide reli

Data Contamination in Neural Hieroglyphic Translation: A Reproducibility Study

Model ReleasesDGX agent

arXiv:2605.07453v1 Announce Type: new Abstract: Ancient and endangered languages pose a unique challenge for NLP: their datasets are inherently scarce, difficult to expand, and built from formulaic co

DiffRetriever: Parallel Representative Tokens for Retrieval with Diffusion Language Models

ResearchDGX agent

arXiv:2605.07210v1 Announce Type: cross Abstract: PromptReps showed that an autoregressive language model can be used directly as a retriever by prompting it to generate dense and sparse representatio

Don't Ignore the Tail: Decoupling top-K Probabilities for Efficient Language Model Distillation

ResearchDGX agent

arXiv:2602.20816v3 Announce Type: replace Abstract: The core learning signal used in language model distillation is the standard Kullback-Leibler (KL) divergence between the student and teacher distri

ExpThink: Experience-Guided Reinforcement Learning for Adaptive Chain-of-Thought Compression

ResearchDGX agent

arXiv:2605.07501v1 Announce Type: cross Abstract: Large reasoning models (LRMs) achieve strong performance via extended chain-of-thought (CoT) reasoning, yet suffer from excessive token consumption an

FinReasoning: A Hierarchical Benchmark for Reliable Financial Research Reporting

Model ReleasesDGX agent

arXiv:2603.19254v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in financial research workflows, where their role is evolving from single-model assistance fo

From 0-Order Selection to 2-Order Judgment: Combinatorial Hardening Exposes Compositional Failures in Frontier LLMs

SafetyDGX agent

arXiv:2605.07268v1 Announce Type: new Abstract: Multiple-choice reasoning benchmarks face dual challenges: rapid saturation from advancing models and data contamination that undermines static evaluati

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems

ResearchDGX agent

arXiv:2506.04565v2 Announce Type: replace-cross Abstract: Compound AI Systems (CAIS) are an emerging paradigm that integrates large language models (LLMs) with external components, including retriever

Generating training datasets for legal chatbots in Korean

Local AiDGX agent

arXiv:2605.07432v1 Announce Type: new Abstract: Chatbots are robots that can communicate with humans using text or voice signals. Legal chatbots improve access to justice, since legal representation a

GLiGuard: Schema-Conditioned Classification for LLM Safeguard

Model ReleasesDGX agent

arXiv:2605.07982v1 Announce Type: new Abstract: Ensuring safe, policy-compliant outputs from large language models requires real-time content moderation that can scale across multiple safety dimension

Gradient-Based LoRA Rank Allocation Under GRPO: An Empirical Study

Model ReleasesDGX agent

arXiv:2605.07366v1 Announce Type: new Abstract: Adaptive rank allocation for LoRA, allocating more parameters to important layers and fewer to unimportant ones, consistently improves efficiency under

GRaSp: Automatic Example Optimization for In-Context Learning in Low-Data Tasks

ResearchDGX agent

arXiv:2605.07454v1 Announce Type: new Abstract: In-context learning enables large language models to adapt to new tasks, but their performance is highly sensitive to the selected examples. Finding eff

Guidance Is Not a Hyperparameter: Learning Dynamic Control in Diffusion Language Models

SafetyDGX agent

arXiv:2605.07701v1 Announce Type: new Abstract: Classifier-Free Guidance (CFG) is a widely used mechanism for controlling diffusion-based generative models, yet its guidance scale is typically treated

How to Train Your Latent Diffusion Language Model Jointly With the Latent Space

TutorialsDGX agent

arXiv:2605.07933v1 Announce Type: new Abstract: Latent diffusion models offer an attractive alternative to discrete diffusion for non-autoregressive text generation by operating on continuous text rep

How Value Induction Reshapes LLM Behaviour

SafetyDGX agent

arXiv:2605.07925v1 Announce Type: new Abstract: Conversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and em

Human-like fleeting memory improves language learning but impairs reading time prediction in transformer language models

TutorialsDGX agent

arXiv:2508.05803v2 Announce Type: replace Abstract: Human memory is fleeting. As words are processed, the exact wordforms that make up incoming sentences are rapidly lost. Cognitive scientists have lo

Hybrid TF--IDF Logistic Regression and MLP Neural Baseline for Indonesian Three-Class Sentiment Analysis on Social Media Text

ApplicationsDGX agent

arXiv:2605.07793v1 Announce Type: new Abstract: This paper presents a compact three-class sentiment analysis study for Indonesian social media text. The task is formulated with positive, negative, and

Intent-Driven Semantic ID Generation for Grounded Conversational News Recommendation

Model ReleasesDGX agent

arXiv:2605.07613v1 Announce Type: new Abstract: Conversational news recommendation requires grounding each suggestion in a rapidly evolving article corpus while addressing implicit user intents that l

InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search

Model ReleasesDGX agent

arXiv:2605.07510v1 Announce Type: cross Abstract: Existing benchmarks for multimodal agentic search evaluate multimodal search and visual browsing, but visual evidence is either confined to the input

Interpreting Speaker Characteristics in the Dimensions of Self-Supervised Speech Features

ResearchDGX agent

arXiv:2603.03096v2 Announce Type: replace-cross Abstract: How do speech models trained through self-supervised learning structure their representations? Previous studies have looked at how information

Is She Even Relevant? When BERT Ignores Explicit Gender Cues

Local AiDGX agent

arXiv:2605.07622v1 Announce Type: new Abstract: Gender bias in large language models has primarily been investigated for English, while languages with grammatical or morphological gender remain compar

LaTER: Efficient Test-Time Reasoning via Latent Exploration and Explicit Verification

ResearchDGX agent

arXiv:2605.07315v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning improves large language models (LLMs) on difficult tasks, but it also makes inference expensive because every intermedi

Learning Agent Routing From Early Experience

Model ReleasesDGX agent

arXiv:2605.07180v1 Announce Type: new Abstract: LLM agents achieve strong performance on complex reasoning tasks but incur high latency and compute cost. In practice, many queries fall within the capa

LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction

TutorialsDGX agent

arXiv:2605.06676v1 Announce Type: cross Abstract: Long-context inference in Large Language Models (LLMs) is bottlenecked by the linear growth of Key-Value (KV) cache memory. Existing KV cache compress

LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling

AgentsDGX agent

arXiv:2605.08083v1 Announce Type: new Abstract: Test-time scaling (TTS) has become an effective approach for improving large language model performance by allocating additional computation during infe

Mean-Pooled Cosine Similarity is Not Length-Invariant: Theory and Cross-Domain Evidence for a Length-Invariant Alternative

Model ReleasesDGX agent

arXiv:2605.07345v1 Announce Type: new Abstract: Mean-pooled cosine similarity is the default metric for comparing neural representations across languages, modalities, and tasks. We establish that this

Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors

ResearchDGX agent

arXiv:2605.07847v1 Announce Type: new Abstract: As user simulators are increasingly used for interactive training and evaluation of AI assistants, it is essential that they represent the diverse behav

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM

ResearchDGX agent

arXiv:2605.05927v2 Announce Type: replace Abstract: Speech large language models (SLMs) are typically built from text large language model (TLM) checkpoints, yet they still suffer from a substantial m

MIPIAD: Multilingual Indirect Prompt Injection Attack Defense with Qwen -- TF-IDF Hybrid and Meta-Ensemble Learning

Model ReleasesDGX agent

arXiv:2605.07269v1 Announce Type: new Abstract: Indirect prompt injection remains a persistent weakness in retrieval-augmented and tool-using LLM systems, and the problem becomes harder to characteris

Multi-Dimensional Evaluation of LLMs for Grammatical Error Correction

ResearchDGX agent

arXiv:2605.07635v1 Announce Type: new Abstract: Automated assistants for Grammatical Error Correction are now embedded in educational platforms serving millions of learners, yet three critical gaps re

MultiSoc-4D: A Benchmark for Diagnosing Instruction-Induced Label Collapse in Closed-Set LLM Annotation of Bengali Social Media

Model ReleasesDGX agent

arXiv:2605.06940v1 Announce Type: new Abstract: Annotation automation via Large Language Models (LLMs) is the core approach for scaling NLP datasets; however, LLM behavior with respect to closed-set i

NCL-UoR at SemEval-2026 Task 5: Embedding-Based Methods, Fine-Tuning, and LLMs for Word Sense Plausibility Rating

Model ReleasesDGX agent

arXiv:2603.08256v2 Announce Type: replace Abstract: Word sense plausibility rating requires predicting the human-perceived plausibility of a given word sense on a 1-5 scale in the context of short nar

Neural Neural Scaling Laws

Model ReleasesDGX agent

arXiv:2601.19831v2 Announce Type: replace-cross Abstract: Neural scaling laws predict how language model performance improves with increased training inputs. While aggregate metrics like validation lo

Not All Tokens Learn Alike: Attention Entropy Reveals Heterogeneous Signals in RL Reasoning

TutorialsDGX agent

arXiv:2605.07660v1 Announce Type: new Abstract: Reinforcement-learning-based post-training has become a key approach for improving the reasoning ability of large language models, but its token-level l

NSMQ Riddles: A Benchmark of Scientific and Mathematical Riddles for Quizzing Large Language Models

Model ReleasesDGX agent

arXiv:2605.07051v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown good performance on various science educational benchmarks, demonstrating their potential for use in science and

On the Complexity of the Matching Problem of Regular Expressions with Backreferences

ResearchDGX agent

arXiv:2605.07289v1 Announce Type: cross Abstract: ReDoS is a well-known type of algorithmic complexity attack, where an adversary supplies maliciously crafted strings to a regular expression matching

OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling

Model ReleasesDGX agent

arXiv:2605.07815v1 Announce Type: cross Abstract: Muon improves neural-network training by orthogonalizing matrix-valued updates, but it leaves each layer's update magnitude controlled mostly by a glo

Overview of the TREC 2025 RAGTIME Track

ResearchDGX agent

arXiv:2602.10024v2 Announce Type: replace-cross Abstract: The principal goal of the RAG TREC Instrument for Multilingual Evaluation (RAGTIME) track at TREC is to study report generation from multiling

PaT: Planning-after-Trial for Efficient Test-Time Code Generation

SafetyDGX agent

arXiv:2605.07248v1 Announce Type: new Abstract: Beyond training-time optimization, scaling test-time computation has emerged as a key paradigm to extend the reasoning capabilities of Large Language Mo

PolySQL: Scaling Text-to-SQL Evaluation Across SQL Dialects via Automated Backend Isomorphism

ResearchDGX agent

arXiv:2605.07796v1 Announce Type: new Abstract: SQL dialects vary in syntax, types, and functions across database engines. Text-to-SQL benchmarks, however, predominantly support only SQLite. This crea

ProtSent: Protein Sentence Transformers

ResearchDGX agent

arXiv:2605.06830v1 Announce Type: cross Abstract: Protein language models (pLMs) produce per-residue representations that capture evolutionary and structural information, yet their mean-pooled sequenc

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory

HardwareDGX agent

arXiv:2605.06675v1 Announce Type: cross Abstract: Large language models cache all previously computed key-value (KV) pairs during generation, and this KV cache grows linearly with sequence length, mak

Reflections and New Directions for Human-Centered Large Language Models

SafetyDGX agent

arXiv:2605.06901v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly shaping the private and professional lives of users, with numerous applications in business, education, fi

Reliable Chain-of-Thought via Prefix Consistency

ResearchDGX agent

arXiv:2605.07654v1 Announce Type: cross Abstract: Large Language Models often improve accuracy on reasoning tasks by sampling multiple Chain-of-Thought (CoT) traces and aggregating them with majority

ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards

Model ReleasesDGX agent

arXiv:2510.00568v3 Announce Type: replace Abstract: Search agents powered by Large Language Models (LLMs) have demonstrated significant potential in tackling knowledge-intensive tasks. Reinforcement l

Rethinking Dense Sequential Chains: Reasoning Language Models Can Extract Answers from Sparse, Order-Shuffling Chain-of-Thoughts

ResearchDGX agent

arXiv:2605.07307v1 Announce Type: new Abstract: Modern reasoning language models generate dense, sequential chain-of-thought traces implicitly assuming that every token contributes and that steps must

Rethinking Experience Utilization in Self-Evolving Language Model Agents

ResearchDGX agent

arXiv:2605.07164v1 Announce Type: new Abstract: Self-evolving agents improve by accumulating and reusing experience from past interactions. Existing work has largely focused on how experience is const

Rethinking State Tracking in Recurrent Models Through Error Control Dynamics

TutorialsDGX agent

arXiv:2605.07755v1 Announce Type: cross Abstract: The theory of state tracking in recurrent architectures has predominantly focused on expressive capacity: whether a fixed architecture can theoretical

Rethinking Weight Tying: Pseudo-Inverse Tying for LM Stable Training and Updates

Model ReleasesDGX agent

arXiv:2602.04556v2 Announce Type: replace Abstract: Weight tying is widely used in compact language models to reduce parameters by sharing the token table between the input embedding and the output pr

Retrieval Heads are Dynamic

ResearchDGX agent

arXiv:2602.11162v2 Announce Type: replace Abstract: Recent studies have identified 'retrieval heads' in Large Language Models (LLMs) responsible for extracting information from input contexts. However

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning

ResearchDGX agent

arXiv:2605.07106v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have made remarkable progress on vision-language reasoning, yet most methods still compress visual evidence int

S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models

Model ReleasesDGX agent

arXiv:2503.05085v2 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have fundamentally reshaped speech-to-speech (S2S) systems, enabling increasingly natural spoken int

SAGE: Hierarchical LLM-Based Literary Evaluation through Ontology-Grounded Interpretive Dimensions

SafetyDGX agent

arXiv:2605.07102v1 Announce Type: new Abstract: Evaluating literary quality requires assessing interpretive dimensions such as cultural representation, emotional depth, and philosophical sophisticatio

SCENE: Recognizing Social Norms and Sanctioning in Group Chats

Model ReleasesDGX agent

arXiv:2605.07823v1 Announce Type: new Abstract: Online group chats are social spaces with implicit behavior patterns that, when broken, are often met with social sanctioning from the group. The abilit

Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability

AgentsDGX agent

arXiv:2605.07110v1 Announce Type: new Abstract: Computer-use agents(CUAs)are moving frombounded benchmarks toward real software environments, wherethey operate browsers, desktops, mobile applications,

SEIF: Self-Evolving Reinforcement Learning for Instruction Following

ResearchDGX agent

arXiv:2605.07465v1 Announce Type: new Abstract: Instruction following is a fundamental capability of large language models (LLMs), yet continuously improving this capability remains challenging. Exist

Self-Consolidating Language Models: Continual Knowledge Incorporation from Context

ResearchDGX agent

arXiv:2605.07076v1 Announce Type: new Abstract: Large language models (LLMs) increasingly receive information as streams of passages, conversations, and long-context workflows. While longer context wi

← Previous
1…8182838485…129
Next →