AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
28 May 2026

Syllabic-Structure Decoder for Automatic Speech Recognition in Vietnamese

ResearchDGX agent

arXiv:2605.27874v1 Announce Type: new Abstract: Most Automatic Speech Recognition (ASR) systems formulate transcription as a prediction problem over orthographic units such as characters, subwords, or

TARQ: Tail-Aware Reconstruction Quantization for Rare-Word Robust Automatic Speech Recognition

ResearchDGX agent

arXiv:2605.27808v1 Announce Type: new Abstract: Data-aware post-training quantization (PTQ) minimizes a per-token reconstruction loss on a small calibration corpus, implicitly weighting positions by t

The Abstraction Gap in Vision-Language Causal Reasoning

Model ReleasesDGX agent

arXiv:2605.28779v1 Announce Type: new Abstract: Vision-language models (VLMs) generate fluent causal explanations, but current evaluations cannot distinguish linguistic plausibility from faithful caus


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness

Model ReleasesDGX agent

arXiv:2605.28190v1 Announce Type: new Abstract: Embedding benchmarks like MTEB report a single score per model, implicitly treating robustness as a static, scalar property. We argue that embedding rob

The Missing Piece in Pre-trained Model Evaluation: Reward-Guided Decoding Unlocks Task-Oriented Behavior Without Parameter Updates

Model ReleasesDGX agent

arXiv:2605.28020v1 Announce Type: new Abstract: With the rapid progress of large language models (LLMs), reliably evaluating the capabilities of pre-trained LLMs has become increasingly important. The

TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling

SafetyDGX agent

arXiv:2605.27690v1 Announce Type: new Abstract: LLM agents increasingly operate through multi-turn tool use and environment interaction, where safety risks often emerge from intermediate steps long be

Trust Me, I'm an Expert: Decoding and Steering Authority Bias in Large Language Models

SafetyDGX agent

arXiv:2601.13433v3 Announce Type: replace Abstract: Prior research demonstrates that performance of language models on reasoning tasks can be influenced by suggestions, hints and endorsements. However

UNIQUE: Universal Top-k Sparse Attention for Training-free Inference and Sparsity-aware Training

ResearchDGX agent

arXiv:2605.27740v1 Announce Type: new Abstract: Long-context inference in large language models (LLMs) is bottlenecked by the linear growth of the self-attention key-value (KV) cache. Top-k sparse att

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention

ResearchDGX agent

arXiv:2602.07574v2 Announce Type: replace-cross Abstract: Modern multimodal large language models (MLLMs) adopt a unified self-attention design that processes visual and textual tokens at every Transf

VLMs May Not Globally Enhance Human Alignment over LLMs During Natural Reading

SafetyDGX agent

arXiv:2605.28818v1 Announce Type: new Abstract: Large language models (LLMs) have become increasingly useful computational models of human language processing, but it remains unclear whether vision-la

What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025

SafetyDGX agent

arXiv:2601.07648v2 Announce Type: replace Abstract: As Natural Language Generation (NLG) dominates modern NLP, scalable evaluation remains a critical bottleneck. Consequently, LLM-as-a-judge (LaaJ) ad

When Confidence Misleads: Suffix Anchoring and Anchor-Proximity Confidence Modulation for Diffusion Language Models

ResearchDGX agent

arXiv:2605.28181v1 Announce Type: new Abstract: Diffusion language models decode text by iteratively denoising masked token sequences, making the choice of which positions to decode a central inferenc

When Discourse Pressures Conflict: Information Structure in Vision-Language Model Outputs

ResearchDGX agent

arXiv:2605.28346v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly evaluated for whether they identify the right visual content, but little is known about whether they expr

When Helpful Context Leaks: Privacy Risks in Domain-Adapted ASR

ResearchDGX agent

arXiv:2605.28211v1 Announce Type: new Abstract: SpeechLLMs are increasingly deployed in professional settings where domain customisation is standard practice: users supply context in prompts with sens

When Seekers Are Hard to Help: Evaluating Emotional Support Dialogue Systems in Worst-Case Interactions

ResearchDGX agent

arXiv:2605.28228v1 Announce Type: new Abstract: Emotional Support Dialogue Systems (ESDSes) are increasingly evaluated and trained with LLM-simulated seekers. However, such simulated seekers often beh

Which Heads Matter for Reasoning? RL-Guided KV Cache Compression

ResearchDGX agent

arXiv:2510.08525v3 Announce Type: replace Abstract: Reasoning large language models exhibit complex reasoning behaviors via extended chain-of-thought generation that are highly fragile to information

Why Gaussian Diffusion Models Fail on Discrete Data and How to Prevent It?

TutorialsDGX agent

arXiv:2604.02028v2 Announce Type: replace Abstract: Diffusion models have become a standard approach for generative modeling in continuous domains, yet their application to discrete data remains chall

Why We Need Speech to Evaluate Speech Translation

ResearchDGX agent

arXiv:2605.28227v1 Announce Type: new Abstract: Speech translation models are increasingly capable of preserving speech-specific information (e.g., speaker gender, prosody, and emphasis), yet evaluati

xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction

SafetyDGX agent

arXiv:2503.18893v2 Announce Type: replace Abstract: Long-context Large Language Models (LLMs) enable powerful applications but incur high memory costs due to the key-value states (KV-Cache). Recent st

You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents

Model ReleasesDGX agent

arXiv:2605.27586v1 Announce Type: cross Abstract: Ensuring agent behaviors in distributed open multi-agent systems remains challenging, especially as populations grow and unaligned agents may exist. W

27 May 2026

A Method for Learning Large-Scale Computational Construction Grammars from Semantically Annotated Corpora

ResearchDGX agent

arXiv:2603.12754v2 Announce Type: replace Abstract: We present a method for learning large-scale, broad-coverage construction grammars from corpora of language use. Starting from utterances annotated

Accountable Human-AI Deliberation with LLMs: Scaling Collective Intelligence through Symbiotic Scaffolding

ResearchDGX agent

arXiv:2605.26940v1 Announce Type: new Abstract: Large language models (LLMs) can support democratic deliberation at scales previously constrained by turn-taking and facilitation bandwidth. Recent work

AdaSD: Adaptive Speculative Decoding for Efficient Language Model Inference

Model ReleasesDGX agent

arXiv:2512.11280v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved remarkable performance across a wide range of tasks, but their increasing parameter sizes significantly s

ADRD-Bench: A Preliminary LLM Benchmark for Alzheimer's Disease and Related Dementias

Model ReleasesDGX agent

arXiv:2602.11460v2 Announce Type: replace Abstract: Large language models (LLMs) have shown great potential for healthcare applications. However, existing evaluation benchmarks provide minimal coverag

Agreement Between Large Language Models and Human Raters in Essay Scoring: A Research Synthesis

ResearchDGX agent

arXiv:2512.14561v2 Announce Type: replace Abstract: Despite the growing promise of large language models (LLMs) in automated essay scoring (AES), empirical findings regarding their reliability compare

AIDG: A Formal Decomposition of Information Extraction and Containment Asymmetries in Multi-Turn LLM Dialogue

Local AiDGX agent

arXiv:2602.17443v2 Announce Type: replace Abstract: Multi-turn LLM evaluation is typically reported as a single win-rate scalar, conflating distinct capabilities. We introduce AIDG (Adversarial Inform

AlbanianLLMSafety: A Safety Evaluation Dataset for Large Language Models in Albanian

Model ReleasesDGX agent

arXiv:2605.26954v1 Announce Type: new Abstract: Safety evaluation of Large Language Models (LLMs) has largely focused on high-resource languages, leaving low-resource languages critically underserved.

AlignEvoSkill: Towards Knowledge-Aware and Task-Aligned Agent Skill Evolution

SafetyDGX agent

arXiv:2506.23149v2 Announce Type: replace Abstract: Reusable skills play a key role in improving LLM-based agents, but existing skill-evolution methods often fail to ensure that evolved skills both co

Anchored Decoding: Provably Reducing Copyright Risk for Any Language Model

ResearchDGX agent

arXiv:2602.07120v2 Announce Type: replace Abstract: Language models (LMs) tend to memorize portions of their training data and emit verbatim spans. When the underlying sources are sensitive or copyrig

Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation

Model ReleasesDGX agent

arXiv:2605.26918v1 Announce Type: new Abstract: Video generation models (VGMs) are rapidly entering classrooms, yet existing benchmarks evaluate only perceptual quality, intrinsic faithfulness, generi

Attribute-Based Diagnosis of LLM Alignment with Hate Speech Annotations

Model ReleasesDGX agent

arXiv:2605.27025v1 Announce Type: new Abstract: Hate speech annotation is costly, subjective, and prone to annotator disagreement, making large-scale dataset construction challenging. We systematicall

BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning

SafetyDGX agent

arXiv:2605.27110v1 Announce Type: cross Abstract: In this work, we propose BAIT (Boundary-Aware Iterative Trap), a three-step jailbreak framework that approaches malicious goals through internal discl

BESPOKE: Benchmark for Search-Augmented Large Language Model Personalization via Diagnostic Feedback

Model ReleasesDGX agent

arXiv:2509.21106v2 Announce Type: replace Abstract: Search-augmented large language models (LLMs) have advanced information-seeking tasks by integrating retrieval into generation, reducing users' cogn

Beyond Binary: Speech Representations Across the Cognitive Score Hierarchy

ResearchDGX agent

arXiv:2605.27189v1 Announce Type: new Abstract: This study examines the relationship between speech representations and the hierarchical structure of cognitive assessment in mild cognitive impairment.

Beyond Self-Talk: A Communication-Centric Survey of LLM-Based Multi-Agent Systems

AgentsDGX agent

arXiv:2502.14321v3 Announce Type: replace-cross Abstract: Large language model-based multi-agent systems have recently gained significant attention due to their potential for complex, collaborative, a

BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

Model ReleasesDGX agent

arXiv:2603.03194v2 Announce Type: replace Abstract: Current code-agent benchmarks primarily evaluate localized issue resolution within a single target repository, leaving under-tested many software en

BhashaSetu: A Data-Centric Approach to Low-Resource Machine Translation

Model ReleasesDGX agent

arXiv:2605.27050v1 Announce Type: new Abstract: We present BhashaSetu, a linguistically enriched English--Marathi parallel dataset addressing persistent data limitations in low-resource neural machine

Bounded Path Context: A Controlled Study of Visible Path History in LLM-Based Knowledge Graph Question Answering

Local AiDGX agent

arXiv:2605.26645v1 Announce Type: new Abstract: LLM-based knowledge-graph question answering (KGQA) delegates graph traversal to language models, turning each question into a sequence of local relatio

Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models

ResearchDGX agent

arXiv:2605.27311v1 Announce Type: new Abstract: Chart question-answering (QA) benchmarks aim to pose questions that require visual reasoning to correctly answer, but models can often reach solutions t

Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze Surprisal

ResearchDGX agent

arXiv:2601.09886v2 Announce Type: replace Abstract: How predictable a word is can be quantified in two ways: using human responses to the cloze task or using probabilities from language models (LMs).W

Conceptual Steganography

ResearchDGX agent

arXiv:2605.26537v1 Announce Type: new Abstract: Language Models (LMs) emit Chains-of-Thought (CoTs) that drive much of their capability. However, the same sequence that carries useful reasoning can al

Conv-to-Bench: Evaluating Language Models Via User-Assistant Dialogues In Code Tasks

SafetyDGX agent

arXiv:2605.26440v1 Announce Type: new Abstract: The rapid advancement of Large Language Models (LLMs) has outpaced the scalability of traditional evaluation benchmarks, which remain heavily dependent

Cultural Value Alignment Via Latent Activation Steering in Large Language Models

SafetyDGX agent

arXiv:2605.26365v1 Announce Type: new Abstract: Large Language Models (LLMs) often exhibit homogenized cultural perspectives. While the World Values Survey (WVS) provides a gold standard for mapping h

Curation and Extraction of Drug-Related Entities from Reddit Platform

Model ReleasesDGX agent

arXiv:2605.26445v1 Announce Type: new Abstract: Physicians learn primarily about illicit drugs from clinical overdose cases, limiting their understanding of real-world usage. Meanwhile, drug users sha

Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers

TutorialsDGX agent

arXiv:2601.20796v2 Announce Type: replace Abstract: Transformer-based multimodal large language models often exhibit in-context learning (ICL) abilities. Motivated by this phenomenon, we ask: how do t

DunbaaBERT: From Sacrifice to Semantics

Model ReleasesDGX agent

arXiv:2605.26935v1 Announce Type: new Abstract: Large language models have achieved strong performance across many NLP tasks, yet Urdu remains comparatively underexplored due to limited resources and

Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancement

SafetyDGX agent

arXiv:2605.26952v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) has proven effective for training LLM-based agents with external tool-use capabilities. However, we identify that ag

Energy-Gated Attention and Wavelet Positional Encoding: Complementary Inductive Biases for Transformer Attention

ResearchDGX agent

arXiv:2605.26355v1 Announce Type: cross Abstract: Standard transformer attention computes pairwise token similarity but treats all tokens as equally salient and all positions as equally local, regardl

ENPMR-Bench: Benchmarking Proactive Memory Retrieval for Emotional Support Agents

Model ReleasesDGX agent

arXiv:2605.27240v1 Announce Type: new Abstract: Memory-augmented language agents are increasingly deployed in affective applications such as emotional support, where understanding and responding to us

Entropy Sentinel: Continuous LLM Accuracy Monitoring from Decoding Entropy Traces in STEM

Model ReleasesDGX agent

arXiv:2601.09001v4 Announce Type: replace Abstract: Deploying LLMs raises two coupled challenges: (1) monitoring--estimating where a model underperforms as traffic and domains drift--and (2) improveme

EpiCurveBench: Evaluating VLMs on Epidemic Curve Digitization

Model ReleasesDGX agent

arXiv:2605.27195v1 Announce Type: new Abstract: Chart-to-data extraction with vision-language models (VLMs) is increasingly evaluated on benchmarks that show diminishing headroom (frontier VLMs exceed

Evi-Steer: Learning to Steer Biomedical Vision-Language Models through Efficient and Generalizable Evidential Tuning

Model ReleasesDGX agent

arXiv:2605.26292v1 Announce Type: cross Abstract: Parameter-efficient adaptation of vision-language foundation models is crucial for precise multimodal understanding of biomedical images, yet existing

Evidence Absence Is Not Evidence Insufficiency: Diagnosing NEI Construction Artifacts in Fact Verification

ResearchDGX agent

arXiv:2605.26663v1 Announce Type: new Abstract: Evidence absence is not evidence insufficiency, but fact verification benchmarks can make them observationally similar. The Not Enough Information (NEI)

ExTax: Explainable Disinformation Detection via Persuasion, Emotion, and Narrative Role Taxonomies

ResearchDGX agent

arXiv:2605.27045v1 Announce Type: new Abstract: The democratization of LLMs has accelerated the generation and circulation of highly fluent disinformation, making traditional syntax-semantic verificat

FAB-Bench: A Framework for Adaptive RAG Benchmarking in Semiconductor Manufacturing

Model ReleasesDGX agent

arXiv:2605.26476v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has become critical for knowledge-intensive applications, yet evaluating its performance in vertical domains remain

FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions

SafetyDGX agent

arXiv:2605.27062v1 Announce Type: new Abstract: State-of-the-art performance for Automatic Speech Recognition (ASR) largely depends on the availability of large-scale labeled corpora. This creates a d

FinHarness: An Inline Lifecycle Safety Harness for Finance LLM Agents

SafetyDGX agent

arXiv:2605.27333v1 Announce Type: new Abstract: Finance LLM agents must simultaneously block prompt-induced unauthorized actions and approve legitimate multi-step business workflows. However, boundary

Formalization of Malagasy conjugation

ResearchDGX agent

arXiv:2605.27161v1 Announce Type: new Abstract: This paper reports the core linguistic work performed to construct a dictionary-based morphological analyser for Malagasy simple verbs. It uses the Unit

From Snippets to Semantics: Rethinking Evidence Granularity for Multilingual Fact Verification

ResearchDGX agent

arXiv:2605.26755v1 Announce Type: new Abstract: Multilingual fact verification requires evidence that is both relevant and sufficiently complete for reliable factuality prediction. However, existing s

Generating Logically Consistent Synthetic Supply Chain Data with LLM-Driven Knowledge Graph Reasoning

TutorialsDGX agent

arXiv:2605.26823v1 Announce Type: new Abstract: Synthetic data offers a promising solution to two persistent barriers in supply chain analytics: data scarcity and data privacy. However, for synthetic

← Previous
1…5758596061…129
Next →