AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
10 Jun 2026

SpenseGPT: Practical One-shot Pruning Enabling Sparse and Dense GEMMs for LLM Inference

HardwareDGX agent

arXiv:2606.10445v1 Announce Type: cross Abstract: Semi-structured 2:4 sparsity is widely supported by modern accelerators, providing up to a 2x theoretical speedup. However, its strict 50% sparsity co

Standard Language Ideology in AI-Generated Language

SafetyDGX agent

arXiv:2406.08726v3 Announce Type: replace Abstract: Large language models (LLMs) generate text that reinforces standard language ideology: a bias towards certain language varieties that are granted mo

Streaming Knowledge Compilation: Proactive Materiality-Scored Pinning for Time-Evolving LLM Wikis

Model ReleasesDGX agent

arXiv:2606.09877v1 Announce Type: cross Abstract: LLM wiki systems compile knowledge into pre-filled KV caches for efficient inference, but assume a static corpus -- an assumption that fails whenever


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Swivuriso: The South African Next Voices Multilingual Speech Dataset

ApplicationsDGX agent

arXiv:2512.02201v3 Announce Type: replace Abstract: This paper introduces Swivuriso, a 3000-hour multilingual speech dataset developed as part of the African Next Voices project, to support the develo

TabClaw: An Interactive and Self-Evolving Agent for Spreadsheet Manipulation and Table Reasoning

AgentsDGX agent

arXiv:2606.10316v1 Announce Type: new Abstract: Spreadsheets and tables are widely used representations for structured data analysis, but effective analysis still requires substantial manual effort an

The Order Matters: Sequential Fine-Tuning of LLaMA for Coherent Automated Essay Scoring

Model ReleasesDGX agent

arXiv:2606.10327v1 Announce Type: new Abstract: Automated Essay Scoring (AES) systems must judge interdependent discourse elements (e.g., lead, claim, evidence, conclusion), yet most approaches treat

The Shibboleth Effect: Auditing the Cross-Lingual Distributional Skew of Large Language Models

Model ReleasesDGX agent

arXiv:2606.11082v1 Announce Type: new Abstract: This study investigates cross-lingual distributional skew (the Shibboleth Effect) in frontier large language models (LLMs) subjected to sustained advers

Trace Only What You Need: Structure-Aware On-Demand Hypergraph Memory for Long-Document Question Answering

AgentsDGX agent

arXiv:2606.10921v1 Announce Type: new Abstract: Long-document question answering (QA) requires large language models (LLMs) to reason over evidence scattered across lengthy documents, where answers of

Training LLMs to Enforce Multi-Level Instruction Hierarchies via Gravity-Weighted Direct Preference Optimization

Model ReleasesDGX agent

arXiv:2606.10860v1 Announce Type: cross Abstract: Production LLMs receive instructions from sources with very different levels of trust, yet attend to every token with uniform architectural privilege.

UniSVQ: 2-bit Unified Scalar-Vector Quantization

ResearchDGX agent

arXiv:2606.10520v1 Announce Type: new Abstract: Post-training quantization at the 2-bit level enables low-cost deployment and inference acceleration for large language models (LLMs). Scalar quantizati

UXBench: Benchmarking User Experience in AI Assistants

Model ReleasesDGX agent

arXiv:2606.09570v2 Announce Type: replace Abstract: As AI assistants serve millions of users daily, evaluating user experience (UX) beyond general model capability has become increasingly important. W

VISTA: A Versatile Interactive User Simulation Toolkit for Agent Evaluation

AgentsDGX agent

arXiv:2606.11079v1 Announce Type: new Abstract: Evaluation remains a critical bottleneck for interactive agent development. Existing evaluation methods often rely on static benchmarks, which fail to c

WebChallenger: A Reliable and Efficient Generalist Web Agent

Model ReleasesDGX agent

arXiv:2606.10423v1 Announce Type: new Abstract: Autonomous web navigation remains challenging for LLM agents, and the strongest generalist systems rely on proprietary reasoning models whose inference

What Really Matters for Table LLMs? A Meta-Evaluation of Model and Data Effects

ResearchDGX agent

arXiv:2501.14717v2 Announce Type: replace Abstract: Table modeling has progressed for decades. In this work, we revisit this trajectory and highlight emerging challenges in the LLM era, particularly t

What Should a Skill Remember? Quality--Cost Trade-offs in Cost-Aware Skill Rewriting for Language Model Agents

SafetyDGX agent

arXiv:2606.09421v2 Announce Type: replace Abstract: Large language model agents increasingly rely on skills: reusable procedural documents encoding workflows, tool use, implementation patterns, valida

When Metrics Disagree: A Meta-Analysis of Knowledge-Graph-Completion Model Benchmarking

ResearchDGX agent

arXiv:2606.10287v1 Announce Type: cross Abstract: Evaluating Knowledge Graph Completion (KGC) models remains challenging because standard assessment relies on isolated rank-based metrics such as MRR,

Where You Inject Diversity Matters: A Unified Framework for Diverse Generation

ResearchDGX agent

arXiv:2606.10302v1 Announce Type: new Abstract: Open-ended generation tasks often require a set of meaningfully different outputs, yet large language models often produce similar generations. Existing

Which LoRA? An Empirical Study on the Effectiveness of LoRA Techniques During Multilingual Instruction Tuning

ResearchDGX agent

arXiv:2606.10428v1 Announce Type: new Abstract: We investigate whether commonly available LoRA variants have an advantage over basic LoRA in multilingual instruction tuning. Experiments involving LoRA

Who Brought Easter Eggs to Eid? Auditing Cultural Translation of Math Word Problems Across Diverse Languages and Regions

Model ReleasesDGX agent

arXiv:2606.11009v1 Announce Type: new Abstract: Large language models are increasingly used to adapt math word problems for personalized learning at scale, but it remains an open question whether thos

Who Wrote the Book? Detecting and Attributing LLM Ghostwriters

ResearchDGX agent

arXiv:2603.28054v2 Announce Type: replace Abstract: In this paper, we introduce GhostWriteBench, a dataset for LLM authorship attribution. It comprises long-form texts (50K+ words per book) generated

8 Jun 2026

A Dynamic Self-Evolving Extraction System

ApplicationsDGX agent

arXiv:2603.06915v2 Announce Type: replace Abstract: The extraction of structured information from raw text is a fundamental component of many NLP applications, including document retrieval, ranking, a

A Four-Condition Diagnostic Protocol for Evidence Utilization in Long-Context and Retrieval-Augmented Language Models

Model ReleasesDGX agent

arXiv:2606.06758v1 Announce Type: new Abstract: Final-answer accuracy, retrieval recall, and citation overlap do not by themselves identify whether a long-context or retrieval-augmented language model

AdaJudge: Adaptive Multi-Perspective Judging for Reward Modeling

SafetyDGX agent

arXiv:2601.08097v2 Announce Type: replace Abstract: Reward modeling is essential for aligning large language models with human preferences, yet predominant architectures rely on a static pooling strat

Adversarial Creation and Detection of AI-Generated Social Bot Content

ApplicationsDGX agent

arXiv:2606.07219v1 Announce Type: new Abstract: The convergence of large language models and social bots allows malicious actors to manipulate the information ecosystem by generating human-like conten

Agentopia: Long-Term Life Simulation and Learning in Agent Societies

AgentsDGX agent

arXiv:2606.07513v1 Announce Type: new Abstract: Humans learn from social life. Simulating this process with LLM-powered agents represents a promising research direction, raising a natural question: wh

An Expanded Synthetic Conversation Dataset for Multi-Turn Smishing Detection

ResearchDGX agent

arXiv:2606.06879v1 Announce Type: new Abstract: Our prior work introduced COVA, a synthetically generated multi-turn conversational smishing dataset of 3,201 labeled conversations, establishing baseli

Are Large Language Models Suitable for Graph Computation? Progress and Prospects

ResearchDGX agent

arXiv:2606.06865v1 Announce Type: new Abstract: Large language models (LLMs) have been increasingly explored for graph computation, where tasks require reasoning over structured relationships and algo

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning

AgentsDGX agent

arXiv:2512.13278v2 Announce Type: replace Abstract: Agentic reinforcement learning has advanced large language models (LLMs) to reason through long chain-of-thought trajectories while interleaving ext

Beyond Rubrics: Exploration-Guided Evaluation Skills for Reward Modeling

Model ReleasesDGX agent

arXiv:2606.07040v1 Announce Type: new Abstract: Open-ended reward modeling requires judges that can follow subtle, domain-specific preferences when verifiable answers are unavailable. Existing rubric-

Contrastive Training with LLM-generated Near-Misses for Robust Code-Switching Speech Recognition

ResearchDGX agent

arXiv:2606.06985v1 Announce Type: new Abstract: Code-switching (CS), the alternation between multiple languages within a single utterance, remains challenging for Automatic Speech Recognition (ASR). T

CRAFT: A Unified Counterfactual Reasoning Framework for Tabular Question Answering and Fact Verification

ResearchDGX agent

arXiv:2606.06842v1 Announce Type: new Abstract: Table reasoning remains challenging for large language models (LLMs), particularly in tasks that require multi-step inference over long and structured t

Creation of the Estonian Subjectivity Dataset: Assessing the Degree of Subjectivity on a Scale

Model ReleasesDGX agent

arXiv:2512.09634v2 Announce Type: replace Abstract: This article presents the creation of an Estonian-language dataset for document-level subjectivity, analyzes the resulting annotations, and reports

DialDefer: A Framework for Detecting and Mitigating LLM Dialogic Deference

ResearchDGX agent

arXiv:2601.10896v2 Announce Type: replace Abstract: LLMs are increasingly used as third-party judges, yet their reliability when evaluating speakers in dialogue remains poorly understood. We show that

DirectAudioEdit: Inversion-Free Text-Guided Audio Editing via Diffusion Prediction Contrast

ResearchDGX agent

arXiv:2606.07356v1 Announce Type: cross Abstract: Text-guided audio editing aims to modify the language-specified acoustic content while preserving edit-irrelevant source components. Existing training

Explain Like I'm 5 or Whatever I Choose: Evaluating the Interactive Potential of Language Model Responses

Model ReleasesDGX agent

arXiv:2606.06788v1 Announce Type: new Abstract: Evaluations of large language models (LLMs) in scientific information seeking tasks have become increasingly use-centric, such as conducting live or mul

Explicit Evidence Grounding via Structured Inline Citation Generation

SafetyDGX agent

arXiv:2606.07130v1 Announce Type: new Abstract: As AI systems become more widely adopted, the demand for factual and faithful generation grows. Properly attributing information through citations becom

From Correctness to Utility: Gain-Based Prefix Evaluation for LLM Reasoning

Local AiDGX agent

arXiv:2606.07190v1 Announce Type: new Abstract: Reasoning prefixes shape the future trajectory of LLM problem solving, yet existing process reward models usually evaluate them through local step corre

Geometry of Semantic Space: Comparative Study of Discrete and Continuous Models

TutorialsDGX agent

arXiv:2606.07183v1 Announce Type: new Abstract: This work examines the semantic geometry underlying NLP models. We compare supervised vector embeddings, such as CamemBERT, with lexical co-occurrence g

HKVM-RAG: Key-Value-Separated Hypergraph Evidence Organization for Multi-Hop RAG

ResearchDGX agent

arXiv:2606.07218v1 Announce Type: cross Abstract: Multi-hop RAG poses a data-engineering problem beyond passage matching: under fixed retrieval budgets, a system must organize retrieved text into evid

Improving Cross-Lingual Factual Recall via Consistency-Driven Reinforcement Learning

Model ReleasesDGX agent

arXiv:2606.06586v1 Announce Type: new Abstract: Large language models (LLMs) trained predominantly on English data encode substantial world knowledge, yet often fail to express it reliably in other la

Interpreting Brain Responses to Language with Sparse Features from Language Models

SafetyDGX agent

arXiv:2606.06857v1 Announce Type: new Abstract: A central goal of cognitive neuroscience is to characterize the features that are represented by human language cortex. Artificial language models (LMs)

KIT's Submission to Cross-Lingual Voice Cloning in IWSLT 2026

ResearchDGX agent

arXiv:2606.07240v1 Announce Type: new Abstract: Cross-lingual voice cloning aims to generate speech in a target language while preserving speaker identity from a source-language reference. This task i

Korean Culture into LLM Alignment: Toward Cultural Coherence

SafetyDGX agent

arXiv:2606.06797v1 Announce Type: new Abstract: Cultural-aspect work on large language models is dominated by a negative target: which outputs to suppress. We argue that a constructive counterpart is

Learning Perspectivist Social Meaning via Demographic-Conditioned Fusion Embeddings

Model ReleasesDGX agent

arXiv:2606.07123v1 Announce Type: new Abstract: Social meaning in language is inherently perspectival, varying across annotator backgrounds, demographics, and ideological positions. However, most NLP

LLM-Guided Evolution for Medical Decision Pipelines

Model ReleasesDGX agent

arXiv:2606.07342v1 Announce Type: new Abstract: Adapting large language models (LLMs) to clinical workflows often requires costly fine-tuning or manual prompt and pipeline engineering. We study LLM-gu

M^3Exam: Benchmarking Multimodal Memory for Realistic User-Agent Interactions

Model ReleasesDGX agent

arXiv:2606.07402v1 Announce Type: new Abstract: Language agents are increasingly deployed over accumulating multimodal information, yet existing benchmarks assume a human-human form with sparse visual

MADE: Beyond Scoring via a Multilingual Agentic Diagnosing Engine for Fine-Grained Evaluation Insights

Model ReleasesDGX agent

arXiv:2606.07020v1 Announce Type: new Abstract: Multilingual and multicultural benchmarks now cover dozens of languages and model families, but the resulting score landscapes remain metric-rich and in

MADRAG: Multi-Agent Debate with Retrieval-Augmented Generation for Training-Free Analytic Essay Scoring

SafetyDGX agent

arXiv:2606.06754v1 Announce Type: cross Abstract: We present MADRAG, a training-free framework for analytic essay scoring that combines multi-agent reasoning with retrieval-augmented grounding. Unlike

MAGE: All-[MASK] Block Already Knows Where to Look in Block Diffusion LLM

ResearchDGX agent

arXiv:2602.14209v2 Announce Type: replace-cross Abstract: Block diffusion LLMs are an emerging paradigm for parallel language generation, but their KV caching makes memory access the dominant bottlene

Meaning in Order, Order in Meaning: Semantic R-precision for Keyphrase Evaluation

ResearchDGX agent

arXiv:2606.07057v1 Announce Type: cross Abstract: Evaluating the quality of automatically generated keyphrases remains a complex challenge. Traditional metrics either rely on exact lexical matching or

Mechanistic Evidence for Faithfulness Decay in Chain-of-Thought Reasoning

TutorialsDGX agent

arXiv:2602.11201v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) explanations are widely used to interpret how language models solve complex problems, yet it remains unclear whether these st

Mining Useful General Data for Low-Resource Domain Adaptation

SafetyDGX agent

arXiv:2511.07380v2 Announce Type: replace Abstract: Adapting large language models (LLMs) to low-resource domains remains challenging due to the scarcity of domain-specific data. While in-domain data

MMAE: A Massive Multitask Audio Editing Benchmark

Model ReleasesDGX agent

arXiv:2606.07229v1 Announce Type: cross Abstract: We introduce MMAE, a Massive Multitask Audio Editing benchmark, serving as the first comprehensive evaluation testbed designed for general-purpose ins

mmPISA-bench: Do LLMs Reason Equally Well Across 43 Languages?

Model ReleasesDGX agent

arXiv:2606.07069v1 Announce Type: new Abstract: We introduce mmPISA-bench, a compact high-quality multilingual reasoning benchmark derived from the OECD Programme for International Student Assessment

Modeling semantic association in self-paced reading with language model embeddings

ResearchDGX agent

arXiv:2606.07066v1 Announce Type: new Abstract: Semantic association between a word and its context has been identified as an important component of reading comprehension, even when word predictabilit

Modular Monolingual Adaptation using Pretrained Language Models

ResearchDGX agent

arXiv:2606.06738v1 Announce Type: new Abstract: Building monolingual language models (LMs) for low-resource languages typically relies on adapting pretrained language models (PLMs) by finetuning the w

Multiscale POD of Transformer Attention Fields: Scale-Selective Analysis via Morlet Scalogram

ResearchDGX agent

arXiv:2606.06573v1 Announce Type: cross Abstract: We introduce scale-selective Proper Orthogonal Decomposition (POD) for transformer attention fields, inspired by the use of POD for extracting energet

Phun-Bench: Evaluating LLMs on Phonological Understanding in Chinese

Model ReleasesDGX agent

arXiv:2606.07300v1 Announce Type: new Abstract: Language is a vehicle for thought, intricately tied to sounds, symbols, and meaning. However, most large language model (LLM) research focuses on meanin

PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration

ResearchDGX agent

arXiv:2502.00527v2 Announce Type: replace-cross Abstract: The KV cache in large language models is a dominant factor in memory usage, limiting their broader applicability. Quantizing the cache to lowe

Principles of Concept Representation in Sentence Encoders

Model ReleasesDGX agent

arXiv:2606.06994v1 Announce Type: new Abstract: What makes a sentence encoder produce good concept representations? We approach this through the lens of representational compositionality: an encoder s

← Previous
1…3940414243…129
Next →