AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Hardware

SpenseGPT: Practical One-shot Pruning Enabling Sparse and Dense GEMMs for LLM Inference

DGX agent

arXiv:2606.10445v1 Announce Type: cross Abstract: Semi-structured 2:4 sparsity is widely supported by modern accelerators, providing up to a 2x theoretical speedup. However, its strict 50% sparsity co

hardwarearxiv-cs-cl
10 Jun 2026
Safety
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Standard Language Ideology in AI-Generated Language

DGX agent

arXiv:2406.08726v3 Announce Type: replace Abstract: Large language models (LLMs) generate text that reinforces standard language ideology: a bias towards certain language varieties that are granted mo

safetyarxiv-cs-cl
10 Jun 2026
Model Releases

Streaming Knowledge Compilation: Proactive Materiality-Scored Pinning for Time-Evolving LLM Wikis

DGX agent

arXiv:2606.09877v1 Announce Type: cross Abstract: LLM wiki systems compile knowledge into pre-filled KV caches for efficient inference, but assume a static corpus -- an assumption that fails whenever

model-releasesarxiv-cs-cl
10 Jun 2026
Applications

Swivuriso: The South African Next Voices Multilingual Speech Dataset

DGX agent

arXiv:2512.02201v3 Announce Type: replace Abstract: This paper introduces Swivuriso, a 3000-hour multilingual speech dataset developed as part of the African Next Voices project, to support the develo

applicationsarxiv-cs-cl
10 Jun 2026
Agents

TabClaw: An Interactive and Self-Evolving Agent for Spreadsheet Manipulation and Table Reasoning

DGX agent

arXiv:2606.10316v1 Announce Type: new Abstract: Spreadsheets and tables are widely used representations for structured data analysis, but effective analysis still requires substantial manual effort an

agentsarxiv-cs-cl
10 Jun 2026
Model Releases

The Order Matters: Sequential Fine-Tuning of LLaMA for Coherent Automated Essay Scoring

DGX agent

arXiv:2606.10327v1 Announce Type: new Abstract: Automated Essay Scoring (AES) systems must judge interdependent discourse elements (e.g., lead, claim, evidence, conclusion), yet most approaches treat

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

The Shibboleth Effect: Auditing the Cross-Lingual Distributional Skew of Large Language Models

DGX agent

arXiv:2606.11082v1 Announce Type: new Abstract: This study investigates cross-lingual distributional skew (the Shibboleth Effect) in frontier large language models (LLMs) subjected to sustained advers

model-releasesarxiv-cs-cl
10 Jun 2026
Agents

Trace Only What You Need: Structure-Aware On-Demand Hypergraph Memory for Long-Document Question Answering

DGX agent

arXiv:2606.10921v1 Announce Type: new Abstract: Long-document question answering (QA) requires large language models (LLMs) to reason over evidence scattered across lengthy documents, where answers of

agentsarxiv-cs-cl
10 Jun 2026
Model Releases

Training LLMs to Enforce Multi-Level Instruction Hierarchies via Gravity-Weighted Direct Preference Optimization

DGX agent

arXiv:2606.10860v1 Announce Type: cross Abstract: Production LLMs receive instructions from sources with very different levels of trust, yet attend to every token with uniform architectural privilege.

model-releasesarxiv-cs-cl
10 Jun 2026
Research

UniSVQ: 2-bit Unified Scalar-Vector Quantization

DGX agent

arXiv:2606.10520v1 Announce Type: new Abstract: Post-training quantization at the 2-bit level enables low-cost deployment and inference acceleration for large language models (LLMs). Scalar quantizati

researcharxiv-cs-cl
10 Jun 2026
Model Releases

UXBench: Benchmarking User Experience in AI Assistants

DGX agent

arXiv:2606.09570v2 Announce Type: replace Abstract: As AI assistants serve millions of users daily, evaluating user experience (UX) beyond general model capability has become increasingly important. W

model-releasesarxiv-cs-cl
10 Jun 2026
Agents

VISTA: A Versatile Interactive User Simulation Toolkit for Agent Evaluation

DGX agent

arXiv:2606.11079v1 Announce Type: new Abstract: Evaluation remains a critical bottleneck for interactive agent development. Existing evaluation methods often rely on static benchmarks, which fail to c

agentsarxiv-cs-cl
10 Jun 2026
Model Releases

WebChallenger: A Reliable and Efficient Generalist Web Agent

DGX agent

arXiv:2606.10423v1 Announce Type: new Abstract: Autonomous web navigation remains challenging for LLM agents, and the strongest generalist systems rely on proprietary reasoning models whose inference

model-releasesarxiv-cs-cl
10 Jun 2026
Research

What Really Matters for Table LLMs? A Meta-Evaluation of Model and Data Effects

DGX agent

arXiv:2501.14717v2 Announce Type: replace Abstract: Table modeling has progressed for decades. In this work, we revisit this trajectory and highlight emerging challenges in the LLM era, particularly t

researcharxiv-cs-cl
10 Jun 2026
Safety

What Should a Skill Remember? Quality--Cost Trade-offs in Cost-Aware Skill Rewriting for Language Model Agents

DGX agent

arXiv:2606.09421v2 Announce Type: replace Abstract: Large language model agents increasingly rely on skills: reusable procedural documents encoding workflows, tool use, implementation patterns, valida

safetyarxiv-cs-cl
10 Jun 2026
Research

When Metrics Disagree: A Meta-Analysis of Knowledge-Graph-Completion Model Benchmarking

DGX agent

arXiv:2606.10287v1 Announce Type: cross Abstract: Evaluating Knowledge Graph Completion (KGC) models remains challenging because standard assessment relies on isolated rank-based metrics such as MRR,

researcharxiv-cs-cl
10 Jun 2026
Research

Where You Inject Diversity Matters: A Unified Framework for Diverse Generation

DGX agent

arXiv:2606.10302v1 Announce Type: new Abstract: Open-ended generation tasks often require a set of meaningfully different outputs, yet large language models often produce similar generations. Existing

researcharxiv-cs-cl
10 Jun 2026
Research

Which LoRA? An Empirical Study on the Effectiveness of LoRA Techniques During Multilingual Instruction Tuning

DGX agent

arXiv:2606.10428v1 Announce Type: new Abstract: We investigate whether commonly available LoRA variants have an advantage over basic LoRA in multilingual instruction tuning. Experiments involving LoRA

researcharxiv-cs-cl
10 Jun 2026
Model Releases

Who Brought Easter Eggs to Eid? Auditing Cultural Translation of Math Word Problems Across Diverse Languages and Regions

DGX agent

arXiv:2606.11009v1 Announce Type: new Abstract: Large language models are increasingly used to adapt math word problems for personalized learning at scale, but it remains an open question whether thos

model-releasesarxiv-cs-cl
10 Jun 2026
Research

Who Wrote the Book? Detecting and Attributing LLM Ghostwriters

DGX agent

arXiv:2603.28054v2 Announce Type: replace Abstract: In this paper, we introduce GhostWriteBench, a dataset for LLM authorship attribution. It comprises long-form texts (50K+ words per book) generated

researcharxiv-cs-cl
10 Jun 2026
Applications

A Dynamic Self-Evolving Extraction System

DGX agent

arXiv:2603.06915v2 Announce Type: replace Abstract: The extraction of structured information from raw text is a fundamental component of many NLP applications, including document retrieval, ranking, a

applicationsarxiv-cs-cl
8 Jun 2026
Model Releases

A Four-Condition Diagnostic Protocol for Evidence Utilization in Long-Context and Retrieval-Augmented Language Models

DGX agent

arXiv:2606.06758v1 Announce Type: new Abstract: Final-answer accuracy, retrieval recall, and citation overlap do not by themselves identify whether a long-context or retrieval-augmented language model

model-releasesarxiv-cs-cl
8 Jun 2026
Safety

AdaJudge: Adaptive Multi-Perspective Judging for Reward Modeling

DGX agent

arXiv:2601.08097v2 Announce Type: replace Abstract: Reward modeling is essential for aligning large language models with human preferences, yet predominant architectures rely on a static pooling strat

safetyarxiv-cs-cl
8 Jun 2026
Applications

Adversarial Creation and Detection of AI-Generated Social Bot Content

DGX agent

arXiv:2606.07219v1 Announce Type: new Abstract: The convergence of large language models and social bots allows malicious actors to manipulate the information ecosystem by generating human-like conten

applicationsarxiv-cs-cl
8 Jun 2026
Agents

Agentopia: Long-Term Life Simulation and Learning in Agent Societies

DGX agent

arXiv:2606.07513v1 Announce Type: new Abstract: Humans learn from social life. Simulating this process with LLM-powered agents represents a promising research direction, raising a natural question: wh

agentsarxiv-cs-cl
8 Jun 2026
Research

An Expanded Synthetic Conversation Dataset for Multi-Turn Smishing Detection

DGX agent

arXiv:2606.06879v1 Announce Type: new Abstract: Our prior work introduced COVA, a synthetically generated multi-turn conversational smishing dataset of 3,201 labeled conversations, establishing baseli

researcharxiv-cs-cl
8 Jun 2026
Research

Are Large Language Models Suitable for Graph Computation? Progress and Prospects

DGX agent

arXiv:2606.06865v1 Announce Type: new Abstract: Large language models (LLMs) have been increasingly explored for graph computation, where tasks require reasoning over structured relationships and algo

researcharxiv-cs-cl
8 Jun 2026
Agents

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning

DGX agent

arXiv:2512.13278v2 Announce Type: replace Abstract: Agentic reinforcement learning has advanced large language models (LLMs) to reason through long chain-of-thought trajectories while interleaving ext

agentsarxiv-cs-cl
8 Jun 2026
Model Releases

Beyond Rubrics: Exploration-Guided Evaluation Skills for Reward Modeling

DGX agent

arXiv:2606.07040v1 Announce Type: new Abstract: Open-ended reward modeling requires judges that can follow subtle, domain-specific preferences when verifiable answers are unavailable. Existing rubric-

model-releasesarxiv-cs-cl
8 Jun 2026
Research

Contrastive Training with LLM-generated Near-Misses for Robust Code-Switching Speech Recognition

DGX agent

arXiv:2606.06985v1 Announce Type: new Abstract: Code-switching (CS), the alternation between multiple languages within a single utterance, remains challenging for Automatic Speech Recognition (ASR). T

researcharxiv-cs-cl
8 Jun 2026
Research

CRAFT: A Unified Counterfactual Reasoning Framework for Tabular Question Answering and Fact Verification

DGX agent

arXiv:2606.06842v1 Announce Type: new Abstract: Table reasoning remains challenging for large language models (LLMs), particularly in tasks that require multi-step inference over long and structured t

researcharxiv-cs-cl
8 Jun 2026
Model Releases

Creation of the Estonian Subjectivity Dataset: Assessing the Degree of Subjectivity on a Scale

DGX agent

arXiv:2512.09634v2 Announce Type: replace Abstract: This article presents the creation of an Estonian-language dataset for document-level subjectivity, analyzes the resulting annotations, and reports

model-releasesarxiv-cs-cl
8 Jun 2026
Research

DialDefer: A Framework for Detecting and Mitigating LLM Dialogic Deference

DGX agent

arXiv:2601.10896v2 Announce Type: replace Abstract: LLMs are increasingly used as third-party judges, yet their reliability when evaluating speakers in dialogue remains poorly understood. We show that

researcharxiv-cs-cl
8 Jun 2026
Research

DirectAudioEdit: Inversion-Free Text-Guided Audio Editing via Diffusion Prediction Contrast

DGX agent

arXiv:2606.07356v1 Announce Type: cross Abstract: Text-guided audio editing aims to modify the language-specified acoustic content while preserving edit-irrelevant source components. Existing training

researcharxiv-cs-cl
8 Jun 2026
Model Releases

Explain Like I'm 5 or Whatever I Choose: Evaluating the Interactive Potential of Language Model Responses

DGX agent

arXiv:2606.06788v1 Announce Type: new Abstract: Evaluations of large language models (LLMs) in scientific information seeking tasks have become increasingly use-centric, such as conducting live or mul

model-releasesarxiv-cs-cl
8 Jun 2026
Safety

Explicit Evidence Grounding via Structured Inline Citation Generation

DGX agent

arXiv:2606.07130v1 Announce Type: new Abstract: As AI systems become more widely adopted, the demand for factual and faithful generation grows. Properly attributing information through citations becom

safetyarxiv-cs-cl
8 Jun 2026
Local Ai

From Correctness to Utility: Gain-Based Prefix Evaluation for LLM Reasoning

DGX agent

arXiv:2606.07190v1 Announce Type: new Abstract: Reasoning prefixes shape the future trajectory of LLM problem solving, yet existing process reward models usually evaluate them through local step corre

local-aiarxiv-cs-cl
8 Jun 2026
Tutorials

Geometry of Semantic Space: Comparative Study of Discrete and Continuous Models

DGX agent

arXiv:2606.07183v1 Announce Type: new Abstract: This work examines the semantic geometry underlying NLP models. We compare supervised vector embeddings, such as CamemBERT, with lexical co-occurrence g

tutorialsarxiv-cs-cl
8 Jun 2026
Research

HKVM-RAG: Key-Value-Separated Hypergraph Evidence Organization for Multi-Hop RAG

DGX agent

arXiv:2606.07218v1 Announce Type: cross Abstract: Multi-hop RAG poses a data-engineering problem beyond passage matching: under fixed retrieval budgets, a system must organize retrieved text into evid

researcharxiv-cs-cl
8 Jun 2026
Model Releases

Improving Cross-Lingual Factual Recall via Consistency-Driven Reinforcement Learning

DGX agent

arXiv:2606.06586v1 Announce Type: new Abstract: Large language models (LLMs) trained predominantly on English data encode substantial world knowledge, yet often fail to express it reliably in other la

model-releasesarxiv-cs-cl
8 Jun 2026
Safety

Interpreting Brain Responses to Language with Sparse Features from Language Models

DGX agent

arXiv:2606.06857v1 Announce Type: new Abstract: A central goal of cognitive neuroscience is to characterize the features that are represented by human language cortex. Artificial language models (LMs)

safetyarxiv-cs-cl
8 Jun 2026
Research

KIT's Submission to Cross-Lingual Voice Cloning in IWSLT 2026

DGX agent

arXiv:2606.07240v1 Announce Type: new Abstract: Cross-lingual voice cloning aims to generate speech in a target language while preserving speaker identity from a source-language reference. This task i

researcharxiv-cs-cl
8 Jun 2026
Safety

Korean Culture into LLM Alignment: Toward Cultural Coherence

DGX agent

arXiv:2606.06797v1 Announce Type: new Abstract: Cultural-aspect work on large language models is dominated by a negative target: which outputs to suppress. We argue that a constructive counterpart is

safetyarxiv-cs-cl
8 Jun 2026
Model Releases

Learning Perspectivist Social Meaning via Demographic-Conditioned Fusion Embeddings

DGX agent

arXiv:2606.07123v1 Announce Type: new Abstract: Social meaning in language is inherently perspectival, varying across annotator backgrounds, demographics, and ideological positions. However, most NLP

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

LLM-Guided Evolution for Medical Decision Pipelines

DGX agent

arXiv:2606.07342v1 Announce Type: new Abstract: Adapting large language models (LLMs) to clinical workflows often requires costly fine-tuning or manual prompt and pipeline engineering. We study LLM-gu

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

M^3Exam: Benchmarking Multimodal Memory for Realistic User-Agent Interactions

DGX agent

arXiv:2606.07402v1 Announce Type: new Abstract: Language agents are increasingly deployed over accumulating multimodal information, yet existing benchmarks assume a human-human form with sparse visual

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

MADE: Beyond Scoring via a Multilingual Agentic Diagnosing Engine for Fine-Grained Evaluation Insights

DGX agent

arXiv:2606.07020v1 Announce Type: new Abstract: Multilingual and multicultural benchmarks now cover dozens of languages and model families, but the resulting score landscapes remain metric-rich and in

model-releasesarxiv-cs-cl
8 Jun 2026
Safety

MADRAG: Multi-Agent Debate with Retrieval-Augmented Generation for Training-Free Analytic Essay Scoring

DGX agent

arXiv:2606.06754v1 Announce Type: cross Abstract: We present MADRAG, a training-free framework for analytic essay scoring that combines multi-agent reasoning with retrieval-augmented grounding. Unlike

safetyarxiv-cs-cl
8 Jun 2026
← Previous
1…4950515253…161
Next →