AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Model Releases

TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models

DGX agent

arXiv:2504.20605v2 Announce Type: replace Abstract: Moral stories are a time-tested vehicle for transmitting values, yet modern NLP lacks a large, structured corpus that couples coherent narratives wi

model-releasesarxiv-cs-cl
5 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

The 2026 ACII Dyadic Conversations (DaiKon) Workshop & Challenge

DGX agent

arXiv:2605.02672v1 Announce Type: cross Abstract: The 2026 ACII Dyadic Conversations (ACII-DaiKon) Workshop & Challenge introduces a benchmark for modeling interpersonal affect and social dynamics in

model-releasesarxiv-cs-cl
5 May 2026
Research

The Company You Keep: How LLMs Respond to Dark Triad Traits

DGX agent

arXiv:2603.04299v3 Announce Type: replace Abstract: Large Language Models (LLMs) often exhibit highly agreeable and reinforcing conversational styles, also known as AI-sycophancy. Although this behavi

researcharxiv-cs-cl
5 May 2026
Model Releases

The Compliance Gap: Why AI Systems Promise to Follow Process Instructions but Don't

DGX agent

arXiv:2605.01771v1 Announce Type: new Abstract: An auditor instructs an AI assistant: 'open each file individually using the Read tool -- no scripts, no agents.' The AI replies 'Yes' -- then issues a

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

The Compliance Trap: How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure

DGX agent

arXiv:2605.02398v1 Announce Type: cross Abstract: As frontier AI models are deployed in high-stakes decision pipelines, their ability to maintain metacognitive stability -- knowing what they do not kn

model-releasesarxiv-cs-cl
5 May 2026
Research

The Cylindrical Representation Hypothesis for Language Model Steering

DGX agent

arXiv:2605.01844v1 Announce Type: new Abstract: Steering is a widely used technique for controlling large language models, yet its effects are often unstable and hard to predict. Existing theoretical

researcharxiv-cs-cl
5 May 2026
Applications

The grip of grammar on meaning uncertainty: cross-linguistic evidence, neural correlates, and clinical relevance

DGX agent

arXiv:2605.01537v1 Announce Type: new Abstract: Isolated word meanings are inherently uncertain. This uncertainty reduces when they are combined and anchored in context. We propose that grammar compre

applicationsarxiv-cs-cl
5 May 2026
Research

The Pre-Training Study of Expanded-SPLADE Models on Web Document Titles

DGX agent

arXiv:2605.01407v1 Announce Type: cross Abstract: Masked Language Modeling (MLM) pre-training is one of the primary ways to initialize Neural Information Retrieval (IR) models prior to retrieval fine-

researcharxiv-cs-cl
5 May 2026
Agents

The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning

DGX agent

arXiv:2605.01704v1 Announce Type: new Abstract: When copies of the same language model are prompted to debate, they produce diverse phrasings of one perspective rather than diverse perspectives. Multi

agentsarxiv-cs-cl
5 May 2026
Research

Tibetan-TTS:Low-Resource Tibetan Speech Synthesis with Large Model Adaptation

DGX agent

arXiv:2605.02496v1 Announce Type: cross Abstract: Tibetan text-to-speech (TTS) has long been challenged by scarce speech resources, significant dialectal variation, and the complex mapping between wri

researcharxiv-cs-cl
5 May 2026
Model Releases

TIME: Temporally Intelligent Meta-reasoning Engine for Context-Triggered Explicit Reasoning

DGX agent

arXiv:2601.05300v2 Announce Type: replace-cross Abstract: Reasoning-oriented language models typically expose explicit reasoning as a long, front-loaded chain of 'thinking' tokens before the main outp

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

TokenTiming: A Dynamic Alignment Method for Universal Speculative Decoding Model Pairs

DGX agent

arXiv:2510.15545v4 Announce Type: replace Abstract: Accelerating the inference of large language models (LLMs) has been a critical challenge in generative AI. Speculative decoding (SD) substantially i

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Toward Culturally Grounded Natural Language Processing

DGX agent

arXiv:2603.26013v2 Announce Type: replace Abstract: Multilingual NLP is often treated as a route to global inclusion, but linguistic coverage and cultural competence frequently diverge. This paper syn

model-releasesarxiv-cs-cl
5 May 2026
Applications

VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation

DGX agent

arXiv:2602.21054v2 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) frequently hallucinate, limiting their safe deployment in real-world applications. Existing LLM self-eval

applicationsarxiv-cs-cl
5 May 2026
Agents

Verbal-R3: Verbal Reranker as the Missing Bridge between Retrieval and Reasoning

DGX agent

arXiv:2605.01399v1 Announce Type: new Abstract: The conventional Retrieval-Augmented Generation (RAG) paradigm of injecting raw retrieved texts into the Large Language Model (LLM)'s context often resu

agentsarxiv-cs-cl
5 May 2026
Local Ai

VERGE: Formal Refinement and Guidance Engine for Verifiable LLM Reasoning

DGX agent

arXiv:2601.20055v2 Announce Type: replace Abstract: Despite the syntactic fluency of Large Language Models (LLMs), ensuring their logical correctness in high-stakes domains remains a fundamental chall

local-aiarxiv-cs-cl
5 May 2026
Model Releases

VeRO: An Evaluation Harness for Agents to Optimize Agents

DGX agent

arXiv:2602.22480v2 Announce Type: replace-cross Abstract: An important emerging application of coding agents is agent optimization: the iterative improvement of a target agent through edit-execute-eva

model-releasesarxiv-cs-cl
5 May 2026
Safety

Virtual Speech Therapist: A Clinician-in-the-Loop AI Speech Therapy Agent for Personalized and Supervised Therapy

DGX agent

arXiv:2605.01101v1 Announce Type: cross Abstract: This paper develops Virtual Speech Therapist (VST), an intelligent agent-based platform that streamlines stuttering assessment and delivers customized

safetyarxiv-cs-cl
5 May 2026
Model Releases

Watermarking LLM Agent Trajectories

DGX agent

arXiv:2602.18700v2 Announce Type: replace-cross Abstract: LLM agents rely heavily on high-quality trajectory data to guide their problem-solving behaviors, yet producing such data requires substantial

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

What Single-Prompt Accuracy Misses: A Multi-Variant Reliability Audit of Language Models

DGX agent

arXiv:2605.02038v1 Announce Type: new Abstract: Single-prompt accuracy is the dominant way to benchmark language models, but it can miss reliability failures that matter. We evaluate a 15-model open-w

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition

DGX agent

arXiv:2605.02782v1 Announce Type: cross Abstract: Automatic speech recognition (ASR) systems remain brittle on dysarthric and other atypical speech. Recent audio-language models raise the possibility

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models

DGX agent

arXiv:2605.02363v1 Announce Type: new Abstract: Deployed language models must produce outputs that are both correct and format-compliant. We study this structured-output reliability gap using two math

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answering

DGX agent

arXiv:2601.19827v3 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) extends large language models (LLMs) beyond parametric knowledge, yet it is unclear when iterative retrieval-re

model-releasesarxiv-cs-cl
5 May 2026
Research

When Less is Enough: Efficient Inference via Collaborative Reasoning

DGX agent

arXiv:2605.01111v1 Announce Type: cross Abstract: In this work, we introduce DUET (Dual-model Efficient Two-stage inference), a collaborative inference framework in which a capable model and a lightwe

researcharxiv-cs-cl
5 May 2026
Research

Where Do Prompt Perturbations Break Generation? A Segment-Level View of Robustness in LoRA-Tuned Language Models

DGX agent

arXiv:2605.01605v1 Announce Type: new Abstract: Large language models are sensitive to minor prompt perturbations, yet existing robustness methods usually enforce consistency at the whole-sequence lev

researcharxiv-cs-cl
5 May 2026
Safety

Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework

DGX agent

arXiv:2605.01416v1 Announce Type: cross Abstract: The increasing scale and complexity of online platforms raises critical policy questions around harmful content, digital well-being, and user autonomy

safetyarxiv-cs-cl
5 May 2026
Hardware

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization

DGX agent

arXiv:2605.02262v1 Announce Type: cross Abstract: Recently, video language models (VLMs) have been applied in various fields. However, the visual token sequence of the VLM is too long, which may cause

hardwarearxiv-cs-cl
5 May 2026
Local Ai

Zero-Shot Confidence Estimation for Small LLMs: When Supervised Baselines Aren't Worth Training

DGX agent

arXiv:2605.02241v1 Announce Type: cross Abstract: How reliably can a small language model estimate its own correctness? The answer determines whether local-to-cloud routing-escalating queries a cheap

local-aiarxiv-cs-cl
5 May 2026
Model Releases

A11y-Compressor: A Framework for Enhancing the Efficiency of GUI Agent Observations through Visual Context Reconstruction and Redundancy Reduction

DGX agent

arXiv:2605.00551v1 Announce Type: new Abstract: AI agents that interact with graphical user interfaces (GUIs) require effective observation representations for reliable grounding. The accessibility tr

model-releasesarxiv-cs-cl
4 May 2026
Research

Adaptive Querying with AI Persona Priors

DGX agent

arXiv:2605.00696v1 Announce Type: cross Abstract: We study adaptive querying for learning user-dependent quantities of interest, such as responses to held-out items and psychometric indicators, within

researcharxiv-cs-cl
4 May 2026
Research

ADVICE: Answer-Dependent Verbalized Confidence Estimation

DGX agent

arXiv:2510.10913v3 Announce Type: replace Abstract: Recent progress in large language models (LLMs) has enabled them to communicate their confidence in natural language, improving transparency and rel

researcharxiv-cs-cl
4 May 2026
Safety

Agent Capsules: Quality-Gated Granularity Control for Multi-Agent LLM Pipelines

DGX agent

arXiv:2605.00410v1 Announce Type: new Abstract: A multi-agent pipeline with N agents typically issues N LLM calls per run. Merging agents into fewer calls (compound execution) promises token savings,

safetyarxiv-cs-cl
4 May 2026
Model Releases

AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go?

DGX agent

arXiv:2605.00334v1 Announce Type: cross Abstract: Production agentic systems make many model calls per user request, and most of those calls are short, structured, and routine. This raises a practical

model-releasesarxiv-cs-cl
4 May 2026
Agents

Agentic AI for Substance Use Education: Integrating Regulatory and Scientific Knowledge Sources

DGX agent

arXiv:2605.00383v1 Announce Type: new Abstract: The delivery of traditional substance education has remained problematic due to challenges in scalability, personalization, and the currency of informat

agentsarxiv-cs-cl
4 May 2026
Model Releases

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs

DGX agent

arXiv:2605.00539v1 Announce Type: new Abstract: Quantization is a key method for reducing the GPU memory requirement of training large language models (LLMs). Yet, current approaches are ineffective f

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

Alethia: A Foundational Encoder for Voice Deepfakes

DGX agent

arXiv:2605.00251v1 Announce Type: cross Abstract: Existing voice deepfake detection and localization models rely heavily on representations extracted from speech foundation models (SFMs). However, dow

model-releasesarxiv-cs-cl
4 May 2026
Applications

Are You the A-hole? A Fair, Multi-Perspective Ethical Reasoning Framework

DGX agent

arXiv:2605.00270v1 Announce Type: new Abstract: Standard methods for aggregating natural language judgments, such as majority voting, often fail to produce logically consistent results when applied to

applicationsarxiv-cs-cl
4 May 2026
Model Releases

BanglaSocialBench: A Benchmark for Evaluating Sociopragmatic and Cultural Alignment of LLMs in Bangladeshi Social Interaction

DGX agent

arXiv:2603.15949v3 Announce Type: replace Abstract: Large Language Models have demonstrated strong multilingual fluency, yet fluency alone does not guarantee socially appropriate language use. In high

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

DGX agent

arXiv:2605.00674v1 Announce Type: new Abstract: Large language models (LLMs) are becoming increasingly capable mathematical collaborators, but static benchmarks are no longer sufficient for evaluating

model-releasesarxiv-cs-cl
4 May 2026
Research

Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe

DGX agent

arXiv:2605.00607v1 Announce Type: new Abstract: Probing is widely used to study which features can be decoded from language model representations. However, the common decoding probe approach has two l

researcharxiv-cs-cl
4 May 2026
Safety

Bias in Large Language Models: Origin, Evaluation, and Mitigation

DGX agent

arXiv:2411.10915v2 Announce Type: replace Abstract: Large Language Models (LLMs) have revolutionized natural language processing, but their susceptibility to biases poses significant challenges. This

safetyarxiv-cs-cl
4 May 2026
Local Ai

Block-wise Codeword Embedding for Reliable Multi-bit Text Watermarking

DGX agent

arXiv:2605.00348v1 Announce Type: cross Abstract: Recent multi-bit watermarking methods for large language models (LLMs) prioritize capacity over reliability, often conflating decoding with detection.

local-aiarxiv-cs-cl
4 May 2026
Model Releases

Borrowed Geometry: Computational Reuse of Frozen Text-Pretrained Transformer Weights Across Modalities

DGX agent

arXiv:2605.00333v1 Announce Type: cross Abstract: Frozen Gemma 4 31B weights pretrained exclusively on text tokens, unmodified, transfer across modality boundaries through a thin trainable interface.

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

Bring Your Own Prompts: Use-Case-Specific Bias and Fairness Evaluation for LLMs

DGX agent

arXiv:2407.10853v5 Announce Type: replace Abstract: Bias and fairness risks in Large Language Models (LLMs) vary substantially across deployment contexts, yet existing approaches lack systematic guida

model-releasesarxiv-cs-cl
4 May 2026
Research

Budget-Aware Routing for Long Clinical Text

DGX agent

arXiv:2605.00336v1 Announce Type: new Abstract: A key challenge for large language models is token cost per query and overall deployment cost. Clinical inputs are long, heterogeneous, and often redund

researcharxiv-cs-cl
4 May 2026
Agents

Build, Judge, Optimize: A Blueprint for Continuous Improvement of Multi-Agent Consumer Assistants

DGX agent

arXiv:2603.03565v2 Announce Type: replace-cross Abstract: Conversational shopping assistants (CSAs) represent a compelling application of agentic AI, but moving from prototype to production reveals tw

agentsarxiv-cs-cl
4 May 2026
Model Releases

Can Coding Agents Reproduce Findings in Computational Materials Science?

DGX agent

arXiv:2605.00803v1 Announce Type: cross Abstract: Large language models are increasingly deployed as autonomous coding agents and have achieved remarkably strong performance on software engineering be

model-releasesarxiv-cs-cl
4 May 2026
Safety

Can Small Language Models Handle Context-Summarized Multi-Turn Customer-Service QA? A Synthetic Data-Driven Comparative Evaluation

DGX agent

arXiv:2602.00665v3 Announce Type: replace Abstract: Customer-service question answering (QA) systems increasingly rely on conversational language understanding. While Large Language Models (LLMs) achi

safetyarxiv-cs-cl
4 May 2026
← Previous
1…112113114115116…161
Next →