AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
12 Aug 2026

The Parser Already Knows: Lightweight Bias Correction in Constrained Decoding

SafetyDGX agent

arXiv:2608.10137v1 Announce Type: new Abstract: Grammar Constrained Decoding (GCD) forces Language Models (LMs) to produce syntactically valid outputs by masking out non-conforming tokens at each step

The Signal Rail: A Deterministic Motion Grammar for Communicating Conversational Agent State in Terminal Interfaces

AgentsDGX agent

arXiv:2608.10689v1 Announce Type: cross Abstract: Terminal interfaces to conversational agents report rich internal state (listening, thinking, executing tools, awaiting input, failing) almost entirel

UT-ACA: Uncertainty-Triggered Adaptive Context Allocation for Long-Context Inference

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2603.18446v2 Announce Type: replace Abstract: Long-context inference remains challenging for large language models due to attention dilution and out-of-distribution degradation. Context selectio

VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?

Model ReleasesDGX agent

arXiv:2608.10408v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world

VoxSumm: A Multilingual Corpus of Long-Form Spoken News for Joint Summarization and Translation

Model ReleasesDGX agent

arXiv:2608.10359v1 Announce Type: cross Abstract: As information increasingly traverses linguistic boundaries, users require concise cross-lingual representations of long-form content. Nevertheless, l

What Iterated Self-Feeding Probes of Language Models Measure, and a test that separates the construction from the model

AgentsDGX agent

arXiv:2608.10986v1 Announce Type: new Abstract: A growing class of methods probes a language model by feeding it its own output: self-consistency, iterated refinement, agentic loops. We ask what such

When Is a General Factor Distinguishable? Non-Proportionality, Stable Structure, and the Bifactor Decision

ResearchDGX agent

arXiv:2608.10731v1 Announce Type: cross Abstract: Whether an additional general dimension is necessary beyond correlated first-order factors is a property of the population covariance matrix, not of a

Who Gets Heeded? An Obligation-Level Audit of Responsiveness in EPA Rulemaking

ApplicationsDGX agent

arXiv:2608.10329v1 Announce Type: cross Abstract: Notice-and-comment rulemaking gives any affected party the same formal right to influence federal regulation, but formal access is not substantive cap

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction

ResearchDGX agent

arXiv:2608.10878v1 Announce Type: new Abstract: Accurate and responsive turn-taking is essential for spoken dialogue systems, which must distinguish in real time between user interruptions, backchanne

11 Aug 2026

Accurate but Natural? Diagnosing Grammatical and Idiomatic Gaps in Japanese EFL Writing

ApplicationsDGX agent

arXiv:2608.09289v1 Announce Type: new Abstract: Second language writing research distinguishes grammatical accuracy from native-like idiomaticity, yet automated writing evaluation often conflates thes

Accurate Ensembles, Fragile Narratives: Multi-Scale Stacking and a Fidelity Audit of LLM-Generated Explanations for Credit Risk

ResearchDGX agent

arXiv:2608.08126v1 Announce Type: cross Abstract: Credit scoring increasingly relies on models whose decision logic cannot be read off their parameters, in tension with supervisory expectations that a

An Agentic Generative Large Language Model for Treatment Planning of Colorectal Cancer

SafetyDGX agent

arXiv:2608.09142v1 Announce Type: new Abstract: Treatment planning in precision oncology requires synthesizing heterogeneous patient information with rapidly evolving clinical guidelines to ensure gui

AnchorFold: A Focus-Then-Fold Framework via Recursive Attention Propagation for Efficient Multi-Vector Visual Document Retrieval

ResearchDGX agent

arXiv:2608.08732v1 Announce Type: cross Abstract: Multi-vector vision-language retrievers enable fine-grained Visual Document Retrieval (VDR) through late interaction, but storing and scoring hundreds

APEX-VW: A Document-Level English-Spanish Post-Editing Dataset in the Healthcare Domain

Model ReleasesDGX agent

arXiv:2608.08059v1 Announce Type: new Abstract: Post-Editing (PE) of Machine Translation (MT) output often involves repeating the same lexical and terminological corrections across many segments, espe

AraSSM: A bidirectional state-space encoder for Arabic masked language modeling

HardwareDGX agent

arXiv:2608.08256v1 Announce Type: new Abstract: Pretrained Transformer encoders such as AraBERT, MARBERT, and CAMeLBERT have become the standard backbone for Arabic natural language understanding, but

Archer: Adaptive Reuse of Cached Hidden States for Efficient Rollback in Diffusion Language Models

ResearchDGX agent

arXiv:2608.08086v1 Announce Type: new Abstract: Diffusion language models (DLMs) iteratively refine a sequence, allowing earlier predictions to be revised as context evolves. This rollback capability

ATLAS: Agentic Taxonomy of Large-Scale Software Ecosystems

Model ReleasesDGX agent

arXiv:2606.21597v2 Announce Type: replace-cross Abstract: The open-source ecosystem on GitHub lacks a systematic hierarchical taxonomy of software repositories. GitHub Topics, the dominant organizatio

Beyond cognacy

SafetyDGX agent

arXiv:2507.03005v3 Announce Type: replace Abstract: Computational phylogenetics has become an established tool in historical linguistics, with many language families now analyzed using likelihood-base

Beyond Direct Identifiers: Probabilistic Privacy Risk Estimation for Privacy-Conscious LLM Query Delegation

Model ReleasesDGX agent

arXiv:2608.09140v1 Announce Type: cross Abstract: Recent work on protecting privacy during user-LLM interactions often focuses on direct, explicit identifiers: the personally-identifiable information

Beyond the Capability Boundary: Zeroth-Order Optimization for Self-Evolving LLM Agents

Model ReleasesDGX agent

arXiv:2608.09292v1 Announce Type: cross Abstract: Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. H

BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and Mitigation

Model ReleasesDGX agent

arXiv:2604.03159v2 Announce Type: replace-cross Abstract: Large language models with web search are increasingly used in scientific publishing agents, yet they produce BibTeX entries with pervasive fi

Commitment Before Realization: When Classifier-Free Guidance Becomes Unnecessary in Masked Diffusion Language Models

ResearchDGX agent

arXiv:2608.08082v1 Announce Type: new Abstract: Classifier-free guidance (CFG) is usually kept on throughout masked diffusion language model decoding, although its benefit varies across prompts and ov

Comparing British and American Audio Description of Movies

ResearchDGX agent

arXiv:2608.09792v1 Announce Type: new Abstract: Narrating the visual component of movies is known as audio description. It is a narrative technique designed to enable blind and visually impaired indiv

Consilience for Verifier-Free Test-Time Scaling

ApplicationsDGX agent

arXiv:2608.09898v1 Announce Type: new Abstract: Test-time scaling often uses an external verifier, such as compilers and test cases in coding or trained value functions in robotics applications, to ob

Conversation as Measurement in Clinical Encounters: Observable Phase Structure, Partially Observable Patient State

Model ReleasesDGX agent

arXiv:2608.08868v1 Announce Type: new Abstract: Many modern AI systems analyze conversational traces to infer aspects of human interaction and state, implicitly assuming that such information is recov

Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning

ResearchDGX agent

arXiv:2602.11149v2 Announce Type: replace Abstract: Supervised fine-tuning (SFT) on chain-of-thought data is an essential post-training step for reasoning language models. Standard machine learning in

Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness

Model ReleasesDGX agent

arXiv:2608.09900v1 Announce Type: new Abstract: Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably wa

Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching

HardwareDGX agent

arXiv:2608.09444v1 Announce Type: cross Abstract: A main promise of looped language models (LMs) is depth-adaptive inference. By iterating a block of shared layers a variable number of times, the mode

Detection of Self-Introductions in Legislative Testimony

ResearchDGX agent

arXiv:2608.07891v1 Announce Type: new Abstract: Self-introductions are common in legislative committee testimonies. Successfully detecting them and extracting the speaker's name is enormously helpful

DevIntent: How Much Does LLM-Generated Code Violate Developer Intent?

Model ReleasesDGX agent

arXiv:2608.07614v1 Announce Type: cross Abstract: Code generated by LLMs can violate a developer's implicit intentions when given an ambiguous prompt, yet standard benchmarks measure only whether code

Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models

SafetyDGX agent

arXiv:2601.03115v2 Announce Type: replace Abstract: Emotion is a central dimension of spoken communication, yet, we still lack a mechanistic account of how modern large audio-language models (LALMs) e

DS@GT ARC at Touche: Large Language Models for Retrieval-Augmented Debate

ResearchDGX agent

arXiv:2608.08143v1 Announce Type: cross Abstract: We extend the DS@GT ARC working-note submission to the Touche 2025 Retrieval-Augmented Debate task. The task has two subtasks: generating the next utt

ELICITED: EHR-grounded Longitudinal Interactive Conversations for Information-seeking Triage Evaluation and Decision-making

Model ReleasesDGX agent

arXiv:2608.09024v1 Announce Type: new Abstract: Emergency-department (ED) triage requires clinicians to rapidly identify patients who need immediate attention, determine who can safely wait, and prior

Embedding Initialization for Unseen Low-resource Languages in Multilingual NMT: A Case Study on Limbum-English Translation

ApplicationsDGX agent

arXiv:2608.07629v1 Announce Type: new Abstract: Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu la

EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models

Model ReleasesDGX agent

arXiv:2608.09189v1 Announce Type: new Abstract: Despite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Model

EvalConvoLearn: An Open-Source Framework for Evaluating Grounded Learner Simulations in Tutoring Conversations

ResearchDGX agent

arXiv:2608.07497v1 Announce Type: cross Abstract: Conversational learner simulations are valuable tools for testing learning theories, evaluating instructional materials and automated tutors, or power

Evaluating Dedicated Monolingual and Joint Multilingual Causal Models for Dravidian Languages

ResearchDGX agent

arXiv:2608.07727v1 Announce Type: new Abstract: Dravidian languages, mainly Tamil, Telugu, Kannada, and Malayalam make up only a small part of the data used to train multilingual language models, so i

Evidence-Calibrated Runtime Reconstruction for Agent Skills Across Heterogeneous Coding Agents

Local AiDGX agent

arXiv:2608.08793v1 Announce Type: new Abstract: Agent Skills package reusable instructions and assets for tool-using language-model agents. Progressive loading creates failure boundaries poorly repres

Evo-Bench: Can Language Models Improve Agent Harness?

Model ReleasesDGX agent

arXiv:2608.09096v1 Announce Type: new Abstract: Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emergi

EvoTrustRAG: Evolution-Aware Conflict Attribution and Evidence Handling for Reliable Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2608.07933v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) improves the factuality of large language models with external knowledge, yet conflicting evidence remains a fund

Explicit Boundary Markers for Subword Vocabularies

ResearchDGX agent

arXiv:2608.08847v1 Announce Type: new Abstract: Subword tokenizers represent many common words twice in space-using writing systems, once with a leading space and once without. The two entries have se

Failure-Aware Long-Form Translation: Design and Implementation of a Recoverable LLM Translation System

ResearchDGX agent

arXiv:2608.09187v1 Announce Type: new Abstract: A long-form translation request can succeed at the API layer and still produce an unusable result. The output may be empty, truncated, filtered, dominat

FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models

Model ReleasesDGX agent

arXiv:2511.18852v2 Announce Type: replace Abstract: Content moderation filters are a critical safeguard against alignment failures in language models. Yet most existing filters focus narrowly on gener

Focus particles and scalar inferences across humans and language models

ResearchDGX agent

arXiv:2608.08227v1 Announce Type: new Abstract: Focus particles such as 'even' and 'only' are central to formal semantic theories that posit structured representations over sets of alternatives. 'Even

From Chains to DAGs: Probing the Graph Structure of Reasoning in LLMs

ResearchDGX agent

arXiv:2601.17593v3 Announce Type: replace Abstract: Recent progress in large language models has renewed interest in how multi-step reasoning is represented internally. While prior work often treats r

From Rebound to Remedy: Understanding and Mitigating Reward Hacking via Representation Engineering

SafetyDGX agent

arXiv:2604.01476v2 Announce Type: replace-cross Abstract: Reinforcement learning for LLMs is vulnerable to reward hacking, where models exploit shortcuts to maximize reward without solving the intende

From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios

ResearchDGX agent

arXiv:2608.08510v1 Announce Type: new Abstract: Humans have the remarkable ability to engage in spontaneous informal conversations and selectively attend to individual speakers while filtering out com

From token probabilities to calibrated confidence: An empirical study of mathematical question answering

ResearchDGX agent

arXiv:2608.07827v1 Announce Type: cross Abstract: Confidence estimation for large language models (LLMs) aims to estimate the probability that a generated answer is correct, while calibration aligns t

GraphWalker: Agentic Knowledge Graph Question Answering via Synthetic Trajectory Curriculum

AgentsDGX agent

arXiv:2603.28533v3 Announce Type: replace Abstract: Agentic knowledge graph question answering (KGQA) requires an agent to iteratively interact with knowledge graphs (KGs), posing challenges in both t

High-Layer Attention Pruning with Rescaling

Model ReleasesDGX agent

arXiv:2507.01900v3 Announce Type: replace Abstract: Pruning is a highly effective approach for compressing large language models (LLMs), significantly reducing inference latency. However, conventional

HindsightBench: A Black-Box Behavioral Audit Protocol for Parametric Hindsight in Time-Indexed LLM Decision Tasks

ResearchDGX agent

arXiv:2607.18867v2 Announce Type: replace-cross Abstract: Large language models leak parametric knowledge of what followed a historical date into decision tasks indexed by that date -- not necessarily

Human-Level Text-to-SQL via Reinforcement Learning on Verified Data, Without Pipeline Engineering

ResearchDGX agent

arXiv:2603.20004v3 Announce Type: replace-cross Abstract: Translating natural language questions to SQL queries (Text-to-SQL) is a long-standing problem in database research. Recent efforts have focus

IDRAAK: From Multi-Agent NLP to Few-Shot Prompting for Semantic Drift Detection in Technical Requirements

AgentsDGX agent

arXiv:2608.08801v1 Announce Type: new Abstract: Translating technical requirements across languages can introduce semantic drift, altering numerical constraints, polarities, modalities, or other speci

InfMem: Learning System-2 Memory Control for Long-Context Agent

AgentsDGX agent

arXiv:2602.02704v2 Announce Type: replace Abstract: Reasoning over ultra-long documents requires synthesizing sparse evidence scattered across distant segments under strict memory constraints. While s

Instability of LLM Pre-Pretraining: It Doesn't Always Help. An Investigation on Multiple Languages

Model ReleasesDGX agent

arXiv:2608.08800v1 Announce Type: new Abstract: Pretraining LLMs on artificial languages ('pre-pretraining') is a technique that could reportedly increase token efficiency by 33%, i.e., save up to 33%

Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation

SafetyDGX agent

arXiv:2608.09420v1 Announce Type: new Abstract: User simulators are widely used as scalable environments for training and evaluating interactive assistants. Generating the next user turn is inherently

Investigating Multimodal Informativity under Different Partner Visibility Conditions in Video-Mediated Dialogue

SafetyDGX agent

arXiv:2608.08915v1 Announce Type: new Abstract: Situated language use is multimodal and embodied. For example, gestures can carry information that is absent or underspecified in the speech signal, yet

Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family

SafetyDGX agent

arXiv:2604.05971v2 Announce Type: replace-cross Abstract: Recent research has shown that contrastive vision-language models such as CLIP often lack fine-grained understanding of visual content. While

Is the ACL Responsible NLP Checklist a Box-Ticking Exercise? A Large-Scale Analysis of EMNLP 2025

Model ReleasesDGX agent

arXiv:2608.09280v1 Announce Type: new Abstract: Responsible NLP practice includes a) transparency, b) ethics, and c) societal impacts. The Responsible NLP Checklist aims to push these goals, and promo

Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation

Model ReleasesDGX agent

arXiv:2608.07763v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong performance on tasks such as image captioning, visual question answering, and image-to-text generatio

← Previous
1234…128
Next →