AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
21 May 2026

Mem-pi: Adaptive Memory through Learning When and What to Generate

AgentsDGX agent

arXiv:2605.21463v1 Announce Type: new Abstract: We present Mem-pi, a framework for adaptive memory in large language model (LLM) agents, where useful guidance is generated on demand rather than retrie

MemGym: a Long-Horizon Memory Environment for LLM Agents

Model ReleasesDGX agent

arXiv:2605.20833v1 Announce Type: new Abstract: Memory is a central capability for LLM agents operating across long-horizon tasks. Existing memory benchmarks predominantly evaluate retention of person

Memory Grafting: Scaling Language Model Pre-training via Offline Conditional Memory

Model ReleasesDGX agent

arXiv:2605.20948v1 Announce Type: new Abstract: Scaling conditional memory offers a promising way to increase language-model capacity, but existing methods such as Engram learn large memory tables fro


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Metaphors in Literary Post-Editing: Opening Pandora's Box?

ResearchDGX agent

arXiv:2605.21178v1 Announce Type: new Abstract: This paper investigates how post-editors of literary texts react and respond to the way metaphors have been translated by Neu ral Machine Translation (N

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs

AgentsDGX agent

arXiv:2605.20315v1 Announce Type: new Abstract: LLM agents have recently emerged as a powerful paradigm for solving complex tasks through planning, tool use, memory retrieval, and multi-step interacti

Most Transformer Modifications Still Do Not Transfer at 1-3B: A 2020-2026 Update to Narang et al. (2021) with Downstream Evaluation and a Noise Floor

ResearchDGX agent

arXiv:2605.20798v1 Announce Type: cross Abstract: Narang et al. (2021) evaluated 40+ Transformer modifications at T5-base scale and concluded that most did not transfer. Five years later, the typical

MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks

Model ReleasesDGX agent

arXiv:2605.20729v1 Announce Type: new Abstract: Accurate evaluation of conversational retrieval is pivotal for advancing Retrieval-Augmented Generation (RAG) systems. However, existing conversational

Multi-agent Collaboration with State Management

AgentsDGX agent

arXiv:2605.20563v1 Announce Type: cross Abstract: Recent advances in multi-agent systems have shown great potential for solving complex tasks. However, when multiple agents edit a shared codebase conc

NeuroQA: A Large-Scale Image-Grounded Benchmark for 3D Brain MRI Understanding

Model ReleasesDGX agent

arXiv:2605.20525v1 Announce Type: cross Abstract: We present NeuroQA, a large-scale benchmark for visual question answering in 3D brain magnetic resonance imaging (MRI), with 56,953 QA pairs from 12,9

On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists

Model ReleasesDGX agent

arXiv:2605.20668v1 Announce Type: new Abstract: With the advancement of AI capabilities, AI reviewers are beginning to be deployed in scientific peer review, yet their capability and credibility remai

Optimal Query Allocation in Extractive QA with LLMs: A Learning-to-Defer Framework with Theoretical Guarantees

ResearchDGX agent

arXiv:2410.15761v4 Announce Type: replace Abstract: Large Language Models excel in generative tasks but exhibit inefficiencies in structured text selection, particularly in extractive question answeri

Parallel LLM Reasoning for Bias-Resilient, Robust Conceptual Abstraction

SafetyDGX agent

arXiv:2605.20194v1 Announce Type: new Abstract: Large language models (LLMs) have been increasingly used to analyze text. However, they are often plagued with contextual reasoning limitations when ana

Playing Devil's Advocate: Off-the-Shelf Persona Vectors Rival Targeted Steering for Sycophancy

ResearchDGX agent

arXiv:2605.21006v1 Announce Type: cross Abstract: We study the effect of different persona on extbf{sycophancy}: model's agreement with users even when the user is incorrect. The standard mitigation,

Post-Hoc Understanding of Metaphor Processing in Decoder-Only Language Models via Conditional Scale Entropy

Model ReleasesDGX agent

arXiv:2605.21391v1 Announce Type: new Abstract: Metaphor requires a language model to resolve a token whose contextual meaning diverges from its basic literal sense. Understanding how transformer mode

Pseudo-Siamese Network for Planning in Target-Oriented Proactive Dialogues

TutorialsDGX agent

arXiv:2605.20195v1 Announce Type: new Abstract: A target-oriented proactive dialogue system is designed to steer conversations toward predefined targets while actively providing suggestions. The core

PulseCol: Periodically Refreshed Column-Sparse Attention for Accelerating Diffusion Language Models

HardwareDGX agent

arXiv:2605.20813v1 Announce Type: new Abstract: Inference in diffusion large language models (dLLMs) is computationally expensive, as full self-attention must be repeatedly executed at each step of th

Puzzled By ChatGPT? No more! A Jigsaw Puzzle to Promote AI Literacy and Awareness

TutorialsDGX agent

arXiv:2605.20404v1 Announce Type: new Abstract: The rapid adoption of Generative AI, including LLM-based chatbots like ChatGPT, has highlighted the need for accessible ways to support public understan

Quantifying the cross-linguistic effects of syncretism on agreement attraction

ResearchDGX agent

arXiv:2605.21403v1 Announce Type: new Abstract: Agreement attraction errors, in which a verb erroneously agrees with an intervening noun rather than its grammatical head, are amplified by morphologica

Refining and Reusing Annotation Guidelines for LLM Annotation

Model ReleasesDGX agent

arXiv:2605.20809v1 Announce Type: new Abstract: While Large Language Models (LLMs) demonstrate remarkable performance on zero-shot annotation tasks, they often struggle with the specialized convention

Reinforcing Human Behavior Simulation via Verbal Feedback

Model ReleasesDGX agent

arXiv:2605.20506v1 Announce Type: cross Abstract: Humans learn social norms and behaviors from verbal feedback (e.g., a parent saying 'that was rude' or a friend explaining 'here's why that hurt'). Ye

Reliable Automated Triage in Spanish Clinical Notes: A Hybrid Framework for Risk-Aware HIV Suspicion Identification

ResearchDGX agent

arXiv:2605.21256v1 Announce Type: new Abstract: Standard clinical Natural Language Processing (NLP) benchmarks often yield inflated metrics by forcing deterministic classification on ambiguous instanc

Retrieval-Augmented Code Generation: A Survey with Focus on Repository-Level Approaches

AgentsDGX agent

arXiv:2510.04905v3 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) have significantly improved automated code generation. While existing approaches have achieved

Retrieval-Augmented Long-Context Translation for Cultural Image Captioning: Gators submission for AmericasNLP 2026 shared task

Model ReleasesDGX agent

arXiv:2605.20626v1 Announce Type: new Abstract: We present the University of Florida Gators submission to the AmericasNLP 2026 shared task on cultural image captioning for Indigenous languages. Our tw

Retrospective Sparse Attention for Efficient Long-Context Generation

ResearchDGX agent

arXiv:2508.09001v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly deployed in long-context tasks such as reasoning, code generation, and multi-turn dialogue. However, i

SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR

SafetyDGX agent

arXiv:2605.20712v1 Announce Type: new Abstract: Automatic speech recognition replaces typing only when correction costs less than manual entry, a threshold determined by error types, not counts: fixin

Self-Training Doesn't Flatten Language -- It Restructures It: Surface Markers Amplify While Deep Syntax Dies

ResearchDGX agent

arXiv:2605.20602v1 Announce Type: new Abstract: Successive self-training on a language model's own outputs is widely characterized as a process of flattening: diversity drops, distributions narrow, an

SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass

Model ReleasesDGX agent

arXiv:2602.06358v2 Announce Type: replace Abstract: We propose SHINE (Scalable Hyper In-context NEtwork), a scalable hypernetwork that can map diverse meaningful contexts into high-quality LoRA adapte

Shiny Stories, Hidden Struggles: Investigating the Representation of Disability Through the Lens of LLMs

SafetyDGX agent

arXiv:2605.20191v1 Announce Type: new Abstract: Modern Large Language Models (LLMs) have recently attracted much attention for their ability to simulate human behavior and generate text that reflects

Single-Pass, Depth-Selective Reading for Multi-Aspect Sentiment Analysis

ResearchDGX agent

arXiv:2605.20998v1 Announce Type: new Abstract: Aspect-Term Sentiment Analysis (ATSA) in multi-aspect sentences faces a fundamental tradeoff between efficiency and expressiveness. Existing models eith

Smarter edits? Post-editing with error highlights and translation suggestions

ResearchDGX agent

arXiv:2605.21135v1 Announce Type: new Abstract: As MT quality increases, interest in enhanced post-editing features such as QE-derived error highlights is growing, yet evidence for their usefulness re

SMoA: Spectrum Modulation Adapter for Parameter-Efficient Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.21147v1 Announce Type: cross Abstract: As the number of model parameters increases, parameter-efficient fine-tuning (PEFT) has become the go-to choice for tailoring pre-trained large langua

SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents

Model ReleasesDGX agent

arXiv:2605.21384v1 Announce Type: cross Abstract: As long-horizon coding agents produce more code than any developer can review, oversight collapses onto a single surface: the automated test suite. Re

Stage-Audit: Auditable Source-Frontier Discovery for Cross-Wiki Tables

SafetyDGX agent

arXiv:2605.20478v1 Announce Type: new Abstract: LLM-curated tables can appear source-grounded while containing unsupported rows: the curator may recall entries from parametric memory and retroactively

Strategy-Induct: Task-Level Strategy Induction for Instruction Generation

ResearchDGX agent

arXiv:2605.20924v1 Announce Type: new Abstract: Designing effective task-level prompts is crucial for improving the performance of Large Language Models (LLMs). While prior work on instruction inducti

SymbolicLight V1: Spike-Gated Dual-Path Language Modeling with High Activation Sparsity and Sub-Billion-Scale Pre-Training Evidence

Model ReleasesDGX agent

arXiv:2605.21333v1 Announce Type: new Abstract: Natively trained spiking language models struggle to combine Transformer-like language quality, stable multi-domain pre-training, and high activation sp

Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models

SafetyDGX agent

arXiv:2605.20356v1 Announce Type: new Abstract: Full-duplex spoken dialogue models (SDMs) can listen and speak simultaneously, enabling interaction dynamics closer to human conversation than turn-base

Task-conditioned probing of instruction-tuned multimodal LLMs: Region-specific brain alignment patterns under naturalistic stimuli

SafetyDGX agent

arXiv:2506.08277v3 Announce Type: replace-cross Abstract: Recent voxel-wise multimodal brain encoding studies have shown that multimodal large language models (MLLMs) exhibit a higher degree of brain

Task-Routed Mixture-of-Experts with Cognitive Appraisal for Implicit Sentiment Analysis

TutorialsDGX agent

arXiv:2605.20916v1 Announce Type: new Abstract: Implicit sentiment analysis is challenging because sentiment toward an aspect is often inferred from events rather than expressed through explicit opini

Terminal-World: Scaling Terminal-Agent Environments via Agent Skills

Model ReleasesDGX agent

arXiv:2605.20876v1 Announce Type: new Abstract: Terminal agents extend Large Language Models with the ability to execute tasks directly in command-line environments, but their progress is bottlenecked

Text Analytics Evaluation Framework: A Case Study on LLMs and Social Media

Model ReleasesDGX agent

arXiv:2605.21338v1 Announce Type: new Abstract: LLMs have demonstrated exceptional proficiency in a wide range of NLP tasks. However, a notable gap remains in practical data analysis scenarios, partic

TextReg: Mitigating Prompt Distributional Overfitting via Regularized Text-Space Optimization

ResearchDGX agent

arXiv:2605.21318v1 Announce Type: new Abstract: Large language models (LLMs) are highly sensitive to the prompts used to specify task objectives and behavioral constraints. Many recent prompt optimiza

The Generation-Recognition Asymmetry: Six Dimensions of a Fundamental Divide in Formal Language Theory

ApplicationsDGX agent

arXiv:2603.10139v2 Announce Type: replace Abstract: Every formal grammar defines a language and can in principle be used in three ways: to generate strings (production), to recognize them (parsing), o

The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering

ResearchDGX agent

arXiv:2605.20745v1 Announce Type: cross Abstract: Generative verifiers have emerged as a promising paradigm for step-wise verification, but their verification behavior is often poorly calibrated: they

The Illusion of Intervention: Your LLM-Simulated Experiment is an Observational Study

SafetyDGX agent

arXiv:2605.20767v1 Announce Type: new Abstract: Large language models (LLMs) show potential as simulators of human behavior, offering a scalable way to study responses to interventions. However, becau

The Visual Iconicity Challenge: Evaluating Vision-Language Models on Sign Language Form-Meaning Mapping

Model ReleasesDGX agent

arXiv:2510.08482v3 Announce Type: replace-cross Abstract: Iconicity, the resemblance between linguistic form and meaning, is pervasive in signed languages, offering a natural testbed for visual ground

Thinking-while-speaking: A Controlled, Interleaved Reasoning Method for Real-Time Speech Generation

ResearchDGX agent

arXiv:2605.20946v1 Announce Type: new Abstract: The thinking-while-speaking paradigm aims to make AI communication more human. A key challenge is maintaining fluent speech while performing deep reason

Towards Context-Invariant Safety Alignment for Large Language Models

SafetyDGX agent

arXiv:2605.20994v1 Announce Type: new Abstract: Preference-based post-training aligns LLMs with human intent, yet safety behavior often remains brittle. A model may refuse a harmful request in a stand

Towards the Anonymization of the Language Modeling

ApplicationsDGX agent

arXiv:2501.02407v3 Announce Type: replace Abstract: Rapid advances in Natural Language Processing (NLP) have revolutionized many fields, including healthcare. However, these advances raise significant

Toxic Subword Pruning for Dialogue Response Generation on Large Language Models

Model ReleasesDGX agent

arXiv:2410.04155v2 Announce Type: replace Abstract: How to defend large language models (LLMs) from generating toxic content is an important research area. Yet, most research focused on various model

Tracing the ongoing emergence of human-like reasoning in Large Language Models

ResearchDGX agent

arXiv:2605.21299v1 Announce Type: new Abstract: Humans effortlessly go beyond literal meanings: If you mow the lawn, I will give you fifty dollars, is typically understood as implying that the speaker

Training Language Agents to Learn from Experience

Model ReleasesDGX agent

arXiv:2605.20477v1 Announce Type: cross Abstract: Language agents can adapt from experience in interactive environments, but current reflection-based methods can only self-correct within a single task

Under Pressure: Emotional Framing Induces Measurable Behavioral Shifts and Structured Internal Geometry in Small Language Models

Model ReleasesDGX agent

arXiv:2605.20202v1 Announce Type: new Abstract: I study whether emotionally framed evaluation follow-ups change both the behavior and the calm-relative internal representations of small, locally deplo

Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs

Model ReleasesDGX agent

arXiv:2505.19075v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated remarkable general capabilities, but enhancing skills such as reasoning often demands substanti

WCXB: A Multi-Type Web Content Extraction Benchmark

Model ReleasesDGX agent

arXiv:2605.21097v1 Announce Type: new Abstract: Web content extraction - isolating a page's main content from surrounding boilerplate - is a prerequisite for search indexing, retrieval-augmented gener

What Do Biomedical NER and Entity Linking Benchmarks Measure? A Corpus-Centric Diagnostic Framework

Model ReleasesDGX agent

arXiv:2605.20537v1 Announce Type: new Abstract: Biomedical named entity recognition (NER) and entity linking (EL) strongly depend on annotated corpora, but the utility of these resources for benchmark

When Irregularity Helps: A Subclass Analysis of Inductive Bias in Neural Morphology

Model ReleasesDGX agent

arXiv:2605.20558v1 Announce Type: new Abstract: Neural morphological generation systems often achieve high aggregate accuracy on benchmark datasets, yet such performance can conceal systematic errors

When Reasoning Supervision Hurts: TTCW-Based Long-Form Literary Review Generation

ResearchDGX agent

arXiv:2605.20364v1 Announce Type: new Abstract: Automatic evaluation of long-form literary writing remains challenging, as generic LLM-as-Judge approaches may not fully capture creativity-related dime

You Are What You Say: Exploiting Linguistic Content for VoicePrivacy Attacks

ResearchDGX agent

arXiv:2506.09521v2 Announce Type: replace-cross Abstract: Speaker anonymization systems hide the identity of speakers while preserving other information such as linguistic content and emotions. To eva

You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories

Model ReleasesDGX agent

arXiv:2605.21468v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a dominant paradigm for improving reasoning in large language models (LLMs), yet the

20 May 2026

A Data-Driven Approach to Idiomaticity Based on Experts' Criteria in Theoretical Linguistics

ResearchDGX agent

arXiv:2605.19575v1 Announce Type: new Abstract: The article observes data analysis of 286 multi-word expressions (MWEs) based on 16 lexical, grammatical and other criteria described in theoretical boo

← Previous
1…6768697071…129
Next →