AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
27 May 2026

SONAR-LLM: Autoregressive Transformer that Thinks in Sentence Embeddings and Speaks in Tokens

Model ReleasesDGX agent

arXiv:2508.05305v2 Announce Type: replace Abstract: The recently proposed Large Concept Model (LCM) generates text by predicting a sequence of sentence-level embeddings and training with either mean-s

SPEAR: Code-Augmented Agentic Prompt Optimization

AgentsDGX agent

arXiv:2605.26275v1 Announce Type: new Abstract: Automatic prompt engineering (APE) rewrites prompts to improve downstream task performance, but existing APE loops treat the optimizer itself as a fixed

Stylistic Evolution and LLM Neutrality in Singlish Language

ResearchDGX agent

arXiv:2601.06580v2 Announce Type: replace Abstract: Singlish is a creole rooted in Singapore's multilingual environment that continues to evolve alongside social and technological change. We examine d


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SWE-Adept: An LLM-Based Agentic Framework for Deep Codebase Analysis and Structured Issue Resolution

Model ReleasesDGX agent

arXiv:2603.01327v2 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit strong performance on self-contained programming tasks. However, they still struggle with repository-leve

Telenor Nordics Customer Service self-help corpus

AgentsDGX agent

arXiv:2605.26891v1 Announce Type: new Abstract: This paper presents a multilingual customer service self-help corpus comprising 1,122 manually validated documents in Finnish, Danish, Norwegian, and Sw

Temporal Simultaneity Predicts Annotation Quality in Sentiment Corpora

Model ReleasesDGX agent

arXiv:2605.27239v1 Announce Type: new Abstract: Annotation quality is difficult to sustain when campaigns span weeks or months with small annotator pools. We present a Setswana sentiment dataset of 3,

The Coverage Illusion: From Pre-retrieval Routing Failure to Post-retrieval Cascades in a Production RAG System

SafetyDGX agent

arXiv:2605.27220v1 Announce Type: new Abstract: In modern RAG pipelines, query augmentation methods such as HyDE and query expansion are applied to every query, resulting in substantial LLM inference

The Daily Dose: Workflow-Integrated Large Language Model Automation for Clinical Summarization and Trial Identification in Radiation Oncology

ResearchDGX agent

arXiv:2605.26346v1 Announce Type: new Abstract: Objective: To describe the design and early clinical evaluation of The Daily Dose (TDD), an LLM-driven, automated clinical summarization and clinical-tr

The Need for an External Observer Formalizing the Sufficiency Gap: A Mathematical Extension of Mixture Identifiability and Contextual Grounding in Sequence Models

AgentsDGX agent

arXiv:2605.26711v1 Announce Type: new Abstract: We construct a binary mixed-regime process with one deterministic textual regime and one random regime governed by an unobserved latent state. Even an i

To model human linguistic prediction, make LLMs less superhuman

SafetyDGX agent

arXiv:2510.05141v2 Announce Type: replace Abstract: When we read, we make predictions about upcoming words; these predictions influence our reading behavior. The success of large language models (LLMs

Towards Just-in-Time Adaptive Feedback: Enhancing Student Learning via Knowledge-Grounded LLM

TutorialsDGX agent

arXiv:2605.26405v1 Announce Type: new Abstract: Educational interventions are effective tools for enhancing student learning. While Large Language Models (LLMs) allow for generating adaptive feedback

UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action

ApplicationsDGX agent

arXiv:2510.17790v3 Announce Type: replace-cross Abstract: Computer-use agents face a fundamental limitation. They rely exclusively on primitive GUI actions (click, type, scroll), creating brittle exec

Uncertainty-Aware Budget Allocation for Adaptive Test-Time Reasoning

ResearchDGX agent

arXiv:2605.26849v1 Announce Type: new Abstract: Sampling multiple responses improves language model reasoning, but uniform compute allocation is inefficient: easy questions are over-sampled while hard

Vectors Are Not Neutral: Sensitive-Information Inference from Exported LLM Representations in Summarization

Model ReleasesDGX agent

arXiv:2605.26433v1 Announce Type: new Abstract: Large language model (LLM) summarization systems may pass compact vector representations of private inputs to downstream retrieval, monitoring, audit, o

VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models

SafetyDGX agent

arXiv:2510.17759v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) extend large language models with visual reasoning, but their multimodal design also introduces new, underexplor

Verilog-Evolve: Feedback-Driven and Skill-Evolving Verilog Generation

ResearchDGX agent

arXiv:2605.26498v1 Announce Type: new Abstract: Large language models (LLMs) have improved Verilog generation from natural-language specifications, but most pipelines still treat generation as isolate

When Does Demographic Information Help? Data and Modeling Regimes for Perspective-Aware Hate Speech Detection

ResearchDGX agent

arXiv:2605.27313v1 Announce Type: new Abstract: Demographic information is often used to model annotator perspectives in subjective tasks such as hate speech detection, but its benefit is inconsistent

Why Prompt Optimization Works, and Why It Sometimes Doesn't: A Causal-Inspired Edit-Level Analysis

Model ReleasesDGX agent

arXiv:2605.26655v1 Announce Type: new Abstract: Automated prompt optimization methods (e.g., DSpy, TextGrad) can substantially improve the performance of large language model (LLM), however, their gen

26 May 2026

A Comprehensive Dataset for Human vs. AI Generated Text Detection

Model ReleasesDGX agent

arXiv:2510.22874v3 Announce Type: replace Abstract: The rapid advancement of large language models (LLMs) has led to increasingly human-like AI-generated text, raising concerns about content authentic

A Lightweight Hybrid Transformer-CRF Architecture for Multi-Type Bangla Medical Entity Recognition

ApplicationsDGX agent

arXiv:2605.25463v1 Announce Type: new Abstract: MedER refers to the identification of medical entities. It is crucial for extracting structured clinical information from unstructured medical text. Man

A Multi-Probe Audit of Clinical-Interview Depression Detection Benchmarks

Model ReleasesDGX agent

arXiv:2605.23977v1 Announce Type: new Abstract: This paper audits benchmark evaluation in clinical-interview depression detection through four complementary probes across DAIC/E-DAIC, CMDC, ANDROIDS,

A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays

Model ReleasesDGX agent

arXiv:2605.25652v1 Announce Type: new Abstract: Free-form legal essay evaluation in NLP treats expert inter-rater stability as a single ceiling number, and treats LLM-judge agreement with that ceiling

Act or Clarify? Modeling Sensitivity to Uncertainty and Cost in Communication

AgentsDGX agent

arXiv:2602.02843v3 Announce Type: replace Abstract: When deciding how to act under uncertainty, agents may choose to act to reduce uncertainty or they may act despite that uncertainty. In communicativ

Adaptive Preference Optimization with Uncertainty-aware Utility Anchor

SafetyDGX agent

arXiv:2509.10515v1 Announce Type: cross Abstract: Offline preference optimization methods are efficient for large language models (LLMs) alignment. Direct Preference optimization (DPO)-like learning,

AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue

Model ReleasesDGX agent

arXiv:2605.23974v1 Announce Type: new Abstract: Current language models create two safety challenges: risk must be detected early enough to avoid exposing harmful continuation, and the harmfulness its

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks

Model ReleasesDGX agent

arXiv:2505.24876v2 Announce Type: replace-cross Abstract: Deep reasoning is fundamental for solving complex tasks, especially in vision-centric scenarios that demand sequential, multimodal understandi

AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios

Model ReleasesDGX agent

arXiv:2508.19988v3 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved high accuracy on complex commonsense and mathematical problems that involve the composition of multiple r

AgentIR: A Workload-Adaptive Cascade Retrieval Substrate for Long-Term Conversational Memory

HardwareDGX agent

arXiv:2605.25092v1 Announce Type: cross Abstract: Long-term conversational memory is a retrieval workload classical IR was not built for: the index grows during the query stream, query types shift int

An Effective-Rank Audit of Alignment-Induced Activation Shifts: Confound Control, Constructive Calibration, and Limits

Model ReleasesDGX agent

arXiv:2605.24583v1 Announce Type: cross Abstract: We audit alignment-induced shifts in residual-stream activations of three open-weight instruction-tuned LLMs (Llama-3.1-8B-Instruct, Gemma-2-9B-it, Qw

Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agents

Model ReleasesDGX agent

arXiv:2605.25971v1 Announce Type: new Abstract: While AI agents demonstrate remarkable capabilities in reasoning and tool use, they remain fundamentally reactive: they compute responses only after exp

AstroMind: A High-Fidelity Benchmark for Spacecraft Behavior Reasoning Based on Large Language Models

Model ReleasesDGX agent

arXiv:2605.24573v1 Announce Type: new Abstract: Understanding why a spacecraft maneuvers -- rather than simply that it did -- is an increasingly important problem for space domain awareness as Earth o

AuthTrace: Diagnosing Evidence Construction in Thematically Dense Single-Author Corpora

Model ReleasesDGX agent

arXiv:2605.25382v1 Announce Type: new Abstract: Evidence construction systems--chunk retrieval, agent memory, knowledge-graph traversal, and thematic indexing--are evaluated on separate benchmarks wit

Automated Benchmark Auditing for AI Agents and Large Language Models

Model ReleasesDGX agent

arXiv:2605.26079v1 Announce Type: new Abstract: Modern AI benchmarks operate at a complexity that outpaces traditional verification methods. Tasks authored by domain experts often contain implicit ass

AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery

AgentsDGX agent

arXiv:2604.05550v2 Announce Type: replace Abstract: Artificial intelligence research increasingly depends on prolonged cycles of reproduction, debugging, and iterative refinement to achieve State-Of-T

Axis-Aligned Semantics for ODRL: Resolving Dimensional Ambiguity in Policy Constraints

Model ReleasesDGX agent

arXiv:2602.19878v3 Announce Type: replace Abstract: The Open Digital Rights Language (ODRL) represents policy constraints as triples of a left operand, an operator, and a value. Several spatial operan

Benchmarking and Learning Real-World Customer Service Dialogue

Model ReleasesDGX agent

arXiv:2510.22143v3 Announce Type: replace Abstract: Existing benchmarks and training pipelines for industrial intelligent customer service (ICS) remain misaligned with real-world dialogue requirements

Better, Faster: Harnessing Self-Improvement in Large Reasoning Models

ResearchDGX agent

arXiv:2605.24998v1 Announce Type: new Abstract: Self-improvement training enables the large reasoning models (LRMs) to improve themselves by self-generating reasoning trajectories as training data wit

Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC

Model ReleasesDGX agent

arXiv:2605.25626v1 Announce Type: new Abstract: Social media platforms enable large-scale cross-lingual communication, but translating user-generated content (UGC) remains challenging due to its infor

Beyond the Target: From Imitation to Collaboration in Speculative Decoding

SafetyDGX agent

arXiv:2605.24793v1 Announce Type: new Abstract: Speculative decoding (SPD) accelerates large language model (LLM) inference by letting a smaller draft model propose multiple future tokens that are ver

Can Large Language Models Resolve Semantic Discrepancy in Self-Destructive Subcultures? Evidence from Jirai Kei

SafetyDGX agent

arXiv:2601.05004v2 Announce Type: replace Abstract: Self-destructive behaviors are linked to complex psychological states and can be challenging to diagnose. These behaviors may be even harder to iden

Can LoRA Fusion Support Cross-Domain Tasks in Cloud-Edge Collaboration?

Model ReleasesDGX agent

arXiv:2605.23913v1 Announce Type: cross Abstract: Cloud-hosted large language models (LLMs) commonly rely on LoRA for domain adaptation, yet domain data are distributed across multiple edge devices an

Clarification Is Not Enough: Post-Clarification Answering Remains the Bottleneck in Multi-Turn QA

Model ReleasesDGX agent

arXiv:2605.25204v1 Announce Type: new Abstract: Pluralistic alignment requires systems to adapt to diverse user values, communication styles, and contextual assumptions. We believe that a foundational

CMAP: Cross-Modal Adaptive Prompting for Multi-Domain Task-Incremental Learning

Model ReleasesDGX agent

arXiv:2605.25708v1 Announce Type: cross Abstract: Multi-domain task-incremental learning requires a model to sequentially acquire knowledge across visually diverse domains without forgetting prior tas

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions

Model ReleasesDGX agent

arXiv:2605.24279v1 Announce Type: new Abstract: A frontier language model's acknowledged 'helpful programming assistant' persona does not survive long agentic-coding sessions in the deployment regime

CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming

AgentsDGX agent

arXiv:2605.24693v1 Announce Type: new Abstract: Large language models still struggle with contest-level programming, while many agentic remedies rely on massive inference-time sampling or expensive mu

CRPO: Character-centric Group Relative Policy Optimization for Role-aware Reasoning in Role-playing Agents

SafetyDGX agent

arXiv:2605.25511v1 Announce Type: new Abstract: Recent advancements in Reinforcement Learning (RL), particularly Group Relative Policy Optimization (GRPO), have significantly enhanced the reasoning ca

CSP-Atlas: Concept-Specific Neural Circuits in a Sparse Python Transformer

Model ReleasesDGX agent

arXiv:2605.24603v1 Announce Type: new Abstract: A sparse 8-layer code transformer develops dedicated neural circuitry for every Python construct tested, and that circuitry is organised by a clean comp

CUNY at CLPsych 2026: A Pipeline Approach to Classification and Summarization of Mental Health Changes

ResearchDGX agent

arXiv:2605.24164v1 Announce Type: new Abstract: We describe our submission to the CLPsych~2026 Shared Task on capturing and characterizing mental health changes through social media timeline dynamics.

Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval

Model ReleasesDGX agent

arXiv:2605.24454v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong performance in the legal domain, demonstrating notable potential in Legal Question Answering (LQA). Howev

DeIDClinic: A Risk-Aware Pseudonymization Framework for Clinical Text De-identification and Re-identification Risk Assessment

ApplicationsDGX agent

arXiv:2410.01648v2 Announce Type: replace Abstract: The increasing availability of sensitive textual data has created an urgent need for robust de-identification methods that enable compliant data sha

Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs

SafetyDGX agent

arXiv:2605.23975v1 Announce Type: new Abstract: Audio large language models (Audio LLMs) exhibit systematic failures in transcribing code-switching speech despite strong multilingual capabilities. Foc

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models

Model ReleasesDGX agent

arXiv:2605.25189v1 Announce Type: cross Abstract: Reward hacking arises when a model improves a proxy reward by exploiting shortcuts rather than solving the intended task. We study this failure mode t

Discovering Lexical Gaps Using Embeddings from Multilingual LLMs

SafetyDGX agent

arXiv:2605.24310v1 Announce Type: new Abstract: Lexical gaps are words that do not exist in certain languages. They pose challenges for building multilingual lexical resources, for machine translation

Distinguishing Right from Wrong in Debates: Attribution Analysis of Chinese Harmful Memes

ResearchDGX agent

arXiv:2605.24344v1 Announce Type: new Abstract: Research on harmful meme detection has garnered significant attention, resulting in the development of numerous datasets and methods. However, progress

Does Continued Pretraining on a Learner Corpus Improve Automated Essay Scoring on English Proficiency Tests? Evidence from EFCAMDAT

ResearchDGX agent

arXiv:2605.25924v1 Announce Type: new Abstract: Recent automated essay scoring (AES) studies increasingly use pretrained transformer models, but these models are usually pretrained on general-domain E

Double Triangle Annotation: A Scalable Human-in-the-Loop Framework for High-Precision Historical Document Annotation

Model ReleasesDGX agent

arXiv:2605.25781v1 Announce Type: new Abstract: Evaluating structured-information extraction from historical documents at scale requires high-precision ground-truth annotations, yet traditional manual

DRInQ: Evaluating Conversational Implicature with Controlled Context Variation

Model ReleasesDGX agent

arXiv:2605.24267v1 Announce Type: new Abstract: Human conversation relies heavily on conversational implicature, in which speakers convey meanings that are suggested rather than explicitly stated. Alt

DTO: a Differentiable Training Objective for Effective Counterfactual Story Rewriting

Local AiDGX agent

arXiv:2605.24885v1 Announce Type: new Abstract: Counterfactual story rewriting is a natural language processing task that requires updating an existing story to reflect a chosen alternative event, yet

DUEL: Adversarial Self-Play for Multimodal Reasoning

ResearchDGX agent

arXiv:2605.24794v1 Announce Type: cross Abstract: Reinforcement learning (RL) has emerged as an effective paradigm for improving the reasoning capability of vision-language models (VLMs). However, RL-

DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning

SafetyDGX agent

arXiv:2605.25604v1 Announce Type: new Abstract: Reinforcement Learning has become a standard paradigm for aligning Large Language Models with human intent and task requirements. While Group Relative P

← Previous
1…5960616263…129
Next →