AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
30 Apr 2026

Differentially-Private Text Rewriting reshapes Linguistic Style

ResearchDGX agent

arXiv:2604.26656v1 Announce Type: new Abstract: Differential Privacy (DP) for text matured from disjointed word-level substitutions to contiguous sentence-level rewriting by leveraging the generative

EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses

AgentsDGX agent

arXiv:2604.26417v1 Announce Type: new Abstract: Emotion perception and adaptive expression are fundamental capabilities in human-agent interaction. While recent advances in speech emotion captioning (

Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing

ResearchDGX agent

arXiv:2604.25923v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have prompted a growing body of work that questions the methodology of prevailing evaluation practices.


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

EvoSelect: Data-Efficient LLM Evolution for Targeted Task Adaptation

SafetyDGX agent

arXiv:2604.26170v1 Announce Type: new Abstract: Adapting large language models (LLMs) to a targeted task efficiently and effectively remains a fundamental challenge. Such adaptation often requires ite

Failure Modes of Maximum Entropy RLHF

ResearchDGX agent

arXiv:2509.20265v3 Announce Type: replace-cross Abstract: In this paper, we show that Simple Preference Optimization (SimPO) can be derived as Maximum Entropy Reinforcement Learning, providing a theor

Flashback: A Reversible Bilateral Run-Peeling Decomposition of Strings

ResearchDGX agent

arXiv:2604.26190v1 Announce Type: cross Abstract: We introduce Flashback, a reversible string decomposition that repeatedly peels the maximal leading and trailing character runs from a sentinel-wrappe

FlowBot: Inducing LLM Workflows with Bilevel Optimization and Textual Gradients

ApplicationsDGX agent

arXiv:2604.26258v1 Announce Type: new Abstract: LLM workflows, which coordinate structured calls to individual LLMs (each augmented with varying instructions and tools) to achieve a particular goal, o

Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference

Model ReleasesDGX agent

arXiv:2604.26294v1 Announce Type: new Abstract: We present tensor and sequence parallelism (TSP), a parallel execution strategy that folds tensor parallelism and sequence parallelism onto a single dev

From Prompt Risk to Response Risk: Paired Analysis of Safety Behavior of Large Language Model

SafetyDGX agent

arXiv:2604.26052v1 Announce Type: new Abstract: Safety evaluations of large language models (LLMs) typically report binary outcomes such as attack success rate, refusal rate, or harmful/not-harmful re

HealthNLP_Retrievers at ArchEHR-QA 2026: Cascaded LLM Pipeline for Grounded Clinical Question Answering

Model ReleasesDGX agent

arXiv:2604.26880v1 Announce Type: new Abstract: Patient portals now give individuals direct access to their electronic health records (EHRs), yet access alone does not ensure patients understand or ac

Heterogeneous Adaptive Policy Optimization: Tailoring Optimization to Every Token's Nature

SafetyDGX agent

arXiv:2509.16591v2 Announce Type: replace Abstract: Using entropy as a measure of heterogeneity to guide optimization has emerged as a crucial research direction in Reinforcement Learning for LLMs. Ho

HIVE: Hidden-Evidence Verification for Hallucination Detection in Diffusion Large Language Models

ResearchDGX agent

arXiv:2604.26139v1 Announce Type: new Abstract: Diffusion large language models generate text through multi-step denoising, where hallucination signals may emerge throughout the trajectory rather than

Information Extraction from Electricity Invoices with General-Purpose Large Language Models

Model ReleasesDGX agent

arXiv:2604.25927v1 Announce Type: new Abstract: Information extraction from semi-structured business documents remains a critical challenge for enterprise management. This study evaluates the capabili

LIT-RAGBench: Benchmarking Generator Capabilities of Large Language Models in Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2603.06198v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) is a framework in which a Generator, such as a Large Language Model (LLM), produces answers by retrieving docum

LLMs Generate Kitsch

ResearchDGX agent

arXiv:2604.25929v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used to generate pictures, texts, music, videos, and other works that have traditionally required human cr

LogitSpec: Accelerating Retrieval-based Speculative Decoding via Next Next Token Speculation

ResearchDGX agent

arXiv:2507.01449v3 Announce Type: replace Abstract: Speculative decoding (SD), where a small draft model is employed to propose draft tokens in advance and then the target model validates them in para

Mapping the maturation of TCM as an adjuvant to radiotherapy

SafetyDGX agent

arXiv:2601.11923v2 Announce Type: replace Abstract: The integration of complementary medicine into oncology represents a paradigm shift that has seen to increasing adoption of Traditional Chinese Medi

MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese

Model ReleasesDGX agent

arXiv:2604.25926v1 Announce Type: new Abstract: The use of large language models (LLMs) for complex mathematical reasoning is an emergent area of research, with fast progress in methods, models, and b

MoRFI: Monotonic Sparse Autoencoder Feature Identification

Model ReleasesDGX agent

arXiv:2604.26866v1 Announce Type: new Abstract: Large language models (LLMs) acquire most of their factual knowledge during the pre-training stage, through next token prediction. Subsequent stages of

Multimodal LLMs are not all you need for Pediatric Speech Language Pathology

Model ReleasesDGX agent

arXiv:2604.26568v1 Announce Type: new Abstract: Speech Sound Disorders (SSD) affect roughly five percent of children, yet speech-language pathologists face severe staffing shortages and unmanageable c

OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory

AgentsDGX agent

arXiv:2604.26622v1 Announce Type: new Abstract: Autonomous LLM agents increasingly operate in long-horizon, interactive settings where success depends on reusing experience accumulated over extended h

One Voice, Many Tongues: Cross-Lingual Voice Cloning for Scientific Speech

ResearchDGX agent

arXiv:2604.26136v1 Announce Type: cross Abstract: Preserving a speaker's voice identity while generating speech in a different language remains a fundamental challenge in spoken language technology, p

One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety

SafetyDGX agent

arXiv:2604.25921v1 Announce Type: new Abstract: Large Language Models (LLMs) are trained to refuse harmful requests, yet they remain vulnerable to jailbreak attacks that exploit weaknesses in conversa

Pointer-CAD: Unifying B-Rep and Command Sequences via Pointer-based Edges & Faces Selection

ApplicationsDGX agent

arXiv:2603.04337v2 Announce Type: replace-cross Abstract: Constructing computer-aided design (CAD) models is labor-intensive but essential for engineering and manufacturing. Recent advances in Large L

Reasoning Gets Harder for LLMs Inside A Dialogue

Model ReleasesDGX agent

arXiv:2603.20133v2 Announce Type: replace Abstract: Large Language Models (LLMs) achieve strong performance on many reasoning benchmarks, yet these evaluations typically focus on isolated tasks that d

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs

ResearchDGX agent

arXiv:2507.21420v3 Announce Type: replace-cross Abstract: The computational cost of training multimodal large language models (MLLMs) grows rapidly with the number of processed tokens. Existing effici

SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts

ResearchDGX agent

arXiv:2604.26506v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly integrated into academic peer review, their vulnerability to adversarial prompts -- adversarial instruc

SAGE: A Strategy-Aware Graph-Enhanced Generation Framework For Online Counseling

SafetyDGX agent

arXiv:2604.26630v1 Announce Type: new Abstract: Effective mental health counseling is a complex, theory-driven process requiring the simultaneous integration of psychological frameworks, real-time dis

Select to Think: Unlocking SLM Potential with Local Sufficiency

Local AiDGX agent

arXiv:2604.26940v1 Announce Type: new Abstract: Small language models (SLMs) offer computational efficiency for scalable deployment, yet they often fall short of the reasoning power exhibited by their

Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training

Model ReleasesDGX agent

arXiv:2510.20956v2 Announce Type: replace-cross Abstract: We discover a novel and surprising phenomenon of unintentional misalignment in reasoning language models (RLMs), which we call self-jailbreaki

Semantic Embeddings of Chemical Elements for Enhanced Materials Inference and Discovery

ResearchDGX agent

arXiv:2502.14912v2 Announce Type: replace Abstract: We present a framework for generating universal semantic embeddings of chemical elements to advance materials inference and discovery. This framewor

Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens

Model ReleasesDGX agent

arXiv:2604.26355v1 Announce Type: new Abstract: Reasoning in Large Language Models incurs significant inference-time compute, yet the token-level information structure of reasoning traces remains unde

SpecTr-GBV: Multi-Draft Block Verification Accelerating Speculative Decoding

ResearchDGX agent

arXiv:2604.25925v1 Announce Type: new Abstract: Autoregressive language models suffer from high inference latency due to their sequential decoding nature. Speculative decoding (SD) mitigates this by e

StarDrinks: An English and Korean Test Set for SLU Evaluation in a Drink Ordering Scenario

Model ReleasesDGX agent

arXiv:2604.26500v1 Announce Type: new Abstract: LLMs and speech assistants are increasingly used for task-oriented interactions, yet their evaluation often relies on controlled scenarios that fail to

Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images

ResearchDGX agent

arXiv:2510.21828v2 Announce Type: replace-cross Abstract: Understanding and reasoning with abstractive information from the visual modality presents significant challenges for current multi-modal larg

Swap distance minimization shapes the order of subject, object and verb in languages of the world

ResearchDGX agent

arXiv:2604.26726v1 Announce Type: new Abstract: Languages of the world vary concerning the order of subject, object and verb. The most frequent dominant orders are SOV and SVO, and researchers have ta

SWE-Edit: Rethinking Code Editing for Efficient SWE-Agent

Model ReleasesDGX agent

arXiv:2604.26102v1 Announce Type: cross Abstract: Large language model agents have achieved remarkable progress on software engineering tasks, yet current approaches suffer from a fundamental context

Talent or Luck? Evaluating Attribution Bias in Large Language Models

SafetyDGX agent

arXiv:2505.22910v2 Announce Type: replace Abstract: When a student fails an exam, do we tend to blame their effort or the test's difficulty? Attribution, defined as how reasons are assigned to event o

Teaching LLM to be Persuasive: Reward-Enhanced Policy Optimization for Alignment from Heterogeneous Rewards

SafetyDGX agent

arXiv:2510.04214v3 Announce Type: replace Abstract: We deploy large language models (LLMs) as business development (BD) agents for persuasive price negotiation in online travel agencies (OTAs). The ag

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation

SafetyDGX agent

arXiv:2604.26347v1 Announce Type: cross Abstract: Objective metrics for emotional expressiveness are vital for speech generation, particularly in expressive synthesis and voice conversion requiring em

The Prompt Engineering Report Distilled: Quick Start Guide for Life Sciences

Model ReleasesDGX agent

arXiv:2509.11295v2 Announce Type: replace Abstract: Developing effective prompts demands significant cognitive investment to generate reliable, high-quality responses from Large Language Models (LLMs)

Theory-Grounded Evaluation Exposes the Authorship Gap in LLM Personalization

Model ReleasesDGX agent

arXiv:2604.26460v1 Announce Type: new Abstract: Stylistic personalization - making LLMs write in a specific individual's style, rather than merely adapting to task preferences - lacks evaluation groun

Thinking with Drafting: Optical Decompression via Logical Reconstruction

Model ReleasesDGX agent

arXiv:2602.11731v2 Announce Type: replace Abstract: Existing multimodal large language models have achieved high-fidelity visual perception and exploratory visual generation. However, a precision para

Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall

ResearchDGX agent

arXiv:2505.13963v3 Announce Type: replace Abstract: Quantization methods are widely used to accelerate inference and streamline the deployment of large language models (LLMs). Although quantization's

Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact Match

Model ReleasesDGX agent

arXiv:2511.22972v3 Announce Type: replace Abstract: Large language models (LLMs) achieve strong performance across diverse tasks but suffer from high inference latency due to their autoregressive gene

Verified Critical Step Optimization for LLM Agents

SafetyDGX agent

arXiv:2602.03412v2 Announce Type: replace Abstract: As large language model agents tackle increasingly complex long-horizon tasks, effective post-training becomes critical. Prior work faces fundamenta

VIGNETTE: Socially Grounded Bias Evaluation for Vision-Language Models

Model ReleasesDGX agent

arXiv:2505.22897v2 Announce Type: replace Abstract: While bias in large language models (LLMs) is well-studied, similar concerns in vision-language models (VLMs) have received comparatively less atten

WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation Models

Model ReleasesDGX agent

arXiv:2510.14438v2 Announce Type: replace Abstract: The hallmark of Deep Research agents lies in compositional reasoning, the capacity to aggregate distributed, heterogeneous information into coherent

What Kind of Language is Easy to Language-Model Under Curriculum Learning?

SafetyDGX agent

arXiv:2604.26844v1 Announce Type: new Abstract: Many of the thousands of attested languages share common configurations of features, creating a spectrum from typologically very rare (e.g., object-verb

When Annotators Disagree, Topology Explains: Mapper, a Topological Tool for Exploring Text Embedding Geometry and Ambiguity

SafetyDGX agent

arXiv:2510.17548v2 Announce Type: replace Abstract: Language models are often evaluated with scalar metrics like accuracy, but such measures fail to capture how models internally represent ambiguity,

When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding?

ResearchDGX agent

arXiv:2604.26412v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference, but SOTA hidden-state-based drafters suffer from long-range decay: draft accuracy degrades as the specul

Zero-Shot to Full-Resource: Cross-lingual Transfer Strategies for Aspect-Based Sentiment Analysis

ResearchDGX agent

arXiv:2604.26619v1 Announce Type: new Abstract: Aspect-based Sentiment Analysis (ABSA) extracts fine-grained opinions toward specific aspects within text but remains largely English-focused despite ma

29 Apr 2026

A Blueprint for AI-Driven Software Quality: Integrating LLMs with Established Standards

SafetyDGX agent

arXiv:2505.13766v5 Announce Type: replace-cross Abstract: Software Quality Assurance (SQA) is critical for delivering reliable, secure, and efficient software products. The Software Quality Assurance

A paradox of AI fluency

ResearchDGX agent

arXiv:2604.25905v1 Announce Type: new Abstract: How much does a user's skill with AI shape what AI actually delivers for them? This question is critical for users, AI product builders, and society at

A Survey on LLM-based Conversational User Simulation

ResearchDGX agent

arXiv:2604.24977v1 Announce Type: new Abstract: User simulation has long played a vital role in computer science due to its potential to support a wide range of applications. Language, as the primary

ADE: Adaptive Dictionary Embeddings -- Scaling Multi-Anchor Representations to Large Language Models

Model ReleasesDGX agent

arXiv:2604.24940v1 Announce Type: new Abstract: Word embeddings are fundamental to natural language processing, yet traditional approaches represent each word with a single vector, creating representa

Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation

Model ReleasesDGX agent

arXiv:2602.11224v3 Announce Type: replace-cross Abstract: We present Agent-Diff, a novel benchmarking framework for evaluating agentic Large Language Models (LLMs) on real-world productivity software

Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses

Model ReleasesDGX agent

arXiv:2604.25850v1 Announce Type: new Abstract: Harnesses have become a central determinant of coding-agent performance, shaping how models interact with repositories, tools, and execution environment

An Investigation of Linguistic Biases in LLM-Based Recommendations

Model ReleasesDGX agent

arXiv:2604.25456v1 Announce Type: new Abstract: We investigate linguistic biases in LLM-based restaurant and product recommendations given prompts varying across Southern American English (AE), Indian

Analyzing LLM Reasoning to Uncover Mental Health Stigma

Model ReleasesDGX agent

arXiv:2604.25053v1 Announce Type: new Abstract: While large language models (LLMs) are increasingly being explored for mental health applications, recent studies reveal that they can exhibit stigma to

← Previous
1…9394959697…129
Next →