AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
6 May 2026

LLM-XTM: Enhancing Cross-Lingual Topic Models with Large Language Models

SafetyDGX agent

arXiv:2605.03299v1 Announce Type: new Abstract: Cross-lingual topic modeling aims to discover shared semantic structures across languages, yet existing models depend on sparse bilingual resources and

Logical Consistency as a Bridge: Improving LLM Hallucination Detection via Label Constraint Modeling between Responses and Self-Judgments

ApplicationsDGX agent

arXiv:2605.03971v1 Announce Type: new Abstract: Large Language Models (LLMs) are prone to factual hallucinations, risking their reliability in real-world applications. Existing hallucination detectors

MAGE: Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.03228v1 Announce Type: cross Abstract: As large language model (LLM)-powered agents are increasingly deployed to perform complex, real-world tasks, they face a growing class of attacks that

Maximizing mutual information between prompts and responses improve LLM personalization with no additional data or human oversight

Model ReleasesDGX agent

arXiv:2603.19294v2 Announce Type: replace-cross Abstract: While post-training has successfully improved large language models (LLMs) across a variety of domains, these gains heavily rely on human-labe

MCJudgeBench: A Benchmark for Constraint-Level Judge Evaluation in Multi-Constraint Instruction Following

Model ReleasesDGX agent

arXiv:2605.03858v1 Announce Type: new Abstract: Multi-constraint instruction following requires verifying whether a response satisfies multiple individual requirements, yet LLM judges are often assess

Mechanism-Faithful Queueing Simulation Model Translation with Large Language Model Support

ResearchDGX agent

arXiv:2601.06543v2 Announce Type: replace Abstract: Queueing simulation studies often require substantial manual effort to translate conceptual system descriptions into executable programs and to veri

MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports

Model ReleasesDGX agent

arXiv:2605.03103v1 Announce Type: new Abstract: Semi-structured information extraction (IE) from OCR-derived clinical reports is crucial for efficiently reconstructing patients' longitudinal medical h

MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue

ResearchDGX agent

arXiv:2603.06194v2 Announce Type: replace Abstract: Reinforcement learning (RL) for large language models (LLMs) has shown strong performance in single-turn tasks, but extending it to multi-turn inter

MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier

HardwareDGX agent

arXiv:2603.03756v3 Announce Type: replace-cross Abstract: While large language models (LLMs) show promise in scientific discovery, existing research focuses on inference or feedback-driven training, l

Multilingual Safety Alignment via Self-Distillation

SafetyDGX agent

arXiv:2605.02971v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit severe multilingual safety misalignment: they possess strong safeguards in high-resource languages but remain hig

Natural Language Processing: A Comprehensive Practical Guide from Tokenisation to RLHF

TutorialsDGX agent

arXiv:2605.03799v1 Announce Type: new Abstract: This preprint presents a systematic, research-oriented practicum that guides the reader through the entire modern NLP pipeline: from tokenisation and ve

Not that Groove: Zero-Shot Symbolic Music Editing

Model ReleasesDGX agent

arXiv:2505.08203v2 Announce Type: replace-cross Abstract: While recent advancements in AI music generation have predominantly focused on direct audio synthesis, these systems suffer from inherent rigi

OCRR: A Benchmark for Online Correction Recovery under Distribution Shift

Model ReleasesDGX agent

arXiv:2605.03153v1 Announce Type: cross Abstract: Static benchmarks measure a model frozen at training time. Real systems face distribution shift: new categories, paraphrased queries, drift: and must

On-Device Fine-Tuning via Backprop-Free Zeroth-Order Optimization

Local AiDGX agent

arXiv:2511.11362v2 Announce Type: replace-cross Abstract: On-device fine-tuning is a critical capability for edge AI systems, which must support adaptation to different agentic tasks under stringent m

On Verbalized Confidence Scores for LLMs

Model ReleasesDGX agent

arXiv:2412.14737v2 Announce Type: replace Abstract: The rise of large language models (LLMs) and their tight integration into our daily life make it essential to dedicate efforts towards their trustwo

OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty Trajectories

AgentsDGX agent

arXiv:2605.04036v1 Announce Type: cross Abstract: Deep search capabilities have become an indispensable competency for frontier Large Language Model (LLM) agents, yet their development remains dominat

PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examination

Model ReleasesDGX agent

arXiv:2605.03571v1 Announce Type: new Abstract: Patent examination is a complex, multi-stage process requiring both technical expertise and legal reasoning, increasingly challenged by rising applicati

Permutation-Consensus Listwise Judging for Robust Factuality Evaluation

ResearchDGX agent

arXiv:2603.20562v2 Announce Type: replace Abstract: Large language models (LLMs) are now widely used as judges, yet their decisions can change under presentation choices that should be irrelevant. We

PIIGuard: Mitigating PII Harvesting under Adversarial Sanitization

Model ReleasesDGX agent

arXiv:2605.03129v1 Announce Type: cross Abstract: Browsing-enabled LLM assistants can fetch webpages and answer contact-seeking queries, creating a practical channel for scraping contact-style persona

psifx -- Psychological and Social Interactions Feature Extraction Package

ResearchDGX agent

arXiv:2407.10266v5 Announce Type: replace Abstract: psifx is a plug-and-play multi-modal feature extraction toolkit, aiming to facilitate and democratize the use of state-of-the-art machine learning t

RAG over Thinking Traces Can Improve Reasoning Tasks

Model ReleasesDGX agent

arXiv:2605.03344v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) has proven effective for knowledge-intensive tasks, but is widely believed to offer limited benefit for reasoning

Rational Communication Shapes Morphological Composition

ApplicationsDGX agent

arXiv:2605.03510v1 Announce Type: new Abstract: Human languages expand vocabularies by combining existing morphemes rather than inventing arbitrary forms. Communicative efficiency shapes lexical syste

ReCode: Reinforcing Code Generation with Reasoning-Process Rewards

Model ReleasesDGX agent

arXiv:2508.05170v3 Announce Type: replace-cross Abstract: In practice, rigorous reasoning is often a key driver of correct code, while Reinforcement Learning (RL) for code generation often neglects op

Reproducing Complex Set-Compositional Information Retrieval

Model ReleasesDGX agent

arXiv:2605.03824v1 Announce Type: new Abstract: Complex information needs may involve set-compositional queries using conjunction, disjunction, and exclusion, yet it remains unclear whether current re

Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems

Model ReleasesDGX agent

arXiv:2605.04018v1 Announce Type: new Abstract: Reasoning-intensive retrieval aims to surface evidence that supports downstream reasoning rather than merely matching topical similarity. This capabilit

Retrieving Floods without Floodlights: Topic Models as Binary Classifiers for Extreme Climate Events in German News

TutorialsDGX agent

arXiv:2605.03450v1 Announce Type: new Abstract: In studies of media coverage of extreme climate events, NLP methods have become indispensable for identifying relevant texts in large news databases. St

Revisiting Graph-Tokenizing Large Language Models: A Systematic Evaluation of Graph Token Understanding

ResearchDGX agent

arXiv:2605.03514v1 Announce Type: new Abstract: The remarkable success of large language models (LLMs) has motivated researchers to adapt them as universal predictors for various graph tasks. As a wid

Robust Language Identification for Romansh Varieties

Model ReleasesDGX agent

arXiv:2603.15969v2 Announce Type: replace Abstract: The Romansh language has several regional varieties, called idioms, which sometimes have limited mutual intelligibility. Despite this linguistic div

Rose-SQL: Role-State Evolution Guided Structured Reasoning for Multi-Turn Text-to-SQL

ResearchDGX agent

arXiv:2605.03720v1 Announce Type: new Abstract: Recent advances in Large Reasoning Models (LRMs) trained with Long Chain-of-Thought have demonstrated remarkable capabilities in code generation and mat

S^2tory: Story Spine Distillation for Movie Script Summarization

AgentsDGX agent

arXiv:2605.03244v1 Announce Type: new Abstract: Movie scripts pose a fundamental challenge for automatic summarization due to their non-linear, cross-cut narrative structure, which makes surface-level

Safety and accuracy follow different scaling laws in clinical large language models

Model ReleasesDGX agent

arXiv:2605.04039v1 Announce Type: new Abstract: Clinical LLMs are often scaled by increasing model size, context length, retrieval complexity, or inference-time compute, with the implicit expectation

SAM-NER: Semantic Archetype Mediation for Zero-Shot Named Entity Recognition

Model ReleasesDGX agent

arXiv:2605.03706v1 Announce Type: new Abstract: Zero-shot Named Entity Recognition (ZS-NER) remains brittle under domain and schema shifts, where unseen label definitions often misalign with a large l

Scoring Edit Impact in Grammatical Error Correction via Embedded Association Graphs

ResearchDGX agent

arXiv:2604.06573v2 Announce Type: replace Abstract: A Grammatical Error Correction (GEC) system produces a sequence of edits to correct an erroneous sentence. The quality of these edits is typically e

SEAD: Self-Evolving Agent for Multi-Turn Service Dialogue

AgentsDGX agent

arXiv:2602.03548v3 Announce Type: replace Abstract: Large Language Models have demonstrated remarkable capabilities in open-domain dialogues. However, current methods exhibit suboptimal performance in

Segmenting Human-LLM Co-authored Text via Change Point Detection

Local AiDGX agent

arXiv:2605.03723v1 Announce Type: new Abstract: The rise of large language models (LLMs) has created an urgent need to distinguish between human-written and LLM-generated text to ensure authenticity a

Semantically Enriching Investor Micro-blogs for Opinion-Aware Emotion Analysis: A Practical Approach

ResearchDGX agent

arXiv:2605.03092v1 Announce Type: new Abstract: While sentiment analysis is the staple of financial NLP, capturing the nuances of 'why' behind that sentiment remains a challenge. There have been attem

Sentiment Analysis of Indonesian Spotify Reviews Using Machine Learning and BiLSTM

ResearchDGX agent

arXiv:2605.03443v1 Announce Type: new Abstract: This paper benchmarks classical machine learning and deep learning approaches for three-class sentiment classification of Indonesian Spotify reviews. Us

SERE: Structural Example Retrieval for Enhancing LLMs in Event Causality Identification

SafetyDGX agent

arXiv:2605.03701v1 Announce Type: new Abstract: Event Causality Identification (ECI) requires models to determine whether a given pair of events in a context exhibits a causal relationship. While Larg

SHIELD: A Diverse Clinical Note Dataset and Distilled Small Language Models for Enterprise-Scale De-identification

Local AiDGX agent

arXiv:2605.03301v1 Announce Type: new Abstract: De-identification of clinical text remains essential for secondary use of electronic health records (EHRs), yet public benchmarks such as i2b2 2006/2014

Should We Still Pretrain Encoders with Masked Language Modeling?

ResearchDGX agent

arXiv:2507.00994v4 Announce Type: replace Abstract: Learning high-quality text representations is fundamental to a wide range of NLP tasks. While encoder pretraining has traditionally relied on Masked

Simulated Students in Tutoring Dialogues: Substance or Illusion?

Model ReleasesDGX agent

arXiv:2601.04025v2 Announce Type: replace Abstract: Advances in large language models (LLMs) enable many new innovations in education. However, evaluating the effectiveness of new technology requires

Sparse Memory Finetuning as a Low-Forgetting Alternative to LoRA and Full Finetuning

Model ReleasesDGX agent

arXiv:2605.03229v1 Announce Type: new Abstract: Adapting a pretrained language model to a new task often hurts the general capabilities it already had, a problem known as catastrophic forgetting. Spar

Steer Like the LLM: Activation Steering that Mimics Prompting

ResearchDGX agent

arXiv:2605.03907v1 Announce Type: new Abstract: Large language models can be steered at inference time through prompting or activation interventions, but activation steering methods often underperform

Stochastic Attention: Connectome-Inspired Randomized Routing for Expressive Linear-Time Attention

Local AiDGX agent

arXiv:2604.00754v2 Announce Type: replace Abstract: The whole-brain connectome of a fruit fly comprises over 130K neurons connected with a probability of merely 0.02%, yet achieves an average shortest

SURE-RAG: Sufficiency and Uncertainty-Aware Evidence Verification for Selective Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2605.03534v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) grounds answers in retrieved passages, but retrieval is not verification: a passage can be topical and still fail t

Task Vector Geometry Underlies Dual Modes of Task Inference in Transformers

ResearchDGX agent

arXiv:2605.03780v1 Announce Type: cross Abstract: Transformers are effective at inferring the latent task from context via two inference modes: recognizing a task seen during training, and adapting to

TeamUp: Semantic Project Matching and Team Formation for Learning at Scale

SafetyDGX agent

arXiv:2605.03237v1 Announce Type: cross Abstract: Project-based learning improves student engagement and learning outcomes, yet allocating students to appropriately challenging projects while forming

The Counterexample Game: Iterated Conceptual Analysis and Repair in Language Models

ResearchDGX agent

arXiv:2605.03936v1 Announce Type: new Abstract: Conceptual analysis -- proposing definitions and refining them through counterexamples -- is central to philosophical methodology. We study whether lang

The Enemy from Within: A Study of Political Delegitimization Discourse in Israeli Political Speech

ResearchDGX agent

arXiv:2508.15524v3 Announce Type: replace Abstract: We present the first large-scale computational study of political delegitimization discourse (PDD), defined as symbolic attacks on the normative val

The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm

HardwareDGX agent

arXiv:2505.16932v5 Announce Type: replace-cross Abstract: Computing the polar decomposition and the related matrix sign function has been a well-studied problem in numerical analysis for decades. Rece

The Right Answer, the Wrong Direction: Why Transformers Fail at Counting and How to Fix It

Model ReleasesDGX agent

arXiv:2605.03258v1 Announce Type: cross Abstract: Large language models often fail at simple counting tasks, even when the items to count are explicitly present in the prompt. We investigate whether t

The TTS-STT Flywheel: Synthetic Entity-Dense Audio Closes the Indic ASR Gap Where Commercial and Open-Source Systems Fail

Model ReleasesDGX agent

arXiv:2605.03073v1 Announce Type: new Abstract: Niche-domain Indic ASR -- digit strings, currency amounts, addresses, brand names, English/Indic codemix -- is under-served by both open-source SOTA and

To Write or to Automate Linguistic Prompts, That Is the Question

ResearchDGX agent

arXiv:2603.25169v2 Announce Type: replace Abstract: LLM performance is highly sensitive to prompt design, yet whether automatic prompt optimization can replace expert prompt engineering in linguistic

TRACE: A Metrologically-Grounded Engineering Framework for Trustworthy Agentic AI Systems in Operationally Critical Domains

SafetyDGX agent

arXiv:2605.03838v1 Announce Type: new Abstract: We introduce TRACE, a cross-domain engineering framework for trustworthy agentic AI in operationally critical domains. TRACE combines a four-layer refer

Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection

Local AiDGX agent

arXiv:2605.02958v1 Announce Type: cross Abstract: Representation Engineering typically relies on static refusal vectors derived from terminal representations. We move beyond this paradigm, demonstrati

Transformers with Selective Access to Early Representations

ResearchDGX agent

arXiv:2605.03953v1 Announce Type: cross Abstract: Several recent Transformer architectures expose later layers to representations computed in the earliest layers, motivated by the observation that low

TriBench-Ko: Evaluating LLM Risks in Judicial Workflows

Model ReleasesDGX agent

arXiv:2605.03792v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into legal workflows. However, existing benchmarks primarily address proxy tasks, such as bar e

Two Calls, Two Moments, and the Vote-Accuracy Curve of Repeated LLM Inference

ResearchDGX agent

arXiv:2605.03379v1 Announce Type: cross Abstract: Repeated sampling is a standard way to spend test-time compute, but its benefit is controlled by the latent distribution of correctness across example

Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks

SafetyDGX agent

arXiv:2502.04419v3 Announce Type: replace-cross Abstract: Generating synthetic datasets via large language models (LLMs) has emerged as a promising approach to improve LLM performance. However, LLMs i

Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development

Model ReleasesDGX agent

arXiv:2603.04601v2 Announce Type: replace-cross Abstract: Code generation has emerged as one of AI's highest-impact use cases, yet existing benchmarks measure isolated tasks rather than the complete '

← Previous
1…8586878889…129
Next →