AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
20 May 2026

A Multi-Agent Framework for Feature-Constrained Difficulty Control in Reading Comprehension Item Generation

AgentsDGX agent

arXiv:2605.19316v1 Announce Type: new Abstract: Recent studies in difficulty-controlled reading comprehension item generation have leveraged large language models (LLMs) to produce items by adjusting

Acoustic scattering AI for non-invasive object classifications: A case study on hair assessment

Model ReleasesDGX agent

arXiv:2506.14148v2 Announce Type: replace-cross Abstract: This paper presents a novel non-invasive object classification approach using acoustic scattering, demonstrated through a case study on hair a

Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.19149v1 Announce Type: new Abstract: Agents operating with computer and Web use inevitably encounter errors: inaccessible webpages, missing files, local and remote misconfigurations, etc. T

An LLM-Based System for Argument Mining

Model ReleasesDGX agent

arXiv:2605.13793v2 Announce Type: replace Abstract: Arguments are a fundamental aspect of human reasoning, in which claims are supported, challenged, and weighed against one another. We present an end

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning

Model ReleasesDGX agent

arXiv:2605.19852v1 Announce Type: new Abstract: Tool-augmented reasoning has emerged as a promising direction for enhancing the reasoning capabilities of multimodal large language models (MLLMs). Howe

CAIT: A Syntactic Parsing Toolkit for Child-Adult InTeractions

ApplicationsDGX agent

arXiv:2605.19718v1 Announce Type: new Abstract: CHILDES is a paramount resource for language acquisition studies -- yet computational tools for analyzing its syntactic structure remain limited. Levera

Can Large Language Models Reliably Correct Errors in Low-Resource ASR? A Contamination-Aware Case Study on West Frisian

Model ReleasesDGX agent

arXiv:2605.19711v1 Announce Type: new Abstract: Automatic speech recognition (ASR) has improved substantially in recent years, yet performance remains limited for low-resource languages. Large languag

Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items?

ResearchDGX agent

arXiv:2510.25064v2 Announce Type: replace Abstract: Estimating the cognitive complexity of reading comprehension (RC) items is crucial for assessing item difficulty before it is administered to learne

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization

SafetyDGX agent

arXiv:2605.19436v1 Announce Type: cross Abstract: When a model produces a correct solution under reinforcement learning with verifiable rewards (RLVR), every token receives the same reward signal rega

CLIF: Concept-Level Influence Functions for Transparent Bottleneck Models

ApplicationsDGX agent

arXiv:2605.19848v1 Announce Type: new Abstract: In recent years, the black-box nature of deep learning models has limited their application in high-stakes domains such as medical diagnosis and finance

ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning

Model ReleasesDGX agent

arXiv:2605.20176v1 Announce Type: new Abstract: Large language models (LLMs) and agentic systems have shown promise for clinical decision support, but existing works largely assume that evidence has a

Context-Aware Detection and Victim-Centered Response Generation for Online Harassment in Private Messaging

ResearchDGX agent

arXiv:2512.14700v2 Announce Type: replace-cross Abstract: Online harassment is a widespread social and public health concern, yet most computational approaches for detecting and addressing harassment

Critique-Guided Distillation for Robust Reasoning via Refinement

TutorialsDGX agent

arXiv:2505.11628v4 Announce Type: replace Abstract: Supervised fine-tuning with expert demonstrations often produces models that imitate outputs without internalizing the reasoning processes needed fo

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models

SafetyDGX agent

arXiv:2510.13293v3 Announce Type: replace Abstract: While Text-to-Speech (TTS) systems enable emotional control via natural-language instructions, expressiveness, naturalness, and speech quality degra

Cubit: Token Mixer with Kernel Ridge Regression

ResearchDGX agent

arXiv:2605.06501v2 Announce Type: replace-cross Abstract: Since its introduction in 2017, the Transformer has become one of the most widely adopted architectures in modern deep learning. Despite exten

DECOR: Auditing LLM Deception via Information Manipulation Theory

AgentsDGX agent

arXiv:2605.19270v1 Announce Type: new Abstract: Large language models can deceive by subtly manipulating truthful information -- omitting key facts, shifting focus, or obscuring meaning -- making such

Difficulty-Controllable Cloze Question Distractor Generation

ResearchDGX agent

arXiv:2511.01526v2 Announce Type: replace Abstract: Multiple-choice cloze questions are commonly used to assess linguistic proficiency and comprehension. However, generating high-quality distractors r

Drifting Objectives for Refining Discrete Diffusion Language Models

TutorialsDGX agent

arXiv:2605.19470v1 Announce Type: new Abstract: Discrete diffusion language models (DDLMs) generate text by iteratively denoising categorical token sequences, while recent drifting methods for continu

ECG-R1: Protocol-Guided and Modality-Agnostic MLLM for Reliable ECG Interpretation

ResearchDGX agent

arXiv:2602.04279v2 Announce Type: replace Abstract: Electrocardiography (ECG) serves as an indispensable diagnostic tool in clinical practice, yet existing multimodal large language models (MLLMs) rem

Efficient Pre-Training with Token Superposition

ResearchDGX agent

arXiv:2605.06546v2 Announce Type: replace Abstract: Pre-training of Large Language Models is often prohibitively expensive and inefficient at scale, requiring complex and invasive modifications in ord

EnsemHalDet: Robust VLM Hallucination Detection via Ensemble of Internal State Detectors

ResearchDGX agent

arXiv:2604.02784v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) excel at multimodal tasks, but they remain vulnerable to hallucinations that are factually incorrect or unground

Federated Learning for ICD Classification with Lightweight Models and Pretrained Embeddings

Model ReleasesDGX agent

arXiv:2507.03122v2 Announce Type: replace-cross Abstract: This study investigates the feasibility and performance of federated learning (FL) for multi-label ICD code classification using clinical note

FedMental: Evaluating Federated Learning for Mental Health Detection from Social Media Data

Model ReleasesDGX agent

arXiv:2605.18936v1 Announce Type: cross Abstract: Social media text data are often used to train Machine Learning (ML) models to identify users exhibiting high-risk mental health behaviors. However, s

Fine-tuning language encoding models on slow fMRI improves prediction for fast ECoG

ResearchDGX agent

arXiv:2605.19224v1 Announce Type: new Abstract: Neuroscientists have recently turned to intracranial brain recording methods, like electrocorticography (ECoG), for human experiments because of the fin

Fingerprinting LLMs via Prompt Injection

ResearchDGX agent

arXiv:2509.25448v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are often modified after release through post-processing such as post-training or quantization, which makes it ch

FlexDraft: Flexible Speculative Decoding via Attention Tuning and Bonus-Guided Calibration

ResearchDGX agent

arXiv:2605.20022v1 Announce Type: new Abstract: Speculative decoding accelerates memory-bound LLM inference without quality degradation by using a fast drafter to propose multiple candidate tokens and

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models

ResearchDGX agent

arXiv:2605.20177v1 Announce Type: new Abstract: Recent advances in vision-language models (VLMs) emphasize long chain-of-thought reasoning; yet, we find that their performance on visual tasks is prima

GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment

Model ReleasesDGX agent

arXiv:2605.19577v1 Announce Type: new Abstract: We present GoLongRL, a fully open-source, capability-oriented post-training recipe for long-context reinforcement learning with verifiable rewards (RLVR

GRAB: A Risk Taxonomy--Grounded Benchmark for Unsupervised Topic Discovery in Financial Disclosures

Model ReleasesDGX agent

arXiv:2509.21698v2 Announce Type: replace Abstract: Risk categorization in 10-K risk disclosures matters for oversight and investment, yet no public benchmark evaluates unsupervised topic models for t

HALvest-Contrastive: Retrieval-Like Authorship Attribution with Patch-Level Late Interaction

ResearchDGX agent

arXiv:2407.20595v4 Announce Type: replace-cross Abstract: Deciding whether two pieces of text share an author is made difficult by topical confound: two writers covering the same topic often look more

How Do Document Parsers Break? Auditing Structural Vulnerability in Document Intelligence

SafetyDGX agent

arXiv:2605.19309v1 Announce Type: new Abstract: Document Layout Analysis (DLA) pipelines provide structured page representations for retrieval-augmented generation, long-document question answering, a

K-Quantization and its Impact on Output Performance

Model ReleasesDGX agent

arXiv:2605.19645v1 Announce Type: new Abstract: Recent advancements in large language models (LLMs) have shown their remarkable capacities in many NLP tasks. However, their substantial size often pres

KoRe: Compact Knowledge Representations for Large Language Models

ResearchDGX agent

arXiv:2605.20170v1 Announce Type: new Abstract: Modern Large Language Models (LLMs) have shown impressive performances in user-facing tasks such as question answering, as well as consistent improvemen

LambdaPO: A Lambda Style Policy Optimization for Reasoning Language Models

SafetyDGX agent

arXiv:2605.19416v1 Announce Type: new Abstract: Group Relative Policy Optimization(GRPO) has become a cornerstone of modern reinforcement learning alignment, prized for its efficacy in foregoing an ex

Language models struggle with compartmentalization

TutorialsDGX agent

arXiv:2605.19284v1 Announce Type: new Abstract: In the training data used by large language models (LLMs), the same latent concept is often presented in multiple distinct ways: the same facts appear i

Language Mutations Sustain the Persistences of Conspiracy Theories on Social Media

ResearchDGX agent

arXiv:2605.20050v1 Announce Type: new Abstract: This study investigates how language mutations affect the persistent diffusion of conspiracy theories on social media. Drawing on a three-year dataset o

Library Hallucinations in LLM-Generated Code: A Risk Analysis Grounded in Developer Queries

Model ReleasesDGX agent

arXiv:2509.22202v3 Announce Type: replace-cross Abstract: Large language models (LLMs) now play a central role in code generation, yet they continue to hallucinate, frequently inventing non-existent l

LLM-Based Financial Sentiment Analysis in Arabic: Evidence from Saudi Markets

ResearchDGX agent

arXiv:2605.19714v1 Announce Type: new Abstract: Investor sentiment shapes financial markets, yet modeling sentiment in Arabic financial contexts remains challenging due to linguistic complexity and li

LLM-MC-Affect: LLM-Based Monte Carlo Modeling of Affective Trajectories and Latent Ambiguity for Interpersonal Dynamic Insight

ApplicationsDGX agent

arXiv:2601.03645v2 Announce Type: replace Abstract: Emotional coordination is a core property of human interaction that shapes how relational meaning is constructed in real time. While text-based affe

LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening

Model ReleasesDGX agent

arXiv:2605.19597v1 Announce Type: new Abstract: Evaluating large language models (LLMs) on natural-language logical reasoning is essential because rule-governed tasks require conclusions to follow str

Lost in Interpretation: The Plausibility-Faithfulness Trade-off in Cross-Lingual Explanations

ResearchDGX agent

arXiv:2605.19274v1 Announce Type: new Abstract: LLMs deployed multilingually are often audited via English explanations for non-English inputs. We evaluate extractive explanations ''where the model id

m3BERT: A Modern, Multi-lingual, Matryoshka Bidirectional Encoder

Model ReleasesDGX agent

arXiv:2605.19568v1 Announce Type: new Abstract: Embedding models are pivotal in industrial information retrieval systems like search and advertising. However, existing pretrained models often exhibit

Measuring Stereotype and Deviation Biases in Large Language Models

SafetyDGX agent

arXiv:2508.06649v3 Announce Type: replace Abstract: Large language models (LLMs) are widely applied across diverse domains, raising concerns about their limitations and potential risks. In this study,

Mind Your Moras: Orthography-Aware Error Analysis of Neural Japanese Morphological Generation

ResearchDGX agent

arXiv:2605.20043v1 Announce Type: new Abstract: We present an orthography-aware error analysis of Japanese past-tense morphological inflection, treating hiragana not merely as a transcriptional medium

MixRea: Benchmarking Explicit-Implicit Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2605.20128v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into high-stakes decision-making. Inspired by the theory of inattentional blindness in human co

MMoA: An AI-Agent framework with recurrence for Memoried Mixure-of-Agent

AgentsDGX agent

arXiv:2605.19194v1 Announce Type: new Abstract: The Mixture-of-Agents (MoA) framework has shown promise in improving large language model (LLM) performance by aggregating outputs from multiple agents.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training

Model ReleasesDGX agent

arXiv:2510.18830v2 Announce Type: replace Abstract: The adoption of long context windows has become a standard feature in Large Language Models (LLMs), as extended contexts significantly enhance their

OpenCompass: A Universal Evaluation Platform for Large Language Models

Model ReleasesDGX agent

arXiv:2605.19276v1 Announce Type: new Abstract: In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large lang

OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond

HardwareDGX agent

arXiv:2605.19660v1 Announce Type: cross Abstract: The rapid advancement toward long-context reasoning and multi-modal intelligence has made the memory footprint of the Key-Value (KV) cache a dominant

PASC: Pipeline-Aware Conformal Prediction with Joint Coverage Guarantees for Multi-Stage NLP and LLM Pipelines

AgentsDGX agent

arXiv:2605.18812v1 Announce Type: cross Abstract: Modern NLP and LLM systems are pipelines: named entity recognition (NER) -> entity disambiguation (NED) -> entity typing, retrieval-augmented generati

Prompting language influences diagnostic reasoning and accuracy of large language models

Model ReleasesDGX agent

arXiv:2605.19173v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly explored for clinical decision support, yet most evaluations are conducted in English, leaving their relia

Qayyem: A Real-time Platform for Scoring Proficiency of Arabic Essays

ResearchDGX agent

arXiv:2603.01009v2 Announce Type: replace Abstract: Over the past years, Automated Essay Scoring (AES) systems have gained increasing attention as scalable and consistent solutions for assessing the p

Rethinking How to Remember: Beyond Atomic Facts in Lifelong LLM Agent Memory

Model ReleasesDGX agent

arXiv:2605.19952v1 Announce Type: new Abstract: To enable reliable long-term interaction, LLM agents require a memory system that can faithfully store, efficiently retrieve, and deeply reason over acc

Retrieval-Augmented Generation for Natural Language Processing: A Survey

Model ReleasesDGX agent

arXiv:2407.13193v4 Announce Type: replace Abstract: Large language models (LLMs) have achieved strong empirical performance in various fields, benefiting from their huge amount of parameters that stor

Retrieval-Augmented Linguistic Calibration

ResearchDGX agent

arXiv:2605.19344v1 Announce Type: new Abstract: Linguistic cues such as 'I believe' and 'probably' offer an intuitive interface for communicating confidence, yet a generalisable, principled calibratio

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents

SafetyDGX agent

arXiv:2605.20061v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) is a promising paradigm for improving large language model (LLM) agents on long-horizon interactiv

Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior

Model ReleasesDGX agent

arXiv:2510.14261v2 Announce Type: replace Abstract: We present an experimental recipe for studying the relationship between training data and language model (LM) behavior. We outline steps for interve

Scaling Evaluation-time Compute with Reasoning Models as Evaluators

ResearchDGX agent

arXiv:2503.19877v2 Announce Type: replace Abstract: As language model (LM) outputs get more and more natural, it is becoming more difficult than ever to evaluate their quality. Simultaneously, increas

SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models

Model ReleasesDGX agent

arXiv:2605.19357v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly applied to scientific research, yet existing evaluations often fail to reflect the fine-grained capabiliti

Self-Filtered Distillation with LLMs-generated Trust Indicators for Reliable Patent Classification

Model ReleasesDGX agent

arXiv:2510.05431v4 Announce Type: replace Abstract: Organizing large-scale patent corpora according to classification schemes is a core information management task that determines the accuracy and eff

← Previous
1…6869707172…129
Next →