AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
27 Apr 2026

Source-Modality Monitoring in Vision-Language Models

AgentsDGX agent

arXiv:2604.22038v1 Announce Type: new Abstract: We define and investigate source-modality monitoring -- the ability of multimodal models to track and communicate the input source from which pieces of

STEM: Structure-Tracing Evidence Mining for Knowledge Graphs-Driven Retrieval-Augmented Generation

ResearchDGX agent

arXiv:2604.22282v1 Announce Type: new Abstract: Knowledge Graph-based Question Answering (KGQA) plays a pivotal role in complex reasoning tasks but remains constrained by two persistent challenges: th

Survey Response Generation: Generating Closed-Ended Survey Responses In-Silico with Large Language Models

SafetyDGX agent

arXiv:2510.11586v2 Announce Type: replace Abstract: Many in-silico simulations of human survey responses with large language models (LLMs) focus on generating closed-ended survey responses, whereas LL


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

System-Mediated Attention Imbalances Make Vision-Language Models Say Yes

SafetyDGX agent

arXiv:2601.12430v2 Announce Type: replace Abstract: Vision-language model (VLM) hallucination is commonly linked to imbalanced allocation of attention across input modalities: system, image and text.

The Bitter Lesson of Diffusion Language Models for Agentic Workflows: A Comprehensive Reality Check

AgentsDGX agent

arXiv:2601.12979v3 Announce Type: replace Abstract: The pursuit of real-time agentic interaction has driven interest in Diffusion-based Large Language Models (dLLMs) as alternatives to auto-regressive

Thinking Without Words: Efficient Latent Reasoning with Abstract Chain-of-Thought

SafetyDGX agent

arXiv:2604.22709v1 Announce Type: new Abstract: While long, explicit chains-of-thought (CoT) have proven effective on complex reasoning tasks, they are costly to generate during inference. Non-verbal

Toward Automated Robustness Evaluation of Mathematical Reasoning

Model ReleasesDGX agent

arXiv:2506.05038v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in various reasoning-intensive tasks. However, these models exhibit unexpecte

Tracing the complexity profiles of different linguistic phenomena through the intrinsic dimension of LLM representations

ResearchDGX agent

arXiv:2601.03779v2 Announce Type: replace Abstract: We explore intrinsic dimension (ID) of LLM representations as a marker of linguistic complexity. Specifically, we test whether ID differences across

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis

SafetyDGX agent

arXiv:2604.22225v1 Announce Type: new Abstract: While generative text-to-speech (TTS) models approach human-level quality, monolithic metrics fail to diagnose fine-grained acoustic artifacts or explai

UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents

Model ReleasesDGX agent

arXiv:2602.07038v2 Announce Type: replace-cross Abstract: Key Information Extraction (KIE) from real-world documents remains challenging due to substantial variations in layout structures, visual qual

Using Embedding Models to Improve Probabilistic Race Prediction

ResearchDGX agent

arXiv:2604.22555v1 Announce Type: new Abstract: Estimating racial disparity requires individual-level race data, which are often unavailable due to the sensitivity of collecting such information. To a

Voice Under Revision: Large Language Models and the Normalization of Personal Narrative

ResearchDGX agent

arXiv:2604.22142v1 Announce Type: new Abstract: This study examines how large language model rewriting alters the style and narrative texture of personal narratives. It analyzes 300 personal narrative

When Cow Urine Cures Constipation on YouTube: Limits of LLMs in Detecting Culture-specific Health Misinformation

Model ReleasesDGX agent

arXiv:2604.22002v1 Announce Type: new Abstract: Social media platforms have become primary channels for health information in the Global South. Using gomutra (cow urine) discourse on YouTube in India

Where Should LoRA Go? Component-Type Placement in Hybrid Language Models

ResearchDGX agent

arXiv:2604.22127v1 Announce Type: new Abstract: Hybrid language models that interleave attention with recurrent components are increasingly competitive with pure Transformers, yet standard LoRA practi

Zero-Shot Morphological Discovery in Low-Resource Bantu Languages via Cross-Lingual Transfer and Unsupervised Clustering

ResearchDGX agent

arXiv:2604.22723v1 Announce Type: cross Abstract: We present a method for discovering morphological features in low-resource Bantu languages by combining cross-lingual transfer learning with unsupervi

24 Apr 2026

AFRILANGTUTOR: Advancing Language Tutoring and Culture Education in Low-Resource Languages with Large Language Models

Model ReleasesDGX agent

arXiv:2604.20996v1 Announce Type: new Abstract: How can language learning systems be developed for languages that lack sufficient training resources? This challenge is increasingly faced by developers

AgentGL: Towards Agentic Graph Learning with LLMs via Reinforcement Learning

SafetyDGX agent

arXiv:2604.05846v2 Announce Type: replace Abstract: Large Language Models (LLMs) increasingly rely on agentic capabilities-iterative retrieval, tool use, and decision-making-to overcome the limits of

AgenticQwen: Training Small Agentic Language Models with Dual Data Flywheels for Industrial-Scale Tool Use

AgentsDGX agent

arXiv:2604.21590v1 Announce Type: new Abstract: Modern industrial applications increasingly demand language models that act as agents, capable of multi-step reasoning and tool use in real-world settin

AITP: Traffic Accident Responsibility Allocation via Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2604.20878v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress in Traffic Accident Detection (TAD) and Traffic Accident Understanding (TAU).

AUDITA: A New Dataset to Audit Humans vs. AI Skill at Audio QA

Model ReleasesDGX agent

arXiv:2604.21766v1 Announce Type: new Abstract: Existing audio question answering benchmarks largely emphasize sound event classification or caption-grounded queries, often enabling models to succeed

Automating Computational Reproducibility in Social Science: Comparing Prompt-Based and Agent-Based Approaches

AgentsDGX agent

arXiv:2602.08561v3 Announce Type: replace-cross Abstract: Reproducing computational research is often assumed to be as simple as rerunning the original code with provided data. In practice, missing pa

Back to the Future: The Role of Past and Future Context Predictability in Incremental Language Production

ApplicationsDGX agent

arXiv:2511.07752v3 Announce Type: replace Abstract: Contextual predictability shapes how we choose and encode words in production. The effects of a word's predictability given preceding or past contex

Beyond N-gram: Data-Aware X-GRAM Extraction for Efficient Embedding Parameter Scaling

Model ReleasesDGX agent

arXiv:2604.21724v1 Announce Type: new Abstract: Large token-indexed lookup tables provide a compute-decoupled scaling path, but their practical gains are often limited by poor parameter efficiency and

Beyond Pixels: Introspective and Interactive Grounding for Visualization Agents

Model ReleasesDGX agent

arXiv:2604.21134v1 Announce Type: new Abstract: Vision-Language Models (VLMs) frequently misread values, hallucinate details, and confuse overlapping elements in charts. Current approaches rely solely

Capabilities and Evaluation Biases of Large Language Models in Classical Chinese Poetry Generation: A Case Study on Tang Poetry

ApplicationsDGX agent

arXiv:2510.15313v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly applied to creative domains, yet their performance in classical Chinese poetry generation and evaluati

CARE: Counselor-Aligned Response Engine for Online Mental-Health Support

SafetyDGX agent

arXiv:2604.21352v1 Announce Type: new Abstract: Mental health challenges are increasing worldwide, straining emotional support services and leading to counselor overload. This can result in delayed re

CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning

SafetyDGX agent

arXiv:2509.20712v5 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a powerful paradigm for optimizing large language models (LLMs) to handle complex reasoning tasks. A co

CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents

Model ReleasesDGX agent

arXiv:2604.21308v1 Announce Type: cross Abstract: Enterprise LLM agents can dramatically improve workplace productivity, but their core capability, retrieving and using internal context to act on a us

Cross-Domain Data Selection and Augmentation for Automatic Compliance Detection

ApplicationsDGX agent

arXiv:2604.21469v1 Announce Type: new Abstract: Automating the detection of regulatory compliance remains a challenging task due to the complexity and variability of legal texts. Models trained on one

Decoupled DiLoCo for Resilient Distributed Pre-training

Model ReleasesDGX agent

arXiv:2604.21428v1 Announce Type: new Abstract: Modern large-scale language model pre-training relies heavily on the single program multiple data (SPMD) paradigm, which requires tight coupling across

DMAP: A Distribution Map for Text

ResearchDGX agent

arXiv:2602.11871v2 Announce Type: replace Abstract: Large Language Models (LLMs) are a powerful tool for statistical text analysis, with derived sequences of next-token probability distributions offer

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models

Model ReleasesDGX agent

arXiv:2507.04023v3 Announce Type: replace Abstract: Large language models (LLMs) achieve impressive performance on complex mathematical benchmarks yet sometimes fail on basic math reasoning while gene

Dr. Assistant: Enhancing Clinical Diagnostic Inquiry via Structured Diagnostic Reasoning Data and Reinforcement Learning

Model ReleasesDGX agent

arXiv:2601.13690v2 Announce Type: replace Abstract: Clinical Decision Support Systems (CDSSs) provide reasoning and inquiry guidance for physicians, yet they face notable challenges, including high ma

DWTSumm: Discrete Wavelet Transform for Document Summarization

Local AiDGX agent

arXiv:2604.21070v1 Announce Type: new Abstract: Summarizing long, domain-specific documents with large language models (LLMs) remains challenging due to context limitations, information loss, and hall

EduCoder: An Open-Source Annotation System for Education Transcript Data

ApplicationsDGX agent

arXiv:2507.05385v4 Announce Type: replace Abstract: We introduce EduCoder, a domain-specialized tool designed to support utterance-level annotation of educational dialogue. While general-purpose text

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning

SafetyDGX agent

arXiv:2512.05591v2 Announce Type: replace-cross Abstract: Large language model post-training relies on reinforcement learning to improve model capability and alignment quality. However, the off-policy

Evaluation of Automatic Speech Recognition Using Generative Large Language Models

ResearchDGX agent

arXiv:2604.21928v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) is traditionally evaluated using Word Error Rate (WER), a metric that is insensitive to meaning. Embedding-based sema

EVENT5Ws: A Large Dataset for Open-Domain Event Extraction from Documents

Model ReleasesDGX agent

arXiv:2604.21890v1 Announce Type: new Abstract: Event extraction identifies the central aspects of events from text. It supports event understanding and analysis, which is crucial for tasks such as in

Explainable Disentangled Representation Learning for Generalizable Authorship Attribution in the Era of Generative AI

TutorialsDGX agent

arXiv:2604.21300v1 Announce Type: new Abstract: Learning robust representations of authorial style is crucial for authorship attribution and AI-generated text detection. However, existing methods ofte

Exploring Continual Fine-Tuning for Enhancing Language Ability in Large Language Model

TutorialsDGX agent

arXiv:2410.16006v3 Announce Type: replace Abstract: A common challenge towards the adaptability of Large Language Models (LLMs) is their ability to learn new languages over time without hampering the

Finding Meaning in Embeddings: Concept Separation Curves

ResearchDGX agent

arXiv:2604.21555v1 Announce Type: new Abstract: Sentence embedding techniques aim to encode key concepts of a sentence's meaning in a vector space. However, the majority of evaluation approaches for s

Fixation Sequences as Time Series: A Topological Approach to Dyslexia Detection

ResearchDGX agent

arXiv:2604.21698v1 Announce Type: new Abstract: Persistent homology, a method from topological data analysis, extracts robust, multi-scale features from data. It produces stable representations of tim

Flipping Against All Odds: Reducing LLM Coin Flip Bias via Verbalized Rejection Sampling

SafetyDGX agent

arXiv:2506.09998v2 Announce Type: replace-cross Abstract: Large language models (LLMs) can often accurately describe probability distributions using natural language, yet they still struggle to genera

From If-Statements to ML Pipelines: Revisiting Bias in Code-Generation

SafetyDGX agent

arXiv:2604.21716v1 Announce Type: new Abstract: Prior work evaluates code generation bias primarily through simple conditional statements, which represent only a narrow slice of real-world programming

From Past To Path: Masked History Learning for Next-Item Prediction in Generative Recommendation

SafetyDGX agent

arXiv:2509.23649v2 Announce Type: replace-cross Abstract: Generative recommendation, which directly generates item identifiers, has emerged as a promising paradigm for recommendation systems. However,

From Tokens to Concepts: Leveraging SAE for SPLADE

ResearchDGX agent

arXiv:2604.21511v1 Announce Type: cross Abstract: Learned Sparse IR models, such as SPLADE, offer an excellent efficiency-effectiveness tradeoff. However, they rely on the underlying backbone vocabula

GerAV: Towards New Heights in German Authorship Verification using Fine-Tuned LLMs on a New Benchmark

Model ReleasesDGX agent

arXiv:2601.13711v2 Announce Type: replace Abstract: Authorship verification (AV) is the task of determining whether two texts were written by the same author and has been studied extensively, predomin

GRISP: Guided Recurrent IRI Selection over SPARQL Skeletons

ResearchDGX agent

arXiv:2604.21133v1 Announce Type: new Abstract: We present GRISP (Guided Recurrent IRI Selection over SPARQL Skeletons), a novel SPARQL-based question-answering method over knowledge graphs based on f

Hierarchical Policy Optimization for Simultaneous Translation of Unbounded Speech

SafetyDGX agent

arXiv:2604.21045v1 Announce Type: new Abstract: Simultaneous speech translation (SST) generates translations while receiving partial speech input. Recent advances show that large language models (LLMs

How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models

ResearchDGX agent

arXiv:2604.21106v1 Announce Type: cross Abstract: We measure how much one extra recurrence is worth to a looped (depth-recurrent) language model, in equivalent unique parameters. From an iso-depth swe

Hyperloop Transformers

Model ReleasesDGX agent

arXiv:2604.21254v1 Announce Type: cross Abstract: LLM architecture research generally aims to maximize model quality subject to fixed compute/latency budgets. However, many applications of interest su

Improving Clinical Diagnosis with Counterfactual Multi-Agent Reasoning

AgentsDGX agent

arXiv:2603.27820v2 Announce Type: replace Abstract: Clinical diagnosis is a complex reasoning process in which clinicians gather evidence, form hypotheses, and test them against alternative explanatio

It's High Time: A Survey of Temporal Question Answering

Model ReleasesDGX agent

arXiv:2505.20243v4 Announce Type: replace Abstract: Time plays a critical role in how information is generated, retrieved, and interpreted. In this survey, we provide a comprehensive overview of Tempo

Job Skill Extraction via LLM-Centric Multi-Module Framework

ResearchDGX agent

arXiv:2604.21525v1 Announce Type: new Abstract: Span-level skill extraction from job advertisements underpins candidate-job matching and labor-market analytics, yet generative large language models (L

Kernel-Smith: A Unified Recipe for Evolutionary Kernel Optimization

Model ReleasesDGX agent

arXiv:2603.28342v2 Announce Type: replace Abstract: We present Kernel-Smith, a framework for high-performance GPU kernel and operator generation that combines a stable evaluation-driven evolutionary a

Language as a Latent Variable for Reasoning Optimization

Model ReleasesDGX agent

arXiv:2604.21593v1 Announce Type: new Abstract: As LLMs reduce English-centric bias, a surprising trend emerges: non-English responses sometimes outperform English on reasoning tasks. We hypothesize t

Learning Dynamic Representations and Policies from Multimodal Clinical Time-Series with Informative Missingness

SafetyDGX agent

arXiv:2604.21235v1 Announce Type: cross Abstract: Multimodal clinical records contain structured measurements and clinical notes recorded over time, offering rich temporal information about the evolut

Learning State-Tracking from Code Using Linear RNNs

ResearchDGX agent

arXiv:2602.14814v2 Announce Type: replace-cross Abstract: Over the last years, state-tracking tasks, particularly permutation composition, have become a testbed to understand the limits of sequence mo

Listen and Chant Before You Read: The Ladder of Beauty in LM Pre-Training

ResearchDGX agent

arXiv:2604.21265v1 Announce Type: new Abstract: We show that pre-training a Transformer on music before language significantly accelerates language acquisition. Using piano performances (MAESTRO datas

Losing our Tail, Again: (Un)Natural Selection & Multilingual LLMs

ResearchDGX agent

arXiv:2507.03933v3 Announce Type: replace Abstract: Multilingual Large Language Models considerably changed how technologies influence language. While previous technologies could mediate or assist hum

← Previous
1…99100101102103…129
Next →